Wrap-up 2 min
The Price of Each Step
What you must prove to move up →
- 1Read only
- · no evidence required
- 2Suggests
- · baseline measured
- · self-verification with evidence
- 3Acts with approval
- · eval suite running
- · golden dataset with hard cases
- · human gate defined
- 4Acts alone
- · 6 weeks above baseline
- · monitoring with alerts
- · kill switch tested
- · named DRI
- · rollback possible
The final product of this module is a delegation decision with a date attached. The test suite is the means.
When you have a baseline, evals, and monitoring, the conversation about autonomy stops being political and becomes operational:
"The agent has been above the human baseline for 6 weeks across all five criteria, with 0 LGPD incidents. Move it from level 3 to level 4 for letters in Portuguese. Keep it at level 3 for Spanish, where we still have 40 cases in the dataset."
A Cognitive Enterprise can say that sentence because it can prove what it says. A traditional company, even with agents, cannot.
Every time someone asks, "Can we trust the agent?" answer with another question: "Compared with what, measured how, and who saw the number?"
