The frameworks disagree on labels but converge on shape: six stages, ordered by who owns the loop. The buyer's question is not "how smart is it?" but "who approves the action?"
The sophistication ladder
Height and color depth encode the same ordering. Select a stage for a summary, or open the full cards below.
Autonomy and scope increase to the right. Reliability and auditability tend to fall.
View as table
| Stage | What defines it | The human's role | In insurance |
|---|
A stage is a property of the deployment, not just the model. The same product can behave as Stage 2 or Stage 3 depending on its permissions, tools, and approval gates. Ask vendors how the system is wired in, not just what the model can do.
This page tracks the technology; the companion page tracks your practice. See the AI fluency ladder. One tells you what vendors can ship; the other, what your team can responsibly use.
The stages, in full
Where the experts disagree
- Model or deployment? OpenAI's ladder treats each rung as a model capability; Anthropic and DeepMind show autonomy emerges from the deployment around it. Both are right, so evaluate the wiring, not the demo.
- Agent teams: next stage or trap? Orchestrated teams post large gains on parallel, high-value work at roughly 15x the token cost, and fail on tightly coupled work. Task shape decides.
- Scaling or architecture? Lab leaders expect the upper rungs from continued scaling. Skeptics argue genuine innovation needs architectural breakthroughs, and even lab leaders concede continuous learning is still missing.
- Frontier claims vs. agent washing. The capability is real in verifiable domains; most marketing is not. Gartner counts about 130 real agentic vendors among thousands, and regulators now treat humans-behind-the-curtain sold as autonomy as a legal liability.
- Is more autonomy even the goal? Engineering and governance bodies treat autonomy as a dial, not a virtue: reliability and auditability fall as autonomy rises, so mature deployments cap it at conditional autonomy with human sign-off.
Source links
- Technical frameworks: DeepMind, "Levels of AGI" · METR task-horizon measurements · Anthropic agent taxonomy
- Industry models: Gartner AI maturity model · NVIDIA AI agents glossary · Hugging Face agents course · LangChain architecture guidance · Cloud Security Alliance agentic work
- Related pages: the companion fluency ladder carries the practitioner sources; the AI roadmap carries numbered market, carrier, and regulatory references.