Levels of AGI: DeepMind's Framework Explained
Stop debating whether AI is broadly "smart" and start measuring its specific capacity for autonomous execution. For enterprise engineering teams, the subjective Turing Test is dead.
As models move from simple chatbots to complex reasoning engines, we need strict taxonomies to evaluate their capabilities. Understanding the architectural shift toward artificial general intelligence requires standard rubrics that separate rote memorization from novel, cross-domain problem-solving.
If you want to grasp the full scope of this intelligence transition, you must first establish a baseline understanding of how these systems will scale toward cognitive parity.
- Dual-Axis Evaluation: DeepMind categorizes intelligence by measuring both performance (skill level) and generality (breadth of tasks).
- Autonomy is Distinct: A model's raw intellectual capability is measured independently from its ability to act autonomously in an environment.
- Current Status: The September 2026 frontier generation (GPT-6 Astra, Claude Opus 5 and Claude Fable 5.1, Gemini 3.8) still sits primarily at Level 1 (Emerging AGI), showing broad capability but lacking consistent expert-level reasoning across all fields, a pattern researchers call "jagged intelligence."
- A Rival Yardstick: A newer CHC-based scoring framework puts GPT-5 at roughly 57-58% of the way to human-level cognitive versatility, up from GPT-4's 27%, a useful number, though critics say it can overstate generality.
- The Superhuman Ceiling: The framework culminates at Level 5, where a system outperforms 100% of humans across 100% of cognitive tasks.
- Risk Tracks Level: Higher levels generally mean higher deployment risk, which is why oversight needs tend to scale alongside the level.
Performance vs. Generality: The Core Metrics
DeepMind's researchers, led by Morris et al., argued that binary definitions of AGI were impractical for tracking real-world AI development. To map progress more precisely, they introduced a matrix based on two distinct axes.
Performance measures how well an AI executes a specific task compared to a human baseline. Generality measures the range of completely different cognitive tasks the system can handle without fine-tuning.
A traditional chess bot possesses superhuman performance but zero generality. It is highly competent but strictly narrow.
Conversely, an early Large Language Model possesses high generality but relatively low performance on complex logic. True AGI requires moving up and to the right on both axes simultaneously.
The 5 Levels of AGI
DeepMind's taxonomy, laid out in the original Levels of AGI paper, classifies artificial intelligence into five distinct tiers, tracking the progression from basic competence to absolute mastery. The percentile thresholds below (50th, 90th, and 99th) are the paper's own benchmarks for comparing systems to skilled human adults, not independently verified performance claims about any specific model.
Level 1: Emerging
At this stage, the AI equals or slightly exceeds an unskilled human across a wide variety of tasks.
These models can draft emails, summarize documents, and generate basic code. They possess high generality but lack the deep, reliable reasoning required for specialized professional work without human intervention.
Level 2: Competent
Level 2 models perform at the level of at least the 50th percentile of skilled adults in a given domain.
They can execute multi-step workflows, troubleshoot standard engineering problems, and maintain contextual consistency over longer horizons. They require less prompting and exhibit stronger baseline reasoning.
Level 3: Expert
An Expert AGI matches the 90th percentile of skilled human professionals.
At this tier, the system is indistinguishable from a senior engineer, doctor, or legal analyst. It can synthesize massive amounts of complex data, identify novel patterns, and generate high-value strategic outputs independently.
Level 4: Virtuoso
A Virtuoso system performs at the 99th percentile of human capability.
This model competes with the absolute best human minds in the world. It can invent new methodologies, solve previously intractable scientific problems, and optimize systems with a level of insight that only elite human specialists possess.
Level 5: Superhuman
The final tier represents a system that outperforms 100% of humans across all possible cognitive tasks.
Level 5 AI can learn, adapt, and execute faster and more flawlessly than any human biological brain. It represents a fundamental shift in cognitive architecture and problem-solving capability on a global scale.
Visual: The Five Levels at a Glance
Bar height is illustrative of relative capability ceiling per level, not a precise metric. The dashed marker shows where consensus currently places frontier models: solidly Level 1, with narrow flashes into Levels 2-3.
Where Today's Frontier Models Actually Sit
As of early September 2026, the current frontier generation is still best classified as Level 1: Emerging AGI on DeepMind's scale, despite producing far more capable outputs than the ChatGPT, GPT-4 and Bard systems the November 2023 paper originally assessed. That generation turns over fast: the first week of September 2026 alone brought OpenAI's GPT-6 Astra, Anthropic's Claude Fable 5.1 and Mythos 5.1, Google's Gemini 3.8 and Meta's Muse Spark 1.3, with Claude Opus 5 having shipped in July. Treat the specific names below as a snapshot rather than a fixed list, and check release dates before relying on any capability claim tied to a model version.
They occasionally flash Level 2 (Competent) or even Level 3 (Expert) capabilities on specific benchmark tests: models from Google DeepMind and OpenAI reached gold-medal level at the 2025 International Mathematical Olympiad, and Gemini separately reached gold-medal level at the 2025 ICPC World Finals coding contest, but overall reliability remains notoriously brittle. This uneven profile is now commonly called jagged intelligence — a phrase popularised by Andrej Karpathy and since formalised by researchers in a 2026 technical report on characterizing model jaggedness: a model can out-reason a specialist on one problem and fail a simple logic puzzle in the same session. Separately, AI safety research group METR has tracked the length of tasks frontier models can complete autonomously, finding that this "time horizon" has roughly doubled every seven months since 2019, evidence of fast-growing autonomy even while overall reliability stays uneven.
OpenAI's GPT-6 Astra launch post, published in that same first week of September, is the sharpest recent illustration of the pattern. OpenAI reports Astra reaching 99.9% on ARC-AGI-3, a benchmark purpose-built to test reasoning in novel interactive environments rather than memorized patterns, and quotes ARC Prize Foundation president Greg Kamradt describing the result as effective human parity on the benchmark's efficiency baseline. That benchmark was administered by ARC Prize rather than by OpenAI, so the milestone is externally confirmed — but the number splits in a way that matters for anyone placing a model on this scale. ARC Prize's own analysis reports roughly 63% on ARC Prize's standard harness and 99.9% through a new provider-adapter harness that preserves reasoning state between requests, letting the model reuse earlier work. What is being levelled, in other words, is a system: a model plus its scaffolding, effort setting and token budget. ARC Prize is explicit that the result does not establish AGI, and the parity run cost roughly $360 per game in tokens against a fraction of a cent for the human baseline. In the same post, OpenAI reports Astra scoring 57.2% on Humanity's Last Exam with tools, trailing the 63-65% range the post credits to rival models, and its cybersecurity capability is separately rated "Critical" under OpenAI's own Preparedness Framework. A near-saturated score on one flagship reasoning benchmark alongside a mid-pack score on another, from the same model in the same week, is jaggedness in miniature. It's also worth noting these are OpenAI's self-reported numbers rather than independently reproduced scores, so treat the specific percentages as provisional until outside evaluators weigh in.
The industry is actively working on architectures to bridge the gap to a consistent Level 2, and the race to reach the higher tiers is drawing enormous capital investment. The field is also split on approach: most labs are betting that scaling current transformer-based architectures gets there, while a smaller but well-funded camp, led by Yann LeCun's new venture AMI Labs, which raised $1.03 billion in March 2026, argues an entirely different architecture (world models rather than next-token prediction) is required. Predictions about how soon these gaps close vary widely: Anthropic told U.S. policymakers in its OSTP submission that it anticipates powerful, broadly capable AI systems could emerge as soon as late 2026 or 2027, while Microsoft AI's Mustafa Suleyman said in a Financial Times interview that most computer-based white-collar tasks could be automated within 12 to 18 months. Both are forward-looking predictions from interested parties, not independently verified results, and should be read as such.
Enterprise teams looking to leverage these systems effectively must prioritize continuous education and robust engineering practices rather than waiting for a "Level 2 announcement" that may arrive unevenly, capability by capability, rather than all at once.
A Competing Framework: The CHC-Based AGI Score
DeepMind's ordinal levels aren't the only serious attempt to operationalize AGI. A widely discussed 2025 paper, "A Definition of AGI" (Hendrycks et al.), takes a different approach: instead of five discrete tiers, it borrows the Cattell-Horn-Carroll (CHC) theory of human cognition, a widely used psychometric model of intelligence, and scores AI systems across ten broad abilities: general knowledge, reading and writing, mathematics, on-the-spot reasoning, working memory, long-term memory storage, long-term memory retrieval, visual processing, auditory processing, and speed.
Averaging those ten scores produces a single AGI percentage. The paper reports GPT-4 at roughly 27% and GPT-5 at roughly 57-58% (sources for the underlying study cite both 57% and 58% depending on the version), a genuine doubling in about two years. But critics of the methodology point out that a simple average can flatter a system with an uneven profile: exceptional knowledge and math scores can mathematically offset a near-zero score in long-term memory storage, producing a composite that looks more "generally capable" than the system actually is in practice. That's precisely the failure mode DeepMind's dual-axis (performance × generality) design was built to avoid, since it refuses to let strength in one narrow area stand in for genuine breadth.
The two frameworks are best read as complementary rather than competing: DeepMind's levels are the field's most-cited ordinal taxonomy for policy and risk conversations, while the CHC-based score gives a concrete, trackable number for year-over-year progress on the underlying cognitive ingredients.
Test Your Understanding: Levels of AGI Quiz
Five questions on DeepMind's taxonomy and where today's models actually sit.
Conclusion
DeepMind's Levels of AGI framework strips away the hype and provides a sober, metric-driven taxonomy for evaluating the future of artificial intelligence.
By explicitly decoupling raw intelligence from autonomy, and measuring performance against generality, engineering teams can accurately assess where frontier models currently sit and what architectural breakthroughs are required to reach the next tier. Read next: our comparison of AGI timeline predictions for when the industry expects the next tier to arrive.
Frequently Asked Questions (FAQ)
The Levels of AGI is a framework published by DeepMind that categorizes AI systems based on a matrix of performance (skill level compared to humans) and generality (the breadth of cognitive tasks the system can handle).
The framework was introduced in a research paper authored by a team of researchers at Google DeepMind (Morris et al.) to establish a standardized taxonomy for evaluating progress toward human-level AI.
Performance measures the depth of skill (how well a model executes a task compared to a human), while generality measures the breadth of skill (how many fundamentally different cognitive tasks the model can perform).
As of September 2026, the frontier generation (GPT-6 Astra, Claude Opus 5 and Claude Fable 5.1, Gemini 3.8) is generally classified as Level 1 (Emerging AGI). While these models show high generality and increasingly hit Level 2 or 3 on specific benchmarks — including gold-medal math and coding olympiad results, and GPT-6 Astra's 99.9% ARC-AGI-3 score under a provider-adapter harness, against roughly 63% on ARC Prize's standard harness — their overall reliability and cross-domain reasoning remain below competent human professionals; that same Astra launch post reported a comparatively middling 57.2% on Humanity's Last Exam with tools. Researchers call this uneven profile "jagged intelligence." Frontier releases have run close to monthly through 2026, so treat the model names here as a snapshot.
Emerging AGI (Level 1) refers to an AI system that equals or slightly exceeds an unskilled human across a wide variety of tasks. It is highly versatile but lacks deep, reliable, and autonomous expert-level reasoning.
Superhuman AGI (Level 5) describes a system that completely outperforms 100% of human beings across 100% of cognitive tasks, demonstrating flawless cross-domain reasoning, invention, and execution.
Raw intellectual capability (solving a math equation) is different from agentic autonomy (identifying a problem, writing the code to solve it, and deploying the solution without human prompting). Measuring them separately clarifies risk and deployment readiness.
A single binary threshold is subjective and ignores the gradual, component-based progression of AI capabilities. A leveled taxonomy allows researchers and enterprises to track incremental architectural breakthroughs accurately.
As models ascend the levels, their capacity for autonomous action and complex reasoning increases. This escalating capability directly correlates to an expanded blast radius for failures, demanding exponentially stronger governance and oversight mechanisms.
Yes. A 2025 paper by Hendrycks et al. proposes a CHC-based AGI score, adapting the Cattell-Horn-Carroll theory of human cognition to score AI systems across ten broad abilities and averaging them into a single percentage. It reported GPT-4 at roughly 27% and GPT-5 at roughly 57-58%. Critics note that a simple average can overstate generality by letting strong scores in one domain offset near-zero scores elsewhere, exactly what DeepMind's dual-axis design was built to prevent.
While various AI labs use their own internal evaluation rubrics, the DeepMind framework is one of the most widely cited taxonomies in academic and enterprise discussions of AGI progression, alongside newer alternatives such as the CHC-based AGI score.