Position: Levels of AGI for Operationalizing Progress on the Path to AGI
Meredith Ringel MorrisJascha Sohl-DicksteinNoah FiedelTris WarkentinAllan DafoeAleksandra FaustClément FarabetShane Legg
Establishes a structured framework that classifies artificial general intelligence across distinct tiers of capability, generality, and autonomy, providing clear criteria to evaluate current models and safely track progress toward human-level systems.
Rapid advances in machine learning and large language models have transformed Artificial General Intelligence (AGI) from an abstract theoretical discussion into an immediate operational priority. However, the artificial intelligence field lacks clear, shared definitions of what constitutes AGI, leading to inconsistent evaluations, unclear risk assessments, and fractured public policy debates. The article addresses this challenge by evaluating existing definitions and establishing a concrete, standardized framework to classify capabilities, measure progress, and assess risks across the development path to AGI.
To develop this framework, the article analyzes nine prominent historical and contemporary conceptions of AGI, ranging from the Turing Test to modern language models and economic criteria. From this review, the authors distill six core principles for defining AGI: prioritizing capabilities over internal cognitive processes, evaluating both generality and performance, focusing on cognitive and metacognitive abilities rather than physical embodiment, assessing inherent potential rather than real-world deployment, using ecologically valid real-world tasks, and treating AGI as an incremental continuum rather than a single endpoint.
The resulting ontology establishes a matrix structured across two dimensions: generality (Narrow versus General) and five tiered performance levels. Level 1 (Emerging) represents performance roughly equal to an unskilled human, Level 2 (Competent) reaches at least the 50th percentile of skilled adults, Level 3 (Expert) meets the 90th percentile, Level 4 (Virtuoso) matches the 99th percentile, and Level 5 (Superhuman) outperforms 100% of humans. The analysis finds that current frontier models, such as ChatGPT, Bard, and Llama 2, fit into Level 1 General AI ("Emerging AGI"). While these models achieve Competent or Expert performance in isolated tasks like simple coding or short essay generation, their lower performance on tasks requiring factuality and complex mathematics keeps them at the Emerging level overall. Notably, no system has yet achieved Level 2 General AI ("Competent AGI"), which represents the baseline matching most historic definitions of general intelligence.
The article demonstrates that system capabilities and operational autonomy are distinct concepts that must not be conflated. While higher capability levels unlock higher autonomy—spanning six interaction tiers from "No AI" and "AI as a Tool" up to "AI as an Agent"—increasing capability does not require relinquishing human control. Risk profiles evolve distinctly across these levels: lower tiers are dominated by human misuse and hallucination errors, middle tiers introduce structural economic and labor disruption, and the highest tiers introduce catastrophic alignment and geopolitical risks. Consequently, safety and governance depend heavily on designing appropriate human-machine interaction paradigms rather than evaluating raw model intelligence in isolation.
Decision-makers and developers should adopt this leveled taxonomy to improve model reporting, policy formulation, and safety evaluations. Instead of relying on static, easily saturated metrics, the research community should collaborate across disciplines to construct dynamic, living benchmarks that measure ecologically valid cognitive and metacognitive capabilities, such as learning new skills and knowing when to ask for human assistance. These evaluations should evaluate dangerous dual-use capabilities within sandboxed environments without requiring autonomous open-world deployment.
The framework carries inherent uncertainties, as the authors intentionally do not define the exact scope or percentage threshold of tasks required to fulfill the generality criteria at each tier. Furthermore, theoretical benchmark performance may overestimate real-world impact due to user interface friction and prompt-engineering barriers. Nevertheless, the conceptual framework provides a robust, highly confident foundation for operationalizing progress, managing deployment risks, and aligning AI capabilities with human oversight.
- Paper: Universal Intelligence: A Definition of Machine Intelligence, Shane Legg et al. (2007). It provides the foundational formal and mathematical definition of general machine intelligence across broad environments, which directly underpins the theoretical framing of AGI.
- Paper: Sparks of Artificial General Intelligence: Early experiments with GPT-4, Sébastien Bubeck et al. (2023). It offers the early empirical investigation into frontier large language model capabilities across diverse domains that motivated operationalizing continuous levels of general intelligence.
- Paper: Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy, Ben Shneiderman (2020). It establishes the crucial conceptual distinction between levels of machine automation and levels of human control, which informs the source's separation of capability from autonomy.
- Paper: Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench authors (2022). It demonstrates large-scale multi-task benchmarking against human baselines, highlighting the challenge of benchmark saturation and the need for standardized capability measurement.
- Paper: On the Opportunities and Risks of Foundation Models, Rishi Bommasani et al. (2021). It articulates the emergence, generality, and sociotechnical risk profile of broad foundation models that serve as the contemporary baseline for emerging AGI.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). It introduced the paradigm shift to broad, few-shot generalist language models that transitioned AGI evaluation from narrow benchmarks to generalized task performance.
- Paper: Intelligence Without Reason, Rodney A. Brooks (1991). It presents the seminal physical embodiment perspective of intelligence, which the source explicitly contrasts by focusing its core ontology on cognitive and metacognitive capabilities.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). It synthesizes the architectural paradigms and operational capabilities of LLM-based autonomous agents, providing context for the source's interaction and autonomy tiers.
- Paper: From AGI to ASI, Tim Genewein et al. (2026). It directly extends the developmental trajectory beyond human-level AGI to examine pathways, frictions, and governance challenges associated with achieving Artificial Superintelligence.
- Paper: Metacognition in LLMs: Foundations, Progress, and Opportunities, Gabrielle Kaili-May Liu et al. (2026). It details empirical methods and benchmarks for assessing metacognition in LLMs, directly advancing the source's call for dynamic evaluations of self-monitoring and cognitive boundaries.
- Paper: CogBench: a large language model walks into a psychology lab, Julian Coda-Forno et al. (2024). It applies cognitive psychology experiments to benchmark behavioral and metacognitive traits in language models, operationalizing the source's criteria for cognitive capability assessment.
- Paper: Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision, Zhiqing Sun et al. (2024). It explores scalable oversight mechanisms that enable supervision and alignment when model capabilities reach superhuman performance levels.
- Paper: Intelligent AI Delegation, Nenad Tomašev et al. (2026). It operationalizes the practical transfer of authority and accountability in multi-agent workflows, extending the source's framework on agent autonomy and human oversight.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). It introduces multi-step interactive evaluation frameworks for autonomous agents, implementing the ecologically valid trajectory benchmarks proposed in the source.
- Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). It develops diagnostic guardrails for evaluating and mitigating safety risks across full agent execution trajectories, aligning with the source's tier-based risk governance.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). It synthesizes agentic reasoning techniques that transition language models from passive tools to autonomous goal-directed agents.
- Paper: Visual General Intelligence: A White Paper, Hirokatsu Kataoka et al. (2026). It investigates whether visual experience and world modeling can offer an alternative or complementary foundation for achieving broad general intelligence.
