Position: Levels of AGI for Operationalizing Progress on the Path to AGI

Meredith Ringel MorrisJascha Sohl-DicksteinNoah FiedelTris WarkentinAllan DafoeAleksandra FaustClément FarabetShane Legg

article2024ICML128 citations

Establishes a structured framework that classifies artificial general intelligence across distinct tiers of capability, generality, and autonomy, providing clear criteria to evaluate current models and safely track progress toward human-level systems.

Listen

Rapid advances in machine learning and large language models have transformed Artificial General Intelligence (AGI) from an abstract theoretical discussion into an immediate operational priority. However, the artificial intelligence field lacks clear, shared definitions of what constitutes AGI, leading to inconsistent evaluations, unclear risk assessments, and fractured public policy debates. The article addresses this challenge by evaluating existing definitions and establishing a concrete, standardized framework to classify capabilities, measure progress, and assess risks across the development path to AGI.

To develop this framework, the article analyzes nine prominent historical and contemporary conceptions of AGI, ranging from the Turing Test to modern language models and economic criteria. From this review, the authors distill six core principles for defining AGI: prioritizing capabilities over internal cognitive processes, evaluating both generality and performance, focusing on cognitive and metacognitive abilities rather than physical embodiment, assessing inherent potential rather than real-world deployment, using ecologically valid real-world tasks, and treating AGI as an incremental continuum rather than a single endpoint.

The resulting ontology establishes a matrix structured across two dimensions: generality (Narrow versus General) and five tiered performance levels. Level 1 (Emerging) represents performance roughly equal to an unskilled human, Level 2 (Competent) reaches at least the 50th percentile of skilled adults, Level 3 (Expert) meets the 90th percentile, Level 4 (Virtuoso) matches the 99th percentile, and Level 5 (Superhuman) outperforms 100% of humans. The analysis finds that current frontier models, such as ChatGPT, Bard, and Llama 2, fit into Level 1 General AI ("Emerging AGI"). While these models achieve Competent or Expert performance in isolated tasks like simple coding or short essay generation, their lower performance on tasks requiring factuality and complex mathematics keeps them at the Emerging level overall. Notably, no system has yet achieved Level 2 General AI ("Competent AGI"), which represents the baseline matching most historic definitions of general intelligence.

The article demonstrates that system capabilities and operational autonomy are distinct concepts that must not be conflated. While higher capability levels unlock higher autonomy—spanning six interaction tiers from "No AI" and "AI as a Tool" up to "AI as an Agent"—increasing capability does not require relinquishing human control. Risk profiles evolve distinctly across these levels: lower tiers are dominated by human misuse and hallucination errors, middle tiers introduce structural economic and labor disruption, and the highest tiers introduce catastrophic alignment and geopolitical risks. Consequently, safety and governance depend heavily on designing appropriate human-machine interaction paradigms rather than evaluating raw model intelligence in isolation.

Decision-makers and developers should adopt this leveled taxonomy to improve model reporting, policy formulation, and safety evaluations. Instead of relying on static, easily saturated metrics, the research community should collaborate across disciplines to construct dynamic, living benchmarks that measure ecologically valid cognitive and metacognitive capabilities, such as learning new skills and knowing when to ask for human assistance. These evaluations should evaluate dangerous dual-use capabilities within sandboxed environments without requiring autonomous open-world deployment.

The framework carries inherent uncertainties, as the authors intentionally do not define the exact scope or percentage threshold of tasks required to fulfill the generality criteria at each tier. Furthermore, theoretical benchmark performance may overestimate real-world impact due to user interface friction and prompt-engineering barriers. Nevertheless, the conceptual framework provides a robust, highly confident foundation for operationalizing progress, managing deployment risks, and aligning AI capabilities with human oversight.

arXiv: 2311.02462
  • Paper: From AGI to ASI, Tim Genewein et al. (2026). It directly extends the developmental trajectory beyond human-level AGI to examine pathways, frictions, and governance challenges associated with achieving Artificial Superintelligence.
  • Paper: Metacognition in LLMs: Foundations, Progress, and Opportunities, Gabrielle Kaili-May Liu et al. (2026). It details empirical methods and benchmarks for assessing metacognition in LLMs, directly advancing the source's call for dynamic evaluations of self-monitoring and cognitive boundaries.
  • Paper: CogBench: a large language model walks into a psychology lab, Julian Coda-Forno et al. (2024). It applies cognitive psychology experiments to benchmark behavioral and metacognitive traits in language models, operationalizing the source's criteria for cognitive capability assessment.
  • Paper: Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision, Zhiqing Sun et al. (2024). It explores scalable oversight mechanisms that enable supervision and alignment when model capabilities reach superhuman performance levels.
  • Paper: Intelligent AI Delegation, Nenad Tomašev et al. (2026). It operationalizes the practical transfer of authority and accountability in multi-agent workflows, extending the source's framework on agent autonomy and human oversight.
  • Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). It introduces multi-step interactive evaluation frameworks for autonomous agents, implementing the ecologically valid trajectory benchmarks proposed in the source.
  • Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). It develops diagnostic guardrails for evaluating and mitigating safety risks across full agent execution trajectories, aligning with the source's tier-based risk governance.
  • Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). It synthesizes agentic reasoning techniques that transition language models from passive tools to autonomous goal-directed agents.
  • Paper: Visual General Intelligence: A White Paper, Hirokatsu Kataoka et al. (2026). It investigates whether visual experience and world modeling can offer an alternative or complementary foundation for achieving broad general intelligence.
Cover for Position: Levels of AGI for Operationalizing Progress on the Path to AGI

Abstract

We propose a framework for classifying the capabilities and behavior of Artificial General Intelligence (AGI) models and their precursors. This framework introduces levels of AGI performance, generality, and autonomy, providing a common language to compare models, assess risks, and measure progress along the path to AGI. To develop our framework, we analyze existing definitions of AGI, and distill six principles that a useful ontology for AGI should satisfy. With these principles in mind, we propose “Levels of AGI” based on depth (performance) and breadth (generality) of capabilities, and reflect on how current systems fit into this ontology. We discuss the challenging requirements for future benchmarks that quantify the behavior and capabilities of AGI models against these levels. Finally, we discuss how these levels of AGI interact with deployment considerations such as autonomy and risk, and emphasize the importance of carefully selecting Human-AI Interaction paradigms for responsible and safe deployment of highly capable AI systems.

Table of Contents

  • 1. Introduction
  • 2. Defining AGI: Case Studies
  • 3. Defining AGI: Six Principles
  • 4. Levels of AGI
  • 5. Testing for AGI
  • 6. Risk, Autonomy, and Interaction
  • 6.1. Levels of AGI as a Framework for Risk Assessment
  • 6.2. Capabilities vs. Autonomy
  • 6.3. Human-AI Interaction and Risk Assessment
  • 7. Conclusion
  • Impact Statement
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Two-dimensional ontology for classifying AGI

    model/method

    The paper defines progress toward Artificial General Intelligence (AGI) using two independent dimensions:

    • Performance measures the depth of capability: how well an AI system performs a task relative to humans who possess the relevant skill. Above the Emerging level, human comparisons are made against skilled adults rather than the entire adult population.
    • Generality measures the breadth of capability: the range of cognitive and metacognitive tasks on which the system reaches a target performance level.

    The ontology distinguishes Narrow AI, which is restricted to a clearly scoped task or set of tasks, from General AI, which handles a wide range of non-physical tasks, including metacognitive abilities such as learning new skills. A system’s level is determined by the minimum performance it achieves across most relevant tasks, although it may perform at higher levels on a narrow subset of tasks. The framework evaluates what a system can do rather than whether it uses human-like reasoning, possesses consciousness, or has been deployed in the real world.

  2. Knowl 2 — Performance and generality levels in the AGI matrix

    model/method

    The proposed matrix combines six performance levels with two generality categories. The performance thresholds and representative classifications are:

    • Level 0 — No AI: No AI capability. Narrow examples include calculator software and compilers; general examples include human-in-the-loop computing such as Amazon Mechanical Turk.
    • Level 1 — Emerging: Performance equal to or somewhat better than an unskilled human. Narrow examples include GOFAI and simple rule-based systems such as SHRDLU. General examples include ChatGPT, Bard, Llama 2, and Gemini as characterized by the paper.
    • Level 2 — Competent: At least the 50th percentile of skilled adults. Narrow examples include toxicity detectors, smart speakers, visual-question-answering systems such as PaLI, IBM Watson, and frontier language models on selected tasks such as short essay writing or simple coding. Competent AGI had not been achieved by any public system at the time of writing.
    • Level 3 — Expert: At least the 90th percentile of skilled adults. Narrow examples include Grammarly and image-generation systems such as Imagen and DALL-E 2. Expert AGI had not been achieved.
    • Level 4 — Virtuoso: At least the 99th percentile of skilled adults. Narrow examples include Deep Blue and AlphaGo. Virtuoso AGI had not been achieved.
    • Level 5 — Superhuman: Outperforms 100% of humans. Narrow examples include AlphaFold, AlphaZero, and Stockfish. A general system at this level is called Artificial Superintelligence (ASI), and had not been achieved.

    The cell assignments are approximate. General systems may perform at higher levels on particular tasks than their overall matrix position suggests, and progression in performance or generality may be nonlinear. The matrix is intended to describe a path toward AGI rather than a single binary endpoint.

  3. Knowl 3 — Six principles for an operational definition of AGI

    model/method

    The paper proposes six principles for making AGI definitions clear and operationalizable:

    1. Capabilities rather than processes: Define AGI by what a system can accomplish, not by whether it thinks like a human, uses a brain-like mechanism, or possesses consciousness or sentience.
    2. Generality and performance: Require both breadth across tasks and adequate depth of performance; generality without reliable performance is insufficient.
    3. Cognitive and metacognitive rather than necessarily physical tasks: Cognitive abilities, including learning and knowing when to request assistance, are central. Physical embodiment may increase generality and may help with some tasks, but it is not a necessary prerequisite for AGI.
    4. Potential rather than deployment: Demonstrated capability to perform the relevant tasks should be sufficient. Actual open-world deployment, labor substitution, or realized economic value should not be required because legal, social, ethical, and safety factors can prevent deployment despite technical capability.
    5. Ecological validity: Benchmarks should measure tasks that correspond to real-world activities people value, including social and artistic value rather than only economic value or metrics that are easy to automate.
    6. The path to AGI rather than a single endpoint: AGI should be represented using levels with associated benchmarks, risks, and human-AI interaction implications, allowing progress to be discussed before a system reaches a single all-or-nothing threshold.
  4. Knowl 4 — Requirements for a living, ecologically valid AGI benchmark

    model/method

    The paper does not propose a finished AGI benchmark; instead, it specifies what such a benchmark should measure. It should contain a broad and difficult suite of cognitive and metacognitive tasks covering linguistic intelligence, mathematical and logical reasoning, spatial reasoning, interpersonal and intrapersonal social intelligence, creativity, and the ability to learn new skills.

    The benchmark should favor ecological and construct validity over traditional metrics that are easy to score but may not represent valued real-world abilities. It should include open-ended and interactive tasks, some requiring qualitative assessment, because these may better approximate real use than fixed-answer tests. Whether tools, including AI-powered tools, are allowed should be decided by task context and by the relevant real-world counterfactual; for example, a driving-safety evaluation may appropriately compare people with access to driver-assistance technology rather than unaided drivers.

    Because no finite list can enumerate every task available to a sufficiently general intelligence, the benchmark should be living: it should include a process for generating, evaluating, and agreeing upon new tasks. The paper deliberately leaves open the set of tasks, the proportion that must be mastered for a generality rating, and whether some tasks—especially metacognitive tasks—must be mandatory.

  5. Knowl 5 — Metacognition as a core requirement for general intelligence

    model/method

    The paper argues that an AGI benchmark should explicitly test metacognitive abilities because a system cannot be optimized in advance for every possible use case. At minimum, it should assess:

    • the ability to learn new skills and select appropriate learning strategies;
    • the ability to recognize when assistance is needed and ask a human for help;
    • calibration, meaning awareness of the system’s limits and the ability to anticipate or evaluate its own likely or actual task performance; and
    • social metacognition, including theory-of-mind abilities needed to model users accurately.

    These capabilities are treated as prerequisites for useful generality and for aligned human-AI interaction. Knowing when to request help can reduce inappropriate autonomous action, while accurate user modeling supports alignment with human goals. Strong metacognition is also expected to be particularly important for collaborative, expert, and agentic interaction paradigms.

  6. Knowl 6 — Six levels of human-AI autonomy

    model/method

    The paper proposes six autonomy levels that characterize the style of human-AI interaction rather than merely the amount of control relinquished to a computer:

    • Autonomy Level 0 — No AI: The human performs everything. Examples include analogue work and non-AI digital workflows. The relevant risks are status-quo risks.
    • Autonomy Level 1 — AI as a Tool: The human fully controls the task and uses AI for mundane subtasks, such as search, grammar correction, or translation. Emerging Narrow AI may enable this level, while Competent Narrow AI is more likely. Risks include deskilling through over-reliance and disruption of established industries.
    • Autonomy Level 2 — AI as a Consultant: The AI performs a substantive role only when invoked by a human, such as summarizing documents, generating code, or recommending entertainment. Competent Narrow AI may enable it; Expert Narrow AI and Emerging AGI make it more likely. Risks include over-trust, radicalization, and targeted manipulation.
    • Autonomy Level 3 — AI as a Collaborator: Human and AI cooperate as co-equals, interactively coordinating goals and tasks. Emerging AGI may enable it, while Expert Narrow AI or Competent AGI make it more likely. Risks include anthropomorphization, parasocial relationships, and rapid societal change.
    • Autonomy Level 4 — AI as an Expert: The AI drives the interaction while the human supplies guidance, feedback, or subtasks, such as using AI to advance scientific discovery. Virtuoso Narrow AI may enable it, while Expert AGI makes it more likely. Risks include societal-scale ennui, mass labor displacement, and decline of human exceptionalism.
    • Autonomy Level 5 — AI as an Agent: The AI acts fully autonomously, as in an autonomous AI-powered personal assistant. This level was not yet unlocked; it is expected to require Virtuoso AGI or ASI. Risks include misalignment and concentration of power.
  7. Knowl 7 — AGI capability does not determine the chosen autonomy level

    theoretical result

    The paper’s framework separates an AI system’s capability level from its deployed autonomy level. Increasing performance and generality can unlock additional interaction paradigms, but it does not require organizations or users to select the maximum available autonomy. A highly capable AGI may deliberately be deployed as a tool, consultant, or collaborator rather than as an autonomous agent when education, enjoyment, assessment, safety, or contextual constraints favor human control.

    The deployment context—including the interface, task, scenario, and end user—substantially affects risk. General systems may also have uneven capabilities: an Emerging AGI can reach Competent or Expert performance on selected tasks and therefore support higher autonomy for those tasks without supporting that autonomy generally. Higher autonomy levels additionally depend on capabilities such as knowing when to ask for help, modeling users, and possessing social-emotional competence. A fully autonomous agent implicitly needs to act in an aligned manner without continuous oversight while still knowing when human consultation is necessary.

  8. Knowl 8 — Risk profiles along the AGI capability path

    theoretical result

    The paper argues that risk should be assessed at each combination of performance, generality, and autonomy rather than only at a hypothetical final AGI endpoint. As systems advance, misuse risks, alignment risks, and structural risks can change in different directions.

    • Below Expert AGI, including Emerging AGI, Competent AGI, and Narrow AI systems, risks are expected to arise primarily from human actions such as accidental, incidental, or malicious misuse.
    • Expert AGI is likely to introduce major structural risks such as economic disruption and job displacement as more industries reach thresholds for substituting machine intelligence for human labor. Its greater reliability may simultaneously reduce some risks of incorrect task execution associated with less capable systems.
    • Virtuoso AGI and ASI are the levels at which concerns about broad misalignment and existential or extreme risks are expected to become especially salient, including possible deception of human operators in pursuit of a mis-specified goal.
    • Rapid progression between levels could create systemic risks if regulation or diplomacy cannot keep pace, including international destabilization and large geopolitical or military advantages for an early developer of ASI.

    These are proposed risk associations rather than claims that every system at a given level will exhibit every listed risk.

  9. Knowl 9 — Implications for evaluating dangerous capabilities

    limitation

    The paper recommends considering dangerous or dual-use capabilities—such as deception, persuasion, and advanced biochemistry—in AGI evaluation because these abilities can support both socially beneficial and harmful applications. Evaluations should focus on potential rather than deployment and should sandbox dangerous tasks so that measuring capability does not require releasing the system into the world.

    Publicly releasing such evaluations creates a countervailing risk: malicious actors could use the benchmark to optimize systems for dangerous abilities. The paper therefore identifies mitigation of benchmark-related risks as an open research problem for AI safety, ethics, and governance. It also suggests that capabilities in the AGI ontology could be connected to risk and containment levels in responsible-scaling policies.

  10. Knowl 10 — Assessment of frontier systems under the proposed ontology

    empirical result

    Using the state of public systems in September 2023 as an illustrative assessment, the paper characterizes frontier language models such as ChatGPT, Bard, and Llama 2 as having Competent-level performance on selected tasks, including short essay writing and simple coding, but Emerging-level performance on most tasks, including mathematical ability and factuality. Their overall classification is therefore Level 1, Emerging AGI, rather than Competent AGI, because generality requires adequate performance across most relevant tasks.

    The paper also uses DALL-E 2 to illustrate the distinction between theoretical and deployed performance. Its image quality is estimated to be at the Expert level relative to most people’s drawing ability, but recurrent failures such as incorrect hands and illegible text prevent a Virtuoso designation. In practical use, complex prompting interfaces may reduce performance to Competent level for ordinary users, showing why benchmarks should measure ecologically valid deployed performance rather than only idealized capability under expert prompting.

  11. Knowl 11 — Open limitations of the ontology and benchmark proposal

    limitation

    The framework leaves several load-bearing questions unresolved. It does not specify the complete set of tasks defining generality, the fraction of tasks that must be passed for a generality rating, or whether particular tasks must always be included. The authors expect the required fraction to be high but not necessarily 100%, since humans are broadly intelligent despite imperfect performance across all tasks.

    The proposed level assignments are approximate until standardized, diverse benchmarks exist, and a system’s measured capability may exceed its practical deployed performance because of interface limitations. The ontology is also intended to evolve as technical advances—such as improved interpretability—change how progress can be measured and as society adapts to new human-AI interaction paradigms.

Coverage note — The nine source definitions of AGI are not extracted as separate knowls because their substantive conclusions are consolidated into the six principles, ontology, benchmark requirements, and system classifications.

References

  1. 1.Aguera y Arcas, B. and Norvig, P. Artificial General Intelligence is Already Here. Noema, October 2023. URL https://www.noemamag.com/artificial-general-intelligence-is-already-here/.
  2. 2.Amazon. Amazon Alexa. URL https://alexa.amazon.com/. accessed on October 20, 2023.
  3. 3.Andreas, J., Begus, G., Bronstein, M. M., Diamant, R., Delaney, D., Gero, S., Goldwasser, S., Gruber, D. F., de Haas, S., Malkin, P., Pavlov, N., Payne, R., Petri, G., Rus, D., Sharma, P., Tchernov, D., Tønnesen, P., Torralba, A., Vogt, D., and Wood, R. J. Toward understanding the communication in sperm whales. iScience, 25(6):104393, 2022. ISSN 2589-0042. doi: https://doi.org/10.1016/j.isci.2022.104393. URL https://www.sciencedirect.com/science/article/pii/S2589004222006642.
  4. 4.Anil, R., Dai, A. M., Firat, O., and et al. PaLM 2 Technical Report. CoRR, abs/2305.10403, 2023. doi: 10.48550/arXiv.2305.10403. URL https://arxiv.org/abs/2305.10403.
  5. 5.Anthropic. Company: Anthropic, 2023a. URL https://www.anthropic.com/company. Accessed October 12, 2023.
  6. 6.Anthropic. Anthropic’s Responsible Scaling Policy, September 2023b. URL https://www-files.anthropic.com/production/files/responsible-scaling-policy-1.0.pdf. accessed on October 20, 2023.
  7. 7.Apple. Siri. URL https://www.apple.com/siri/. accessed on October 20, 2023.
  8. 8.Bellier, L., Llorens, A., Marciano, D., Gunduz, A., Schalk, G., Brunner, P., and Knight, R. T. Music can be reconstructed from human auditory cortex activity using nonlinear decoding models. PLOS Biology, 21(8):1–27, 08 2023. doi: 10.1371/journal.pbio.3002176. URL https://doi.org/10.1371/journal.pbio.3002176.
  9. 9.Bengio, Y., Hinton, G., Yao, A., Song, D., Abbeel, P., Harari, Y. N., Zhang, Y.-Q., Xue, L., Shalev-Shwartz, S., Hadfield, G., Clune, J., Maharaj, T., Hutter, F., Baydin, A. G., McIlraith, S., Gao, Q., Acharya, A., Krueger, D., Dragan, A., Torr, P., Russell, S., Kahneman, D., Brauner, J., and Mindermann, S. Managing AI Risks in an Era of Rapid Progress. CoRR, abs/2310.17688, 2023. doi: 10.48550/arXiv.2310.17688. URL https://arxiv.org/abs/2310.17688.
  10. 10.Boden, M. A. GOFAI, pp. 89–107. Cambridge University Press, 2014.
  11. 11.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., Florence, P., Fu, C., Arenas, M. G., Gopalakrishnan, K., Han, K., Hausman, K., Herzog, A., Hsu, J., Ichter, B., Irpan, A., Joshi, N., Julian, R., Kalashnikov, D., Kuang, Y., Leal, I., Lee, L., Lee, T.-W. E., Levine, S., Lu, Y., Michalewski, H., Mordatch, I., Pertsch, K., Rao, K., Reymann, K., Ryoo, M., Salazar, G., Sanketi, P., Sermanet, P., Singh, J., Singh, A., Soricut, R., Tran, H., Vanhoucke, V., Vuong, Q., Wahid, A., Welker, S., Wohlhart, P., Wu, J., Xia, F., Xiao, T., Xu, P., Xu, S., Yu, T., and Zitkovich, B. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. CoRR, abs/2307.15818, 2023. doi: 10.48550/arXiv.2307.15818. URL https://arxiv.org/abs/2307.15818.
  12. 12.Brynjolfsson, E. The Turing Trap: The Promise & Peril of Human-Like Artificial Intelligence. CoRR, abs/2201.04200, 2022. doi: 10.48550/arXiv.2201.04200. URL https://arxiv.org/abs/2201.04200.
  13. 13.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y. Sparks of Artificial General Intelligence: Early experiments with GPT-4. CoRR, abs/2303.12712, 2023. doi: 10.48550/arXiv:2303.12712. URL https://arxiv.org/abs/2303.12712.
  14. 14.Butlin, P., Long, R., Elmoznino, E., Bengio, Y., Birch, J., Constant, A., Deane, G., Fleming, S. M., Frith, C., Ji, X., Kanai, R., Klein, C., Lindsay, G., Michel, M., Mudrik, L., Peters, M. A. K., Schwitzgebel, E., Simon, J., and VanRullen, R. Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. CoRR, abs/2308.08708, 2023. doi: 10.48550/arXiv.2308.08708. URL https://arxiv.org/abs/2308.08708.
  15. 15.Campbell, M., Hoane, A. J., and Hsu, F.-h. Deep Blue. Artif. Intell., 134(1–2):57–83, jan 2002. ISSN 0004-3702. doi: 10.1016/S0004-3702(01)00129-1. URL https://doi.org/10.1016/S0004-3702(01)00129-1.
  16. 16.Chen, X., Wang, X., Changpinyo, S., and et al. PaLI: A Jointly-Scaled Multilingual Language-Image Model. CoRR, abs/2209.06794, 2023. doi: 10.48550/arXiv.2209.06794. URL https://arxiv.org/abs/2209.06794.
  17. 17.Chollet, F. On the measure of intelligence, 2019.
  18. 18.Christian, B. The Alignment Problem. W. W. Norton & Company, 2020.
  19. 19.Das, M. M., Saha, P., and Das, M. Which One is More Toxic? Findings from Jigsaw Rate Severity of Toxic Comments. CoRR, abs/2206.13284, 2022. doi: 10.48550/arXiv.2206.13284. URL https://arxiv.org/abs/2206.13284.
  20. 20.Dell’Acqua, F., McFowland, E., Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., and Lakhani, K. R. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School Technology & Operations Management Unit Working Paper Number 24-013, September 2023.
  21. 21.Demetriou, A. and Kazi, S. Self-awareness in g (with processing efficiency and reasoning). Intelligence, 34:297–317, 2006. doi: https://doi.org/10.1016/j.intell.2005.10.002.
  22. 22.Ellingrud, K., Sanghvi, S., Dandona, G. S., Madgavkar, A., Chui, M., White, O., and Hasebe, P. Generative AI and the future of work in America. McKinsey Institute Global Report, July 2023. URL https://www.mckinsey.com/mgi/our-research/generative-ai-and-the-future-of-work-in-america.
  23. 23.Eloundou, T., Manning, S., Mishkin, P., and Rock, D. Gpts are gpts: An early look at the labor market impact potential of large language models, 2023.
  24. 24.Englebart, D. Augmenting human intellect: A conceptual framework. October 1962. URL https://www.dougengelbart.org/pubs/papers/scanned/Doug_Engelbart-AugmentingHumanIntellect.pdf.
  25. 25.for AI Safety, C. Statement on AI Risk, 2023. URL https://www.safe.ai/statement-on-ai-risk.
  26. 26.Gardner, H. E. Frames of Mind: The Theory of Multiple Intelligences. Basic Books, 2011.
  27. 27.Goertzel, B. Artificial General Intelligence: Concept, State of the Art, and Future Prospects. Journal of Artificial General Intelligence, 01 2014. doi: 10.2478/jagi-2014-0001.
  28. 28.Goldwasser, S., Gruber, D. F., Kalai, A. T., and Paradise, O. A theory of unsupervised translation motivated by understanding animal communication, 2023.
  29. 29.Google. Google Assistant, your own personal Google. URL https://assistant.google.com/. accessed on October 20, 2023.
  30. 30.Grammarly, 2023. URL https://www.grammarly.com/.
  31. 31.Gubrud, M. Nanotechnology and International Security. Fifth Foresight Conference on Molecular Nanotechnology, November 1997.
  32. 32.IBM. IBM Watson. URL https://www.ibm.com/watson. accessed on October 20, 2023.
  33. 33.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Zˇ´ıdek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P., and Hassabis, D. Highly Accurate Protein Structure Prediction with AlphaFold. Nature, 596:583–589, 2021. doi: 10.1038/s41586-021-03819-2.
  34. 34.Kenton, Z., Everitt, T., Weidinger, L., Gabriel, I., Mikulik, V., and Irving, G. Alignment of Language Agents. CoRR, abs/2103.14659, 2021. doi: 10.48550/arXiv.2103.14659. URL https://arxiv.org/abs/2103.14659.
  35. 35.Kissinger, H., Schmidt, E., and Huttenlocher, D. The Age of AI. Back Bay Books, November 2022.
  36. 36.Legg, S. Machine Super Intelligence. Doctoral Dissertation submitted to the Faculty of Informatics of the University of Lugano, June 2008.
  37. 37.Legg, S. Twitter (now ”X”), May 2022. URL https://twitter.com/ShaneLegg/status/1529483168134451201. Accessed on October 12, 2023.
  38. 38.Liang, P., Bommasani, R., Lee, T., and et al. Holistic Evaluation of Language Models. CoRR, abs/2211.09110, 2023. doi: 10.48550/arXiv.2211.09110. URL https://arxiv.org/abs/2211.09110.
  39. 39.Lynch, S. AI Benchmarks Hit Saturation. Stanford Human-Centered Artificial Intelligence Blog, April 2023. URL https://hai.stanford.edu/news/ai-benchmarks-hit-saturation.
  40. 40.Marcus, G. Dear Elon Musk, here are five things you might want to consider about AGI. ”Marcus on AI” Substack, May 2022a. URL https://garymarcus.substack.com/p/dear-elon-musk-here-are-five-things?s=r.
  41. 41.Marcus, G. Twitter (now ”X”), May 2022b. URL https://twitter.com/GaryMarcus/status/1529457162811936768. Accessed on October 12, 2023.
  42. 42.McCarthy, J., Minsky, M., Rochester, N., and Shannon, C. A Proposal for The Dartmouth Summer Research Project on Artificial Intelligence. Dartmouth Workshop, 1955.
  43. 43.Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., and Gebru, T. Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency. ACM, jan 2019. doi: 10.1145/3287560.3287596. URL https://doi.org/10.1145%2F3287560.3287596.
  44. 44.Morris, M. R. Scientists’ Perspectives on the Potential for Generative AI in their Fields. CoRR, abs/2304.01420, 2023. doi: 10.48550/arXiv.2304.01420. URL https://arxiv.org/abs/2304.01420.
  45. 45.Morris, M. R., Cai, C. J., Holbrook, J., Kulkarni, C., and Terry, M. The Design Space of Generative Models. CoRR, abs/2304.10547, 2023. doi: 10.48550/arXiv.2304.10547. URL https://arxiv.org/abs/2304.10547.
  46. 46.Mustafa Suleyman and Michael Bhaskar. The Coming Wave: Technology, Power, and the 21st Century’s Greatest Dilemma. Crown, September 2023.
  47. 47.OpenAI. OpenAI Charter, 2018. URL https://openai.com/charter. Accessed October 12, 2023.
  48. 48.OpenAI. OpenAI: About, 2023. URL https://openai.com/about. Accessed October 12, 2023.
  49. 49.OpenAI. GPT-4 Technical Report. CoRR, abs/2303.08774, 2023. doi: 10.48550/arXiv.2303.08774. URL https://arxiv.org/abs/2303.08774.
  50. 50.Papakyriakopoulos, O., Watkins, E. A., Winecoff, A., Jazwi´nska, K., and Chattopadhyay, T. Qualitative Analysis for Human Centered AI. CoRR, abs/2112.03784, 2021. doi: 10.48550/arXiv.2112.03784. URL https://arxiv.org/abs/2112.03784.
  51. 51.Parasuraman, R., Sheridan, T., and Wickens, C. A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans, 30(3):286–297, 2000. doi: 10.1109/3468.844354.
  52. 52.Pichai, S. and Hassabis, D. Introducing gemini: our largest and most capable ai model, December 2023. URL https://blog.google/technology/ai/google-gemini-ai/.
  53. 53.Pressley, M., Borkowski, J., and Schneider, W. Cognitive strategies: Good strategy users coordinate metacognition and knowledge. Annals of Child Development, 4:89–129, 1987.
  54. 54.PromptBase. PromptBase: Prompt Marketplace. URL https://promptbase.com/. accessed on October 20, 2023.
  55. 55.Raji, I. D., Bender, E. M., Paullada, A., Denton, E., and Hanna, A. AI and the Everything in the Whole Wide World Benchmark. CoRR, abs/2111.15366, 2021. doi: 10.48550/arXiv.2111.15366. URL https://arxiv.org/abs/2111.15366.
  56. 56.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical Text-Conditional Image Generation with CLIP Latents. April 2022. URL https://cdn.openai.com/papers/dall-e-2.pdf.
  57. 57.Rauker, T., Ho, A., Casper, S., and Hadfield-Menell, D. Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks. CoRR, abs/2207.13243, 2023. doi: 10.48550/arXiv.2207.13243. URL https://arxiv.org/abs/2207.13243.
  58. 58.Richmond, J. Y. and McKinney, R. W. Biosafety in microbiological and biomedical laboratories, 2009.
  59. 59.Roy, N., Posner, I., Barfoot, T., Beaudoin, P., Bengio, Y., Bohg, J., Brock, O., Depatie, I., Fox, D., Koditschek, D., Lozano-Perez, T., Mansinghka, V., Pal, C., Richards, B., Sadigh, D., Schaal, S., Sukhatme, G., Therien, D., Toussaint, M., and de Panne, M. V. From Machine Learning to Robotics: Challenges and Opportunities for Embodied Intelligence. CoRR, abs/2110.15245, 2021. doi: 10.48550/arXiv.2110.15245. URL https://arxiv.org/abs/2110.15245.
  60. 60.SAE International. Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles, April 2021. URL https://www.sae.org/standards/content/j3016_202104. Accessed October 12, 2023.
  61. 61.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. CoRR, abs/2205.11487, 2022. doi: 10.48550/arXiv:2205.11487. URL https://arxiv.org/abs/2205.11487.
  62. 62.Schoenegger, P. and Park, P. S. Large language model prediction capabilities: Evidence from a real-world forecasting tournament, 2023.
  63. 63.Searle, J. R. Minds, Brains, and Programs. Behavioral and Brain Sciences, 3:417–424, 1980. doi: 10.1017/S0140525X00005756.
  64. 64.Serapio-Garc´ıa, G., Safdari, M., Crepy, C., Sun, L., Fitz, S., Romero, P., Abdulhai, M., Faust, A., and Mataric, M. Personality Traits in Large Language Models. CoRR, abs/2307.00184, 2023. doi: 10.48550/arXiv.2307.00184. URL https://arxiv.org/abs/2307.00184.
  65. 65.Shah, R., Freire, P., Alex, N., Freedman, R., Krasheninnikov, D., Chan, L., Dennis, M. D., Abbeel, P., Dragan, A., and Russell, S. Benefits of Assistance over Reward Learning, 2021. URL https://openreview.net/forum?id=DFIoGDZejIB.
  66. 66.Shanahan, M. Embodiment and the Inner Life. Oxford University Press, 2010.
  67. 67.Shanahan, M. The Technological Singularity. MIT Press, August 2015.
  68. 68.Sheridan, T. B. and Parasuraman, R. Human-automation interaction. Reviews of Human Factors and Ergonomics, 1(1):89–129, 2005. doi: 10.1518/155723405783703082. URL https://doi.org/10.1518/155723405783703082.
  69. 69.Sheridan, T. B., Verplank, W. L., and Brooks, T. Human/computer control of undersea teleoperators. In NASA. Ames Res. Center The 14th Ann. Conf. on Manual Control, 1978.
  70. 70.Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., Ho, L., Siddarth, D., Avin, S., Hawkins, W., Kim, B., Gabriel, I., Bolina, V., Clark, J., Bengio, Y., Christiano, P., and Dafoe, A. Model evaluation for extreme risks. CoRR, abs/2305.15324, 2023. doi: 10.48550/arXiv.2305.15324. URL https://arxiv.org/abs/2305.15324.
  71. 71.Shneiderman, B. Human-centered artificial intelligence: Reliable, safe & trustworthy, 2020. URL https://arxiv.org/abs/2002.04087v1.
  72. 72.Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. Mastering the Game of Go with Deep Neural Networks and Tree Search. Nature, 529:484–489, 2016. doi: 10.1038/nature16961.
  73. 73.Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D. Mastering the Game of Go Without Human Knowledge. Nature, 550:354–359, 2017. doi: 10.1038/nature24270.
  74. 74.Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D. A General Reinforcement Learning Algorithm that Masters Chess, Shogi, and Go through Self-play. Science, 362(6419):1140–1144, 2018. doi: 10.1126/science.aar6404. URL https://www.science.org/doi/abs/10.1126/science.aar6404.
  75. 75.Srivastava, A., Rastogi, A., Rao, A., and et al. Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models. CoRR, abs/2206.04615, 2023. doi: 10.48550/arXiv.2206.04615. URL https://arxiv.org/abs/2206.04615.
  76. 76.Stockfish. Stockfish - Open Source Chess Engine, 2023. URL https://stockfishchess.org/.
  77. 77.Tang, J., LeBel, A., Jain, S., and Huth, A. G. Semantic Reconstruction of Continuous Language from Non-invasive Brain Recordings. Nature Neuroscience, 26:858–866, 2023. doi: 10.1038/s41593-023-01304-9.
  78. 78.Terry, M., Kulkarni, C., Wattenberg, M., Dixon, L., and Morris, M. R. AI Alignment in the Design of Interactive AI: Specification Alignment, Process Alignment, and Evaluation Support. CoRR, abs/2311.00710, 2023. doi: 10.48550/arXiv.2311.00710. URL https://arxiv.org/abs/2311.00710.
  79. 79.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open Foundation and Fine-Tuned Chat Models, 2023.
  80. 80.Tullis, J. and Fraundorf, S. Predicting others’ memory performance: The accuracy and bases of social metacognition. Journal of Memory and Language, 95:124–137, 2017. doi: https://doi.org/10.1016/j.jml.2017.03.003.
  81. 81.Turing, A. Computing Machinery and Intelligence. Mind, LIX:433–460, October 1950. URL https://doi.org/10.1093/mind/LIX.236.433.
  82. 82.Varadi, M., Anyango, S., Deshpande, M., Nair, S., Natassia, C., Yordanova, G., Yuan, D., Stroe, O., Wood, G., Laydon, A., Zˇ´ıdek, A., Green, T., Tunyasuvunakool, K., Petersen, S., Jumper, J., Clancy, E., Green, R., Vora, A., Lutfi, M., Figurnov, M., Cowie, A., Hobbs, N., Kohli, P., Kleywegt, G., Birney, E., Hassabis, D., and Velankar, S. AlphaFold Protein Structure Database: Massively Expanding the Structural Coverage of Protein-Sequence Space with High-Accuracy Models. Nucleic Acids Research, 50:D439–D444, 11 2021. ISSN 0305-1048. doi: 10.1093/nar/gkab1061. URL https://doi.org/10.1093/nar/gkab1061.
  83. 83.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention Is All You Need. CoRR, abs/1706.03762, 2023. doi: 10.48550/arXiv.1706.03762. URL https://arxiv.org/abs/1706.03762.
  84. 84.Veerabadran, V., Goldman, J., Shankar, S., and et al. Subtle Adversarial Image Manipulations Influence Both Human and Machine Perception. Nature Communications, 14, 2023. doi: 10.1038/s41467-023-40499-0.
  85. 85.Webb, T., Holyoak, K. J., and Lu, H. Emergent Analogical Reasoning in Large Language Models. Nature Human Behavior, 7:1526–1541, 2023. URL https://doi.org/10.1038/s41562-023-01659-w.
  86. 86.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. Emergent Abilities of Large Language Models. CoRR, abs/2206.07682, 2022. doi: 10.48550/arXiv.2206.07682. URL https://arxiv.org/abs/2206.07682.
  87. 87.Weizenbaum, J. ELIZA—a Computer Program for the Study of Natural Language Communication between Man and Machine. Commun. ACM, 9(1):36–45, jan 1966. ISSN 0001-0782. doi: 10.1145/365153.365168. URL https://doi.org/10.1145/365153.365168.
  88. 88.Wiggers, K. OpenAI Disbands its Robotics Research Team. VentureBeat, July 2021.
  89. 89.Wikipedia. Eugene Goostman - Wikipedia, The Free Encyclopedia. https://en.wikipedia.org/wiki/Eugene Goostman, 2023a. Accessed October 12, 2023.
  90. 90.Wikipedia. Turing Test: Weaknesses — Wikipedia, The Free Encyclopedia. https://en.wikipedia.org/wiki/Turing test, 2023b. Accessed October 12, 2023.
  91. 91.Winograd, T. Procedures as a Representation for Data in a Computer Program for Understanding Natural Language. MIT AI Technical Reports, 1971.
  92. 92.Wozniak, S. Could a Computer Make a Cup of Coffee? Fast Company interview: https://www.youtube.com/watch?v=MowergwQR5Y, 2010.
  93. 93.Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.-C., Liu, Z., and Wang, L. The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision). CoRR, abs/2309.17421, 2023. doi: 10.48550/arXiv.2309.17421. URL https://arxiv.org/abs/2309.17421.
  94. 94.Zamfirescu-Pereira, J., Wong, R. Y., Hartmann, B., and Yang, Q. Why johnny can’t prompt: How non-ai experts try (and fail) to design llm prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9781450394215. doi: 10.1145/3544548.3581388. URL https://doi.org/10.1145/3544548.3581388.
  95. 95.Zwetsloot, R. and Dafoe, A. Thinking about Risks from AI: Accidents, Misuse and Structure. Lawfare, 11:2019, 2019.

Citation

MLA
Morris, M. R., et al. “Levels of AGI for Operationalizing Progress on the Path to AGI”. Proceedings of ICML 2024, 2023, http://arxiv.org/abs/2311.02462v5.
APA
Morris, M. R., Sohl-Dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., & Legg, S. (2023). Levels of AGI for Operationalizing Progress on the Path to AGI. Proceedings of ICML 2024. http://arxiv.org/abs/2311.02462v5
Chicago
Morris, M. R., J. Sohl-Dickstein, N. Fiedel, et al. 2023. “Levels of AGI for Operationalizing Progress on the Path to AGI”. Proceedings of ICML 2024. http://arxiv.org/abs/2311.02462v5.
Harvard
Morris, M.R. et al. (2023) “Levels of AGI for Operationalizing Progress on the Path to AGI”, Proceedings of ICML 2024 [Preprint]. Available at: http://arxiv.org/abs/2311.02462v5.
Vancouver
1. Morris MR, Sohl-Dickstein J, Fiedel N, Warkentin T, Dafoe A, Faust A, Farabet C, Legg S (2023) Levels of AGI for Operationalizing Progress on the Path to AGI. Proceedings of ICML 2024

BibTeX

@article{morris2023levels,
  title = {Levels of AGI for Operationalizing Progress on the Path to AGI},
  author = {Morris, Meredith Ringel and Sohl-Dickstein, Jascha and Fiedel, Noah and Warkentin, Tris and Dafoe, Allan and Faust, Aleksandra and Farabet, Clement and Legg, Shane},
  year = {2023},
  journal = {Proceedings of ICML 2024},
  url = {http://arxiv.org/abs/2311.02462v5},
  eprint = {2311.02462}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/