Built independently by an author, for readers. Read the story and support ChapterPal

keyword

AGI benchmarks

AGI benchmarks are tests or sets of tasks used to evaluate how well AI systems perform across a broad range of abilities, rather than on a single narrow task. They help measure both the level of performance and the breadth of capabilities, offering a way to compare systems and track progress toward artificial general intelligence.

1 item

Position: Levels of AGI for Operationalizing Progress on the Path to AGI

Position: Levels of AGI for Operationalizing Progress on the Path to AGI

Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clément Farabet, Shane Legg

OrganizationsGoogle

Why you should read this

Establishes a structured framework that classifies artificial general intelligence across distinct tiers of capability, generality, and autonomy, providing clear criteria to evaluate current models and safely track progress toward human-level systems.

We propose a framework for classifying the capabilities and behavior of Artificial General Intelligence (AGI) models and their precursors. This framework introduces levels of AGI performance, generality, and autonomy, providing a common language to compare models, assess risks, and measure progress along the path to AGI. To develop our framework, we analyze existing definitions of AGI, and distill six principles that a useful ontology for AGI should satisfy. With these principles in mind, we propose “Levels of AGI” based on depth (performance) and breadth (generality) of capabilities, and reflect on how current systems fit into this ontology. We discuss the challenging requirements for future benchmarks that quantify the behavior and capabilities of AGI models against these levels. Finally, we discuss how these levels of AGI interact with deployment considerations such as autonomy and risk, and emphasize the importance of carefully selecting Human-AI Interaction paradigms for responsible and safe deployment of highly capable AI systems.

Added

2026-09-26