General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin ZhouLorenzo PacchiardiFernando Martínez-PlumedKatherine M. CollinsYael Moros-DavalSeraphina ZhangQinlin ZhaoYitian HuangLuning SunJonathan E. Prunty
Introduces a rubric-based measurement framework that maps task demands onto non-saturating general scales, enabling researchers to interpret common benchmarks and accurately predict large language model performance on novel, out-of-distribution tasks.
Current artificial intelligence evaluation relies heavily on aggregate benchmark scores that report average percentage accuracy across collections of test questions. However, these aggregate numbers do not explain why large language models fail at basic tasks while solving complex ones, nor do they predict whether a system will succeed on a specific new problem in real-world deployment. As modern AI benchmarks rapidly saturate and suffer from data contamination, decision-makers lack reliable, interpretable tools to assess AI capabilities and risks before deployment.
The article demonstrates a new evaluation methodology based on general, absolute demand scales that map the cognitive, knowledge, and structural requirements of tasks. The main objective is to establish an automated measurement framework that explains what benchmarks actually measure, profiles the distinct abilities of AI systems independently of other models, and accurately predicts performance on individual task instances both within and outside familiar test distributions.
To achieve this, the authors established 18 open-ended demand scales spanning cognitive abilities, scientific and everyday knowledge, and extraneous task factors such as atypicality, volume, and unguessability. Using automated large language model annotators validated against expert human consensus, the authors scored 16,108 curated task instances across 20 established benchmarks and 63 tasks. They then evaluated 15 commercial and open-weight models—ranging from small distilled networks to frontier reasoning systems—by mapping their success rates against task demands to extract individual capability profiles and train instance-level predictive models called assessors.
The analysis revealed several critical findings. First, existing benchmarks frequently lack specificity and sensitivity; many tests incorporate heavy extraneous demands or fail to span the difficulty levels needed to test the capabilities they claim to assess. Second, model capability profiles show clear architectural divergences: scaling parameter size primarily expands domain knowledge, whereas chain-of-thought reasoning models dramatically boost quantitative reasoning, logical deduction, and social cognition even at smaller parameter sizes. Third, lightweight predictive assessors trained on the 19 demand dimensions accurately forecast instance-level AI success, achieving an average discriminative score of 0.84 and near-perfect calibration error of 0.01 in-distribution. In out-of-distribution tests on entirely unseen benchmarks, demand-based predictors substantially outperformed complex baseline methods like text fine-tuning and embeddings, dropping only moderately to a score of 0.75 while baselines collapsed.
These findings indicate that task difficulty can be treated as an absolute, measurable property rather than a shifting statistical artifact. By decoupling capability measurement from specific benchmark populations, organizations can identify exact operational boundaries, implement automated routing between smaller and larger models, and enforce proactive rejection rules when task demands exceed system capabilities. This provides an interpretable mechanism to manage deployment costs, safety boundaries, and compliance risks without relying on black-box heuristics.
Organizations evaluating or deploying frontier AI should adopt capability profiling to audit both internal test suites and third-party systems. For operational pipelines, deploying demand-based assessors offers a cost-effective way to route queries and block high-risk failures. Future efforts should expand the rubric library to cover multimodal tasks and agent workflows, while curating more balanced datasets that include higher difficulty levels beyond current ceiling thresholds.
The methodology currently focuses on text-only tasks and exhibits some predictive degradation in out-of-distribution settings due to sparse coverage in extreme difficulty tiers and agent-oriented domains. Additionally, the approach relies on automated model grading, which introduces minor label noise. Nevertheless, the high agreement between human experts and automated annotators provides strong confidence in using general demand scales as a robust foundation for modern AI evaluation.
- Paper: Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench authors (2022). BIG-bench establishes the foundational multi-task benchmarking landscape and empirical limits of aggregate scaling metrics that general scales explicitly seek to decompose and predict.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). MMLU serves as a primary standard multitask benchmark whose sensitivity, specificity, and underlying task demands are directly dissected by the proposed general scales.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). This survey provides a comprehensive taxonomy of existing LLM evaluation methodologies and their lack of cross-task transferability, framing the core challenge the source paper solves.
- Paper: Position: Levels of AGI for Operationalizing Progress on the Path to AGI, Meredith Ringel Morris et al. (2024). This paper formalizes tiered capability and generality levels for evaluating AI progress, providing conceptual motivation for unsaturating ability profiles.
- Paper: CogBench: a large language model walks into a psychology lab, Julian Coda-Forno et al. (2024). CogBench introduces cognitive and metacognitive profiling of LLMs using psychology-based metrics, laying groundwork for characterising model ability beyond raw accuracy.
- Paper: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, Yubo Wang et al. (2024). MMLU-Pro highlights the saturation and reasoning limitations of classic benchmarks, directly motivating the need for demand-calibrated general evaluation scales.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). This survey identifies systemic flaws and poor explanatory power in legacy generation evaluations, establishing the methodological deficit addressed by rubric-based scales.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). This work analyzes how chain-of-thought prompting unlocks latent model reasoning on hard benchmark tasks, a phenomenon quantified and modeled via the source's general scales.
- Paper: Dynamic Evaluation of Large Language Models by Meta Probing Agents, Kaijie Zhu et al. (2024). This study introduces psychometric meta-probing across cognitive dimensions, providing a key precedent for extracting multidimensional demand profiles from benchmark instances.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey outlines the mechanics and rubric structures of LLM-as-a-judge, which underpin the automated grading and demand-leveling rubrics deployed in the source paper.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). BiGGen Bench applies instance-specific, rubric-grounded evaluator models across nine core capabilities, offering a concrete implementation of structured evaluation rubrics.
- Paper: Are We Done with MMLU?, Aryo Pradipta Gema et al. (2025). This paper investigates ground-truth errors and noise across MMLU subjects, complementing the source's analysis of benchmark sensitivity and specificity.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). JudgeBench establishes an objective benchmark to evaluate automated LLM judges on complex reasoning, validating the reliability of automated scoring mechanisms.
- Paper: MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations, Kaixuan Huang et al. (2025). MATH-Perturb measures the out-of-distribution robustness of mathematical reasoning under structural shifts, aligning with the predictive evaluation of instance demands.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). Agent-as-a-Judge extends automated multidimensional rubric evaluation to multi-step agent trajectories and software engineering workflows.
- Paper: Benchmarking AI Agents for Addressing Scientific Challenges Across Scales, Tianyu Liu et al. (2026). SciAgentArena operationalizes multi-scale capability profiling for scientific AI agents across diverse multi-step research domains.
- Paper: SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations, Shuaiqi Wang et al. (2026). SynAE formulates multidimensional validity and fidelity metrics for synthetic agent evaluation, extending ability and demand profiling to tool-calling agents.
