General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Lexin ZhouLorenzo PacchiardiFernando Martínez-PlumedKatherine M. CollinsYael Moros-DavalSeraphina ZhangQinlin ZhaoYitian HuangLuning SunJonathan E. Prunty

article2025arXiv38 citations

Introduces a rubric-based measurement framework that maps task demands onto non-saturating general scales, enabling researchers to interpret common benchmarks and accurately predict large language model performance on novel, out-of-distribution tasks.

Listen

Current artificial intelligence evaluation relies heavily on aggregate benchmark scores that report average percentage accuracy across collections of test questions. However, these aggregate numbers do not explain why large language models fail at basic tasks while solving complex ones, nor do they predict whether a system will succeed on a specific new problem in real-world deployment. As modern AI benchmarks rapidly saturate and suffer from data contamination, decision-makers lack reliable, interpretable tools to assess AI capabilities and risks before deployment.

The article demonstrates a new evaluation methodology based on general, absolute demand scales that map the cognitive, knowledge, and structural requirements of tasks. The main objective is to establish an automated measurement framework that explains what benchmarks actually measure, profiles the distinct abilities of AI systems independently of other models, and accurately predicts performance on individual task instances both within and outside familiar test distributions.

To achieve this, the authors established 18 open-ended demand scales spanning cognitive abilities, scientific and everyday knowledge, and extraneous task factors such as atypicality, volume, and unguessability. Using automated large language model annotators validated against expert human consensus, the authors scored 16,108 curated task instances across 20 established benchmarks and 63 tasks. They then evaluated 15 commercial and open-weight models—ranging from small distilled networks to frontier reasoning systems—by mapping their success rates against task demands to extract individual capability profiles and train instance-level predictive models called assessors.

The analysis revealed several critical findings. First, existing benchmarks frequently lack specificity and sensitivity; many tests incorporate heavy extraneous demands or fail to span the difficulty levels needed to test the capabilities they claim to assess. Second, model capability profiles show clear architectural divergences: scaling parameter size primarily expands domain knowledge, whereas chain-of-thought reasoning models dramatically boost quantitative reasoning, logical deduction, and social cognition even at smaller parameter sizes. Third, lightweight predictive assessors trained on the 19 demand dimensions accurately forecast instance-level AI success, achieving an average discriminative score of 0.84 and near-perfect calibration error of 0.01 in-distribution. In out-of-distribution tests on entirely unseen benchmarks, demand-based predictors substantially outperformed complex baseline methods like text fine-tuning and embeddings, dropping only moderately to a score of 0.75 while baselines collapsed.

These findings indicate that task difficulty can be treated as an absolute, measurable property rather than a shifting statistical artifact. By decoupling capability measurement from specific benchmark populations, organizations can identify exact operational boundaries, implement automated routing between smaller and larger models, and enforce proactive rejection rules when task demands exceed system capabilities. This provides an interpretable mechanism to manage deployment costs, safety boundaries, and compliance risks without relying on black-box heuristics.

Organizations evaluating or deploying frontier AI should adopt capability profiling to audit both internal test suites and third-party systems. For operational pipelines, deploying demand-based assessors offers a cost-effective way to route queries and block high-risk failures. Future efforts should expand the rubric library to cover multimodal tasks and agent workflows, while curating more balanced datasets that include higher difficulty levels beyond current ceiling thresholds.

The methodology currently focuses on text-only tasks and exhibits some predictive degradation in out-of-distribution settings due to sparse coverage in extreme difficulty tiers and agent-oriented domains. Additionally, the approach relies on automated model grading, which introduces minor label noise. Nevertheless, the high agreement between human experts and automated annotators provides strong confidence in using general demand scales as a robust foundation for modern AI evaluation.

arXiv: 2503.06378Kinds-of-Intelligence-CFI/ADeLe-AIEvaluation
Cover for General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Abstract

Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activities. So far, benchmarking has guided progress in AI, but it has offered limited explanatory and predictive power for general-purpose AI systems, given the low transferability across diverse tasks. In this paper, we introduce general scales for AI evaluation that can explain what common AI benchmarks really measure, extract ability profiles of AI systems, and predict their performance for new task instances, in- and out-of-distribution. Our fully-automated methodology builds on 18 newly-crafted rubrics that place instance demands on general scales that do not saturate. Illustrated for 15 large language models and 63 tasks, high explanatory power is unleashed from inspecting the demand and ability profiles, bringing insights on the sensitivity and specificity exhibited by different benchmarks, and how knowledge, metacognition and reasoning are affected by model size, chain-of-thought and distillation. Surprisingly, high predictive power at the instance level becomes possible using these demand levels, providing superior estimates over black-box baseline predictors based on embeddings or finetuning, especially in out-of-distribution settings (new tasks and new benchmarks). The scales, rubrics, battery, techniques and results presented here represent a major step for AI evaluation, underpinning the reliable deployment of AI in the years ahead. (Collaborative platform: this https URL.)

Table of Contents

  • 1 Introduction
  • 2 AI Evaluation at Scale
  • 2.1 General Scales and Automated Annotation
  • 2.2 Slicing the Demand-Ability Space
  • 3 Results
  • 3.1 Annotation and Scales Analysis: Distinguishing Levels and Dimensions
  • 3.2 Explanatory Power Analysis: Profiling Benchmark Demands
  • 3.3 Explanatory Power Analysis: Profiling LLM Abilities
  • 3.4 Predictive Power Analysis: Anticipating Performance with Assessors
  • 4 Discussion
  • References
  • 5 Methods
  • 5.1 Scales and Rubrics
  • 5.2 LLM Annotators
  • 5.3 Inter-rater Analysis
  • 5.4 Benchmark Battery: Instance Selection and Curation
  • 5.5 Subject LLMs and Grading
  • 5.6 Assessors and Metrics
  • 5.7 Slicing Methods for Characteristic Curves
  • 5.8 ADeLe-Light
  • 6 Acknowledgements
  • 7 Cost, Ethical and Safety Implications
  • 8 Appendix
  • 8.1 Related Work
  • 8.2 Scaling Curves of Model Abilities
  • 8.3 Calibration of Assessors
  • 8.4 Sources of Unpredictability
  • 8.5 Other Predictive Models
  • 8.5.1 Feature Importance
  • 8.5.2 Assessor Only with AT, UG, and VO
  • 8.5.3 Assessor without AT, UG, and VO
  • 8.5.4 Feature Grouping: Assessor with 11 Broad Dimensions
  • 8.5.5 Demand-based Assessor with Logistic Regression
  • 8.5.6 A Universal Assessor
  • 8.5.7 Algebraic Assessor
  • 8.6 SCCs for all models
  • 8.7 Glossary
  • 9 DeLeAn Rubric Set v.1.0
  • 9.1 Primordial
  • 9.2 Knowledge
  • 9.3 Extraneous

Knowls

  1. Knowl 1 — DeLeAn Rubric Taxonomy and 19 Demand Dimensions

    model/method

    The Demand-Level-Annotation (DeLeAn) Rubric Set v.1.0 defines a 19-dimensional system to quantify the cognitive demands and structural properties of task instances. The dimensions are divided into three categories:

    1. Primordial Cognitive Dimensions (11 dimensions, evaluated on a discrete scale from 0 to 5+):

      • Attention and Scan (AS): The requirement to focus on, scan for, or track specific target elements among distractors.
      • Verbal Comprehension (CEc): Understanding semantic content, implicit meanings, and complex concepts in text or other representations.
      • Verbal Expression (CEe): Generating and articulating ideas, coherent narratives, and domain-appropriate expression.
      • Conceptualisation, Learning and Abstraction (CL): Real-time inductive/analogical reasoning, pattern synthesis, and forming abstractions across domains.
      • Identifying Relevant Information (MCr): Metacognitive filtering to distinguish relevant information from noise or distractors during problem solving.
      • Critical Thinking Processes (MCt): Monitoring, regulating, and analyzing multi-step thought processes and evaluating competing arguments.
      • Calibrating Knowns and Unknowns (MCu): Metacognitive recognition of the boundaries of one's own knowledge and certainty.
      • Mind Modelling and Social Cognition (MS): Attributing mental states (beliefs, desires, intentions) to other agents and higher-order Theory of Mind.
      • Logical Reasoning (QLl): Deductive reasoning, rule application, algorithmic steps, and premise chaining.
      • Quantitative Reasoning (QLq): Numerical operations, mathematical transformations, and reasoning with quantities.
      • Spatio-physical Reasoning (SNs): Understanding spatial relationships, 2D/3D mental rotations, and predicting physical dynamics.
    2. Knowledge Dimensions (5 dimensions, evaluated on a discrete scale from 0 to 5+ calibrated against formal education tiers from primary school to postgraduate levels):

      • Applied Sciences (KNa): Medicine, law, engineering, agriculture, business, and education.
      • Customary Everyday Knowledge (KNc): Societal, cultural, and daily-life knowledge acquired outside formal schooling.
      • Formal Sciences (KNf): Mathematics, logic, statistics, and computer science.
      • Natural Sciences (KNn): Physics, chemistry, biology, astronomy, and earth sciences.
      • Social Sciences and Humanities (KNs): History, sociology, psychology, philosophy, literature, and art.
    3. Extraneous Dimensions (3 dimensions capturing task design artifacts independent of cognitive demands):

      • Atypicality (AT, scale 0 to 5+): The degree of novelty/uniqueness of the instance, diagnosing data contamination and memorization.
      • Volume (VO, scale 0 to 5+): Proportional to the logarithm of the time an expert human needs to complete the task, diagnosing task amalgamation.
      • Unguessability (UG, scale 0% to 100%): One minus the probability of answering correctly by random guess or naive choice (e.g., 75%75\% for a 4-option multiple-choice item, 100%100\% for open-ended items), diagnosing task funnelling.
  2. Knowl 2 — Ratio Scale Calibration of Cognitive Demand Levels

    definition

    The demand levels in the DeLeAn framework are defined on an absolute ratio scale in the range [0,∞)[0, \infty), discretized for practical annotation into levels l∈{0,1,2,3,4,5+}l \in \{0, 1, 2, 3, 4, 5+\}. A task instance is assigned demand level ll if ll is the highest integer such that, in at least 95%95\% of random samples of n=10ln = 10^l individuals from the human population, there is at least one correct response:

    • Level 0 (None): n=100=1n = 10^0 = 1 (trivial or baseline task requiring no domain capability).
    • Level 1 (Very Low): n=101=10n = 10^1 = 10 individuals (elementary school tier).
    • Level 2 (Low): n=102=100n = 10^2 = 100 individuals (middle school tier).
    • Level 3 (Intermediate): n=103=1,000n = 10^3 = 1{,}000 individuals (high school tier).
    • Level 4 (High): n=104=10,000n = 10^4 = 10{,}000 individuals (undergraduate tier).
    • Level 5+ (Very High): n≥105=100,000n \ge 10^5 = 100{,}000 individuals (postgraduate or domain-expert tier).

    Under a logistic response model, this formulation ensures ratio scale properties where a doubling of the demand level halves the log-odds of a correct response.

  3. Knowl 3 — Dominant Slices Procedure for Extracting Subject Characteristic Curves

    algorithm

    To evaluate an AI system along a single cognitive demand dimension without being confounded by other capabilities, the dominant slices procedure filters multi-dimensional task evaluations to construct a 1D Subject Characteristic Curve (SCC):

    Input: Evaluation dataset D={(xk,yk)}k=1ND = \{(x_k, y_k)\}_{k=1}^N where xk=(dk,1,…,dk,18)∈{0,1,2,3,4,5}18x_k = (d_{k,1}, \dots, d_{k,18}) \in \{0, 1, 2, 3, 4, 5\}^{18} is the demand vector and yk∈{0,1}y_k \in \{0, 1\} is the system binary correctness; target dimension index i∈{1,…,18}i \in \{1, \dots, 18\}.
    Output: Subject Characteristic Curve parameters (β0,β1)(\beta_0, \beta_1) and dimension ability aia_i.
    for each demand level l∈{1,2,3,4,5}l \in \{1, 2, 3, 4, 5\} do
        Select dominant slice instances: Di,l={(xk,yk)∈D∣dk,i=l and max⁡j≠idk,j≤l}D_{i,l} = \{ (x_k, y_k) \in D \mid d_{k,i} = l \text{ and } \max_{j \neq i} d_{k,j} \le l \}
        Calculate empirical success probability: pl=1∣Di,l∣∑(xk,yk)∈Di,lykp_l = \frac{1}{|D_{i,l}|} \sum_{(x_k, y_k) \in D_{i,l}} y_k
        Assign bin weight wl=1.0w_l = 1.0 if ∣Di,l∣≥100|D_{i,l}| \ge 100, else wl=∣Di,l∣100w_l = \frac{|D_{i,l}|}{100}
    end for
    Add anchor point (lanchor,panchor)=(20,0)(l_{\text{anchor}}, p_{\text{anchor}}) = (20, 0) with total weight wanchor=∑l=15wlw_{\text{anchor}} = \sum_{l=1}^5 w_l
    Fit a 2-parameter logistic curve P(y=1∣l)=11+exp⁡(−(β0+β1l))P(y=1 \mid l) = \frac{1}{1 + \exp(-(\beta_0 + \beta_1 l))} to points {(l,pl)}l=15∪{(20,0)}\{(l, p_l)\}_{l=1}^5 \cup \{(20, 0)\} via weighted least squares
    Compute ability ai=−β0β1a_i = -\frac{\beta_0}{\beta_1} (the demand level where P(y=1∣l)=0.5P(y=1 \mid l) = 0.5, which equals the area under the curve for l≥0l \ge 0)
    return (β0,β1),ai(\beta_0, \beta_1), a_i
  4. Knowl 4 — Non-Populational Cognitive Ability Metric

    definition

    The cognitive ability ai∈[0,∞)a_i ∈ [0, \infty) of an AI system for a given demand dimension ii is defined as the demand level at which the system's Subject Characteristic Curve achieves a success probability of 0.50.5, assuming no other demand dimension dominates the task. Mathematically, for a logistic characteristic curve P(success∣l)=11+exp⁡(−(β0+β1l))P(\text{success} \mid l) = \frac{1}{1 + \exp(-(\beta_0 + \beta_1 l))}, the ability is:

    ai=−β0β1=∫0∞P(success∣l) dla_i = -\frac{\beta_0}{\beta_1} = \int_{0}^{\infty} P(\text{success} \mid l) \, dl

    This metric is strictly non-populational: unlike Item Response Theory (IRT), Principal Component Analysis (PCA), or Factor Analysis, aia_i is computed solely from the evaluation of the individual AI system on an absolute scale and remains completely invariant to the scores, capabilities, or presence of other AI systems in the evaluation pool.

  5. Knowl 5 — Demand-Based Algebraic Assessor Equation

    equation

    Given an AI system's 18-dimensional ability profile vector a=(a1,…,a18)\mathbf{a} = (a_1, \dots, a_{18}) and a task instance characterized by an 18-dimensional demand profile vector d=(d1,…,d18)\mathbf{d} = (d_1, \dots, d_{18}) and unguessability UG∈[0,100]\text{UG} \in [0, 100], the predicted success probability S∈[0,1]S \in [0, 1] of the system on that instance is computed algebraically without machine learning training via the generalized mean:

    S=(119(∑i=118[σ(ai−di)]r+(100−UG100)r))1/rS = \left( \frac{1}{19} \left( \sum_{i=1}^{18} [\sigma(a_i - d_i)]^r + \left(\frac{100 - \text{UG}}{100}\right)^r \right) \right)^{1/r}

    where σ(z)=11+exp⁡(−z)\sigma(z) = \frac{1}{1 + \exp(-z)} is the standard logistic function, and rr is the generalized mean exponent controlling dimensional compensatoriness:

    • As r→0r \to 0, SS converges to the geometric mean, which emphasizes the weakest capability relative to demands.
    • r=0.25r = 0.25 yields the highest discriminative performance (AUROC).
    • r=0r = 0 (geometric mean) yields optimal calibration (lowest Expected Calibration Error).
  6. Knowl 6 — Demand-Based Instance-Level Performance Predictors (Assessors)

    model/method

    An assessor is an external meta-model trained to predict the binary success or failure (y∈{0,1}y \in \{0, 1\}) of a subject LLM on individual task instances before query execution. The paper introduces and compares three assessor architectures:

    1. Demand-Based Assessor (Random Forest): A Random Forest classifier trained on the 19-dimensional vector consisting of the 18 DeLeAn demand levels plus the Unguessability score (UG). The minimum samples per split is tuned over {2,50,200}\{2, 50, 200\} using in-distribution cross-validation.
    2. Embeddings-Based Assessor (Random Forest): A Random Forest classifier trained on dense representations obtained by averaging pre-computed 300-dimensional GloVe word embeddings over the raw text of each question.
    3. Fine-Tuned Language Model Assessor (LLaMA): An end-to-end LLaMA-3.1-8B model equipped with a linear classification head, fine-tuned on raw question text using LoRA, NF4 quantization, learning rate ∈[2×10−5,1×10−4]\in [2\times 10^{-5}, 1\times 10^{-4}], batch size 16, weight decay 0.01, for 3 epochs.
  7. Knowl 7 — Instance-Level Performance Predictability Across ID, Task OOD, and Benchmark OOD

    empirical result

    Evaluation of 15 LLMs across the ADeLe battery under 10-fold In-Distribution (ID), Task Out-of-Distribution (Task OOD), and Benchmark Out-of-Distribution (Benchmark OOD) settings demonstrates that the 19-dimensional demand-based Random Forest assessor matches or exceeds black-box predictors while exhibiting superior calibration and computational efficiency:

    Evaluation Setting Demands (RF) Embeddings (GloVe RF) Finetuning (LLaMA-8B)
    AUROC ↑\uparrow ECE ↓\downarrow AUROC ↑\uparrow ECE ↓\downarrow AUROC ↑\uparrow ECE ↓\downarrow
    In-Distribution (ID) 0.839 0.011 0.805 0.032 0.840 0.043
    Task OOD 0.810 0.022 0.740 0.047 0.788 0.075
    Benchmark OOD 0.747 0.037 0.480 0.114 0.692 0.121
    Compute Cost (per subject) 4 seconds (CPU) 160 seconds (CPU) 300 hours (V100 GPU)

    Values represent accuracy-weighted averages across all 15 evaluated LLMs. In Benchmark OOD, embedding baselines collapse to near-random discrimination (AUROC 0.480), and fine-tuned LLaMA drops significantly (AUROC 0.692, ECE 0.121), while demand-based Random Forest maintains robust discrimination (AUROC 0.747) and superior calibration (ECE 0.037).

  8. Knowl 8 — Feature Importance of Demand Dimensions in Instance-Level Prediction

    empirical result

    Permutation feature importance analysis across all trained demand-based Random Forest assessors shows that all 19 dimensions contribute positively (importance >0.02> 0.02). Applying the elbow method identifies the six most influential dimensions for performance prediction:

    1. Calibrating Knowns and Unknowns (MCu) (importance ≈0.111\approx 0.111)
    2. Unguessability (UG) (importance ≈0.100\approx 0.100)
    3. Knowledge of Formal Sciences (KNf) (importance ≈0.084\approx 0.084)
    4. Conceptualisation, Learning and Abstraction (CL) (importance ≈0.078\approx 0.078)
    5. Logical Reasoning (QLl) (importance ≈0.074\approx 0.074)
    6. Identifying Relevant Information (MCr) (importance ≈0.066\approx 0.066)

    A reduced Random Forest assessor trained exclusively on these six dimensions retains high predictive power across all splits:

    • In-Distribution: weighted AUROC = 0.811, ECE = 0.012
    • Task OOD: weighted AUROC = 0.784, ECE = 0.031
    • Benchmark OOD: weighted AUROC = 0.742, ECE = 0.056
  9. Knowl 9 — ADeLe and ADeLe-Light Benchmark Battery Curation

    experimental setup

    The Annotated Demand Levels (ADeLe) battery v.1.0 was curated from 2024 proceedings of top machine learning (ICML, NeurIPS, ICLR) and NLP (ACL, EMNLP, NAACL) conferences across 20 benchmarks and 63 tasks:

    1. Inclusion Criteria: Frontier LLM accuracy <75%< 75\% (preventing triviality), objective verifiable ground truth, no AI-generated benchmark text, and multiple-choice (≥4\ge 4 options) or open-ended format.
    2. Instance Sampling & GPT-4o Quality Filtering: 500 instances per task were sampled (21,996 total). GPT-4o scored factual accuracy, objectivity, and ambiguity on 1–5 Likert scales. Items scoring 1 on any metric were pruned (16% removed), leaving 18,291 instances.
    3. Dual-LLM Grading Verification: Outputs of subject models were graded on a 1–5 scale by both GPT-4o and Claude-3.5-Sonnet. Instances where graders disagreed (one giving ≥4\ge 4 while the other ≤2\le 2) were removed (12% reduction), resulting in the final ADeLe v.1.0 battery of 16,108 instances (98% human-verified accuracy on sample).
    4. ADeLe-Light v.1.0: A compact subset constructed by finding the k=10k=10 nearest neighbors in 19D demand space. Instances with Average Squared Distance (ASD) <0.21< 0.21 were subsampled by removing 90%90\% of redundant points, yielding 6,179 instances while maintaining the demand distributions and model ability profiles.
  10. Knowl 10 — LLM Ability Profile Disentanglement Across Scale, Distillation, and Reasoning

    empirical result

    Profiling 15 LLMs across the 18 DeLeAn demand scales reveals clear capability disentanglement across model design paradigms:

    • Pre-training Parameter Scale: Model parameter scaling (e.g., LLaMA-3 from 1B to 405B) predominantly increases performance on Knowledge dimensions (KNa, KNc, KNf, KNn, KNs). However, ratio-scale ability curves reveal diminishing returns when scaling between the second-largest (90B) and largest (405B) models.
    • Chain-of-Thought (CoT) and Inference-Time Compute: Dedicated reasoning models (OpenAI o1, o1-mini, DeepSeek-R1-Distill-Qwen) exhibit sharp performance shifts specifically on Quantitative Reasoning (QLq), Logical Reasoning (QLl), Identifying Relevant Information (MCr), and Mind Modelling / Social Cognition (MS), maintaining substantial success even at demand levels ≥5\ge 5.
    • Distillation Effects: Distilling reasoning into smaller models (e.g., DeepSeek-R1-Distill-Qwen down to 7B) successfully preserves MCr, QL, and MS gains, but drops significantly on Knowledge dimensions and Spatio-physical Reasoning (SNs), showing that SNs requires both reasoning procedures and parameter scale.
  11. Knowl 11 — Reliability and Delphi Consensus of Automated Demand Annotations

    empirical result

    The DeLeAn rubrics were validated through inter-rater agreement analysis comparing 5 human experts and GPT-4o across 900 sampled instances (50 instances per demand dimension, stratified across levels):

    • Human-to-Human Agreement: Prior to consensus discussions, the within-group inter-rater reliability index (rWGr_{WG}) among three independent human annotators per instance averaged 0.830.83 across all 18 demand dimensions (ranging from 0.700.70 on KNs to 0.910.91 on AS, CEc, and KNn).
    • Delphi Consensus: A Delphi consensus process resolved disagreements of ≥2\ge 2 points, mitigating individual rater fatigue, vocabulary misinterpretations, and knowledge gaps.
    • Human-LLM Agreement: Agreement between the human Delphi consensus and GPT-4o annotations achieved an average rWGr_{WG} of 0.860.86 (ranging from 0.750.75 on KNa to 0.940.94 on CEe and KNn) and Spearman rank correlations between 0.750.75 and 0.940.94, establishing the reliability of fully automated LLM rubric annotation.
  12. Knowl 12 — Diagnostic Profiling of Standard AI Benchmarks for Specificity and Sensitivity

    empirical result

    Applying DeLeAn demand profiling across 20 widely used AI benchmarks reveals severe structural validity deficits in conventional benchmark design:

    1. Lack of Specificity: Benchmarks intended to evaluate a single construct impose heavy incidental demands on unrelated capabilities. For example, AGIEval Civil Service Examination (claimed logical reasoning) and LiveBench Reasoning require high competence in Verbal Comprehension (CEc), Knowledge, and Metacognition, confounding aggregate accuracy.
    2. Lack of Sensitivity: Specialized benchmarks often exhibit narrow demand distributions clustered exclusively at low difficulty (e.g., TempReason, TimeDial, and TimeQA for temporal reasoning), rendering them incapable of differentiating advanced models.
    3. Extraneous Confounders: Many benchmarks rely heavily on high Atypicality (AT, measuring memorization), high Volume (VO, inflating difficulty through long multi-step collages), or low Unguessability (UG, multiple-choice funnelling), which obscure true underlying cognitive capabilities.

Coverage note — None omitted. All key contributed constructs, mathematical formulations, calibration procedures, algorithms, benchmark datasets, empirical evaluation tables, and capability profiling findings have been completely covered.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023). Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Ackerman, T. A. (1989). Unidimensional irt calibration of compensatory and noncompensatory multidimensional items. Applied Psychological Measurement, 13(2):113–127.
  3. 3.Adams, S., Arel, I., Bach, J., Coop, R., Furlan, R., Goertzel, B., Hall, J. S., Samsonovich, A., Scheutz, M., Schlesinger, M., et al. (2012). Mapping the landscape of human-level artificial general intelligence. AI magazine, 33(1):25–42.
  4. 4.Ahuja, K., Dandapat, S., Sitaram, S., and Choudhury, M. (2022). Beyond static models and test sets: Benchmarking the potential of pre-trained models across tasks and languages. arXiv preprint arXiv:2205.06356.
  5. 5.Andrews, K., Spaulding, S., and Westra, E. (2020). Introduction to folk psychology: Pluralistic approaches. Synthese, 199(1-2):1685–1700.
  6. 6.Baker, F. B. (2001). The basics of item response theory. ERIC.
  7. 7.Balachandran, V., Chen, J., Joshi, N., Nushi, B., Palangi, H., Salinas, E., Vineet, V., Woffinden-Luey, J., and Yousefi, S. (2024). Eureka: Evaluating and understanding large foundation models. arXiv preprint arXiv:2409.10566.
  8. 8.Balepur, N., Rudinger, R., and Boyd-Graber, J. L. (2025). Which of these best describes multiple choice evaluation with llms? a) forced b) flawed c) fixable d) all of the above. arXiv preprint arXiv:2502.14127.
  9. 9.Balloccu, S., Schmidtová, P., Lango, M., and Dušek, O. (2024). Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms. arXiv preprint arXiv:2402.03927.
  10. 10.Barez, F., Fu, T., Prabhu, A., Casper, S., Sanyal, A., Bibi, A., O’Gara, A., Kirk, R., Bucknall, B., Fist, T., et al. (2025). Open problems in machine unlearning for ai safety. arXiv preprint arXiv:2501.04952.
  11. 11.Bergman, A. S., Hendricks, L. A., Rauh, M., Wu, B., Agnew, W., Kunesch, M., Duan, I., Gabriel, I., and Isaac, W. (2023). Representation in AI evaluations. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 519–533.
  12. 12.Biderman, S., Schoelkopf, H., Sutawika, L., Gao, L., Tow, J., Abbasi, B., Aji, A. F., Ammanamanchi, P. S., Black, S., Clive, J., DiPofi, A., Etxaniz, J., Fattori, B., Forde, J. Z., Foster, C., Jaiswal, M., Lee, W. Y., Li, H., Lovering, C., Muennighoff, N., Pavlick, E., Phang, J., Skowron, A., Tan, S., Tang, X., Wang, K. A., Winata, G. I., Yvon, F., and Zou, A. (2024). Lessons from the trenches on reproducible evaluation of language models. ArXiv, abs/2405.14782.
  13. 13.Bock, R. D. and Gibbons, R. D. (2021). Item response theory. John Wiley & Sons.
  14. 14.Bonifay, W. (2019). Multidimensional item response theory. Sage Publications.
  15. 15.Bowman, S. R. and Dahl, G. E. (2021). What will it take to fix benchmarking in natural language understanding? ArXiv, abs/2104.02145.
  16. 16.Breiman, L. (2001). Random forests. Machine learning, 45:5–32.
  17. 17.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  18. 18.Burden, J. (2024). Evaluating AI evaluation: Perils and prospects. arXiv preprint arXiv:2407.09221.
  19. 19.Burden, J., Tešic, M., Pacchiardi, L., and Hernández-Orallo, J. (2025). Paradigms of AI evaluation: Mapping goals, methodologies and culture.
  20. 20.Burden, J., Voudouris, K., Burnell, R., Rutar, D., Cheke, L., and Hernández-Orallo, J. (2023). Inferring capabilities from task performance with Bayesian triangulation. arXiv preprint arXiv:2309.11975.
  21. 21.Burnell, R., Hao, H., Conway, A. R., and Orallo, J. H. (2023a). Revealing the structure of language model capabilities. arXiv preprint arXiv:2306.10062.
  22. 22.Burnell, R., Schellaert, W., Burden, J., Ullman, T. D., Martinez-Plumed, F., Tenenbaum, J. B., Rutar, D., Cheke, L. G., Sohl-Dickstein, J., Mitchell, M., et al. (2023b). Rethink reporting of evaluation results in ai. Science, 380(6641):136–138.
  23. 23.Cao, B., Ren, M., Lin, H., Han, X., Zhang, F., Zhan, J., and Sun, L. (2024). StructEval: Deepen and broaden large language model assessment via structured evaluation. In Annual Meeting of the Association for Computational Linguistics.
  24. 24.Carlini, N. (2024). A GPT-4 capability forecasting challenge. https://nicholas.carlini.com/writing/llm-forecast/question/Capital-of-Paris. Accessed: 2024-09-08.
  25. 25.Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., et al. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3):1–45.
  26. 26.Chen, F. F. and Zhang, Z. (2018). Bifactor models in psychometric test development. The Wiley handbook of psychometric testing: A multidisciplinary reference on survey, scale and test development, pages 325–345.
  27. 27.Chu, Z., Chen, J., Chen, Q., Yu, W., Wang, H., Liu, M., and Qin, B. (2023). Timebench: A comprehensive evaluation of temporal reasoning abilities in large language models. arXiv preprint arXiv:2311.17667.
  28. 28.Coda-Forno, J., Binz, M., Wang, J. X., and Schulz, E. (2024). CogBench: a large language model walks into a psychology lab. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net.
  29. 29.Cohn, A. G. and Hernandez-Orallo, J. (2023). Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of llms. arXiv preprint arXiv:2304.11164.
  30. 30.Collins, K. M., Jiang, A. Q., Frieder, S., Wong, L., Zilka, M., Bhatt, U., Lukasiewicz, T., Wu, Y., Tenenbaum, J. B., Hart, W., et al. (2024). Evaluating language models for mathematics through interactions. Proceedings of the National Academy of Sciences, 121(24):e2318124121.
  31. 31.Crisp, V. and Novakovic, N. z. d. (2009). Is this year’s exam as demanding as last year’s? using a pilot method to evaluate the consistency of examination demands over time. Evaluation & Research in Education, 22(1):3–15.
  32. 32.De Ayala, R. J. (2009). Theory and practice of item response theory. Guilford Publications.
  33. 33.De Boeck, P. (2004). Explanatory item response models: A generalized linear and nonlinear approach. Springer Science & Business Media.
  34. 34.De Boeck, P. and Wilson, M. (2014). Multidimensional explanatory item response modeling. In Handbook of item response theory modeling, pages 252–271. Routledge.
  35. 35.DiBello, L. V., Henson, R. A., and Stout, W. F. (2015). A family of generalized diagnostic classification models for multiple choice option-based scoring. Applied psychological measurement, 39(1):62–79.
  36. 36.Dominguez-Olmedo, R., Dorner, F. E., and Hardt, M. (2024). Training on the test task confounds evaluation and emergence. arXiv preprint arXiv:2407.07890.
  37. 37.Drapal, P., Prudãncio, R. B., and Filho, T. M. S. (2024a). Towards explainable evaluation: Explaining predicted performance using local performance regions. Applied Soft Computing, 167:112351.
  38. 38.Drapal, P., Silva-Filho, T., and Prudãncio, R. B. C. (2024b). Meta-Learning and Novelty Detection for Machine Learning with Reject Option. In 2024 International Joint Conference on Neural Networks (IJCNN), pages 1–8.
  39. 39.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. (2024). The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  40. 40.Edwards, J. and Dall’Alba, G. (1981). Development of a scale of cognitive demand for analysis of printed secondary science materials. Research in Science Education, 11(1):158–170.
  41. 41.Embretson, S. and Reise, S. (2000). Item response theory for psychologists. Mahwah, NJ: Erlbaum.
  42. 42.Embretson, S. E. and Reise, S. P. (2013). Item response theory. Psychology Press.
  43. 43.Eriksson, M., Purificato, E., Noroozian, A., Vinagre, J., Chaslot, G., Gomez, E., and Fernandez-Llorca, D. (2025). Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation. arXiv preprint arXiv:2502.06559.
  44. 44.Estévez Almenzar, M., Fernández Llorca, D., Gómez, E., and Martinez Plumed, F. (2022). Glossary of human-centric artificial intelligence. Sevilla: Joint Research Centre (Seville Site).
  45. 45.European Union (2024). EU Artificial Intelligence Act. Regulation (EU) 2024/1689, Official Journal. Interinstitutional File: 2021/0106(COD).
  46. 46.Fang, Q., Oberski, D. L., and Nguyen, D. (2024). PATCH - psychometrics-assisted benchmarking of large language models: A case study of mathematics proficiency. ArXiv, abs/2404.01799.
  47. 47.Federiakin, D. (2025). Improving LLM leaderboards with psychometrical methodology. arXiv preprint arXiv:2501.17200.
  48. 48.Ferrando, P. J. (2014). A general approach for assessing person fit and person reliability in typical-response measurement. Applied Psychological Measurement, 38(2):166–183.
  49. 49.Fischer, G. H. (1973). The linear logistic test model as an instrument in educational research. Acta psychologica, 37(6):359–374.
  50. 50.Flesch, R. (1943). Marks of readable style; a study in adult education. Teachers College Contributions to Education.
  51. 51.Fountas, Z., Benfeghoul, M. A., Oomerjee, A., Christopoulou, F., Lampouras, G., Bou-Ammar, H., and Wang, J. (2024). Human-like episodic memory for infinite context llms. arXiv preprint arXiv:2407.09450.
  52. 52.François, T. and Miltsakaki, E. (2012). Do nlp and machine learning improve traditional readability formulas? In Proceedings of the First Workshop on Predicting and Improving Text Readability for target reader populations, pages 49–57.
  53. 53.Fränken, J.-P., Gandhi, K., Qiu, T., Khawaja, A., Goodman, N. D., and Gerstenberg, T. (2024). Procedural dilemma generation for evaluating moral reasoning in humans and language models. ArXiv, abs/2404.10975.
  54. 54.Freund, R. (2019). Rasch and rationality: Scale typologies as applied to item response theory. PhD thesis, UC Berkeley.
  55. 55.Gao, B., Song, F., Yang, Z., Cai, Z., Miao, Y., Dong, Q., Li, L., Ma, C., Chen, L., Xu, R., et al. (2024). Omni-math: A universal olympiad level mathematic benchmark for large language models. arXiv preprint arXiv:2410.07985.
  56. 56.Göllner, S. and Tropmann-Frick, M. (2023). Bridging the gap between theory and practice: Towards responsible AI evaluation. In CHAI@KI, pages 68–76.
  57. 57.Gorman, K. and Bedrick, S. (2019). We need to talk about standard splits. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2786–2791. Association for Computational Linguistics.
  58. 58.Graesser, A. C., McNamara, D. S., and Kulikowich, J. M. (2011). Coh-metrix: Providing multilevel analyses of text characteristics. Educational researcher, 40(5):223–234.
  59. 59.Guinet, G., Omidvar-Tehrani, B., Deoras, A., and Callot, L. (2024). Automated evaluation of retrieval-augmented language models with task-specific exam generation. ArXiv, abs/2405.13622.
  60. 60.Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. (2025). Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
  61. 61.Guo, T., Nan, B., Liang, Z., Guo, Z., Chawla, N., Wiest, O., Zhang, X., et al. (2023). What can large language models do in chemistry? a comprehensive benchmark on eight tasks. Advances in Neural Information Processing Systems, 36:59662–59688.
  62. 62.Hand, D. J. (2016). Measurement: A very short introduction. Oxford University Press.
  63. 63.Hardy, A., Reuel, A., Meimandi, K. J., Soder, L., Griffith, A., Asmar, D. M., Koyejo, S., Bernstein, M. S., and Kochenderfer, M. J. (2024). More than marketing? on the information value of ai benchmarks for practitioners. arXiv preprint arXiv:2412.05520.
  64. 64.He, X., Lin, Z., Gong, Y., Jin, A.-L., Zhang, H., Lin, C., Jiao, J., Yiu, S. M., Duan, N., and Chen, W. (2023). Annollm: Making large language models to be better crowdsourced annotators.
  65. 65.Hendrickx, K., Perini, L., Van der Plas, D., Meert, W., and Davis, J. (2024). Machine learning with a reject option: A survey. Machine Learning, 113(5):3073–3110.
  66. 66.Hernández-Orallo, J. (2017a). Evaluation in artificial intelligence: from task-oriented to ability-oriented measurement. Artificial Intelligence Review, 48:397–447.
  67. 67.Hernández-Orallo, J. (2017b). The measure of all minds: evaluating natural and artificial intelligence. Cambridge University Press.
  68. 68.Hernandez-Orallo, J. (2020). Ai evaluation: On broken yardsticks and measurement scales. In Workshop on evaluating evaluation of AI systems at AAAI.
  69. 69.Hernandez-Orallo, J. (2024). Caveats and solutions for characterising general-purpose AI. In ECAI 2024, pages 2–9. IOS Press.
  70. 70.Hernández-Orallo, J., Loe, B. S., Cheke, L., Martínez-Plumed, F., and Ó hÉigeartaigh, S. (2021). General intelligence disentangled via a generality metric for natural and artificial intelligence. Scientific reports, 11(1):22822.
  71. 71.Hernández-Orallo, J., Schellaert, W., and Martínez-Plumed, F. (2022). Training on the test set: Mapping the system-problem space in ai. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 12256–12261.
  72. 72.Hernández-Orallo, J. and Vold, K. (2019). AI extenders: the ethical and societal implications of humans cognitively extended by AI. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, pages 507–513.
  73. 73.Hu, J. and Frank, M. (2024). Auxiliary task demands mask the capabilities of smaller language models. In First Conference on Language Modeling.
  74. 74.Hughes, S., Pollitt, A., and Ahmed, A. (1998). The development of a tool for gauging the demands of gcse and a level exam questions. BERA, Queen’s University Belfast.
  75. 75.Hüllermeier, E. and Waegeman, W. (2021). Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Machine learning, 110(3):457–506.
  76. 76.Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al. (2024). Gpt-4o system card. arXiv preprint arXiv:2410.21276.
  77. 77.Ilic, D. and Gignac, G. E. (2024). Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement? Intelligence, 106:101858.
  78. 78.Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al. (2024). Openai o1 system card. arXiv preprint arXiv:2412.16720.
  79. 79.James, L. R., Demaree, R. G., and Wolf, G. (1984). Estimating within-group interrater reliability with and without response bias. Journal of applied psychology, 69(1):85.
  80. 80.Jiang, M., Liu, K., Zhong, M., Schaeffer, R., Ouyang, S., Han, J., and Koyejo, S. (2024a). Does data contamination make a difference? insights from intentionally contaminating pre-training data for language models. In ICLR 2024 Workshop on Navigating and Addressing Data Problems for Foundation Models.
  81. 81.Jiang, M., Liu, K. Z., Zhong, M., Schaeffer, R., Ouyang, S., Han, J., and Koyejo, S. (2024b). Investigating data contamination for pre-training language models. arXiv preprint arXiv:2401.06059.
  82. 82.Jin, Z., Chen, Y., Leeb, F., Gresele, L., Kamal, O., Lyu, Z., Blin, K., Gonzalez, F., Kleiman-Weiner, M., Sachan, M., and Schölkopf, B. (2023). CLadder: Assessing causal reasoning in language models. In NeurIPS.
  83. 83.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., et al. (2022). Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  84. 84.Kadian, A., Truong, J., Gokaslan, A., Clegg, A., Wijmans, E., Lee, S., Savva, M., Chernova, S., and Batra, D. (2020). Sim2real predictivity: Does evaluation in simulation predict real-world performance? IEEE Robotics and Automation Letters, 5(4):6670–6677.
  85. 85.Kalyuga, S. (2011). Cognitive load theory: How many types of load does it really need? Educational psychology review, 23:1–19.
  86. 86.Kazemi, M., Fatemi, B., Bansal, H., Palowitch, J., Anastasiou, C., Mehta, S. V., Jain, L. K., Aglietti, V., Jindal, D., Chen, P., et al. (2025). Big-bench extra hard. arXiv preprint arXiv:2502.19187.
  87. 87.Khandekar, N., Jin, Q., Xiong, G., Dunn, S., Applebaum, S. S., Anwar, Z., Sarfo-Gyamfi, M., Safranek, C. W., Anwar, A. A., Zhang, A., et al. (2024). Medcalc-bench: Evaluating large language models for medical calculations. arXiv preprint arXiv:2406.12036.
  88. 88.Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., et al. (2021). Dynabench: Rethinking benchmarking in nlp. arXiv preprint arXiv:2104.14337.
  89. 89.Kipnis, A., Voudouris, K., Buschoff, L. M. S., and Schulz, E. (2024). metabench–a sparse benchmark to measure general ability in large language models. arXiv preprint arXiv:2407.12844.
  90. 90.Krathwohl, D. R. (2002). A revision of Bloom’s taxonomy: An overview. Theory Into Practice, 41(4):212–218.
  91. 91.Lalor, J. P., Rodriguez, P., Sedoc, J., and Hernandez-Orallo, J. (2024). Item response theory for natural language processing. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Tutorial Abstracts, pages 9–13.
  92. 92.Lalor, J. P., Wu, H., and Yu, H. (2016). Building an evaluation scale using item response theory. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing, volume 2016, page 648. NIH Public Access.
  93. 93.LeBreton, J. M. and Senter, J. L. (2008). Answers to 20 questions about interrater reliability and interrater agreement. Organizational research methods, 11(4):815–852.
  94. 94.Lee, M., Srivastava, M., Hardy, A., Thickstun, J., Durmus, E., Paranjape, A., Gerard-Ursin, I., Li, X. L., Ladhak, F., Rong, F., et al. (2024). Evaluating human-language model interaction. Transactions on Machine Learning Research.
  95. 95.Lei, Z., Liang, T., Hu, H., Zhang, J., Zhou, Y., Shao, Y., Li, L., Li, C., Wang, C., Yan, H., and Guo, Q. (2024). GAOKAO-eval: Does high scores truly reflect strong capabilities in LLMs? ArXiv, abs/2412.10056.
  96. 96.Levy, M., Jacoby, A., and Goldberg, Y. (2024). Same task, more tokens: the impact of input length on the reasoning performance of large language models. In Annual Meeting of the Association for Computational Linguistics.
  97. 97.Liao, T., Taori, R., Raji, D., and Schmidt, L. (2021). Are we learning yet? a meta review of evaluation failures across machine learning. In Vanschoren, J. and Yeung, S., editors, Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1.
  98. 98.Linstone, H. A., Turoff, M., et al. (1975). The delphi method, volume 1975. Addison-Wesley Reading, MA.
  99. 99.Liu, Q., Gong, Z., Huang, Z., Liu, C., Zhu, H., Li, Z., Chen, E., and Xiong, H. (2023). Multi-dimensional ability diagnosis for machine learning algorithms. ArXiv, abs/2307.07134.
  100. 100.Liu, Y. L., Blodgett, S. L., Cheung, J., Liao, Q. V., Olteanu, A., and Xiao, Z. (2024). ECBD: Evidence-centered benchmark design for NLP. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16349–16365. Association for Computational Linguistics.
  101. 101.Lord, F. M. (1975). The ‘ability’scale in item characteristic curve theory. Psychometrika, 40(2):205–217.
  102. 102.Lumsden, J. (1977). Person reliability. Applied Psychological Measurement, 1(4):477–482.
  103. 103.Lyu, Y. and Du, Y. (2025). The ethical evaluation of large language models and its optimization. AI and Ethics, pages 1–14.
  104. 104.Martínez-Plumed, F., Prudãncio, R. B., Martínez-Usó, A., and Hernández-Orallo, J. (2019). Item response theory in ai: Analysing machine learning classifiers at the instance level. Artificial intelligence, 271:18–42.
  105. 105.Masry, A., Rodriguez, J. A., Zhang, T., Wang, S., Wang, C., Feizi, A., Suresh, A. K., Puri, A., Jian, X., Noãl, P.-A., et al. (2025). Alignvlm: Bridging vision and language latent spaces for multimodal understanding. arXiv preprint arXiv:2502.01341.
  106. 106.McGrew, K. S. (2005). The cattell-horn-carroll theory of cognitive abilities: Past, present, and future. In Flanagan, D. P. and Harrison, P. L., editors, Contemporary intellectual assessment: Theories, tests, and issues, pages 136–181. The Guilford Press, 2 edition.
  107. 107.McIntosh, T. R., Susnjak, T., Liu, T., Watters, P., and Halgamuge, M. N. (2024). Inadequacies of large language model benchmarks in the era of generative artificial intelligence. ArXiv, abs/2402.09880.
  108. 108.Michell, J. (1999). Measurement in psychology: A critical history of a methodological concept, volume 53. Cambridge University Press.
  109. 109.Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., and Farajtabar, M. (2024). Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models. arXiv preprint arXiv:2410.05229.
  110. 110.Momennejad, I., Hasanbeig, H., Frujeri, F. V., Sharma, H., Jojic, N., Palangi, H., Ness, R., and Larson, J. (2023). Evaluating cognitive maps and planning in large language models with CogEval. In Thirty-seventh Conference on Neural Information Processing Systems.
  111. 111.Mondorf, P. and Plank, B. (2024). Liar, liar, logical mire: A benchmark for suppositional reasoning in large language models. arXiv preprint arXiv:2406.12546.
  112. 112.Moros-Daval, Y., Martínez-Plumed, F., and Hernández-Orallo, J. (2024). Language task difficulty prediction through LLM-annotated meta-features. In ECAI 2024, pages 2434–2441. IOS Press.
  113. 113.Morris, M. R., Sohl-Dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., and Legg, S. (2023). Levels of agi for operationalizing progress on the path to agi. arXiv preprint arXiv:2311.02462.
  114. 114.OECD (2024). Education at a glance 2024: Oecd indicators. OECD Publishing.
  115. 115.Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., and Stoica, I. (2024). Routellm: Learning to route llms from preference data. In The Thirteenth International Conference on Learning Representations.
  116. 116.Pacchiardi, L., Cheke, L. G., and Hernández-Orallo, J. (2024). 100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances.
  117. 117.Pacchiardi, L., Voudouris, K., Slater, B., Martínez-Plumed, F., Hernández-Orallo, J., Zhou, L., and Schellaert, W. (2025). PredictaBoard: Benchmarking LLM score predictability.
  118. 118.Pangakis, N., Wolken, S., and Fasching, N. (2023). Automated annotation with generative ai requires validation. arXiv preprint arXiv:2306.00176.
  119. 119.Partridge, G. E. (1910). An outline of individual study. Sturgis and Walton Company.
  120. 120.Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al. (2011). Scikit-learn: Machine learning in python. the Journal of machine Learning research, 12:2825–2830.
  121. 121.Pennington, J., Socher, R., and Manning, C. D. (2014). Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532–1543.
  122. 122.Phan, L., Gatti, A., Han, Z., Li, N., Hu, J., Zhang, H., Shi, S., Choi, M., Agrawal, A., Chopra, A., et al. (2025). Humanity’s last exam. arXiv preprint arXiv:2501.14249.
  123. 123.Phuong, M., Aitchison, M., Catt, E., Cogan, S., Kaskasoli, A., Krakovna, V., Lindner, D., Rahtz, M., Assael, Y., Hodkinson, S., et al. (2024). Evaluating frontier models for dangerous capabilities. arXiv preprint arXiv:2403.13793.
  124. 124.Pollitt, A., Ahmed, A., and Crisp, V. (2007). The demands of examination syllabuses and question papers. Techniques for monitoring the comparability of examination standards, pages 166–206.
  125. 125.Polo, F. M., Weber, L., Choshen, L., Sun, Y., Xu, G., and Yurochkin, M. (2024). tinybenchmarks: evaluating LLMs with fewer examples. In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models.
  126. 126.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2018). Improving language understanding by generative pre-training. Technical report, OpenAI.
  127. 127.Rahwan, I., Cebrian, M., Obradovich, N., Bongard, J., Bonnefon, J.-F., Breazeal, C., Crandall, J. W., Christakis, N. A., Couzin, I. D., Jackson, M. O., et al. (2019). Machine behaviour. Nature, 568(7753):477–486.
  128. 128.Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., et al. (2024). Gaps in the safety evaluation of generative ai. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 1200–1217.
  129. 129.Reckase, M. D. (2006). Chapter 18 - multidimensional item response theory. Handbook of statistics, 26:607–642.
  130. 130.Ren, R., Basart, S., Khoja, A., Gatti, A., Phan, L., Yin, X., Mazeika, M., Pan, A., Mukobi, G., Kim, R. H., Fitz, S., and Hendrycks, D. (2024). Safetywashing: Do AI safety benchmarks actually measure safety progress? In Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J. M., and Zhang, C., editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024.
  131. 131.Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., and Kochenderfer, M. (2024). BetterBench: Assessing AI benchmarks, uncovering issues, and establishing best practices. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  132. 132.Rhemtulla, M., Brosseau-Liard, P. É., and Savalei, V. (2012). When can categorical variables be treated as continuous? a comparison of robust continuous and categorical sem estimation methods under suboptimal conditions. Psychological methods, 17(3):354.
  133. 133.Roberts, M., Thakur, H., Herlihy, C., White, C., and Dooley, S. (2023). Data contamination through the lens of time. arXiv preprint arXiv:2310.10628.
  134. 134.Ruan, Y., Maddison, C. J., and Hashimoto, T. (2024). Observational scaling laws and the predictability of language model performance. arXiv preprint arXiv:2405.10938.
  135. 135.Rust, J., Kosinski, M., and Stillwell, D. (2021). Modern psychometrics: The science of psychological assessment. 4th Edition, Routledge.
  136. 136.Saparov, A. and He, H. (2023). Language models are greedy reasoners: A systematic formal analysis of chain-of-thought. In The Eleventh International Conference on Learning Representations.
  137. 137.Savelka, J. and Ashley, K. D. (2023). The unreasonable effectiveness of large language models in zero-shot semantic annotation of legal texts. Frontiers in Artificial Intelligence, 6.
  138. 138.Schellaert, W. (2025). The Evaluation of Artificial Intelligence as a Prediction Problem. PhD thesis, Universitat Politecnica de Valencia.
  139. 139.Schellaert, W., Martínez-Plumed, F., and Hernández-Orallo, J. (2025). Analysing the predictability of language model performance. ACM Transactions on Intelligent Systems and Technology, 16(2):1–26.
  140. 140.Schlangen, D. (2019). Language tasks and language games: On methodology in current natural language processing research. arXiv preprint arXiv:1908.10747.
  141. 141.Schulze Buschoff, L. M., Akata, E., Bethge, M., and Schulz, E. (2025). Visual cognition in multimodal large language models. Nature Machine Intelligence, pages 1–11.
  142. 142.Shorinwa, O., Mei, Z., Lidard, J., Ren, A. Z., and Majumdar, A. (2024). A survey on uncertainty quantification of large language models: Taxonomy, open research challenges, and future directions. arXiv preprint arXiv:2412.05563.
  143. 143.Siska, C., Marazopoulou, K., Ailem, M., and Bono, J. (2024). Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10406–10421. Association for Computational Linguistics.
  144. 144.Srinivasan, A., Sitaram, S., Ganu, T., Dandapat, S., Bali, K., and Choudhury, M. (2021). Predicting the performance of multilingual nlp models. arXiv preprint arXiv:2110.08875.
  145. 145.Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. (2023). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Journal of Machine Learning Research.
  146. 146.Srivastava, S., PV, A., Menon, S., Sukumar, A., Philipose, A., Prince, S., Thomas, S., et al. (2024). Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap. arXiv preprint arXiv:2402.19450.
  147. 147.Stahl, B. C., Antoniou, J., Bhalla, N., Brooks, L., Jansen, P., Lindqvist, B., Kirichenko, A., Marchal, S., Rodrigues, R., Santiago, N., et al. (2023). A systematic review of artificial intelligence impact assessments. Artificial Intelligence Review, 56(11):12799–12831.
  148. 148.Stevens, S. S. (1946). On the theory of scales of measurement. Science, 103(2684):677–680.
  149. 149.Steyvers, M., Tejeda, H., Kumar, A., Belem, C., Karny, S., Hu, X., Mayer, L. W., and Smyth, P. (2025). What large language models know and what people think they know. Nature Machine Intelligence, pages 1–11.
  150. 150.Stucky, B. D. and Edelen, M. O. (2014). Using hierarchical irt models to create unidimensional measures from multidimensional data. Handbook of item response theory modeling, pages 183–206.
  151. 151.Subramonian, A., Yuan, X., Daumé III, H., and Blodgett, S. L. (2023). It takes two to tango: Navigating conceptualizations of NLP tasks and measurements of performance. In Rogers, A., Boyd-Graber, J., and Okazaki, N., editors, Findings of the Association for Computational Linguistics: ACL 2023, pages 3234–3279. Association for Computational Linguistics.
  152. 152.Suhara, Y., Li, J., Li, Y., Zhang, D., Demiralp, Ç., Chen, C., and Tan, W.-C. (2022). Annotating columns with pre-trained language models. In Proceedings of the 2022 International Conference on Management of Data, pages 1493–1503.
  153. 153.Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al. (2022). Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261.
  154. 154.Sweller, J. (2011). Cognitive load theory. In Psychology of learning and motivation, volume 55, pages 37–76. Elsevier.
  155. 155.Tang, F., Gao, W., Peng, L., and Zhan, J. (2023). Agibench: A multi-granularity, multimodal, human-referenced, auto-scoring benchmark for large language models. ArXiv, abs/2309.06495.
  156. 156.Thissen, D. and Wainer, H. (2002). Test scoring.
  157. 157.Thurstone, L. L. (1937). Ability, motivation, and speed. Psychometrika, 2(4):249–254.
  158. 158.Thurstone, L. L. (1938). Primary mental abilities: Psychometric monographs no. 1. In The measurement of intelligence, pages 131–136. Springer.
  159. 159.Tolan, S., Pesole, A., Martínez-Plumed, F., Fernández-Macías, E., Hernández-Orallo, J., and Gómez, E. (2021). Measuring the occupational impact of AI: tasks, cognitive abilities and AI benchmarks. Journal of Artificial Intelligence Research, 71:191–236.
  160. 160.Trabin, T. E. and Weiss, D. J. (1983). The person response curve: Fit of individuals to item response theory models. In New horizons in testing, pages 83–108. Elsevier.
  161. 161.Vafa, K., Rambachan, A., and Mullainathan, S. (2024). Do large language models perform the way people expect? measuring the human generalization function. In Forty-first International Conference on Machine Learning.
  162. 162.Vale, C. D. and Weiss, D. J. (1975). A study of computer-administered stradaptive ability testing. Technical report, Minnesota Univ. Minneapolis Dept. of Psychology.
  163. 163.Van der Linden, W. J. and van der Linden, W. (2016). Handbook of item response theory, volume 1. CRC press New York.
  164. 164.van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S. F., and Ward, F. R. (2024). Ai sandbagging: Language models can strategically underperform on evaluations. arXiv preprint arXiv:2406.07358.
  165. 165.Vania, C., Htut, P. M., Huang, W., Mungra, D., Pang, R. Y., Phang, J., Liu, H., Cho, K., and Bowman, S. R. (2021). Comparing test sets with item response theory. In Zong, C., Xia, F., Li, W., and Navigli, R., editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1141–1158. Association for Computational Linguistics.
  166. 166.Von Davier, M. (2008). A general diagnostic model applied to language testing data. British Journal of Mathematical and Statistical Psychology, 61(2):287–307.
  167. 167.von Davier, M. and Yamamoto, K. (2004). A class of models for cognitive diagnosis. In 4th spearman conference, Philadelphia, PA.
  168. 168.Wainer, H. and Braun, H. I. (2013). Test validity. Routledge.
  169. 169.Wallmark, J., Josefsson, M., and Wiberg, M. (2024). Introducing flexible monotone multiple choice item response theory models and bit scales. arXiv preprint arXiv:2410.01480.
  170. 170.Wang, C. J., Lee, D., Menghini, C., Mols, J., Doughty, J., Khoja, A., Lynch, J., Hendryx, S., Yue, S., and Hendrycks, D. (2025a). Enigmaeval: A benchmark of long multimodal reasoning challenges. arXiv preprint arXiv:2502.08859.
  171. 171.Wang, F., Gao, W., Liu, Q., Li, J., Zhao, G., Zhang, Z., Huang, Z., Zhu, M., Wang, S., Tong, W., et al. (2024a). A survey of models for cognitive diagnosis: New developments and future directions. arXiv preprint arXiv:2407.05458.
  172. 172.Wang, H., Shi, H., Tan, S., Qin, W., Wang, W., Zhang, T., Nambi, A. U., Ganu, T., and Wang, H. (2024b). Multimodal needle in a haystack: Benchmarking long-context capability of multimodal large language models. ArXiv, abs/2406.11230.
  173. 173.Wang, H., Zhao, S., Qiang, Z., Xi, N., Qin, B., and Liu, T. (2025b). LLMs may perform MCQA by selecting the least incorrect option. In Proceedings of the 31st International Conference on Computational Linguistics, pages 5852–5862.
  174. 174.Wang, X., Hu, Z., Lu, P., Zhu, Y., Zhang, J., Subramaniam, S., Loomba, A. R., Zhang, S., Sun, Y., and Wang, W. (2023a). Scibench: Evaluating college-level scientific problem-solving abilities of large language models. arXiv preprint arXiv:2307.10635.
  175. 175.Wang, X., Jiang, L., Hernandez-Orallo, J., Stillwell, D., Sun, L., Luo, F., and Xie, X. (2023b). Evaluating general-purpose AI with psychometrics. arXiv preprint arXiv:2310.16379.
  176. 176.Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., et al. (2024c). Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. arXiv preprint arXiv:2406.01574.
  177. 177.Wang, Y. and Zhao, Y. (2024). RUPBench: Benchmarking reasoning under perturbations for robustness evaluation in large language models. ArXiv, abs/2406.11020.
  178. 178.Wasserman, E. A. and Zentall, T. R. (2006). Comparative cognition: Experimental explorations of animal intelligence. Oxford University Press.
  179. 179.White, C., Dooley, S., Roberts, M., Pal, A., Feuer, B., Jain, S., Shwartz-Ziv, R., Jain, N., Saifullah, K., Naidu, S., et al. (2024). Livebench: A challenging, contamination-free llm benchmark. arXiv preprint arXiv:2406.19314.
  180. 180.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2020). Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45.
  181. 181.Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al. (2024). Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115.
  182. 182.Yao, J., Yi, X., Duan, S., Wang, J., Bai, Y., Huang, M., Zhang, P., Lu, T., Dou, Z., Sun, M., et al. (2025). Value compass leaderboard: A platform for fundamental and validated evaluation of llms values. arXiv preprint arXiv:2501.07071.
  183. 183.Ye, Z., Liu, P., Fu, J., and Neubig, G. (2021). Towards more fine-grained and reliable nlp performance prediction. arXiv preprint arXiv:2102.05486.
  184. 184.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019). Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830.
  185. 185.Zeng, Y. (2024). Quantifying risk propensities of large language models: Ethical focus and bias detection through role-play. arXiv preprint arXiv:2411.08884.
  186. 186.Zhang, J., Huang, W., Ma, Z., Michel, O., He, D., Gupta, T., Ma, W.-C., Farhadi, A., Kembhavi, A., and Krishna, R. (2024). Task me anything. arXiv preprint arXiv:2406.11775.
  187. 187.Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N. (2023). Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364.
  188. 188.Zhou, L., Farag, Y., and Vlachos, A. (2024a). An llm feature-based framework for dialogue constructiveness assessment. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 5389–5409.
  189. 189.Zhou, L., Martínez-Plumed, F., Hernández-Orallo, J., Ferri, C., and Schellaert, W. (2022). Reject before you run: Small assessors anticipate big language models. Proceedings of the Workshop on AI Evaluation Beyond Metrics co-located with the 31st International Joint Conference on Artificial Intelligence (IJCAI-ECAI 2022).
  190. 190.Zhou, L., Moreno-Casares, P. A., Martínez-Plumed, F., Burden, J., Burnell, R., Cheke, L., Ferri, C., Marcoci, A., Mehrbakhsh, B., Moros-Daval, Y., et al. (2023). Predictable artificial intelligence. arXiv preprint arXiv:2310.06167.
  191. 191.Zhou, L., Schellaert, W., Martínez-Plumed, F., Moros-Daval, Y., Ferri, C., and Hernández-Orallo, J. (2024b). Larger and more instructable language models become less reliable. Nature, 634(8032):61–68.
  192. 192.Zhu, K., Wang, J., Zhao, Q., Xu, R., and Xie, X. (2024). Dynamic evaluation of large language models by meta probing agents. In Forty-first International Conference on Machine Learning.
  193. 193.Zhuang, Y., Liu, Q., Ning, Y., Huang, W., Pardos, Z. A., Kyllonen, P. C., Zu, J., Mao, Q., Lv, R., Huang, Z., Zhao, G., Zhang, Z., Wang, S., and Chen, E. (2024). From static benchmarks to adaptive testing: Psychometrics in AI evaluation.

Citation

MLA
Zhou, L., et al. “General Scales Unlock AI Evaluation with Explanatory and Predictive Power”. arXiv, 2025, https://doi.org/10.48550/arxiv.2503.06378.
APA
Zhou, L., Pacchiardi, L., Martínez-Plumed, F., Collins, K. M., Moros-Daval, Y., Zhang, S., Zhao, Q., Huang, Y., Sun, L., Prunty, J. E., Li, Z., Sánchez-García, P., Chen, K. J., Casares, P. A. M., Zu, J., Burden, J., Mehrbakhsh, B., Stillwell, D., Cebrian, M., … Hernández-Orallo, J. (2025). General Scales Unlock AI Evaluation with Explanatory and Predictive Power. arXiv. https://doi.org/10.48550/arxiv.2503.06378
Chicago
Zhou, L., L. Pacchiardi, F. Martínez-Plumed, et al. 2025. “General Scales Unlock AI Evaluation with Explanatory and Predictive Power”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2503.06378.
Harvard
Zhou, L. et al. (2025) “General Scales Unlock AI Evaluation with Explanatory and Predictive Power”. arXiv. Available at: https://doi.org/10.48550/arxiv.2503.06378.
Vancouver
1. Zhou L, Pacchiardi L, Martínez-Plumed F, et al (2025) General Scales Unlock AI Evaluation with Explanatory and Predictive Power. https://doi.org/10.48550/arxiv.2503.06378

BibTeX

@misc{https://doi.org/10.48550/arxiv.2503.06378,
  doi = {10.48550/ARXIV.2503.06378},
  url = {https://arxiv.org/abs/2503.06378},
  author = {Zhou, Lexin and Pacchiardi, Lorenzo and Martínez-Plumed, Fernando and Collins, Katherine M. and Moros-Daval, Yael and Zhang, Seraphina and Zhao, Qinlin and Huang, Yitian and Sun, Luning and Prunty, Jonathan E. and Li, Zongqian and Sánchez-García, Pablo and Chen, Kexin Jiang and Casares, Pablo A. M. and Zu, Jiyun and Burden, John and Mehrbakhsh, Behzad and Stillwell, David and Cebrian, Manuel and Wang, Jindong and Henderson, Peter and Wu, Sherry Tongshuang and Kyllonen, Patrick C. and Cheke, Lucy and Xie, Xing and Hernández-Orallo, José},
  keywords = {Artificial Intelligence (cs.AI), Computation and Language (cs.CL), Computers and Society (cs.CY), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {General Scales Unlock AI Evaluation with Explanatory and Predictive Power},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/