tinyBenchmarks: evaluating LLMs with fewer examples
Felipe Maia PoloLucas WeberLeshem ChoshenYuekai SunGongjun XuMikhail Yurochkin
Proposes Item Response Theory and clustering techniques to drastically cut large language model evaluation costs by estimating full benchmark performance on datasets like MMLU and HELM using only 100 representative examples per scenario within an average 2% error margin.
Assessing large language models (LLMs) requires testing them against standard multi-task benchmarks containing tens of thousands of questions. This process has become financially, computationally, and environmentally prohibitive. Evaluating a single model on comprehensive benchmark suites can demand thousands of graphics processing unit hours or tens of thousands of dollars in commercial interface fees. Because engineering workflows require frequent evaluations across training checkpoints and prompt variations, organizations face significant resource bottlenecks.
The article evaluates whether statistical psychometrics, specifically Item Response Theory (IRT), can drastically reduce the number of evaluation items needed to assess LLMs without sacrificing accuracy. The authors demonstrate a framework to select representative test subsets and accurately predict full-benchmark performance using only a small fraction of the original data.
The authors analyzed historical evaluation records across four leading benchmarks: HuggingFace's Open LLM Leaderboard, MMLU, HELM Lite, and AlpacaEval 2.0, spanning over 400 models. They applied multidimensional IRT models to represent question difficulty and model ability, comparing IRT-guided example selection against stratified random sampling and direct correctness clustering. To validate practicality, the methods were tested under challenging conditions, including temporal shifts where models were trained on older systems and tested on newer releases, as well as evaluations on specialized domain models.
The primary finding is that evaluating an LLM on just 100 curated examples per benchmark scenario accurately estimates full benchmark performance within an average error margin of approximately 2%. On the 14,000-question MMLU benchmark, this represents a 140-fold reduction in evaluation items. The best-performing approach, an IRT-adjusted estimator termed IRT++, proved robust against temporal distribution shifts and domain specialization, maintaining a worst-case error under 4% across almost all tested models. Furthermore, the statistical adjustment runs on standard hardware in a matter of seconds, providing immediate computational efficiency.
These results demonstrate that organizations can reduce LLM evaluation budgets, energy consumption, and turnaround cycles by over 90% without compromising assessment quality. Rapid, low-cost benchmarking enables practitioners to run iterative tests during model pre-training, fine-tuning, and prompt optimization that were previously cost-prohibitive. Consequently, teams can accelerate deployment timelines while mitigating development risks.
Organizations should adopt curated sub-benchmarks and statistical estimation tools for routine tracking, model selection, and prompt engineering, reserving full benchmark evaluations only for final regulatory or release milestones. Looking forward, evaluation pipelines can incorporate adaptive testing to select questions dynamically, though current implementations require software optimization to reduce execution latency.
Confidence in these findings is high for standard model architectures within tested difficulty ranges. However, caution is warranted if a model exhibits anomalous capability profiles, such as solving difficult questions while failing simple ones, or during major architectural shifts. Benchmark curators should periodically update IRT parameter estimates using recent models to maintain predictive stability as LLM capabilities expand.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). tinyBenchmarks directly studies reducing evaluation effort for MMLU, so reading the original MMLU benchmark first clarifies the task, scale, and performance estimates being compressed.
- Paper: From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder Pipeline, Tianle Li et al. (2025). BenchBuilder continues tinyBenchmarks’ efficiency agenda by turning crowdsourced data into smaller, curated benchmarks designed to preserve reliable model rankings at lower evaluation cost.
