tinyBenchmarks: evaluating LLMs with fewer examples

Felipe Maia PoloLucas WeberLeshem ChoshenYuekai SunGongjun XuMikhail Yurochkin

article2024ICML321 citations

Proposes Item Response Theory and clustering techniques to drastically cut large language model evaluation costs by estimating full benchmark performance on datasets like MMLU and HELM using only 100 representative examples per scenario within an average 2% error margin.

Listen

Assessing large language models (LLMs) requires testing them against standard multi-task benchmarks containing tens of thousands of questions. This process has become financially, computationally, and environmentally prohibitive. Evaluating a single model on comprehensive benchmark suites can demand thousands of graphics processing unit hours or tens of thousands of dollars in commercial interface fees. Because engineering workflows require frequent evaluations across training checkpoints and prompt variations, organizations face significant resource bottlenecks.

The article evaluates whether statistical psychometrics, specifically Item Response Theory (IRT), can drastically reduce the number of evaluation items needed to assess LLMs without sacrificing accuracy. The authors demonstrate a framework to select representative test subsets and accurately predict full-benchmark performance using only a small fraction of the original data.

The authors analyzed historical evaluation records across four leading benchmarks: HuggingFace's Open LLM Leaderboard, MMLU, HELM Lite, and AlpacaEval 2.0, spanning over 400 models. They applied multidimensional IRT models to represent question difficulty and model ability, comparing IRT-guided example selection against stratified random sampling and direct correctness clustering. To validate practicality, the methods were tested under challenging conditions, including temporal shifts where models were trained on older systems and tested on newer releases, as well as evaluations on specialized domain models.

The primary finding is that evaluating an LLM on just 100 curated examples per benchmark scenario accurately estimates full benchmark performance within an average error margin of approximately 2%. On the 14,000-question MMLU benchmark, this represents a 140-fold reduction in evaluation items. The best-performing approach, an IRT-adjusted estimator termed IRT++, proved robust against temporal distribution shifts and domain specialization, maintaining a worst-case error under 4% across almost all tested models. Furthermore, the statistical adjustment runs on standard hardware in a matter of seconds, providing immediate computational efficiency.

These results demonstrate that organizations can reduce LLM evaluation budgets, energy consumption, and turnaround cycles by over 90% without compromising assessment quality. Rapid, low-cost benchmarking enables practitioners to run iterative tests during model pre-training, fine-tuning, and prompt optimization that were previously cost-prohibitive. Consequently, teams can accelerate deployment timelines while mitigating development risks.

Organizations should adopt curated sub-benchmarks and statistical estimation tools for routine tracking, model selection, and prompt engineering, reserving full benchmark evaluations only for final regulatory or release milestones. Looking forward, evaluation pipelines can incorporate adaptive testing to select questions dynamically, though current implementations require software optimization to reduce execution latency.

Confidence in these findings is high for standard model architectures within tested difficulty ranges. However, caution is warranted if a model exhibits anomalous capability profiles, such as solving difficult questions while failing simple ones, or during major architectural shifts. Benchmark curators should periodically update IRT parameter estimates using recent models to maintain predictive stability as LLM capabilities expand.

  • Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). tinyBenchmarks directly studies reducing evaluation effort for MMLU, so reading the original MMLU benchmark first clarifies the task, scale, and performance estimates being compressed.
Cover for tinyBenchmarks: evaluating LLMs with fewer examples

Abstract

The versatility of large language models (LLMs) led to the creation of diverse benchmarks that thoroughly test a variety of language models’ abilities. These benchmarks consist of tens of thousands of examples making evaluation of LLMs very expensive. In this paper, we investigate strategies to reduce the number of evaluations needed to assess the performance of an LLM on several key benchmarks. For example, we show that to accurately estimate the performance of an LLM on MMLU, a popular multiple-choice QA benchmark consisting of 14K examples, it is sufficient to evaluate this LLM on 100 curated examples. We release evaluation tools and tiny versions of popular benchmarks: Open LLM Leaderboard, MMLU, HELM, and AlpacaEval 2.0. Our empirical analysis demonstrates that these tools and tiny benchmarks are sufficient to reliably and efficiently reproduce the original evaluation results¹.

Table of Contents

  • 1. Introduction
  • 1.1. Related work
  • 2. Problem statement
  • 3. Selecting evaluation examples
  • 3.1. Stratified random sampling
  • 3.2. Clustering
  • 4. Better performance estimation with IRT
  • 4.1. The IRT model
  • 4.2. IRT-based LLM performance estimation
  • 4.3. Using IRT when Y il is not binary
  • 4.4. Fitting the IRT model
  • 5. Assessing evaluation strategies
  • 6. Conclusion
  • 6.1. Extensions
  • 6.2. Limitations
  • Acknowledgements
  • Impact Statement
  • References
  • A. Evaluation when subscenarios have different number of samples
  • B. tinyMMLU
  • C. Proof of Proposition 4.1
  • D. More details about benchmarks
  • E. Extra results
  • E.1. Robustness in predicting performance in a longer time horizon
  • E.2. How costly is it for stratified random sampling beat IRT++ with larger samples?
  • E.3. Running time
  • E.4. Rank correlation results
  • E.5. Adaptive testing
  • F. Individual performances per scenario
  • F.1. Open LLM Leaderboard
  • F.2. HELM

Knowls

  1. Knowl 1 — Multidimensional Item Response Theory Formulation for LLM Evaluation

    model/method

    To evaluate large language models (LLMs) on benchmark datasets, an item response theory (IRT) framework models each benchmark example i∈Ii \in I as an item and each LLM ll as a testee. Under the two-parameter multidimensional IRT model, the probability that model ll correctly answers item ii is given by:

    pil≜P(Yil=1∣θl,αi,βi)=11+exp⁡(−αi⊤θl+βi)p_{il} \triangleq \mathbb{P}(Y_{il} = 1 \mid \theta_l, \alpha_i, \beta_i) = \frac{1}{1 + \exp(-\alpha_i^\top \theta_l + \beta_i)}

    where θl∈Rd\theta_l \in \mathbb{R}^d represents the unobserved latent ability vector of LLM ll, αi∈Rd\alpha_i \in \mathbb{R}^d is the item discrimination vector indicating which latent abilities are required to answer item ii correctly, and βi∈R\beta_i \in \mathbb{R} is the item bias/difficulty parameter regulating the probability of correctness when θl=0\theta_l = \mathbf{0}.

    Model parameters are fitted via variational Bayesian inference on a reference set of models Ltr\mathcal{L}_{\text{tr}} using hierarchical priors:

    θl∼N(μθ1d,uθ−1Id),αi∼N(μα1d,uα−1Id),βi∼N(μβ,uβ−1)\theta_l \sim \mathcal{N}(\mu_\theta \mathbf{1}_d, u_\theta^{-1} \mathbf{I}_d), \quad \alpha_i \sim \mathcal{N}(\mu_\alpha \mathbf{1}_d, u_\alpha^{-1} \mathbf{I}_d), \quad \beta_i \sim \mathcal{N}(\mu_\beta, u_\beta^{-1})

    with hyperpriors μθ,μα,μβ∼N(0,10)\mu_\theta, \mu_\alpha, \mu_\beta \sim \mathcal{N}(0, 10) and uθ,uα,uβ∼Γ(1,1)u_\theta, u_\alpha, u_\beta \sim \Gamma(1, 1). Point estimates (α^i,β^i)(\hat{\alpha}_i, \hat{\beta}_i) and θ^l\hat{\theta}_l are taken from the means of the variational posterior distributions. The latent dimensionality d∈{2,5,10,15}d \in \{2, 5, 10, 15\} is selected via validation on training LLM evaluation data.

  2. Knowl 2 — Performance-IRT (p-IRT) Estimator

    model/method

    Let scenario jj of a benchmark comprise a set of items IjI_j. The true mean performance of model ll on scenario jj is:

    Zjl≜1∣Ij∣∑i∈IjYilZ_{jl} \triangleq \frac{1}{|I_j|} \sum_{i \in I_j} Y_{il}

    where Yil∈{0,1}Y_{il} \in \{0, 1\} is the correctness of model ll on item ii. When model ll is evaluated only on a small subset I^j⊂Ij\hat{I}_j \subset I_j of items, the Performance-IRT (p-IRT) estimator approximates ZjlZ_{jl} by estimating the conditional expectation E[Zjl∣{Yil}i∈I^j]\mathbb{E}[Z_{jl} \mid \{Y_{il}\}_{i \in \hat{I}_j}]:

    Z^jlp-IRT=λ^∣I^j∣∑i∈I^jYil+1−λ^∣Ij∖I^j∣∑i∈Ij∖I^jp^il\hat{Z}_{jl}^{\text{p-IRT}} = \frac{\hat{\lambda}}{|\hat{I}_j|} \sum_{i \in \hat{I}_j} Y_{il} + \frac{1 - \hat{\lambda}}{|I_j \setminus \hat{I}_j|} \sum_{i \in I_j \setminus \hat{I}_j} \hat{p}_{il}

    where λ^=∣I^j∣/∣Ij∣∈[0,1]\hat{\lambda} = |\hat{I}_j| / |I_j| \in [0, 1] weights the contribution of observed items, and p^il=[1+exp⁡(−α^i⊤θ^l+β^i)]−1\hat{p}_{il} = [1 + \exp(-\hat{\alpha}_i^\top \hat{\theta}_l + \hat{\beta}_i)]^{-1} is the IRT model's predicted probability of correctness for unobserved item ii. The item parameters (α^i,β^i)(\hat{\alpha}_i, \hat{\beta}_i) are pre-computed on a training set of LLMs, while the ability vector θ^l\hat{\theta}_l of a new LLM ll is estimated by maximizing the log-likelihood over the newly observed responses {Yil}i∈I^j\{Y_{il}\}_{i \in \hat{I}_j} via logistic regression M-estimation while keeping item parameters fixed.

  3. Knowl 3 — Generalized p-IRT (gp-IRT) Estimator

    model/method

    The generalized p-IRT (gp-IRT) estimator combines the weighted empirical evaluation on a sample I^j⊂Ij\hat{I}_j \subset I_j with the model-based p-IRT estimator Z^jlp-IRT\hat{Z}_{jl}^{\text{p-IRT}} via a convex combination:

    Z^jlgp-IRT≜λ∑i∈I^jwiYil+(1−λ)Z^jlp-IRT\hat{Z}_{jl}^{\text{gp-IRT}} \triangleq \lambda \sum_{i \in \hat{I}_j} w_i Y_{il} + (1 - \lambda) \hat{Z}_{jl}^{\text{p-IRT}}

    where {wi}i∈I^j\{w_i\}_{i \in \hat{I}_j} are non-negative weights summing to 1, and λ∈[0,1]\lambda \in [0, 1] optimizes the trade-off between the variance of the sample average and the potential bias of the IRT model due to misspecification.

    For random sampling, λ\lambda is chosen as:

    λ=b^2σ^2/∣I^j∣+b^2\lambda = \frac{\hat{b}^2}{\hat{\sigma}^2 / |\hat{I}_j| + \hat{b}^2}

    where σ^2\hat{\sigma}^2 is the average empirical variance of YilY_{il} (i∈Iji \in I_j) across LLMs in the training set, and b^2\hat{b}^2 is an estimate of IRT bias obtained by splitting the training set of LLMs into two halves, fitting the IRT model on the first half, estimating ability parameters on half the examples for the second half of LLMs, and calculating the squared mean absolute error between IRT predicted scenario scores and actual scores on unseen examples. When using lower-variance anchor points, σ^2\hat{\sigma}^2 is divided by 4 (equivalent to halving the standard deviation).

  4. Knowl 4 — Asymptotic Convergence of the Conditional Expectation in p-IRT

    theoretical result

    Let II denote the set of all benchmark items across all scenarios, and let I^\hat{I} denote the subset of items evaluated for LLM ll. Assume that:

    1. The ability estimate θ^l\hat{\theta}_l converges in probability to the true ability parameter θl\theta_l as ∣I^∣→∞|\hat{I}| \to \infty (i.e., θ^l→pθl\hat{\theta}_l \xrightarrow{p} \theta_l).
    2. The true item parameters (αi,βi)(\alpha_i, \beta_i) are known for all i∈Ii \in I, and there exists a universal constant c<∞c < \infty such that sup⁡i∈I∥αi∥2≤c\sup_{i \in I} \|\alpha_i\|_2 \le c.

    Then, for scenario jj with item set IjI_j and true mean score Zjl=1∣Ij∣∑i∈IjYilZ_{jl} = \frac{1}{|I_j|} \sum_{i \in I_j} Y_{il}, the estimated conditional expectation E^[Zjl∣{Yil}i∈I^j]\hat{\mathbb{E}}[Z_{jl} \mid \{Y_{il}\}_{i \in \hat{I}_j}] (the p-IRT estimator) converges in probability to the true conditional expectation E[Zjl∣{Yil}i∈I^j]\mathbb{E}[Z_{jl} \mid \{Y_{il}\}_{i \in \hat{I}_j}] as ∣I^∣→∞|\hat{I}| \to \infty:

    ∣E^[Zjl∣{Yil}i∈I^j]−E[Zjl∣{Yil}i∈I^j]∣→p0\left| \hat{\mathbb{E}}\left[Z_{jl} \mid \{Y_{il}\}_{i \in \hat{I}_j}\right] - \mathbb{E}\left[Z_{jl} \mid \{Y_{il}\}_{i \in \hat{I}_j}\right] \right| \xrightarrow{p} 0

  5. Knowl 5 — IRT Anchor Point Selection for Benchmark Subsampling

    algorithm

    To select a representative subset I^j⊂Ij\hat{I}_j \subset I_j of KK items for evaluating an LLM on scenario jj, items are clustered in the latent parameter space learned by an IRT model.

    Input: Scenario item set IjI_j, pre-fitted IRT parameter estimates (α^i,β^i)(\hat{\alpha}_i, \hat{\beta}_i) for each item i∈Iji \in I_j, target subset size KK, normalized balance weights {ωˉi}i∈Ij\{\bar{\omega}_i\}_{i \in I_j}
    Output: Anchor subset I^j\hat{I}_j, anchor weights {wi}i∈I^j\{w_i\}_{i \in \hat{I}_j}
    for each item i∈Iji \in I_j do
        Ei←(α^i,β^i)∈Rd+1E_i \leftarrow (\hat{\alpha}_i, \hat{\beta}_i) \in \mathbb{R}^{d+1}
    end for
    Run K-Means clustering on embeddings {Ei}i∈Ij\{E_i\}_{i \in I_j} with KK clusters, yielding cluster assignments C1,…,CKC_1, \dots, C_K and centroids μ1,…,μK\mu_1, \dots, \mu_K
    for each cluster k∈{1,…,K}k \in \{1, \dots, K\} do
        Find anchor item ik∗=arg⁡min⁡i∈Ck∥Ei−μk∥2i_k^* = \arg\min_{i \in C_k} \|E_i - \mu_k\|_2
        Compute anchor weight wik∗=∑i∈Ckωˉiw_{i_k^*} = \sum_{i \in C_k} \bar{\omega}_i
    end for
    I^j←{i1∗,…,iK∗}\hat{I}_j \leftarrow \{i_1^*, \dots, i_K^*\}
    return I^j\hat{I}_j, {wik∗}k=1K\{w_{i_k^*}\}_{k=1}^K

    When all subscenarios are equally sized or absent, ωˉi=1/∣Ij∣\bar{\omega}_i = 1 / |I_j|; when subscenarios have differing sizes, ωˉi=1sj∣Ijk∣\bar{\omega}_i = \frac{1}{s_j |I_{jk}|} for item ii belonging to subscenario kk.

  6. Knowl 6 — Score Binarization for IRT Modeling of Continuous and Bounded Benchmarks

    model/method

    When a benchmark produces non-binary bounded correctness scores Yil∈[0,1]Y_{il} \in [0, 1] (such as win rates in AlpacaEval 2.0 or F1 scores in HELM), standard binary IRT estimation cannot be applied directly. The continuous scores are binarized into an indicator variable Y~il=1[Yil≥c]\tilde{Y}_{il} = \mathbf{1}[Y_{il} \ge c] for a scenario-specific threshold cc.

    For each scenario jj, the constant c∈[0,1]c \in [0, 1] is selected to equalize the aggregate sum of continuous scores and indicator values across all items in scenario jj and all models in the training set Ltr\mathcal{L}_{\text{tr}}:

    ∑i∈Ij,l∈LtrYil≈∑i∈Ij,l∈Ltr1[Yil≥c]\sum_{i \in I_j, l \in \mathcal{L}_{\text{tr}}} Y_{il} \approx \sum_{i \in I_j, l \in \mathcal{L}_{\text{tr}}} \mathbf{1}[Y_{il} \ge c]

    The resulting binary variables Y~il∈{0,1}\tilde{Y}_{il} \in \{0, 1\} are then modeled using standard binary IRT variational inference and logistic regression ability estimation.

  7. Knowl 7 — Weighted Scenario Accuracy and p-IRT Estimation with Subscenario Balancing

    equation

    When a benchmark scenario jj contains sjs_j subscenarios Ij1,…,IjsjI_{j1}, \dots, I_{j s_j} of differing sample sizes, the unweighted average across subscenarios defines the overall scenario score for model ll:

    Zjl=1sj∑k=1sj1∣Ijk∣∑i∈IjkYil=∑i∈IjωˉiYilZ_{jl} = \frac{1}{s_j} \sum_{k=1}^{s_j} \frac{1}{|I_{jk}|} \sum_{i \in I_{jk}} Y_{il} = \sum_{i \in I_j} \bar{\omega}_i Y_{il}

    where ωˉi=1sj∣Ijk∣\bar{\omega}_i = \frac{1}{s_j |I_{jk}|} is the normalized balance weight for item i∈Ijki \in I_{jk}, satisfying ∑i∈Ijωˉi=1\sum_{i \in I_j} \bar{\omega}_i = 1, and ωi≜∣Ij∣ωˉi\omega_i \triangleq |I_j| \bar{\omega}_i is the unnormalized balance weight. The subscenario-balanced p-IRT estimator for an evaluated subset I^j⊂Ij\hat{I}_j \subset I_j with IRT predictions p^il\hat{p}_{il} is:

    Z^jlp-IRT=λ^∣I^j∣∑i∈I^jωiYil+1−λ^∣Ij∖I^j∣∑i∈Ij∖I^jωip^il\hat{Z}_{jl}^{\text{p-IRT}} = \frac{\hat{\lambda}}{|\hat{I}_j|} \sum_{i \in \hat{I}_j} \omega_i Y_{il} + \frac{1 - \hat{\lambda}}{|I_j \setminus \hat{I}_j|} \sum_{i \in I_j \setminus \hat{I}_j} \omega_i \hat{p}_{il}

    where λ^=∣I^j∣/∣Ij∣\hat{\lambda} = |\hat{I}_j| / |I_j|.

  8. Knowl 8 — LLM Benchmark Downsampling Performance of tinyBenchmarks

    empirical result

    Evaluating LLMs using tinyBenchmarks—comprising 100 curated IRT anchor examples per scenario combined with the generalized p-IRT estimator (IRT++)—achieves an average performance estimation error within approximately 2% (0.02) of the true full-benchmark performance across four standard benchmarks:

    • MMLU (14,000 total examples across 57 subjects): 100 curated items achieve an average accuracy estimation error of 1.9% on test models evaluated chronologically after the training set, with maximum error ≤4%\le 4\% across test models, reducing evaluation cost by a factor of 140.
    • HuggingFace Open LLM Leaderboard (29,000 total examples across 6 scenarios): Evaluating 30 curated items per scenario (180 items total, a 160x reduction) or 100 items per scenario (600 items total) achieves average performance error below 2%.
    • HELM Lite v1.0.0 (10,000 total examples across 10 scenarios): 100 items per scenario (1,000 items total) predict mean win rate across scenarios within 2% error on unseen organizational model splits.
    • AlpacaEval 2.0 (805 total examples): 100 curated items predict GPT-4 win rate with an average error of approximately 1.5%, saving substantial judge API costs.
  9. Knowl 9 — Robustness of IRT Anchor Selection to Domain Specialization

    empirical result

    On MMLU, selecting anchor items via IRT parameter clustering maintains an average accuracy estimation error of approximately 2% when evaluated on a test set of 40 domain-specialized LLMs (models specifically fine-tuned for mathematics, coding, biology, or finance).

    In contrast, anchor selection via correctness clustering (grouping examples by the raw binary prediction vector across training models) degrades significantly on domain-specialized models, yielding substantially higher estimation error. This occurs because domain-specialized models possess localized strengths that violate the empirical correctness correlation patterns observed among general-purpose models, whereas IRT anchor points yield higher effective sample size (ESS = 0.85 for IRT vs. ESS = 0.53 for correctness clustering) and more uniformly distributed subscenario weights.

  10. Knowl 10 — Prompt Template Evaluation and Adaptive Testing Extensions

    empirical result

    The IRT evaluation framework extends effectively to prompt template sensitivity and adaptive testing:

    • Prompt Template Evaluation: When evaluated on ANLI (750 examples wrapped in 15 prompt templates from PromptSource) using 8 LLaMA models, an IRT model trained on smaller models (7B, 13B, 30B) accurately predicts the performance of larger models (65B) and unseen prompt templates within 4% error using 100 examples.
    • Adaptive Testing (adaptIRT++): Dynamically selecting successive evaluation examples based on computerized adaptive testing principles on MMLU reduces estimation error further than fixed anchor sets at small sample sizes (20 to 60 examples), though it requires approximately 5 minutes of online inference per model compared to negligible CPU time for static gp-IRT.
  11. Knowl 11 — Vulnerability of IRT Benchmarking to Severe Architectural Shifts and Capability Extrapolation

    limitation

    The accuracy of IRT-based benchmark subsampling relies on the assumption that item difficulty (βi\beta_i) and discrimination (αi\alpha_i) parameters remain stable across model generations. The framework exhibits higher estimation error under two types of severe distribution shifts:

    1. Inverted correctness patterns: If a novel model architecture or pre-training distribution causes an LLM to fail on simple items while succeeding on complex items, the model violates the IRT monotonicity assumption.
    2. Capability extrapolation: When evaluating models whose capabilities substantially exceed all reference models in the training set Ltr\mathcal{L}_{\text{tr}}, the ability parameter estimate θ^l\hat{\theta}_l suffers from extrapolation errors.

    To prevent error accumulation over time, curated anchor points and IRT parameter estimates must be periodically re-fitted using evaluation data from contemporary models.

Coverage note — None was omitted; all key theoretical formulations, estimators, algorithms, empirical benchmarks, extensions, and limitations were extracted into standalone knowls.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.An, X. and Yung, Y.-F. Item response theory: What it is and how you can use the irt procedure to apply it. SAS Institute Inc, 10(4):364–2014, 2014.
  3. 3.Bach, S., Sanh, V., Yong, Z. X., Webson, A., Raffel, C., Nayak, N. V., Sharma, A., Kim, T., Bari, M. S., Fevry, T., Alyafeai, Z., Dey, M., Santilli, A., Sun, Z., Ben-david, S., Xu, C., Chhablani, G., Wang, H., Fries, J., Al-shaibani, M., Sharma, S., Thakker, U., Almubarak, K., Tang, X., Radev, D., Jiang, M. T.-j., and Rush, A. PromptSource: An integrated development environment and repository for natural language prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pp. 93–104, Dublin, Ireland, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-demo.9. URL https://aclanthology.org/2022.acl-demo.9.
  4. 4.Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T. Open llm leaderboard. https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard, 2023.
  5. 5.Biderman, S., Prashanth, U. S., Sutawika, L., Schoelkopf, H., Anthony, Q., Purohit, S., and Raf, E. Emergent and predictable memorization in large language models. arXiv preprint arXiv:2304.11158, 2023a.
  6. 6.Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and van der Wal, O. Pythia: A suite for analyzing large language models across training and scaling. ArXiv, abs/2304.01373, 2023b. URL https://api.semanticscholar.org/CorpusID:257921893.
  7. 7.Bojar, O., Buck, C., Federmann, C., Haddow, B., Koehn, P., Leveling, J., Monz, C., Pecina, P., Post, M., Saint-Amand, H., et al. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the ninth workshop on statistical machine translation, pp. 12–58, 2014.
  8. 8.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  9. 9.Brzezińska, J. Item response theory models in the measurement theory. Communications in Statistics-Simulation and Computation, 49(12):3299–3313, 2020.
  10. 10.Cai, L., Choi, K., Hansen, M., and Harrell, L. Item response theory. Annual Review of Statistics and Its Application, 3:297–321, 2016.
  11. 11.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  12. 12.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  13. 13.Ein-Dor, L., Halfon, A., Gera, A., Shnarch, E., Dankin, L., Choshen, L., Danilevsky, M., Aharonov, R., Katz, Y., and Slonim, N. Active Learning for BERT: An Empirical Study. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 7949–7962, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.638. URL https://aclanthology.org/2020.emnlp-main.638.
  14. 14.Elvira, V., Martino, L., and Robert, C. P. Rethinking the effective sample size. International Statistical Review, 90 (3):525–550, 2022.
  15. 15.Fahrmeir, L. and Kaufmann, H. Consistency and asymptotic normality of the maximum likelihood estimator in generalized linear models. The Annals of Statistics, 13 (1):342–368, 1985.
  16. 16.Guha, N., Nyarko, J., Ho, D., Re, C., Chilton, A., Chohlas-Wood, A., Peters, A., Waldon, B., Rockmore, D., Zambrano, D., et al. Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models. Advances in Neural Information Processing Systems, 36, 2024.
  17. 17.Hastie, T., Tibshirani, R., Friedman, J. H., and Friedman, J. H. The elements of statistical learning: data mining, inference, and prediction, volume 2. Springer, 2009.
  18. 18.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  19. 19.Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.
  20. 20.Ji, D., Logan, R. L., Smyth, P., and Steyvers, M. Active bayesian assessment of black-box classifiers. Proceedings of the AAAI Conference on Artificial Intelligence, 35(9):7935–7944, May 2021. doi: 10.1609/aaai.v35i9.16968. URL https://ojs.aaai.org/index.php/AAAI/article/view/16968.
  21. 21.Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421, 2021.
  22. 22.Katariya, N., Iyer, A., and Sarawagi, S. Active evaluation of classifiers on large datasets. In 2012 IEEE 12th International Conference on Data Mining, pp. 329–338, 2012. doi: 10.1109/ICDM.2012.161.
  23. 23.Kingston, N. M. and Dorans, N. J. The feasibility of using item response theory as a psychometric model for the gre aptitude test. ETS Research Report Series, 1982(1):i–148, 1982.
  24. 24.Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E. The narrativeqa reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328, 2018.
  25. 25.Kossen, J., Farquhar, S., Gal, Y., and Rainforth, T. Active testing: Sample-efficient model evaluation. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 5753–5763. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/kossen21a.html.
  26. 26.Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466, 2019.
  27. 27.Lalor, J. P. and Rodriguez, P. py-irt: A scalable item response theory library for python. INFORMS Journal on Computing, 35(1):5–13, 2023.
  28. 28.Lalor, J. P., Wu, H., and Yu, H. Building an evaluation scale using item response theory. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing, volume 2016, pp. 648. NIH Public Access, 2016.
  29. 29.Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval, 2023.
  30. 30.Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022.
  31. 31.Lin, S., Hilton, J., and Evans, O. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958, 2021.
  32. 32.Liu, Z., Qiao, A., Neiswanger, W., Wang, H., Tan, B., Tao, T., Li, J., Wang, Y., Sun, S., Pangarkar, O., et al. Llm360: Towards fully transparent open-source llms. arXiv preprint arXiv:2312.06550, 2023.
  33. 33.Lord, F., Novick, M., and Birnbaum, A. Statistical theories of mental test scores. 1968.
  34. 34.Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8086–8098, Dublin, Ireland, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.556. URL https://aclanthology.org/2022.acl-long.556.
  35. 35.Maia Polo, F. and Vicente, R. Effective sample size, dimensionality, and generalization in covariate shift adaptation. Neural Computing and Applications, 35(25):18187–18199, 2023.
  36. 36.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018.
  37. 37.Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 11048–11064, Abu Dhabi, United Arab Emirates, 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.759.
  38. 38.Mishra, S., Khashabi, D., Baral, C., Choi, Y., and Hajishirzi, H. Reframing instructional prompts to GPTk’s language. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 589–612, Dublin, Ireland, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.findings-acl.50. URL https://aclanthology.org/2022.findings-acl.50.
  39. 39.Mizrahi, M., Kaplan, G., Malkin, D., Dror, R., Shahaf, D., and Stanovsky, G. State of what art? a call for multi-prompt llm evaluation. arXiv preprint arXiv:2401.00595, 2023.
  40. 40.Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., and Kiela, D. Adversarial nli: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4885–4901, 2020.
  41. 41.Perlitz, Y., Bandel, E., Gera, A., Arviv, O., Ein-Dor, L., Shnarch, E., Slonim, N., Shmueli-Scheuer, M., and Choshen, L. Efficient benchmarking (of language models). arXiv preprint arXiv:2308.11696, 2023.
  42. 42.Petersen, N. S. et al. Using item response theory to equate scholastic aptitude test scores. 1982.
  43. 43.Rodriguez, P., Barrow, J., Hoyle, A. M., Lalor, J. P., Jia, R., and Boyd-Graber, J. Evaluation examples are not equally informative: How should that change NLP leaderboards? In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 4486–4503, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.346. URL https://aclanthology.org/2021.acl-long.346.
  44. 44.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  45. 45.Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. arXiv preprint arXiv:2310.11324, 2023.
  46. 46.Song, W. T. Minimal-mse linear combinations of variance estimators of the sample mean. In 1988 Winter Simulation Conference Proceedings, pp. 414–421. IEEE, 1988.
  47. 47.Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.
  48. 48.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
  49. 49.Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  50. 50.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  51. 51.Van der Linden, W. J. Handbook of item response theory: Three volume set. CRC Press, 2018.
  52. 52.Vania, C., Htut, P. M., Huang, W., Mungra, D., Pang, R. Y., Phang, J., Liu, H., Cho, K., and Bowman, S. R. Comparing test sets with item response theory. arXiv preprint arXiv:2106.00840, 2021.
  53. 53.Vivek, R., Ethayarajh, K., Yang, D., and Kiela, D. Anchor points: Benchmarking models with much fewer examples. arXiv preprint arXiv:2309.08638, 2023.
  54. 54.Voronov, A., Wolf, L., and Ryabinin, M. Mind your format: Towards consistent evaluation of in-context learning improvements. arXiv preprint arXiv:2401.06766, 2024.
  55. 55.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  56. 56.Weber, L., Bruni, E., and Hupkes, D. The icl consistency test. arXiv preprint arXiv:2312.04945, 2023a.
  57. 57.Weber, L., Bruni, E., and Hupkes, D. Mind the instructions: a holistic evaluation of consistency and interactions in prompt-based learning. arXiv preprint arXiv:2310.13486, 2023b.
  58. 58.Wei, J., Wei, J., Tay, Y., Tran, D., Webson, A., Lu, Y., Chen, X., Liu, H., Huang, D., Zhou, D., et al. Larger language models do in-context learning differently. ArXiv preprint, abs/2303.03846, 2023. URL https://arxiv.org/abs/2303.03846.
  59. 59.Ye, Q., Fu, H. Y., Ren, X., and Jia, R. How predictable are large language model capabilities? a case study on big-bench. arXiv preprint arXiv:2305.14947, 2023.
  60. 60.Yoo, K. M., Kim, J., Kim, H. J., Cho, H., Jo, H., Lee, S.-W., Lee, S.-g., and Kim, T. Ground-truth labels matter: A deeper look into input-label demonstrations. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 2422–2437, Abu Dhabi, United Arab Emirates, 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.155.
  61. 61.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  62. 62.Zhuang, Y., Liu, Q., Ning, Y., Huang, W., Lv, R., Huang, Z., Zhao, G., Zhang, Z., Mao, Q., Wang, S., et al. Efficiently measuring the cognitive ability of llms: An adaptive testing perspective. arXiv preprint arXiv:2306.10512, 2023.

Citation

MLA
Polo, F. M., et al. “tinyBenchmarks: Evaluating LLMs with Fewer Examples”. arXiv, 2024, http://arxiv.org/abs/2402.14992v2.
APA
Polo, F. M., Weber, L., Choshen, L., Sun, Y., Xu, G., & Yurochkin, M. (2024). tinyBenchmarks: evaluating LLMs with fewer examples. arXiv. http://arxiv.org/abs/2402.14992v2
Chicago
Polo, F. M., L. Weber, L. Choshen, Y. Sun, G. Xu, and M. Yurochkin. 2024. “tinyBenchmarks: Evaluating LLMs with Fewer Examples”. arXiv. http://arxiv.org/abs/2402.14992v2.
Harvard
Polo, F.M. et al. (2024) “tinyBenchmarks: evaluating LLMs with fewer examples”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.14992v2.
Vancouver
1. Polo FM, Weber L, Choshen L, Sun Y, Xu G, Yurochkin M (2024) tinyBenchmarks: evaluating LLMs with fewer examples. arXiv

BibTeX

@article{polo2024tinybenchmarks,
  title = {tinyBenchmarks: evaluating LLMs with fewer examples},
  author = {Polo, Felipe Maia and Weber, Lucas and Choshen, Leshem and Sun, Yuekai and Xu, Gongjun and Yurochkin, Mikhail},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.14992v2},
  eprint = {2402.14992}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/