AutoEval Done Right: Using Synthetic Data for Model Evaluation

Pierre BoyeauAnastasios Nikolas AngelopoulosTianle LiNir YosefJitendra MalikMichael I. Jordan

article2025ICML63 citations

Develops a statistically rigorous autoevaluation framework using prediction-powered inference to combine limited human annotations with abundant synthetic data, delivering unbiased model performance estimates and tight confidence intervals at a fraction of standard labeling costs.

Listen

Evaluating modern artificial intelligence systems requires extensive validation data to ensure accuracy, fairness, and safety. Relying solely on human-annotated datasets is expensive and time-consuming, while using automated artificial intelligence annotators alone can introduce hidden biases that lead to unreliable performance assessments.

The article demonstrates an automated evaluation framework that integrates a small set of human-labeled data with large volumes of synthetic, model-generated labels. The objective is to produce mathematically unbiased performance estimates and valid confidence intervals while significantly reducing the need for costly human annotations.

The approach applies prediction-powered inference, using human labels to quantify and subtract the systematic bias of artificial intelligence annotators. The authors validated this methodology across three distinct domains: classifying images using computer vision architectures on ImageNet, evaluating zero-shot protein fitness regression models on biological assay benchmarks, and ranking twenty large language models using pairwise comparison preferences from human and automated judges.

The analysis yielded several key findings. First, incorporating synthetic data increased the effective sample size by 20% to 50% compared to human-only evaluation, matching the precision of much larger human datasets. Second, the optimized estimator consistently produced lower error rates and narrower confidence intervals while remaining strictly unbiased. Third, model ranking accuracy improved significantly across all domains, yielding up to a five-fold correlation improvement with ground-truth benchmarks in the protein study. Finally, when annotator models were weak or uninformative, the framework adaptively discounted the synthetic data, performing at least as well as traditional human-only testing.

These findings indicate that organizations can drastically reduce annotation expenses and evaluation timelines without sacrificing statistical rigor or safety guarantees. Rather than discarding artificial intelligence judges or trusting them blindly, decision-makers can leverage them to stretch limited human validation budgets further.

Organizations should adopt debiasing frameworks when using synthetic labels for model evaluation and benchmarking. In practice, teams should also apply importance reweighting techniques whenever validation data might suffer from distribution shifts between human samples and production workloads. The methodology relies on the assumption that sampled data reflects target distributions, but with appropriate shift adjustments, stakeholders can have high confidence in the resulting evaluations.

No sufficiently relevant recommendations were found.

Cover for AutoEval Done Right: Using Synthetic Data for Model Evaluation

Abstract

The evaluation of machine learning models using human-labeled validation data can be expensive and time-consuming. AI-labeled synthetic data can be used to decrease the number of human annotations required for this purpose in a process called autoevaluation. We suggest efficient and statistically principled algorithms for this purpose that improve sample efficiency while remaining unbiased.

Table of Contents

  • 1 Introduction
  • 1.1 Related Work
  • 2 Autoevaluating Accuracy and other Metrics
  • 2.1 Defining the Goal
  • 2.2 The Algorithm
  • 2.3 Application to Rank Computer Vision Models
  • 2.4 Application to Evaluate Protein Fitness Prediction Models
  • 3 Evaluating Model Performance from Pairwise Comparisons
  • 3.1 A Model to Assess Relative Performance
  • 3.2 Autoevaluation of Relative Performance
  • 3.3 Autoevaluation of LLMs from Pairwise Preferences
  • References
  • A Code snippets
  • B Handling covariate shifts
  • C Experimental details
  • C.1 Data acquisition and preprocessing
  • C.2 Methodological details
  • C.3 Experimental setup
  • D Additional experiments
  • D.1 Running times
  • D.2 Larger sample sizes
  • D.3 Coverage analysis

Knowls

  1. Knowl 1 — Bias-corrected estimation with synthetic labels

    model/method

    Let nn human-labeled examples (Xi,Yi)(X_i,Y_i) and NN unlabeled inputs XjuX_j^u be independent samples from the same target population, with NN potentially much larger than nn. For each of MM evaluated models fmf_m and a metric function ϕ\phi, the target is μm=E[ϕ(fm(X),Y)]\mu_m=\mathbb{E}[\phi(f_m(X),Y)]. Suppose an annotator supplies a synthetic label distribution P~i,m\widetilde P_{i,m} for each labeled input and P~j,mu\widetilde P^u_{j,m} for each unlabeled input. Define the annotator-imputed metric by E^i,m=∫ϕ(fm(Xi),y) dP~i,m(y)\widehat E_{i,m}=\int \phi(f_m(X_i),y)\,d\widetilde P_{i,m}(y), with E^j,mu\widehat E^u_{j,m} defined analogously. AutoEval estimates the target as

    μ^m=λN∑j=1NE^j,mu+1n∑i=1n[ϕ(fm(Xi),Yi)−λE^i,m],0≤λ≤1.\widehat\mu_m=\frac{\lambda}{N}\sum_{j=1}^{N}\widehat E^u_{j,m}+\frac{1}{n}\sum_{i=1}^{n}\left[\phi(f_m(X_i),Y_i)-\lambda\widehat E_{i,m}\right],\qquad 0\leq\lambda\leq1.

    The first term averages the synthetic metric over the large unlabeled sample; the second corrects its bias using the human-labeled sample. Under the matched-sampling assumption, this estimator is unbiased for every fixed λ\lambda, even if the annotator is inaccurate. A point-valued synthetic label is covered by taking P~\widetilde P to be a point mass at that label. The construction applies to accuracy, losses, and other metrics expressible as expectations.

  2. Knowl 2 — Variance optimization and the role of PPI++

    theoretical result

    For the bias-corrected estimator, let Lju=(E^j,1u,…,E^j,Mu)L^u_j=(\widehat E^u_{j,1},\ldots,\widehat E^u_{j,M}) be the vector of synthetic metrics on unlabeled example jj, and let Δi=(ϕ(f1(Xi),Yi)−λE^i,1,…,ϕ(fM(Xi),Yi)−λE^i,M)\Delta_i=(\phi(f_1(X_i),Y_i)-\lambda\widehat E_{i,1},\ldots,\phi(f_M(X_i),Y_i)-\lambda\widehat E_{i,M}). With independent labeled and unlabeled samples, its covariance is

    Cov⁡(μ^)=λ2NCov⁡(Lju)+1nCov⁡(Δi).\operatorname{Cov}(\widehat\mu)=\frac{\lambda^2}{N}\operatorname{Cov}(L^u_j)+\frac{1}{n}\operatorname{Cov}(\Delta_i).

    The weight λ\lambda discounts the synthetic contribution: λ=0\lambda=0 recovers the classical human-only estimator, while λ=1\lambda=1 gives the basic PPI estimator. PPI++ chooses λ\lambda to minimize the estimated variance, with limited additional computation. The paper reports that this variance-optimized estimator is asymptotically no less efficient than the classical estimator; when synthetic predictions carry no useful information, optimization can set λ\lambda to zero.

  3. Knowl 3 — Asymptotic confidence intervals for multiple metrics

    theoretical result

    For MM metrics estimated from nn labeled and NN unlabeled examples under the matched-sampling assumptions, define V=nCov⁡(μ^)V=n\operatorname{Cov}(\widehat\mu) and estimate it with the plug-in covariance V^=(nλ2/N)Cov⁡^(Lju)+Cov⁡^(Δi)\widehat V=(n\lambda^2/N)\widehat{\operatorname{Cov}}(L^u_j)+\widehat{\operatorname{Cov}}(\Delta_i). As n,N→∞n,N\to\infty, the standardized estimator satisfies n V^−1/2(μ^−μ)⇒N(0,IM)\sqrt n\,\widehat V^{-1/2}(\widehat\mu-\mu)\Rightarrow\mathcal N(0,I_M). Thus an asymptotically marginal 1−α1-\alpha interval for metric mm is

    μ^m  ±  z1−α/2V^mm/n,\widehat\mu_m\;\pm\;z_{1-\alpha/2}\sqrt{\widehat V_{mm}/n},

    where z1−α/2z_{1-\alpha/2} is the corresponding standard-normal quantile. A simultaneous confidence set for the full metric vector is {μ:n(μ^−μ)TV^−1(μ^−μ)≤χ1−α,M2}\{\mu:n(\widehat\mu-\mu)^T\widehat V^{-1}(\widehat\mu-\mu)\leq\chi^2_{1-\alpha,M}\}, where χ1−α,M2\chi^2_{1-\alpha,M} is the 1−α1-\alpha quantile of a chi-squared distribution with MM degrees of freedom. These results provide uncertainty estimates for both individual metrics and joint comparisons.

  4. Knowl 4 — Autoevaluation of Bradley–Terry model rankings

    model/method

    For pairwise comparisons among MM models, let ci∈RMc_i\in\mathbb R^M encode a comparison, with −1-1 at the first model's coordinate and +1+1 at the second's; let Yi=1Y_i=1 mean the second model is preferred. The Bradley–Terry model assigns preference probability σ(ciTζ)\sigma(c_i^T\zeta), where σ(t)=1/(1+e−t)\sigma(t)=1/(1+e^{-t}) and the coefficient vector ζ\zeta is identifiable by setting ζ1=0\zeta_1=0. Let ℓζ(c,y)=−ylog⁡σ(cTζ)−(1−y)log⁡[1−σ(cTζ)]\ell_\zeta(c,y)=-y\log\sigma(c^T\zeta)-(1-y)\log[1-\sigma(c^T\zeta)] be binary cross-entropy. With human labels YiY_i on nn comparisons, AI-judge preferences Y^i\widehat Y_i on those comparisons, and judge preferences Y^ju\widehat Y^u_j on NN additional unlabeled comparisons cjuc^u_j, the AutoEval estimator minimizes

    ζ^=arg⁡min⁡ζ: ζ1=0{1n∑i=1n[ℓζ(ci,Yi)−λℓζ(ci,Y^i)]+λN∑j=1Nℓζ(cju,Y^ju)}.\widehat\zeta=\arg\min_{\zeta:\,\zeta_1=0}\left\{\frac1n\sum_{i=1}^n\left[\ell_\zeta(c_i,Y_i)-\lambda\ell_\zeta(c_i,\widehat Y_i)\right]+\frac{\lambda}{N}\sum_{j=1}^N\ell_\zeta(c^u_j,\widehat Y^u_j)\right\}.

    The human-label term corrects the judge-based loss, while λ∈[0,1]\lambda\in[0,1] controls the synthetic preferences' contribution. The paper presents this as an unbiased, lower-variance loss construction and reports that PPI/PPI++ inference provides confidence intervals for the Bradley–Terry coefficients.

  5. Knowl 5 — Importance-weighted estimators for covariate shift

    model/method

    The paper also treats covariate shift in which labeled inputs follow QXQ_X but the target and unlabeled inputs follow PXP_X, while the conditional label distribution PY∣XP_{Y\mid X} is shared. If the density ratio w(x)=dPX/dQXw(x)=dP_X/dQ_X is known, the stated reweighted metric estimator is

    μ^mw=λN∑j=1NE^j,mu+1n∑i=1nΔi,mλ,w,Δi,mλ,w=w(Xi)ϕ(fm(Xi),Yi)−λE^i,m.\widehat\mu^w_m=\frac{\lambda}{N}\sum_{j=1}^{N}\widehat E^u_{j,m}+\frac1n\sum_{i=1}^{n}\Delta^{\lambda,w}_{i,m},\qquad \Delta^{\lambda,w}_{i,m}=w(X_i)\phi(f_m(X_i),Y_i)-\lambda\widehat E_{i,m}.

    For pairwise comparisons, the corresponding stated objective replaces the human loss by its importance-weighted version:

    ζ^=arg⁡min⁡ζ: ζ1=0{1n∑i=1n[w(Xi)ℓζ(ci,Yi)−λℓζ(ci,Y^i)]+λN∑j=1Nℓζ(cju,Y^ju)}.\widehat\zeta=\arg\min_{\zeta:\,\zeta_1=0}\left\{\frac1n\sum_{i=1}^{n}\left[w(X_i)\ell_\zeta(c_i,Y_i)-\lambda\ell_\zeta(c_i,\widehat Y_i)\right]+\frac{\lambda}{N}\sum_{j=1}^{N}\ell_\zeta(c^u_j,\widehat Y^u_j)\right\}.

    In an ImageNet covariate-shift experiment, labeled examples were sampled with probabilities based on ResNet-101's predicted probability for one of 1,000 classes, and importance weights were estimated by self-normalized importance sampling. For 90% confidence intervals, coverage was:

    Could not parse LaTeX table

    The unweighted estimator lost coverage under this shift, whereas the reweighted estimator maintained coverage near the nominal 0.9 level.

  6. Knowl 6 — ImageNet accuracy estimation and ranking

    empirical result

    The ImageNet experiment evaluated ResNet-18, ResNet-34, ResNet-50, ResNet-101, and ResNet-152 on the validation set, treating a small random subset as human-labeled and the remainder as unlabeled. Synthetic label distributions were the models' softmax outputs. Results were averaged over 250 random splits. Across the tested labeled-set sizes, both PPI (λ=1\lambda=1) and PPI++ had lower mean-squared error than the classical accuracy estimate; the paper reports little difference between their point-estimation errors. Their confidence intervals were calibrated and tighter than the classical intervals. PPI/PPI++ effective sample size was approximately 50% above the classical method. For model ranking, models with overlapping Bonferroni-corrected confidence intervals were tied; PPI++ produced substantially higher rank correlation with the ground-truth ordering than the alternatives.

    At n=10,000n=10{,}000 labeled examples, the reported comparison was:

    Could not parse LaTeX table

    Thus, in this larger-sample experiment, PPI++ had the lowest reported MSE and interval width, with coverage near the nominal 90% level and an efficiency ratio of 1.29.

  7. Knowl 7 — Protein-fitness correlation estimation

    empirical result

    The protein-fitness experiment ranked seven protein language models on 536,962 pairwise mutations in the IgG-binding domain of protein G. The target metric was Pearson correlation between a model's predicted fitness and experimental fitness; the scores were standardized to zero mean and unit variance, so the metric was E[Yfm(X)]\mathbb E[Yf_m(X)]. VESPA supplied point-valued synthetic fitness labels. PPI++ increased effective sample size by approximately 50% over the classical estimator, and at n=1,000n=1{,}000 its model-ranking correlation with the ground-truth ranking was reported to improve by a factor of five. Its confidence intervals were slightly tighter than the classical intervals and remained calibrated. Basic PPI performed worse than the classical estimator in this experiment, a weakness the variance-optimized PPI++ method mitigated.

  8. Knowl 8 — LLM preference-based ranking

    empirical result

    The pairwise-comparison experiment used 16,000 Chatbot Arena observations involving 20 language models. Each observation contained a human prompt, two model responses, and a human preference; GPT-4o-mini judged the same response pairs to provide synthetic preferences. Across random labeled/unlabeled splits, PPI++ estimated Bradley–Terry coefficients with lower reported MSE than the classical estimator and produced calibrated, tighter confidence intervals. Its effective sample size was 20%–25% higher than the classical approach, and its estimated model rankings had higher correlation with the ground-truth rankings. Ranking used Bonferroni-corrected 90% confidence intervals, assigning tied ranks when intervals overlapped.

  9. Knowl 9 — Effect of annotator quality on efficiency

    empirical result

    Across the protein-fitness and LLM-preference experiments, PPI++ effective sample size increased as the synthetic annotator's predictions became more correlated with the ground-truth labels. In the protein experiment, even the relatively weak CARP annotator yielded about a 10% effective-sample-size increase over the classical estimator; in the LLM experiment, the tested judges yielded gains of 20%–35%, depending on judge quality. When synthetic predictions have no correlation with the true labels, variance optimization can set λ=0\lambda=0, reducing PPI++ to the classical estimator rather than relying on uninformative synthetic data.

  10. Knowl 10 — Representativeness limits the standard guarantees

    limitation

    The standard AutoEval estimators and their confidence intervals rely on the human-labeled and unlabeled examples representing the same target population. If labeled inputs are not sampled representatively from the unlabeled pool, or the two sets come from different populations, the standard metric and Bradley–Terry estimators need not retain their statistical guarantees, and their confidence intervals may be invalid. The paper's importance-weighted extension addresses the specified covariate-shift setting when the density ratio between labeled and target input distributions is known; the standard unweighted procedure does not by itself handle such shifts.

Coverage note — The supplementary runtime comparison and code snippets are omitted because they provide implementation and execution details rather than additional core methodology or scientific findings.

References

  1. 1.Chatbot Arena: Benchmarking LLMs in the wild. https://lmarena.ai/. Accessed: 2025-01-21.
  2. 2.Agarwal, M., Kalia, R., Bahel, V., and Thomas, A. Autoeval: An NLP approach for automatic test evaluation system. In 2021 IEEE 4th International Conference on Computing, Power and Communication Technologies (GUCON), pp. 1–6. IEEE, 2021.
  3. 3.Al Mohamad, D., van Zwet, E., Solari, A., and Goeman, J. Simultaneous confidence intervals for ranks using the partitioning principle. Electronic Journal of Statistics, 15:2608–2646, 2021.
  4. 4.Angelopoulos, A. N., Bates, S., Fannjiang, C., Jordan, M. I., and Zrnic, T. Prediction-powered inference. Science, 382:669–674, 2023a.
  5. 5.Angelopoulos, A. N., Duchi, J. C., and Zrnic, T. Ppi++: Efficient prediction-powered inference. arXiv preprint arXiv:2311.01453, 2023b.
  6. 6.Arazo, E., Ortego, D., Albert, P., O’Connor, N. E., and McGuinness, K. Pseudo-labeling and confirmation bias in deep semi-supervised learning. In 2020 International joint conference on neural networks (IJCNN), pp. 1–8. IEEE, 2020.
  7. 7.Bachman, P., Alsharif, O., and Precup, D. Learning with pseudo-ensembles. Advances in neural information processing systems, 27, 2014.
  8. 8.Bradley, R. and Terry, M. E. Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.
  9. 9.Chaganty, A. T., Mussman, S., and Liang, P. The price of debiasing automatic metrics in natural language evaluation. arXiv preprint arXiv:1807.02202, 2018.
  10. 10.Chatzi, I., Straitouri, E., Thejaswi, S., and Rodriguez, M. G. Prediction-powered ranking of large language models. arXiv preprint arXiv:2402.17826, 2024.
  11. 11.Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhang, H., Zhu, B., Jordan, M., Gonzalez, J. E., and Stoica, I. Chatbot arena: An open platform for evaluating LLMs by human preference. arXiv preprint arXiv:2403.04132, 2024.
  12. 12.Corston-Oliver, S., Gamon, M., and Brockett, C. A machine learning approach to the automatic evaluation of machine translation. In Proceedings of the 39th Annual Meeting of the Association for Computational Linguistics, pp. 148–155, 2001.
  13. 13.Ford Jr, L. R. Solution of a ranking problem from binary comparisons. The American Mathematical Monthly, 64 (8P2):28–33, 1957.
  14. 14.Garg, S., Balakrishnan, S., Lipton, Z. C., Neyshabur, B., and Sedghi, H. Leveraging unlabeled data to predict out-of-distribution performance. arXiv preprint arXiv:2201.04234, 2022.
  15. 15.Hastie, T., Tibshirani, R., Friedman, J. H., and Friedman, J. H. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer, 2009.
  16. 16.Huang, H., Qu, Y., Bu, X., Zhou, H., Liu, J., Yang, M., Xu, B., and Zhao, T. An empirical study of llm-as-a-judge for llm evaluation: Fine-tuned judge model is not a general substitute for gpt-4. arXiv preprint arXiv:2403.02839, 2024.
  17. 17.Hunter, D. R. MM algorithms for generalized Bradley-Terry models. The annals of statistics, 32(1):384–406, 2004.
  18. 18.Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, 2023.
  19. 19.Laine, E., Karami, Y., and Carbone, A. GEMME: A simple and fast global epistatic model predicting mutational effects. Molecular Biology and Evolution, 36(11):2604–2619, August 2019.
  20. 20.Laine, S. and Aila, T. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242, 2016.
  21. 21.Lee, D.-H. et al. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on challenges in representation learning, ICML, volume 3, pp. 896. Atlanta, 2013.
  22. 22.Li, H., Dong, Q., Chen, J., Su, H., Zhou, Y., Ai, Q., Ye, Z., and Liu, Y. Llms-as-judges: A comprehensive survey on llm-based evaluation methods. arXiv preprint arXiv:2412.05579, 2024a.
  23. 23.Li, T., Chiang, W.-L., Frick, E., Dunlap, L., Wu, T., Zhu, B., Gonzalez, J. E., and Stoica, I. From crowdsourced data to high-quality benchmarks: Arena-hard and bench-builder pipeline. arXiv preprint arXiv:2406.11939, 2024b.
  24. 24.Lin, C.-Y. and Och, F. Looking for a few good metrics: Rouge and its evaluation. In Ntcir workshop, 2004.
  25. 25.Marquet, C., Heinzinger, M., Olenyi, T., Dallago, C., Erckert, K., Bernhofer, M., Nechaev, D., and Rost, B. Embeddings from protein language models predict conservation and variant effects. Human Genetics, 141(10):1629–1647, October 2022.
  26. 26.Meier, J., Rao, R., Verkuil, R., Liu, J., Sercu, T., and Rives, A. Language models enable zero-shot prediction of the effects of mutations on protein function. bioRxiv, pp. 2021.07.09.450648, November 2021.
  27. 27.Notin, P., Kollasch, A. W., Ritter, D., van Niekerk, L., Paul, S., Spinner, H., Rollins, N., Shaw, A., Weitzman, R., Frazer, J., Dias, M., Franceschi, D., Orenbuch, R., Gal, Y., and Marks, D. S. ProteinGym: Large-Scale benchmarks for protein design and fitness prediction. bioRxiv, pp. 2023.12.07.570727, December 2023.
  28. 28.Olson, C. A., Wu, N. C., and Sun, R. A comprehensive biophysical description of pairwise epistasis throughout an entire protein domain. Current Biology, 24(22):2643–2651, November 2014.
  29. 29.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318, 2002.
  30. 30.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. Pytorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems, 32, 2019.
  31. 31.Riesselman, A. J., Ingraham, J. B., and Marks, D. S. Deep generative models of genetic variation capture the effects of mutations. Nature Methods, 15(10):816–822, September 2018.
  32. 32.Robins, J. M. and Rotnitzky, A. Semiparametric efficiency in multivariate regression models with missing data. Journal of the American Statistical Association, 90(429):122–129, 1995.
  33. 33.Saad-Falcon, J., Khattab, O., Potts, C., and Zaharia, M. Ares: An automated evaluation framework for retrieval-augmented generation systems. arXiv preprint arXiv:2311.09476, 2023.
  34. 34.Tsiatis, A. A. Semiparametric Theory and Missing Data. Springer, 2006.
  35. 35.van Breugel, B., Seedat, N., Imrie, F., and van der Schaar, M. Can you rely on your model evaluation? improving model evaluation with synthetic test data. Advances in Neural Information Processing Systems, 36:1889–1904, 2023.
  36. 36.Wei, J. T.-Z. and Jia, R. The statistical advantage of automatic NLG metrics at the system level. arXiv preprint arXiv:2105.12437, 2021.
  37. 37.Zermelo, E. Die berechnung der turnier-ergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung. Mathematische Zeitschrift, 29(1):436–460, 1929.
  38. 38.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging LLM-as-a-Judge with MT-Bench and chatbot arena. Advances in Neural Information Processing Systems, 36, 2024.

Citation

MLA
Boyeau, P., et al. “AutoEval Done Right: Using Synthetic Data for Model Evaluation”. arXiv, 2024, http://arxiv.org/abs/2403.07008v3.
APA
Boyeau, P., Angelopoulos, A. N., Yosef, N., Malik, J., & Jordan, M. I. (2024). AutoEval Done Right: Using Synthetic Data for Model Evaluation. arXiv. http://arxiv.org/abs/2403.07008v3
Chicago
Boyeau, P., A. N. Angelopoulos, N. Yosef, J. Malik, and M. I. Jordan. 2024. “AutoEval Done Right: Using Synthetic Data for Model Evaluation”. arXiv. http://arxiv.org/abs/2403.07008v3.
Harvard
Boyeau, P. et al. (2024) “AutoEval Done Right: Using Synthetic Data for Model Evaluation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.07008v3.
Vancouver
1. Boyeau P, Angelopoulos AN, Yosef N, Malik J, Jordan MI (2024) AutoEval Done Right: Using Synthetic Data for Model Evaluation. arXiv

BibTeX

@article{boyeau2024autoeval,
  title = {AutoEval Done Right: Using Synthetic Data for Model Evaluation},
  author = {Boyeau, Pierre and Angelopoulos, Anastasios N. and Yosef, Nir and Malik, Jitendra and Jordan, Michael I.},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.07008v3},
  eprint = {2403.07008}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/