Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Seungone KimJuyoung SukShayne LongpreBill Yuchen LinJamin ShinSean WelleckGraham NeubigMoontae LeeKyungjae LeeMinjoon Seo

article2024EMNLP531 citations

Presents Prometheus 2, an open-source evaluator language model that handles both direct assessment and pairwise ranking with custom criteria by merging models trained on separate evaluation formats, closely mirroring human and GPT-4 judgments.

Listen

Assessing the quality of text generated by artificial intelligence increasingly relies on using large language models as automated judges. While commercial systems like GPT-4 perform this task effectively and closely align with human judgment, their proprietary nature introduces substantial barriers around cost, lack of transparency, and limited user control. Existing open-source alternatives address these transparency concerns but historically suffer from significant drawbacks: their ratings diverge widely from human evaluations, and they lack the flexibility to handle both direct scoring (assigning an absolute numerical score) and pairwise comparison (choosing the better of two outputs) across custom evaluation criteria.

The article aims to resolve these limitations by introducing Prometheus 2, an open-source evaluation model designed to bridge the performance gap with proprietary models while supporting both direct assessment and pairwise ranking across custom evaluation standards.

To build this system, the researchers created a new dataset comprising 200,000 pairwise comparison instances featuring 1,000 custom evaluation rubrics and detailed written explanations generated by GPT-4. They then trained specialized models based on open base architectures for both direct assessment and pairwise ranking. Rather than training a single model simultaneously on both tasks—which often leads to negative interference between objectives—the team applied a weight-merging technique called DARE-Linear. This method mathematically combines the parameters of two separately trained models into a single unified evaluator. The resulting systems were tested across four direct assessment benchmarks and four pairwise ranking benchmarks, comparing their outputs to evaluations from humans and leading commercial language models.

The findings show that Prometheus 2 significantly outperforms all existing open-source evaluator models and cuts the performance gap with proprietary models in half. On direct scoring benchmarks, the model achieved a correlation with human and commercial judges that surpassed baseline open models by more than 0.2 correlation units, consistently maintaining strong alignment above 0.5 across diverse datasets. In pairwise comparisons, the larger variant achieved top accuracy among open models, reaching an 85.5% agreement rate on standard alignment benchmarks. Furthermore, the analysis confirmed that merging weights from models trained on distinct formats creates positive cross-task transfer, whereas merging models trained on the same format or using traditional joint training yielded inferior results.

These results demonstrate that organizations do not need to rely exclusively on costly, closed proprietary models to conduct reliable, granular evaluations of language model outputs. By deploying Prometheus 2, enterprises and researchers can establish transparent, reproducible, and cost-effective internal evaluation pipelines tailored to specific domain guidelines rather than generic helpfulness standards.

For practitioners seeking to deploy automated evaluation pipelines, the article supports adopting weight-merged open models alongside reference answers, as providing reference answers significantly improves scoring correlation. Decision-makers can choose between smaller (7B parameter) and larger (Mixtral 8x7B) variants depending on their computing budgets and latency constraints. Future efforts should focus on extending this unified evaluation approach to additional scoring formats, such as ten-point Likert scales, multi-response rankings, and checklist assessments.

The primary limitations of this study involve its indirect validation approach, which relies on matching proxy human judgments and proprietary model outputs rather than end-to-end task auditing. The model is also restricted to English text, five-point Likert direct scales, and binary pairwise rankings. Nevertheless, given the consistent outperformance across eight diverse evaluation benchmarks and robust consistency metrics, there is high confidence in Prometheus 2 as a state-of-the-art open evaluation solution within its designated operational scope.

Cover for Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Abstract

Proprietary LMs such as GPT-4 are often employed to assess the quality of responses from various LMs. However, concerns including transparency, controllability, and affordability strongly motivate the development of open-source LMs specialized in evaluations. On the other hand, existing open evaluator LMs exhibit critical shortcomings: 1) they issue scores that significantly diverge from those assigned by humans, and 2) they lack the flexibility to perform both direct assessment and pairwise ranking, the two most prevalent forms of assessment. Additionally, they often do not possess the ability to evaluate based on custom evaluation criteria, focusing instead on general attributes like helpfulness and harmlessness. To address these issues, we introduce Prometheus 2. Prometheus 2 is more powerful than its predecessor, and closely mirrors human and GPT-4 judgements. Moreover, it is capable of processing both direct assessment and pair-wise ranking formats grouped with a user-defined evaluation criteria. On four direct assessment benchmarks and four pairwise ranking benchmarks, PROMETHEUS 2 scores the highest correlation and agreement with humans and proprietary LM judges among all tested open evaluator LMs. Our models, code, and data are all publicly available. 1

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Language Model-based Evaluation
  • 2.2 Weight Merging
  • 3 Methodology
  • 3.1 Direct Assessment
  • 3.2 Pairwise Ranking
  • 3.3 The Preference Collection
  • 3.4 Training Methods & Baselines
  • 4 Experimental Setup
  • 5 Experimental Results
  • 5.1 Direct Assessment Results
  • 5.2 Pairwise Ranking Results
  • 6 Analyses of Weight Merging
  • 6.1 Weight Merging vs Joint Training
  • 6.2 Is the Effectiveness of Weight Merging due to Model Ensembling?
  • 6.3 Quantifying Positive Transfer across Evaluation Formats
  • 7 Conclusion
  • Acknowledgements
  • Limitations
  • References
  • A Quality Verification of the PREFERENCE COLLECTION
  • B Training and Inference Details
  • C Direct Assessment Results: Extended
  • D License
  • E Consistency of Evaluator LMs
  • F Reference-free Evaluation in Direct Assessment Formats
  • G Merging Method Ablation
  • H PREFERENCE COLLECTION Augmentation Prompt
  • I Direct Assessment Prompt
  • J Pairwise Ranking Prompt

Knowls

  1. Knowl 1 — Weight Merging Training Recipe for Prometheus 2 Evaluator Models

    model/method

    PROMETHEUS 2 is an open-source evaluator language model developed to perform both criteria-guided direct assessment (1–5 Likert scoring) and pairwise ranking. To avoid the negative task interference that occurs during joint multi-task fine-tuning, PROMETHEUS 2 is trained using a weight merging paradigm.

    Given a base pretrained model with parameters θinit\theta_{\text{init}} (specifically Mistral-7B-Instruct-v0.2 or Mixtral-8x7B-Instruct-v0.1), two separate models are fine-tuned via supervised fine-tuning (SFT):

    1. A direct assessment evaluator θd\theta_d trained on the FEEDBACK COLLECTION (DdD_d, comprising 100,000 instances).
    2. A pairwise ranking evaluator θp\theta_p trained on the PREFERENCE COLLECTION (DpD_p, comprising 200,000 instances).

    The final model weights θfinal\theta_{\text{final}} are obtained by merging θd\theta_d and θp\theta_p using Drop And RE-scale (DARE)-Linear merging. DARE randomly drops delta parameters (θd−θinit)(\theta_d - \theta_{\text{init}}) and (θp−θinit)(\theta_p - \theta_{\text{init}}) with drop probability p=0.1p = 0.1 and rescales the remaining weights by a factor λ=1.95\lambda = 1.95 before performing linear parameter combination. For the Mixtral-8x7B model, parameter-efficient fine-tuning (PEFT) is used with LoRA rank r=256r = 256, α=512\alpha = 512, dropout 0.10.1, targeting projection modules Q,K,V,O,WQ, K, V, O, W, and the LM head.

  2. Knowl 2 — Formulation of Criteria-Guided Direct Assessment and Pairwise Ranking

    model/method

    PROMETHEUS 2 unifies two evaluation schemes, both conditioned on explicit user-defined criteria and generating natural language verbal feedback before emitting final scoring decisions:

    1. Direct Assessment: Maps an instruction ii, candidate response rr, reference answer aa, and score rubric ee to a verbal feedback explanation vrv_r and an integer scalar score s∈{1,2,3,4,5}s \in \{1, 2, 3, 4, 5\}: fdirect:(i,r,a,e)↦(vr,s)f_{\text{direct}} : (i, r, a, e) \mapsto (v_r, s) The score rubric ee explicitly specifies the criterion description alongside granular textual descriptors defining each score level from 1 to 5.

    2. Pairwise Ranking: Maps an instruction ii, candidate response pair (rm,rn)(r_m, r_n), reference answer aa (which can be omitted during reference-free evaluation), and evaluation criterion ee to a comparative verbal feedback vrm,rnv_{r_m, r_n} and a preference decision s∈{m,n}s \in \{m, n\}: fpair:(i,rm,rn,a,e)↦(vrm,rn,s)f_{\text{pair}} : (i, r_m, r_n, a, e) \mapsto (v_{r_m, r_n}, s) In pairwise ranking, ee describes the evaluation criterion itself without score-level gradations, and vrm,rnv_{r_m, r_n} explicitly contrasts the commonalities and differences between rmr_m and rnr_n relative to ee before stating which response is superior.

  3. Knowl 3 — The Preference Collection Dataset

    definition

    The PREFERENCE COLLECTION is a fine-grained pairwise ranking feedback dataset constructed to train evaluator language models under custom criteria beyond generic helpfulness and harmlessness. It contains 200,000 instances derived from the FEEDBACK COLLECTION.

    Key structural properties include:

    • Instructions and Criteria: Contains 20,000 unique instructions, 20,000 reference answers, and 1,000 custom instance-wise evaluation criteria.
    • Pair Construction: Each instruction in the FEEDBACK COLLECTION originally possessed five candidate responses scoring from 1 to 5. All (52)=10\binom{5}{2} = 10 response pairings per instruction were formed, generating 20,000×10=200,00020,000 \times 10 = 200,000 training instances.
    • Balanced Targets: The ground-truth preferred response is determined by comparing original score assignments, resulting in an exact 50/50 balance between labels (100,000 instances where Response A is preferred and 100,000 instances where Response B is preferred).
    • Comparative Verbal Feedback: For each pair, GPT-4-1106 generated structured verbal feedback vrm,rnv_{r_m, r_n} explaining the commonalities, differences, and criterion-specific justifications for the preference.
    • Quality Verification: In human validation over 200 sampled instances, the generated feedback achieved 99.5% decision coherence, 98.5% criterion suitability, and an 88% win rate in analytical criticality over concatenating independent direct assessment feedbacks.
  4. Knowl 4 — Direct Assessment Benchmark Performance of Prometheus 2

    empirical result

    PROMETHEUS 2 (7B and 8x7B) achieves state-of-the-art correlation with human evaluators and proprietary LLM judges on direct assessment benchmarks (Vicuna Bench, MT Bench, FLASK, and Feedback Bench), outperforming all open-source evaluator baselines.

    Evaluator LM Vicuna Bench MT Bench FLASK Feedback Bench
    GPT-4 Claude-3 GPT-4 Claude-3 GPT-4 Claude-3 Humans GPT-4-0613
    Llama2-Chat-7B 0.205 0.243 0.036 0.055 0.317 0.256 0.299 0.523
    Llama2-Chat-70B 0.350 0.463 0.178 0.228 0.388 0.402 0.317 0.592
    Mistral-Instruct-7B 0.486 0.561 0.284 0.396 0.448 0.437 0.377 0.586
    Mixtral-Instruct-8x7B 0.566 0.579 0.551 0.539 0.483 0.495 0.420 0.673
    Prometheus-13B 0.492 0.534 0.404 0.477 0.462 0.470 0.449 0.860
    Auto-J (13B) 0.351 0.262 0.432 0.375 0.430 0.370 0.473 0.637
    Prometheus-2-7B 0.666 0.654 0.548 0.517 0.617 0.561 0.545 0.882
    Prometheus-2-8x7B 0.685 0.635 0.665 0.614 0.659 0.626 0.555 0.898
    GPT-3.5-Turbo-0613 0.335 0.349 0.183 0.194 0.437 0.396 0.450 0.594
    GPT-4-1106 – 0.694 – 0.717 – 0.736 0.679 0.753
    Claude-3-Opus 0.694 – 0.717 – 0.736 – 0.573 0.788

    PROMETHEUS 2 scores exceed previous open evaluator baselines by over 0.2 Pearson correlation units across datasets. On FLASK human correlation, PROMETHEUS-2-8x7B reaches 0.555 (compared to 0.449 for Prometheus-13B), halving the performance gap to GPT-4 (0.679).

  5. Knowl 5 — Pairwise Ranking Benchmark Performance of Prometheus 2

    empirical result

    PROMETHEUS 2 (7B and 8x7B) achieves the highest agreement accuracy with human judgments on four pairwise ranking benchmarks: HHH Alignment (221 test instances), MT Bench Human Judgment (3,360 pairs), Auto-J Eval (1,392 pairs), and Preference Bench (2,000 pairs).

    Evaluator LM HHH Total Avg. MT Bench Human Auto-J Eval Preference Bench
    w/ Tie w/o Tie w/ Tie w/o Tie Instance Criteria
    Llama2-Chat-70B 68.78% 55.14% 60.88% 53.38% 50.64% 64.70%
    Mistral-Instruct-7B 67.42% 53.81% 63.82% 53.88% 60.94% 79.40%
    Mixtral-Instruct-8x7B 77.38% 51.85% 71.42% 53.81% 73.50% 84.00%
    Pair RM (0.4B) 84.62% – 59.00% – 59.05% 81.80%
    Ultra RM (13B) 83.71% – 56.00% – 59.85% 86.97%
    Auto-J (13B) 75.57% 42.56% 69.12% 43.46% 76.64% 81.35%
    Prometheus-2-7B 74.66% 50.45% 70.78% 54.96% 75.07% 93.25%
    Prometheus-2-8x7B 85.52% 55.07% 71.96% 58.41% 79.98% 90.65%
    GPT-3.5-Turbo-0613 76.47% 54.65% 69.41% 45.98% 72.13% 75.05%
    GPT-4-1106-Preview 90.95% 60.38% 79.90% 52.80% 83.12% 85.50%
    Claude-3-Opus 94.57% 55.35% 77.65% 60.70% 82.92% 89.85%

    PROMETHEUS-2-8x7B outperforms domain-specific reward models such as PairRM (84.62% on HHH) and Auto-J on its own in-domain benchmark (79.98% vs. 76.64% on Auto-J Eval w/o tie), halving the gap between open models and proprietary judges such as GPT-4.

  6. Knowl 6 — Weight Merging versus Joint Training for Multi-Format Evaluator Models

    empirical result

    Joint multi-task training on direct assessment (DdD_d) and pairwise ranking (DpD_p) datasets leads to negative task interference, whereas weight merging yields positive transfer across both formats.

    Training Method Direct Assessment (Pearson) Pairwise Ranking (Accuracy %)
    Vicuna MT Bench FLASK Avg. HHH MT H.J. Auto-J Avg.
    Base: Mistral-Instruct-7B
    Prompting (Zero-shot) 0.486 0.284 0.480 0.417 67.42 63.82 60.94 64.06
    Direct Assessment Only 0.537 0.561 0.519 0.539 73.33 56.76 64.38 64.82
    Pairwise Ranking Only – – – – 78.73 67.06 72.03 72.61
    Joint Training 0.548 0.450 0.457 0.485 80.09 65.49 73.60 73.06
    Weight Merging 0.666 0.548 0.659 0.624 74.66 70.78 75.07 73.50
    Base: Mixtral-Instruct-8x7B
    Prompting (Zero-shot) 0.566 0.551 0.507 0.541 77.38 71.42 73.55 74.56
    Direct Assessment Only 0.625 0.664 0.587 0.625 74.21 53.14 65.85 64.40
    Pairwise Ranking Only – – – – 84.16 66.27 75.66 75.36
    Joint Training 0.628 0.560 0.596 0.595 82.35 68.73 74.78 75.29
    Weight Merging 0.685 0.665 0.659 0.670 85.52 71.96 79.98 79.15

    Joint training reduces average direct assessment correlation from 0.539 to 0.485 on Mistral-7B and from 0.625 to 0.595 on Mixtral-8x7B relative to single-format direct training. In contrast, weight merging elevates average direct assessment correlation to 0.624 (7B) and 0.670 (8x7B), while simultaneously boosting average pairwise accuracy above both single-task and joint baselines.

  7. Knowl 7 — Cross-Format Knowledge Integration versus Same-Format Model Ensembling

    empirical result

    To establish whether the gains of weight merging stem from generic multi-seed model ensembling or from cross-format task integration, an ablation was conducted using Mistral-7B-Instruct by merging models trained on the same format versus models trained on different formats:

    Merged Formats Direct Assessment (Pearson) Pairwise Ranking (Accuracy %)
    Vicuna MT Bench FLASK Avg. HHH MT H.J. Auto-J Avg.
    Direct Only (1 seed) 0.537 0.561 0.519 0.539 73.33 56.76 64.38 64.82
    Pairwise Only (1 seed) – – – – 78.73 67.06 72.03 72.61
    Direct Direct (2 seeds) 0.552 0.493 0.505 0.517 73.30 55.00 63.69 64.13
    Pairwise Pairwise (2 seeds) – – – – 78.70 65.20 72.72 72.21
    Direct Pairwise 0.666 0.548 0.659 0.624 74.66 70.78 75.07 73.50

    Merging two distinct random seed models trained on the exact same format degrades performance (direct assessment average drops from 0.539 to 0.517; pairwise average drops from 72.61% to 72.21%). The performance gain occurs exclusively when combining models across heterogeneous evaluation formats (direct assessment + pairwise ranking), demonstrating functional cross-format knowledge transfer rather than variance reduction from ensembling.

  8. Knowl 8 — Merging Ratio Dynamics and Directional Task Transfer

    empirical result

    When linearly merging direct assessment weights θd\theta_d and pairwise ranking weights θp\theta_p via θfinal=αθd+(1−α)θp\theta_{\text{final}} = \alpha \theta_d + (1-\alpha) \theta_p for α∈[0.1,0.9]\alpha \in [0.1, 0.9], task transfer exhibits asymmetry:

    1. Direct Assessment Optimal Ratio: Direct assessment correlation is maximized at an exact balance α=0.5\alpha = 0.5 (a 5:5 ratio between direct assessment and pairwise ranking weights).
    2. Pairwise Ranking Optimal Ratio: Pairwise ranking accuracy is maximized when α=0.3\alpha = 0.3 (a 3:7 ratio, allocating 70% of the weight to the pairwise model).
    3. Asymmetry of Transfer: Incorporating relative pairwise ranking weights into an absolute direct assessment model provides a substantially larger boost to direct scoring correlation than incorporating direct assessment weights provides to pairwise ranking accuracy.
  9. Knowl 9 — Comparative Performance of Weight Merging Techniques for Evaluator Models

    empirical result

    An empirical comparison of model merging algorithms on Mistral-7B shows distinct behaviors across direct assessment and pairwise ranking tasks:

    Merging Method Direct Avg. (Pearson) Pairwise Avg. (%) Combined Overall Score
    Linear (α=0.5\alpha = 0.5) 0.652 78.06 82.93
    Slerp 0.649 77.44 82.67
    Task Arithmetic 0.582 78.93 81.01
    TIES 0.614 78.56 80.58
    DARE-TIES 0.655 78.55 83.27
    DARE-Linear 0.660 78.44 83.32

    Task Arithmetic merging frequently produces generation format corruptions (e.g., generating integers instead of selection identifiers during pairwise ranking). DARE-Linear merging achieves the highest direct assessment Pearson correlation (0.660) and highest combined score (83.32), leading to its selection as the default merging algorithm for PROMETHEUS 2.

  10. Knowl 10 — Cross-Format Consistency and Decision Transitivity of Evaluator Models

    empirical result

    Evaluator model robustness is assessed through cross-format prediction stability and transitivity in pairwise preferences:

    1. Cross-Format Evaluation Gap (∥Direct2Pair−Pair2Pair∥\|\text{Direct2Pair} - \text{Pair2Pair}\|): Pairwise preferences can be evaluated either directly (fpairf_{\text{pair}}) or by scoring each candidate independently via direct assessment (fdirectf_{\text{direct}}) and comparing scores. The accuracy discrepancy Δ\Delta measures format consistency:

      • Auto-J (13B) exhibits severe discrepancy: Δ=28.96\Delta = 28.96 on HHH, 20.9820.98 on MT Bench Human, and 29.2429.24 on Auto-J Eval.
      • GPT-4-1106-Preview exhibits Δ=7.24\Delta = 7.24 on HHH, 11.8611.86 on MT Bench, and 28.8528.85 on Auto-J Eval.
      • PROMETHEUS-2-7B maintains high consistency: Δ=0.45\Delta = 0.45 on HHH, 7.547.54 on MT Bench, and 6.966.96 on Auto-J Eval.
      • PROMETHEUS-2-8x7B exhibits Δ=4.07\Delta = 4.07 on HHH, 10.2910.29 on MT Bench, and 13.4413.44 on Auto-J Eval.
    2. Pairwise Transitivity: Transitivity measures whether (rB≻rA)∧(rC≻rB)  ⟹  (rC≻rA)(r_B \succ r_A) \wedge (r_C \succ r_B) \implies (r_C \succ r_A) across candidate triplets on the PREFERENCE COLLECTION:

      • PROMETHEUS-2-7B achieves 97.60% transitivity.
      • PROMETHEUS-2-8x7B achieves 96.75% transitivity.
      • Proprietary models achieve 95.70% (GPT-4-1106) and 96.20% (Claude-3-Opus), while base unmerged open models achieve 87.10% (Mistral-7B) and 89.65% (Auto-J 13B).
  11. Knowl 11 — Impact of Reference Answer Conditioning on Direct Assessment Accuracy

    empirical result

    Conditioning evaluator language models on a high-quality reference answer significantly increases scoring correlation with human judgments across direct assessment benchmarks.

    Evaluator LM BiGGen Bench FLASK
    Ref-Free Ref-Based Δ\Delta Ref-Free Ref-Based Δ\Delta
    Mistral-Instruct-7B 0.305 0.310 +0.005 0.331 0.374 +0.043
    Mixtral-Instruct-8x7B 0.320 0.322 +0.002 0.377 0.386 +0.009
    Prometheus-2-7B 0.403 0.455 +0.052 0.425 0.545 +0.120
    Prometheus-2-8x7B 0.424 0.472 +0.048 0.411 0.555 +0.144
    GPT-3.5-Turbo-0613 0.236 0.252 +0.016 0.354 0.374 +0.020
    GPT-4-1106 0.554 0.599 +0.045 0.616 0.679 +0.063

    Removing the reference answer degrades human correlation across all tested models (e.g., dropping FLASK correlation by 0.144 for PROMETHEUS-2-8x7B and by 0.063 for GPT-4). PROMETHEUS 2 models retain higher correlation in reference-free settings (0.424 on BiGGen Bench) than unaligned base models in reference-based settings (0.322 on Mixtral-8x7B).

  12. Knowl 12 — Operational and Methodological Limitations of Prometheus 2

    limitation

    The PROMETHEUS 2 evaluator models are subject to several operational constraints:

    1. Fixed Evaluation Formats: Models are specialized exclusively for 1–5 integer Likert scoring in direct assessment and binary choice ('Response A is better' vs. 'Response B is better') in pairwise ranking. They lack support for 1–10 continuous scales, multi-candidate listwise ranking (N>2N > 2), or rubric-free checklist evaluation.
    2. Loss of Arbitrary Format Generalization: Unlike frontier proprietary models that adapt to varied zero-shot prompt instructions, fine-tuning open-source models on structured templates restricts their evaluation flexibility when presented with out-of-format prompts.
    3. Evaluation Capability Proxies: Evaluation efficacy is benchmarked via proxy correlation with human judges and proprietary models (GPT-4, Claude-3-Opus), which may inherit biases from the underlying reference evaluators.
    4. Lack of Theoretical Foundation for Merging Efficacy: The elimination of negative task transfer via weight merging is established empirically, but a formal mathematical characterization of why parameter merging circumvents multi-task optimization failure in LLMs remains an open question.

Coverage note — None. All core contributions—including model architecture, merging recipes, dataset construction, experimental benchmarks, consistency ablations, reference grounding analyses, and stated limitations—are fully covered.

References

  1. 1.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan. 2021. A general language assistant as a laboratory for alignment.
  2. 2.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  3. 3.Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. Chateval: Towards better llm-based evaluators through multi-agent debate. arXiv preprint arXiv:2308.07201.
  4. 4.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An opensource chatbot impressing gpt-4 with 90%* chatgpt quality.
  5. 5.Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023. Ultrafeedback: Boosting language models with high-quality feedback. arXiv preprint arXiv:2310.01377.
  6. 6.Shachar Don-Yehiya, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen. 2022. Cold fusion: Collaborative descent for distributed multitask finetuning. arXiv preprint arXiv:2212.01378.
  7. 7.Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Alpacafarm: A simulation framework for methods that learn from human feedback. arXiv preprint arXiv:2305.14387.
  8. 8.Markus Freitag, David Grangier, and Isaac Caswell. 2020. Bleu might be guilty but references are not innocent. arXiv preprint arXiv:2004.06063.
  9. 9.Mingqi Gao, Xinyu Hu, Jie Ruan, Xiao Pu, and Xiaojun Wan. 2024. Llm-based nlg evaluation: Current status and challenges. arXiv preprint arXiv:2402.01383.
  10. 10.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh D Dhole, et al. 2021. The gem benchmark: Natural language generation, its evaluation and metrics. arXiv preprint arXiv:2102.01672.
  11. 11.Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran, Alex Wang, Alexandros Papangelis, Aman Madaan, Angelina McMillan-Major, Anna Shvets, Ashish Upadhyay, Bingsheng Yao, et al. 2022. Gemv2: Multilingual nlg benchmarking in a single line of code. arXiv preprint arXiv:2206.11249.
  12. 12.Charles Goddard, Shamane Siriwardhana, Malikeh Ehghaghi, Luke Meyers, Vlad Karpukhin, Brian Benedict, Mark McQuade, and Jacob Solawetz. 2024. Arcee’s mergekit: A toolkit for merging large language models. arXiv preprint arXiv:2403.13257.
  13. 13.Suchin Gururangan, Margaret Li, Mike Lewis, Weijia Shi, Tim Althoff, Noah A Smith, and Luke Zettlemoyer. 2023. Scaling expert language models with unsupervised domain discovery. arXiv preprint arXiv:2303.14177.
  14. 14.Michael Hanna and Ondřej Bojar. 2021. A fine-grained analysis of bertscore. In Proceedings of the Sixth Conference on Machine Translation, pages 507–517.
  15. 15.Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2022. Editing models with task arithmetic. arXiv preprint arXiv:2212.04089.
  16. 16.Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, and Prithviraj Ammanabrolu. 2023a. Personalized soups: Personalized large language model alignment via post-hoc parameter merging. arXiv preprint arXiv:2310.11564.
  17. 17.Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2023b. Exploring the benefits of training expert language models over instruction tuning. arXiv preprint arXiv:2302.03202.
  18. 18.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023a. Mistral 7b. arXiv preprint arXiv:2310.06825.
  19. 19.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  20. 20.Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen. 2023b. Tigerscore: Towards building explainable metric for all text generation tasks. arXiv preprint arXiv:2310.00752.
  21. 21.Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. 2023c. Llm-blender: Ensembling large language models with pairwise ranking and generative fusion. arXiv preprint arXiv:2306.02561.
  22. 22.Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. 2023. Prometheus: Inducing fine-grained evaluation capability in language models. arXiv preprint arXiv:2310.08491.
  23. 23.Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, et al. 2024. The biggen bench: A principled benchmark for fine-grained evaluation of language models with language models. arXiv preprint arXiv:2406.05761.
  24. 24.Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, et al. 2024. Rewardbench: Evaluating reward models for language modeling. arXiv preprint arXiv:2403.13787.
  25. 25.Seongyun Lee, Seungone Kim, Sue Hyun Park, Geewook Kim, and Minjoon Seo. 2024. Prometheusvision: Vision-language model as a judge for fine-grained evaluation. arXiv preprint arXiv:2401.06591.
  26. 26.Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. 2023a. Generative judge for evaluating alignment. arXiv preprint arXiv:2310.05470.
  27. 27.Margaret Li, Suchin Gururangan, Tim Dettmers, Mike Lewis, Tim Althoff, Noah A Smith, and Luke Zettlemoyer. 2022. Branch-train-merge: Embarrassingly parallel training of expert language models. arXiv preprint arXiv:2208.03306.
  28. 28.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023b. Alpacaeval: An automatic evaluator of instructionfollowing models. https://github.com/tatsu-lab/alpaca_eval.
  29. 29.Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, and Chongyang Tao. 2024. Leveraging large language models for nlg evaluation: A survey. arXiv preprint arXiv:2401.07103.
  30. 30.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  31. 31.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023a. G-eval: Nlg evaluation using gpt-4 with better human alignment.
  32. 32.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b. Gpteval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  33. 33.Michael S Matena and Colin A Raffel. 2022. Merging models with fisher-weighted averaging. Advances in Neural Information Processing Systems, 35:17703–17716.
  34. 34.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  35. 35.Alexandre Rame, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya, Mustafa Shukor, Laure Soulier, and Matthieu Cord. 2024. Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards. Advances in Neural Information Processing Systems, 36.
  36. 36.Natalie Schluter. 2017. The limits of automatic summarisation according to rouge. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, pages 41–45. Association for Computational Linguistics.
  37. 37.Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma, Hu Xu, Xi Victoria Lin, Baptiste Rozière, Jacob Kahn, Daniel Li, Wen-tau Yih, Jason Weston, et al. 2024. Branch-train-mix: Mixing expert llms into a mixture-of-experts llm. arXiv preprint arXiv:2403.07816.
  38. 38.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and finetuned chat models.
  39. 39.Haoxiang Wang, Yong Lin, Wei Xiong, Rui Yang, Shizhe Diao, Shuang Qiu, Han Zhao, and Tong Zhang. 2024. Arithmetic control of llms for diverse user preferences: Directional preference alignment with multi-objective rewards. arXiv preprint arXiv:2402.18571.
  40. 40.Peiyi Wang, Lei Li, Zhihong Shao, RX Xu, Damai Dai, Yifei Li, Deli Chen, Y Wu, and Zhifang Sui. 2023a. Math-shepherd: A label-free step-by-step verifier for llms in mathematical reasoning. arXiv preprint arXiv:2312.08935.
  41. 41.Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, et al. 2023b. Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization. arXiv preprint arXiv:2306.05087.
  42. 42.Prateek Yadav, Derek Tam, Leshem Choshen, Colin A Raffel, and Mohit Bansal. 2024. Ties-merging: Resolving interference when merging models. Advances in Neural Information Processing Systems, 36.
  43. 43.Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim, and Minjoon Seo. 2023. Flask: Fine-grained language model evaluation based on alignment skill sets. arXiv preprint arXiv:2307.10928.
  44. 44.Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. 2023. Language models are super mario: Absorbing abilities from homologous models as a free lunch. arXiv preprint arXiv:2311.03099.
  45. 45.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
  46. 46.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685.
  47. 47.Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2023. Judgelm: Fine-tuned large language models are scalable judges. arXiv preprint arXiv:2310.17631.

Citation

MLA
Kim, S., et al. “Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 4334–53, https://doi.org/10.18653/v1/2024.emnlp-main.248.
APA
Kim, S., Suk, J., Longpre, S., Lin, B. Y., Shin, J., Welleck, S., Neubig, G., Lee, M., Lee, K., & Seo, M. (2024). Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 4334–4353. https://doi.org/10.18653/v1/2024.emnlp-main.248
Chicago
Kim, S., J. Suk, S. Longpre, et al. 2024. “Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 4334–53. https://doi.org/10.18653/v1/2024.emnlp-main.248.
Harvard
Kim, S. et al. (2024) “Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4334–4353. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.248.
Vancouver
1. Kim S, Suk J, Longpre S, Lin BY, Shin J, Welleck S, Neubig G, Lee M, Lee K, Seo M (2024) Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4334–4353

BibTeX

@inproceedings{kim-etal-2024-prometheus,
    title = "Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models",
    author = "Kim, Seungone  and
      Suk, Juyoung  and
      Longpre, Shayne  and
      Lin, Bill Yuchen  and
      Shin, Jamin  and
      Welleck, Sean  and
      Neubig, Graham  and
      Lee, Moontae  and
      Lee, Kyungjae  and
      Seo, Minjoon",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.248/",
    doi = "10.18653/v1/2024.emnlp-main.248",
    pages = "4334--4353"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/