Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Seungone KimJuyoung SukShayne LongpreBill Yuchen LinJamin ShinSean WelleckGraham NeubigMoontae LeeKyungjae LeeMinjoon Seo
Presents Prometheus 2, an open-source evaluator language model that handles both direct assessment and pairwise ranking with custom criteria by merging models trained on separate evaluation formats, closely mirroring human and GPT-4 judgments.
Assessing the quality of text generated by artificial intelligence increasingly relies on using large language models as automated judges. While commercial systems like GPT-4 perform this task effectively and closely align with human judgment, their proprietary nature introduces substantial barriers around cost, lack of transparency, and limited user control. Existing open-source alternatives address these transparency concerns but historically suffer from significant drawbacks: their ratings diverge widely from human evaluations, and they lack the flexibility to handle both direct scoring (assigning an absolute numerical score) and pairwise comparison (choosing the better of two outputs) across custom evaluation criteria.
The article aims to resolve these limitations by introducing Prometheus 2, an open-source evaluation model designed to bridge the performance gap with proprietary models while supporting both direct assessment and pairwise ranking across custom evaluation standards.
To build this system, the researchers created a new dataset comprising 200,000 pairwise comparison instances featuring 1,000 custom evaluation rubrics and detailed written explanations generated by GPT-4. They then trained specialized models based on open base architectures for both direct assessment and pairwise ranking. Rather than training a single model simultaneously on both tasks—which often leads to negative interference between objectives—the team applied a weight-merging technique called DARE-Linear. This method mathematically combines the parameters of two separately trained models into a single unified evaluator. The resulting systems were tested across four direct assessment benchmarks and four pairwise ranking benchmarks, comparing their outputs to evaluations from humans and leading commercial language models.
The findings show that Prometheus 2 significantly outperforms all existing open-source evaluator models and cuts the performance gap with proprietary models in half. On direct scoring benchmarks, the model achieved a correlation with human and commercial judges that surpassed baseline open models by more than 0.2 correlation units, consistently maintaining strong alignment above 0.5 across diverse datasets. In pairwise comparisons, the larger variant achieved top accuracy among open models, reaching an 85.5% agreement rate on standard alignment benchmarks. Furthermore, the analysis confirmed that merging weights from models trained on distinct formats creates positive cross-task transfer, whereas merging models trained on the same format or using traditional joint training yielded inferior results.
These results demonstrate that organizations do not need to rely exclusively on costly, closed proprietary models to conduct reliable, granular evaluations of language model outputs. By deploying Prometheus 2, enterprises and researchers can establish transparent, reproducible, and cost-effective internal evaluation pipelines tailored to specific domain guidelines rather than generic helpfulness standards.
For practitioners seeking to deploy automated evaluation pipelines, the article supports adopting weight-merged open models alongside reference answers, as providing reference answers significantly improves scoring correlation. Decision-makers can choose between smaller (7B parameter) and larger (Mixtral 8x7B) variants depending on their computing budgets and latency constraints. Future efforts should focus on extending this unified evaluation approach to additional scoring formats, such as ten-point Likert scales, multi-response rankings, and checklist assessments.
The primary limitations of this study involve its indirect validation approach, which relies on matching proxy human judgments and proprietary model outputs rather than end-to-end task auditing. The model is also restricted to English text, five-point Likert direct scales, and binary pairwise rankings. Nevertheless, given the consistent outperformance across eight diverse evaluation benchmarks and robust consistency metrics, there is high confidence in Prometheus 2 as a state-of-the-art open evaluation solution within its designated operational scope.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This foundational study establishes the LLM-as-a-judge paradigm and characterizes key biases in automated evaluators that Prometheus 2 directly targets and aims to resolve in an open-source framework.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). Reading this paper provides essential background on using rubric-based, criterion-driven prompting for LLM-based evaluation, a core evaluation format supported and refined by Prometheus 2.
- Paper: Can Large Language Models Be an Alternative to Human Evaluations?, David Cheng-Han Chiang et al. (2023). This work provides empirical motivation for replacing costly human annotations with LLM evaluators across direct assessment tasks.
- Paper: Large Language Models are not Fair Evaluators, Peiyi Wang et al. (2024). This paper offers critical context on positional biases in pairwise evaluations that Prometheus 2 must address to achieve high agreement with human judgments.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). JudgeBench directly evaluates the objective accuracy and robustness of LLM-based evaluators like Prometheus 2 on complex reasoning and code tasks.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). BiGGen Bench extends fine-grained, instance-specific evaluation rubrics to benchmark large suites of language models using evaluator LMs.
- Paper: From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder Pipeline, Tianle Li et al. (2025). This paper builds automated pipelines using LLM-as-a-judge mechanisms to create human-aligned, open-ended benchmark datasets.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This survey offers a comprehensive taxonomy of techniques, applications, and challenges in the LLM-as-a-judge ecosystem, contextualizing specialized evaluator models.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey synthesizes modern optimization strategies and bias mitigation techniques across automated LLM evaluator frameworks.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). This work investigates the vulnerabilities and adversarial robustness of automated LLM evaluators when judging complex instruction adherence.
- Paper: Split and Merge: Aligning Position Biases in LLM-based Evaluators, Zongjie Li et al. (2024). This paper introduces lightweight calibration methods to systematically reduce pairwise position bias in automated model judges.
