JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

Hossein A. RahmaniEmine YilmazNick CraswellBhaskar Mitra

article2025arXiv8 citations

Proposes JudgeBlender, an ensembling framework that combines judgments from smaller open-source language models across multiple architectures and prompts to match proprietary models in retrieval evaluation at lower cost and with reduced bias.

Listen

Evaluating and training search engines and information retrieval systems requires massive datasets of relevance judgments that determine how well a passage answers a query. Traditionally, these judgments depend on manual assessments by human evaluators, a process that is both prohibitively expensive and time-consuming. While recent efforts employ commercial large language models such as GPT-4 to automate scoring, relying on a single large proprietary model introduces high financial costs, reproducibility challenges, and inherent model biases that can unfairly favor systems built on similar architectures.

The article demonstrates and evaluates JudgeBlender, an automated evaluation framework that uses ensembles of smaller, open-source language models to generate accurate, robust, and cost-effective relevance judgments. The framework operates through two distinct variations: PromptBlender, which queries a single open-source model using multiple distinct prompt formulations, and LLMBlender, which queries a diverse panel of different open-source models using tailored prompts, subsequently aggregating their judgments into a unified score.

To establish empirical credibility, the evaluation used the benchmark dataset from the LLMJudge challenge, based on the TREC 2023 Deep Learning track. This test collection comprises thousands of query-passage pairs evaluated against expert human judgments on a four-point relevance scale. The experiments tested open-source models with approximately 7 to 8 billion parameters (Meta-Llama-3-8B, Mistral-7B, and Gemma-7B) using aggregator functions such as average voting and majority voting with various tie-breaking strategies, comparing them directly against individual models, specialized fine-tuned models, and leading commercial baselines like GPT-4o.

The findings show that ensembling smaller open-source models outperforms single-model baselines and matches or exceeds commercial systems. First, the LLMBlender framework achieved the highest correlation with human assessors across standard inter-rater reliability metrics (Cohen's Kappa and Krippendorff's Alpha), surpassing both standalone open-source models and commercial systems like GPT-4o. Second, LLMBlender demonstrated superior ranking preservation for retrieval systems, yielding the highest agreement with human-ranked benchmarks on standard search metrics. Third, the ensemble approach produced balanced accuracy across all relevance tiers, maintaining consistent performance rather than skewing toward solely relevant or non-relevant content. Finally, the blended methods mitigated evaluation bias, avoiding the systemic overestimation and underestimation of specific search architectures seen in single-model judges.

These results demonstrate that organizations do not need to rely on massive, expensive proprietary language models to perform reliable relevance evaluations. Deploying ensembles of small, open-source models substantially reduces operational API costs, avoids proprietary data-leakage risks, and enhances auditability while achieving equal or better alignment with human experts. In addition, the balanced perspective of a model jury prevents evaluation skew, ensuring fairer comparisons between competing retrieval technologies.

Organizations evaluating search and retrieval pipelines should consider adopting open-source ensemble methods like JudgeBlender as an alternative to single commercial models. Practical implementation should prioritize multi-model ensembles with consensus aggregation over single-model prompting where infrastructure allows. Because these evaluations were conducted on a single passage-ranking benchmark dataset, decision-makers should pilot the ensemble framework on internal domain-specific data to optimize panel composition, prompting designs, and fusion strategies before full-scale operational rollout.

Cover for JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

Abstract

The effective training and evaluation of retrieval systems require a substantial amount of relevance judgments, which are traditionally collected from human assessors -- a process that is both costly and time-consuming. Large Language Models (LLMs) have shown promise in generating relevance labels for search tasks, offering a potential alternative to manual assessments. Current approaches often rely on a single LLM, such as GPT-4, which, despite being effective, are expensive and prone to intra-model biases that can favour systems leveraging similar models. In this work, we introduce JudgeBlender, a framework that employs smaller, open-source models to provide relevance judgments by combining evaluations across multiple LLMs (LLMBlender) or multiple prompts (PromptBlender). By leveraging the LLMJudge benchmark [18], we compare JudgeBlender with state-of-the-art methods and the top performers in the LLMJudge challenge. Our results show that JudgeBlender achieves competitive performance, demonstrating that very large models are often unnecessary for reliable relevance assessments.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 JudgeBlender
  • 4 Experimental Setup
  • 4.1 Dataset
  • 4.2 Aggregator Function
  • 4.3 Models and Prompt Families
  • 4.4 Evaluation Measurement
  • Correlation to Human Judgments.
  • System Ranking Correlation.
  • 4.5 Comparison Methods
  • 5 Results
  • 5.1 Correlation to Human Judgments
  • 5.2 System Ranking Correlation
  • 5.3 Inter-Judge Agreement Analysis
  • 5.4 Bias in System Evaluation
  • 6 Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — JudgeBlender ensembles relevance judgments across prompts or models

    model/method

    JudgeBlender estimates the relevance of a passage to a query by collecting several judgments and combining them into a final score. PromptBlender uses one language model with multiple distinct prompts, aiming to elicit different assessments from the same model. LLMBlender uses multiple language models, each with its own prompt, to obtain judgments from different model families. Both variants apply an aggregation function to the resulting relevance scores.

  2. Knowl 2 — Majority and average voting are the tested score aggregators

    model/method

    For a panel of relevance judges, average voting (AV) returns the arithmetic mean of their scores. Majority voting (MV) returns the most frequently assigned relevance level; when multiple levels tie for the most votes, the experiments resolve the tie by selecting a random tied level, the maximum tied level, the minimum tied level, or the average of the tied levels. The paper denotes these MV tie-breaking variants as Rnd, Max, Min, and Avg, respectively. Since averaging can yield a non-integer, AV and MV(Avg) need not return one of the original integer relevance levels.

  3. Knowl 3 — Evaluation uses the LLMJudge passage-judgment dataset

    experimental setup

    The experiments use the LLMJudge challenge dataset, based on the TREC Deep Learning 2023 passage-ranking task and its human-assigned relevance labels. Relevance is on a four-level scale: irrelevant (0), related (1), highly relevant (2), and perfectly relevant (3). The development split has 25 queries, 7,224 passages, and 7,263 query–passage judgments: 4,538 irrelevant, 1,403 related, 625 highly relevant, and 697 perfectly relevant. The test split has 25 queries, 4,414 passages, and 4,423 judgments: 2,005 irrelevant, 1,233 related, 808 highly relevant, and 377 perfectly relevant.

  4. Knowl 4 — The ensemble panels use three small open models and three prompt families

    experimental setup

    PromptBlender uses Meta-Llama-3-8B with three prompt strategies: direct relevance grading based on Thomas et al., criteria-based grading that decomposes relevance, and a two-step strategy that first makes a binary relevance decision and then assigns a score to passages judged relevant. LLMBlender uses Mistral-7B with the criteria-based prompt, Gemma-7B with the direct-grading prompt, and Llama-3-8B with the two-step prompt. The resulting judgments are combined with majority voting under each of the four tie-breaking rules or with average voting.

  5. Knowl 5 — Judgment agreement and system-ranking agreement are evaluated separately

    experimental setup

    Agreement between generated relevance judgments and human judgments is measured with Cohen’s κ and Krippendorff’s α. Agreement between system rankings induced by generated judgments and rankings based on human judgments is measured with Kendall’s τ and Spearman’s ρ, separately for NDCG@10 and MAP. The paper describes κ and α above 0.8 as strong agreement and above 0.6 as moderate agreement; τ and ρ range from −1 for complete disagreement to 1 for perfect ranking agreement, with 0 indicating no correlation.

  6. Knowl 6 — Blending achieves the strongest reported correlation with human relevance labels

    empirical result

    On the LLMJudge evaluation, the highest Cohen’s κ among the reported methods is 0.2619 for LLMBlender with majority voting and random tie-breaking; the strongest non-ensemble baseline is GPT-4o RelExp at 0.2519. The highest Krippendorff’s α is 0.4887 for PromptBlender with average voting, narrowly above 0.4877 for the Thomas direct-grading prompt with GPT-4-32k. Thus, different JudgeBlender variants lead on the two label-agreement metrics, and their best reported values exceed the best baseline values on those metrics.

  7. Knowl 7 — LLMBlender leads on NDCG@10 system ranking agreement, but not on MAP

    empirical result

    For system rankings evaluated with NDCG@10, LLMBlender with majority voting and average tie-breaking has the highest Kendall’s τ, 0.9612. Its Spearman’s ρ is 0.9940, tied with LLMBlender using random tie-breaking. For MAP, the strongest reported ranking correlations are instead from GenRE-dev, a fine-tuned Llama-3-8B method: Kendall’s τ is 0.9312 and Spearman’s ρ is 0.9879. The results therefore do not identify one evaluator as best across both system-ranking metrics.

  8. Knowl 8 — Per-level agreement shows different strengths across the ensemble and baselines

    empirical result

    In the reported four-level comparison with human labels, the percentage correct for PromptBlender with MV(Avg.) is 9.81% for perfectly relevant, 63.11% for highly relevant, 12.73% for related, and 72.51% for irrelevant judgments. For LLMBlender with MV(Avg.), the corresponding values are 26.79%, 59.28%, 28.46%, and 59.75%. MultiCriteria has higher agreement than LLMBlender on highly relevant judgments (73.76% versus 59.28%), while RelExp has higher agreement on irrelevant judgments (77.30% versus 59.75%). LLMBlender’s percentages exceed PromptBlender’s for perfectly relevant and related judgments but are lower for highly relevant and irrelevant judgments.

  9. Knowl 9 — Ensembles are reported to reduce system-type bias in NDCG@10 estimates

    empirical result

    The paper compares NDCG@10 for TREC Deep Learning 2023 runs under human judgments and four automatic assessors, grouping runs as GPT-based, T5-based, GPT+T5, or other systems. The authors report that RelExp tends to overestimate the effectiveness of top-performing GPT+T5 and T5 systems, while MultiCriteria notably overestimates GPT-based systems with lower human-measured effectiveness. PromptBlender with MV(Avg.) and LLMBlender with MV(Avg.) are described as more balanced across system categories, with estimates distributed more closely around the human-judgment diagonal. This is a qualitative finding from the plotted comparisons; the paper does not report a numerical bias measure.

  10. Knowl 10 — The evaluation is limited to one dataset and a small set of panels and prompts

    limitation

    The authors state that resource and budget constraints limited the experiments to one dataset and a small number of evaluator settings, panel compositions, and prompts. They identify evaluation on additional datasets and models, optimization of panel membership and prompting for quality–cost trade-offs, and exploration of more advanced aggregation strategies as future work.

Coverage note — The full per-method judgment confusion matrices and every baseline row are omitted; the main per-level comparisons and the best reported label- and ranking-agreement results are retained.

References

  1. 1.Abbasiantaeb, Z., Meng, C., Azzopardi, L., Aliannejadi, M.: Can we use large language models to fill relevance judgment holes? arXiv preprint arXiv:2405.05600 (2024)
  2. 2.Aniol, A., Pietron, M., Duda, J.: Ensemble approach for natural language question answering problem. In: 2019 Seventh International Symposium on Computing and Networking Workshops (CANDARW). pp. 180–183. IEEE (2019)
  3. 3.Breiman, L.: Bagging predictors. Machine learning 24, 123–140 (1996)
  4. 4.Craswell, N., Mitra, B., Yilmaz, E., Rahmani, H.A., Campos, D., Lin, J., Voorhees, E.M., Soboroff, I.: Overview of the trec 2023 deep learning track. In: Text REtrieval Conference (TREC). NIST, TREC (February 2024)
  5. 5.Dietterich, T.G.: Ensemble methods in machine learning. In: International workshop on multiple classifier systems. pp. 1–15. Springer (2000)
  6. 6.Faggioli, G., Dietz, L., Clarke, C.L., Demartini, G., Hagen, M., Hauff, C., Kando, N., Kanoulas, E., Potthast, M., Stein, B., et al.: Perspectives on large language models for relevance judgment. In: Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval. pp. 39–50 (2023)
  7. 7.Farzi, N., Dietz, L.: Best in tau@ llmjudge: Criteria-based relevance evaluation with llama3. arXiv preprint arXiv:2410.14044 (2024)
  8. 8.Farzi, N., Dietz, L.: Pencils down! automatic rubric-based evaluation of retrieve/generate systems. In: Proceedings of the 2024 ACM SIGIR International Conference on Theory of Information Retrieval. pp. 175–184 (2024)
  9. 9.Izacard, G., Grave, E.: Leveraging passage retrieval with generative models for open domain question answering. arXiv preprint arXiv:2007.01282 (2020)
  10. 10.Jahrer, M., Töscher, A., Legenstein, R.: Combining predictions for accurate recommender systems. In: Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining. pp. 693–702 (2010)
  11. 11.Jiang, D., Ren, X., Lin, B.Y.: Llm-blender: Ensembling large language models with pairwise comparison and generative fusion. In: Proceedings of the 61th Annual Meeting of the Association for Computational Linguistics (ACL 2023) (2023)
  12. 12.Meng, C., Arabzadeh, N., Askari, A., Aliannejadi, M., de Rijke, M.: Query performance prediction using relevance judgments generated by large language models. arXiv preprint arXiv:2404.01012 (2024)
  13. 13.Polikar, R.: Ensemble learning. Ensemble machine learning: Methods and applications pp. 1–34 (2012)
  14. 14.Rahmani, H.A., Craswell, N., Yilmaz, E., Mitra, B., Campos, D.: Synthetic test collections for retrieval evaluation. In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. pp. 2647–2651 (2024)
  15. 15.Rahmani, H.A., Siro, C., Aliannejadi, M., Craswell, N., Clarke, C.L.A., Faggioli, G., Mitra, B., Thomas, P., Yilmaz, E.: Llm4eval: Large language model for evaluation in ir. In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. p. 3040–3043. SIGIR ’24, Association for Computing Machinery, New York, NY, USA (2024)
  16. 16.Rahmani, H.A., Siro, C., Aliannejadi, M., Craswell, N., Clarke, C.L., Faggioli, G., Mitra, B., Thomas, P., Yilmaz, E.: Report on the 1st workshop on large language model for evaluation in information retrieval (llm4eval 2024) at sigir 2024. arXiv preprint arXiv:2408.05388 (2024)
  17. 17.Rahmani, H.A., Wang, X., Yilmaz, E., Craswell, N., Mitra, B., Thomas, P.: Syndl: A large-scale synthetic test collection for passage retrieval. arXiv preprint arXiv:2408.16312 (2024)
  18. 18.Rahmani, H.A., Yilmaz, E., Craswell, N., Mitra, B., Thomas, P., Clarke, C.L., Aliannejadi, M., Siro, C., Faggioli, G.: Llmjudge: Llms for relevance judgments. arXiv preprint arXiv:2408.08896 (2024)
  19. 19.Ravaut, M., Joty, S., Chen, N.F.: Towards summary candidates fusion. arXiv preprint arXiv:2210.08779 (2022)
  20. 20.Sagi, O., Rokach, L.: Ensemble learning: A survey. Wiley interdisciplinary reviews: data mining and knowledge discovery 8(4), e1249 (2018)
  21. 21.Sun, W., Yan, L., Ma, X., Wang, S., Ren, P., Chen, Z., Yin, D., Ren, Z.: Is chatgpt good at search? investigating large language models as re-ranking agents. arXiv preprint arXiv:2304.09542 (2023)
  22. 22.Thomas, P., Spielman, S., Craswell, N., Mitra, B.: Large language models can accurately predict searcher preferences. In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. pp. 1930–1940 (2024)
  23. 23.Upadhyay, S., Kamalloo, E., Lin, J.: Llms can patch up missing relevance judgments in evaluation. arXiv preprint arXiv:2405.04727 (2024)
  24. 24.Upadhyay, S., Pradeep, R., Thakur, N., Craswell, N., Lin, J.: Umbrela: Umbrela is the (open-source reproduction of the) bing relevance assessor. arXiv preprint arXiv:2406.06519 (2024)
  25. 25.Verga, P., Hofstatter, S., Althammer, S., Su, Y., Piktus, A., Arkhangorodsky, A., Xu, M., White, N., Lewis, P.: Replacing judges with juries: Evaluating llm generations with a panel of diverse models. arXiv preprint arXiv:2404.18796 (2024)
  26. 26.Xu, J., Li, H.: Adarank: a boosting algorithm for information retrieval. In: Proceedings of the 30th annual international ACM SIGIR conference on Research and development in information retrieval. pp. 391–398 (2007)

Citation

MLA
Rahmani, H. A., et al. “JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment”. arXiv, 2024, http://arxiv.org/abs/2412.13268v1.
APA
Rahmani, H. A., Yilmaz, E., Craswell, N., & Mitra, B. (2024). JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment. arXiv. http://arxiv.org/abs/2412.13268v1
Chicago
Rahmani, H. A., E. Yilmaz, N. Craswell, and B. Mitra. 2024. “JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment”. arXiv. http://arxiv.org/abs/2412.13268v1.
Harvard
Rahmani, H.A. et al. (2024) “JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2412.13268v1.
Vancouver
1. Rahmani HA, Yilmaz E, Craswell N, Mitra B (2024) JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment. arXiv

BibTeX

@article{rahmani2024judgeblender,
  title = {JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment},
  author = {Rahmani, Hossein A. and Yilmaz, Emine and Craswell, Nick and Mitra, Bhaskar},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2412.13268v1},
  eprint = {2412.13268}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/