JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
Hossein A. RahmaniEmine YilmazNick CraswellBhaskar Mitra
Proposes JudgeBlender, an ensembling framework that combines judgments from smaller open-source language models across multiple architectures and prompts to match proprietary models in retrieval evaluation at lower cost and with reduced bias.
Evaluating and training search engines and information retrieval systems requires massive datasets of relevance judgments that determine how well a passage answers a query. Traditionally, these judgments depend on manual assessments by human evaluators, a process that is both prohibitively expensive and time-consuming. While recent efforts employ commercial large language models such as GPT-4 to automate scoring, relying on a single large proprietary model introduces high financial costs, reproducibility challenges, and inherent model biases that can unfairly favor systems built on similar architectures.
The article demonstrates and evaluates JudgeBlender, an automated evaluation framework that uses ensembles of smaller, open-source language models to generate accurate, robust, and cost-effective relevance judgments. The framework operates through two distinct variations: PromptBlender, which queries a single open-source model using multiple distinct prompt formulations, and LLMBlender, which queries a diverse panel of different open-source models using tailored prompts, subsequently aggregating their judgments into a unified score.
To establish empirical credibility, the evaluation used the benchmark dataset from the LLMJudge challenge, based on the TREC 2023 Deep Learning track. This test collection comprises thousands of query-passage pairs evaluated against expert human judgments on a four-point relevance scale. The experiments tested open-source models with approximately 7 to 8 billion parameters (Meta-Llama-3-8B, Mistral-7B, and Gemma-7B) using aggregator functions such as average voting and majority voting with various tie-breaking strategies, comparing them directly against individual models, specialized fine-tuned models, and leading commercial baselines like GPT-4o.
The findings show that ensembling smaller open-source models outperforms single-model baselines and matches or exceeds commercial systems. First, the LLMBlender framework achieved the highest correlation with human assessors across standard inter-rater reliability metrics (Cohen's Kappa and Krippendorff's Alpha), surpassing both standalone open-source models and commercial systems like GPT-4o. Second, LLMBlender demonstrated superior ranking preservation for retrieval systems, yielding the highest agreement with human-ranked benchmarks on standard search metrics. Third, the ensemble approach produced balanced accuracy across all relevance tiers, maintaining consistent performance rather than skewing toward solely relevant or non-relevant content. Finally, the blended methods mitigated evaluation bias, avoiding the systemic overestimation and underestimation of specific search architectures seen in single-model judges.
These results demonstrate that organizations do not need to rely on massive, expensive proprietary language models to perform reliable relevance evaluations. Deploying ensembles of small, open-source models substantially reduces operational API costs, avoids proprietary data-leakage risks, and enhances auditability while achieving equal or better alignment with human experts. In addition, the balanced perspective of a model jury prevents evaluation skew, ensuring fairer comparisons between competing retrieval technologies.
Organizations evaluating search and retrieval pipelines should consider adopting open-source ensemble methods like JudgeBlender as an alternative to single commercial models. Practical implementation should prioritize multi-model ensembles with consensus aggregation over single-model prompting where infrastructure allows. Because these evaluations were conducted on a single passage-ranking benchmark dataset, decision-makers should pilot the ensemble framework on internal domain-specific data to optimize panel composition, prompting designs, and fusion strategies before full-scale operational rollout.
- Paper: LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval, Hossein A. Rahmani et al. (2025). This overview introduces the LLMJudge challenge and its relevance-labeling evaluation setting that JudgeBlender directly uses to compare its ensemble strategies.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). Its survey of judge design, model and prompt selection, and reliability biases supplies the framework for understanding why JudgeBlender ensembles relevance assessments.
- Paper: LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion, Dongfu Jiang et al. (2023). Its LLM-Blender establishes the multi-model ranking and combination paradigm that JudgeBlender adapts from response generation to relevance judging.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This early study of LLMs as judges explains the human-alignment and bias concerns that motivate JudgeBlender’s alternative to a single evaluator.
- Paper: Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents, Weiwei Sun et al. (2023). Its experiments on LLM passage reranking provide context for using language models to make relevance judgments in information retrieval.
- Paper: From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder Pipeline, Tianle Li et al. (2025). BenchBuilder carries ensemble judging into refreshed, human-aligned retrieval-style benchmarks, using judge ensembles to reduce bias in pairwise model evaluation.
