LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval
Hossein A. RahmaniClemencia SiroMohammad AliannejadiNick CraswellCharles L A ClarkeGuglielmo FaggioliBhaskar MitraPaul ThomasEmine Yilmaz
Presents the scope and shared task for the LLM4Eval workshop at WSDM 2025, detailing how researchers use large language models to automate relevance judgments, evaluate retrieval-augmented generation pipelines, and replace or support human assessments.
Modern information systems increasingly rely on automated methods to evaluate search and text generation quality, replacing or augmenting expensive human review. Large language models (LLMs) have emerged as capable tools for these evaluation tasks, yet deploying them introduces questions around reliability, consistency, and alignment with human judgment. The article outlines the organization and objectives of the second LLM4Eval workshop at WSDM 2025, which evaluates and benchmarks the effectiveness, trustworthiness, and robustness of large language models when used as evaluators in information retrieval.
The workshop's approach combines interactive academic and industry discussions with an empirical competition called the LLMJudge challenge. In this shared evaluation task, participants use large language models to generate relevance labels for search datasets, aiming to maximize statistical correlation with ground-truth human judgments. This builds directly upon the first workshop iteration held at SIGIR 2024, which featured 18 accepted papers and 39 labeler submissions from seven university and industry research groups.
The primary findings and focal areas synthesized across these initiatives highlight that advanced models can effectively predict relevance and searcher preferences, often rivaling human evaluators. Prompt structuring techniques, such as chain-of-thought prompting, substantially increase the alignment between model outputs and human assessments compared to traditional automated evaluation metrics. However, critical operational risks remain, including evaluation validity, intrinsic randomness introduced by parameter tuning and prompt engineering, and challenges in maintaining reproducibility across evaluation runs.
These findings suggest that using generative models for evaluation can significantly reduce costs and speed up development timelines for search and retrieval systems, but uncalibrated adoption risks producing inconsistent or biased benchmarks. To mitigate these risks, organizations should establish rigorous validation pipelines and evaluate model stability before replacing human annotators. Future efforts must focus on standardizing benchmarks and resolving randomness to ensure trustworthy, replicable evaluation frameworks across academic and industrial deployments.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey provides a comprehensive foundation on the LLM-as-a-Judge paradigm, detailing automated judgment methodologies, biases, and evaluation frameworks central to the workshop's theme.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This foundational work establishes the core methodology for using LLMs as evaluators against human preferences and examines key evaluator biases directly relevant to automated judgment tasks.
- Paper: ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems, Jon Saad-Falcon et al. (2024). It introduces a foundational automated evaluation framework tailored specifically to retrieval-augmented generation pipelines, a central focus area of the workshop.
- Paper: Ragas: Automated Evaluation of Retrieval Augmented Generation, Shahul Es et al. (2024). This work establishes key metrics for the automated, LLM-based evaluation of retrieval-augmented generation systems, framing the context and relevance assessment discussed at the workshop.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). It introduces standard form-filling and chain-of-thought techniques for aligning LLM-based evaluators with human judgments across generative tasks.
- Paper: Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models, Seungone Kim et al. (2024). It presents an open-source model trained specifically to serve as an evaluator across custom rubrics and judgment tasks, directly addressing the workshop's focus on automated evaluation models.
- Paper: Humans or LLMs as the Judge? A Study on Judgement Bias, Guiming Chen et al. (2024). This paper analyzes vulnerabilities and systematic biases in LLM judgments compared to humans, providing crucial background on the trustworthiness of automated evaluators.
- Paper: Split and Merge: Aligning Position Biases in LLM-based Evaluators, Zongjie Li et al. (2024). It details mechanisms to mitigate position bias in LLM-based evaluators, addressing a primary obstacle in building reliable automated judgment systems.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This survey extends the workshop's discussions by providing a broad, structured taxonomy of emerging opportunities, architectures, and meta-evaluation benchmarks for LLM-as-a-judge systems.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). It advances the evaluation of automated judges by introducing a rigorous benchmark specifically designed to test their ability to detect subtle factual and logical reasoning errors.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). It generalizes LLM evaluation to autonomous multi-step agents by developing an Agent-as-a-Judge framework that evaluates execution traces and complex environments.
- Paper: Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation, Abubakarr Jaye et al. (2025). It extends automated LLM judging to agentic workflows by systematically validating intermediate reasoning steps and logs during task execution.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). It operationalizes fine-grained automated judging across diverse capabilities using instance-specific rubrics to evaluate over a hundred language models.
- Paper: From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder Pipeline, Tianle Li et al. (2025). It translates the principles of automated judging into an automated pipeline for generating challenging, human-aligned benchmarks from live crowdsourced data.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). It applies LLM-based automated evaluation to step-by-step reasoning processes, measuring validity and redundancy rather than just final outcome accuracy.
