LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval

Hossein A. RahmaniClemencia SiroMohammad AliannejadiNick CraswellCharles L A ClarkeGuglielmo FaggioliBhaskar MitraPaul ThomasEmine Yilmaz

article2025WSDM2 citations

Presents the scope and shared task for the LLM4Eval workshop at WSDM 2025, detailing how researchers use large language models to automate relevance judgments, evaluate retrieval-augmented generation pipelines, and replace or support human assessments.

Listen

Modern information systems increasingly rely on automated methods to evaluate search and text generation quality, replacing or augmenting expensive human review. Large language models (LLMs) have emerged as capable tools for these evaluation tasks, yet deploying them introduces questions around reliability, consistency, and alignment with human judgment. The article outlines the organization and objectives of the second LLM4Eval workshop at WSDM 2025, which evaluates and benchmarks the effectiveness, trustworthiness, and robustness of large language models when used as evaluators in information retrieval.

The workshop's approach combines interactive academic and industry discussions with an empirical competition called the LLMJudge challenge. In this shared evaluation task, participants use large language models to generate relevance labels for search datasets, aiming to maximize statistical correlation with ground-truth human judgments. This builds directly upon the first workshop iteration held at SIGIR 2024, which featured 18 accepted papers and 39 labeler submissions from seven university and industry research groups.

The primary findings and focal areas synthesized across these initiatives highlight that advanced models can effectively predict relevance and searcher preferences, often rivaling human evaluators. Prompt structuring techniques, such as chain-of-thought prompting, substantially increase the alignment between model outputs and human assessments compared to traditional automated evaluation metrics. However, critical operational risks remain, including evaluation validity, intrinsic randomness introduced by parameter tuning and prompt engineering, and challenges in maintaining reproducibility across evaluation runs.

These findings suggest that using generative models for evaluation can significantly reduce costs and speed up development timelines for search and retrieval systems, but uncalibrated adoption risks producing inconsistent or biased benchmarks. To mitigate these risks, organizations should establish rigorous validation pipelines and evaluate model stability before replacing human annotators. Future efforts must focus on standardizing benchmarks and resolving randomness to ensure trustworthy, replicable evaluation frameworks across academic and industrial deployments.

Cover for LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval

Abstract

Large language models (LLMs) have demonstrated increasing task-solving abilities not present in smaller models. Utilizing the capabilities and responsibilities of LLMs for automated evaluation (LLM4Eval) has recently attracted considerable attention in multiple research communities. For instance, LLM4Eval models have been studied in the context of automated judgments, natural language generation, and retrieval augmented generation systems. We believe that the information retrieval community can significantly contribute to this growing research area by designing, implementing, analyzing, and evaluating various aspects of LLMs with applications to LLM4Eval tasks. The main goal of LLM4Eval workshop is to bring together researchers from industry and academia to discuss various aspects of LLMs for evaluation in information retrieval, including automated judgments, retrieval-augmented generation pipeline evaluation, altering human evaluation, robustness, and trustworthiness of LLMs for evaluation in addition to their impact on real-world applications. We also plan to run an automated judgment challenge prior to the workshop, where participants will be asked to generate labels for a given dataset while maximising correlation with human judgments. The format of the workshop is interactive, including roundtable and keynote sessions and tends to avoid the one-sided dialogue of a mini-conference. This is the second iteration of the workshop. The first version was held in conjunction with SIGIR 2024, attracting over 50 participants.

Table of Contents

  • 1 Title
  • 2 Motivation
  • 3 Format and Planned Activities
  • 3.1 Pre-workshop: LLMJudge Challenge
  • 3.2 Synchronous Workshop
  • 4 Related Workshop
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Scope and Objectives of LLM4Eval in Information Retrieval

    definition

    The LLM4Eval paradigm investigates the utilization, capabilities, and responsibilities of Large Language Models (LLMs) for automated evaluation within information retrieval (IR) and retrieval-augmented generation (RAG) systems. Its primary research directions encompass:

    1. Automated Relevance Judgment: Using LLMs to estimate query-document relevance for ranking and for generating ground-truth labels that can train and benchmark downstream rankers.
    2. RAG Pipeline Evaluation: Assessing end-to-end performance, retrieval precision, and generation fidelity in retrieval-augmented text generation systems.
    3. Prompt and Alignment Engineering: Designing prompts (such as Chain-of-Thought prompts) and scoring functions that maximize alignment and correlation between automated LLM judgments and human assessments.
    4. Trustworthiness and Robustness: Examining the reliability, stability, and failure modes of LLM-based evaluators across diverse query types and domain distributions.
  2. Knowl 2 — Key Methodological Challenges for LLM-Based IR Evaluation

    limitation

    Employing Large Language Models (LLMs) as evaluators in information retrieval introduces four fundamental methodological challenges:

    • Evaluation Validity: Formulating verifiable methods to prove that LLM-derived evaluation scores and relevance labels reliably measure true system quality and reflect actual searcher satisfaction.
    • Intrinsic Randomness: Mitigating the non-deterministic variability introduced by prompt phrasing, parameter tuning (such as sampling temperature), and decoding heuristics, which can lead to unstable evaluation outcomes.
    • Replicability and Reproducibility: Ensuring consistency across differing model checkpoints, proprietary API updates, and experimental setups over time.
    • Human–LLM Parallelism and Divergence: Understanding how closely LLM judgments mirror human judgment distributions, including characterizing systematic biases and discrepancies where LLMs diverge from human evaluators.
  3. Knowl 3 — LLMJudge Shared Task Design for Automated Relevance Labeling

    experimental setup

    The LLMJudge Challenge is a shared evaluation task focused on benchmarking the capability of LLMs to generate automated relevance judgments for information retrieval (IR) workloads.

    • Objective: Participants design LLM-based labeling methodologies (including prompting strategies, model ensembles, and scoring heuristics) to predict relevance labels for a given evaluation collection of query-document pairs.
    • Evaluation Criterion: The primary performance metric is the correlation between the LLM-generated relevance labels and human relevance judgments.
    • Utility: Validated LLM labelers serve as reference-free relevance evaluators, facilitating cost-effective dataset curation and evaluation of retrieval algorithms without exhaustive human annotation.

Coverage note — All substantial contributions from this workshop proposal—namely the scope of LLM4Eval, the core methodological challenges of LLM-based evaluation, and the LLMJudge challenge setup—have been converted into knowls. Background literature citations and past edition attendance statistics were omitted.

References

  1. 1.Garbiel Bénédict, Ruqing Zhang, and Donald Metzler. 2023. Gen-ir@ sigir 2023: The first workshop on generative information retrieval. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 3460–3463.
  2. 2.Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations?. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 15607–15631. https://doi.org/10.18653/v1/2023.acl-long.870
  3. 3.Guglielmo Faggioli, Laura Dietz, Charles Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023. Perspectives on large language models for relevance judgment. arXiv:2304.09161 [cs.IR]
  4. 4.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110 (2022).
  5. 5.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 2511–2522. https://doi.org/10.18653/v1/2023.emnlp-main.153
  6. 6.Zheng Liu, Yujia Zhou, Yutao Zhu, Jianxun Lian, Chaozhuo Li, Zhicheng Dou, Defu Lian, and Jian-Yun Nie. 2024. Information Retrieval Meets Large Language Models. In Companion Proceedings of the ACM Web Conference 2024 (Singapore, Singapore) (WWW ’24). Association for Computing Machinery, New York, NY, USA, 1586–1589. https://doi.org/10.1145/3589335.3641299
  7. 7.Hossein A Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, and Daniel Campos. 2024. Synthetic test collections for retrieval evaluation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2647–2651.
  8. 8.Hossein A Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles LA Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, and Emine Yilmaz. 2024. Report on the 1st Workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) at SIGIR 2024. arXiv preprint arXiv:2408.05388 (2024).
  9. 9.Hossein A. Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, and Emine Yilmaz. 2024. LLM4Eval: Large Language Model for Evaluation in IR. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) (SIGIR ’24). Association for Computing Machinery, New York, NY, USA, 3040–3043. https://doi.org/10.1145/3626772.3657992
  10. 10.Hossein A Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, Paul Thomas, Charles LA Clarke, Mohammad Aliannejadi, Clemencia Siro, and Guglielmo Faggioli. 2024. LLMJudge: LLMs for Relevance Judgments. arXiv preprint arXiv:2408.08896 (2024).
  11. 11.Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2023. Large language models can accurately predict searcher preferences. arXiv preprint arXiv:2309.10621 (2023).
  12. 12.Jiaan Wang, Yunlong Liang, Fandong Meng, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023. Is chatgpt a good nlg evaluator? a preliminary study. arXiv preprint arXiv:2303.04048 (2023).
  13. 13.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675 (2019).

Citation

MLA
Rahmani, H. A., et al. “LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval”. Proceedings of the Eighteenth ACM International Conference on Web Search and Data Mining, 2025, pp. 1120–21, https://doi.org/10.1145/3701551.3705706.
APA
Rahmani, H. A., Siro, C., Aliannejadi, M., Craswell, N., Clarke, C. L. A., Faggioli, G., Mitra, B., Thomas, P., & Yilmaz, E. (2025). LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval. Proceedings of the Eighteenth ACM International Conference on Web Search and Data Mining, 1120–1121. https://doi.org/10.1145/3701551.3705706
Chicago
Rahmani, H. A., C. Siro, M. Aliannejadi, et al. 2025. “LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval”. Proceedings of the Eighteenth ACM International Conference on Web Search and Data Mining, 1120–21. https://doi.org/10.1145/3701551.3705706.
Harvard
Rahmani, H.A. et al. (2025) “LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval”, Proceedings of the Eighteenth ACM International Conference on Web Search and Data Mining. ACM, pp. 1120–1121. Available at: https://doi.org/10.1145/3701551.3705706.
Vancouver
1. Rahmani HA, Siro C, Aliannejadi M, Craswell N, Clarke CLA, Faggioli G, Mitra B, Thomas P, Yilmaz E (2025) LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval. In: Proceedings of the Eighteenth ACM International Conference on Web Search and Data Mining. ACM, pp 1120–1121

BibTeX

@inproceedings{Rahmani_2025, series={WSDM ’25}, title={LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval}, url={http://dx.doi.org/10.1145/3701551.3705706}, DOI={10.1145/3701551.3705706}, booktitle={Proceedings of the Eighteenth ACM International Conference on Web Search and Data Mining}, publisher={ACM}, author={Rahmani, Hossein A. and Siro, Clemencia and Aliannejadi, Mohammad and Craswell, Nick and Clarke, Charles L.A. and Faggioli, Guglielmo and Mitra, Bhaskar and Thomas, Paul and Yilmaz, Emine}, year={2025}, month=Mar, pages={1120–1121}, collection={WSDM ’25} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF