HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Junyi LiXiaoxue ChengXin ZhaoJian-Yun NieJi-Rong Wen

article2023EMNLP560 citations

Presents HaluEval, a benchmark of 35,000 generated and human-annotated samples across question answering, dialogue, and summarization, revealing that large language models struggle to recognize factual fabrications without explicit external knowledge or step-by-step reasoning.

Listen

Large language models have rapidly expanded across consumer and enterprise applications, yet they frequently generate hallucinations—plausible-sounding text that is factually incorrect or unverifiable against real-world knowledge. This issue introduces serious operational, compliance, and reputational risks for organizations deploying automated language tools. The article introduces HaluEval, a large-scale evaluation benchmark comprising 35,000 samples designed to assess how frequently language models generate hallucinations and how effectively they can detect them across general queries and specific tasks.

To construct this benchmark, the researchers combined automated sample generation with rigorous human annotation. The automated pipeline used a two-stage approach: generating candidate incorrect samples via ChatGPT using structured instruction styles, followed by a filtering step to select the most plausible and difficult-to-detect errors. This yielded 30,000 task-specific examples across question answering, conversational dialogue, and text summarization. In parallel, thirty trained human annotators evaluated 5,000 divergent ChatGPT responses to user queries, achieving high inter-annotator agreement.

The investigation produced several critical findings regarding model reliability. First, human evaluation revealed that ChatGPT generated unverifiable or hallucinated information in 19.5% of its responses to divergent user queries, particularly around technology, climate, and language topics. Second, existing language models struggle significantly to detect hallucinations in text; ChatGPT achieved only 58.53% accuracy on summarization tasks and 62.59% on question answering, while several smaller open-source models performed below random chance (50%). Third, the majority of detection failures stemmed from subtle factual contradictions where statements appeared correct on the surface but conflicted with the underlying context. Finally, providing models with external reference knowledge substantially improved detection accuracy—boosting ChatGPT's performance in question answering from 62.59% to 76.83%—whereas intermediate reasoning steps yielded mixed results and comparing samples directly caused further confusion.

These findings demonstrate that language models cannot reliably self-police or catch factual errors through pure reasoning alone. Deploying language models in high-stakes environments without external factual grounding creates substantial operational risk. Consequently, organizations should integrate retrieval-augmented generation systems that supply verified, external domain knowledge directly to models, rather than relying solely on internal model weights or prompt-based reasoning. Additional pre-deployment testing against challenging benchmark datasets should be instituted before releasing automated text-generation workflows.

While the study provides high confidence through large sample sizes and rigorous human validation, readers should consider key limitations. The automated generation process relied on ChatGPT itself, meaning the complexity of benchmark errors is bounded by that model's capabilities. Additionally, the study focuses on evaluating error recognition rather than pinpointing root training causes, suggesting organizations should conduct domain-specific testing before relying entirely on these findings.

Cover for HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Abstract

Large language models (LLMs), such as ChatGPT, are prone to generate hallucinations, i.e., content that conflicts with the source or cannot be verified by the factual knowledge. To understand what types of content and to which extent LLMs are apt to hallucinate, we introduce the Hallucination Evaluation benchmark for Large Language Models (HaluEval), a large collection of generated and human-annotated hallucinated samples for evaluating the performance of LLMs in recognizing hallucination. To generate these samples automatically, we propose a two-stage framework, i.e., sampling-then-filtering. Besides, we hire some human labelers to annotate the hallucinations in ChatGPT responses. The empirical results suggest that ChatGPT is likely to generate hallucinated content related to specific topics by fabricating unverifiable information (i.e., about 19.5% responses). Moreover, existing LLMs face great challenges in recognizing the hallucinations in texts. However, our experiments also prove that providing external knowledge or adding reasoning steps can help LLMs recognize hallucinations. Our benchmark can be accessed at https://github.com/RUCAIBox/HaluEval.

Table of Contents

  • 1 Introduction
  • 2 The HaluEval Benchmark
  • 2.1 Automatic Generation
  • 2.2 Human Annotation
  • 2.3 Benchmark Analysis and Usage
  • 3 Experiments
  • 3.1 Experimental Setup
  • 3.2 Results and Analysis
  • 3.2.1 Hallucination Recognition
  • 3.2.2 Improvement Strategies
  • 3.3 Case Study
  • 4 Related Work
  • 5 Conclusion
  • 6 Limitations
  • Acknowledgments
  • References
  • Appendix
  • A Hallucination Sampling
  • B Hallucination Filtering
  • C Hallucination Recognition
  • D Details of HaluEval

Knowls

  1. Knowl 1 — HaluEval Benchmark Dataset Composition

    definition

    The Hallucination Evaluation benchmark for Large Language Models (HaluEval) is a collection of 35,000 evaluated text instances designed to analyze and evaluate the ability of large language models to generate and recognize hallucinations.

    The benchmark consists of two primary subsets:

    1. Task-Specific Hallucination Dataset (30,000 instances): Comprises 10,000 multi-hop question answering examples derived from HotpotQA, 10,000 knowledge-grounded dialogue examples derived from OpenDialKG, and 10,000 abstractive text summarization examples derived from CNN/DailyMail. For each example, an automated sampling-then-filtering pipeline pairs the reference ground-truth text with a plausible, subtle hallucinated counterpart.

    2. General User Queries Dataset (5,000 instances): Comprises 5,000 instruction-following queries selected from the 52K Alpaca dataset based on high response divergence across multiple generations by ChatGPT. Each query and its corresponding ChatGPT response are labeled by three human annotators for hallucinated content spans across three dimensions: unverifiable information, non-factual information, and irrelevant information.

  2. Knowl 2 — Sampling-then-Filtering Framework for Generating Hallucinated Text

    model/method

    The sampling-then-filtering framework is a two-stage automated generation method that leverages an instruction-following large language model (specifically ChatGPT) to create difficult and diverse hallucinated samples from factual seed data:

    1. Diverse Hallucination Sampling: Given a reference context and ground-truth text, candidate hallucinated texts are generated using two distinct instruction schemas:

      • One-pass Instruction: Feeds the complete task specification (system role, hallucination pattern definition, few-shot demonstrations, and length constraints) directly into the model to produce Candidate #1.
      • Conversational Instruction: Interactively presents the instruction rules step-by-step across conversational dialogue turns to ensure rule comprehension before prompting the model to produce Candidate #2.
    2. High-Quality Hallucination Filtering: To select the candidate that is hardest to distinguish from factual text, ChatGPT is prompted as an answer/summary/response judge. The model is provided with few-shot demonstrations comparing correct texts against hallucinated texts and is tasked with selecting the "best and correct" text between Candidate #1 and Candidate #2. The chosen candidate is retained as the final hallucinated counterpart for the benchmark.

  3. Knowl 3 — Task-Specific Hallucination Typology in HaluEval

    definition

    HaluEval categorizes hallucinations into distinct task-specific patterns across three benchmark tasks, generating 10,000 hallucinated samples per task:

    • Question Answering (HotpotQA):

      • Comprehension (Pattern I, 2,280 samples): The answer is factually true in the real world but contradicts or fails to follow from the provided context knowledge.
      • Factualness (Pattern II, 1,378 samples): The answer contains fabricated, non-existent, or incorrect factual entities.
      • Specificity (Pattern III, 5,102 samples): The answer is either overly general or overly specific relative to the question scope.
      • Inference (Pattern IV, 1,240 samples): The answer relies on faulty logical deduction or incorrect reasoning steps given the context.
    • Knowledge-Grounded Dialogue (OpenDialKG):

      • Extrinsic-Soft (Pattern I, 8,330 samples): The response replaces the ground-truth entity with a highly similar semantic entity.
      • Extrinsic-Hard (Pattern II, 1,196 samples): The response replaces the ground-truth entity with a dissimilar semantic entity within the same entity type.
      • Extrinsic-Grouped (Pattern III, 474 samples): The response replaces the ground-truth entity with an entity belonging to a completely different entity category.
    • Text Summarization (CNN/DailyMail):

      • Factual (Pattern I, 2,614 samples): The summary contains information that is factually true in the real world but cannot be directly entailed from the source document.
      • Non-factual (Pattern II, 3,562 samples): The summary introduces fabricated, incorrect, or unverifiable claims not present in the document.
      • Intrinsic (Pattern III, 3,824 samples): The summary presents direct factual contradictions or incorrect entity relationships relative to the source document.
  4. Knowl 4 — Human Annotation Pipeline and Hallucination Prevalence on General User Queries

    experimental setup

    To evaluate hallucinations on general user queries, 5,000 queries were sampled from the Alpaca 52K instruction dataset using semantic divergence filtering. For each query, three independent responses were generated using ChatGPT (gpt-3.5-turbo), and their pairwise semantic similarity was evaluated using BERTScore. The 5,000 queries exhibiting the lowest inter-response similarity were selected, based on the principle that hallucinations frequently occur when generative outputs diverge.

    Thirty undergraduate-educated annotators who achieved high label agreement with researchers on a validation set annotated the selected ChatGPT responses. For each sample, three annotators labeled whether the response contained hallucinations (categorized as unverifiable, non-factual, or irrelevant content) and marked the corresponding text spans. Final labels were determined via majority voting.

    Inter-annotator reliability reached a Fleiss's Kappa of κ=0.811\kappa = 0.811 (0.80≤κ≤1.000.80 \le \kappa \le 1.00), indicating perfect agreement. Among the 5,000 annotated ChatGPT responses, 977 responses were identified as containing hallucinated information, representing a baseline hallucination rate of 19.54% on divergent instruction queries.

  5. Knowl 5 — Hallucination Recognition Performance Across Large Language Models

    data/table

    Ten large language models were evaluated in a zero-shot binary classification setting (determining whether a given text contains hallucinated information: "Yes" vs "No"). The models were evaluated at temperature 0 across Question Answering (QA), Dialogue, Summarization, and General user queries.

    Models QA Dialogue Summarization General
    ChatGPT (gpt-3.5-turbo) 62.59 72.40 58.53 79.44
    Claude 2 69.78 64.73 57.75 75.00
    Claude 67.60 64.83 53.76 73.88
    Davinci002 (text-davinci-002) 60.05 60.81 47.77 80.42
    Davinci003 (text-davinci-003) 49.65 68.37 48.07 80.40
    GPT-3 (davinci) 49.21 50.02 51.23 72.72
    Llama 2-Chat (7B) 49.60 43.99 49.55 20.46
    ChatGLM (7B) 47.93 44.41 48.57 30.92
    Falcon (7B) 39.66 29.08 42.71 18.98
    Vicuna (7B) 60.34 46.35 45.62 19.48
    Alpaca (7B) 6.68 17.55 20.63 9.54

    The results demonstrate that LLMs struggle to detect subtle hallucinations in generated text. GPT-3 performs at near-random chance (~50%) across task-specific benchmarks. Smaller open-source 7B parameter models (such as Alpaca, Falcon, and Vicuna) achieve accuracies well below random chance due to severe classification biases toward predicting non-hallucination. Advanced instruction-tuned models (ChatGPT and Claude variants) achieve higher performance but struggle on text summarization, where ChatGPT obtains 58.53% accuracy.

  6. Knowl 6 — Pattern Distribution and Topic Sensitivity of LLM Hallucination Recognition Failures

    empirical result

    An error analysis of ChatGPT's failures in zero-shot hallucination recognition reveals strong pattern and topic biases across evaluated domains:

    Tasks #Failed P-I P-II P-III P-IV
    QA 3109 1559 245 278 1027
    Dialogue 891 465 344 82 -
    Summarization 3868 3106 705 57 -
    1. Pattern Distribution: Across QA, Dialogue, and Summarization, over 50% of ChatGPT's recognition failures stem from Pattern I (Comprehension in QA: 1,559 / 3,109; Extrinsic-soft in Dialogue: 465 / 891; Factual hallucination in Summarization: 3,106 / 3,868). These cases represent text spans that are factually true in isolation but conflict with or cannot be inferred from the provided context knowledge.

    2. Topic Sensitivity: Latent Dirichlet Allocation (LDA) topic clustering demonstrates that failure rates are topic-dependent. In QA, failures cluster predominantly in topics concerning film, company, and band. In general user queries, ChatGPT's hallucination recognition errors cluster heavily in topics related to technology, climate, and language.

  7. Knowl 7 — Effects of Knowledge Retrieval, Chain-of-Thought, and Contrastive Prompting on Hallucination Recognition

    data/table

    The effectiveness of three prompt-enhancement strategies was evaluated on ChatGPT (gpt-3.5-turbo) for hallucination recognition across QA, Dialogue, Summarization, and General user query datasets:

    Variants QA Dialogue Summarization General
    ChatGPT (Base) 62.59 72.40 58.53 86.22
    w/ Knowledge Retrieval 76.83 73.80 - 90.73
    w/ Chain-of-Thought (CoT) 59.58 71.39 61.21 86.50
    w/ Sample Contrast 49.19 68.67 49.46 -
    • Knowledge Retrieval: Providing relevant factual context retrieved from Wikipedia significantly improves hallucination detection accuracy in QA (+14.24%) and General queries (+4.51%), with minor gains in Dialogue (+1.40%).
    • Chain-of-Thought (CoT): Generating intermediate reasoning steps before issuing a judgment improves detection on Summarization (+2.68%), where logical checking between source document and summary is required, but degrades performance in QA (-3.01%) and Dialogue (-1.01%) due to reasoning interference.
    • Sample Contrast: Presenting both the correct ground-truth text and the hallucinated candidate simultaneously to force pairwise discrimination results in substantial performance degradation (dropping to near-random chance: 49.19% on QA and 49.46% on Summarization) because the high lexical and semantic similarity between candidate pairs creates severe model confusion.
  8. Knowl 8 — Instruction Design for Hallucination Generation, Filtering, and Recognition

    model/method

    The generation, filtering, and evaluation prompts in HaluEval are built with standardized functional components:

    1. Hallucination Generation Instruction: Contains four sections:

      • Intention Description: Assigns the persona (e.g., "hallucination answer generator") and states the goal of generating plausible yet factually incorrect output.
      • Hallucination Pattern: Explicitly states the failure mode to inject (e.g., factual contradiction, entity replacement, or unsupported inference).
      • Few-Shot Demonstrations: Ground-truth context, queries/dialogue/documents, and annotated hallucinated outputs.
      • Length Constraint: Strictly restricts the length of the hallucinated output (e.g., #Hallucinated Answer# can only have about 5 more words than #Right Answer#) to prevent models from using response length as a spurious classification shortcut.
    2. Hallucination Filtering Instruction: Assigns the role of an "answer judge" and provides reference knowledge, demonstrations comparing ground truth with hallucinated text, and two candidate hallucinated texts, prompting the model to choose the candidate that is subtlest and closest to the truth.

    3. Hallucination Recognition Instruction: Assigns the role of a judge, provides seed knowledge/dialogue history/document, the candidate text, definitions of hallucination types, few-shot demonstrations with "Yes"/"No" labels, and forces a strict binary "Yes" or "No" classification.

  9. Knowl 9 — Limitations of the HaluEval Benchmark

    limitation

    The HaluEval benchmark has three primary limitations noted by its design:

    1. Generator Capacity Constraints: The automated generation pipeline relies entirely on ChatGPT (gpt-3.5-turbo) to follow complex negative instructions for hallucination injection. Generated hallucinated samples are bounded by the instruction-following and reasoning abilities of this underlying model.
    2. Evaluation Scope vs. Causal Discovery: HaluEval benchmarks zero-shot recognition and discrimination capabilities of LLMs on hallucinated texts, but it does not analyze or evaluate the internal mechanistic causes, training dynamics, or representation failures that induce hallucination during autoregressive generation.
    3. Dual-Use Risk: Because the benchmark synthesizes fluent, highly plausible, and subtle factual errors that are difficult for existing models to distinguish from factual truth, the generated samples present potential misuse risks if repurposed for automated misinformation generation.

Coverage note — Specific domain-adapted prompt text variations from Appendix A-C are summarized within the instruction design knowl rather than extracted as repetitive individual knowls.

References

  1. 1.
    1. Introducing Falcon LLM . https://falconllm.tii.ae.
  2. 2.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. CoRR, abs/2302.04023.
  3. 3.David M Blei, Andrew Y Ng, and Michael I Jordan. 2003. Latent dirichlet allocation. Journal of machine Learning research, 3(Jan):993–1022.
  4. 4.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165.
  5. 5.Meng Cao, Yue Dong, and Jackie Chi Kit Cheung. 2022. Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3340–3354.
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  7. 7.Souvik Das, Sougata Saha, and Rohini K Srihari. 2023. Diving deep into modes of fact hallucinations in dialogue systems. arXiv preprint arXiv:2301.04449.
  8. 8.Bhuwan Dhingra, Manaal Faruqui, Ankur P. Parikh, Ming-Wei Chang, Dipanjan Das, and William W. Cohen. 2019. Handling divergent reference texts when evaluating table-to-text generation. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 4884–4895. Association for Computational Linguistics.
  9. 9.Nouha Dziri, Ehsan Kamalloo, Sivan Milton, Osmar R. Zaïane, Mo Yu, Edoardo Maria Ponti, and Siva Reddy. 2022a. Faithdial: A faithful benchmark for information-seeking dialogue. Trans. Assoc. Comput. Linguistics, 10:1473–1490.
  10. 10.Nouha Dziri, Andrea Madotto, Osmar Zaïane, and Avishek Joey Bose. 2021. Neural path hunter: Reducing hallucination in dialogue systems via path grounding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 2197–2214. Association for Computational Linguistics.
  11. 11.Nouha Dziri, Hannah Rashkin, Tal Linzen, and David Reitter. 2022b. Evaluating attribution in dialogue systems: The BEGIN benchmark. Trans. Assoc. Comput. Linguistics, 10:1066–1083.
  12. 12.Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022. Dialfact: A benchmark for fact-checking in dialogue. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 3785–3801. Association for Computational Linguistics.
  13. 13.Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022. TRUE: re-evaluating factual consistency evaluation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 3905–3920. Association for Computational Linguistics.
  14. 14.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  15. 15.Jinhao Jiang, Kun Zhou, Keming Ye Zican Dong, Wayne Xin Zhao, and Ji-Rong Wen. 2023. Structgpt: A general framework for large language model to reason on structured data.
  16. 16.Yunshi Lan, Gaole He, Jinhao Jiang, Jing Jiang, Wayne Xin Zhao, and Ji-Rong Wen. 2022. Complex knowledge base question answering: A survey. IEEE Transactions on Knowledge & Data Engineering, (01):1–20.
  17. 17.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  18. 18.Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jingyuan Wang, Jian-Yun Nie, and Ji-Rong Wen. 2023a. The web can be your oyster for improving language models. In Findings of the Association for Computational Linguistics: ACL 2023, pages 728–746.
  19. 19.Junyi Li, Tianyi Tang, Wayne Xin Zhao, and Ji-Rong Wen. 2021. Pretrained language model for text generation: A survey. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021, Virtual Event / Montreal, Canada, 19-27 August 2021, pages 4492–4499. ijcai.org.
  20. 20.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. 2023b. Evaluating object hallucination in large vision-language models.
  21. 21.Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. CoRR, abs/2303.08896.
  22. 22.Seungwhan Moon, Pararth Shah, Anuj Kumar, and Rajen Subba. 2019. Opendialkg: Explainable conversational reasoning with attention-based walks over knowledge graphs. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
  23. 23.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In NeurIPS.
  24. 24.Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2021. Measuring attribution in natural language generation models. CoRR, abs/2112.12870.
  25. 25.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1073–1083. Association for Computational Linguistics.
  26. 26.Weiwei Sun, Zhengliang Shi, Shen Gao, Pengjie Ren, Maarten de Rijke, and Zhaochun Ren. 2023. Contrastive learning reduces hallucination in conversations. In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7-14, 2023, pages 13618–13626. AAAI Press.
  27. 27.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  28. 28.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  29. 29.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS.
  30. 30.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Conference on Empirical Methods in Natural Language Processing (EMNLP).
  31. 31.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Peng Zhang, Yuxiao Dong, and Jie Tang. 2022. GLM-130B: an open bilingual pre-trained model. abs/2210.02414.
  32. 32.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  33. 33.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models. CoRR, abs/2303.18223.
  34. 34.Zheng Zhao, Shay B. Cohen, and Bonnie Webber. 2020. Reducing quantity hallucinations in abstractive summarization. CoRR, abs/2009.13312.
  35. 35.Shen Zheng, Jie Huang, and Kevin Chen-Chuan Chang. 2023. Why does chatgpt fall short in answering questions faithfully? CoRR, abs/2304.10513.

Citation

MLA
Li, J., et al. “HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 6449–64, https://doi.org/10.18653/v1/2023.emnlp-main.397.
APA
Li, J., Cheng, X., Zhao, X., Nie, J.-Y., & Wen, J.-R. (2023). HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6449–6464. https://doi.org/10.18653/v1/2023.emnlp-main.397
Chicago
Li, J., X. Cheng, X. Zhao, J.-Y. Nie, and J.-R. Wen. 2023. “HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6449–64. https://doi.org/10.18653/v1/2023.emnlp-main.397.
Harvard
Li, J. et al. (2023) “HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6449–6464. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.397.
Vancouver
1. Li J, Cheng X, Zhao X, Nie J-Y, Wen J-R (2023) HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6449–6464

BibTeX

@inproceedings{li-etal-2023-halueval,
    title = "{H}alu{E}val: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models",
    author = "Li, Junyi  and
      Cheng, Xiaoxue  and
      Zhao, Xin  and
      Nie, Jian-Yun  and
      Wen, Ji-Rong",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.397/",
    doi = "10.18653/v1/2023.emnlp-main.397",
    pages = "6449--6464"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/