HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Junyi LiXiaoxue ChengXin ZhaoJian-Yun NieJi-Rong Wen
Presents HaluEval, a benchmark of 35,000 generated and human-annotated samples across question answering, dialogue, and summarization, revealing that large language models struggle to recognize factual fabrications without explicit external knowledge or step-by-step reasoning.
Large language models have rapidly expanded across consumer and enterprise applications, yet they frequently generate hallucinations—plausible-sounding text that is factually incorrect or unverifiable against real-world knowledge. This issue introduces serious operational, compliance, and reputational risks for organizations deploying automated language tools. The article introduces HaluEval, a large-scale evaluation benchmark comprising 35,000 samples designed to assess how frequently language models generate hallucinations and how effectively they can detect them across general queries and specific tasks.
To construct this benchmark, the researchers combined automated sample generation with rigorous human annotation. The automated pipeline used a two-stage approach: generating candidate incorrect samples via ChatGPT using structured instruction styles, followed by a filtering step to select the most plausible and difficult-to-detect errors. This yielded 30,000 task-specific examples across question answering, conversational dialogue, and text summarization. In parallel, thirty trained human annotators evaluated 5,000 divergent ChatGPT responses to user queries, achieving high inter-annotator agreement.
The investigation produced several critical findings regarding model reliability. First, human evaluation revealed that ChatGPT generated unverifiable or hallucinated information in 19.5% of its responses to divergent user queries, particularly around technology, climate, and language topics. Second, existing language models struggle significantly to detect hallucinations in text; ChatGPT achieved only 58.53% accuracy on summarization tasks and 62.59% on question answering, while several smaller open-source models performed below random chance (50%). Third, the majority of detection failures stemmed from subtle factual contradictions where statements appeared correct on the surface but conflicted with the underlying context. Finally, providing models with external reference knowledge substantially improved detection accuracy—boosting ChatGPT's performance in question answering from 62.59% to 76.83%—whereas intermediate reasoning steps yielded mixed results and comparing samples directly caused further confusion.
These findings demonstrate that language models cannot reliably self-police or catch factual errors through pure reasoning alone. Deploying language models in high-stakes environments without external factual grounding creates substantial operational risk. Consequently, organizations should integrate retrieval-augmented generation systems that supply verified, external domain knowledge directly to models, rather than relying solely on internal model weights or prompt-based reasoning. Additional pre-deployment testing against challenging benchmark datasets should be instituted before releasing automated text-generation workflows.
While the study provides high confidence through large sample sizes and rigorous human validation, readers should consider key limitations. The automated generation process relied on ChatGPT itself, meaning the complexity of benchmark errors is bounded by that model's capabilities. Additionally, the study focuses on evaluating error recognition rather than pinpointing root training causes, suggesting organizations should conduct domain-specific testing before relying entirely on these findings.
- Paper: Survey of Hallucination in Natural Language Generation, Ziwei Ji et al. (2022). Provides a comprehensive taxonomy and foundational framework for understanding intrinsic versus extrinsic hallucinations in natural language generation that HaluEval directly operationalizes.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Introduces the core benchmark and methodology for evaluating factual falsehoods and misconceptions in language models that underpins subsequent hallucination evaluation suites.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). Establishes foundational distinctions between factuality and faithfulness in generation tasks that guide HaluEval's benchmark design.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). Analyzes parametric versus non-parametric knowledge limits in language models, explaining the mechanisms behind the factual fabrications HaluEval seeks to evaluate.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). Demonstrates the power of chain-of-thought prompting to improve reasoning, directly motivating HaluEval's finding that multi-step reasoning aids hallucination recognition.
- Paper: Confabulation: The Surprising Value of Large Language Model Hallucinations, Peiqi Sui et al. (2024). Directly employs the HaluEval benchmark to analyze the narrative and communicative properties of model confabulations.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). Synthesizes modern hallucination evaluation benchmarks like HaluEval into an end-to-end taxonomy and comprehensive survey of detection and mitigation methods.
- Paper: Factuality of Large Language Models: A Survey, Yuxia Wang et al. (2024). Builds on general hallucination benchmarks to categorize modern factuality evaluations, detection metrics, and mitigation frameworks across textual and multimodal domains.
- Paper: RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models, Cheng Niu et al. (2024). Extends the evaluation of factual errors to fine-grained, word-level hallucinations specifically within retrieval-augmented generation systems.
- Paper: Med-HALT: Medical Domain Hallucination Test for Large Language Models, Ankit Pal et al. (2023). Adapts large-scale hallucination testing protocols to high-stakes clinical and biomedical reasoning scenarios.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). Applies insights from hallucination benchmarks to train preference-based reinforcement learning pipelines that actively reduce factual errors.
- Paper: SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models, Potsawee Manakul et al. (2023). Develops a zero-resource, black-box consistency checking method to detect hallucinations identified by large-scale benchmarks without requiring external databases.
- Paper: Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?, Zorik Gekhman et al. (2024). Investigates whether fine-tuning on new knowledge actively induces the factual hallucinations quantified by benchmarks like HaluEval.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). Introduces multi-turn cross-examination between language models as a specialized detection mechanism for the types of hallucinations benchmarked in HaluEval.
- Paper: Evaluating Object Hallucination in Large Vision-Language Models, Yifan Li et al. (2023). Generalizes benchmark-driven hallucination evaluation from pure text models to multimodal vision-language models.
