RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models

Cheng NiuYuanhao WuJuno ZhuSiliang XuKashun ShumRandy ZhongJuntong SongTong Zhang

article2024ACL403 citations

Presents RAGTruth, a large-scale word-level hallucination benchmark of nearly 18,000 manually annotated retrieval-augmented generations that enables smaller language models to match GPT-4 in detecting and mitigating factual errors.

Listen

Retrieval-augmented generation is widely deployed to reduce factual errors in large language models by supplying them with relevant reference materials. However, models still generate unsupported or contradictory claims relative to these references, posing significant reliability risks in high-stakes operational environments. To address this issue, the article introduces a large-scale evaluation benchmark named RAGTruth to measure, detect, and mitigate word-level hallucinations in retrieval-augmented workflows.

The research team constructed an annotated corpus of nearly 18,000 natural language responses generated by six major commercial and open-source models across three tasks: question answering, structured data-to-text writing, and news summarization. Professional annotators manually labeled and categorized 14,289 hallucinated text segments into four distinct categories: evident conflicts, subtle conflicts, evident baseless additions, and subtle baseless additions. The authors then evaluated existing automated detection methods and fine-tuned an open-source model using the dataset to detect and filter erroneous outputs.

The evaluation revealed several critical findings. First, introducing baseless information was substantially more common than directly contradicting the context, and hallucinations occurred most frequently in structured data-to-text generation (affecting 68.6% of responses) compared to question answering (29.1%) and news summarization (29.7%). Second, hallucination rates increased with response length and tended to cluster toward the end of generated answers. Third, leading prompt-based detection using top-tier models achieved limited accuracy, with GPT-4-turbo attaining a response-level F1 score of 63.4% and a low span-level precision of 18.4%. In contrast, fine-tuning a smaller, 13-billion-parameter open-source model on the RAGTruth dataset achieved an overall response-level F1 score of 78.7% and a span-level F1 score of 52.7%. Applying this fine-tuned detector to filter model outputs reduced response hallucination rates by up to 51.0% for commercial models and up to 63.2% for smaller open-source models.

These findings indicate that general prompting of frontier models is insufficient for reliable error catching, and organizations cannot assume retrieval grounding alone eliminates factual inaccuracies. Instead, deploying smaller, specialized detection models fine-tuned on high-quality task data provides a more accurate and cost-effective quality assurance layer for operational pipelines.

Organizations deploying retrieval-based systems should implement targeted post-generation detection filters and format input data clearly to avoid common pitfalls, such as models misinterpreting missing structured values as explicit negatives. Moving forward, teams should benchmark the transferability of specialized detectors across their unique business domains and evaluate cost-effective combinations of human annotation and synthetic training data.

arXiv: 2401.00396
Cover for RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models

Abstract

Retrieval-augmented generation (RAG) has become a main technique for alleviating hallucinations in large language models (LLMs). Despite the integration of RAG, LLMs may still present unsupported or contradictory claims to the retrieved contents. In order to develop effective hallucination prevention strategies under RAG, it is important to create benchmark datasets that can measure the extent of hallucination. This paper presents RAGTruth, a corpus tailored for analyzing word-level hallucinations in various domains and tasks within the standard RAG frameworks for LLM applications. RAGTruth comprises nearly 18,000 naturally generated responses from diverse LLMs using RAG. These responses have undergone meticulous manual annotations at both the individual case and word levels, incorporating evaluations of hallucination intensity. We not only benchmark hallucination frequencies across different LLMs, but also critically assess the effectiveness of several existing hallucination detection methodologies. We show that using a high-quality dataset such as RAGTruth, it is possible to finetune a relatively small LLM and achieve a competitive hallucination detection performance when compared to the existing prompt-based approaches using state-of-the-art LLMs such as GPT-4. Furthermore, the finetuned model can effectively mitigate hallucination in LLM responses.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Hallucination of Large Language Models
  • 2.2 Hallucination Evaluation Datasets
  • 2.3 Hallucination Detection Methods
  • 3 Construction Process of RAGTruth
  • 3.1 Hallucination Taxonomy
  • 3.2 Response Generation
  • 3.3 Human Annotation
  • 3.4 Annotations for Adaptive Evaluation
  • 4 Hallucination Benchmark Analysis
  • 4.1 Basic Statistics
  • 4.2 Hallucination Statistics
  • 5 Experimental Setup
  • 5.1 Hallucination Detection Algorithms
  • 5.2 Data Split
  • 5.3 Evaluation Metrics
  • 6 Experimental Results
  • 6.1 Response-level Detection
  • 6.2 Span-level Detection
  • 6.3 Hallucination Suppression
  • 7 Conclusion
  • 8 Limitations
  • 9 Ethical considerations
  • Acknowledgement
  • References
  • A Hallucination Examples
  • EVIDENT CONFLICT
  • SUBTLE CONFLICT
  • EVIDENT INTRODUCTION OF BASELESS INFORMATION
  • SUBTLE INTRODUCTION OF BASELESS INFORMATION
  • B Response Generation Prompts
  • QUESTION ANSWERING
  • DATA-TO-TEXT WRITING
  • SUMMARIZATION
  • C Annotation Details
  • D Hallucination Detection Prompts
  • SUMMARIZATION
  • QUESTION ANSWERING
  • DATA-TO-TEXT WRITING

Knowls

  1. Knowl 1 — Four-Way Taxonomy and Adaptive Annotations for RAG Hallucinations

    definition

    In retrieval-augmented generation (RAG), text generation is conditioned on retrieved reference contexts. Hallucinations in this setting are categorized into four fine-grained types across two main dimensions (conflict vs. baseless information, and evident vs. subtle severity):

    1. Evident Conflict: Generated content directly contradicts or opposes facts in the retrieved reference context. These errors are readily verifiable without complex contextual reasoning (e.g., conflicting numbers, inverted dates, or misspelled entity names).
    2. Subtle Conflict: Generated content diverges from the reference context in a way that alters the intended contextual meaning, nuance, or severity (e.g., substituting legal terms or altering descriptive intensity).
    3. Evident Introduction of Baseless Information: Generated content includes facts, entities, or claims entirely absent from and unsupported by the reference context (e.g., inventing unmentioned events, reconciliations, or business services).
    4. Subtle Introduction of Baseless Information: Generated content extrapolates beyond the reference context by introducing inferred details, subjective commentary, sentiments, or common-sense assumptions not explicitly stated in the source.

    Additionally, two special annotation tags support adaptive evaluation:

    • implicit_true: Flags spans that introduce factual world knowledge not contained in the provided reference context, but which may be true in reality. While open-domain generation benefits from parametric knowledge, strict RAG requires strict grounding on provided contexts.
    • due_to_null: Flags spans in structured-data tasks where a JSON null or None attribute is incorrectly interpreted by the language model as a negative assertion (e.g., interpreting "RestaurantsReservations": null as "does not take reservations").
  2. Knowl 2 — RAGTruth Corpus Composition and Basic Statistics

    data/table

    The RAGTruth corpus comprises 2,965 input instances, 17,790 naturally generated responses (6 model outputs per instance), and 14,289 human-annotated hallucination spans across three core RAG tasks:

    1. Question Answering (QA): Sampled from the MS MARCO training set, filtered to daily-life questions with 3 retrieved reference passages per question.
    2. Data-to-Text Writing (D2T): Sampled from the Yelp Open Dataset (restaurant and nightlife categories), using structured JSON business attributes and up to 3 user reviews.
    3. News Summarization: Sampled from the CNN/Daily Mail dataset and recent platform news articles.

    Responses were generated using six large language models: GPT-3.5-turbo-0613, GPT-4-0613, Mistral-7B-Instruct, Llama-2-7B-chat, Llama-2-13B-chat, and Llama-2-70B-chat (4-bit AWQ quantized). Each response was independently labeled by two professional annotators (achieving 91.8% response-level agreement and 78.8% span-level agreement), with discrepancies adjudicated by a third reviewer.

    Task # Instance # Resp. Context Length Resp. Length Hallucination
    Mean Max Mean Max # Resp. (%) # Span
    Question Answering 989 5,934 243 509 119 381 1,724 (29.1%) 2,927
    Data-to-text Writing 1,033 6,198 354 1,253 159 369 4,254 (68.6%) 9,290
    Summarization (CNN/DM) 628 3,768 648 1,749 124 632 1,165 (30.9%) 1,474
    Summarization (Recent News) 315 1,890 369 481 89 240 521 (27.6%) 598
    Overall 2,965 17,790 381 1,749 131 632 7,664 (43.1%) 14,289

    Baseless information is significantly more prevalent than conflicting information across all tasks. Overall, 13.5% (1,928 spans) of all hallucinated spans are annotated as implicit_true, and 11.5% (1,642 spans) are annotated as due_to_null (accounting for 17.7% of all hallucinated spans in data-to-text writing).

  3. Knowl 3 — Model-Level Hallucination Density and Rates in RAG Tasks

    data/table

    Hallucination frequency varies substantially across language model architectures and tasks. Hallucination density is defined as the average number of hallucinated spans per 100 words in generated responses:

    Hallucination Density=Total Hallucinated SpansTotal Words Generated×100\text{Hallucination Density} = \frac{\text{Total Hallucinated Spans}}{\text{Total Words Generated}} \times 100

    Model QA Data-to-Text Summarization Overall
    # Resp # Span Dens. # Resp # Span Dens. # Resp # Span Dens. # Resp # Span
    GPT-3.5-turbo-0613 75 89 0.12 272 384 0.18 54 60 0.05 401 533
    GPT-4-0613 48 51 0.06 290 354 0.27 74 80 0.08 406 485
    Llama-2-7B-chat 510 1,010 0.59 888 1,775 1.27 434 517 0.58 1,832 3,302
    Llama-2-13B-chat 399 654 0.48 983 2,803 1.53 295 342 0.41 1,677 3,799
    Llama-2-70B-chat (4-bit) 320 529 0.40 863 1,834 1.15 212 245 0.26 1,395 2,608
    Mistral-7B-Instruct 378 594 0.59 958 2,140 1.51 617 828 0.86 1,953 3,562

    Key empirical findings:

    1. Proprietary models (GPT-4 and GPT-3.5-turbo) exhibit significantly lower hallucination rates and densities than open-source models across all tasks, with GPT-4 achieving the fewest hallucinated spans overall (485 spans across 406 responses).
    2. Across open-source models in the Llama-2 family, scaling model parameters from 7B to 70B leads to a consistent decrease in hallucination density in QA (0.59 to 0.40) and summarization (0.58 to 0.26).
    3. Mistral-7B-Instruct produces the highest overall count of responses containing hallucinations (1,953 responses), showing elevated hallucination rates in summarization (density 0.86) and data-to-text writing (density 1.51).
  4. Knowl 4 — Hallucination Distribution Across Response Positions and Text Lengths

    empirical result

    Hallucination occurrences in RAG responses exhibit distinct length and positional dependencies:

    1. Response Length Dependence: Partitioning responses into three equal-sized length buckets (RLB 1, 2, and 3) reveals that the average number of hallucinated spans per response increases monotonically with response length across all tasks:

      • Question Answering: bucket 1 (19–93 words) = 0.21 spans; bucket 2 (93–138 words) = 0.37 spans; bucket 3 (138–257 words) = 0.87 spans.
      • Data-to-Text Writing: bucket 1 (93–131 words) = 1.20 spans; bucket 2 (131–175 words) = 1.59 spans; bucket 3 (175–258 words) = 1.69 spans.
      • Summarization: bucket 1 (44–87 words) = 0.34 spans; bucket 2 (87–119 words) = 0.32 spans; bucket 3 (119–245 words) = 0.44 spans.
    2. Context Length Dependence: In news summarization, longer source contexts lead to more hallucinations (0.29 spans in bucket 1 [176–368 words] versus 0.44 spans in bucket 3 [587–1422 words]). For QA and data-to-text writing, context length variations show minimal impact on hallucination frequency (e.g., QA remains stable at 0.50, 0.51, and 0.49 spans across context buckets).

    3. Positional Bias: In question answering and news summarization, hallucinations concentrate heavily toward the end of generated responses (the final 25%–30% of normalized token positions). In data-to-text generation, hallucinations are elevated in the first half of the response where business attributes and metadata claims are initially introduced.

  5. Knowl 5 — Supervised Fine-Tuning Setup of Llama-2-13B for Hallucination Detection

    model/method

    Hallucination detection is formulated as a conditioned sequence generation task where the model directly outputs the exact spans containing hallucinated text.

    Input Format: The model receives a structured prompt containing the reference context (e.g., news article, retrieved passages, or JSON business attributes), the candidate response generated by a target LLM, and task-specific instructions defining conflict and baseless information.

    Output Format: The model is trained to output a JSON object with the key "hallucination list" mapping to a list of exact hallucinated substring spans: {"hallucination list": ["<span1>", "<span2>", ...]}. If no hallucinations exist, it outputs {"hallucination list": []}.

    Training Configuration:

    • Base model: Llama-2-13B (chat/base).
    • Dataset split: 2,515 training instances (15,090 response examples) from RAGTruth; 450 instances (150 per task) held out for testing.
    • Optimization: Full parameter supervised fine-tuning for 1 epoch with an initial learning rate of 2×10−52 \times 10^{-5}.
    • Hardware: 4 NVIDIA A100 GPUs.
  6. Knowl 6 — Response-Level Hallucination Detection Performance

    data/table

    Response-level hallucination detection evaluates whether a detector correctly classifies an entire model response as containing at least one hallucination. Evaluations are conducted on the 450-instance RAGTruth test set.

    Baselines include:

    • Prompt Baseline (GPT-3.5-turbo / GPT-4-turbo): Zero-shot prompt instructing the model to identify conflicts and baseless info.
    • SelfCheckGPT (GPT-3.5-turbo): Zero-resource black-box fact checking sampling 3 stochastic alternative responses per query to verify sentence-level consistency.
    • LMvLM (GPT-4-turbo): Multi-turn cross-examination between two LLMs to expose factual inconsistencies.
    • Finetuned Llama-2-13B: Supervised fine-tuning on RAGTruth train split.
    Methods QA Data-to-Text Summarization Overall
    Prec. Rec. F1 Prec. Rec. F1 Prec. Rec. F1 Prec. Rec. F1
    Prompt (GPT-3.5-turbo) 18.8 84.4 30.8 65.1 95.5 77.4 23.4 89.2 37.1 37.1 92.3 52.9
    Prompt (GPT-4-turbo) 33.2 90.6 45.6 64.3 100.0 78.3 31.5 97.6 47.6 46.9 97.9 63.4
    SelfCheckGPT (GPT-3.5-t) 35.0 58.0 43.7 68.2 82.8 74.8 31.1 56.5 40.1 49.7 71.9 58.8
    LMvLM (GPT-4-turbo) 18.7 76.9 30.1 68.0 76.7 72.1 23.3 81.9 36.2 36.2 77.8 49.4
    Finetuned Llama-2-13B 61.6 76.3 68.2 85.4 91.0 88.1 64.0 54.9 59.1 76.9 80.7 78.7

    Prompting GPT-4-turbo achieves very high recall (97.9%) but suffers from low precision (46.9%) due to over-flagging non-hallucinated content. The specialized Llama-2-13B model fine-tuned on RAGTruth achieves the highest overall precision (76.9%) and F1 score (78.7%), outperforming all general LLM prompting and interactive baselines.

  7. Knowl 7 — Span-Level Character Overlap Hallucination Detection Evaluation

    data/table

    Span-level evaluation measures fine-grained localization accuracy by calculating character-level precision, recall, and F1 between predicted hallucination character spans and ground-truth human-annotated spans on the RAGTruth test set (450 instances):

    Methods QA Data-to-Text Summarization Overall
    Prec. Rec. F1 Prec. Rec. F1 Prec. Rec. F1 Prec. Rec. F1
    Prompt (GPT-3.5-turbo) 7.9 25.1 12.1 8.7 45.1 14.6 6.1 33.7 10.3 7.8 35.3 12.8
    Prompt (GPT-4-turbo) 23.7 52.0 32.6 17.9 66.4 28.2 14.7 65.4 24.1 18.4 60.9 28.3
    Finetuned Llama-2-13B 55.8 60.8 58.2 56.5 50.7 53.5 52.4 30.8 38.8 55.6 50.2 52.7

    GPT-4-turbo achieves an overall span-level F1 score of only 28.3% with an extremely low precision of 18.4%. Finetuning Llama-2-13B on RAGTruth improves span-level F1 to 52.7% (55.6% precision and 50.2% recall).

    Across hallucination sub-types, models consistently achieve higher recall on evident hallucinations than subtle hallucinations:

    • GPT-4-turbo: 63.4% recall on evident conflict vs. 49.8% on subtle conflict; 66.3% on evident baseless info vs. 60.4% on subtle baseless info.
    • Finetuned Llama-2-13B: 38.3% recall on evident conflict vs. 2.5% on subtle conflict; 55.8% on evident baseless info vs. 52.9% on subtle baseless info.
  8. Knowl 8 — Post-Generation Hallucination Mitigation via Detector-Based Response Selection

    data/table

    A fine-tuned hallucination detection model can be deployed as a post-generation filter to reduce hallucinations by selecting between two candidate responses generated by different LLMs with comparable baseline hallucination tendencies. Two selection strategies are evaluated against a random-selection baseline on the 450 test instances:

    1. Strategy 1 (Fewer Spans): Select the candidate response with the smaller number of predicted hallucinated spans.
    2. Strategy 2 (Zero Spans): Select the candidate response that has zero predicted hallucinated spans (applied to the subset of instances where at least one candidate has zero predicted spans).
    Model Pair (Baseline Rate) Selection Strategy Valid Instances Hallucination Rate (%)
    Llama-2-7B-chat (51.8%) Random Selection 450 52.4 (-)
    Mistral-7B-Instruct (57.6%) Select response with fewer spans 450 41.1 (↓\downarrow 21.6%)
    Select response with zero spans 328 19.3 (↓\downarrow 63.2%)
    GPT-3.5-Turbo-0613 (10.9%) Random Selection 450 9.8 (-)
    GPT-4-0613 (9.3%) Select response with fewer spans 450 5.6 (↓\downarrow 42.9%)
    Select response with zero spans 448 4.8 (↓\downarrow 51.0%)

    Detector-guided selection achieves relative hallucination reductions of 21.6% to 63.2% for open-source models, and 42.9% to 51.0% for high-performing proprietary models (GPT-3.5-turbo and GPT-4).

Coverage note — None was omitted; all key contributions—including the RAGTruth dataset construction, annotation taxonomy, benchmark analysis across models/tasks/lengths/positions, detection baseline evaluations, supervised fine-tuning setup, and mitigation experiments—are fully represented.

References

  1. 1.Ayush Agrawal, Mirac Suzgun, Lester Mackey, and Adam Kalai. 2024. Do language models know when they're hallucinating references? In Findings of the Association for Computational Linguistics: EACL 2024, pages 912–928.
  2. 2.Amos Azaria and Tom Mitchell. 2023. The internal state of an LLM knows when it's lying. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 967–976.
  3. 3.Mario Barrantes, Benedikt Herudek, and Richard Wang. 2020. Adversarial nli for factual correctness in text summarisation models.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901.
  5. 5.Zouying Cao, Yifei Yang, and Hai Zhao. 2023. Autohall: Automated hallucination dataset generation for large language models.
  6. 6.Canyu Chen and Kai Shu. 2024. Can LLM-generated misinformation be detected? In The Twelfth International Conference on Learning Representations.
  7. 7.Shiqi Chen, Yiran Zhao, Jinghan Zhang, I-Chun Chern, Siyang Gao, Pengfei Liu, and Junxian He. 2023. Felm: Benchmarking factuality evaluation of large language models. In Advances in Neural Information Processing Systems, volume 36, pages 44502–44523.
  8. 8.I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu. 2023. Factool: Factuality detection in generative ai – a tool augmented framework for multi-task and multi-domain scenarios.
  9. 9.Roi Cohen, May Hamri, Mor Geva, and Amir Globerson. 2023. LM vs LM: Detecting factual errors via cross examination. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12621–12640.
  10. 10.Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023. Chain-of-verification reduces hallucination in large language models.
  11. 11.Esin Durmus, He He, and Mona Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5055–5070, Online. Association for Computational Linguistics.
  12. 12.Nouha Dziri, Hannah Rashkin, Tal Linzen, and David Reitter. 2022. Evaluating attribution in dialogue systems: The begin benchmark. Transactions of the Association for Computational Linguistics, 10:1066–1083.
  13. 13.Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019. Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2214–2220.
  14. 14.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, and Kelvin Guu. 2023. RARR: Researching and revising what language models say, using language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16477–16508.
  15. 15.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 1321–1330.
  16. 16.Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. Transactions of the Association for Computational Linguistics, 10:178–206.
  17. 17.Xiangkun Hu, Dongyu Ru, Qipeng Guo, Lin Qiu, and Zheng Zhang. 2023. Refchecker for fine-grained hallucination detection.
  18. 18.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12).
  19. 19.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b.
  20. 20.Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. 2023. Challenges and applications of large language models.
  21. 21.Ryo Kamoi, Tanya Goyal, Juan Diego Rodriguez, and Greg Durrett. 2023. WiCE: Real-world entailment for claims in Wikipedia. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7561–7583.
  22. 22.Tom Kocmi and Christian Federmann. 2023. Large language models are state-of-the-art evaluators of translation quality. In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, pages 193–203.
  23. 23.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474.
  24. 24.Junyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023. HaluEval: A large-scale hallucination evaluation benchmark for large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6449–6464.
  25. 25.Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2024. Pre-trained language models for text generation: A survey. ACM Comput. Surv., 56(9).
  26. 26.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  27. 27.Alisa Liu and Jiacheng Liu. 2023. The memo-trap dataset. https://github.com/liujch1998/memo-trap.
  28. 28.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7052–7063.
  29. 29.Andrey Malinin and Mark Gales. 2021. Uncertainty estimation in autoregressive structured prediction. In International Conference on Learning Representations.
  30. 30.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9802–9822.
  31. 31.Potsawee Manakul, Adian Liusie, and Mark Gales. 2023. SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9004–9017.
  32. 32.Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang, Graham Neubig, Yulia Tsvetkov, and Hannaneh Hajishirzi. 2024. Fine-grained hallucination detection and editing for language models.
  33. 33.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A human generated machine reading comprehension dataset. CoRR, abs/1611.09268.
  34. 34.OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O'Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2024. Gpt-4 technical report.
  35. 35.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744.
  36. 36.Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2023. Med-HALT: Medical domain hallucination test for large language models. In Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), pages 314–334.
  37. 37.Vipula Rawte, Amit Sheth, and Amitava Das. 2023. A Survey of Hallucination in Large Foundation Models.
  38. 38.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1073–1083.
  39. 39.Jiaming Shen, Jialu Liu, Dan Finnie, Negar Rahmati, Mike Bendersky, and Marc Najork. 2023. "why is this misleading?": Detecting news headline hallucinations with explanations. In Proceedings of the ACM Web Conference 2023, WWW '23, page 1662–1672.
  40. 40.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3784–3803.
  41. 41.Maxim Tkachenko, Mikhail Malyuk, Andrey Holmanyuk, and Nikolai Liubimov. 2020-2022. Label Studio: Data labeling software. Open source software available from https://github.com/heartexlabs/label-studio.
  42. 42.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  43. 43.Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. 2023. A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation.
  44. 44.Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, and Thang Luong. 2023. Freshllms: Refreshing large language models with search engine augmentation.
  45. 45.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020.
  46. 46.Yijun Xiao and William Yang Wang. 2021. On hallucination and predictive uncertainty in conditional language generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2734–2744.
  47. 47.Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2024. Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs. In The Twelfth International Conference on Learning Representations.
  48. 48.Yelp. 2021. Yelp open dataset. https://www.yelp.com/dataset. Accessed: 2023-11-03.
  49. 49.Hanning Zhang, Shizhe Diao, Yong Lin, Yi R. Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang. 2024. R-tuning: Instructing large language models to say 'i don't know'.
  50. 50.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, Shuming Shi, and Tencent AI Lab. 2023. Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.
  51. 51.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models.
  52. 52.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems, volume 36, pages 46595–46623.
  53. 53.Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev. 2021. QMSum: A new benchmark for query-based multi-domain meeting summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5905–5921, Online. Association for Computational Linguistics.
  54. 54.Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, LILI YU, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. LIMA: Less Is More for Alignment. In Advances in Neural Information Processing Systems, volume 36, pages 55006–55021.

Citation

MLA
Niu, C., et al. “RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 10862–78, https://doi.org/10.18653/v1/2024.acl-long.585.
APA
Niu, C., Wu, Y., Zhu, J., Xu, S., Shum, K., Zhong, R., Song, J., & Zhang, T. (2024). RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10862–10878. https://doi.org/10.18653/v1/2024.acl-long.585
Chicago
Niu, C., Y. Wu, J. Zhu, et al. 2024. “RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10862–78. https://doi.org/10.18653/v1/2024.acl-long.585.
Harvard
Niu, C. et al. (2024) “RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10862–10878. Available at: https://doi.org/10.18653/v1/2024.acl-long.585.
Vancouver
1. Niu C, Wu Y, Zhu J, Xu S, Shum K, Zhong R, Song J, Zhang T (2024) RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10862–10878

BibTeX

@inproceedings{niu-etal-2024-ragtruth,
    title = "{RAGT}ruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models",
    author = "Niu, Cheng  and
      Wu, Yuanhao  and
      Zhu, Juno  and
      Xu, Siliang  and
      Shum, KaShun  and
      Zhong, Randy  and
      Song, Juntong  and
      Zhang, Tong",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.585/",
    doi = "10.18653/v1/2024.acl-long.585",
    pages = "10862--10878"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/