Factuality of Large Language Models: A Survey

Yuxia WangMinghan WangMuhammad Arslan ManzoorFei LiuGeorgi Nenkov GeorgievRocktim Jyoti DasPreslav Nakov

article2024EMNLP77 citations

Synthesizes recent advances in large language model factuality across text and vision modalities by categorizing evaluation benchmarks, clarifying distinctions between factuality and hallucination, and analyzing error-mitigation techniques alongside calibration strategies.

Listen

Large language models are rapidly being integrated into digital assistants and everyday workflows to synthesize information and answer questions directly. However, these systems regularly fabricate ungrounded claims and present incorrect information as fact. Because inaccurate outputs introduce serious operational risks, misinformation, and compliance concerns, establishing reliable factual integrity in model generation has become a critical priority.

The article systematically analyzes the current research landscape surrounding model factuality across text and visual domains. It examines the fundamental causes of factual errors, categorizes benchmarks and evaluation techniques, assesses mitigation strategies across all development stages, and outlines the major obstacles to automated verification.

The authors conducted a comprehensive synthesis of recent literature, organizing evaluation datasets into four operational formats: open-domain generation, binary decisions, short phrases, and multiple-choice questions. They assessed mitigation methods across the entire model lifecycle—including pre-training corpus curation, supervised fine-tuning, preference optimization, decoding adjustments, inference prompting, and post-generation automatic fact-checking—while also examining factuality challenges in multimodal systems.

The findings highlight several core challenges. First, language models are fundamentally trained to optimize sentence probability and fluency rather than objective truth, meaning fluent outputs frequently contain factual inaccuracies. Second, human-annotated error rates on open-ended outputs are notably high, reaching between 42.6% and 68.0% on several major open-ended benchmarks. Third, automated fact-checking systems struggle with accuracy; top-tier verifiers achieve F1 scores of only 0.53 to 0.63 when identifying false claims. Fourth, retrieval augmentation substantially improves factual grounding but introduces latency, requires roughly 25% more computation during pre-training, and remains vulnerable to noisy internet sources. Finally, models are prone to error snowballing, where early inaccuracies compound across subsequent generated text.

These findings indicate that factual errors are not rare edge cases but systemic structural issues inherent in current architectures. Relying on current models for mission-critical, legal, medical, or automated decision-making poses considerable risk without external verification. Furthermore, standard multiple-choice evaluation benchmarks obscure the true severity of errors encountered in realistic, open-ended tasks.

To mitigate these issues, decision-makers should implement layered strategies. Organizations deploying these systems should integrate real-time retrieval-augmented generation and intermediate-sentence verification to prevent error compounding, while accepting modest latency trade-offs. Model developers should adopt behavioral fine-tuning and refusal-aware training so models decline to answer questions outside their knowledge base. Additionally, automated verification pipelines should transition toward smaller, specialized natural language inference models to improve verification reliability and reduce operational costs.

These conclusions should be interpreted in light of current research constraints. Evaluating open-ended text remains inherently uncertain because dividing text into atomic claims and scoring search quality lack universal standards. While confidence is high regarding the conceptual trade-offs and structural causes of hallucinations, rigorous automated measurement of open-ended model factuality remains an active, evolving area of investigation.

arXiv: 2402.02420
Cover for Factuality of Large Language Models: A Survey

Abstract

Large language models (LLMs), especially when instruction-tuned for chat, have become part of our daily lives, freeing people from the process of searching, extracting, and integrating information from multiple sources by offering a straightforward answer to a variety of questions in a single place. Unfortunately, in many cases, LLM responses are factually incorrect, which limits their applicability in real-world scenarios. As a result, research on evaluating and improving the factuality of LLMs has attracted a lot of attention recently. In this survey, we critically analyze existing work with the aim to identify the major challenges and their associated causes, pointing out to potential solutions for improving the factuality of LLMs, and analyzing the obstacles to automated factuality evaluation for open-ended text generation. We further offer an outlook on where future research should go.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Evaluating Factuality
  • 3.1 Datasets and Metrics
  • 3.2 Other Metrics
  • 4 Improving Factuality
  • 4.1 Pre-training
  • 4.2 Tuning and RLXF
  • 4.3 Inference
  • 4.3.1 Decoding Strategy
  • 4.3.2 ICL and Self-reasoning
  • 4.4 Automatic Fact Checkers
  • 5 Factuality of Multimodal LLMs
  • 6 Challenges and Future Directions
  • 7 Conclusion
  • Limitations
  • References

Knowls

  1. Knowl 1 — Conceptual Disambiguation of LLM Factuality, Hallucination, and Trustworthiness

    definition

    In large language model (LLM) research, factuality, hallucination, and trustworthiness represent distinct but interrelated evaluation axes:

    1. Factuality: The ability of an LLM to generate statements that accurately reflect verifiable real-world knowledge and established facts (e.g., drawn from encyclopedias, reference corpora, or domain knowledge bases). Improving factuality corresponds to maximizing the conditional probability P(truth∣prompt)P(\text{truth} \mid \text{prompt}). Factuality errors occur when a model fails to learn, store, or retrieve correct factual knowledge.

    2. Hallucination: The generation of content that diverges from the user prompt, contradicts provided input context, contradicts previously generated output within the same response, or diverges from objective reality. Hallucinations are split into:

      • Faithfulness hallucination: Input-conflicting, context-conflicting, or logically inconsistent generations.
      • Factuality hallucination: Discrepancies with verifiable external reality, subdivided into factual inconsistency (direct contradiction of real-world facts) and factual fabrication (making up non-existent entities or events).

    A generation can be non-hallucinatory but factually inaccurate (when faithfully reproducing an inaccurate source document), or factually accurate while being a hallucination (producing off-topic or unprompted but true statements).

    1. Trustworthiness: A multi-dimensional reliability paradigm extending beyond factual accuracy across eight distinct axes: truthfulness, safety, fairness, robustness, privacy, ethics, transparency, and accountability.
  2. Knowl 2 — Benchmark Categorization and Taxonomy for LLM Factuality Evaluation

    data/table

    LLM factuality benchmarks are categorized into four distinct types based on their output space and the feasibility of automated evaluation:

    • Type I (Open-domain, free-form, long-form text generation): Most realistic yet hardest to evaluate automatically. Models generate multi-sentence passages (e.g., biographies) evaluated via atomic claim decomposition and verification.
    • Type II (Yes/No response): Constrained binary classification evaluated by Exact Match, Accuracy, and F1 score.
    • Type III (Short-form phrase or entity list generation): Constrained generation evaluated by Exact Match, Accuracy, or Precision/Recall@KK.
    • Type IV (Multiple-choice QA): Automation-friendly discrimination benchmarks evaluated by standard multi-class classification accuracy or F1 score.
    Type Dataset Topic / Task Size Error Rate (%) Evaluation Metric / Method
    I FactScore-Bio Biography 549 42.6 Human annotation auto fact-checkers
    I Factcheck-GPT Open-ended questions 94 64.9 Human annotation
    I FacTool-QA Knowledge-based QA 50 54.0 Human annotation auto fact-checkers
    I FELM-WK Knowledge-based QA 184 46.2 Human annotation, Accuracy, F1
    I HaluEval Open-ended questions 5000 12.3 Human annotation, AUROC, LLM judge, PARENT
    I FreshQA Open-ended questions 499 68.0 Human annotation
    I SelfAware Open-ended questions 3369 – Awareness of unknown by F1-score
    II Snowball Yes/No questions 1500 9.4 Exact match, Accuracy, F1-score
    III Wiki-category List Name entity lists 55 – Precision@5, Recall@5
    III Multispan QA Short-term answer 428 – Exact match, F1-score
    IV TruthfulQA False beliefs / misconceptions 817 – Accuracy
    IV HotpotQA Multi-step reasoning 113,000 – Exact match, F1-score
    IV StrategyQA Multi-step reasoning 2780 – Recall@10
    IV MMLU Multi-domain knowledge 15,700 – Accuracy

    Human-annotated error rates on open-ended generation benchmarks (Type I) reveal high factual error frequencies in commercial LLMs, ranging from 12.3% on HaluEval up to 68.0% on FreshQA.

  3. Knowl 3 — Quantitative Evaluation Metrics for Open-Ended LLM Factuality

    model/method

    Evaluating open-ended LLM generations relies on specific metrics tailored for claim-level and document-level veracity:

    1. FactScore: Decomposes a generated text document DD into a set of NN discrete atomic claims C={c1,c2,…,cN}C = \{c_1, c_2, \dots, c_N\}. Each claim cic_i is individually checked against reliable external sources (e.g., Wikipedia). The document-level score is the proportion of supported claims: FactScore(D)=1N∑i=1NI(ci is supported by evidence)\text{FactScore}(D) = \frac{1}{N} \sum_{i=1}^N \mathbb{I}(c_i \text{ is supported by evidence}) where I(⋅)\mathbb{I}(\cdot) is the indicator function.

    2. Hallucinated Named Entities Error: Measures entity fabrication relative to a reference document DrefD_{\text{ref}}: ErrorNE=∣Entities(Dgen)∖Entities(Dref)∣∣Entities(Dgen)∣\text{Error}_{\text{NE}} = \frac{|\text{Entities}(D_{\text{gen}}) \setminus \text{Entities}(D_{\text{ref}})|}{|\text{Entities}(D_{\text{gen}})|}

    3. Entailment Ratio: Computes the fraction of generated units (sentences or claims) that are logically entailed by ground-truth reference material: Entailment Ratio=∑s∈DgenI(Dref⊨s)∣Dgen∣\text{Entailment Ratio} = \frac{\sum_{s \in D_{\text{gen}}} \mathbb{I}(D_{\text{ref}} \models s)}{|D_{\text{gen}}|}

    4. Hallucination Vulnerability Index (HVI): A composite scoring metric used to quantify and rank an LLM's overall susceptibility to producing ungrounded assertions across diverse prompt formulations.

  4. Knowl 4 — Pre-Training and Fine-Tuning Interventions for Factuality Enhancement

    model/method

    LLM factuality can be improved during training and alignment stages using specific data curation and parameter-tuning techniques:

    1. Pre-Training Corpus Engineering:

      • Quality Filtering: Automated filtering of web scrapes (e.g., CommonCrawl) based on similarity to high-quality reference corpora.
      • Reliability Up-sampling: Increasing the sampling probability of verifiable documents from authoritative sources.
      • Synthetic Textbooks: Generating dense, reasoning-rich synthetic datasets (e.g., phi-1.5 pre-training data).
      • Retrieval-Augmented Pre-training: Integrating an explicit retrieval database during autoregressive pre-training (e.g., RETRO). While this improves factual accuracy and lowers hallucination rates, it incurs approximately a 25%25\% increase in training compute and remains vulnerable to outdated or biased external database content.
    2. Supervised Fine-Tuning (SFT) and RLXF:

      • Knowledge Injection (KI): Intermediate or combined tuning on entity triplets and summaries.
      • Refusal-Aware Tuning (R-Tuning): Identifies the knowledge boundary between the model's parametric knowledge and instruction data, training the LLM to refuse queries outside its parametric scope, mitigating behavior cloning where models guess answers to unfamiliar queries.
      • Behavioral Fine-Tuning (BeInfo): Trains for selectivity (extracting valid info) and response adequacy (acknowledging lack of evidence).
      • Sycophancy Mitigation: Using synthetic preference interventions to decouple factual correctness from user opinion alignment.
      • Direct Preference Optimization (DPO): Optimizing LLM parameters using preference rankings generated from automatic fact-checker veracity scores or model confidence metrics.
      • Hallucination-Augmented Recitations (HAR): Fine-tuning on counterfactual datasets to force attribution to context rather than parametric memory.
  5. Knowl 5 — Inference-Time Decoding and Reasoning Strategies for Factuality Enhancement

    model/method

    At inference time, factuality can be enhanced without parameter updates through decoding modifications and structured prompting:

    1. Decoding Strategies:

      • Factual-Nucleus Sampling: Standard top-pp nucleus sampling increases randomness and factual errors toward the end of generated sequences. Factual-nucleus sampling dynamically decreases the top-pp probability threshold as sequence generation progresses, constraining randomness in later tokens.
      • Context-Aware Decoding (CAD): Amplifies context reliance by contrasting token output distributions conditioned with context against unconditioned prior distributions using contrastive ensemble logits: PCAD(yt∣y<t,x)∝P(yt∣y<t,x)P(yt∣y<t)αP_{\text{CAD}}(y_t \mid y_{<t}, x) \propto \frac{P(y_t \mid y_{<t}, x)}{P(y_t \mid y_{<t})^{\alpha}} where α\alpha is a weighting hyperparameter.
      • Decoding by Contrasting Layers (DoLa): Selects premature intermediate transformer layers containing weaker factual representations and contrasts their logits against the final layer output, isolating factual knowledge without external retrieval. This introduces a 1.01×1.01\times to 1.08×1.08\times decoding latency increase.
    2. Prompting and Reasoning Strategies:

      • In-Context Knowledge Editing (ICL): Injects updated facts via few-shot demonstrations to override outdated parametric memory.
      • Multi-Agent Debate: Multiple LLM instances debate and critique draft candidate answers across successive rounds until consensus is reached, where increasing agent count and debate length correlates with improved factuality.
      • Iterative Generation-Verification: Sentence-by-sentence verification and rectification during generation (e.g., EVER), preventing downstream hallucination snowballing at the cost of increased real-time response latency.
  6. Knowl 6 — Architectural Pipeline and Operational Bottlenecks of Automated Fact-Checkers

    model/method

    Automated fact-checking pipelines evaluate and rectify generated text through a three-component architecture:

    1. Claim Processor: Takes raw generated text and executes document decomposition, decontextualization (resolving coreference and contextual dependencies), and check-worthiness filtering to extract a set of independent, verifiable atomic claims.
    2. Retriever: Queries external web search engines or internal corpora (e.g., Wikipedia), reranking and optionally summarizing the retrieved passages into relevant evidence sets.
    3. Verifier: Employs Natural Language Inference (NLI) models or LLMs (e.g., GPT-4) to classify each claim as true or false based on the retrieved evidence, generating optional natural language explanations and rectifying false claims via targeted editing.
    [Document] -> [Claim Processor: Decompose -> Decontextualize -> Checkworthiness]
                       |
                       v  (List of checkworthy atomic claims)
                  [Retriever: Search -> Rerank -> Summarize]
                       |
                       v  (Set of relevant evidence)
                  [Verifier: Verify -> Explain -> Edit] -> [Document True/False / Corrected Text]
    

    Operational Bottlenecks:

    • Absence of intermediate gold standards: No objective ground truth exists for optimal atomic claim granularity or open-web retrieved evidence sets, preventing standalone optimization of claim processors and retrievers.
    • Verifier performance upper bounds: Automated verifiers remain imperfect; state-of-the-art verifiers using GPT-4 with Google Search achieve only an F1 score of 0.630.63 in detecting false claims (F1=0.53F_1 = 0.53 with PerplexityAI) relative to human ground truth.
    • Error propagation: Errors in claim extraction or evidence retrieval cascade directly into verifier misclassifications.
  7. Knowl 7 — Taxonomy, Evaluation, and Mitigation of Multimodal LLM (MLLM) Factuality

    definition

    Factuality in Multimodal Large Language Models (MLLMs) measures the consistency between generated textual responses and associated visual inputs. Multimodal factuality errors are categorized into three classes:

    1. Existence Factuality: Falsely claiming that non-existent objects, persons, or entities are present in an image.
    2. Attribute Factuality: Correctly identifying an object's presence but misidentifying its visual attributes (e.g., color, size, shape, material).
    3. Relationship Factuality: Correctly identifying objects but incorrectly describing spatial relations, interactions, or structural connections between them.

    Evaluation Benchmarks:

    • CHAIR: Evaluates object hallucination in image captions against predefined MS COCO objects.
    • POPE: Probes visual hallucinations via balanced binary choice (Yes/No) questions regarding object presence.
    • GAVIE: Uses GPT-4-assisted visual instruction evaluation to score visual hallucinations without human intervention.

    Mitigation Techniques:

    • Supervised Alignment: Robust instruction tuning using positive/negative instruction pairs (LRV-Instruction) or preference optimization with factually augmented RLHF (LLaVA-RLHF).
    • Post-Generation Rectification: Specialized visual correction models (e.g., Woodpecker, LURE) that parse text, query visual detectors, and rewrite inconsistent spans.
    • Representation Learning / Decoding: Contrastive visual decoding and feature regularization (HallE-Switch, VCD, HACL) to suppress language prior bias over visual tokens.
  8. Knowl 8 — Core Technical Challenges and Bottlenecks in LLM Factuality

    limitation

    Three fundamental technical limitations constrain the factuality of current large language models:

    1. Mismatched Training Objective: Autoregressive language modeling maximizes conditional token log-likelihood over a linguistic corpus: max⁡θ∑tlog⁡Pθ(wt∣w<t)\max_\theta \sum_t \log P_\theta(w_t \mid w_{<t}). The objective optimizes fluency and statistical likelihood rather than truth value or factual correctness, ensuring that convergence does not imply factual veracity.

    2. Evaluation Ground-Truth Circularity: Evaluating open-ended text factuality requires automated fact-checkers, yet evaluating and tuning automated fact-checkers requires annotated open-ended veracity datasets, creating a circular dependency. Furthermore, leading LLM verifiers are computationally expensive, sensitive to minor prompt variations, and lack systematic evaluation on intermediate processing steps.

    3. Retrieval Bottlenecks in RAG: Retrieval-Augmented Generation suffers from:

      • Information contamination: Susceptibility to false, outdated, or adversarial content retrieved from open web environments.
      • Multi-hop reasoning deficits: Failure of dense and sparse retrievers to aggregate scattered, interdependent evidence across multiple documents.
      • Inference latency: Multi-step retrieval and real-time generation-verification cycles (e.g., sentence-level error checking) introduce computational and latency overheads that hinder deployment in high-throughput interactive systems.

Coverage note — Omitted Table 1 from the paper (meta-comparison of prior hallucination surveys) as it represents a high-level review of existing literature rather than a substantive technical contribution of this survey. All original taxonomies, lifecycle analyses, benchmark classifications, fact-checker mechanics, multimodal frameworks, and limitation analyses are fully covered.

References

  1. 1.Sebastian Borgeaud, Arthur Mensch, and Jordan Hoffmann et al. 2021. Improving language models by retrieving from trillions of tokens. In ICML.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, and Melanie Subbiah et al. 2020. Language models are few-shot learners. In NeurIPS 2020.
  3. 3.Shiqi Chen, Yiran Zhao, Jinghan Zhang, and et al. 2023. Felm: Benchmarking factuality evaluation of large language models. arXiv preprint arXiv:2310.00741.
  4. 4.I-Chun Chern, Steffi Chern, and Shiqi Chen et al. 2023. Factool: Factuality detection in generative AI - A tool augmented framework for multi-task and multi-domain scenarios. CoRR, abs/2307.13528.
  5. 5.Yung-Sung Chuang, Yujia Xie, and Hongyin Luo et al. 2023. Dola: Decoding by contrasting layers improves factuality in large language models. CoRR, abs/2309.03883.
  6. 6.Shehzaad Dhuliawala, Mojtaba Komeili, and et al. 2023. Chain-of-verification reduces hallucination in large language models. arXiv preprint arXiv:2309.11495.
  7. 7.Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2023. Improving factuality and reasoning in language models through multiagent debate. CoRR, abs/2305.14325.
  8. 8.Mohamed Elaraby, Mengyin Lu, and Jacob Dunn et al. 2023. Halo: Estimation and reduction of hallucinations in open-source weak large language models. CoRR, abs/2308.11764.
  9. 9.Luyu Gao, Zhuyun Dai, and Panupong et al. Pasupat. 2022. Attributed text generation via post-hoc research and revision. arXiv preprint arXiv:2210.08726.
  10. 10.Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023a. Enabling large language models to generate text with citations. In EMNLP, pages 6465–6488.
  11. 11.Yunfan Gao, Yun Xiong, and et al. 2023b. Retrieval-augmented generation for large language models: A survey. CoRR, abs/2312.10997.
  12. 12.Mor Geva, Daniel Khashabi, and et al. 2021. Did aristotle use a laptop? A question answering benchmark with implicit reasoning strategies. TACL, 9:346–361.
  13. 13.Anish Gunjal, Jihan Yin, and Erhan Bas. 2023. Detecting and preventing hallucinations in large vision language models. In AAAI Conference on Artificial Intelligence.
  14. 14.Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. TACL, 10:178–206.
  15. 15.Dan Hendrycks, Collin Burns, and et al. 2021. Measuring massive multitask language understanding. In ICLR 2021.
  16. 16.Ari Holtzman, Jan Buys, and Li et al. 2020. The curious case of neural text degeneration. In ICLR.
  17. 17.Lei Huang, Weijiang Yu, Weitao Ma, and Weihong Zhong et al. 2023a. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. CoRR, abs/2311.05232.
  18. 18.Lei Huang, Weijiang Yu, Weitao Ma, and Weihong Zhong et al. 2023b. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. CoRR, abs/2311.05232.
  19. 19.Gautier Izacard and Edouard Grave. 2020. Leveraging passage retrieval with generative models for open domain question answering. ArXiv, abs/2007.01282.
  20. 20.Ziwei Ji, Nayeon Lee, and Rita Frieske et al. 2023. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12):248:1–248:38.
  21. 21.Chaoya Jiang, Haiyang Xu, Mengfan Dong, Jiaxing Chen, Wei Ye, Mingshi Yan, Qinghao Ye, Ji Zhang, Fei Huang, and Shikun Zhang. 2023. Hallucination augmented contrastive learning for multimodal large language model. ArXiv, abs/2312.06968.
  22. 22.Haoqiang Kang, Juntong Ni, and Huaxiu Yao. 2023. Ever: Mitigating hallucination in large language models through real-time verification and rectification. CoRR, abs/2311.09114.
  23. 23.Vladimir Karpukhin, Barlas Oguz, and Sewon Min et al. 2020. Dense passage retrieval for open-domain question answering. ArXiv, abs/2004.04906.
  24. 24.Abdullatif Köksal, Renat Aksitov, and Chung-Ching Chang. 2023. Hallucination augmented recitations for language models. arXiv preprint arXiv:2311.07424.
  25. 25.Nayeon Lee, Wei Ping, and Peng et al. Xu. 2022. Factuality enhanced language models for open-ended text generation. NeuralPS, 35:34586–34599.
  26. 26.Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Li Bing. 2023. Mitigating object hallucinations in large vision-language models through visual contrastive decoding. ArXiv, abs/2311.16922.
  27. 27.Patrick Lewis, Ethan Perez, and et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. ArXiv, abs/2005.11401.
  28. 28.Junyi Li and Xiaoxue Cheng et al. 2023a. Halueval: A large-scale hallucination evaluation benchmark for large language models. CoRR, abs/2305.11747.
  29. 29.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji rong Wen. 2023. Evaluating object hallucination in large vision-language models. In Conference on Empirical Methods in Natural Language Processing.
  30. 30.Yuanzhi Li and Sébastien Bubeck et al. 2023b. Textbooks are all you need II: phi-1.5 technical report. CoRR, abs/2309.05463.
  31. 31.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Truthfulqa: Measuring how models mimic human falsehoods. In ACL, pages 3214–3252.
  32. 32.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European Conference on Computer Vision.
  33. 33.Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. 2023. Mitigating hallucination in large multi-modal models via robust instruction tuning.
  34. 34.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in gpt. In Neural Information Processing Systems.
  35. 35.Sewon Min and Kalpesh Krishna et al. 2023. Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. CoRR, abs/2305.14251.
  36. 36.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. 2021. Fast model editing at scale. ArXiv, abs/2110.11309.
  37. 37.Dor Muhlgay, Ori Ram, and Inbal Magar et al. 2023. Generating benchmarks for factuality evaluation of language models. CoRR, abs/2307.06908.
  38. 38.Reiichiro Nakano, Jacob Hilton, and et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. ArXiv, abs/2112.09332.
  39. 39.Long Ouyang, Jeff Wu, and Xu Jiang et al. 2022. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155.
  40. 40.Rafael Rafailov, Archit Sharma, and et al. 2023. Direct preference optimization: Your language model is secretly a reward model. CoRR, abs/2305.18290.
  41. 41.Vipula Rawte, Swagata Chakraborty, and Agnibh et al. Pathak. 2023a. The troubling emergence of hallucination in large language models - an extensive definition, quantification, and prescriptive remediations. In EMNLP 2023, pages 2541–2573.
  42. 42.Vipula Rawte, Amit P. Sheth, and Amitava Das. 2023b. A survey of hallucination in large foundation models. CoRR, abs/2309.05922.
  43. 43.Evgeniia Razumovskaia, Ivan Vulic, and Pavle Markovic et al. 2023. Dial beinfo for faithfulness: Improving factuality of information-seeking dialogue via behavioural fine-tuning. CoRR, abs/2311.09800.
  44. 44.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018. Object hallucination in image captioning. In Conference on Empirical Methods in Natural Language Processing.
  45. 45.Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, and Amanda Askell et al. 2023. Towards understanding sycophancy in language models. CoRR, abs/2310.13548.
  46. 46.Weijia Shi, Xiaochuang Han, and et al. 2023. Trusting your evidence: Hallucinate less with context-aware decoding. arXiv preprint arXiv:2305.14739.
  47. 47.Noah Shinn, Federico Cassano, and Gopinath et al. 2023. Reflexion: Language agents with verbal reinforcement learning. In NeuralPS.
  48. 48.Lichao Sun, Yue Huang, and Haoran Wang et al. 2024. Trustllm: Trustworthiness in large language models. ArXiv, abs/2401.05561.
  49. 49.Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liangyan Gui, Yu-Xiong Wang, Yiming Yang, Kurt Keutzer, and Trevor Darrell. 2023. Aligning large multimodal models with factually augmented rlhf. ArXiv, abs/2309.14525.
  50. 50.James Thorne, Andreas Vlachos, and et al. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In NAACL, pages 809–819.
  51. 51.Katherine Tian, Eric Mitchell, and et al. 2023. Fine-tuning language models for factuality. arXiv preprint arXiv:2311.08401.
  52. 52.S. M. Towhidul Islam Tonmoy, S. M. Mehedi Zaman, and et al. 2024. A comprehensive survey of hallucination mitigation techniques in large language models. CoRR, abs/2401.01313.
  53. 53.Faraz Torabi, Garrett Warnell, and Peter Stone. 2018. Behavioral cloning from observation. In IJCAI, pages 4950–4957. ijcai.org.
  54. 54.Hugo Touvron, Louis Martin, and et al. 2023. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  55. 55.Neeraj Varshney, Wenlin Yao, and et al. 2023. A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation. CoRR, abs/2307.03987.
  56. 56.Tu Vu, Mohit Iyyer, and et al. 2023. Freshllms: Refreshing large language models with search engine augmentation. arXiv preprint arXiv:2310.03214.
  57. 57.Boxin Wang, Wei Ping, and et al. 2023a. Shall we pre-train autoregressive language models with retrieval? a comprehensive study. In EMNLP.
  58. 58.Cunxiang Wang, Xiaoze Liu, and et al. 2023b. Survey on factuality in large language models: Knowledge, retrieval and domain-specificity. ArXiv, abs/2310.07521.
  59. 59.Yuxia Wang, Revanth Gangi Reddy, and et al. 2023c. Factcheck-gpt: End-to-end fine-grained document-level fact-checking and correction of LLM output. CoRR, abs/2311.09000.
  60. 60.Jason Wei, Xuezhi Wang, and Dale et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS 2022.
  61. 61.Jerry W. Wei, Da Huang, and Yifeng Lu et al. 2023. Simple synthetic data reduces sycophancy in large language models. CoRR, abs/2308.03958.
  62. 62.Zhilin Yang and Peng Qi et al. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In EMNLP 2018, pages 2369–2380.
  63. 63.Shunyu Yao, Jeffrey Zhao, and Dian et al. 2023. React: Synergizing reasoning and acting in language models. In ICLR.
  64. 64.Shukang Yin, Chaoyou Fu, and et al. 2023a. Woodpecker: Hallucination correction for multimodal large language models. CoRR, abs/2310.16045.
  65. 65.Zhangyue Yin, Qiushi Sun, and Qipeng Guo et al. 2023b. Do large language models know what they don’t know? In ACL, pages 8653–8665.
  66. 66.Bohan Zhai, Shijia Yang, Xiangchen Zhao, Chenfeng Xu, Sheng Shen, Dongdi Zhao, Kurt Keutzer, Manling Li, Tan Yan, and Xiangjun Fan. 2023. Halle-switch: Rethinking and controlling object existence hallucinations in large vision language models for detailed caption. ArXiv, abs/2310.01779.
  67. 67.Hanning Zhang, Shizhe Diao, and et al. 2023a. R-tuning: Teaching large language models to refuse unknown questions. CoRR, abs/2311.09677.
  68. 68.Muru Zhang, Ofir Press, and et al. 2023b. How language model hallucinations can snowball. CoRR, abs/2305.13534.
  69. 69.Yue Zhang, Yafu Li, and et al. 2023c. Siren’s song in the AI ocean: A survey on hallucination in large language models. CoRR, abs/2309.01219.
  70. 70.Ce Zheng, Lei Li, and et al. 2023. Can we edit factual knowledge by in-context learning? In EMNLP, pages 4862–4876.
  71. 71.Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, and Huaxiu Yao. 2023. Analyzing and mitigating object hallucination in large vision-language models. ArXiv, abs/2310.00754.

Citation

MLA
Wang, Y., et al. “Factuality of Large Language Models: A Survey”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 19519–29, https://doi.org/10.18653/v1/2024.emnlp-main.1088.
APA
Wang, Y., Wang, M., Manzoor, M. A., Liu, F., Georgiev, G. N., Das, R. J., & Nakov, P. (2024). Factuality of Large Language Models: A Survey. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 19519–19529. https://doi.org/10.18653/v1/2024.emnlp-main.1088
Chicago
Wang, Y., M. Wang, M. A. Manzoor, et al. 2024. “Factuality of Large Language Models: A Survey”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 19519–29. https://doi.org/10.18653/v1/2024.emnlp-main.1088.
Harvard
Wang, Y. et al. (2024) “Factuality of Large Language Models: A Survey”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 19519–19529. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.1088.
Vancouver
1. Wang Y, Wang M, Manzoor MA, Liu F, Georgiev GN, Das RJ, Nakov P (2024) Factuality of Large Language Models: A Survey. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 19519–19529

BibTeX

@inproceedings{wang-etal-2024-factuality,
    title = "Factuality of Large Language Models: A Survey",
    author = "Wang, Yuxia  and
      Wang, Minghan  and
      Manzoor, Muhammad Arslan  and
      Liu, Fei  and
      Georgiev, Georgi Nenkov  and
      Das, Rocktim Jyoti  and
      Nakov, Preslav",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1088/",
    doi = "10.18653/v1/2024.emnlp-main.1088",
    pages = "19519--19529"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/