Survey of Hallucination in Natural Language Generation

Ziwei JiNayeon LeeRita FrieskeTiezheng YuDan SuYan XuEtsuko IshiiYe Jin BangAndrea MadottoPascale Fung

article2022CSUR5,830 citations

Systematizes the evaluation metrics and mitigation techniques for hallucinations across diverse natural language generation tasks and large language models to guide the development of factually reliable systems.

Listen

The article addresses the growing problem of hallucination in natural language generation systems, where models produce fluent but unfaithful or nonsensical text. This issue degrades performance in applications such as summarization, dialogue, and machine translation while raising safety risks in domains like medicine and privacy. The problem has become more pressing with advances in transformer-based models that improve fluency yet remain prone to generating unsupported content.

The survey sets out to evaluate the full scope of hallucination research across NLG, offering definitions, causes, metrics, and mitigation strategies while covering both general principles and task-specific progress. It reviews literature on abstractive summarization, dialogue generation, question answering, data-to-text, translation, vision-language tasks, and large language models.

The authors conduct a systematic literature review organized into general sections on definitions, data and training contributors, evaluation metrics, and mitigation methods, followed by detailed task analyses. They draw on dozens of studies, highlight common patterns, and update sections on large language models to reflect recent developments.

Three core findings stand out: hallucinations divide into intrinsic types that contradict the source and extrinsic types unverifiable from it; they arise from source-reference mismatches in data as well as from model architecture, decoding strategies, and exposure bias during training; and existing metrics such as ROUGE or BLEU correlate poorly with human judgments of faithfulness, prompting new statistical, model-based, and human evaluation approaches. Mitigation succeeds most when combining data cleaning, architectural changes, reinforcement learning, and post-processing.

These results indicate that unchecked hallucinations increase operational risk, limit deployment in high-stakes settings, and require coordinated progress across metrics and methods. They also show that extrinsic hallucinations remain especially difficult to detect and correct because they may draw on external facts.

Next steps include development of fine-grained metrics that separate hallucination types, automated fact-checking pipelines for extrinsic cases, and methods that explicitly model numerals and long-context reasoning. Task-specific work should prioritize dialogue summarization and logical data-to-text generation, while controllability techniques can help balance faithfulness against diversity.

The survey relies on published studies through mid-2024 and notes that definitions and evaluation standards still vary across tasks. Its conclusions are therefore most reliable for well-studied tasks and should be treated cautiously for rapidly evolving large language model applications where new evidence continues to emerge.

Cover for Survey of Hallucination in Natural Language Generation

Abstract

Natural Language Generation (NLG) has improved exponentially in recent years thanks to the development of sequence-to-sequence deep learning technologies such as Transformer-based language models. This advancement has led to more fluent and coherent NLG, leading to improved development in downstream tasks such as abstractive summarization, dialogue generation and data-to-text generation. However, it is also apparent that deep learning based generation is prone to hallucinate unintended text, which degrades the system performance and fails to meet user expectations in many real-world scenarios. To address this issue, many studies have been presented in measuring and mitigating hallucinated texts, but these have never been reviewed in a comprehensive manner before. In this survey, we thus provide a broad overview of the research progress and challenges in the hallucination problem in NLG. The survey is organized into two parts: (1) a general overview of metrics, mitigation methods, and future directions; (2) an overview of task-specific research progress on hallucinations in the following downstream tasks, namely abstractive summarization, dialogue generation, generative question answering, data-to-text generation, machine translation, and visual-language generation; and (3) hallucinations in large language models (LLMs). This survey serves to facilitate collaborative efforts among researchers in tackling the challenge of hallucinated texts in NLG.

Table of Contents

  • 1 Introduction
  • 2 Definitions
  • 2.1 Categorization
  • 2.2 Task Comparison
  • 2.3 Terminology Clarification
  • 3 Contributors to Hallucination in NLG
  • 3.1 Hallucination from Data
  • 3.2 Hallucination from Training and Inference
  • 4 Metrics Measuring Hallucination
  • 4.1 Statistical Metric
  • 4.2 Model-based Metric
  • 4.2.1 Information Extraction (IE)-based
  • 4.2.2 QA-based
  • 4.2.3 Natural Language Inference (NLI) Metrics
  • 4.2.4 Faithfulness Classification Metrics
  • 4.2.5 LM-based Metrics
  • 4.3 Human Evaluation
  • 5 Hallucination Mitigation Methods
  • 5.1 Data-Related Methods
  • 5.1.1 Building a Faithful Dataset
  • 5.1.2 Cleaning Data Automatically
  • 5.1.3 Information Augmentation
  • 5.2 Modeling and Inference Methods
  • 5.2.1 Architecture
  • 5.2.2 Training
  • 5.2.3 Post-Processing
  • 6 Future Directions
  • 6.1 Future Directions in Metrics Design
  • 6.2 Future Directions in Mitigation Methods
  • 7 Hallucination in Abstractive Summarization
  • 7.1 Hallucination Definition in Abstractive Summarization
  • 7.2 Hallucination Metrics in Abstractive Summarization
  • 7.2.1 Unsupervised Metrics
  • 7.2.2 Semi-Supervised Metrics
  • 7.3 Hallucination Mitigation in Abstractive Summarization
  • 7.3.1 Architecture Method.
  • 7.3.2 Training Method
  • 7.3.3 Post-Processing Method
  • 7.4 Future Directions in Abstractive Summarization
  • 8 Hallucination in Dialogue Generation
  • 8.1 Hallucination Definition in Dialogue Generation
  • 8.2 Open-domain Dialogue Generation
  • 8.2.1 Self-Consistency
  • 8.2.2 External Consistency
  • 8.2.3 Hallucination Metrics
  • 8.2.4 Mitigation Methods
  • 8.3 Task-oriented Dialogue Generation
  • 8.3.1 Hallucination Metrics
  • 8.3.2 Mitigation Methods
  • 8.4 Future Directions in Dialogue Generation
  • 9 Hallucination in Generative Question Answering
  • 9.1 Hallucination Definition in GQA
  • 9.2 Hallucination-related Metrics in GQA
  • 9.3 Hallucination Mitigation in GQA
  • 9.4 Future Directions in GQA
  • 10 Hallucination in Data-to-Text Generation
  • 10.1 Hallucination Definition in Data-to-Text Generation
  • 10.2 Hallucination Metrics in Data-to-Text Generation
  • 10.3 Hallucination Mitigation in Data-to-Text Generation
  • 10.4 Future Directions in Data-to-Text Generation
  • 11 Hallucinations in Neural Machine Translation
  • 11.1 Hallucinations Definition and Categories in NMT
  • 11.2 Hallucination Metrics in NMT
  • 11.2.1 Statistical Metrics
  • 11.2.2 Model-Based Metrics
  • 11.3 Hallucination Mitigation Methods in NMT
  • 11.3.1 Data-Related
  • 11.3.2 Modeling and Inference
  • 11.4 Future Directions in NMT
  • 12 Hallucination in Vision-Language Generation
  • 12.1 Object Hallucination in Image Captioning
  • 12.2 Hallucination in Other VL Tasks
  • 12.3 Future Directions in VL
  • 13 Hallucination in Large Language Models
  • 13.1 Hallucination Definition in LLMs
  • 13.2 Hallucination Metrics for LLMs
  • 13.2.1 Reference-dependent Metrics
  • 13.2.2 Reference-free Metrics
  • 13.3 Hallucination Mitigation in LLMs
  • 13.3.1 Data-Related Methods
  • 13.3.2 Modelling and Inference Methods
  • 13.4 Future Directions of Hallucination Mitigation in LLMs
  • 13.4.1 Hallucination in Large Multimodal Models
  • 13.4.2 Hallucination in Long-tail and Low-resource Domains
  • 13.4.3 Estimating the Knowledge Boundary and Expressing the Uncertainty
  • 13.4.4 Minimizing the Alignment Tax During Hallucination Mitigation
  • 13.4.5 Understanding Hallucination in LLMs
  • 14 Conclusion
  • References

Knowls

  1. Knowl 1 — Categorization of Hallucination in Natural Language Generation

    definition

    In Natural Language Generation (NLG), hallucination is defined as generated text that is nonsensical or unfaithful to the provided source content. Hallucinations are divided into two primary categories:

    • Intrinsic Hallucination: Generated text that directly contradicts information present in the source context. For example, if the source states that a vaccine was approved in 2019, generating an output stating the vaccine was approved in 2021 constitutes an intrinsic hallucination.
    • Extrinsic Hallucination: Generated text containing information that can neither be verified nor contradicted by the source context because the content is entirely unmentioned. Extrinsic hallucinations may still be factually true in the real world (representing background knowledge), but they remain unverifiable with respect to the input source and introduce factual safety risks.

    Terminology distinction:

    • Faithfulness: Consistency and truthfulness of the generated output strictly with respect to the provided source text.
    • Factuality: Truthfulness of the generated output with respect to world knowledge (actual real-world facts).
  2. Knowl 2 — Causes and Contributors to Hallucination in NLG

    definition

    Hallucination in natural language generation arises from four interrelated factors spanning datasets, modeling representations, training dynamics, and decoding mechanisms:

    1. Data-Related Contributors:

      • Source-Reference Divergence: Heuristic data collection frequently pairs target text that includes information unsupported by the source input (for instance, in the WikiBio dataset, 62% of introductory sentences contain extra information absent from the infobox tables).
      • Corpus Duplication: Duplicated sequences in pre-training data bias the language model toward memorizing and re-generating specific phrases regardless of source context.
      • Innate Task Divergence: Conversational or open-ended tasks naturally permit subjectivity and chit-chat, embedding reference targets that diverge from source inputs.
    2. Training and Inference Contributors:

      • Defective Representation Learning: Encoders that learn spurious correlations between tokens fail to construct faithful semantic representations of input documents.
      • Erroneous Attention and Decoding: Decoders that misallocate cross-attention weights confuse relations among entities. Furthermore, stochastic decoding algorithms designed to increase diversity (such as top-kk sampling) increase the unexpectedness of token sequences and correlate positively with higher hallucination rates.
      • Exposure Bias: The discrepancy between training with teacher forcing (conditioning next-token prediction on ground-truth prefixes) and inference generation (conditioning on model-generated tokens) causes compounding errors over longer sequences.
      • Parametric Knowledge Bias: High-capacity pre-trained models prioritize memorized parametric knowledge from pre-training over the explicit information provided in the input prompt.
  3. Knowl 3 — Taxonomy of Automatic Hallucination Evaluation Metrics

    model/method

    Automatic metrics for evaluating hallucination in NLG fall into two overarching paradigms: statistical metrics and model-based metrics.

    1. Statistical Metrics: Evaluate lexical overlap (nn-grams) and token-level alignments between the generated output and reference or source texts (e.g., PARENT, PARENT-T, Knowledge F1, BVSS). While simple and computationally efficient, they cannot assess semantic paraphrases or syntactic variability.

    2. Model-Based Metrics:

      • Information Extraction (IE)-Based: Information extraction models parse both source and generated texts into relational tuples (subject,relation,object)(subject, relation, object). The metric computes the ratio of generated relational tuples supported by the source knowledge tuples.
      • Question Answering (QA)-Based: A Question Generation (QG) model generates question-answer pairs from the generated candidate text, and a Question Answering (QA) model answers those questions using the source text as context. Faithfulness is scored by the semantic similarity between the candidate answers and the QA model answers (e.g., FEQA, QAGS, QuestEval).
      • Natural Language Inference (NLI)-Based: Formulates hallucination evaluation as a premise-hypothesis entailment task. The source text serves as the premise and the generated text serves as the hypothesis; the metric calculates the entailment probability across sentences or dependency graphs (e.g., SummaC).
      • Language Model (LM)-Based: Measures token-level support by evaluating the cross-entropy loss difference between an unconditional language model and a conditional language model during forced-path decoding.
      • Task-Specific Faithfulness Classifiers: Supervised or weakly supervised classifiers trained on synthetic datasets with injected factual corruptions (e.g., FactCC).
  4. Knowl 4 — Taxonomy of Hallucination Mitigation Methods

    model/method

    Hallucination mitigation methods target four stages of the NLG pipeline:

    1. Data-Related Methods:

      • Faithful Dataset Construction: Manual creation or multi-stage rule-based purification (phrase trimming, decontextualization, and syntax modification) of existing training sets.
      • Automated Data Cleaning: Filtering low-quality training instances via hallucination scores, or repairing meaning representations (MRs) by reverse parsing references.
      • Information Augmentation: Enriching input context with explicit relation triples, knowledge graph embeddings, or retrieved background documents.
    2. Architectural Modifications:

      • Dual/Structured Encoders: Combining sequential document encoders with graph neural networks (GNNs) to represent structured knowledge.
      • Constrained Decoders: Implementing tree-based decoding, focus attention mechanisms that bias output distributions toward source tokens, or multi-branch decoders that isolate fluency from content grounding.
    3. Training Paradigms:

      • Planning and Sketching: Guiding generation using two-step pipelines where explicit content plans or skeleton keywords are predicted prior to surface realization.
      • Reinforcement Learning (RL): Optimizing generation policies with rewards based on semantic consistency, QA cloze-scores, or NLI entailment probabilities.
      • Multi-Task and Contrastive Learning: Jointly training on auxiliary alignment or rationale extraction tasks, or training contrastive objectives (e.g., CLIFF, CONFIT) to differentiate clean references from corrupted negative samples.
    4. Post-Processing:

      • Generate-Then-Refine: Deploying specialized editor modules, QA-guided span correctors (e.g., SpanFact), or contrastive entity replacement models to correct factual inconsistencies in draft outputs without retraining the underlying generator.
  5. Knowl 5 — CHAIR Metric for Object Hallucination in Image Captioning

    equation

    In vision-language image captioning, Caption Hallucination Assessment with Image Relevance (CHAIR) evaluates the proportion of generated object words that do not appear in the image according to ground-truth reference captions.

    CHAIR is evaluated at two granularities:

    CHAIRi=#{hallucinated object instances}#{all object instances mentioned in generated captions}CHAIR_i = \frac{\#\{\text{hallucinated object instances}\}}{\#\{\text{all object instances mentioned in generated captions}\}}

    CHAIRs=#{hallucinated captions}#{all generated captions}CHAIR_s = \frac{\#\{\text{hallucinated captions}\}}{\#\{\text{all generated captions}\}}

    Where:

    • CHAIRiCHAIR_i (instance-level) represents the fraction of mentioned object instances in the generated text that are absent from the ground truth.
    • CHAIRsCHAIR_s (sentence-level) represents the fraction of generated captions that contain at least one hallucinated object.
    • Objects are extracted and aligned by matching lemmatized nouns and synonyms against a pre-defined object segmentation lexicon (such as the 80 object categories in MS COCO).
  6. Knowl 6 — Language Model Loss Differential for Token-Level Hallucination Detection

    model/method

    Token-level hallucination detection can be performed by comparing the cross-entropy loss of a conditional language model against an unconditional language model during forced-path decoding.

    Two language models are used:

    1. An unconditional model LMuncondLM_{\text{uncond}} trained strictly on target reference texts yy.
    2. A conditional model LMcondLM_{\text{cond}} trained on source-target pairs (x,y)(x, y).

    For a target token sequence y=(y1,y2,,yy)y = (y_1, y_2, \dots, y_{|y|}), the per-token loss L(yt)\mathcal{L}(y_t) is evaluated under both models. If L(ytLMuncond)<L(ytLMcond,x)\mathcal{L}(y_t \mid LM_{\text{uncond}}) < \mathcal{L}(y_t \mid LM_{\text{cond}}, x) the token yty_t is flagged as hallucinatory, under the premise that the unconditional prior explains the token better than the source conditioning xx.

    The sequence-level hallucination score is the proportion of hallucinated tokens across the sequence length y|y|: Hallucination Score=1yt=1yI(L(ytLMuncond)<L(ytLMcond,x))\text{Hallucination Score} = \frac{1}{|y|} \sum_{t=1}^{|y|} \mathbb{I}\left(\mathcal{L}(y_t \mid LM_{\text{uncond}}) < \mathcal{L}(y_t \mid LM_{\text{cond}}, x)\right) where I()\mathbb{I}(\cdot) is the indicator function.

  7. Knowl 7 — PARENT and PARENT-T Grounding Metrics in Data-to-Text Generation

    model/method

    PARENT (Precision And Recall of Entailed nn-grams from the Table) measures factual alignment in data-to-text generation by calculating lexical nn-gram entailment of a generated text GG against both the source table TT and the human reference description RR.

    PARENT evaluates:

    • Entailment Precision: The proportion of nn-grams in GG that are verified as entailed either directly by the table TT or by the reference RR. This rewards valid table extractions omitted by the human reference.
    • Entailment Recall: The proportion of nn-grams in RR that are grounded in TT and correctly reproduced in GG.

    The overall PARENT score is the harmonic mean (F1F_1-score) of entailment precision and recall averaged across nn-gram orders n{1,2,3,4}n \in \{1, 2, 3, 4\}.

    PARENT-T is a reference-free variant that evaluates the instance pair (T,G)(T, G) exclusively by checking entailment directly against table TT, ignoring human reference RR entirely to prevent reference divergence errors from distorting evaluation.

  8. Knowl 8 — Hallucination Categories in Neural Machine Translation

    definition

    In Neural Machine Translation (NMT), hallucinations denote target outputs that are completely disconnected or inconsistent with the source input text, categorized as follows:

    • Intrinsic Translation Hallucination: Output that contains incorrect meaning relative to the source (e.g., entity swapping, negation of positive source verbs).
    • Extrinsic Translation Hallucination: Output that generates supplementary content without any semantic grounding in the source sentence.
    • Perturbation-Induced Hallucination: Pathological translations where minimal, non-semantic source perturbations (such as repeated characters or tokens) cause the model to produce entirely disconnected fluent sentences.
    • Detached Hallucination: Natural hallucinations where the target translation is completely semantically decoupled from the source input.
    • Oscillatory Hallucination (Over-Translation): Pathological translation patterns where the decoder produces repetitive nn-grams or loops indefinitely.
    • Under-Translation: The omission or systematic skipping of valid source clauses during generation.
  9. Knowl 9 — Consistency Dimensions in Dialogue Generation: Self-Consistency vs External Consistency

    definition

    Dialogue generation systems distinguish between two independent dimensions of consistency:

    1. Self-Consistency (Persona & Turn Consistency):

      • Measures whether an agent's current utterance contradicts its own prior utterances or its assigned persona profile across multi-turn exchanges.
      • Self-contradiction (such as stating conflicting personal facts or differing answers to semantically equivalent prompts) constitutes an intrinsic hallucination with respect to the dialogue history.
    2. External Consistency (Knowledge Grounding):

      • Measures whether responses remain faithful to provided background knowledge documents or structured knowledge graph triples.
      • In knowledge-grounded dialogue, intrinsic hallucination denotes misrepresenting entities or relations in the provided knowledge (e.g., swapping subject and object), while extrinsic hallucination denotes output assertions that cannot be validated from the provided knowledge corpus.
  10. Knowl 10 — Reference-Dependent and Reference-Free Metrics for LLM Hallucination

    model/method

    Hallucination evaluation in Large Language Models (LLMs) is categorized into reference-dependent and reference-free paradigms:

    1. Reference-Dependent Benchmarking:

      • Long-Tail and Adversarial QA: Structured evaluation sets targeting low-frequency long-tail knowledge, temporal drift, or human misconceptions (e.g., TruthfulQA, PopQA, RealTimeQA).
      • Atomic Claim Decomposition: Segmenting long-form generations into standalone atomic propositions and evaluating precision against verified knowledge graphs or Wikipedia (e.g., FActScore, FactualityPrompt).
      • Automated Retrieval & Verification Pipelines: Chained systems that prompt the LLM to formulate search queries, retrieve external evidence from search engines, and apply NLI or LLM judges to confirm claim entailment (e.g., FactTool, CRITIC, RARR).
    2. Reference-Free Detection:

      • Uncertainty-Based Detection (White-Box): Measuring token probabilities, predictive entropy, or activation states across output tokens, aggregating them across sequence spans to identify low-confidence hallucinations.
      • Consistency-Based Detection (Black-Box): Sampling multiple stochastic responses for a given prompt and measuring consistency across outputs using semantic entropy, BERTScore, or NLI entailment (e.g., SelfCheckGPT).
  11. Knowl 11 — Decoding and Inference-Time Hallucination Mitigation in Large Language Models

    model/method

    Inference-time methods mitigate LLM hallucinations without parameter retraining or architectural redesign:

    1. Logit-Contrastive Decoding:

      • DoLa (Decoding by Contrasting Layers): Contrasts output probability logits between mature higher layers and premature intermediate layers, amplifying factual tokens that crystallize in deeper layers.
      • CAD (Context-Aware Decoding): Subtracts unconditional logits (prompted without context) from conditional logits (prompted with context) to enforce grounding on the provided evidence over parametric priors.
    2. Activation-Level Intervention:

      • Inference-Time Intervention (ITI): Probes attention heads during forward inference to detect truthfulness-correlated activation directions and shifts residual stream activations along those truth vectors.
    3. Reasoning and Verification Workflows:

      • Chain-of-Verification (CoVE): A multi-step pipeline where the model drafts an initial response, generates independent factual verification questions, answers them without conditioning on the draft, and synthesizes a verified final response.
      • Self-Reflection Loops: Prompting the model to perform iterative natural language inference, critique detected errors, and execute automated revision before final output delivery.

Coverage note — Omitted exhaustive literature review citations and granular descriptions of individual dataset benchmarks across subtasks that do not present generalizable methodological contributions.

References

  1. 1.Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, Dana Rezazadegan, Li Liu, Mohammad Ghavamzadeh, Paul W. Fieguth, Xiaochun Cao, Abbas Khosravi, U. Rajendra Acharya, Vladimir Makarenkov, and Saeid Nahavandi. 2020. A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges. CoRR abs/2011.06225 (2020). arXiv:2011.06225 https://arxiv.org/abs/2011.06225
  2. 2.Vaibhav Adlakha, Parishad BehnamGhader, Xing Han Lu, Nicholas Meade, and Siva Reddy. 2023. Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering. CoRR abs/2307.16877 (2023). https://doi.org/10.48550/arXiv.2307.16877 arXiv:2307.16877
  3. 3.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022. Flamingo: a Visual Language Model for Few-Shot Learning. ArXiv abs/2204.14198 (2022).
  4. 4.Joshua Albrecht and Rebecca Hwa. 2007. A Re-examination of Machine Learning Approaches for Sentence-Level MT Evaluation. In Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics. Association for Computational Linguistics, Prague, Czech Republic, 880–887. https://aclanthology.org/P07-1111
  5. 5.Alfonso Amayuelas, Liangming Pan, Wenhu Chen, and William Yang Wang. 2023. Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models. CoRR abs/2305.13712 (2023). https://doi.org/10.48550/arXiv.2305.13712 arXiv:2305.13712
  6. 6.Rahul Aralikatte, Shashi Narayan, Joshua Maynez, Sascha Rothe, and Ryan McDonald. 2021. Focus Attention: Promoting Faithfulness and Diversity in Summarization. ACL (2021).
  7. 7.Amos Azaria and Tom Mitchell. 2023. The internal state of an llm knows when its lying. arXiv preprint arXiv:2304.13734 (2023).
  8. 8.Xiang Bai, Xinggang Wang, Longin Jan Latecki, Wenyu Liu, and Zhuowen Tu. 2009. Active Skeleton for Non-rigid Object Detection. In 2009 IEEE 12th International Conference on Computer Vision. 575–582. https://doi.org/10.1109/ICCV.2009.5459188
  9. 9.S. Baker and T. Kanade. 2000. Hallucinating Faces. In Proceedings Fourth IEEE International Conference on Automatic Face and Gesture Recognition (Cat. No. PR00580). 83–88. https://doi.org/10.1109/AFGR.2000.840616
  10. 10.Anusha Balakrishnan, Jinfeng Rao, Kartikeya Upasani, Michael White, and Rajen Subba. 2019. Constrained Decoding for Neural NLG from Compositional Representations in Task-Oriented Dialogue. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 831–844. https://doi.org/10.18653/v1/P19-1080
  11. 11.Ramy Baly, Georgi Karadzhov, Dimitar Alexandrov, James Glass, and Preslav Nakov. 2018. Predicting Factuality of Reporting and Bias of News Media Sources. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 3528–3539.
  12. 12.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The Long-Document Transformer. arXiv:2004.05150 (2020).
  13. 13.Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks. Advances in neural information processing systems 28 (2015).
  14. 14.Anne Beyer, Sharid Loáiciga, and David Schlangen. 2021. Is Incoherence Surprising? Targeted Evaluation of Coherence Prediction from Language Models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, Online, 4164–4173. https://doi.org/10.18653/v1/2021.naacl-main.328
  15. 15.Bin Bi, Chen Wu, Ming Yan, Wei Wang, Jiangnan Xia, and Chenliang Li. 2019. Incorporating External Knowledge into Machine Reading for Generative Question Answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2521–2530.
  16. 16.Andrzej Białecki, Robert Muir, Grant Ingersoll, and Lucid Imagination. 2012. Apache lucene 4. In SIGIR 2012 workshop on open source information retrieval. 17.
  17. 17.Ali Furkan Biten, Lluís Gómez, and Dimosthenis Karatzas. 2022. Let There Be a Clock on the Beach: Reducing Object Hallucination in Image Captioning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). 1381–1390.
  18. 18.Jan Dirk Blom. [n. d.]. A Dictionary of Hallucinations. Springer.
  19. 19.Eleftheria Briakou and Marine Carpuat. 2021. Beyond Noise: Mitigating the Impact of Fine-grained Semantic Divergences on Neural Machine Translation. CoRR abs/2105.15087 (2021). arXiv:2105.15087 https://arxiv.org/abs/2105.15087
  20. 20.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 1877–1901. https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
  21. 21.Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2022. Discovering Latent Knowledge in Language Models Without Supervision. In The Eleventh International Conference on Learning Representations.
  22. 22.Meng Cao, Yue Dong, Jiapeng Wu, and Jackie Chi Kit Cheung. 2020. Factual Error Correction for Abstractive Summarization Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 6251–6258.
  23. 23.Shuyang Cao and Lu Wang. 2021. CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization. EMNLP (2021).
  24. 24.Yihan Cao, Yanbin Kang, Chi Wang, and Lichao Sun. 2023. INSTRUCTION MINING: WHEN DATA MINING MEETS LARGE LANGUAGE MODEL FINETUNING. arXiv preprint arXiv:2307.06290 (2023).
  25. 25.Ziqiang Cao, Furu Wei, Wenjie Li, and Sujian Li. 2018. Faithful to the Original: Fact Aware Neural Abstractive Summarization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32.
  26. 26.Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B Brown, Dawn Song, Úlfar Erlingsson, et al. 2020. Extracting Training Data from Large Language Models. (2020).
  27. 27.Anthony Chen, Panupong Pasupat, Sameer Singh, Hongrae Lee, and Kelvin Guu. 2023. PURR: Efficiently Editing Language Model Hallucinations by Denoising Language Model Corruptions. arXiv preprint arXiv:2305.14908 (2023).
  28. 28.Delong Chen, Jianfeng Liu, Wenliang Dai, and Baoyuan Wang. 2023. Visual Instruction Tuning with Polite Flamingo. CoRR abs/2307.01003 (2023). https://doi.org/10.48550/ARXIV.2307.01003 arXiv:2307.01003
  29. 29.Jifan Chen, Grace Kim, Aniruddh Sriram, Greg Durrett, and Eunsol Choi. 2023. Complex Claim Verification with Evidence Retrieved in the Wild. CoRR abs/2305.11859 (2023). https://doi.org/10.48550/arXiv.2305.11859 arXiv:2305.11859
  30. 30.Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2023. Benchmarking Large Language Models in Retrieval-Augmented Generation. CoRR abs/2309.01431 (2023). https://doi.org/10.48550/arXiv.2309.01431 arXiv:2309.01431
  31. 31.Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. 2023. AlpaGasus: Training A Better Alpaca with Fewer Data. arXiv preprint arXiv:2307.08701 (7 2023). http://arxiv.org/abs/2307.08701
  32. 32.Sihao Chen, Fan Zhang, Kazoo Sone, and Dan Roth. 2021. Improving Faithfulness in Abstractive Summarization with Contrast Candidate Generation and Selection. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 5935–5941.
  33. 33.Wenqing Chen, Jidong Tian, Yitian Li, Hao He, and Yaohui Jin. 2021. De-Confounded Variational Encoder-Decoder for Logical Table-to-Text Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 5532–5542.
  34. 34.Zhiyu Chen, Wenhu Chen, Hanwen Zha, Xiyou Zhou, Yunkai Zhang, Sairam Sundaresan, and William Yang Wang. 2020. Logic2Text: High-Fidelity Natural Language Generation from Logical Forms. In EMNLP (Findings).
  35. 35.I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu. 2023. FacTool: Factuality Detection in Generative AI - A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios. CoRR abs/2307.13528 (2023). https://doi.org/10.48550/arXiv.2307.13528 arXiv:2307.13528
  36. 36.Andrew Chisholm, Will Radford, and Ben Hachey. 2017. Learning to generate one-sentence biographies from Wikidata. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 633–642.
  37. 37.Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He. 2023. DOLA: DECODING BY CONTRASTING LAYERS IMPROVES FACTUALITY IN LARGE LANGUAGE MODELS. arXiv preprint arXiv:2309.03883 (2023).
  38. 38.Michael Crawshaw. 2020. Multi-Task Learning with Deep Neural Networks: A Survey. arXiv preprint arXiv:2009.09796 (2020).
  39. 39.Leyang Cui, Yu Wu, Shujie Liu, Yue Zhang, and Ming Zhou. 2020. MuTual: A Dataset for Multi-Turn Dialogue Reasoning. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 1406–1416.
  40. 40.Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, and Edoardo M Ponti. 2023. Elastic Weight Removal for Faithful and Abstractive Dialogue Generation. arXiv preprint arXiv:2303.17574 (2023).
  41. 41.Wenliang Dai, Lu Hou, Lifeng Shang, Xin Jiang, Qun Liu, and Pascale Fung. 2022. Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation. In Findings of the Association for Computational Linguistics: ACL 2022. Association for Computational Linguistics, Dublin, Ireland, 2383–2395. https://doi.org/10.18653/v1/2022.findings-acl.187
  42. 42.Wenliang Dai, Zihan Liu, Ziwei Ji, Dan Su, and Pascale Fung. 2022. Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-training. ArXiv abs/2210.07688 (2022).
  43. 43.Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, and Joel Veness. 2023. Language Modeling Is Compression. CoRR abs/2309.10668 (2023). https://doi.org/10.48550/ARXIV.2309.10668 arXiv:2309.10668
  44. 44.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186. https://doi.org/10.18653/v1/N19-1423
  45. 45.Bhuwan Dhingra, Manaal Faruqui, Ankur Parikh, Ming-Wei Chang, Dipanjan Das, and William Cohen. 2019. Handling Divergent Reference Texts when Evaluating Table-to-Text Generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 4884–4895.
  46. 46.Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023. Chain-of-Verification Reduces Hallucination in Large Language Models. arXiv preprint arXiv:2309.11495 (2023).
  47. 47.Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2020. The Second Conversational Intelligence Challenge (ConvAI2). In The NeurIPS’18 Competition. Springer, 187–208.
  48. 48.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. Wizard of Wikipedia: Knowledge-Powered Conversational agents. ICLR (2019).
  49. 49.Georgiana Dinu, Prashant Mathur, Marcello Federico, and Yaser Al-Onaizan. 2019. Training Neural Machine Translation To Apply Terminology Constraints. ACL 2019 - 57th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (6 2019), 3063–3068. https://doi.org/10.18653/v1/p19-1294
  50. 50.Yue Dong, Shuohang Wang, Zhe Gan, Yu Cheng, Jackie Chi Kit Cheung, and Jingjing Liu. 2020. Multi-Fact Correction in Abstractive Text Summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 9320–9331.
  51. 51.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net. https://openreview.net/forum?id=YicbFdNTTy
  52. 52.Li Du, Yequan Wang, Xingrun Xing, Yiqun Ya, Xiang Li, Xin Jiang, and Xuezhi Fang. 2023. Quantifying and Attributing the Hallucination of Large Language Models via Association Analysis. CoRR abs/2309.05217 (2023). https://doi.org/10.48550/arXiv.2309.05217 arXiv:2309.05217
  53. 53.Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. 2023. Improving Factuality and Reasoning in Language Models through Multiagent Debate. In Forty-first International Conference on Machine Learning.
  54. 54.Esin Durmus, He He, and Mona Diab. 2020. FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 5055–5070.
  55. 55.Ondřej Dušek, David M Howcroft, and Verena Rieser. 2019. Semantic Noise Matters for Neural Natural Language Generation. In Proceedings of the 12th International Conference on Natural Language Generation. 421–426.
  56. 56.Ondřej Dušek and Filip Jurčíček. 2016. Sequence-to-Sequence Generation for Spoken Dialogue via Deep Syntax Trees and Strings. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, Berlin, Germany, 45–51. https://doi.org/10.18653/v1/P16-2008
  57. 57.Ondřej Dušek and Zdenĕk Kasner. 2020. Evaluating Semantic Accuracy of Data-to-Text Generation with Natural Language Inference. In Proceedings of the 13th International Conference on Natural Language Generation. Association for Computational Linguistics, Dublin, Ireland, 131–137. https://aclanthology.org/2020.inlg-1.19
  58. 58.Nouha Dziri, Ehsan Kamalloo, Kory Mathewson, and Osmar Zaiane. 2019. Evaluating Coherence in Dialogue Systems using Entailment. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, 3806–3812. https://doi.org/10.18653/v1/N19-1381
  59. 59.Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaïd Harchaoui, and Yejin Choi. 2023. Faith and Fate: Limits of Transformers on Compositionality. CoRR abs/2305.18654 (2023). https://doi.org/10.48550/ARXIV.2305.18654 arXiv:2305.18654
  60. 60.Nouha Dziri, Andrea Madotto, Osmar Zaiane, and Avishek Joey Bose. 2021. Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding. EMNLP (2021).
  61. 61.Nouha Dziri, Hannah Rashkin, Tal Linzen, and David Reitter. 2021. Evaluating Groundedness in Dialogue Systems: The BEGIN Benchmark. Findings of ACL (2021).
  62. 62.Mohamed Elaraby, Mengyin Lu, Jacob Dunn, Xueying Zhang, Yu Wang, and Shizhu Liu. 2023. Halo: Estimation and Reduction of Hallucinations in Open-Source Weak Large Language Models. CoRR abs/2308.11764 (2023). https://doi.org/10.48550/arXiv.2308.11764 arXiv:2308.11764
  63. 63.Mihail Eric and Christopher Manning. 2017. A Copy-Augmented Sequence-to-Sequence Architecture Gives Good Performance on Task-Oriented Dialogue. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers. Association for Computational Linguistics, Valencia, Spain, 468–473. https://aclanthology.org/E17-2075
  64. 64.Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert. 2023. RAGAS: Automated Evaluation of Retrieval Augmented Generation. arXiv preprint arXiv:2309.15217 (9 2023). http://arxiv.org/abs/2309.15217
  65. 65.Oren Etzioni, Michele Banko, Stephen Soderland, and Daniel S. Weld. 2008. Open Information Extraction from the Web. Commun. ACM 51, 12 (Dec. 2008), 68–74. https://doi.org/10.1145/1409360.1409378
  66. 66.Tobias Falke, Leonardo FR Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019. Ranking Generated Summaries by Correctness: An Interesting but Challenging Application for Natural Language Inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2214–2220.
  67. 67.Angela Fan, Claire Gardent, Chloé Braud, and Antoine Bordes. 2019. Using Local Knowledge Graph Construction to Scale Seq2Seq Models to Multi-Document Inputs. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 4186–4196.
  68. 68.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: Long Form Question Answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 3558–3567.
  69. 69.Alhussein Fawzi, Horst Samulowitz, Deepak Turaga, and Pascal Frossard. 2016. Image Inpainting through Neural Networks Hallucinations. In 2016 IEEE 12th Image, Video, and Multidimensional Signal Processing Workshop (IVMSP). Ieee, 1–5.
  70. 70.Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang. 2023. Towards Revealing the Mystery behind Chain of Thought: a Theoretical Perspective. CoRR abs/2305.15408 (2023). https://doi.org/10.48550/ARXIV.2305.15408 arXiv:2305.15408
  71. 71.Yang Feng, Wanying Xie, Shuhao Gu, Chenze Shao, Wen Zhang, Zhengxin Yang, and Dong Yu. 2020. Modeling Fluency and Faithfulness for Diverse Neural Machine Translation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 59–66.
  72. 72.Katja Filippova. 2020. Controlled Hallucinations: Learning to Generate Faithfully from Noisy Data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings. 864–870.
  73. 73.William Fish et al. 2009. Perception, hallucination, and illusion. OUP USA.
  74. 74.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. GPTScore: Evaluate as You Desire. CoRR abs/2302.04166 (2023). https://doi.org/10.48550/arXiv.2302.04166 arXiv:2302.04166
  75. 75.Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao. 2021. GO FIGURE: A Meta Evaluation of Factuality in Summarization. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. Association for Computational Linguistics, Online, 478–487. https://doi.org/10.18653/v1/2021.findings-acl.42
  76. 76.Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark. 2022. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. CoRR abs/2209.07858 (2022). https://doi.org/10.48550/arXiv.2209.07858 arXiv:2209.07858
  77. 77.Jianfeng Gao, Michel Galley, and Lihong Li. 2018. Neural Approaches to Conversational AI. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts. Association for Computational Linguistics, Melbourne, Australia, 2–7. https://doi.org/10.18653/v1/P18-5002
  78. 78.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y. Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, and Kelvin Guu. 2023. RARR: Researching and Revising What Language Models Say, Using Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, 16477–16508. https://doi.org/10.18653/v1/2023.acl-long.910
  79. 79.Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017. Creating Training Corpora for NLG Micro-Planning. In 55th annual meeting of the Association for Computational Linguistics (ACL).
  80. 80.Sarthak Garg, Stephan Peitz, Udhyakumar Nallasamy, and Matthias Paulik. 2019. Jointly Learning to Align and Translate with Transformer Models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 4453–4462.
  81. 81.Albert Gatt and Emiel Krahmer. 2018. Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation. Journal of Artificial Intelligence Research 61 (2018), 65–170.
  82. 82.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer Feed-Forward Layers Are Key-Value Memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). Association for Computational Linguistics, 5484–5495. https://doi.org/10.18653/V1/2021.EMNLP-MAIN.446
  83. 83.Deepanway Ghosal, Pengfei Hong, Siqi Shen, Navonil Majumder, Rada Mihalcea, and Soujanya Poria. 2021. CIDER: Commonsense Inference for Dialogue Explanation and Reasoning. ACL (2021).
  84. 84.Alexandru L Ginsca, Adrian Popescu, and Mihai Lupu. 2015. Credibility in Information Retrieval. Foundations and Trends in Information Retrieval 9, 5 (2015), 355–475.
  85. 85.Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019. SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization. In Proceedings of the 2nd Workshop on New Frontiers in Summarization. Association for Computational Linguistics, Hong Kong, China, 70–79. https://doi.org/10.18653/v1/D19-5409
  86. 86.Silke M Göbel and Matthew FS Rushworth. 2004. Cognitive Neuroscience: Acting on Numbers. Current Biology 14, 13 (2004), R517–R519.
  87. 87.Ben Goodrich, Vinay Rao, Peter J Liu, and Mohammad Saleh. 2019. Assessing the Factual Accuracy of Generated Text. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 166–175.
  88. 88.Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2023. CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing. CoRR abs/2305.11738 (2023). https://doi.org/10.48550/arXiv.2305.11738 arXiv:2305.11738
  89. 89.Kartik Goyal, Chris Dyer, and Taylor Berg-Kirkpatrick. 2017. Differentiable Scheduled Sampling for Credit Assignment. ACL 2017 - 55th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers) 2 (4 2017), 366–371. https://doi.org/10.18653/v1/P17-2058
  90. 90.Tanya Goyal and Greg Durrett. 2020. Evaluating Factuality in Generation with Dependency-level Entailment. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings. 3592–3603.
  91. 91.Jian Guan and Minlie Huang. 2020. UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 9157–9166.
  92. 92.Beliz Gunel, Chenguang Zhu, Michael Zeng, and Xuedong Huang. 2019. Mind the facts: Knowledge-Boosted Coherent Abstractive Text Summarization. NeurIPS, Knowledge Representation & Reasoning Meets Machine Learning (KR2ML workshop) (2019).
  93. 93.Zhijiang Guo, Michael Sejr Schlichtkrull, and Andreas Vlachos. 2022. A Survey on Automated Fact-Checking. Trans. Assoc. Comput. Linguistics 10 (2022), 178–206. https://doi.org/10.1162/tacl_a_00454
  94. 94.Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2021. DialFact: A Benchmark for Fact-Checking in Dialogue. arXiv preprint arXiv:2110.08222 (2021).
  95. 95.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A Smith. 2018. Annotation Artifacts in Natural Language Inference Data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). 107–112.
  96. 96.Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazare, and Jason Weston. 2019. Learning from Dialogue after Deployment: Feed Yourself, Chatbot!. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 3667–3684. https://doi.org/10.18653/v1/P19-1358
  97. 97.Tianxing He, Jingzhao Zhang, Zhiming Zhou, and James Glass. 2021. Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation?. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 5087–5102.
  98. 98.Chris Hokamp and Qun Liu. 2017. Lexically Constrained Decoding for Sequence Generation Using Grid Beam Search. ACL 2017 - 55th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers) 1 (4 2017), 1535–1546. https://doi.org/10.18653/v1/P17-1141
  99. 99.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The Curious Case of Neural Text Degeneration. In International Conference on Learning Representations.
  100. 100.Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022. TRUE: Re-evaluating Factual Consistency Evaluation. In Proceedings of the Second DialDoc Workshop on Document-grounded Dialogue and Conversational Question Answering. 161–175.
  101. 101.Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021. Q2 : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering. EMNLP (2021).
  102. 102.Luyang Huang, Lingfei Wu, and Lu Wang. 2020. Knowledge Graph-Augmented Abstractive Summarization with Semantic-Driven Cloze Reward. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (2020).
  103. 103.Minlie Huang, Xiaoyan Zhu, and Jianfeng Gao. 2020. Challenges in Building Intelligent Open-domain Dialog Systems. ACM Transactions on Information Systems (TOIS) 38, 3 (2020), 1–32.
  104. 104.Yichong Huang, Xiachong Feng, Xiaocheng Feng, and Bing Qin. 2021. The Factual Inconsistency Problem in Abstractive Text Summarization: A Survey. arXiv preprint arXiv:2104.14839 (2021).
  105. 105.Yuheng Huang, Jiayang Song, Zhijie Wang, Huaming Chen, and Lei Ma. 2023. Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models. CoRR abs/2307.10236 (2023). https://doi.org/10.48550/arXiv.2307.10236 arXiv:2307.10236
  106. 106.Siqing Huo, Negar Arabzadeh, and Charles L. A. Clarke. 2023. Retrieving Supporting Evidence for LLMs Generated Answers. CoRR abs/2306.13781 (2023). https://doi.org/10.48550/arXiv.2306.13781 arXiv:2306.13781
  107. 107.Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne. 2017. Imitation Learning: A Survey of Learning Methods. ACM Comput. Surv. 50, 2, Article 21 (apr 2017), 35 pages. https://doi.org/10.1145/3054912
  108. 108.Ziwei Ji, Zihan Liu, Nayeon Lee, Tiezheng Yu, Bryan Wilie, Min Zeng, and Pascale Fung. 2023. RHO: Reducing Hallucination in Open-domain Dialogues with Knowledge Grounding. In Findings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 4504–4522. https://doi.org/10.18653/v1/2023.findings-acl.275
  109. 109.Ziwei Ji, Yan Xu, I-Tsun Cheng, Samuel Cahyawijaya, Rita Frieske, Etsuko Ishii, Min Zeng, Andrea Madotto, and Pascale Fung. 2022. VScript: Controllable Script Generation with Visual Presentation. arXiv:2203.00314
  110. 110.Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung. 2023. Towards Mitigating Hallucination in Large Language Models via Self-Reflection. EMNLP Findings (2023).
  111. 111.Marcin Junczys-Dowmunt. 2018. Dual Conditional Cross-Entropy Filtering of Noisy Parallel Corpora. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers. Association for Computational Linguistics, Belgium, Brussels, 888–895. https://doi.org/10.18653/v1/W18-6478
  112. 112.Daniel Jurafsky and James H. Marin. 2019. Speech and Language Processing. Draft of October 16th, 2019, Website: https://web.stanford.edu/~jurafsky/slp3/26.pdf, Chapter 26.
  113. 113.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221 (2022).
  114. 114.Ehsan Kamalloo, Nouha Dziri, Charles L. A. Clarke, and Davood Rafiei. 2023. Evaluating Open-Domain Question Answering in the Era of Large Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, 5591–5606. https://doi.org/10.18653/v1/2023.acl-long.307
  115. 115.Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023. Large Language Models Struggle to Learn Long-Tail Knowledge. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol. 202), Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 15696–15707. https://proceedings.mlr.press/v202/kandpal23a.html
  116. 116.Daniel Kang and Tatsunori Hashimoto. 2020. Improved Natural Language Generation via Loss Truncation. (4 2020), 718–731. https://arxiv.org/abs/2004.14589v2
  117. 117.Daniel Kang and Tatsunori B Hashimoto. 2020. Improved Natural Language Generation via Loss Truncation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 718–731.
  118. 118.Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir R. Radev, Noah A. Smith, Yejin Choi, and Kentaro Inui. 2022. RealTime QA: What’s the Answer Right Now? CoRR abs/2207.13332 (2022). https://doi.org/10.48550/arXiv.2207.13332 arXiv:2207.13332
  119. 119.Osman Semih Kayhan, Bart Vredebregt, and Jan C van Gemert. 2021. Hallucination In Object Detection—A Study In Visual Part VERIFICATION. In 2021 IEEE International Conference on Image Processing (ICIP). IEEE, 2234–2238.
  120. 120.Daniel Khashabi, Amos Ng, Tushar Khot, Ashish Sabharwal, Hannaneh Hajishirzi, and Chris Callison-Burch. 2021. GooAQ: Open Question Answering with Diverse Answer Types. In Findings of the Association for Computational Linguistics: EMNLP 2021. Association for Computational Linguistics, Punta Cana, Dominican Republic, 421–433. https://doi.org/10.18653/v1/2021.findings-emnlp.38
  121. 121.Philipp Koehn and Rebecca Knowles. 2017. Six Challenges for Neural Machine Translation. In First Workshop on Neural Machine Translation. Association for Computational Linguistics, 28–39.
  122. 122.Philipp Koehn and Rebecca Knowles. 2017. Six Challenges for Neural Machine Translation. (2017), 28–39. http://www.statmt.org/wmt17/
  123. 123.Takeshi Kojima, Shixiang Shane Gu, Machel Reid Google Research, Yutaka Matsuo, and Yusuke Iwasawa. 2023. Large Language Models are Zero-Shot Reasoners. URL https://arxiv. org/abs/2205.11916 (2023).
  124. 124.Xiang Kong, Zhaopeng Tu, Shuming Shi, Eduard Hovy, and Tong Zhang. 2019. Neural Machine Translation with Adequacy-Oriented Learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 6618–6625.
  125. 125.Kalpesh Krishna, Aurko Roy, and Mohit Iyyer. 2021. Hurdles to Progress in Long-form Question Answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4940–4957.
  126. 126.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the Factual Consistency of Abstractive Text Summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 9332–9346.
  127. 127.Karen Kukich. 1983. Design of a Knowledge-based Report Generator. In 21st Annual Meeting of the Association for Computational Linguistics. 145–150.
  128. 128.Ilia Kulikov, Alexander H. Miller, Kyunghyun Cho, and Jason Weston. 2019. Importance of Search and Evaluation Strategies in Neural Dialogue Modeling. In Proceedings of the 12th International Conference on Natural Language Generation, INLG 2019, Tokyo, Japan, October 29 - November 1, 2019, Kees van Deemter, Chenghua Lin, and Hiroya Takamura (Eds.). Association for Computational Linguistics, 76–87. https://doi.org/10.18653/v1/W19-8609
  129. 129.Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju. 2023. Certifying LLM Safety against Adversarial Prompting. CoRR abs/2309.02705 (2023). https://doi.org/10.48550/arXiv.2309.02705 arXiv:2309.02705
  130. 130.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural Questions: A Benchmark for Question Answering Research. Transactions of the Association for Computational Linguistics 7 (2019), 452–466.
  131. 131.Philippe Laban, Tobias Schnabel, Paul N Bennett, and Marti A Hearst. 2022. SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization. Transactions of the Association for Computational Linguistics 10 (2022), 163–177.
  132. 132.Rémi Lebret, David Grangier, and Michael Auli. 2016. Neural Text Generation from Structured Data with Application to the Biography Domain. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (2016).
  133. 133.Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. 2019. Hallucinations in Neural Machine Translation. ICLR (2019).
  134. 134.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021. Deduplicating Training Data Makes Language Models Better. arXiv preprint arXiv:2107.06499 (2021).
  135. 135.Nayeon Lee, Belinda Z Li, Sinong Wang, Wen-Tau Yih, Hao Ma, and Madian Khabsa. 2020. Language Models as Fact Checkers? ACL 2020 (2020), 36.
  136. 136.Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2022. Factuality Enhanced Language Models for Open-Ended Text Generation. arXiv preprint arXiv:2206.04624 (2022).
  137. 137.Nayeon Lee, Chien-Sheng Wu, and Pascale Fung. [n. d.]. Improving Large-Scale Fact-Checking using Decomposable Attention Models and Lexical Tagging. ([n. d.]).
  138. 138.Deren Lei, Yaxi Li, Mingyu Wang, Vincent Yun, Emily Ching, Eslam Kamal, et al. 2023. Chain of natural language inference for reducing large language model ungrounded hallucinations. arXiv preprint arXiv:2310.03951 (2023).
  139. 139.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 7871–7880.
  140. 140.Bohan Li, Yutai Hou, and Wanxiang Che. 2021. Data Augmentation Approaches in Natural Language Processing: A Survey. arXiv preprint arXiv:2110.01852 (2021).
  141. 141.Chenliang Li, Bin Bi, Ming Yan, Wei Wang, and Songfang Huang. 2021. Addressing Semantic Drift in Generative Question Answering with Auxiliary Extraction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). 942–947.
  142. 142.Haoran Li, Junnan Zhu, Jiajun Zhang, and Chengqing Zong. 2018. Ensure the Correctness of the Summary: Incorporate Entailment Knowledge into Abstractive Sentence Summarization. In Proceedings of the 27th International Conference on Computational Linguistics. Association for Computational Linguistics, Santa Fe, New Mexico, USA, 1430–1441. https://aclanthology.org/C18-1121
  143. 143.Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023. HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. CoRR abs/2305.11747 (2023). https://doi.org/10.48550/arXiv.2305.11747 arXiv:2305.11747
  144. 144.Jiwei Li, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao, and Bill Dolan. 2016. A Persona-Based Neural Conversation Model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Berlin, Germany, 994–1003. https://doi.org/10.18653/v1/P16-1094
  145. 145.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. In ICML.
  146. 146.Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2024. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems 36 (2024).
  147. 147.Miaoran Li, Baolin Peng, and Zhu Zhang. 2023. Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models. CoRR abs/2305.14623 (2023). https://doi.org/10.48550/arXiv.2305.14623 arXiv:2305.14623
  148. 148.Margaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck, Y-Lan Boureau, Kyunghyun Cho, and Jason Weston. 2020. Don’t Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Online, 4715–4728. https://doi.org/10.18653/v1/2020.acl-main.428
  149. 149.Tian Li, Ahmad Beirami, Maziar Sanjabi, and Virginia Smith. 2020. Tilted Empirical Risk Minimization. In International Conference on Learning Representations.
  150. 150.Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee. 2023. Textbooks Are All You Need II: phi-1.5 technical report. https://arxiv.org/pdf/2309.05463.pdf
  151. 151.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. 2023. Evaluating Object Hallucination in Large Vision-Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, 292–305. https://aclanthology.org/2023.emnlp-main.20
  152. 152.Yangming Li, Kaisheng Yao, Libo Qin, Wanxiang Che, Xiaolong Li, and Ting Liu. 2020. Slot-consistent NLG for Task-oriented Dialogue Systems with Iterative Rectification Network. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Online, 97–106. https://doi.org/10.18653/v1/2020.acl-main.10
  153. 153.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let’s Verify Step by Step. arXiv preprint arXiv:2305.20050 (5 2023). http://arxiv.org/abs/2305.20050
  154. 154.Chin-Yew Lin. 2004. Rouge: A Package for Automatic Evaluation of Summaries. In Text summarization branches out. 74–81.
  155. 155.Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv preprint arXiv:2109.07958 (2021).
  156. 156.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, 3214–3252. https://doi.org/10.18653/v1/2022.acl-long.229
  157. 157.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. ArXiv abs/1405.0312 (2014).
  158. 158.Yong Lin, Lu Tan, Hangyu Lin, Zeming Zheng, Renjie Pi, Jipeng Zhang, Shizhe Diao, Haoxiang Wang, Han Zhao, Yuan Yao, et al. 2023. Speciality vs generality: An empirical study on catastrophic forgetting in fine-tuning foundation models. arXiv preprint arXiv:2309.06256 (2023).
  159. 159.Ce Liu, Heung-Yeung Shum, and William T Freeman. 2007. Face Hallucination: Theory and Practice. International Journal of Computer Vision 75, 1 (2007), 115–134.
  160. 160.Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. 2021. A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text Generation. arXiv preprint arXiv:2104.08704 (2021).
  161. 161.Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. 2022. A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text Generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, 6723–6737. https://doi.org/10.18653/v1/2022.acl-long.464
  162. 162.Tianyu Liu, Xin Zheng, Baobao Chang, and Zhifang Sui. 2021. Towards Faithfulness in Open Domain Table-to-text Generation from an Entity-centric View. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 13415–13423.
  163. 163.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-Based Knowledge Conflicts in Question Answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 7052–7063.
  164. 164.Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2018. Neural Baby Talk. In CVPR.
  165. 165.Yukun Ma, Khanh Linh Nguyen, Frank Z Xing, and Erik Cambria. 2020. A Survey on Empathetic Dialogue Systems. Information Fusion 64 (2020), 50–70.
  166. 166.Fiona Macpherson and Dimitris Platchias. 2013. Hallucination: Philosophy and psychology. MIT Press.
  167. 167.Andrea Madotto, Samuel Cahyawijaya, Genta Indra Winata, Yan Xu, Zihan Liu, Zhaojiang Lin, and Pascale Fung. 2020. Learning Knowledge Bases with Parameters for Task-Oriented Dialogue Systems. In Findings of the Association for Computational Linguistics: EMNLP 2020. Association for Computational Linguistics, Online, 2372–2394. https://doi.org/10.18653/v1/2020.findings-emnlp.215
  168. 168.Andrea Madotto, Zhaojiang Lin, Chien-Sheng Wu, and Pascale Fung. 2019. Personalizing dialogue agents via meta-learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 5454–5459.
  169. 169.Andrea Madotto, Zihan Liu, Zhaojiang Lin, and Pascale Fung. 2020. Language Models as Few-Shot Learner for Task-Oriented Dialogue Systems. arXiv:2008.06239 [cs.CL]
  170. 170.Andrea Madotto, Chien-Sheng Wu, and Pascale Fung. 2018. Mem2Seq: Effectively Incorporating Knowledge Bases into End-to-End Task-Oriented Dialog Systems. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Melbourne, Australia, 1468–1478. https://doi.org/10.18653/v1/P18-1136
  171. 171.Amr Magdy and Nayer Wanas. 2010. Web-based Statistical Fact Checking of Textual Documents. In Proceedings of the 2nd international workshop on Search and mining user-generated contents. 103–110.
  172. 172.Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. 2023. ExpertQA: Expert-Curated Questions and Attributed Answers. CoRR abs/2309.07852 (2023). https://doi.org/10.48550/arXiv.2309.07852 arXiv:2309.07852
  173. 173.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, 9802–9822. https://doi.org/10.18653/v1/2023.acl-long.546
  174. 174.Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. 2023. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. CoRR abs/2303.08896 (2023). https://doi.org/10.48550/arXiv.2303.08896 arXiv:2303.08896
  175. 175.Marianna J. Martindale, Marine Carpuat, Kevin Duh, and Paul McNamee. 2019. Identifying Fluently Inadequate Output in Neural and Statistical Machine Translation. In MTSummit.
  176. 176.Kim Martineau. 2023. What is retrieval-augmented generation? https://research.ibm.com/blog/retrieval-augmented-generation-RAG
  177. 177.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 1906–1919.
  178. 178.Pierre-Emmanuel Mazaré, Samuel Humeau, Martin Raison, and Antoine Bordes. 2018. Training Millions of Personalized Dialogue Agents. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Brussels, Belgium, 2775–2779. https://doi.org/10.18653/v1/D18-1298
  179. 179.Kathleen McKeown. 1992. Text Generation. Cambridge University Press.
  180. 180.José Mena, Oriol Pujol, and Jordi Vitrià. 2022. A Survey on Uncertainty Estimation in Deep Learning Classification Systems from a Bayesian Perspective. ACM Comput. Surv. 54, 9 (2022), 193:1–193:35. https://doi.org/10.1145/3477140
  181. 181.Mohsen Mesgar, Edwin Simpson, and Iryna Gurevych. 2021. Improving Factual Consistency Between a Response and Persona Facts. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. Association for Computational Linguistics, Online, 549–562. https://doi.org/10.18653/v1/2021.eacl-main.44
  182. 182.Donald Metzler, Yi Tay, Dara Bahri, and Marc Najork. 2021. Rethinking Search: Making Experts out of Dilettantes. arXiv preprint arXiv:2105.02274 (2021).
  183. 183.Sabrina J Mielke, Arthur Szlam, Y-Lan Boureau, and Emily Dinan. 2020. Linguistic calibration through metacognition: aligning dialogue agent responses with expected correctness. arXiv preprint arXiv:2012.14983 (2020).
  184. 184.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. CoRR abs/2305.14251 (2023). https://doi.org/10.48550/arXiv.2305.14251 arXiv:2305.14251
  185. 185.Anshuman Mishra, Dhruvesh Patel, Aparna Vijayakumar, Xiang Lorraine Li, Pavan Kapanipathi, and Kartik Talamadupula. 2021. Looking Beyond Sentence-Level Natural Language Inference for Question Answering and Text Summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1322–1336.
  186. 186.Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, and Yoav Shoham. 2023. Generating Benchmarks for Factuality Evaluation of Language Models. CoRR abs/2307.06908 (2023). https://doi.org/10.48550/arXiv.2307.06908 arXiv:2307.06908
  187. 187.Mathias Müller, Annette Rios, and Rico Sennrich. 2020. Domain Robustness in Neural Machine Translation. In 14th Conference of the Association for Machine Translation in the Americas. Association for Machine Translation in the Americas, AMTA, 151–164.
  188. 188.Niels Mündler, Jingxuan He, Slobodan Jenko, and Martin T. Vechev. 2023. Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation. CoRR abs/2305.15852 (2023). https://doi.org/10.48550/arXiv.2305.15852 arXiv:2305.15852
  189. 189.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. WebGPT: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332 (2021).
  190. 190.Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, and Bing Xiang. 2021. Entity-level Factual Consistency of Abstractive Text Summarization. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2727–2733.
  191. 191.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A human generated machine reading comprehension dataset. In CoCo@ NIPS.
  192. 192.Feng Nie, Jinpeng Wang, Jin-Ge Yao, Rong Pan, and Chin-Yew Lin. 2018. Operation-guided Neural Networks for High Fidelity Data-To-Text Generation. In EMNLP.
  193. 193.Feng Nie, Jin-Ge Yao, Jinpeng Wang, Rong Pan, and Chin-Yew Lin. 2019. A Simple Recipe towards Reducing Hallucination in Neural Surface Realisation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 2673–2679. https://doi.org/10.18653/v1/P19-1256
  194. 194.Franz Josef Och. 2003. Minimum Error Rate Training in Statistical Machine Translation. In Proceedings of the 41st annual meeting of the Association for Computational Linguistics. 160–167.
  195. 195.OpenAI. 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774 (3 2023). http://arxiv.org/abs/2303.08774
  196. 196.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray John Schulman Jacob Hilton Fraser Kelton Luke Miller Maddie Simens Amanda Askell, Peter Welinder Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35 (2022), 27730–27744.
  197. 197.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding Factuality in Abstractive Summarization with FRANK: A Benchmark for Factuality Metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4812–4829.
  198. 198.Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020. ToTTo: A Controlled Table-To-Text Generation Dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 1173–1186.
  199. 199.Prasanna Parthasarathi, Koustuv Sinha, Joelle Pineau, and Adina Williams. 2021. Sometimes We Want Ungrammatical Translations. In Findings of the Association for Computational Linguistics: EMNLP 2021. Association for Computational Linguistics, Punta Cana, Dominican Republic, 3205–3227. https://doi.org/10.18653/v1/2021.findings-emnlp.275
  200. 200.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only The Falcon LLM team The RefinedWeb dataset for Falcon LLM. arXiv preprint arXiv:2306.01116 (2023). https://arxiv.org/pdf/2306.01116.pdf
  201. 201.Fabio Petroni, Tim Rocktàschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language Models as Knowledge Bases?. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics, Hong Kong, China, 2463–2473. https://doi.org/10.18653/v1/D19-1250
  202. 202.Pouya Pezeshkpour. 2023. Measuring and Modifying Factual Knowledge in Large Language Models. CoRR abs/2306.06264 (2023). https://doi.org/10.48550/arXiv.2306.06264 arXiv:2306.06264
  203. 203.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis Only Baselines in Natural Language Inference. In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics. 180–191.
  204. 204.Kashyap Popat, Subhabrata Mukherjee, Jannik Strötgen, and Gerhard Weikum. 2016. Credibility Assessment of Textual Claims on the Web. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management. 2173–2178.
  205. 205.Kashyap Popat, Subhabrata Mukherjee, Andrew Yates, and Gerhard Weikum. 2018. DeClarE: Debunking Fake News and False Claims using Evidence-Aware Deep Learning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Brussels, Belgium, 22–32. https://doi.org/10.18653/v1/D18-1003
  206. 206.Matt Post and David Vilar. 2018. Fast Lexically Constrained Decoding with Dynamic Beam Allocation for Neural Machine Translation. NAACL HLT 2018 - 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference 1 (4 2018), 1314–1324. https://doi.org/10.18653/v1/n18-1119
  207. 207.Ratish Puduppully, Li Dong, and Mirella Lapata. 2019. Data-to-text generation with content selection and planning. In Proceedings of the AAAI conference on artificial intelligence.
  208. 208.Ratish Puduppully and Mirella Lapata. 2021. Data-to-text generation with macro planning. Transactions of the Association for Computational Linguistics (2021).
  209. 209.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. OpenAI Blog 1, 8 (2019), 9.
  210. 210.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research 21, 140 (2020), 1–67. http://jmlr.org/papers/v21/20-074.html
  211. 211.Harsh Raj, Domenic Rosati, and Subhabrata Majumdar. 2022. Measuring Reliability of Large Language Models through Semantic Consistency. CoRR abs/2211.05853 (2022). https://doi.org/10.48550/arXiv.2211.05853 arXiv:2211.05853
  212. 212.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don’t Know: Unanswerable Questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics.
  213. 213.Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016. Sequence Level Training with Recurrent Neural Networks. ICLR (2016).
  214. 214.Hannah Rashkin, David Reitter, Gaurav Singh Tomar, and Dipanjan Das. 2021. Increasing Faithfulness in Knowledge-Grounded Dialogue with Controllable Features. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, Online, 704–718. https://doi.org/10.18653/v1/2021.acl-long.58
  215. 215.Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt. 2021. The Curious Case of Hallucinations in Neural Machine Translation. (4 2021), 1172–1183. https://arxiv.org/abs/2104.06683v1
  216. 216.Clément Rebuffel, Marco Roberti, Laure Soulier, Geoffrey Scoutheeten, Rossella Cancelliere, and Patrick Gallinari. 2022. Controlling hallucinations at word level in data-to-text generation. Data Mining and Knowledge Discovery 36, 1 (2022), 318–354.
  217. 217.Clément Rebuffel, Thomas Scialom, Laure Soulier, Benjamin Piwowarski, Sylvain Lamprier, Jacopo Staiano, Geoffrey Scoutheeten, and Patrick Gallinari. 2021. Data-QuestEval: A Reference-less Metric for Data-to-Text Semantic Evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.
  218. 218.Ehud Reiter. 2018. A Structured Review of the Validity of BLEU. Computational Linguistics 44, 3 (Sept. 2018), 393–401. https://doi.org/10.1162/coli_a_00322
  219. 219.Ehud Reiter and Robert Dale. 1997. Building Applied Natural Language Generation Systems. Natural Language Engineering 3, 1 (1997), 57–87.
  220. 220.Ruiyang Ren, Yuhao Wang, Yingqi Qu, Wayne Xin Zhao, Jing Liu, Hao Tian, Hua Wu, Ji-Rong Wen, and Haifeng Wang. 2023. Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation. CoRR abs/2307.11019 (2023). https://doi.org/10.48550/arXiv.2307.11019 arXiv:2307.11019
  221. 221.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How Much Knowledge Can You Pack Into the Parameters of a Language Model?. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, Online, 5418–5426. https://doi.org/10.18653/v1/2020.emnlp-main.437
  222. 222.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018. Object Hallucination in Image Captioning. In EMNLP.
  223. 223.Stephen Roller, Y-Lan Boureau, Jason Weston, Antoine Bordes, Emily Dinan, Angela Fan, David Gunning, Da Ju, Margaret Li, Spencer Poff, et al. 2020. Open-Domain Conversational Agents: Current Progress, Open Problems, and Future Directions. arXiv preprint arXiv:2006.12442 (2020).
  224. 224.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, et al. 2021. Recipes for Building an Open-Domain Chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 300–325.
  225. 225.Masoud Jalili Sabet, Philipp Dufter, François Yvon, and Hinrich Schütze. 2020. SimAlign: High Quality Word Alignments Without Parallel Training Data Using Static and Contextualized Embeddings. In Findings of the Association for Computational Linguistics: EMNLP 2020. 1627–1643.
  226. 226.Sashank Santhanam, Behnam Hedayatnia, Spandana Gella, Aishwarya Padmakumar, Seokhwan Kim, Yang Liu, and Dilek Hakkani-Tur. 2021. Rome was built in 1776: A Case Study on Factual Correctness in Knowledge-Grounded Response Generation. arXiv preprint arXiv:2110.05456 (2021).
  227. 227.Thomas Scialom, Paul-Alexis Dray, Patrick Gallinari, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, and Alex Wang. 2021. Questeval: Summarization asks for fact-based evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.
  228. 228.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get To The Point: Summarization with Pointer-Generator Networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Vancouver, Canada, 1073–1083. https://doi.org/10.18653/v1/P17-1099
  229. 229.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning Robust Metrics for Text Generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 7881–7892.
  230. 230.Lei Shen, Haolan Zhan, Xin Shen, Hongshen Chen, Xiaofang Zhao, and Xiaodan Zhu. 2021. Identifying Untrustworthy Samples: Data Filtering for Open-domain Dialogues with Bayesian Optimization. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management. 1598–1608.
  231. 231.Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Scott Yih. 2023. Trusting Your Evidence: Hallucinate Less with Context-aware Decoding. arXiv preprint arXiv:2305.14739 (2023).
  232. 232.Noah Shinn, Beck Labash, and Ashwin Gopinath. 2023. Reflexion: an autonomous agent with dynamic memory and self-reflection. arXiv preprint arXiv:2303.11366 (2023).
  233. 233.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval Augmentation Reduces Hallucination in Conversation. EMNLP (2021).
  234. 234.Haoyu Song, Wei-Nan Zhang, Jingwen Hu, and Ting Liu. 2020. Generating Persona Consistent Dialogues by Exploiting Natural Language Inference. Proceedings of the AAAI Conference on Artificial Intelligence 34, 05 (Apr. 2020), 8878–8885. https://doi.org/10.1609/aaai.v34i05.6417
  235. 235.Kaiqiang Song, Logan Lebanoff, Qipeng Guo, Xipeng Qiu, Xiangyang Xue, Chen Li, Dong Yu, and Fei Liu. 2020. Joint Parsing and Generation for Abstractive Summarization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 8894–8901.
  236. 236.Kai Song, Yue Zhang, Heng Yu, Weihua Luo, Kun Wang, and Min Zhang. 2019. Code-Switching for Enhancing NMT with Pre-Specified Translation. NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference 1 (4 2019), 449–459. https://arxiv.org/abs/1904.09107v4
  237. 237.Dan Su, Xiaoguang Li, Jindi Zhang, Lifeng Shang, Xin Jiang, Qun Liu, and Pascale Fung. 2022. Read before Generate! Faithful Long Form Question Answering with Machine Reading. arXiv:2203.00343 [cs.CL]
  238. 238.Hui Su, Xiaoyu Shen, Sanqiang Zhao, Zhou Xiao, Pengwei Hu, Cheng Niu, and Jie Zhou. 2020. Diversifying Dialogue Generation with Non-Conversational Text. In 58th Annual Meeting of the Association for Computational Linguistics. ACL, 7087–7097.
  239. 239.Yixuan Su, David Vandyke, Sihui Wang, Yimai Fang, and Nigel Collier. 2021. Plan-then-Generate: Controlled Data-to-Text Generation via Planning. Findings of EMNLP (2021).
  240. 240.Lya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura, and Hiroya Takamura. 2021. Towards Table-to-text Generation with Numerical Reasoning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 1451–1465.
  241. 241.Kai Sun, Yifan Ethan Xu, Hanwen Zha, Yue Liu, and Xin Luna Dong. 2023. Head-to-Tail: How Knowledgeable are Large Language Models (LLM)? A.K.A. Will LLMs Replace Knowledge Graphs? CoRR abs/2308.10168 (2023). https://doi.org/10.48550/arXiv.2308.10168 arXiv:2308.10168
  242. 242.Yanli Sun. 2010. Mining the Correlation between Human and Automatic Evaluation at Sentence Level.. In LREC.
  243. 243.Raymond Hendy Susanto, Shamil Chollampatt, and Liling Tan. 2020. Lexically Constrained Neural Machine Translation with Levenshtein Transformer. (7 2020), 3536–3543. https://doi.org/10.18653/V1/2020.ACL-MAIN.325
  244. 244.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. Advances in neural information processing systems 27 (2014).
  245. 245.Xiangru Tang, Arjun Nair, Borui Wang, Bingyao Wang, Jai Desai, Aaron Wade, Haoran Li, Asli Celikyilmaz, Yashar Mehdad, and Dragomir Radev. 2021. CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning. arXiv preprint arXiv:2112.08713 (2021).
  246. 246.Avijit Thawani, Jay Pujara, Filip Ilievski, and Pedro Szekely. 2021. Representing Numbers in NLP: a Survey and a Vision. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 644–656.
  247. 247.Craig Thomson and Ehud Reiter. 2020. A Gold Standard Methodology for Evaluating Accuracy in Data-To-Text Systems. In Proceedings of the 13th International Conference on Natural Language Generation. 158–168.
  248. 248.James Thorne and Andreas Vlachos. 2018. Automated Fact Checking: Task Formulations, Methods and Future Directions. In Proceedings of the 27th International Conference on Computational Linguistics, COLING 2018, Santa Fe, New Mexico, USA, August 20-26, 2018, Emily M. Bender, Leon Derczynski, and Pierre Isabelle (Eds.). Association for Computational Linguistics, 3346–3359. https://aclanthology.org/C18-1283/
  249. 249.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a Large-scale Dataset for Fact Extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). Association for Computational Linguistics, New Orleans, Louisiana, 809–819. https://doi.org/10.18653/v1/N18-1074
  250. 250.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2019. Evaluating adversarial attacks against multiple fact verification systems. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics, 2944–2953. https://doi.org/10.18653/v1/D19-1292
  251. 251.Ran Tian, Shashi Narayan, Thibault Sellam, and Ankur P. Parikh. 2020. Sticking to the Facts: Confident Decoding for Faithful Data-to-Text Generation. arXiv:1910.08684 [cs.CL]
  252. 252.Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. 2024. Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs. arXiv preprint arXiv:2401.06209 (2024).
  253. 253.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation Language Models. arXiv preprint arXiv:2302.13971 (2023). https://arxiv.org/pdf/2302.13971.pdf
  254. 254.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv preprint arXiv:2307.09288 (7 2023). http://arxiv.org/abs/2307.09288
  255. 255.Van-Khanh Tran and Le-Minh Nguyen. 2017. Natural Language Generation for Spoken Dialogue System using RNN Encoder-Decoder Networks. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017). Association for Computational Linguistics, Vancouver, Canada, 442–451. https://doi.org/10.18653/v1/K17-1044
  256. 256.Zhaopeng Tu, Yang Liu, Lifeng Shang, Xiaohua Liu, and Hang Li. 2017. Neural Machine Translation with Reconstruction. In Thirty-First AAAI Conference on Artificial Intelligence.
  257. 257.Zhaopeng Tu, Zhengdong Lu, Yang Liu, Xiaohua Liu, and Hang Li. 2016. Modeling Coverage for Neural Machine Translation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 76–85.
  258. 258.Victor Uc-Cetina, Nicolas Navarro-Guerrero, Anabel Martin-Gonzalez, Cornelius Weber, and Stefan Wermter. 2021. Survey on reinforcement learning for language processing. arXiv preprint arXiv:2104.05565 (2021).
  259. 259.Logesh Kumar Umapathi, Ankit Pal, and Malaikannan Sankarasubbu. 2023. Med-HALT: Medical Domain Hallucination Test for Large Language Models. CoRR abs/2307.15343 (2023). https://doi.org/10.48550/arXiv.2307.15343 arXiv:2307.15343
  260. 260.Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. 2023. A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation. CoRR abs/2307.03987 (2023). https://doi.org/10.48550/arXiv.2307.03987 arXiv:2307.03987
  261. 261.Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. 2023. A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation. CoRR abs/2307.03987 (2023). https://doi.org/10.48550/arXiv.2307.03987 arXiv:2307.03987
  262. 262.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. 5998–6008. https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf
  263. 263.Oriol Vinyals and Quoc Le. 2015. A Neural Conversational Model. ICML Deep Learning Workshop (2015).
  264. 264.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and Answering Questions to Evaluate the Factual Consistency of Summaries. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (2020).
  265. 265.Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2023. DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models. CoRR abs/2306.11698 (2023). https://doi.org/10.48550/arXiv.2306.11698 arXiv:2306.11698
  266. 266.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  267. 267.Chaojun Wang and Rico Sennrich. 2020. On Exposure Bias, Hallucination and Domain Shift in Neural Machine Translation. In 2020 Annual Conference of the Association for Computational Linguistics. Association for Computational Linguistics (ACL), 3544–3552.
  268. 268.Hongmin Wang. 2019. Revisiting Challenges in Data-to-Text Generation with Fact Grounding. In Proceedings of the 12th International Conference on Natural Language Generation. 311–322.
  269. 269.Peng Wang, Junyang Lin, An Yang, Chang Zhou, Yichang Zhang, Jingren Zhou, and Hongxia Yang. 2021. Sketch and Refine: Towards Faithful and Informative Table-to-Text Generation. ACL (2021).
  270. 270.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework. CoRR abs/2202.03052 (2022).
  271. 271.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei. 2022. Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks. ArXiv abs/2208.10442 (2022).
  272. 272.Xu Wang, Hainan Zhang, Shuai Zhao, Yanyan Zou, Hongshen Chen, Zhuoye Ding, Bo Cheng, and Yanyan Lan. 2021. FCM: A Fine-grained Comparison Model for Multi-turn Dialogue Reasoning. EMNLP Findings (2021).
  273. 273.Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, and Heng Ji. 2023. Unleashing cognitive synergy in large language models: A task-solving agent through multi-persona selfcollaboration. arXiv preprint arXiv:2307.05300 1, 2 (2023), 3.
  274. 274.Zhenyi Wang, Xiaoyang Wang, Bang An, Dong Yu, and Changyou Chen. 2020. Towards Faithful Neural Table-to-Text Generation with Content-Matching Constraints. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 1072–1086.
  275. 275.Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2022. SimVLM: Simple Visual Language Model Pretraining with Weak Supervision. ArXiv abs/2108.10904 (2022).
  276. 276.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682 (2022).
  277. 277.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2019. Neural Text Generation With Unlikelihood Training. In International Conference on Learning Representations.
  278. 278.Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho. 2019. Dialogue Natural Language Inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 3731–3741. https://doi.org/10.18653/v1/P19-1363
  279. 279.Tsung-Hsien Wen, Milica Gašić, Dongho Kim, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015. Stochastic Language Generation in Dialogue using Recurrent Neural Networks with Convolutional Sentence Reranking. In Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue. Association for Computational Linguistics, Prague, Czech Republic, 275–284. https://doi.org/10.18653/v1/W15-4639
  280. 280.Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Pei-hao Su, David Vandyke, and Steve J Young. 2015. Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems. In EMNLP.
  281. 281.Rongxiang Weng, Heng Yu, Xiangpeng Wei, and Weihua Luo. 2020. Towards Enhancing Faithfulness for Neural Machine Translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2675–2684.
  282. 282.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). Association for Computational Linguistics, New Orleans, Louisiana, 1112–1122. https://doi.org/10.18653/v1/N18-1101
  283. 283.Sam Wiseman, Stuart Shieber, and Alexander Rush. 2017. Challenges in Data-to-Document Generation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Copenhagen, Denmark, 2253–2263. https://doi.org/10.18653/v1/D17-1239
  284. 284.Thomas Wolf, Victor Sanh, Julien Chaumond, and Clement Delangue. 2019. TransferTransfo: A Transfer Learning Approach for Neural Network Based Conversational Agents. CoRR abs/1901.08149 (2019). arXiv:1901.08149 http://arxiv.org/abs/1901.08149
  285. 285.Zeqiu Wu, Michel Galley, Chris Brockett, Yizhe Zhang, Xiang Gao, Chris Quirk, Rik Koncel-Kedziorski, Jianfeng Gao, Hannaneh Hajishirzi, Mari Ostendorf, et al. 2021. A Controllable Model of Grounded Response Generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 14085–14093.
  286. 286.Yijun Xiao and William Yang Wang. 2021. On Hallucination and Predictive Uncertainty in Conditional Language Generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2734–2744.
  287. 287.Jing Xu, Arthur D. Szlam, and Jason Weston. 2021. Beyond Goldfish Memory: Long-Term Open-Domain Conversation. ArXiv abs/2107.07567 (2021).
  288. 288.Weijia Xu and Marine Carpuat. 2020. EDITOR: an Edit-Based Transformer with Repositioning for Neural Machine Translation with Soft Lexical Constraints. Transactions of the Association for Computational Linguistics 9 (11 2020), 311–328. https://doi.org/10.1162/tacl_a_00368d3/2021.
  289. 289.Weijia Xu and Marine Carpuat. 2021. Rule-based Morphological Inflection Improves Neural Terminology Translation. (9 2021), 5902–5914. https://doi.org/10.18653/v1/2021.emnlp-main.477
  290. 290.Weijia Xu, Xing Niu, and Marine Carpuat. 2019. Differentiable Sampling with Flexible Reference Word Order for Neural Machine Translation. NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference 1 (4 2019), 2047–2053. https://doi.org/10.18653/v1/n19-1207
  291. 291.Xinnuo Xu, Ondřej Dušek, Verena Rieser, and Ioannis Konstas. 2021. AGGGEN: Ordering and Aggregating while Generating. roceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL2021) (2021).
  292. 292.Yan Xu, Etsuko Ishii, Samuel Cahyawijaya, Zihan Liu, Genta Indra Winata, Andrea Madotto, Dan Su, and Pascale Fung. 2021. Retrieval-Free Knowledge-Grounded Dialogue Response Generation with Adapters. arXiv preprint arXiv:2105.06232 (2021).
  293. 293.Steve Yadlowsky, Lyric Doshi, and Nilesh Tripuraneni. 2023. Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models. CoRR abs/2311.00871 (2023). https://doi.org/10.48550/ARXIV.2311.00871 arXiv:2311.00871
  294. 294.Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2023. Alignment for Honesty. https://api.semanticscholar.org/CorpusID:266174420
  295. 295.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2369–2380.
  296. 296.Semih Yavuz, Abhinav Rastogi, Guan-Lin Chao, and Dilek Hakkani-Tur. 2019. DEEPCOPY: Grounded Response Generation with Hierarchical Pointer Networks. In Proceedings of SIGdial.
  297. 297.Jun Yin, Xin Jiang, Zhengdong Lu, Lifeng Shang, Hang Li, and Xiaoming Li. 2016. Neural generative question answering. In Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence. 2972–2978.
  298. 298.Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. 2023. Do Large Language Models Know What They Don’t Know?. In Findings of the Association for Computational Linguistics: ACL 2023. Association for Computational Linguistics, Toronto, Canada, 8653–8665. https://doi.org/10.18653/v1/2023.findings-acl.551
  299. 299.Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. 2023. Do Large Language Models Know What They Don’t Know?. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023, Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, 8653–8665. https://doi.org/10.18653/v1/2023.findings-acl.551
  300. 300.Takuma Yoneda, Jeff Mitchell, Johannes Welbl, Pontus Stenetorp, and Sebastian Riedel. 2018. UCL Machine Reading Group: Four Factor Framework For Fact Finding (HexaF). In Proceedings of the First Workshop on Fact Extraction and VERification (FEVER). Association for Computational Linguistics, Brussels, Belgium, 97–102. https://doi.org/10.18653/v1/W18-5515
  301. 301.Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-li, Xin Lv, Hao Peng, Zijun Yao, Xiaohan Zhang, Hanming Li, Chunyang Li, Zheyuan Zhang, Yushi Bai, Yantao Liu, Amy Xin, Nianyi Lin, Kaifeng Yun, Linlu Gong, Jianhui Chen, Zhili Wu, Yunjia Qi, Weikai Li, Yong Guan, Kaisheng Zeng, Ji Qi, Hailong Jin, Jinxin Liu, Yu Gu, Yuan Yao, Ning Ding, Lei Hou, Zhiyuan Liu, Bin Xu, Jie Tang, and Juanzi Li. 2023. KoLA: Carefully Benchmarking World Knowledge of Large Language Models. CoRR abs/2306.09296 (2023). https://doi.org/10.48550/arXiv.2306.09296 arXiv:2306.09296
  302. 302.Tiezheng Yu, Zihan Liu, and Pascale Fung. 2021. AdaptSum: Towards Low-Resource Domain Adaptation for Abstractive Summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 5892–5904.
  303. 303.Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, and Tat-Seng Chua. 2023. RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback. CoRR abs/2312.00849 (2023). https://doi.org/10.48550/ARXIV.2312.00849 arXiv:2312.00849
  304. 304.Wenhao Yu, Zhihan Zhang, Zhenwen Liang, Meng Jiang, and Ashish Sabharwal. 2023. Improving Language Models via Plug-and-Play Retrieval Feedback. arXiv preprint arXiv:2305.14002 (2023).
  305. 305.Xiang Yue, Boshi Wang, Kai Zhang, Ziru Chen, Yu Su, and Huan Sun. 2023. Automatic Evaluation of Attribution by Large Language Models. CoRR abs/2305.06311 (2023). https://doi.org/10.48550/arXiv.2305.06311 arXiv:2305.06311
  306. 306.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020. Big Bird: Transformers for Longer Sequences. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 17283–17297. https://proceedings.neurips.cc/paper/2020/file/c8512d142a2d849725f31a9a7a361ab9-Paper.pdf
  307. 307.Yury Zemlyanskiy and Fei Sha. 2018. Aiming to Know You Better Perhaps Makes Me a More Engaging Dialogue Partner. In Proceedings of the 22nd Conference on Computational Natural Language Learning. Association for Computational Linguistics, Brussels, Belgium, 551–561. https://doi.org/10.18653/v1/K18-1053
  308. 308.Yuexiang Zhai, Shengbang Tong, Xiao Li, Mu Cai, Qing Qu, Yong Jae Lee, and Yi Ma. 2023. Investigating the catastrophic forgetting in multimodal large language models. arXiv preprint arXiv:2309.10313 (2023).
  309. 309.Chen Zhang, Grandee Lee, Luis Fernando D’Haro, and Haizhou Li. 2021. D-score: Holistic dialogue evaluation without reference. IEEE/ACM Transactions on Audio, Speech, and Language Processing 29 (2021), 2502–2516.
  310. 310.Hongguang Zhang, Jing Zhang, and Piotr Koniusz. 2019. Few-Shot Learning via Saliency-Guided Hallucination of Samples. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019. Computer Vision Foundation / IEEE, 2770–2779. https://doi.org/10.1109/CVPR.2019.00288
  311. 311.Jiacheng Zhang, Huanbo Luan, Maosong Sun, Feifei Zhai, Jingfang Xu, and Yang Liu. 2021. Neural Machine Translation with Explicit Phrase Alignment. IEEE/ACM Transactions on Audio, Speech, and Language Processing 29 (2021), 1001–1010.
  312. 312.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing Dialogue Agents: I have a dog, do you have pets too?. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2204–2213.
  313. 313.Shuo Zhang, Liangming Pan, Junzhou Zhao, and William Yang Wang. 2023. Mitigating language model hallucination with interactive question-knowledge alignment. arXiv preprint arXiv:2305.13669 (2023).
  314. 314.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. BERTScore: Evaluating Text Generation with BERT. In International Conference on Learning Representations.
  315. 315.Xikun Zhang, Deepak Ramachandran, Ian Tenney, Yanai Elazar, and Dan Roth. 2020. Do Language Embeddings capture Scales?. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP. 292–299.
  316. 316.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2023. Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. arXiv preprint arXiv:2309.01219 (9 2023). http://arxiv.org/abs/2309.01219
  317. 317.Yuhao Zhang, Derek Merck, Emily Tsai, Christopher D Manning, and Curtis Langlotz. 2020. Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 5108–5120.
  318. 318.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020. DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations. Association for Computational Linguistics, Online, 270–278. https://doi.org/10.18653/v1/2020.acl-demos.30
  319. 319.Jing Zhao, Junwei Bao, Yifan Wang, Yongwei Zhou, Youzheng Wu, Xiaodong He, and Bowen Zhou. 2021. RoR: Read-over-Read for Long Document Machine Reading Comprehension. In Findings of the Association for Computational Linguistics: EMNLP 2021. Association for Computational Linguistics, Punta Cana, Dominican Republic, 1862–1872. https://aclanthology.org/2021.findings-emnlp.160
  320. 320.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023).
  321. 321.Zheng Zhao, Shay B. Cohen, and Bonnie Webber. 2020. Reducing Quantity Hallucinations in Abstractive Summarization. In Findings of the Association for Computational Linguistics: EMNLP 2020. Association for Computational Linguistics, Online, 2237–2249. https://doi.org/10.18653/v1/2020.findings-emnlp.203
  322. 322.Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, et al. 2021. QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 5905–5921.
  323. 323.Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2024. Lima: Less is more for alignment. Advances in Neural Information Processing Systems 36 (2024).
  324. 324.Chunting Zhou, Xuezhe Ma, Di Wang, and Graham Neubig. 2019. Density Matching for Bilingual Word Embedding. In NAACL.
  325. 325.Chunting Zhou, Graham Neubig, Jiatao Gu, Mona Diab, Francisco Guzmán, Luke Zettlemoyer, and Marjan Ghazvininejad. 2021. Detecting Hallucinated Content in Conditional Neural Sequence Generation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. 1393–1404.
  326. 326.Kangyan Zhou, Shrimai Prabhumoye, and Alan W Black. 2018. A Dataset for Document Grounded Conversations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 708–713.
  327. 327.Pei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, and Dilek Hakkani-Tur. 2021. Think Before You Speak: Using Self-talk to Generate Implicit Commonsense Knowledge for Response Generation. arXiv preprint arXiv:2110.08501 (2021).
  328. 328.Chenguang Zhu, William Hinthorn, Ruochen Xu, Qingkai Zeng, Michael Zeng, Xuedong Huang, and Meng Jiang. 2021. Enhancing Factual Consistency of Abstractive Summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 718–733.
  329. 329.Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al. 2023. PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts. arXiv preprint arXiv:2306.04528 (2023).

Citation

MLA
Ji, Z., et al. “Survey of Hallucination in Natural Language Generation”. ACM Computing Surveys, vol. 55, no. 12, 2023, pp. 1–8, https://doi.org/10.1145/3571730.
APA
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12), 1–38. https://doi.org/10.1145/3571730
Chicago
Ji, Z., N. Lee, R. Frieske, et al. 2023. “Survey of Hallucination in Natural Language Generation”. ACM Computing Surveys 55 (12): 1–38. https://doi.org/10.1145/3571730.
Harvard
Ji, Z. et al. (2023) “Survey of Hallucination in Natural Language Generation”, ACM Computing Surveys, 55(12), pp. 1–38. Available at: https://doi.org/10.1145/3571730.
Vancouver
1. Ji Z, Lee N, Frieske R, Yu T, Su D, Xu Y, Ishii E, Bang YJ, Madotto A, Fung P (2023) Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55:1–38

BibTeX

@article{Ji_2023, title={Survey of Hallucination in Natural Language Generation}, volume={55}, ISSN={1557-7341}, url={http://dx.doi.org/10.1145/3571730}, DOI={10.1145/3571730}, number={12}, journal={ACM Computing Surveys}, publisher={Association for Computing Machinery (ACM)}, author={Ji, Ziwei and Lee, Nayeon and Frieske, Rita and Yu, Tiezheng and Su, Dan and Xu, Yan and Ishii, Etsuko and Bang, Ye Jin and Madotto, Andrea and Fung, Pascale}, year={2023}, month=Mar, pages={1–38} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/