Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus

Tianhang ZhangLin QiuQipeng GuoCheng DengYue ZhangZheng ZhangChenghu ZhouXinbing WangLuoyi Fu

article2023EMNLP129 citations

Proposes a reference-free hallucination detection framework that evaluates LLM-generated text without extra sampling or external retrieval by modeling uncertainty through keyword filtering, attention-based error propagation, and entity-specific frequency adjustments.

Listen

Large language models frequently generate factually incorrect or nonsensical text, known as hallucinations. This tendency creates major reliability risks in critical domains such as medicine, finance, and education. Existing detection techniques usually depend on external knowledge retrieval or generate multiple sampled responses to check for consistency. These existing approaches are computationally slow, resource-heavy, and difficult to deploy in low-latency environments.

The article develops and evaluates a reference-free, uncertainty-based method to detect hallucinations directly from generated text. The objective is to determine whether an open proxy language model can accurately evaluate factuality without external knowledge bases, additional sampled responses, or task-specific fine-tuning.

The authors design an approach that models three aspects of human fact-checking: isolating informative keywords and named entities, penalizing subsequent words when they rely on prior untrustworthy text through attention propagation, and adjusting word probabilities based on entity categories and token frequency. The method was evaluated primarily on the WikiBio GPT-3 benchmark consisting of 1,908 annotated sentences across 238 text passages, using 22 different open-source proxy models of varying sizes. Supplementary tests were also performed on standard summarization benchmarks.

The findings show that the proposed framework consistently outperforms existing retrieval-free baselines across all evaluated metrics. When using a 30-billion-parameter open-source proxy model, the approach achieved sentence-level non-factual precision-recall scores of roughly 89.8% and human correlation scores exceeding 73% to 77%, surpassing multi-query sampling systems like SelfCheckGPT. Within model families, detection quality generally increases with model scale, though gains plateau at larger sizes where a 30-billion model matched or slightly exceeded a 65-billion model. Furthermore, relatively small 7-billion parameter models equipped with these focus mechanisms rivaled the raw internal uncertainty metrics of large proprietary systems.

These results demonstrate that factuality evaluation can be performed cost-effectively using standalone proxy models. Organizations deploying generative artificial intelligence can lower operational expenses and latency by replacing costly multi-query consistency checks with single-pass proxy evaluation. This enables more viable real-time monitoring and automated fact-checking pipelines across commercial deployments.

Decision-makers should consider integrating open proxy models into production safeguards rather than relying on expensive multi-sample query strategies. Future work should expand entity categories beyond standard toolkits, refine detection for broader non-entity grammatical errors, and ensure underlying proxy models receive regular knowledge updates so that assessments remain aligned with real-world facts.

arXiv: 2311.13230
Cover for Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus

Abstract

Large Language Models (LLMs) have gained significant popularity for their impressive performance across diverse fields. However, LLMs are prone to hallucinate untruthful or nonsensical outputs that fail to meet user expectations in many real-world applications. Existing works for detecting hallucinations in LLMs either rely on external knowledge for reference retrieval or require sampling multiple responses from the LLM for consistency verification, making these methods costly and inefficient. In this paper, we propose a novel reference-free, uncertainty-based method for detecting hallucinations in LLMs. Our approach imitates human focus in factuality checking from three aspects: 1) focus on the most informative and important keywords in the given text; 2) focus on the unreliable tokens in historical context which may lead to a cascade of hallucinations; and 3) focus on the token properties such as token type and token frequency. Experimental results on relevant datasets demonstrate the effectiveness of our proposed method, which achieves state-of-the-art performance across all the evaluation metrics and eliminates the need for additional information.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Hallucinations in Text Generation
  • 2.2 Hallucination Detection
  • 3 Methods
  • 3.1 Keywords selection
  • 3.2 Hallucination propagation
  • 3.3 Probability correction
  • 3.4 Putting things together
  • 4 Experiments and Results
  • 4.1 Experiment setting
  • 4.2 Main results
  • 4.3 Analysis
  • 4.4 Case study
  • 4.4.1 Non-factual cases detected by hallucination propagation
  • 4.4.2 Failure cases after entity type provision
  • 5 Conclusion
  • References
  • A Results for detecting hallucinations generated by small models
  • B Implementation Details
  • C More attention heat map cases
  • D Examples of passages with entity types provided
  • E Details of the proxy models
  • F Performance comparison of LLaMA family
  • G Hyper parameters analysis
  • H Additional Results

Knowls

  1. Knowl 1 — Focus-Enhanced Reference-Free Hallucination Detection Architecture

    model/method

    A proxy autoregressive language model is used to detect factual hallucinations in generated texts without requiring external reference retrieval or repeated stochastic sampling. Standard token generation probabilities from language models suffer from three structural discrepancies when aligned against human factuality checking:

    1. Varying token informativeness: Unweighted probability aggregation across all tokens incorporates syntactic and function-word noise. The focus framework restricts score aggregation to salient keywords K\mathcal{K}, specifically named entities and nouns.
    2. Overconfidence from exposure bias: Autoregressive self-attention causes models to condition on previous hallucinated tokens, assigning high generation probabilities to downstream cascade errors. The framework propagates uncertainty penalties from preceding keywords across normalized self-attention weights.
    3. Underconfidence from open topic branching and rare token bias: Factual tokens spanning multiple valid semantic continuations or rare vocabulary items receive low raw generation probabilities. The framework conditions probability distributions on extracted entity types via in-context tag insertion and recalibrates them using Inverse Document Frequency (IDF).
  2. Knowl 2 — Local and Global Uncertainty Scoring Formulation

    equation

    For a text sequence r=(t0,t1,…,t∣r∣−1)r = (t_0, t_1, \dots, t_{|r|-1}) over vocabulary V\mathcal{V}, the base hallucination score hih_i of token tit_i combines local negative log-likelihood and vocabulary-level predictive entropy:

    hi=−log⁡(pi(ti))+Hih_i = -\log(p_i(t_i)) + H_i

    where pi(v)p_i(v) denotes the predicted probability of token v∈Vv \in \mathcal{V} at position ii, and the predictive entropy HiH_i is computed as:

    Hi=2−∑v∈Vpi(v)log⁡2(pi(v))H_i = 2^{-\sum_{v \in \mathcal{V}} p_i(v) \log_2(p_i(v))}

    To compute the hallucination score hsh^s for a sentence ss containing ∣s∣|s| tokens, a weighted average is calculated over the set of salient keywords K\mathcal{K} (comprising 18 named entity types and non-entity nouns identified by a linguistic parser):

    hs=1∑i=0∣s∣−1I(ti∈K)∑i=0∣s∣−1I(ti∈K)⋅hih^s = \frac{1}{\sum_{i=0}^{|s|-1} \mathbb{I}(t_i \in \mathcal{K})} \sum_{i=0}^{|s|-1} \mathbb{I}(t_i \in \mathcal{K}) \cdot h_i

    where I(⋅)\mathbb{I}(\cdot) is an indicator function equal to 1 if ti∈Kt_i \in \mathcal{K} and 0 otherwise. Passage-level hallucination scores are computed by averaging hih_i across all keywords in the passage.

  3. Knowl 3 — Hallucination Propagation via Max-Pooled Self-Attention

    model/method

    To penalize tokens generated by attending to earlier hallucinated content (addressing overconfidence caused by exposure bias), the base token score hih_i is updated to an accumulated penalty score h^i\hat{h}_i:

    h^i=hi+I(ti∈K)⋅γ⋅pi\hat{h}_i = h_i + \mathbb{I}(t_i \in \mathcal{K}) \cdot \gamma \cdot p_i

    where K\mathcal{K} denotes the set of identified keywords, γ∈[0,1]\gamma \in [0, 1] is a discount coefficient (set to 0.90.9) that diminishes multi-hop propagation geometrically, and pip_i is the propagated penalty:

    pi=∑j=0i−1wi,jh^jp_i = \sum_{j=0}^{i-1} w_{i,j} \hat{h}_j

    The propagation weight wi,jw_{i,j} between token tit_i and preceding token tjt_j is computed by normalizing the max-pooled cross-layer and cross-head self-attention weights atti,j\text{att}_{i,j} over preceding keywords:

    wi,j=I(tj∈K)⋅atti,j∑k=0i−1I(tk∈K)⋅atti,kw_{i,j} = \frac{\mathbb{I}(t_j \in \mathcal{K}) \cdot \text{att}_{i,j}}{\sum_{k=0}^{i-1} \mathbb{I}(t_k \in \mathcal{K}) \cdot \text{att}_{i,k}}

    Tokens that are not keywords receive zero penalty and do not propagate penalties downstream.

  4. Knowl 4 — Probability Correction via Entity In-Context Conditioning and Token IDF

    model/method

    To mitigate model underconfidence on valid factual tokens, probability estimation is constrained to target semantic categories and adjusted for frequency. For a token sequence t0:n−1t_{0:n-1} with candidate set c(t0:i)c(t_{0:i}), Bayes' rule gives:

    p(ti∣t0:i−1,c(t0:i))=p(ti∣t0:i−1)∑v∈c(t0:i)p(v∣t0:i−1)p(t_i \mid t_{0:i-1}, c(t_{0:i})) = \frac{p(t_i \mid t_{0:i-1})}{\sum_{v \in c(t_{0:i})} p(v \mid t_{0:i-1})}

    To construct this candidate set in autoregressive models, entity type tags (e.g., <DATE>, <PERSON>, <ORG>) identified by a parser are prepended directly before each named entity in the input prompt and sequence. The candidate set c(t0:i)c(t_{0:i}) is approximated by taking tokens whose generation probability under the typed prompt exceeds threshold ρ=0.01\rho = 0.01.

    The resulting conditional probability p~(t)\tilde{p}(t) is subsequently corrected for token frequency bias using token Inverse Document Frequency (idf\text{idf}) derived from a 1,000,000 document sample of the RedPajama dataset:

    p^(t)=p~(t)⋅idf(t)∑v∈Vp~(v)⋅idf(v)\hat{p}(t) = \frac{\tilde{p}(t) \cdot \text{idf}(t)}{\sum_{v \in \mathcal{V}} \tilde{p}(v) \cdot \text{idf}(v)}

    The calibrated probability p^(t)\hat{p}(t) substitutes for standard token probability in negative log-likelihood and entropy calculations.

  5. Knowl 5 — Focus-Enhanced Hallucination Detection Algorithm

    algorithm
    Input: Evaluated text passage r=(t0,t1,…,tn−1)r = (t_0, t_1, \dots, t_{n-1}) partitioned into sentences SS, proxy language model MM, candidate probability threshold ρ=0.01\rho = 0.01, decay factor γ=0.9\gamma = 0.9, token inverse document frequency mapping idf\text{idf} over vocabulary V\mathcal{V}.
    Output: Sentence hallucination scores {hs}s∈S\{h^s\}_{s \in S} and passage hallucination score hrh^r.
    1. Identify keyword set K\mathcal{K} in rr (18 named entity categories and non-entity nouns) using a linguistic entity tagger.
    2. Insert entity type tags <TYPE> immediately before each named entity in rr to produce typed text rtyper_{\text{type}}.
    3. Execute a forward pass of MM on rtyper_{\text{type}} to obtain conditioned token probability distributions p~i\tilde{p}_i and max-pooled attention matrix att\text{att}.
    4. for i=0i = 0 to n−1n - 1 do:
         p^i(t)←p~i(t)⋅idf(t)∑v∈Vp~i(v)⋅idf(v)\hat{p}_i(t) \leftarrow \frac{\tilde{p}_i(t) \cdot \text{idf}(t)}{\sum_{v \in \mathcal{V}} \tilde{p}_i(v) \cdot \text{idf}(v)}
         Hi←2−∑v∈Vp^i(v)log⁡2(p^i(v))H_i \leftarrow 2^{-\sum_{v \in \mathcal{V}} \hat{p}_i(v) \log_2(\hat{p}_i(v))}
         hi←−log⁡(p^i(ti))+Hih_i \leftarrow -\log(\hat{p}_i(t_i)) + H_i
         if ti∈Kt_i \in \mathcal{K} then:
           for j=0j = 0 to i−1i - 1 do:
             wi,j←I(tj∈K)⋅atti,j∑k=0i−1I(tk∈K)⋅atti,kw_{i,j} \leftarrow \frac{\mathbb{I}(t_j \in \mathcal{K}) \cdot \text{att}_{i,j}}{\sum_{k=0}^{i-1} \mathbb{I}(t_k \in \mathcal{K}) \cdot \text{att}_{i,k}}
           pi←∑j=0i−1wi,jh^jp_i \leftarrow \sum_{j=0}^{i-1} w_{i,j} \hat{h}_j
           h^i←hi+γ⋅pi\hat{h}_i \leftarrow h_i + \gamma \cdot p_i
         else:
           h^i←hi\hat{h}_i \leftarrow h_i
    5. for each sentence s∈Ss \in S do:
         hs←∑ti∈s,ti∈Kh^i∑ti∈sI(ti∈K)h^s \leftarrow \frac{\sum_{t_i \in s, t_i \in \mathcal{K}} \hat{h}_i}{\sum_{t_i \in s} \mathbb{I}(t_i \in \mathcal{K})}
    6. hr←∑ti∈r,ti∈Kh^i∑ti∈rI(ti∈K)h^r \leftarrow \frac{\sum_{t_i \in r, t_i \in \mathcal{K}} \hat{h}_i}{\sum_{t_i \in r} \mathbb{I}(t_i \in \mathcal{K})}
    7. return {hs}s∈S,hr\{h^s\}_{s \in S}, h^r
  6. Knowl 6 — Hallucination Detection Performance on WikiBio GPT-3 Benchmark

    data/table

    On the WikiBio GPT-3 dataset (1,908 annotated sentences across 238 Wikipedia passages generated by text-davinci-003), the focus-enhanced detection framework using LLaMA models outperforms internal GPT-3 uncertainty measures and sampling-based SelfCheckGPT baselines across sentence-level AUC-PR and passage-level correlation metrics without generating additional stochastic responses.

    Method Sentence-level Metrics (AUC-PR) Passage-level Metrics
    NonFact NonFact* Factual Pearson Spearman
    GPT-3 Uncertainties
    Avg(−log⁡p)\text{Avg}(-\log p) 83.21 38.89 53.97 57.04 53.93
    Avg(H)\text{Avg}(H) 80.73 37.09 52.07 55.52 50.87
    Max(−log⁡p)\text{Max}(-\log p) 87.51 35.88 50.46 57.83 55.69
    Max(H)\text{Max}(H) 85.75 32.43 50.27 52.48 49.55
    SelfCheckGPT
    BERTScore 81.96 45.96 44.23 58.18 55.90
    QA 84.26 40.06 48.14 61.07 59.29
    Unigram (max) 85.63 41.04 58.47 64.71 64.91
    Combination 87.33 44.37 61.83 69.05 67.77
    Ours (Focus-Enhanced)
    LLaMA-7Bfocus_{\text{focus}} 84.26 40.20 57.04 64.47 54.73
    LLaMA-13Bfocus_{\text{focus}} 87.90 43.84 62.46 70.62 63.03
    LLaMA-30Bfocus_{\text{focus}} 89.79 48.80 65.69 77.15 73.24
    LLaMA-65Bfocus_{\text{focus}} 89.94 48.69 64.90 76.80 73.01

    LLaMA-30Bfocus_{\text{focus}} improves upon SelfCheckGPT-Combination by +2.46%+2.46\% in NonFact AUC-PR, +4.43%+4.43\% in NonFact* AUC-PR, +8.10%+8.10\% in Pearson correlation, and +5.47%+5.47\% in Spearman correlation.

  7. Knowl 7 — Ablation of Focus-Enhanced Components on LLaMA-30B

    data/table

    An ablation study on the WikiBio GPT-3 dataset using LLaMA-30B (with decay γ=0.9\gamma = 0.9 and threshold ρ=0.01\rho = 0.01) demonstrates the incremental effect of adding keyword filtering, attention penalty propagation, entity type in-context conditioning, and token IDF probability calibration.

    Method Configuration NoFac (AUC-PR) NoFac* (AUC-PR) Fact (AUC-PR) Pearson Spearman
    avg(h)\text{avg}(h) 82.07 41.47 47.22 51.03 37.29
    + keyword 83.01 41.57 45.82 56.07 44.77
    + penalty 86.68 45.27 54.93 59.08 55.84
    + entity type 88.89 46.92 65.12 76.82 71.49
    + token idf 89.79 48.80 65.69 77.15 73.24

    Keyword selection primarily improves passage-level metrics (+5.04%+5.04\% Pearson, +7.48%+7.48\% Spearman). Attention propagation penalty boosts Non-Factual AUC-PR by +3.67%+3.67\% by correcting overconfidence. Entity typing and token IDF together raise Factual AUC-PR by +10.76%+10.76\% and passage Pearson correlation by +18.07%+18.07\% over the penalty stage.

  8. Knowl 8 — Sensitivity to Propagation Decay Factor and Candidate Probability Threshold

    empirical result

    Empirical parameter sweeps on LLaMA-30B demonstrate the performance sensitivity to the propagation coefficient γ\gamma and candidate cutoff ρ\rho:

    • Attention decay coefficient γ\gamma: Performance across all metrics increases steadily as γ\gamma increases from 0.00.0 to 0.80.8, demonstrating the utility of multi-hop penalty propagation. Performance degrades when γ>0.8\gamma > 0.8, showing that excessively large decay factors over-penalize valid tokens following historical errors.
    • Candidate threshold ρ\rho: Evaluation across ρ∈[10−5,10−1]\rho \in [10^{-5}, 10^{-1}] reveals an optimum at ρ=0.01\rho = 0.01. Values of ρ>0.01\rho > 0.01 over-constrain the candidate set and discard valid tokens, while values of ρ<0.01\rho < 0.01 admit unrelated vocabulary noise into the denominator.
  9. Knowl 9 — Focus-Enhanced Detection on Small-Model Summarization Benchmarks

    empirical result

    When applied to small-model abstractive summaries from the SummaC benchmark (XSumFaith and FRANK) using LLaMA-30B-SFT as the proxy model:

    • On XSumFaith (984984 samples, 90.40%90.40\% hallucination rate), the full pipeline increases Non-Factual AUC-PR from 92.79%92.79\% (unmodified avg(h)\text{avg}(h)) to 95.13%95.13\%, Factual AUC-PR from 11.75%11.75\% to 18.86%18.86\%, and Balanced Accuracy from 57.65%57.65\% to 64.81%64.81\%.
    • On FRANK (1,2421,242 samples, 57.41%57.41\% hallucination rate, excluding unsupported OutE instances), hallucinations frequently consist of predicate, pronoun, and preposition errors rather than entity inaccuracies. Consequently, utilizing token negative log-probability avg(−log⁡p)\text{avg}(-\log p) across all tokens without keyword filtering or entropy summation achieves optimal results (Non-Factual AUC-PR 90.12%90.12\%, Factual AUC-PR 80.00%80.00\%, Balanced Accuracy 80.70%80.70\%, compared to the raw baseline Balanced Accuracy of 78.79%78.79\%).
  10. Knowl 10 — Limitations in Linguistic Tagging and Static Knowledge Cutoffs

    limitation

    The focus-enhanced hallucination detection approach exhibits two principal limitations:

    1. Parser errors and taxonomy limits: Keyword extraction and entity categorization rely on spaCy. Parsing errors (such as misclassifying creative work titles like 'The Great Ambition' as organizations) distort the conditioned candidate probabilities p~(t)\tilde{p}(t). Furthermore, standard 18-type NER ontologies do not span specialized real-world concepts such as vehicles, foods, or technical terminology.
    2. Static proxy knowledge horizons: Proxy language models rely on frozen pretraining corpora. When evaluating claims involving facts or entities that emerged after the proxy model's pretraining cutoff, the proxy assigns low generation probabilities to truthful statements.

Coverage note — Ablation tables for the other 21 proxy models in Appendix H were omitted because they duplicate the structural findings demonstrated on LLaMA-30B, and example passage texts in Appendix D were omitted as illustrative dataset samples.

References

  1. 1.Hussam Alkaissi and Samy I McFarlane. 2023. Artificial hallucinations in chatgpt: implications in scientific writing. Cureus, 15(2).
  2. 2.Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. Falcon-40B: an open large language model with state-of-the-art performance.
  3. 3.David Baidoo-Anu and Leticia Owusu Ansah. 2023. Education in the era of generative artificial intelligence (ai): Understanding the potential benefits of chatgpt in promoting teaching and learning. Available at SSRN 4337484.
  4. 4.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023.
  5. 5.Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015. Scheduled sampling for sequence prediction with recurrent neural networks. Advances in neural information processing systems, 28.
  6. 6.Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. 2022. Gpt-neox-20b: An open-source autoregressive language model. In Proceedings of BigScience Episode# 5–Workshop on Challenges & Perspectives in Creating Large Language Models, pages 95–136.
  7. 7.Meng Cao, Yue Dong, and Jackie Chi Kit Cheung. 2022. Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3340–3354.
  8. 8.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An opensource chatbot impressing gpt-4 with 90%* chatgpt quality.
  9. 9.Together Computer. 2023. Redpajama: An open source recipe to reproduce llama training dataset.
  10. 10.David Dale, Elena Voita, Loïc Barrault, and Marta R Costa-jussà. 2022. Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity even better. arXiv preprint arXiv:2212.08597.
  11. 11.David Demeter, Gregory Kimmel, and Doug Downey. 2020. Stolen probability: A structural weakness of neural language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2191–2197.
  12. 12.Yue Dong, John Wieting, and Pat Verga. 2022. Faithful to the document or to the world? mitigating hallucinations via entity-linked knowledge in abstractive summarization. arXiv preprint arXiv:2204.13761.
  13. 13.Nouha Dziri, Sivan Milton, Mo Yu, Osmar R Zaiane, and Siva Reddy. 2022. On the origin of hallucinations in conversational models: Is it the datasets or the models? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5271–5285.
  14. 14.Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. Summeval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
  15. 15.Tobias Falke, Leonardo FR Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019. Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2214–2220.
  16. 16.Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023. Trueteacher: Learning factual consistency evaluation with large language models. arXiv preprint arXiv:2305.11171.
  17. 17.Nuno M Guerreiro, Duarte Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, and André FT Martins. 2023. Hallucinations in large multilingual translation models. arXiv preprint arXiv:2303.16104.
  18. 18.Nuno M Guerreiro, Elena Voita, and André FT Martins. 2022. Looking for a needle in a haystack: A comprehensive study of hallucinations in neural machine translation. arXiv preprint arXiv:2208.05309.
  19. 19.Matthew Honnibal and Ines Montani. 2017. spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing. To appear, 7(1):411–420.
  20. 20.Dandan Huang, Leyang Cui, Sen Yang, Guangsheng Bao, Kun Wang, Jun Xie, and Yue Zhang. 2020. What have we achieved on text summarization? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 446–469.
  21. 21.Yichong Huang, Xiachong Feng, Xiaocheng Feng, and Bing Qin. 2021. The factual inconsistency problem in abstractive text summarization: A survey. arXiv preprint arXiv:2104.14839.
  22. 22.Touseef Iqbal and Shaima Qureshi. 2022. The survey: Text generation models in deep learning. Journal of King Saud University-Computer and Information Sciences, 34(6):2515–2528.
  23. 23.Mohd Javaid, Abid Haleem, and Ravi Pratap Singh. 2023. Chatgpt for healthcare services: An emerging stage for an innovative perspective. BenchCouncil Transactions on Benchmarks, Standards and Evaluations, page 100105.
  24. 24.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  25. 25.Zdeněk Kasner, Simon Mille, and Ondřej Dušek. 2021. Text-in-context: Token-level error detection for table-to-text generation. In Proceedings of the 14th International Conference on Natural Language Generation, pages 259–265.
  26. 26.Wojciech Kryściński, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346.
  27. 27.Philippe Laban, Tobias Schnabel, Paul N Bennett, and Marti A Hearst. 2022. Summac: Re-visiting nli-based models for inconsistency detection in summarization. Transactions of the Association for Computational Linguistics, 10:163–177.
  28. 28.Peter Lee, Sebastien Bubeck, and Joseph Petro. 2023. Benefits, limits, and risks of gpt-4 as an ai chatbot for medicine. New England Journal of Medicine, 388(13):1233–1239.
  29. 29.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880.
  30. 30.Jiongnan Liu, Jiajie Jin, Zihan Wang, Jiehan Cheng, Zhicheng Dou, and Ji-Rong Wen. 2023. Reta-llm: A retrieval-augmented large language model toolkit. arXiv preprint arXiv:2306.05212.
  31. 31.Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and William B Dolan. 2022. A token-level reference-free hallucination detection benchmark for free-form text generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6723–6737.
  32. 32.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7052–7063.
  33. 33.Alejandro Lopez-Lira and Yuehua Tang. 2023. Can chatgpt forecast stock price movements? return predictability and large language models. Return Predictability and Large Language Models (April 6, 2023).
  34. 34.Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. arXiv preprint arXiv:2303.08896.
  35. 35.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919.
  36. 36.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. arXiv preprint arXiv:2305.14251.
  37. 37.Niels Mündler, Jingxuan He, Slobodan Jenko, and Martin Vechev. 2023. Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation. arXiv preprint arXiv:2305.15852.
  38. 38.Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen Mckeown, and Bing Xiang. 2021. Entity-level factual consistency of abstractive text summarization. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2727–2733.
  39. 39.Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807.
  40. 40.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  41. 41.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with frank: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4812–4829.
  42. 42.Hannah Rashkin, David Reitter, Gaurav Singh Tomar, and Dipanjan Das. 2021. Increasing faithfulness in knowledge-grounded dialogue with controllable features. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 704–718.
  43. 43.Vikas Raunak, Siddharth Dalmia, Vivek Gupta, and Florian Metze. 2020. On long-tailed phenomena in neural machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3088–3095.
  44. 44.Clément Rebuffel, Marco Roberti, Laure Soulier, Geoffrey Scoutheeten, Rossella Cancelliere, and Patrick Gallinari. 2022. Controlling hallucinations at word level in data-to-text generation. Data Mining and Knowledge Discovery, pages 1–37.
  45. 45.Malik Sallam. 2023. The utility of chatgpt as an example of large language models in healthcare education, research and practice: Systematic review on the future perspectives and potential limitations. medRxiv, pages 2023–02.
  46. 46.Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023a. In chatgpt we trust? measuring and characterizing the reliability of chatgpt. arXiv preprint arXiv:2304.08979.
  47. 47.Yiqiu Shen, Laura Heacock, Jonathan Elias, Keith D Hentel, Beatriu Reig, George Shih, and Linda Moy. 2023b. Chatgpt and other large language models are double-edged swords.
  48. 48.Dan Su, Xiaoguang Li, Jindi Zhang, Lifeng Shang, Xin Jiang, Qun Liu, and Pascale Fung. 2022. Read before generate! faithful long form question answering with machine reading. In Findings of the Association for Computational Linguistics: ACL 2022, pages 744–756.
  49. 49.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  50. 50.Ahmed Tlili, Boulus Shehata, Michael Agyemang Adarkwah, Aras Bozkurt, Daniel T Hickey, Ronghuai Huang, and Brighter Agyemang. 2023. What if the devil is my guardian angel: Chatgpt as a case study of using chatbots in education. Smart Learning Environments, 10(1):15.
  51. 51.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  52. 52.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  53. 53.Liam van der Poel, Ryan Cotterell, and Clara Meister. 2022. Mutual information alleviates hallucinations in abstractive summarization. In EMNLP 2022. arXiv.
  54. 54.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  55. 55.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  56. 56.Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023. Bloomberggpt: A large language model for finance. arXiv preprint arXiv:2303.17564.
  57. 57.Yijun Xiao and William Yang Wang. 2021. On hallucination and predictive uncertainty in conditional language generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2734–2744.
  58. 58.Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J Martindale, and Marine Carpuat. 2023. Understanding and detecting hallucinations in neural machine translation via model introspection. arXiv preprint arXiv:2301.07779.
  59. 59.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.

Citation

MLA
Zhang, T., et al. “Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 915–32, https://doi.org/10.18653/v1/2023.emnlp-main.58.
APA
Zhang, T., Qiu, L., Guo, Q., Deng, C., Zhang, Y., Zhang, Z., Zhou, C., Wang, X., & Fu, L. (2023). Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 915–932. https://doi.org/10.18653/v1/2023.emnlp-main.58
Chicago
Zhang, T., L. Qiu, Q. Guo, et al. 2023. “Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 915–32. https://doi.org/10.18653/v1/2023.emnlp-main.58.
Harvard
Zhang, T. et al. (2023) “Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 915–932. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.58.
Vancouver
1. Zhang T, Qiu L, Guo Q, Deng C, Zhang Y, Zhang Z, Zhou C, Wang X, Fu L (2023) Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 915–932

BibTeX

@inproceedings{zhang-etal-2023-enhancing-uncertainty,
    title = "Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus",
    author = "Zhang, Tianhang  and
      Qiu, Lin  and
      Guo, Qipeng  and
      Deng, Cheng  and
      Zhang, Yue  and
      Zhang, Zheng  and
      Zhou, Chenghu  and
      Wang, Xinbing  and
      Fu, Luoyi",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.58/",
    doi = "10.18653/v1/2023.emnlp-main.58",
    pages = "915--932"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/