Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better

David DaleElena VoitaLoïc BarraultMarta R. Costa-jussà

article2023ACL103 citations

Demonstrates that measuring source token contribution from a translation model's internal mechanics doubles the detection accuracy of severe hallucinations without external tools, while cross-lingual sentence similarity models achieve an 80% precision gain across all hallucination types.

Listen

Machine translation systems occasionally generate hallucinations, which are outputs completely or partially detached from the original source text. While rare, these severe errors undermine user trust and present substantial operational risks in real-world deployment. Detecting and repairing naturally occurring hallucinations in production has proven difficult because standard automated quality estimation metrics frequently fail to flag severe pathologies, leading previous approaches to rely on artificial perturbations or external evaluation tools.

The article evaluates both internal model mechanisms and external tools to determine how effectively they can detect naturally occurring hallucinations and mitigate them at test time without relying on human reference translations. The authors analyze a clean test bed of German-to-English translations generated by a standard neural translation model, evaluating performance across a curated dataset of over 3,400 annotated examples containing naturally occurring translation pathologies.

Key findings show that internal model metrics alone perform remarkably well. Using a method called ALTI to calculate the percentage of source token contribution within the translation model doubles the detection precision for fully detached hallucinations compared to baseline sequence log-probability (reaching 67.4% precision at 90% recall compared to 31.0%). When external tools are permitted, cross-lingual sentence similarity models, specifically LaBSE, deliver the highest overall detection accuracy across all hallucination types, achieving an 80% improvement in precision over the baseline. Furthermore, evaluating hypotheses within a detect-then-rewrite pipeline reveals that generating candidates via Monte Carlo dropout and reranking them with sentence similarity cuts the human-judged hallucination rate from 53% down to 16%, while internal ALTI reranking matches the mitigation performance of complex external quality estimators, cutting the rate to 22%.

These findings demonstrate that organizations do not necessarily need computationally heavy external quality estimation pipelines to mitigate translation hallucinations. Because an internal measure of source contribution achieves comparable mitigation to state-of-the-art external estimators, organizations operating under compute constraints or translating low-resource languages can rely on the translation model's own internal representations to flag and correct detached outputs.

Engineering and product teams deploying machine translation should implement a detect-then-rewrite framework using Monte Carlo dropout for candidate generation. For systems with sufficient compute and multilingual support, cross-lingual sentence similarity models like LaBSE offer the strongest performance for both detection and reranking. Where auxiliary models are impractical, practitioners should implement internal source-attribution tracking to filter hallucinations without adding external dependencies.

Decision-makers should note that the analysis is limited to a single German-to-English news translation model and dataset. While confidence in detecting fully detached translations is high, existing methods remain less effective at isolating partial hallucinations where only a few words are detached. Further testing across additional language pairs, domains, and token-level detection strategies is recommended before enterprise-wide standardization.

Cover for Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better

Abstract

While the problem of hallucinations in neural machine translation has long been recognized, so far the progress on its alleviation is very little. Indeed, recently it turned out that without artificially encouraging models to hallucinate, previously existing methods fall short and even the standard sequence log-probability is more informative. It means that internal characteristics of the model can give much more information than we expect, and before using external models and measures, we first need to ask: how far can we go if we use nothing but the translation model itself ? We propose to use a method that evaluates the percentage of the source contribution to a generated translation. Intuitively, hallucinations are translations “detached” from the source, hence they can be identified by low source contribution. This method improves detection accuracy for the most severe hallucinations by a factor of 2 and is able to alleviate hallucinations at test time on par with the previous best approach that relies on external models. Next, if we move away from internal model characteristics and allow external tools, we show that using sentence similarity from cross-lingual embeddings further improves these results. We release the code of our experiments.

Table of Contents

  • 1 Introduction
  • 2 Background and Setting
  • 2.1 Model
  • 2.2 Hallucination Dataset
  • 3 Hallucination Detection Methods
  • 3.1 Reference-Based Oracles
  • 3.2 Internal Measures
  • 3.3 External models
  • 4 Detection Experiments
  • 4.1 Main results
  • 4.2 Analysing Distributions of the Scores
  • 4.3 Detected Pathology Types
  • 5 Mitigating Hallucinations at Test Time
  • 5.1 Evaluation methodology
  • 5.2 Generation Strategies
  • 5.2.1 The Impact of Generation Strategy
  • 5.2.2 The Impact of Number of Hypotheses
  • 5.3 Reranking Approaches
  • 5.3.1 Automatic Evaluation
  • 5.3.2 Human evaluation
  • 6 Conclusions
  • 7 Limitations
  • 8 Ethical statement
  • References
  • A Implementation and computing
  • B Mitigating Hallucinations at Test Time
  • C Manual Evaluation
  • ACL 2023 Responsible NLP Checklist

Knowls

  1. Knowl 1 — ALTI+ estimates how much a translation depends on its source

    model/method

    The paper uses ALTI+ to score the source contribution to a generated translation. For each target token, ALTI+ estimates the contributions of input tokens to that token’s representation; the contributions from all source tokens are summed, then averaged over target tokens. The resulting average source-contribution score is computed using the same translation model that generated the output. The method’s rationale is that hallucinations detached from the source should have lower source contribution, making this score usable for detection and reranking.

  2. Knowl 2 — Detection results: ALTI helps most on fully detached hallucinations, while LaBSE and XNLI are strongest overall

    data/table

    The table compares reference-based metrics (ChrF and COMET), internal metrics (Seq-Logprob and ALTI), and reference-free external metrics (COMET-QE, LASER, LaBSE, and XNLI) on German-to-English translations. ROC AUC and precision at 90% recall (P@R90) are reported on a percentage scale; higher is better. P@R90 measures the precision obtained while retaining 90% recall of the relevant hallucinations. The results show that ALTI’s main advantage over Seq-Logprob is for fully detached hallucinations: its P@R90 is 67.4 versus 31.0. LaBSE and XNLI have the strongest P@R90 for all hallucinations, and LaBSE has the highest P@R90 for fully detached hallucinations.

    MetricAll hallucinations: AUCAll hallucinations: P@R90Fully detached: AUCFully detached: P@R90
    ChrF75.414.489.616.6
    COMET83.419.287.712.6
    Seq-Logprob83.013.993.531.0
    ALTI84.912.598.767.4
    COMET-QE70.214.266.16.0
    LASER79.414.491.220.8
    LaBSE91.725.998.570.3
    XNLI90.924.198.760.4
  3. Knowl 3 — Cross-lingual similarity and entailment provide strong reference-free detection

    model/method

    The paper evaluates three external methods that compare a source sentence with its translation. LASER and LaBSE compute cosine similarity between cross-lingual sentence embeddings; LaBSE uses a model fine-tuned for translation ranking, whereas LASER uses a multilingual sentence-embedding model. XNLI multiplies the entailment probability from source to translation by the entailment probability in the reverse direction, using a cross-lingual NLI model. In the detection experiments, LaBSE and XNLI substantially outperform the COMET-QE quality-estimation baseline, while LASER performs noticeably worse than LaBSE. The authors note that LaBSE’s translation-ranking objective is aligned with ranking translations by error severity; they also report XNLI as a useful detection approach despite its not having been used for machine-translation hallucination detection in the work they reviewed.

  4. Knowl 4 — Reranking reduces human-judged hallucinations, with LaBSE producing the best result

    empirical result

    In a human evaluation, three annotators labeled deduplicated, shuffled outputs for 200 German source sentences, assigning each translation to correct, error, or hallucination categories. The reported hallucination and correct-translation rates were: original translations, 0.53 and 0.20; COMET-QE reranking, 0.22 and 0.54; ALTI reranking, 0.22 and 0.49; and LaBSE reranking, 0.16 and 0.56. Thus all three rerankers substantially reduced hallucinations, and LaBSE had the lowest hallucination rate and the highest correct rate. The difference in hallucination rate between ALTI and COMET-QE was not statistically significant (paired test, p = 1.00); LaBSE had a lower hallucination rate than COMET-QE and ALTI (p = 0.01 for each comparison), and a higher correct rate than ALTI (p = 0.02).

  5. Knowl 5 — Monte Carlo dropout beam search is the strongest tested candidate-generation strategy

    empirical result

    The paper compares alternative-hypothesis generation strategies for the detect-then-rewrite pipeline, evaluating each with multiple rerankers. The tested strategies include ordinary beam search, sampling from the full or top-p = 0.80 distribution, diverse beam search, diverse decoding, and Monte Carlo (MC) dropout. MC beam search generates candidates through n independent beam-search runs with dropout active, using beam size 10 in each run; the default was n = 10. This strategy clearly outperformed the other tested generation methods across rerankers. The authors emphasize that it introduces variation through model uncertainty while leaving the usual beam-search procedure otherwise intact.

  6. Knowl 6 — Automatic reranking improves COMET scores, with different strengths for ALTI and LaBSE

    data/table

    These average COMET scores compare the original translation (no reranking) with MC-dropout hypotheses reranked by different criteria. The evaluation sampled 150 examples from each of four groups—fully detached hallucinations, strongly detached hallucinations, other translation pathologies, and correct translations—and the table reports scores by pathology group and overall. Any reranking method improved the score over no reranking in each group. LaBSE achieved the highest overall score and improved on COMET-QE for fully detached hallucinations and correct translations, but scored lower than COMET-QE for other pathologies. ALTI exceeded COMET-QE for fully detached hallucinations but was weaker for the other pathology groups.

    RerankerFully detachedStrongly detachedOther pathologiesCorrectAverage
    No reranking-1.23-0.97-0.590.27-0.63
    COMET-QE-0.21-0.13-0.140.35-0.03
    ALTI-0.17-0.24-0.390.25-0.14
    LASER-0.11-0.23-0.350.27-0.11
    LaBSE-0.07-0.12-0.260.39-0.01
    XNLI-0.12-0.18-0.280.30-0.07
  7. Knowl 7 — The evaluation uses naturally occurring pathologies in a German-to-English test set

    experimental setup

    The detection and mitigation experiments use 3,415 manually annotated German-to-English translations selected from 1.8 million translations as likely pathological. The outputs were generated by a Fairseq Transformer-base model trained on WMT 2018 German–English news data excluding Paracrawl; the held-out set was used for analysis, and the model was trained on the remaining two-thirds of the data. The annotations contain 323 hallucinations, 1,044 less severe translation errors, and the remaining correct translations. Hallucinations are errors detached from the source: they may be oscillatory, or fluent but fully detached (the content is unsupported) or strongly detached (a significant portion is unsupported). Other annotated translation errors are classified as not detached from the source.

  8. Knowl 8 — Detector scores expose different patterns across pathology types

    empirical result

    The score-distribution analysis finds that ALTI and Seq-Logprob behave similarly: scores for less severe errors resemble those for correct translations, while strongly detached hallucinations have a bimodal distribution. This suggests that some partial hallucinations resemble fully detached cases to the detectors, while others resemble ordinary errors. LaBSE gives the clearest ordering across fully detached hallucinations, strongly detached hallucinations, and non-hallucinations; LASER’s distributions overlap too much for similarly reliable detection. COMET and COMET-QE do not clearly distinguish hallucinations from less severe errors, and XNLI scores concentrate near binary outcomes, making severity difficult to estimate. When selecting the lowest-scoring 20% of examples, all methods detect the fully detached hallucinations; the analysis also finds that XNLI, and to a lesser extent LaBSE, flags many undertranslations.

  9. Knowl 9 — More than ten candidates continue to improve reranking quality in the tested range

    empirical result

    The authors examine COMET scores as the number of MC-dropout hypotheses increases. Scores continue to rise beyond 10 candidates and do not saturate within the tested range. This indicates that generating more candidates can improve translation quality in settings where quality is prioritized over the additional computation, although the paper does not report a universal optimal candidate count.

  10. Knowl 10 — Generalization and partial-hallucination detection remain limitations

    limitation

    The experiments cover one translation direction (German to English), one dataset, and one Transformer-based translation model, so generalization to other languages, datasets, and models remains unverified. The evaluated methods detect fully detached hallucinations well, but do not reliably separate strongly detached, partial hallucinations from correct translations; the authors suggest that detection at the token level may be more suitable for this case. ALTI requires a Transformer-based translation model, while the strongest external detectors, LaBSE and XNLI, require additional encoders for the relevant languages, limiting their use for lower-resource languages or under computational constraints.

Coverage note — No substantial contributed material was omitted; implementation details and the full annotator instructions were left out because they do not add separate scientific findings.

References

  1. 1.Mikel Artetxe and Holger Schwenk. 2019. Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. Transactions of the Association for Computational Linguistics, 7:597–610.
  2. 2.Alexandre Berard, Ioan Calapodescu, and Claude Roux. 2019. Naver labs Europe’s systems for the WMT19 machine translation robustness task. In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 526–532, Florence, Italy. Association for Computational Linguistics.
  3. 3.Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Philipp Koehn, and Christof Monz. 2018. Findings of the 2018 conference on machine translation (WMT18). In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 272–303, Belgium, Brussels. Association for Computational Linguistics.
  4. 4.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  5. 5.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. Xnli: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  6. 6.Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022. Language-agnostic BERT sentence embedding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 878–891, Dublin, Ireland. Association for Computational Linguistics.
  7. 7.Patrick Fernandes, António Farinhas, Ricardo Rei, José De Souza, Perez Ogayo, Graham Neubig, and Andre Martins. 2022. Quality-aware decoding for neural machine translation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1396–1412, Seattle, United States. Association for Computational Linguistics.
  8. 8.Javier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano, and Marta R. Costa-jussà. 2022. Towards opening the black box of neural machine translation: Source and target interpretations of the transformer. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Online and Abu-Dhabi, UAE. Association for Computational Linguistics.
  9. 9.Marina Fomicheva, Lucia Specia, and Francisco Guzmán. 2020. Multi-hypothesis machine translation evaluation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1218–1232, Online. Association for Computational Linguistics.
  10. 10.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021. Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain. In Proceedings of the Sixth Conference on Machine Translation, pages 733–774, Online. Association for Computational Linguistics.
  11. 11.Yarin Gal and Zoubin Ghahramani. 2016. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 1050–1059, New York, New York, USA. PMLR.
  12. 12.Nuno M. Guerreiro, Elena Voita, and André F. T. Martins. 2022. Looking for a needle in a haystack: A comprehensive study of hallucinations in neural machine translation.
  13. 13.Kevin Heffernan, Onur Çelebi, and Holger Schwenk. 2022. Bitext mining using distilled sentence representations for low-resource languages.
  14. 14.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In International Conference on Learning Representations.
  15. 15.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022. Survey of hallucination in natural language generation.
  16. 16.Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021. To ship or not to ship: An extensive evaluation of automatic metrics for machine translation.
  17. 17.Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. 2019. Hallucinations in neural machine translation.
  18. 18.Jiwei Li, Will Monroe, and Dan Jurafsky. 2016. A simple, fast diverse decoding algorithm for neural generation. arXiv preprint arXiv:1611.08562.
  19. 19.Nitika Mathur, Johnny Wei, Markus Freitag, Qingsong Ma, and Ondřej Bojar. 2020. Results of the WMT20 metrics shared task. In Proceedings of the Fifth Conference on Machine Translation, pages 688–725, Online. Association for Computational Linguistics.
  20. 20.Mathias Müller, Annette Rios, and Rico Sennrich. 2020. Domain robustness in neural machine translation. In Proceedings of the 14th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track), pages 151–164, Virtual. Association for Machine Translation in the Americas.
  21. 21.Mathias Müller and Rico Sennrich. 2021. Understanding the properties of minimum Bayes risk decoding in neural machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 259–272, Online. Association for Computational Linguistics.
  22. 22.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53, Minneapolis, Minnesota. Association for Computational Linguistics.
  23. 23.Maja Popović. 2017. chrF++: words helping character n-grams. In Proceedings of the Second Conference on Machine Translation, pages 612–618, Copenhagen, Denmark. Association for Computational Linguistics.
  24. 24.Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt. 2021. The curious case of hallucinations in neural machine translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1172–1183, Online. Association for Computational Linguistics.
  25. 25.Vikas Raunak, Matt Post, and Arul Menezes. 2022. Salted: A framework for salient long-tail translation error detection.
  26. 26.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020a. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  27. 27.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020b. Unbabel’s participation in the WMT20 metrics shared task. In Proceedings of the Fifth Conference on Machine Translation, pages 911–920, Online. Association for Computational Linguistics.
  28. 28.Katsuhito Sudoh, Kosuke Takahashi, and Satoshi Nakamura. 2021. Is this translation error critical?: Classification-based human and automatic machine translation evaluation focusing on critical errors. In Proceedings of the Workshop on Human Evaluation of NLP Systems (HumEval), pages 46–55, Online. Association for Computational Linguistics.
  29. 29.Kosuke Takahashi, Yoichi Ishibashi, Katsuhito Sudoh, and Satoshi Nakamura. 2021. Multilingual machine translation evaluation metrics fine-tuned on pseudo-negative examples for wmt 2021 metrics task. In Proceedings of the Sixth Conference on Machine Translation, pages 1049–1052, Online. Association for Computational Linguistics.
  30. 30.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  31. 31.Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2016. Diverse beam search: Decoding diverse solutions from neural sequence models. arXiv preprint arXiv:1610.02424.
  32. 32.Elena Voita, Rico Sennrich, and Ivan Titov. 2021. Analyzing the source and target contributions to predictions in neural machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1126–1140, Online. Association for Computational Linguistics.
  33. 33.Chaojun Wang and Rico Sennrich. 2020. On exposure bias, hallucination and domain shift in neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3544–3552, Online. Association for Computational Linguistics.
  34. 34.Chrysoula Zerva, Daan van Stigt, Ricardo Rei, Ana C Farinha, Pedro Ramos, José G. C. de Souza, Taisiya Glushkova, Miguel Vera, Fabio Kepler, and André F. T. Martins. 2021. IST-unbabel 2021 submission for the quality estimation shared task. In Proceedings of the Sixth Conference on Machine Translation, pages 961–972, Online. Association for Computational Linguistics.
  35. 35.Chunting Zhou, Graham Neubig, Jiatao Gu, Mona Diab, Francisco Guzmán, Luke Zettlemoyer, and Marjan Ghazvininejad. 2021. Detecting hallucinated content in conditional neural sequence generation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1393–1404, Online. Association for Computational Linguistics.

Citation

MLA
Dale, D., et al. “Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 36–50, https://doi.org/10.18653/v1/2023.acl-long.3.
APA
Dale, D., Voita, E., Barrault, L., & Costa-jussà, M. R. (2023). Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 36–50. https://doi.org/10.18653/v1/2023.acl-long.3
Chicago
Dale, D., E. Voita, L. Barrault, and M. R. Costa-jussà. 2023. “Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 36–50. https://doi.org/10.18653/v1/2023.acl-long.3.
Harvard
Dale, D. et al. (2023) “Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 36–50. Available at: https://doi.org/10.18653/v1/2023.acl-long.3.
Vancouver
1. Dale D, Voita E, Barrault L, Costa-jussà MR (2023) Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 36–50

BibTeX

@inproceedings{dale-etal-2023-detecting,
    title = "Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity {E}ven Better",
    author = "Dale, David  and
      Voita, Elena  and
      Barrault, Loic  and
      Costa-juss{\`a}, Marta R.",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.3/",
    doi = "10.18653/v1/2023.acl-long.3",
    pages = "36--50"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/