Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better
David DaleElena VoitaLoïc BarraultMarta R. Costa-jussà
Demonstrates that measuring source token contribution from a translation model's internal mechanics doubles the detection accuracy of severe hallucinations without external tools, while cross-lingual sentence similarity models achieve an 80% precision gain across all hallucination types.
Machine translation systems occasionally generate hallucinations, which are outputs completely or partially detached from the original source text. While rare, these severe errors undermine user trust and present substantial operational risks in real-world deployment. Detecting and repairing naturally occurring hallucinations in production has proven difficult because standard automated quality estimation metrics frequently fail to flag severe pathologies, leading previous approaches to rely on artificial perturbations or external evaluation tools.
The article evaluates both internal model mechanisms and external tools to determine how effectively they can detect naturally occurring hallucinations and mitigate them at test time without relying on human reference translations. The authors analyze a clean test bed of German-to-English translations generated by a standard neural translation model, evaluating performance across a curated dataset of over 3,400 annotated examples containing naturally occurring translation pathologies.
Key findings show that internal model metrics alone perform remarkably well. Using a method called ALTI to calculate the percentage of source token contribution within the translation model doubles the detection precision for fully detached hallucinations compared to baseline sequence log-probability (reaching 67.4% precision at 90% recall compared to 31.0%). When external tools are permitted, cross-lingual sentence similarity models, specifically LaBSE, deliver the highest overall detection accuracy across all hallucination types, achieving an 80% improvement in precision over the baseline. Furthermore, evaluating hypotheses within a detect-then-rewrite pipeline reveals that generating candidates via Monte Carlo dropout and reranking them with sentence similarity cuts the human-judged hallucination rate from 53% down to 16%, while internal ALTI reranking matches the mitigation performance of complex external quality estimators, cutting the rate to 22%.
These findings demonstrate that organizations do not necessarily need computationally heavy external quality estimation pipelines to mitigate translation hallucinations. Because an internal measure of source contribution achieves comparable mitigation to state-of-the-art external estimators, organizations operating under compute constraints or translating low-resource languages can rely on the translation model's own internal representations to flag and correct detached outputs.
Engineering and product teams deploying machine translation should implement a detect-then-rewrite framework using Monte Carlo dropout for candidate generation. For systems with sufficient compute and multilingual support, cross-lingual sentence similarity models like LaBSE offer the strongest performance for both detection and reranking. Where auxiliary models are impractical, practitioners should implement internal source-attribution tracking to filter hallucinations without adding external dependencies.
Decision-makers should note that the analysis is limited to a single German-to-English news translation model and dataset. While confidence in detecting fully detached translations is high, existing methods remain less effective at isolating partial hallucinations where only a few words are detached. Further testing across additional language pairs, domains, and token-level detection strategies is recommended before enterprise-wide standardization.
- Paper: Survey of Hallucination in Natural Language Generation, Ziwei Ji et al. (2022). Its survey establishes the key hallucination definitions, categories, and detection approaches—including translation-specific work—that frame this paper’s focused evaluation.
- Paper: Six Challenges for Neural Machine Translation, Philipp Koehn et al. (2017). Its analysis of neural machine translation’s brittleness and source-detached outputs motivates the production hallucination problem this paper detects and repairs.
- Paper: On the Limitations of Reference-Free Evaluations of Generated Text, Daniel Deutsch et al. (2022). Its critique of reference-free generation metrics clarifies why ordinary quality estimation can miss severe translation errors, a central concern of this paper.
- Paper: SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models, Potsawee Manakul et al. (2023). Its comparison of internal probability signals and sampled-output consistency provides useful groundwork for this paper’s evaluation of internal and external hallucination detectors.
- Paper: RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models, Cheng Niu et al. (2024). It extends natural-error hallucination detection to retrieval-augmented generation with a large annotated corpus and word-level detection and filtering.
- Paper: How Language Model Hallucinations Can Snowball, Muru Zhang et al. (2024). It carries the study of naturally occurring hallucinations forward by examining how an initial false answer can trigger further unsupported claims.
- Paper: Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation, Xiaoying Zhang et al. (2024). It extends detect-and-rank mitigation into self-evaluation and preference tuning, reducing hallucinations without external human labels.
