Document-Level Machine Translation with Large Language Models
Longyue WangChenyang LyuTianbo JiZhirui ZhangDian YuShuming ShiZhaopeng Tu
Demonstrates that large language models outperform commercial translation systems in document-level translation quality and establishes an instruction-based benchmark to evaluate discourse-level phenomena such as entity consistency and pronoun resolution.
Machine translation has traditionally evaluated performance on isolated sentences, often producing translations that lack contextual coherence, consistent terminology, and accurate pronoun references across full texts. The rapid rise of large language models presents an opportunity to address these limitations. The article evaluates the capabilities of large language models, specifically GPT-3.5 and GPT-4, in handling document-level machine translation and capturing broader linguistic discourse properties.
To conduct this assessment, the authors tested models across seven domains and three language pairs using both standard benchmarks and recent datasets designed to prevent data contamination. The evaluation compared the language models against established commercial translation tools and specialized document-level neural translation methods. The analysis combined automatic evaluation metrics with rigorous human assessments from professional linguists, alongside targeted tests designed to probe linguistic phenomena such as ellipsis, contextual references, and terminology consistency.
Key findings show that large language models offer strong capabilities for full-text translation. First, human evaluators rated GPT-3.5 and GPT-4 significantly higher than commercial translation products in overall translation quality and discourse awareness (averaging 2.8 to 3.1 out of 5, compared to 1.7 to 2.1 for commercial systems). Second, while commercial systems scored higher on automated n-gram metrics like document-level sacreBLEU in formal news and social media, the language models excelled in human-rated fluency and naturalness. Third, continuous document translation prompts that process multi-sentence context without rigid sentence boundaries yielded the best translation quality and terminology consistency. Fourth, targeted linguistic probing revealed that GPT-4 substantially outperforms GPT-3.5 in identifying and explaining complex discourse phenomena, although both models still occasionally struggle with fine-grained contextual distinctions compared to specialized repair modules. Finally, training techniques like code pre-training, supervised fine-tuning, and reinforcement learning from human feedback substantially improved discourse modeling performance.
These findings suggest that large language models represent a viable and promising paradigm for translating long-form content, particularly where narrative flow, tone, and conversational context are critical. However, automated metrics alone do not fully capture translation quality, meaning organizations relying purely on standard automatic benchmarks may misjudge model performance. Decision-makers should consider large language models for tasks requiring high contextual naturalness while noting trade-offs in computational stability and exact lexical matching.
Organizations evaluating translation solutions should explore hybrid deployment strategies, using continuous context prompting and establishing evaluation frameworks that integrate human review alongside automated metrics. Further research and development should focus on testing emerging evaluation techniques on newer long-form benchmarks, refining training transparency, and addressing known limitations. These limitations include periodic translation instability, potential data contamination from public test sets, and the evolving nature of closed commercial model APIs.
- Paper: Prompting Large Language Model for Machine Translation: A Case Study, Biao Zhang et al. (2023). This paper establishes foundational empirical findings on prompt engineering strategies and demonstration selection for LLM-based machine translation, which the source extends to document-level discourse contexts.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). This foundational work introduces the few-shot in-context learning paradigm of large language models that the source adapts for context-aware document translation.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). This study analyzes how prompt framing and zero-shot directives steer language model behavior on translation tasks, providing key conceptual foundations for the prompt design evaluated in the source.
- Paper: Language Models are Unsupervised Multitask Learners, Alec Radford et al. (2019). This work demonstrates that autoregressive language models inherently learn cross-sentence translation in an unsupervised, zero-shot setting, forming the conceptual baseline for LLM-based translation evaluation.
- Paper: Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate, Tian Liang et al. (2024). This paper extends LLM translation capabilities by using multi-agent debate frameworks to resolve subtle commonsense and contextual translation ambiguities identified in baseline LLM translation evaluations.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). This benchmark broadens the evaluation of LLM translation and generation capabilities across dozens of diverse typological languages and multimodal formats.
- Paper: Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models, Mosh Levy et al. (2024). This paper investigates how increasing input context lengths impacts LLM reasoning and consistency, providing crucial analytical insight into the long-context discourse limitations observed in document-level translation.
