Document-Level Machine Translation with Large Language Models

Longyue WangChenyang LyuTianbo JiZhirui ZhangDian YuShuming ShiZhaopeng Tu

article2023EMNLP219 citations

Demonstrates that large language models outperform commercial translation systems in document-level translation quality and establishes an instruction-based benchmark to evaluate discourse-level phenomena such as entity consistency and pronoun resolution.

Listen

Machine translation has traditionally evaluated performance on isolated sentences, often producing translations that lack contextual coherence, consistent terminology, and accurate pronoun references across full texts. The rapid rise of large language models presents an opportunity to address these limitations. The article evaluates the capabilities of large language models, specifically GPT-3.5 and GPT-4, in handling document-level machine translation and capturing broader linguistic discourse properties.

To conduct this assessment, the authors tested models across seven domains and three language pairs using both standard benchmarks and recent datasets designed to prevent data contamination. The evaluation compared the language models against established commercial translation tools and specialized document-level neural translation methods. The analysis combined automatic evaluation metrics with rigorous human assessments from professional linguists, alongside targeted tests designed to probe linguistic phenomena such as ellipsis, contextual references, and terminology consistency.

Key findings show that large language models offer strong capabilities for full-text translation. First, human evaluators rated GPT-3.5 and GPT-4 significantly higher than commercial translation products in overall translation quality and discourse awareness (averaging 2.8 to 3.1 out of 5, compared to 1.7 to 2.1 for commercial systems). Second, while commercial systems scored higher on automated n-gram metrics like document-level sacreBLEU in formal news and social media, the language models excelled in human-rated fluency and naturalness. Third, continuous document translation prompts that process multi-sentence context without rigid sentence boundaries yielded the best translation quality and terminology consistency. Fourth, targeted linguistic probing revealed that GPT-4 substantially outperforms GPT-3.5 in identifying and explaining complex discourse phenomena, although both models still occasionally struggle with fine-grained contextual distinctions compared to specialized repair modules. Finally, training techniques like code pre-training, supervised fine-tuning, and reinforcement learning from human feedback substantially improved discourse modeling performance.

These findings suggest that large language models represent a viable and promising paradigm for translating long-form content, particularly where narrative flow, tone, and conversational context are critical. However, automated metrics alone do not fully capture translation quality, meaning organizations relying purely on standard automatic benchmarks may misjudge model performance. Decision-makers should consider large language models for tasks requiring high contextual naturalness while noting trade-offs in computational stability and exact lexical matching.

Organizations evaluating translation solutions should explore hybrid deployment strategies, using continuous context prompting and establishing evaluation frameworks that integrate human review alongside automated metrics. Further research and development should focus on testing emerging evaluation techniques on newer long-form benchmarks, refining training transparency, and addressing known limitations. These limitations include periodic translation instability, potential data contamination from public test sets, and the evolving nature of closed commercial model APIs.

  • Paper: Prompting Large Language Model for Machine Translation: A Case Study, Biao Zhang et al. (2023). This paper establishes foundational empirical findings on prompt engineering strategies and demonstration selection for LLM-based machine translation, which the source extends to document-level discourse contexts.
  • Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). This foundational work introduces the few-shot in-context learning paradigm of large language models that the source adapts for context-aware document translation.
  • Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). This study analyzes how prompt framing and zero-shot directives steer language model behavior on translation tasks, providing key conceptual foundations for the prompt design evaluated in the source.
  • Paper: Language Models are Unsupervised Multitask Learners, Alec Radford et al. (2019). This work demonstrates that autoregressive language models inherently learn cross-sentence translation in an unsupervised, zero-shot setting, forming the conceptual baseline for LLM-based translation evaluation.
Cover for Document-Level Machine Translation with Large Language Models

Abstract

Large language models (LLMs) such as ChatGPT can produce coherent, cohesive, relevant, and fluent answers for various natural language processing (NLP) tasks. Taking document-level machine translation (MT) as a testbed, this paper provides an in-depth evaluation of LLMs' ability on discourse modeling. The study focuses on three aspects: 1) Effects of Context-Aware Prompts, where we investigate the impact of different prompts on document-level translation quality and discourse phenomena; 2) Comparison of Translation Models, where we compare the translation performance of ChatGPT with commercial MT systems and advanced document-level MT methods; 3) Analysis of Discourse Modelling Abilities, where we further probe discourse knowledge encoded in LLMs and shed light on impacts of training techniques on discourse modeling. By evaluating on a number of benchmarks, we surprisingly find that LLMs have demonstrated superior performance and show potential to become a new paradigm for document-level translation: 1) leveraging their powerful long-text modeling capabilities, GPT-3.5 and GPT-4 outperform commercial MT systems in terms of human evaluation;2) GPT-4 demonstrates a stronger ability for probing linguistic knowledge than GPT-3.5. This work highlights the challenges and opportunities of LLMs for MT, which we hope can inspire the future design and evaluation of LLMs.2

Table of Contents

  • 1 Introduction
  • 2 Experimental Step
  • 2.1 Dataset
  • 2.2 Evaluation Method
  • 3 Effects of Context-Aware Prompts
  • 3.1 Motivation
  • 3.2 Comparison of Different Prompts
  • 4 Comparison of Translation Models
  • 4.1 ChatGPT vs. Commercial Systems
  • 4.2 ChatGPT vs. Document NMT Methods
  • 5 Analysis of Large Language Models
  • 5.1 Probing Discourse Knowledge in LLM
  • 5.2 Potential Impacts of Training Techniques
  • 6 Conclusion and Future Work
  • Limitations
  • Ethical Considerations
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Significance Testing
  • A.2 Human Evaluation Guidelines
  • A.3 Training Methods in LLMs

Knowls

  1. Knowl 1 — Comparison of Large Language Models and Commercial MT Systems in Document Translation

    empirical result

    When evaluated across diverse domains (News, Social media, Web fiction, Q&A forum) on Chinese-to-English translation tasks, commercial translation systems (Google Translate, DeepL Translate, Tencent TranSmart) and large language models (GPT-3.5 and GPT-4) exhibit a sharp divergence between automatic nn-gram matching metrics and human qualitative assessments.

    While commercial systems achieve comparable or slightly higher document-level sacreBLEU (d-BLEU) scores—particularly in structured domains—human evaluations (rated 0–5 on General Quality and Discourse Awareness) demonstrate that GPT-3.5 and GPT-4 significantly outperform commercial systems across all domains.

    Model Automatic (d-BLEU) Human (General / Discourse)
    News Social Fiction Q A Ave. News Social Fiction Q A Ave.
    Google 27.7 35.4 16.0 12.0 22.8 1.9 / 2.0 1.2 / 1.3 2.1 / 2.4 1.5 / 1.5 1.7 / 1.8
    DeepL 30.3 33.4 16.1 11.9 22.9 2.2 / 2.2 1.3 / 1.1 2.4 / 2.6 1.6 / 1.5 1.9 / 1.9
    Tencent 29.3 38.8 20.7 15.0 26.0 2.3 / 2.2 1.5 / 1.5 2.6 / 2.8 1.8 / 1.7 2.1 / 2.1
    GPT-3.5 29.1 35.5 17.4 17.4 24.9 2.8 / 2.8 2.5 / 2.7 2.8 / 2.9 2.9 / 2.9 2.8 / 2.8
    GPT-4 29.7 34.4 18.8 19.0 25.5 3.3 / 3.4 2.9 / 2.9 2.6 / 2.8 3.1 / 3.2 3.0 / 3.1

    The discrepancy indicates that d-BLEU predominantly measures exact lexical nn-gram overlap against single/few references, whereas human evaluation captures long-range inter-sentential coherence, stylistic naturalness, and discourse consistency where LLMs excel.

  2. Knowl 2 — Discourse Knowledge Probing and Explanation in LLMs via Contrastive Testing

    empirical result

    Probing LLMs using contrastive test sets for English-to-Russian translation (evaluating deixis, lexical consistency, inflectional ellipsis, and verb phrase ellipsis) reveals distinct prediction accuracy and explanation capabilities between model generations.

    In translation prediction on contrastive sets, GPT-3.5 underperforms dedicated context-aware neural post-editing baselines (such as DocRepair), especially in deixis (57.9%57.9\%) and lexical consistency (44.4%44.4\%). GPT-4 substantially improves prediction accuracy (85.9%85.9\% on deixis, 72.4%72.4\% on lexical consistency, and 81.4%81.4\% on verb phrase ellipsis).

    Model Deixis Lexical Consistency Ellipsis (Inflection) Ellipsis (VP)
    Sent2Sent 51.1 45.6 55.4 27.4
    MR-Doc2Doc 64.7 46.3 65.9 53.0
    CADec 81.6 58.1 72.2 80.0
    DocRepair 91.8 80.6 86.4 75.2
    GPT-3.5 57.9 44.4 75.0 71.6
    GPT-4 85.9 72.4 69.8 81.4

    When evaluating explanations provided by LLMs for why a translation choice is correct, human evaluations on 100 instances per subset reveal that GPT-4 generates accurate explanations significantly more often than GPT-3.5 (93.0%93.0\% vs 18.0%18.0\% on deixis; 86.0%86.0\% vs 11.0%11.0\% on lexical consistency; 91.0%91.0\% vs 58.0%58.0\% on inflection ellipsis; 94.0%94.0\% vs 75.0%75.0\% on VP ellipsis). However, the correlation between correct prediction and correct explanation measured by the Phi coefficient (rϕr_\phi) remains weak to moderate (ranging between 0.1840.184 and 0.5390.539 for GPT-4), demonstrating a decoupling between generating correct discourse choices and articulating the underlying linguistic rationale.

  3. Knowl 3 — Impact of Training Techniques on LLM Document Translation and Discourse Modeling

    empirical result

    Evaluating sequential model variants in the GPT lineage illustrates how specific training stages impact document-level translation quality (Chinese-to-English Web Fiction) and discourse probing accuracy (English-to-Russian contrastive prediction and explanation):

    Model Training Method Fiction d-BLEU Fiction Human (Gen/Disc) Probing (Pred/Expl %)
    GPT-3 Pre-training (175B) 3.3 – –
    InstructGPT (+SFT) Supervised Fine-Tuning 7.1 – –
    InstructGPT (+FeedME-1) Demonstration Fine-Tuning 14.1 2.2 / 2.5 30.5 / 28.6
    CodexGPT (+FeedME-2) Code Pre-training + SFT 16.1 2.2 / 2.3 34.4 / 30.1
    CodexGPT (+PPO) SFT + RLHF (PPO) 17.2 2.6 / 2.7 58.0 / 39.4
    GPT-3.5 Conversational SFT / RLHF 17.4 2.8 / 2.9 62.3 / 40.5
    GPT-4 Advanced Multimodal / RLHF 18.8 2.6 / 2.8 78.5 / 91.0

    Key trends include:

    • SFT on quality demonstration examples (+FeedME-1) elevates translation capability to usable levels (14.114.1 d-BLEU vs 7.17.1 for standard SFT).
    • Pre-training on source code (+FeedME-2) improves document translation quality by +2.0+2.0 d-BLEU points and improves discourse probing.
    • Reinforcement Learning from Human Feedback via Proximal Policy Optimization (+PPO) produces the largest relative leap in discourse awareness, increasing probing prediction accuracy from 34.4%34.4\% to 58.0%58.0\%.
    • Subsequent conversational scaling in GPT-3.5 and GPT-4 delivers further gains in discourse probing (78.5%78.5\% prediction and 91.0%91.0\% explanation accuracy in GPT-4).
  4. Knowl 4 — Context-Aware Prompting Strategies for Document Machine Translation

    model/method

    Document-level machine translation using LLMs can be elicited using three distinct prompting strategies:

    1. P1 (Sentence-by-sentence in multi-turn chat): Translating one sentence SS at a time per conversational turn while keeping previous sentences within the chat context (Please provide the TGT translation for the sentence: S).
    2. P2 (Multi-sentence continuous with boundary tags): Concatenating all sentences into a single turn with explicit sentence demarcations (Translate the following SRC sentences into TGT: [S1], [S2] ...).
    3. P3 (Multi-sentence continuous whole document): Supplying continuous document text without internal sentence boundary brackets ((Continue) Translate this document from SRC to TGT: S1 S2 ...).

    Evaluating these prompts with ChatGPT on Chinese-to-English benchmarks yields:

    Prompt News BLEU Fiction BLEU News d-BLEU Fiction d-BLEU Fiction CTT Fiction AZPT
    Base (InstructGPT, no context) 25.5 12.4 28.2 15.4 0.19 0.39
    P1 (Multi-turn per sentence) 25.8 13.9 28.7 17.0 0.29 0.41
    P2 (Document with `[]` tags) 26.2 13.8 28.8 16.5 0.28 0.41
    P3 (Continuous document text) 26.5 14.4 29.1 17.4 0.33 0.44

    Whole-document input without artificial sentence brackets (P3) achieves superior performance in both general translation metrics and specific discourse metrics (consistency of terminology translation, CTT, and accuracy of zero pronoun translation, AZPT), matching human translation behavior where sentence boundaries are flexibly adapted across languages.

  5. Knowl 5 — Comparison of ChatGPT with Dedicated Document-Level NMT Architectures

    empirical result

    Comparing ChatGPT against specialized document-level neural machine translation (NMT) architectures across standard benchmarks demonstrates competitive or state-of-the-art results on spoken and news text, with domain sensitivity on legislative text.

    Model Zh⇒\RightarrowEn TED En⇒\RightarrowDe TED En⇒\RightarrowDe News En⇒\RightarrowDe Europarl
    BLEU d-BLEU BLEU d-BLEU BLEU d-BLEU BLEU d-BLEU
    MCN 19.1 25.7 25.1 29.1 24.9 27.0 30.4 32.6
    G-Trans – – 25.1 27.2 25.5 27.1 32.4 34.1
    Sent2Sent 19.2 25.8 25.2 29.2 25.0 27.0 31.7 33.8
    MR-Doc2Sent 19.4 25.8 25.2 29.2 25.0 26.7 32.1 34.2
    MR-Doc2Doc – 25.9 – 29.3 – 26.7 – 34.5
    Sent2Sent* 21.9 27.9 27.1 30.7 27.9 29.4 32.1 34.2
    MR-Doc2Sent* 22.0 28.1 27.3 31.0 29.5 31.2 32.4 34.5
    MR-Doc2Doc* – 28.4 – 31.4 – 32.6 – 34.9
    ChatGPT – 28.3 – 33.6 – 39.4 – 30.4

    (* indicates models trained with additional sentence-level pre-training data.)

    ChatGPT outperforms the best specialized models (MR-Doc2Doc*) on En⇒\RightarrowDe TED (33.633.6 vs 31.431.4 d-BLEU) and En⇒\RightarrowDe News Commentary (39.439.4 vs 32.632.6 d-BLEU), while matching MR-Doc2Doc* on Zh⇒\RightarrowEn TED (28.328.3 vs 28.428.4 d-BLEU). Conversely, on En⇒\RightarrowDe Europarl v7, ChatGPT achieves only 30.430.4 d-BLEU, falling below the sentence-level baseline Sent2Sent (33.833.8 d-BLEU), attributed to domain distribution shift and occasional generation instabilities (e.g., copying or omissions).

  6. Knowl 6 — Lexical Translation Consistency Metric (CTT)

    equation

    Lexical translation consistency (CTT) measures how consistently repeated terminology words in a source document are translated across the entire target document. Let TT\text{TT} denote the set of source terminology words occurring repeatedly in the document. For each term w∈TTw \in \text{TT} with kk target translations (t1,t2,…,tk)(t_1, t_2, \dots, t_k), CTT is defined as:

    CTT=1∣TT∣∑w∈TT∑i=1k∑j=i+1k1(ti=tj)Ck2\text{CTT} = \frac{1}{|\text{TT}|} \sum_{w \in \text{TT}} \frac{\sum_{i=1}^k \sum_{j=i+1}^k \mathbf{1}(t_i = t_j)}{C_k^2}

    where:

    • Ck2=(k2)=k(k−1)2C_k^2 = \binom{k}{2} = \frac{k(k-1)}{2} is the number of distinct translation pairs for term ww.
    • 1(ti=tj)\mathbf{1}(t_i = t_j) is an indicator function returning 11 if translation tokens tit_i and tjt_j are identical, and 00 otherwise.
    • ∣TT∣|\text{TT}| is the total number of evaluated terminology terms in the document.

    A higher CTT score indicates superior discourse-level lexical cohesion and terminology consistency across sentences.

  7. Knowl 7 — Accuracy of Zero Pronoun Translation Metric (AZPT)

    equation

    The Accuracy of Zero Pronoun Translation (AZPT) quantifies a model's ability to resolve and translate dropped pronouns when translating from a pro-drop language (such as Chinese) into a non-pro-drop target language (such as English). AZPT is defined as:

    AZPT=∑z∈ZPA(tz∣z)∣ZP∣\text{AZPT} = \frac{\sum_{z \in \text{ZP}} A(t_z \mid z)}{|\text{ZP}|}

    where:

    • ZP\text{ZP} is the set of annotated zero pronouns present in the source document.
    • tzt_z is the generated translation token corresponding to zero pronoun zz.
    • A(tz∣z)∈{0,1}A(t_z \mid z) \in \{0, 1\} is a binary evaluation function scoring 11 if tzt_z correctly recovers and translates the referent of zero pronoun zz according to context, and 00 otherwise.
    • ∣ZP∣|\text{ZP}| is the cardinality of zero pronouns in the document.
  8. Knowl 8 — Two-Dimensional Human Evaluation Framework for Document Translation

    experimental setup

    Human evaluation of document-level MT outputs is conducted on a sliding multi-sentence window contextualized by the full document, assigning two independent integer scores from 0 to 5:

    1. General Quality (0–5): Evaluates fluency, grammatical correctness, word choice, and adequacy against the source text (Score 5: fully fluent, accurate, zero mistranslation/omission; Score 0: completely unrelated to source).
    2. Discourse Awareness (0–5):
      • Score 5: Total consistency of named entities, technical terms, and organizations; natural inter-sentential discourse connectives; stable tone, topic, and register conforming to target cultural conventions.
      • Score 4: Fluent logical flow; minor absence of sentence transitions that does not impair contextual comprehension; consistent key terms.
      • Score 3: Mostly readable but lacks smooth linkages; minor abrupt transitions; occasional inconsistency in tone or non-critical terminology.
      • Score 2: Inconsistent key terms requiring frequent re-reading; disjointed passage flow; tonal or topical drift across sentences.
      • Score 1: Pervasive terminology contradictions; severe lack of inter-sentence cohesion heavily obstructing reading comprehension.
      • Score 0: Output disconnected from preceding or succeeding discourse context.

Coverage note — None was omitted; all primary empirical comparisons, prompting formulations, mathematical metrics (CTT, AZPT), probing experiments, and training progression analyses are represented.

References

  1. 1.Guangsheng Bao, Yue Zhang, Zhiyang Teng, Boxing Chen, and Weihua Luo. 2021. G-transformer for document-level machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers).
  2. 2.Rachel Bawden, Rico Sennrich, Alexandra Birch, and Barry Haddow. 2018. Evaluating discourse phenomena in neural machine translation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers).
  3. 3.Welch Bl. 1947. The generalization of ‘student’s’ problem when several different population varlances are involved. Biometrika.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems.
  5. 5.Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. 2020. With little power comes great responsibility. arXiv preprint arXiv:2010.06595.
  6. 6.Sheila Castilho. 2021. Towards document-level human mt evaluation: On the issues of annotator agreement, effort and misevaluation. Association for Computational Linguistics.
  7. 7.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  9. 9.Marjan Ghazvininejad, Hila Gonen, and Luke Zettlemoyer. 2023. Dictionary-based phrase-level prompting of large language models for machine translation. arXiv preprint arXiv:2302.07856.
  10. 10.Yvette Graham, Barry Haddow, and Philipp Koehn. 2020. Statistical power and translationese in machine translation evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing.
  11. 11.Nuno M. Guerreiro, Elena Voita, and André Martins. 2023. Looking for a needle in a haystack: A comprehensive study of hallucinations in neural machine translation. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics.
  12. 12.Junliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei, Boxing Chen, and Enhong Chen. 2020. Incorporating bert into parallel sequence decoding with adapters. Advances in Neural Information Processing Systems.
  13. 13.Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How good are gpt models at machine translation? a comprehensive evaluation. arXiv preprint arXiv:2302.09210.
  14. 14.Guoping Huang, Lemao Liu, Xing Wang, Longyue Wang, Huayang Li, Zhaopeng Tu, Chengyan Huang, and Shuming Shi. 2021. Transmart: A practical interactive machine translation system. arXiv preprint arXiv:2105.13072.
  15. 15.Yuchen Eleanor Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Mrinmaya Sachan, and Ryan Cotterell. 2023. Discourse centric evaluation of machine translation with a densely annotated parallel corpus. arXiv preprint arXiv:2305.11142.
  16. 16.Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu. 2023. Is chatgpt a good translator? a preliminary study. arXiv preprint arXiv:2301.08745.
  17. 17.Marzena Karpinska and Mohit Iyyer. 2023. Large language models effectively leverage document-level context for literary translation, but critical errors persist. arXiv preprint arXiv:2304.03245.
  18. 18.Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Novák, Martin Popel, and Maja Popović. 2022. Findings of the 2022 conference on machine translation (WMT22). In Proceedings of the Seventh Conference on Machine Translation.
  19. 19.Tom Kocmi and Christian Federmann. 2023. Large language models are state-of-the-art evaluators of translation quality. arXiv preprint arXiv:2302.14520.
  20. 20.Klaus Krippendorff. 2011. Computing krippendorff’s alpha-reliability.
  21. 21.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics.
  22. 22.Hongyuan Lu, Haoyang Huang, Dongdong Zhang, Haoran Yang, Wai Lam, and Furu Wei. 2023. Chain-of-dictionary prompting elicits translation in large language models. arXiv preprint arXiv:2305.06575.
  23. 23.Chenyang Lyu, Jitao Xu, and Longyue Wang. 2023. New trends in machine translation using large language models: Case examples with chatgpt. arXiv preprint arXiv:2305.01181.
  24. 24.Xinglin Lyu, Junhui Li, Zhengxian Gong, and Min Zhang. 2021. Encouraging lexical translation consistency for document-level neural machine translation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.
  25. 25.Valentin Macé and Christophe Servan. 2019. Using whole document context in neural machine translation. In Proceedings of the 16th International Conference on Spoken Language Translation, Hong Kong. Association for Computational Linguistics.
  26. 26.Mary L McHugh. 2012. Interrater reliability: the kappa statistic. Biochemia medica.
  27. 27.Graham Neubig and Zhiwei He. 2023. Zeno gpt machine translation report.
  28. 28.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, et al. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems.
  29. 29.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers.
  30. 30.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI.
  31. 31.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research.
  32. 32.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  33. 33.Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul. 2006. A study of translation edit rate with targeted human annotation. In Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers.
  34. 34.Zewei Sun, Mingxuan Wang, Hao Zhou, Chengqi Zhao, Shujian Huang, Jiajun Chen, and Lei Li. 2022. Rethinking document-level neural machine translation. In Findings of the Association for Computational Linguistics: ACL 2022.
  35. 35.Katherine Thai, Marzena Karpinska, Kalpesh Krishna, Bill Ray, Moira Inghilleri, John Wieting, and Mohit Iyyer. 2022. Exploring document-level literary machine translation with parallel paragraphs from world literature. arXiv preprint arXiv:2210.14250.
  36. 36.Zhaopeng Tu, Yang Liu, Shuming Shi, and Tong Zhang. 2018. Learning to remember translation history with a continuous cache. Transactions of the Association for Computational Linguistics.
  37. 37.David Vilar, Markus Freitag, Colin Cherry, Jiaming Luo, Viresh Ratnakar, and George Foster. 2022. Prompting palm for translation: Assessing strategies and performance. arXiv preprint arXiv:2211.09102.
  38. 38.Elena Voita, Rico Sennrich, and Ivan Titov. 2019a. Context-aware monolingual repair for neural machine translation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing.
  39. 39.Elena Voita, Rico Sennrich, and Ivan Titov. 2019b. When a good translation is wrong in context: Context-aware machine translation improves on deixis, ellipsis, and lexical cohesion. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
  40. 40.Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov. 2018. Context-aware neural machine translation learns anaphora resolution. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).
  41. 41.Longyue Wang. 2019. Discourse-aware neural machine translation. Ph.D. thesis, Ph. D. thesis, Dublin City University, Dublin, Ireland.
  42. 42.Longyue Wang, Zefeng Du, Donghuai Liu, Cai Deng, Dian Yu, Haiyun Jiang, Yan Wang, Leyang Cui, Shuming Shi, and Zhaopeng Tu. 2023a. Disco-bench: A discourse-aware evaluation benchmark for language modelling. arXiv preprint arXiv:2307.08074.
  43. 43.Longyue Wang, Siyou Liu, Mingzhou Xu, Linfeng Song, Shuming Shi, and Zhaopeng Tu. 2023b. A survey on zero pronoun translation. arXiv preprint arXiv:2305.10196.
  44. 44.Longyue Wang, Zhaopeng Tu, Yan Gu, Siyou Liu, Dian Yu, Qingsong Ma, Chenyang Lyu, Liting Zhou, Chao-Hong Liu, Yufeng Ma, Weiyu Chen, Yvette Graham, Bonnie Webber, Philipp Koehn, Andy Way, Yulin Yuan, and Shuming Shi. 2023c. Findings of the WMT 2023 shared task on discourse-level literary translation. In Proceedings of the Eighth Conference on Machine Translation (WMT).
  45. 45.Longyue Wang, Zhaopeng Tu, Shuming Shi, Tong Zhang, Yvette Graham, and Qun Liu. 2018a. Translating pro-drop languages with reconstruction models. In Proceedings of the 2018 AAAI Conference on Artificial Intelligence, volume 32.
  46. 46.Longyue Wang, Zhaopeng Tu, Xing Wang, and Shuming Shi. 2019. One model to learn both: Zero pronoun prediction and translation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing.
  47. 47.Longyue Wang, Zhaopeng Tu, Andy Way, and Qun Liu. 2017. Exploiting cross-sentence context for neural machine translation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing.
  48. 48.Longyue Wang, Zhaopeng Tu, Andy Way, and Qun Liu. 2018b. Learning to jointly translate and predict dropped pronouns with a shared reconstruction mechanism. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing.
  49. 49.Longyue Wang, Zhaopeng Tu, Xiaojun Zhang, Hang Li, Andy Way, and Qun Liu. 2016. A novel approach to dropped pronoun translation. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  50. 50.Longyue Wang, Mingzhou Xu, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi, and Zhaopeng Tu. 2022. GuoFeng: A benchmark for zero pronoun recovery and translation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.
  51. 51.RF Woolson. 2007. Wilcoxon signed-rank test. Wiley encyclopedia of clinical trials.
  52. 52.Tong Xiao, Jingbo Zhu, Shujie Yao, and Hao Zhang. 2011. Document-level consistency verification in machine translation. In Proceedings of Machine Translation Summit XIII: Papers.
  53. 53.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  54. 54.Biao Zhang, Ankur Bapna, Melvin Johnson, Ali Dabirmoghaddam, Naveen Arivazhagan, and Orhan Firat. 2022. Multilingual document-level translation enables zero-shot transfer from sentences to documents. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).
  55. 55.Zaixiang Zheng, Xiang Yue, Shujian Huang, Jiajun Chen, and Alexandra Birch. 2020. Toward making the most of context in neural machine translation. ArXiv.
  56. 56.Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu. 2020. Incorporating bert into neural machine translation. arXiv preprint arXiv:2002.06823.

Citation

MLA
Wang, L., et al. “Document-Level Machine Translation with Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 16646–61, https://doi.org/10.18653/v1/2023.emnlp-main.1036.
APA
Wang, L., Lyu, C., Ji, T., Zhang, Z., Yu, D., Shi, S., & Tu, Z. (2023). Document-Level Machine Translation with Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 16646–16661. https://doi.org/10.18653/v1/2023.emnlp-main.1036
Chicago
Wang, L., C. Lyu, T. Ji, et al. 2023. “Document-Level Machine Translation with Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 16646–61. https://doi.org/10.18653/v1/2023.emnlp-main.1036.
Harvard
Wang, L. et al. (2023) “Document-Level Machine Translation with Large Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 16646–16661. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.1036.
Vancouver
1. Wang L, Lyu C, Ji T, Zhang Z, Yu D, Shi S, Tu Z (2023) Document-Level Machine Translation with Large Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 16646–16661

BibTeX

@inproceedings{wang-etal-2023-document-level,
    title = "Document-Level Machine Translation with Large Language Models",
    author = "Wang, Longyue  and
      Lyu, Chenyang  and
      Ji, Tianbo  and
      Zhang, Zhirui  and
      Yu, Dian  and
      Shi, Shuming  and
      Tu, Zhaopeng",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.1036/",
    doi = "10.18653/v1/2023.emnlp-main.1036",
    pages = "16646--16661"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/