When Does Translation Require Context? A Data-driven, Multilingual Exploration

Patrick FernandesKayo YinEmmy LiuAndré F. T. MartinsGraham Neubig

article2023ACL56 citationsResource Award

Introduces an information-theoretic methodology and multilingual benchmark spanning 14 language pairs to systematically identify context-dependent translation phenomena and assess how effectively context-aware models resolve discourse-level ambiguities.

Listen

Standard machine translation systems frequently struggle with document-level discourse features, such as maintaining pronoun gender, preserving consistent levels of politeness, and ensuring lexical cohesion across sentences. Because these context-dependent words represent a small fraction of overall text, standard translation quality metrics like BLEU often obscure whether models are properly using broader document context. Previous efforts to assess discourse translation relied heavily on manual intuition and were restricted to a small number of language pairs.

The article aims to systematically identify when translation requires multi-sentence context and to establish a standardized, data-driven benchmark—the Multilingual Discourse-Aware (MuDA) benchmark—to evaluate how well context-aware machine translation models resolve discourse ambiguities across diverse languages.

The researchers developed an information-theoretic metric called Pointwise Cross-Mutual Information (P-CXMI) to identify words whose translation probability increases substantially when surrounding context is provided. Using parallel transcripts from TED talks covering English to 14 target languages (including German, French, Mandarin, Japanese, Russian, and Arabic), the authors categorized discourse phenomena into five key classes: pronoun choice, formality, verb form consistency, lexical cohesion, and ellipsis. They built automated taggers to label these phenomena across test sets and evaluated both custom models (small and large Transformer architectures) and commercial engines (Google Cloud Translation and DeepL).

The analysis reveals that current context-aware translation architectures provide only marginal and inconsistent improvements over baseline models that translate sentence by sentence. While context-aware models show statistically significant gains on specific phenomena like formality and ellipsis, they fail to reliably resolve complex challenges such as verb form cohesion and lexical consistency. Among commercial systems, DeepL operated with full-document context outperformed its single-sentence baseline and Google Cloud Translation across most discourse metrics, but substantial ambiguity gaps remain across all tested systems.

These findings indicate that simply extending context windows or concatenating adjacent sentences does not resolve core document-level translation issues. Standard corpus-wide evaluation scores can be misleading for organizations that rely on high-fidelity translations, where incorrect formality or mismatched pronouns create reputational and operational risks. Decision-makers should recognize that existing commercial and open-source models cannot yet be trusted to maintain document-wide coherence automatically.

Organizations developing or deploying contextual translation systems should integrate targeted, phenomenon-specific evaluation tools like MuDA rather than relying solely on global performance metrics. Future research and development should focus on specialized architectures that directly target identified weaknesses, particularly verb form consistency and lexical cohesion, before relying on automated pipelines for publication-grade document translation.

Confidence in the tagging framework is high for lexical choice, formality, pronouns, and verb forms (which demonstrated precision near 1.0 in manual evaluations), but results for ellipsis tagging are less precise (0.26 to 0.78 precision) due to word-alignment errors in non-literal translations. Furthermore, the benchmark relies on exact word-matching metrics that may penalize valid synonymous translations, and evaluations did not include the latest generation of massive large language models trained on long contexts.

arXiv: 2109.07446CoderPat/MuDA

No sufficiently relevant recommendations were found.

Cover for When Does Translation Require Context? A Data-driven, Multilingual Exploration

Abstract

Although proper handling of discourse significantly contributes to the quality of machine translation (MT), these improvements are not adequately measured in common translation quality metrics. Recent works in context-aware MT attempt to target a small set of discourse phenomena during evaluation, however not in a fully systematic way. In this paper, we develop the Multilingual Discourse-Aware (MuDA) benchmark, a series of taggers that identify and evaluate model performance on discourse phenomena in any given dataset. The choice of phenomena is inspired by a novel methodology to systematically identify translations requiring context. We confirm the difficulty of previously studied phenomena while uncovering others that were previously unaddressed. We find that common context-aware MT models make only marginal improvements over context-agnostic models, which suggests these models do not handle these ambiguities effectively. We release code and data for 14 language pairs to encourage the MT community to focus on accurately capturing discourse phenomena.

Table of Contents

  • 1 Introduction
  • 2 Measuring Context Usage
  • 2.1 Cross-Mutual Information
  • 2.2 Context Usage Per Sentence and Word
  • 3 Which Translation Phenomena Benefit from Context?
  • 3.1 Data & Model
  • 3.2 Analysis Procedure
  • 3.3 Identified Phenomena
  • 3.3.1 Lexical Cohesion
  • 3.3.2 Formality
  • 3.3.3 Pronoun Choice
  • 3.3.4 Verb Form
  • 3.3.5 Ellipsis
  • 4 Cross-phenomenon MT Evaluation
  • 4.1 MT Evaluation Framework
  • 4.2 Automatic Tagging
  • 4.3 Evaluation of Automatic Tags
  • 4.4 Extension to New Languages
  • 5 Exploring Context-aware MT
  • 5.1 Trained Models
  • 5.2 Commercial Models
  • 5.3 Results and Discussion
  • 6 Related Work
  • 7 Conclusions and Future Work
  • Limitations
  • Acknowledgements
  • References
  • A MuDA Toolkit Usage
  • B Language Properties
  • C P-CXMI Results
  • D Tag Numbers
  • E Tagging other Document-level Datasets
  • F Tagger Details
  • F.1 Formality Words
  • F.2 Ambiguous Pronouns
  • F.3 Ambiguous Verbs
  • F.4 Ellipsis Classifier
  • G Training details
  • H Results Tables

Knowls

  1. Knowl 1 — Pointwise cross-mutual information measures context’s effect on translation

    equation

    For a source sentence xx, its target translation yy, and preceding target-side context CC, Pointwise Cross-Mutual Information (P-CXMI) measures the change in the model-assigned likelihood of the reference when context is provided. Let qAq_A be a context-agnostic translation model and qCq_C a context-aware translation model. Sentence-level P-CXMI is

    P ⁣-CXMI⁡(y,x,C)=−log⁡qA(y∣x)qC(y∣x,C).\operatorname{P\!\text{-}CXMI}(y,x,C)=-\log\frac{q_A(y\mid x)}{q_C(y\mid x,C)}.

    For target token yiy_i at position ii, autoregressive decoding gives the token-level measure

    P ⁣-CXMI⁡(i,y,x,C)=−log⁡qA(yi∣y<i,x)qC(yi∣y<i,x,C),\operatorname{P\!\text{-}CXMI}(i,y,x,C)=-\log\frac{q_A(y_i\mid y_{<i},x)}{q_C(y_i\mid y_{<i},x,C)},

    where y<iy_{<i} is the target prefix before token ii. Positive values indicate that the context-aware model assigns the token or sentence greater likelihood; negative values indicate lower likelihood. The models’ learned distributions determine the score. Averaged over NN held-out source–target–context examples, sentence-level P-CXMI estimates corpus-level CXMI as −1N∑n=1Nlog⁡qA(y(n)∣x(n))qC(y(n)∣x(n),C(n))-\frac{1}{N}\sum_{n=1}^{N}\log\frac{q_A(y^{(n)}\mid x^{(n)})}{q_C(y^{(n)}\mid x^{(n)},C^{(n)})}. The paper uses a single model trained to operate with and without context so that differences are attributable to context rather than other model factors.

  2. Knowl 2 — A data-driven analysis identifies five context-sensitive translation phenomena

    empirical result

    The authors identify context-sensitive translation patterns by examining high-P-CXMI items across 14 English-to-target-language TED-talk translation directions. Their analysis considers mean P-CXMI by part of speech, vocabulary items with high mean P-CXMI, and individual high-scoring tokens, then groups recurring patterns into phenomena. The resulting categories are: lexical cohesion, where an entity should receive a consistent target-language lexical choice throughout a document; formality, including T–V pronoun distinctions and honorifics; pronoun choice, where antecedent information can determine grammatical gender or the appropriate target pronoun; verb form, where context can guide choices reflecting tense, mood, tone, or document cohesion; and ellipsis, where omitted source material must be inferred to produce an appropriate target verb or other expression. The analysis surfaces verb-form consistency as a context-relevant category not previously addressed in the MT work discussed by the authors.

  3. Knowl 3 — MuDA automatically tags target words associated with discourse ambiguity

    algorithm

    The Multilingual Discourse-Aware (MuDA) tagger assigns one or more discourse-phenomenon tags to target-language tokens in parallel documents. It uses word alignment, coreference, linguistic analysis, and language-specific lists; the formality, pronoun, and verb-form lists were checked by native speakers.

    • Lexical cohesion: Lemmatize content words and obtain bidirectional word alignments between source and target documents. Tag a target word if its aligned source–target lemma pair has already occurred at least three times elsewhere in the same document, excluding the current sentence.
    • Formality: For languages with T–V distinctions, tag a target pronoun when an item of the same formality level has appeared earlier in the document. Where formality is expressed through verb inflection rather than an overt subject pronoun, detect a second-person source subject and the corresponding second-person (T) or third-person (V) target verb. For languages with honorific systems, tag words from a language-specific honorific-related list.
    • Pronoun choice: For each target language, provide a list mapping English pronouns to possible target pronouns. Use word alignment and coreference resolution; tag an aligned target pronoun when the source pronoun maps to that target form and its antecedent is outside the current sentence.
    • Verb form: Provide a list of target-language verb forms that have context-dependent alternatives for translating an English verb form. Tag a target verb from that list if the same form has already appeared earlier in its document.
    • Ellipsis: Train an English source-side BERT sentence classifier on Penn Treebank data, labeling sentences containing the treebank ellipsis tag as positive. The training set has 248,596 sentences, of which 2,863 are positive; the positive-class loss is up-weighted by 100. The classifier achieves 0.77 precision and 0.73 recall. For a sentence it predicts contains ellipsis, tag a target verb, noun, proper noun, or pronoun if the word appeared in an earlier target sentence in the document and has no alignment to a source word in the current sentence.

    The paper’s proposed extension procedure is to supply a word aligner for lexical-cohesion and ellipsis tagging, and add suitable language-specific lists for formality, pronouns, and verb forms.

  4. Knowl 4 — MuDA evaluates translation quality separately for each tagged phenomenon

    model/method

    Given parallel source and reference documents, MuDA attaches phenomenon tags to target-reference tokens. For each tag, the benchmark compares a system’s output with the reference and reports mean word-level F-measure over the tagged words, using surface-form matching. This provides a phenomenon-specific score—for example, for pronoun-choice or lexical-cohesion tokens—alongside corpus-level MT scores such as BLEU and COMET. A target token may have more than one tag, and the benchmark’s taggers do not require P-CXMI to be computed on the dataset being evaluated.

  5. Knowl 5 — Evaluation covers 14 English-to-target directions and multiple context conditions

    experimental setup

    The main dataset consists of English paired with Arabic, German, Spanish, French, Hebrew, Italian, Japanese, Korean, Dutch, Portuguese, Romanian, Russian, Turkish, and Mandarin Chinese. For each direction, TED-talk data contains 113,711 training sentence pairs from 1,368 talks, 2,678 development pairs from 41 talks, and 3,385 test pairs from 43 talks. The paper trains small Transformer translation systems with and without context; the context-aware evaluation model uses concatenation with a fixed context of three preceding target sentences. It also trains pretrained large context-aware and context-agnostic models for German, French, Japanese, and Chinese. Evaluations compare predicted-context and reference-context decoding where applicable, using BLEU and COMET corpus-wide and MuDA word F-measure per phenomenon. Commercial comparisons include Google Translate, DeepL’s document-level system, and a sentence-level DeepL ablation. For P-CXMI analysis, a separate small Transformer uses dynamic context sampled from zero to three target sentences. The small Transformer has hidden size 512, feedforward size 1,024, six layers, and eight attention heads; the large Transformer has hidden size 1,024, feedforward size 4,096, six layers, and 16 attention heads.

  6. Knowl 6 — Context-aware systems show selective, not consistent, gains on MuDA phenomena

    empirical result

    For the small trained systems, adding context often improves word F-measure on particular tagged categories, especially ellipsis and formality, but does not produce significant gains consistently across phenomena; lexical cohesion and verb form are examples where significant improvement is often absent. Across all words, the three small-system conditions do not significantly differ in mean word F-measure. Corpus-level metrics alone give a mixed picture: reference-context systems have the highest BLEU for most language pairs, while context-agnostic systems have higher COMET scores. The pretrained large systems generally benefit from context on corpus-level metrics, particularly COMET, and often achieve their best tagged-word scores with reference context, especially for lexical cohesion and pronouns; those tagged-word gains are less pronounced than the corpus-level gains. Among commercial systems, DeepL generally scores above Google on most metrics and language pairs, and its document-level version outperforms its sentence-level ablation on most MuDA tags. The model-comparison visualizations on pages 7–8 display these metric and per-tag patterns, marking statistically significant improvements at p<0.05p<0.05. Overall, the experiments show that the tested context-aware systems improve some discourse phenomena but do not consistently outperform context-agnostic systems on the benchmark.

  7. Knowl 7 — MuDA tags generally have high precision, with ellipsis as the main exception

    empirical result

    Native speakers with computational-linguistics backgrounds manually checked MuDA tags in eight languages, reviewing 50 randomly selected utterances per language as well as every automatically tagged ellipsis instance. Precision was high for lexical-cohesion, formality, pronoun, and verb-form tags where those categories applied; ellipsis precision was lower, in part because word alignment can miss target words in one-to-many or non-literal translations. The authors also report that tagged words have higher P-CXMI overall than untagged words, providing evidence that the tags tend to identify words whose model predictions benefit from context.

    Language Lexical Formality Pronouns Verb form Ellipsis
    Spanish 1.00 0.92 1.00 1.00 0.53
    French 1.00 1.00 1.00 0.94 0.43
    Japanese 1.00 1.00 1.00 – 0.41
    Korean 1.00 0.94 – – 0.26
    Portuguese 0.99 0.88 1.00 – 0.31
    Russian 1.00 1.00 – 1.00 0.50
    Turkish 1.00 1.00 – 1.00 0.57
    Chinese 1.00 1.00 – – 0.78

    Entries are tag precision, not recall; “--” indicates that the tag was not applicable or not reported for that language.

  8. Knowl 8 — The P-CXMI discovery procedure can be reused to find phenomena in new languages

    algorithm

    To investigate context-sensitive translation phenomena in a new language pair, the authors propose: (1) train a translation model with a dynamically sampled context size; (2) use it to calculate token-level P-CXMI on a parallel document-level corpus; (3) inspect high-P-CXMI part-of-speech categories, vocabulary items, and individual tokens; and (4) connect recurring patterns to discourse phenomena using linguistic resources. This method does not require the researchers to specify the phenomena in advance, although identifying and interpreting the patterns involves manual analysis.

  9. Knowl 9 — The benchmark’s scope is limited by tagging and scoring errors and model age

    limitation

    MuDA depends on upstream tools such as word alignment and coreference resolution, whose errors or weak out-of-domain generalization can affect the tags. Surface-form word F-measure can also penalize a contextually correct translation that uses a valid synonym or other equivalent wording. Finally, the evaluated systems may not represent newer state-of-the-art models, particularly systems based on large language models trained on long-context data. The paper also reports using a single random seed per experiment.

Coverage note — The full language-specific word lists and exhaustive per-language score matrices are omitted: the lists instantiate the tagging rules, while the matrices add detail beyond the summarized cross-system findings and tag-precision evidence.

References

  1. 1.Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. Findings of the 2019 conference on machine translation (WMT19). In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 1–61, Florence, Italy. Association for Computational Linguistics.
  2. 2.Rachel Bawden, Rico Sennrich, Alexandra Birch, and Barry Haddow. 2018. Evaluating discourse phenomena in neural machine translation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1304–1313, New Orleans, Louisiana. Association for Computational Linguistics.
  3. 3.Ann Bies, Mark Ferguson, Karen Katz, Robert MacIntyre, Victoria Tredinnick, Grace Kim, Mary Ann Marcinkiewicz, and Britta Schasberger. 1995. Bracketing guidelines for treebank ii style penn treebank project. University of Pennsylvania, 97:100.
  4. 4.Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology, 3(2):77–101.
  5. 5.Emanuele Bugliarello, Sabrina J. Mielke, Antonios Anastasopoulos, Ryan Cotterell, and Naoaki Okazaki. 2020. It’s easier to translate out of English than into it: Measuring neural translation difficulty by cross-mutual information. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1640–1649, Online. Association for Computational Linguistics.
  6. 6.Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012. WIT3: Web inventory of transcribed and translated talks. In Proceedings of the 16th Annual conference of the European Association for Machine Translation, pages 261–268, Trento, Italy. European Association for Machine Translation.
  7. 7.Kenneth Ward Church and Patrick Hanks. 1990. Word association norms, mutual information, and lexicography. Computational Linguistics, 16(1):22–29.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  9. 9.Zi-Yi Dou and Graham Neubig. 2021. Word alignment by fine-tuning embeddings on parallel corpora. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2112–2128, Online. Association for Computational Linguistics.
  10. 10.Miquel Esplà, Mikel Forcada, Gema Ramírez-Sánchez, and Hieu Hoang. 2019. ParaCrawl: Web-scale parallel corpora for the languages of the EU. In Proceedings of Machine Translation Summit XVII Volume 2: Translator, Project and User Tracks, pages 118–119, Dublin, Ireland. European Association for Machine Translation.
  11. 11.Patrick Fernandes, Kayo Yin, Graham Neubig, and André F. T. Martins. 2021. Measuring and increasing context usage in context-aware machine translation. In Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), Virtual.
  12. 12.Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke S. Zettlemoyer. 2017. Allennlp: A deep semantic natural language processing platform.
  13. 13.Liane Guillou, Christian Hardmeier, Ekaterina Lapshinova-Koltunski, and Sharid Loáiciga. 2018. A pronoun test suite evaluation of the English–German MT systems at WMT 2018. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 570–577, Belgium, Brussels. Association for Computational Linguistics.
  14. 14.Christian Hardmeier, Marcello Fondazione, and Bruno Kessler. 2010. Modelling pronominal anaphora in statistical machine translation.
  15. 15.Matthew Honnibal and Ines Montani. 2017. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. To appear.
  16. 16.Prathyusha Jwalapuram, Shafiq Joty, Irina Temnikova, and Preslav Nakov. 2019. Evaluating pronominal anaphora in machine translation: An evaluation measure and a test suite. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2964–2975, Hong Kong, China. Association for Computational Linguistics.
  17. 17.Prathyusha Jwalapuram, Barbara Rychalska, Shafiq R. Joty, and Dominika Basaj. 2020. Can your context-aware MT system pass the dip benchmark tests? : Evaluation benchmarks for discourse phenomena in machine translation. CoRR, abs/2004.14607.
  18. 18.Samuel Läubli, Rico Sennrich, and Martin Volk. 2018. Has machine translation achieved human parity? a case for document-level evaluation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4791–4796, Brussels, Belgium. Association for Computational Linguistics.
  19. 19.António Lopes, M. Amin Farajian, Rachel Bawden, Michael Zhang, and André F. T. Martins. 2020. Document-level neural MT: A systematic comparison. In Proceedings of the 22nd Annual Conference of the European Association for Machine Translation, pages 225–234, Lisboa, Portugal. European Association for Machine Translation.
  20. 20.Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993. Building a large annotated corpus of English: The Penn Treebank. Computational Linguistics, 19(2):313–330.
  21. 21.Sameen Maruf and Gholamreza Haffari. 2018. Document context neural machine translation with memory networks. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1275–1284, Melbourne, Australia. Association for Computational Linguistics.
  22. 22.Sameen Maruf, Fahimeh Saleh, and Gholamreza Haffari. 2021. A survey on document-level neural machine translation: Methods and evaluation. ACM Comput. Surv., 54(2).
  23. 23.Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018. Document-level neural machine translation with hierarchical attention networks. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2947–2954, Brussels, Belgium. Association for Computational Linguistics.
  24. 24.Makoto Morishita, Jun Suzuki, and Masaaki Nagata. 2020. JParaCrawl: A large scale web-based English-Japanese parallel corpus. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 3603–3609, Marseille, France. European Language Resources Association.
  25. 25.Mathias Müller, Annette Rios, Elena Voita, and Rico Sennrich. 2018. A large-scale test set for the evaluation of context-aware pronoun translation in neural machine translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 61–72, Brussels, Belgium. Association for Computational Linguistics.
  26. 26.Graham Neubig, Zi-Yi Dou, Junjie Hu, Paul Michel, Danish Pruthi, and Xinyi Wang. 2019. compare-mt: A tool for holistic comparison of language generation systems. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 35–41, Minneapolis, Minnesota. Association for Computational Linguistics.
  27. 27.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53, Minneapolis, Minnesota. Association for Computational Linguistics.
  28. 28.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  29. 29.Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020. Stanza: A python natural language processing toolkit for many human languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 101–108, Online. Association for Computational Linguistics.
  30. 30.Ye Qi, Devendra Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig. 2018. When and why are pre-trained word embeddings useful for neural machine translation? In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 529–535, New Orleans, Louisiana. Association for Computational Linguistics.
  31. 31.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  32. 32.Jörg Tiedemann and Yves Scherrer. 2017. Neural machine translation with extended context. In Proceedings of the Third Workshop on Discourse in Machine Translation, pages 82–92, Copenhagen, Denmark. Association for Computational Linguistics.
  33. 33.Antonio Toral, Sheila Castilho, Ke Hu, and Andy Way. 2018. Attaining the unattainable? reassessing claims of human parity in neural machine translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 113–123, Brussels, Belgium. Association for Computational Linguistics.
  34. 34.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008.
  35. 35.Elena Voita, Rico Sennrich, and Ivan Titov. 2019a. Context-aware monolingual repair for neural machine translation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 877–886, Hong Kong, China. Association for Computational Linguistics.
  36. 36.Elena Voita, Rico Sennrich, and Ivan Titov. 2019b. When a good translation is wrong in context: Context-aware machine translation improves on deixis, ellipsis, and lexical cohesion. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1198–1212, Florence, Italy. Association for Computational Linguistics.
  37. 37.Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov. 2018. Context-aware neural machine translation learns anaphora resolution. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1264–1274, Melbourne, Australia. Association for Computational Linguistics.
  38. 38.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  39. 39.Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, and Graham Neubig. 2021. Do context-aware translation models pay the right attention? In Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), Virtual.

Citation

MLA
Fernandes, P., et al. “When Does Translation Require Context? A Data-driven, Multilingual Exploration”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 606–26, https://doi.org/10.18653/v1/2023.acl-long.36.
APA
Fernandes, P., Yin, K., Liu, E., Martins, A. F. T., & Neubig, G. (2023). When Does Translation Require Context? A Data-driven, Multilingual Exploration. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 606–626. https://doi.org/10.18653/v1/2023.acl-long.36
Chicago
Fernandes, P., K. Yin, E. Liu, A. F. T. Martins, and G. Neubig. 2023. “When Does Translation Require Context? A Data-driven, Multilingual Exploration”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 606–26. https://doi.org/10.18653/v1/2023.acl-long.36.
Harvard
Fernandes, P. et al. (2023) “When Does Translation Require Context? A Data-driven, Multilingual Exploration”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 606–626. Available at: https://doi.org/10.18653/v1/2023.acl-long.36.
Vancouver
1. Fernandes P, Yin K, Liu E, Martins AFT, Neubig G (2023) When Does Translation Require Context? A Data-driven, Multilingual Exploration. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 606–626

BibTeX

@inproceedings{fernandes-etal-2023-translation,
    title = "When Does Translation Require Context? A Data-driven, Multilingual Exploration",
    author = "Fernandes, Patrick  and
      Yin, Kayo  and
      Liu, Emmy  and
      Martins, Andr{\'e}  and
      Neubig, Graham",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.36/",
    doi = "10.18653/v1/2023.acl-long.36",
    pages = "606--626"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/