The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study
Verna DankersElia BruniDieuwke Hupkes
Demonstrates that neural machine translation models fail to appropriately balance local and global compositionality when evaluated on natural data, exposing the limitations of assessing compositional generalization through rigid synthetic benchmarks alone.
Artificial intelligence systems for language processing are often criticized for failing to generalize like humans. In language, understanding typically relies on compositionality, which is the ability to build the meaning of a complex phrase from the meanings of its individual parts. Previous evaluations have mostly used rigid, synthetic tests that assume language works like arithmetic, where each part is processed independently of its wider context. However, natural language requires a complex balance between local interpretation, where independent parts retain fixed meanings, and global interpretation, where surrounding context resolves ambiguities or idiomatic phrases. The article evaluates how state-of-the-art neural machine translation models handle these competing forms of compositionality when trained on real-world data.
To evaluate this behavior, the authors trained standard Transformer translation models on an English-to-Dutch dataset across three data scales, ranging from one million to sixty-nine million sentence pairs. They adapted three core tests from linguistic literature: systematicity, which checks whether combining familiar components produces consistent translations; substitutivity, which evaluates whether swapping synonyms preserves overall sentence translations; and overgeneralization, which examines whether models correctly recognize non-literal idioms rather than translating them word-for-word. The evaluations tested these capabilities across synthetic sentences, semi-natural sentences extracted from real grammatical patterns, and fully natural text.
The investigation produced four central findings. First, translation models exhibit surprisingly low consistency. Modifying a single word in one part of a sentence frequently caused the model to alter the translation of an entirely unrelated, distant clause. Second, increasing the volume of training data improved local consistency; models trained on the full dataset achieved higher consistency than those trained on smaller subsets. Third, when translating idioms, models progress through three distinct learning phases across training: an initial phase of generating generic high-frequency words, a peak overgeneralization phase where idioms are translated literally and incorrectly, and a final memorization phase where idiomatic translations are learned. Finally, a manual review of nine hundred inconsistent translation pairs revealed that forty percent of inconsistencies in systematicity tests were outright translation errors, while thirty-eight percent were acceptable stylistic rephrasings and sixteen percent stemmed from source text ambiguities.
These findings indicate that translation models struggle to modulate appropriately between local and global processing. The models frequently apply global context to alter translations where strict local rules should apply, introducing unforced grammatical and lexical errors. Conversely, they often translate idioms too locally when the surrounding context provides insufficient support. For decision-makers, this erratic sensitivity poses operational risks for automated systems where output stability, reliability, and predictability are critical, particularly in lower-resource settings where smaller training datasets make models significantly more error-prone.
Organizations deploying neural translation models should not assume standard benchmark metrics guarantee robust contextual reasoning. Technical teams should establish specialized consistency and robustness audits alongside traditional accuracy evaluations before deploying models in production. Future research and development must focus on creating evaluation benchmarks derived from natural language rather than artificial test sets, while investigating model architectures that can systematically separate default local processing from context-driven global adjustments.
These conclusions carry high confidence within the scope of modern Transformer architectures, though readers should note the boundary conditions of the study. The analysis focused specifically on English-to-Dutch translation, and while semi-natural and natural datasets were utilized, certain tests required controlled sentence templates to isolate grammatical variables. Stakeholders should exercise caution when extrapolating these specific error rates to substantially different language pairs or non-generative language tasks.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Introduces the Transformer sequence-to-sequence architecture that forms the foundational model evaluated throughout the compositionality study.
- Paper: Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data, Emily M. Bender et al. (2020). Establishes foundational theoretical distinctions between linguistic form, semantic meaning, and true understanding in data-driven neural language systems.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). Provides diagnostic methodology demonstrating how neural models rely on superficial syntactic heuristics rather than robust compositional rules.
- Paper: Six Challenges for Neural Machine Translation, Philipp Koehn et al. (2017). Identifies key empirical vulnerabilities and brittleness in neural machine translation models under varying data scales and contextual conditions.
- Paper: Neural Machine Translation of Rare Words with Subword Units, Rico Sennrich et al. (2016). Presents subword tokenization via byte pair encoding, which directly governs how neural translation models represent and compose lexical units.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). Analyzes the geometry of contextualized representations and how attention layers modulate between static lexical identity and surrounding context.
- Paper: Neural Network Acceptability Judgments, Alex Warstadt et al. (2018). Introduces benchmarks for evaluating grammatical acceptability and systematic linguistic competence in neural sequence models.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). Extends the study of compositional reasoning failures from neural translation to multi-step question answering in large autoregressive language models.
- Paper: Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs, Kanishka Misra et al. (2024). Investigates how transformer models generalize rare, non-compositional grammatical constructions from frequent linguistic patterns during training.
- Paper: Consistency Analysis of ChatGPT, Myeongjun Jang et al. (2023). Evaluates logical and semantic consistency across modern generative models when subjected to targeted input perturbations and paraphrases.
- Paper: Measuring the Mixing of Contextual Information in the Transformer, Javier Ferrando et al. (2022). Develops layer-wise attribution methods to interpret how transformer attention blocks mix contextual and local token representations.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). Surveys systemic flaws in text generation evaluation practices, emphasizing the need for robust audits beyond surface metrics.
- Paper: On the Limitations of Reference-Free Evaluations of Generated Text, Daniel Deutsch et al. (2022). Critiques reference-free text generation metrics by revealing structural biases and failures to detect critical translation errors.
