The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study

Verna DankersElia BruniDieuwke Hupkes

article2022ACL93 citations

Demonstrates that neural machine translation models fail to appropriately balance local and global compositionality when evaluated on natural data, exposing the limitations of assessing compositional generalization through rigid synthetic benchmarks alone.

Listen

Artificial intelligence systems for language processing are often criticized for failing to generalize like humans. In language, understanding typically relies on compositionality, which is the ability to build the meaning of a complex phrase from the meanings of its individual parts. Previous evaluations have mostly used rigid, synthetic tests that assume language works like arithmetic, where each part is processed independently of its wider context. However, natural language requires a complex balance between local interpretation, where independent parts retain fixed meanings, and global interpretation, where surrounding context resolves ambiguities or idiomatic phrases. The article evaluates how state-of-the-art neural machine translation models handle these competing forms of compositionality when trained on real-world data.

To evaluate this behavior, the authors trained standard Transformer translation models on an English-to-Dutch dataset across three data scales, ranging from one million to sixty-nine million sentence pairs. They adapted three core tests from linguistic literature: systematicity, which checks whether combining familiar components produces consistent translations; substitutivity, which evaluates whether swapping synonyms preserves overall sentence translations; and overgeneralization, which examines whether models correctly recognize non-literal idioms rather than translating them word-for-word. The evaluations tested these capabilities across synthetic sentences, semi-natural sentences extracted from real grammatical patterns, and fully natural text.

The investigation produced four central findings. First, translation models exhibit surprisingly low consistency. Modifying a single word in one part of a sentence frequently caused the model to alter the translation of an entirely unrelated, distant clause. Second, increasing the volume of training data improved local consistency; models trained on the full dataset achieved higher consistency than those trained on smaller subsets. Third, when translating idioms, models progress through three distinct learning phases across training: an initial phase of generating generic high-frequency words, a peak overgeneralization phase where idioms are translated literally and incorrectly, and a final memorization phase where idiomatic translations are learned. Finally, a manual review of nine hundred inconsistent translation pairs revealed that forty percent of inconsistencies in systematicity tests were outright translation errors, while thirty-eight percent were acceptable stylistic rephrasings and sixteen percent stemmed from source text ambiguities.

These findings indicate that translation models struggle to modulate appropriately between local and global processing. The models frequently apply global context to alter translations where strict local rules should apply, introducing unforced grammatical and lexical errors. Conversely, they often translate idioms too locally when the surrounding context provides insufficient support. For decision-makers, this erratic sensitivity poses operational risks for automated systems where output stability, reliability, and predictability are critical, particularly in lower-resource settings where smaller training datasets make models significantly more error-prone.

Organizations deploying neural translation models should not assume standard benchmark metrics guarantee robust contextual reasoning. Technical teams should establish specialized consistency and robustness audits alongside traditional accuracy evaluations before deploying models in production. Future research and development must focus on creating evaluation benchmarks derived from natural language rather than artificial test sets, while investigating model architectures that can systematically separate default local processing from context-driven global adjustments.

These conclusions carry high confidence within the scope of modern Transformer architectures, though readers should note the boundary conditions of the study. The analysis focused specifically on English-to-Dutch translation, and while semi-natural and natural datasets were utilized, certain tests required controlled sentence templates to isolate grammatical variables. Stakeholders should exercise caution when extrapolating these specific error rates to substantially different language pairs or non-generative language tasks.

arXiv: 2108.05885
Cover for The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study

Abstract

Obtaining human-like performance in NLP is often argued to require compositional generalisation. Whether neural networks exhibit this ability is usually studied by training models on highly compositional synthetic data. However, compositionality in natural language is much more complex than the rigid, arithmetic-like version such data adheres to, and artificial compositionality tests thus do not allow us to determine how neural models deal with more realistic forms of compositionality. In this work, we re-instantiate three compositionality tests from the literature and reformulate them for neural machine translation (NMT). Our results highlight that: i) unfavourably, models trained on more data are more compositional; ii) models are sometimes less compositional than expected, but sometimes more, exemplifying that different levels of compositionality are required, and models are not always able to modulate between them correctly; iii) some of the non-compositional behaviours are mistakes, whereas others reflect the natural variation in data. Apart from an empirical study, our work is a call to action: we should rethink the evaluation of compositionality in neural networks and develop benchmarks using real data to evaluate compositionality on natural language, where composing meaning is not as straightforward as doing the math.^1

Table of Contents

  • 1 Introduction
  • 2 Local and global compositionality
  • 3 Setup
  • 3.1 Model and training
  • 3.2 Evaluation data
  • 4 Experiments and results
  • 4.1 Systematicity
  • 4.1.1 Experiments
  • 4.1.2 Results
  • 4.2 Substitutivity
  • 4.2.1 Experiments
  • 4.2.2 Results
  • 4.3 Global compositionality
  • 4.3.1 Experiments
  • 4.3.2 Results
  • 5 Manual analysis
  • 6 Related work
  • 7 Discussion
  • Acknowledgements
  • References
  • Appendix A Semi-natural templates
  • Appendix B Systematicity
  • Appendix C Substitutivity
  • Appendix D Global compositionality
  • Appendix E Reproducibility details
  • E.1 Data
  • E.2 Architecture and training
  • E.3 Compute
  • Appendix F Manual analysis
  • F.1 Setup
  • F.2 Results
  • F.2.1 Rephrasing
  • F.2.2 Source ambiguities
  • F.2.3 Target errors
  • F.2.4 Formatting
  • F.2.5 Inconsistentcies in synonym translations

Knowls

  1. Knowl 1 — Systematicity Evaluation Protocol for Neural Machine Translation

    model/method

    Systematicity evaluates a translation model's capacity to process novel recombinations of known lexical and syntactic constituents consistently. In neural machine translation (NMT), systematicity is evaluated using two sentence-level context-free grammar configurations:

    1. Noun and Verb Phrase Recombinations (S→NP VPS \to NP\ VP):

      • Condition NP→NP′NP \to NP': The subject noun within the noun phrase (NPNP) is substituted with a different noun while preserving grammatical number agreement with the verb phrase (VPVP).
      • Condition VP→VP′VP \to VP': A noun within the verb phrase (VPVP) is substituted with a different noun.
      • Systematicity is quantified by evaluating whether the translation of the unchanged constituent (VPVP in NP→NP′NP \to NP', or the subject NPNP in VP→VP′VP \to VP') remains identical before and after the replacement.
    2. Conjoined Sentence Recombinations (S→S CONJ SS \to S\ \text{CONJ}\ S):

      • Two sentences S1S_1 and S2S_2 are concatenated using the conjunction "and".
      • Condition S1→S1′S_1 \to S_1': A minimal lexical substitution is applied to S1S_1 (specifically, modifying the noun within its verb phrase).
      • Condition S1→S3S_1 \to S_3: Sentence S1S_1 is completely replaced by an independently sampled sentence S3S_3 with a distinct syntactic template.
      • Systematicity is quantified by evaluating whether the translation of the second conjunct S2S_2 remains invariant under perturbations to the first conjunct (S1S_1).

    Consistency is defined as the exact string equality between target translations of the invariant component across paired inputs, after controlling for deterministic morphological adaptations in the target language (such as Dutch definite article selection, de vs. het).

  2. Knowl 2 — Substitutivity Evaluation Protocol for Neural Machine Translation

    model/method

    Substitutivity evaluates whether replacing a word with an exact synonym preserves the translation of the surrounding sentence. The evaluation uses 20 pairs of British and American English synonyms that share identical target translations in Dutch (e.g., doughnut/donut →\to donut, aubergine/eggplant →\to aubergine, whisky/whiskey →\to whisky).

    Each test instance consists of a source sentence containing the British variant and a counterpart sentence containing the American variant across three data modalities:

    • Natural data: Sentences extracted from the OPUS corpus containing the target term.
    • Synthetic and Semi-natural data: Synonyms embedded within relative clauses attached to animate subject nouns (e.g., "The poet criticises the king that eats the {doughnut, donut}").

    Two consistency metrics are evaluated:

    1. Full Sentence Consistency: The proportion of sentence pairs where the model produces identical full target translations for both inputs (treating instances where the synonym is dropped in both as consistent if the rest of the translation matches).
    2. Synonym Translation Consistency: The proportion of sentence pairs where the extracted translation span corresponding to the synonym itself remains identical, independent of changes in the surrounding sentence context. The synonym target span is identified using the top-5 beam hypotheses of the model evaluated on isolated probe sentences of the form "This is the [NOUN]".
  3. Knowl 3 — Idiom Overgeneralisation Protocol for Evaluating Global Compositionality

    model/method

    Idiom overgeneralisation evaluates whether an NMT model erroneously applies local, compositional translation rules to expressions requiring non-compositional (global) translation. The test utilizes 20 English idioms selected from the MAGPIE corpus (Haagsma et al., 2020) having at least 200 occurrences in OPUS and where >80%>80\% of natural corpus target translations do not contain a word-by-word literal translation.

    Idioms are tested across four context conditions:

    1. Natural context: Sentences extracted directly from OPUS containing the idiom.
    2. Synthetic and Semi-natural contexts: Idioms inserted into templates via subordinate clauses attached to animate nouns (e.g., "The poet criticises the king that said 'I knew the formula by heart'").
    3. Unsupportive context: Idioms embedded within a random sequence of ten nouns to test baseline local bias.

    Evaluation Metric: A translation is labeled as an overgeneralised (literal) translation if it contains predefined target-language keywords that indicate a strictly local, word-by-word translation rather than an idiomatic paraphrase. For example, for the idiom "by heart", an idiomatic Dutch translation is "uit het hoofd" ("from the head"); the presence of the literal keyword "hart" ("heart") flags the translation as overgeneralised.

  4. Knowl 4 — Empirical Systematicity Results Across Training Scales and Sentence Types

    data/table

    NMT models evaluated across systematicity tests display sensitivity to distant, unrelated contextual changes, violating strong local compositionality. Consistency increases as training data volume grows, and minimal modifications to a sentence (S1→S1′S_1 \to S_1') yield higher consistency than full sentence replacements (S1→S3S_1 \to S_3).

    Evaluation Data Condition Training Dataset Size
    Small (1M) Medium (8.6M) Full (69M)
    S→NP VPS \to NP\ VP
    Synthetic NP→NP′NP \to NP' 0.73 0.84 0.84
    Synthetic VP→VP′VP \to VP' 0.76 0.87 0.88
    Semi-natural NP→NP′NP \to NP' 0.63 0.66 0.64
    S→S CONJ SS \to S\ \text{CONJ}\ S
    Synthetic S1→S1′S_1 \to S_1' 0.81 0.90 0.92
    Synthetic S1→S3S_1 \to S_3 0.53 0.76 0.82
    Semi-natural S1→S1′S_1 \to S_1' 0.65 0.73 0.76
    Semi-natural S1→S3S_1 \to S_3 0.29 0.49 0.49
    Natural S1→S1′S_1 \to S_1' 0.58 0.67 0.72
    Natural S1→S3S_1 \to S_3 0.25 0.39 0.47

    The consistency of the invariant second conjunct (S2S_2) drops as low as 0.25 in natural data under full first-conjunct replacement (S1→S3S_1 \to S_3) for small models, showing that conjoined clauses are not processed independently.

  5. Knowl 5 — Empirical Substitutivity Results and Contextual Translation Drift

    data/table

    Substitutivity experiments on English-to-Dutch NMT models demonstrate that substituting synonymous terms causes substantial translation divergence in surrounding sentence contexts, even when the model translates the synonym itself consistently.

    Evaluation Data Metric Training Dataset Size
    Small (1M) Medium (8.6M) Full (69M)
    Synthetic Full Sentence Consistency 0.49 0.67 0.76
    Synthetic Synonym-Only Consistency 0.67 0.82 0.93
    Semi-natural Full Sentence Consistency 0.34 0.55 0.62
    Semi-natural Synonym-Only Consistency 0.62 0.84 0.93
    Natural Full Sentence Consistency 0.37 0.52 0.63
    Natural Synonym-Only Consistency 0.61 0.75 0.85

    While synonym-only translation consistency reaches 0.93 on synthetic and semi-natural data for models trained on the full corpus, full sentence consistency remains substantially lower (0.62--0.76), indicating that lexical substitution triggers non-local translation changes across the sentence.

  6. Knowl 6 — Developmental Phases and Context Sensitivity of Idiom Overgeneralisation

    empirical result

    Tracking idiom overgeneralisation across training epochs in Transformer NMT models reveals three distinct developmental phases:

    1. Early Phase: Overgeneralisation is near zero because early checkpoint outputs consist primarily of generic high-frequency target tokens rather than literal lexical translations of the idiom.
    2. Peak Phase: Overgeneralisation spikes sharply as the network internalizes general, local compositional rules, translating constituent words of idioms literally (e.g., translating "by heart" word-for-word into Dutch "door hart" instead of idiomatic "uit het hoofd").
    3. Convergence/Memorisation Phase: Overgeneralisation declines as the model memorizes idiosyncratic global mappings for idiomatic phrases.

    Context Sensitivity:

    • Models trained on smaller training sets exhibit higher residual overgeneralisation at convergence compared to full-data models.
    • Idioms embedded in natural contexts yield lower overgeneralisation than those in synthetic or semi-natural templates.
    • When idioms are placed in unsupportive contexts (surrounded by 10 random words), overgeneralisation approaches 1.01.0 across all training scales, demonstrating that global idiomatic processing requires supporting contextual evidence.
  7. Knowl 7 — Error Taxonomy and Frequency Distribution of Translation Inconsistencies

    empirical result

    A manual analysis of 900 inconsistent translation pairs from systematicity tests and 900 from substitutivity tests classifies non-compositional behaviors into four primary error categories:

    1. Target Errors (40% in systematicity): Cases where one translation is ungrammatical or mistranslated while the other is correct. In substitutivity, 24% of inconsistencies involve the synonym directly (with omission/untranslated words being the most frequent error), while other errors comprise 46% (small model), 18% (medium), and 22% (full model).
    2. Rephrasing (38% in systematicity): Stylistic variations where both translations are equally valid (e.g., adverb repositioning, verb choice between wensen and willen, or noun choice between sporter and atleet).
    3. Source Ambiguities (16% in systematicity): In conjoined sentences (S1 and S2S_1 \text{ and } S_2), verbs in S1S_1 (e.g., wishes, thinks) spuriously take scope over S2S_2, altering Dutch word order from main clause (SVO) to subordinate clause (SOV).
    4. Formatting (6% in systematicity): Differences in punctuation (e.g., comma insertions) or compound spacing.

    Training dataset size shifts this distribution: models trained on small datasets produce mostly target errors (59% in systematicity), whereas models trained on full datasets produce mostly benign rephrasings (27% errors).

  8. Knowl 8 — Semi-Natural Sentence Generation Protocol via Tree-Substitution Grammar

    algorithm

    To evaluate compositionality on syntactically varied and plausible structures without losing experimental control, semi-natural test sentences are generated via data-oriented parsing using the following algorithm:

    Input: English OPUS corpus C, target template frames T
    Output: Corpus of semi-natural test sentences
    1. Sample a random subset of 100,000 English sentences from C
    2. Parse the sampled sentences into a constituent treebank using continuous data-oriented parsing (DiscoDOP)
    3. Decompose the treebank into constituent tree fragments conforming to a Tree-Substitution Grammar
    4. Extract the 100 most frequent NP (noun phrase) and VP (verb phrase) fragments that contain at least 15 non-terminal nodes
    5. Qualitatively select 5 NP fragments and 3 VP fragments from the top candidates to maximize structural diversity
    6. Search the treebank (treesearch) to extract all natural lexical instances matching the 8 selected syntactic fragments
    7. Construct 10 semi-natural sentence templates by embedding the retrieved NP and VP fragments into fixed sentential skeletons and systematically varying lexical slots
  9. Knowl 9 — English-Dutch Transformer Machine Translation Experimental Configuration

    experimental setup

    NMT models are trained on English-to-Dutch translation across three corpus scales derived from the OPUS collection (Tiedemann, 2020):

    • Full: 69,000,000 sentence pairs.
    • Medium: 8,582,811 sentence pairs (1/81/8 of full).
    • Small: 1,072,851 sentence pairs (1/641/64 of full).

    Architecture and Training Parameters:

    • Model: Transformer-base implemented in Fairseq (Vaswani et al., 2017), with 6 encoder layers, 6 decoder layers, hidden dimension dmodel=512d_{\text{model}} = 512, 8 attention heads, and feed-forward dimension dff=2048d_{\text{ff}} = 2048 (~80M trainable parameters).
    • Embeddings: Shared between encoder and decoder.
    • Preprocessing: Moses tokenizer; subword segmentation using Byte Pair Encoding (BPE) with 60,000 merge operations.
    • Optimization: Adam optimizer (β1=0.9,β2=0.98\beta_1 = 0.9, \beta_2 = 0.98), inverse square root learning rate scheduler with peak learning rate 0.00050.0005, linear warmup of 4000 updates from initial learning rate 10−710^{-7}.
    • Regularization: Dropout 0.3, weight decay 0.0001, label smoothing 0.1, clip-norm 0.0.
    • Batch Size: Maximum 3584 tokens per batch with gradient accumulation factor of 8.
    • Stopping Criterion: Early stopping with patience of 10 epochs on FLORES-101 dev set.
    • FLORES-101 devtest SacreBLEU scores (beam size 5, averaged over 5 seeds): Small = 20.6±0.420.6 \pm 0.4, Medium = 24.4±0.324.4 \pm 0.3, Full = 25.8±0.125.8 \pm 0.1.
  10. Knowl 10 — Strong versus Weak Compositionality Framework for Machine Translation

    definition

    In formal semantics and NLP evaluation, compositionality is distinguished along two dimensions:

    1. Strong (Local / Closed) Compositionality: The meaning (or translation) of a compound expression is determined strictly as a bottom-up function of the meanings of its immediate largest subparts, independent of external context (analogous to arithmetic evaluation, e.g., (3+5)=8(3+5) = 8 regardless of context). In machine translation, strong compositionality requires that modifying an unrelated constituent in a sentence or adjacent conjunct must not alter the translation of an invariant subconstituent.

    2. Weak (Global / Open) Compositionality: The meaning of a compound expression is a function of the meanings of its parts and their syntactic combination (Partee, 1984), but allows the interpretation of constituent parts and combining functions to depend on global syntactic structure, contextual disambiguation, or external arguments.

    Natural language requires modulating between local processing for regular compositional phrases and global processing for non-compositional phenomena (such as idioms, homonyms, and scope ambiguities).

Coverage note — None was omitted; all key contributions—including the local vs. global compositionality formulation, the three NMT evaluation protocols (systematicity, substitutivity, overgeneralisation), semi-natural dataset generation, experimental setup, full empirical tables, developmental dynamics, and manual error taxonomy—are covered across the knowls.

References

  1. 1.Giosuè Baggio. 2021. Compositionality in a parallel architecture for language processing. Cognitive Science, 45(5):e12949.
  2. 2.Jasmijn Bastings, Marco Baroni, Jason Weston, Kyunghyun Cho, and Douwe Kiela. 2018. Jump to better conclusions: SCAN both left and right. In Proceedings of the Workshop: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@EMNLP 2018, Brussels, Belgium, November 1, 2018, pages 47–55. Association for Computational Linguistics.
  3. 3.Rahma Chaabouni, Roberto Dessì, and Eugene Kharitonov. 2021. Can transformers jump around right in natural language? assessing performance transfer from scan. In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 136–148.
  4. 4.Verna Dankers, Anna Langedijk, Kate McCurdy, Adina Williams, and Dieuwke Hupkes. 2021. Generalising to German plural noun classes, from the perspective of a recurrent neural network. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 94–108, Online. Association for Computational Linguistics.
  5. 5.Emmanuel Dupoux. 2018. Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language-learner. Cognition, 173:43–59.
  6. 6.Marzieh Fadaee and Christof Monz. 2020. The unreasonable volatility of neural machine translation models. In Proceedings of the Fourth Workshop on Neural Generation and Translation, NGT@ACL 2020, Online, July 5-10, 2020, pages 88–96. Association for Computational Linguistics.
  7. 7.Catherine Finegan-Dollak, Jonathan K Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, and Dragomir Radev. 2018. Improving text-to-SQL evaluation methodology. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 351–360.
  8. 8.Jerry A Fodor and Zenon W Pylyshyn. 1988. Connectionism and cognitive architecture: A critical analysis. Cognition, 28(1-2):3–71.
  9. 9.Eduardo García-Ramírez. 2019. Open Compositionality: Toward a New Methodology of Language. Rowman & Littlefield.
  10. 10.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzman, and Angela Fan. 2021. The FLORES-101 evaluation benchmark for low-resource and multilingual machine translation. CoRR, abs/2106.03193.
  11. 11.Hessel Haagsma, Johan Bos, and Malvina Nissim. 2020. Magpie: A large corpus of potentially idiomatic expressions. In Proceedings of The 12th Language Resources and Evaluation Conference, pages 279–287.
  12. 12.Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni. 2020. Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intellgence Research, 67:757–795.
  13. 13.Pauline Jacobson. 2002. The (dis)organization of the grammar: 25 years. Linguistics and Philosophy, 25(5/6):601–626.
  14. 14.Theo MV Janssen. 1998. Algebraic translations, correctness and algebraic compiler construction. Theoretical Computer Science, 199(1-2):25–56.
  15. 15.Theo MV Janssen and Barbara H Partee. 1997. Compositionality. In Handbook of logic and language, pages 417–473. Elsevier.
  16. 16.Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, et al. 2019. Measuring compositional generalization: A comprehensive method on realistic data. In International Conference on Learning Representations.
  17. 17.Najoung Kim and Tal Linzen. 2020. COGS: a compositional generalization challenge based on semantic interpretation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9087–9105.
  18. 18.Kris Korrel, Dieuwke Hupkes, Verna Dankers, and Elia Bruni. 2019. Transcoding compositionally: Using attention to find more generalizable solutions. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 1–11.
  19. 19.Brenden Lake and Marco Baroni. 2018. Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In International Conference on Machine Learning, pages 2873–2882. PMLR.
  20. 20.Brenden M. Lake, Tal Linzen, and Marco Baroni. 2019. Human few-shot learning of compositional instructions. In Proceedings of the 41th Annual Meeting of the Cognitive Science Society, CogSci 2019: Creativity + Cognition + Computation, Montreal, Canada, July 24-27, 2019, pages 611–617. cognitivesciencesociety.org.
  21. 21.Yair Lakretz, German Kruszewski, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene, and Marco Baroni. 2019. The emergence of number and syntax units in LSTM language models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 11–20.
  22. 22.Yafu Li, Yongjing Yin, Yulong Chen, and Yue Zhang. 2021. On compositional generalization of neural machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4767–4780.
  23. 23.Gary F Marcus. 2003. The algebraic mind: Integrating connectionism and cognitive science. MIT press.
  24. 24.Mathijs Mul and Willem Zuidema. 2019. Siamese recurrent networks learn first-order logic reasoning and exhibit zero-shot compositional generalization. In CoRR, abs/1906.00180.
  25. 25.Ryan M Nefdt. 2020. A puzzle concerning compositionality in machines. Minds & Machines, 30(1).
  26. 26.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53.
  27. 27.Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018. Scaling neural machine translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, WMT 2018, Belgium, Brussels, October 31 - November 1, 2018, pages 1–9. Association for Computational Linguistics.
  28. 28.Peter Pagin and Dag Westerståhl. 2010. Compositionality ii: Arguments and problems. Philosophy Compass, 5(3):265–282.
  29. 29.Barbara Partee. 1984. Compositionality. Varieties of formal semantics, 3:281–311.
  30. 30.Prasanna Parthasarathi, Koustuv Sinha, Joelle Pineau, and Adina Williams. 2021. Sometimes we want ungrammatical translations. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3205–3227.
  31. 31.Ellie Pavlick and Chris Callison-Burch. 2016. Most “babies” are “little” and most “problems” are “huge”: Compositional entailment in adjective-nouns. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2164–2173.
  32. 32.Martina Penke. 2012. The dual-mechanism debate. In The Oxford handbook of compositionality.
  33. 33.Matt Post. 2018. A call for clarity in reporting bleu scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191.
  34. 34.Vikas Raunak, Vaibhav Kumar, Florian Metze, and Jaimie Callan. 2019. On compositionality in neural machine translation. In NeurIPS 2019 Context and Compositionality in Biological and Artificial Neural Systems Workshop.
  35. 35.MT Rosetta. 1994. The rosetta characteristics. In Compositional Translation, pages 85–102. Springer.
  36. 36.D E Rumelhart and J McClelland. 1986. On Learning the Past Tenses of English Verbs. In Parallel distributed processing: Explorations in the microstructure of cognition, pages 216–271. MIT Press, Cambridge, MA.
  37. 37.Naomi Saphra and Adam Lopez. 2020. LSTMs compose—and Learn—Bottom-up. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2797–2809, Online. Association for Computational Linguistics.
  38. 38.Peter Shaw, Ming-Wei Chang, Panupong Pasupat, and Kristina Toutanova. 2021. Compositional generalization and natural language variation: Can a semantic parsing approach handle both? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 922–938, Online. Association for Computational Linguistics.
  39. 39.Paul Smolensky. 1990. Tensor product variable binding and the representation of symbolic structures in connectionist systems. Artificial intelligence, 46(1-2):159–216.
  40. 40.Zoltan Szabó. 2012. The case for compositionality. The Oxford handbook of compositionality, 64:80.
  41. 41.Jörg Tiedemann. 2020. The tatoeba translation challenge – realistic data sets for low resource and multilingual MT. In Proceedings of the Fifth Conference on Machine Translation, pages 1174–1182, Online. Association for Computational Linguistics.
  42. 42.Jörg Tiedemann and Santhosh Thottingal. 2020. OPUS-MT – building open translation services for the world. In Proceedings of the 22nd Annual Conference of the European Association for Machine Translation, pages 479–480.
  43. 43.Andreas Van Cranenburgh, Remko Scha, and Rens Bod. 2016. Data-oriented parsing with discontinuous constituents and function tags. Journal of Language Modelling, 4(1):57–111.
  44. 44.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, pages 5998–6008.
  45. 45.Dag Westerståhl. 2002. On the compositionality of idioms. Proceedings of LLC8. CSLI Publications.

Citation

MLA
Dankers, V., et al. “The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 4154–75, https://doi.org/10.18653/v1/2022.acl-long.286.
APA
Dankers, V., Bruni, E., & Hupkes, D. (2022). The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4154–4175. https://doi.org/10.18653/v1/2022.acl-long.286
Chicago
Dankers, V., E. Bruni, and D. Hupkes. 2022. “The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4154–75. https://doi.org/10.18653/v1/2022.acl-long.286.
Harvard
Dankers, V., Bruni, E. and Hupkes, D. (2022) “The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4154–4175. Available at: https://doi.org/10.18653/v1/2022.acl-long.286.
Vancouver
1. Dankers V, Bruni E, Hupkes D (2022) The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4154–4175

BibTeX

@inproceedings{dankers-etal-2022-paradox,
    title = "The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study",
    author = "Dankers, Verna  and
      Bruni, Elia  and
      Hupkes, Dieuwke",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.286/",
    doi = "10.18653/v1/2022.acl-long.286",
    pages = "4154--4175"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/