Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs
Kanishka MisraKyle Mahowald
Demonstrates through controlled corpus ablation that language models acquire rare syntactic constructions by abstracting structural patterns from more frequent, related linguistic phenomena rather than relying solely on verbatim memorization.
A central debate in artificial intelligence and linguistics is whether statistical language models truly generalize abstract grammar or merely act as memorization engines that mimic massive training datasets. This question is especially critical when evaluating how systems handle rare linguistic expressions that appear infrequently in everyday language. The article evaluates whether language models trained on human-scale data can learn rare grammatical constructions through structural generalization from more common, related phrasing rather than direct memorization.
To test this, the researchers trained 97-million-parameter transformer language models on systematically modified versions of the 100-million-word BabyLM dataset, a corpus matching the scale of human developmental language exposure. They focused on the rare Article + Adjective + Numeral + Noun construction, such as "a beautiful five days," which constitutes only about 0.02% of the training text. The investigation analyzed model performance across carefully controlled conditions: standard training, complete removal of the target construction, deliberate corruption of word orders, systematic removal of related grammatical structures, and variations in input vocabulary diversity. Model evaluations used syntactic acceptability benchmarks compared against human baseline ratings.
Four primary findings emerged from the analysis. First, language models exposed to standard training successfully identified the target construction as grammatical with approximately 70% accuracy, compared to a 20% chance baseline. Second, models trained without ever seeing a single instance of the construction still achieved 47% accuracy—substantially above chance and well above performance on ungrammatical control word orders. Third, removing frequent but structurally related phenomena from the training data, such as plural measure nouns acting as singular units (for example, "a few days" or "five miles is"), reduced the likelihood assigned to unseen target sentences by an average of 36.5%. Fourth, models exposed to higher vocabulary diversity across the adjective, numeral, and noun slots demonstrated significantly stronger generalization than those exposed to repetitive, low-variability examples.
These findings indicate that general-purpose statistical learners can acquire rare and complex grammatical structures by drawing structural abstractions from more frequent, related input. This demonstrates that models do not require trillions of words or verbatim memorization to master long-tail grammatical nuances. For organizations investing in artificial intelligence, this suggests that smaller, highly structured, and diverse training datasets can achieve strong linguistic competence, potentially lowering data collection, computing, and fine-tuning costs while reducing reliance on massive data scale.
Decision-makers and research teams should prioritize data curation that maximizes structural and lexical diversity rather than relying exclusively on expanding corpus volume. Before applying these insights to operational pipelines, organizations should conduct targeted pilot tests to evaluate whether these structural generalization mechanisms transfer across different languages and directly improve downstream natural language understanding tasks.
Confidence in these findings is supported by rigorous control conditions, including artificial noise injections confirming that imperfect detection of the target construction did not distort results. However, the study is bounded by its focus on English syntax within a single autoregressive model architecture and evaluated structural acceptability rather than task-specific semantic comprehension.
- Paper: Neural Network Acceptability Judgments, Alex Warstadt et al. (2018). Introduces the foundational framework and benchmark dataset for evaluating grammatical acceptability judgments in neural language models.
- Paper: A Structural Probe for Finding Syntax in Word Representations, John Hewitt et al. (2019). Establishes structural probing methods to demonstrate that deep contextual language representations implicitly encode hierarchical syntactic trees.
- Paper: What Does BERT Learn about the Structure of Language?, Ganesh Jawahar et al. (2019). Analyzes how transformer representations layer-by-layer encode syntactic and structural properties of natural language.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). Examines whether statistical learners rely on surface heuristics versus genuine abstract syntactic generalizations when processing language.
- Paper: When transformers learn "impossible" languages, what do they learn?, Ram Janarthan et al. (2026). Extends BabyLM-scale evaluation of grammatical sensitivity by investigating how transformers learn systematically perturbed and linguistically impossible word-order variants.
