Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs

Kanishka MisraKyle Mahowald

article2024EMNLP75 citations

Demonstrates through controlled corpus ablation that language models acquire rare syntactic constructions by abstracting structural patterns from more frequent, related linguistic phenomena rather than relying solely on verbatim memorization.

Listen

A central debate in artificial intelligence and linguistics is whether statistical language models truly generalize abstract grammar or merely act as memorization engines that mimic massive training datasets. This question is especially critical when evaluating how systems handle rare linguistic expressions that appear infrequently in everyday language. The article evaluates whether language models trained on human-scale data can learn rare grammatical constructions through structural generalization from more common, related phrasing rather than direct memorization.

To test this, the researchers trained 97-million-parameter transformer language models on systematically modified versions of the 100-million-word BabyLM dataset, a corpus matching the scale of human developmental language exposure. They focused on the rare Article + Adjective + Numeral + Noun construction, such as "a beautiful five days," which constitutes only about 0.02% of the training text. The investigation analyzed model performance across carefully controlled conditions: standard training, complete removal of the target construction, deliberate corruption of word orders, systematic removal of related grammatical structures, and variations in input vocabulary diversity. Model evaluations used syntactic acceptability benchmarks compared against human baseline ratings.

Four primary findings emerged from the analysis. First, language models exposed to standard training successfully identified the target construction as grammatical with approximately 70% accuracy, compared to a 20% chance baseline. Second, models trained without ever seeing a single instance of the construction still achieved 47% accuracy—substantially above chance and well above performance on ungrammatical control word orders. Third, removing frequent but structurally related phenomena from the training data, such as plural measure nouns acting as singular units (for example, "a few days" or "five miles is"), reduced the likelihood assigned to unseen target sentences by an average of 36.5%. Fourth, models exposed to higher vocabulary diversity across the adjective, numeral, and noun slots demonstrated significantly stronger generalization than those exposed to repetitive, low-variability examples.

These findings indicate that general-purpose statistical learners can acquire rare and complex grammatical structures by drawing structural abstractions from more frequent, related input. This demonstrates that models do not require trillions of words or verbatim memorization to master long-tail grammatical nuances. For organizations investing in artificial intelligence, this suggests that smaller, highly structured, and diverse training datasets can achieve strong linguistic competence, potentially lowering data collection, computing, and fine-tuning costs while reducing reliance on massive data scale.

Decision-makers and research teams should prioritize data curation that maximizes structural and lexical diversity rather than relying exclusively on expanding corpus volume. Before applying these insights to operational pipelines, organizations should conduct targeted pilot tests to evaluate whether these structural generalization mechanisms transfer across different languages and directly improve downstream natural language understanding tasks.

Confidence in these findings is supported by rigorous control conditions, including artificial noise injections confirming that imperfect detection of the target construction did not distort results. However, the study is bounded by its focus on English syntax within a single autoregressive model architecture and evaluated structural acceptability rather than task-specific semantic comprehension.

arXiv: 2403.19827kanishkamisra/aannalysis
Cover for Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs

Abstract

Language models learn rare syntactic phenomena, but the extent to which this is attributable to generalization vs. memorization is a major open question. To that end, we iteratively trained transformer language models on systematically manipulated corpora which were human-scale in size, and then evaluated their learning of a rare grammatical phenomenon: the English Article+Adjective+Numeral+Noun (AANN) construction (“a beautiful five days”). We compared how well this construction was learned on the default corpus relative to a counterfactual corpus in which AANN sentences were removed. We found that AANNs were still learned better than systematically perturbed variants of the construction. Using additional counterfactual corpora, we suggest that this learning occurs through generalization from related constructions (e.g., “a few days”). An additional experiment showed that this learning is enhanced when there is more variability in the input. Taken together, our results provide an existence proof that LMs can learn rare grammatical phenomena by generalization from less rare phenomena. Data and code: https://github.com/kanishkamisra/aannanalysis.

Table of Contents

  • 1. Introduction
  • 1.1 Motivation and Prior Work
  • 1.2 Summary of findings
  • 2. General Methods
  • 2.1 Corpus
  • 2.2 Language Model
  • 2.3 Construction Detection
  • 2.4 Acceptability data
  • 2.5 Scoring and Accuracy
  • 2.6 Ablations
  • 3. Experiment 1: LMs learn about AANNs without having seen a single instance
  • 4. Experiment 2: Keys to Learning AANNs
  • 4.1 Analysis and Results
  • 5. Experiment 3: The Role of Variability
  • 6. Conclusion
  • 7. Limitations
  • 8. Acknowledgments
  • 9. Corrigendum
  • References
  • A. Dataset Access and Licensing
  • B. LM training details
  • C. Detecting AANNs and related phenomena
  • C.1 AANNs
  • C.2 DT ANNs
  • C.3 A few/couple/dozen NOUNs
  • C.4 Measure NNS with Singular Verbs
  • D. A/An + ADJ/NUM frequency balancing
  • E. Lexical semantic constraints on AANN slots
  • F. Variability Analysis

Knowls

  1. Knowl 1 — The AANN construction as a rare generalization target

    definition

    The paper studies the English Article–Adjective–Numeral–Noun (AANN) construction, in which an indefinite article is followed by an adjective, a numeral, and a plural noun, as in “a beautiful five days.” This construction is unusual relative to the ordinary order “five beautiful days.” The evaluation tests whether a language model prefers grammatical AANNs to four matched corruptions: adjective–numeral order reversal, omission of the article, omission of the adjective, and omission of the numeral. The test items are novel with respect to the model’s training corpus, so success is treated as evidence of generalization rather than verbatim recall.

  2. Knowl 2 — Controlled corpus interventions for causal analysis of rare-construction learning

    model/method

    The authors train language models from scratch on systematically manipulated versions of the 100-million-token BabyLM-strict corpus. The baseline corpus retains its 2,448 detected AANNs; a NO-AANN corpus removes them; counterfactual corpora replace them with the same lexical material in ungrammatical ANAN or NAAN orders; and further corpora remove hypothesized related phenomena such as “a few days” or “five dollars is a lot.” After an ablation, unaffected utterances are up-sampled so that every model encounters the same total number of tokens. Comparing models trained on these otherwise matched corpora allows the authors to test whether AANN behavior depends on direct AANN exposure, related constructions, or merely corpus size.

  3. Knowl 3 — Training corpus, models, and acceptability evaluation

    experimental setup

    The training data are the human-scale BabyLM-strict corpus, containing approximately 100 million tokens and 11.5 million utterances. The principal models are 97-million-parameter OPT-style autoregressive transformers with 12 layers, 12 attention heads, 768-dimensional embeddings, a 3,072-dimensional feed-forward layer, a 16,384-token vocabulary, maximum sequence length 256, batch size 32, 32,000 warm-up steps, and a maximum of 20 epochs. Reported results average three random-seed runs. A POS-tagging and regular-expression pipeline detected 2,448 AANNs, approximately 0.02% of the utterances; hand annotation of a final validation sample found 17 of 18 instances, giving an estimated 95% recall, with the single miss involving an apparent singular/plural typo.

    The evaluation source contains 12,960 templatically generated AANN sentences, including 3,420 with human acceptability ratings from 190 participants on a 1–10 scale. The main accuracy evaluation retains the 2,027 items rated above 7 and removes four items that occurred in the training corpus. Each retained item is compared with its four corrupted variants, and a model is correct only when it scores the grammatical AANN above all four alternatives.

  4. Knowl 4 — SLOR scoring and five-way accuracy criterion

    equation

    For a target construction CC occurring after a sentence prefix pp, the study uses the Syntactic Log-Odds Ratio (SLOR) rather than raw normalized log probability:

    SLOR⁡m,p(C)=1∣C∣log⁡(Pm(C∣p)Pu(C)).\operatorname{SLOR}_{m,p}(C)=\frac{1}{|C|}\log\left(\frac{P_m(C\mid p)}{P_u(C)}\right).

    Here mm is the autoregressive language model, uu is a unigram model trained on the same manipulated corpus and tokenizer, Pm(C∣p)P_m(C\mid p) is the model probability of the construction after the prefix, Pu(C)P_u(C) is its unigram probability, and ∣C∣|C| is the number of tokenizer tokens in CC. Normalizing by the unigram probability prevents corpus manipulations from being evaluated solely through changed unigram frequencies. An item is counted as correct when the well-formed construction has a higher SLOR than each of its four corrupted versions; because there are five alternatives, chance accuracy is 20%.

  5. Knowl 5 — AANNs are learned above chance even without direct examples

    empirical result

    Transformers trained on the unmodified BabyLM corpus achieved approximately 70% accuracy on novel AANN evaluations, despite AANNs constituting only about 0.02% of the training utterances. Removing all 2,448 detected AANNs reduced accuracy to approximately 47%, but this remained 27 percentage points above the 20% chance level. Thus, the models assigned a non-trivial preference to unseen grammatical AANNs even when no detected positive AANN example was present during training. A 4-gram language model trained on BabyLM reached only about 41%, whereas GPT-2 XL and Llama-2-7B reached 78% and 83%, respectively, indicating that the main result was not reproduced by shallow local n-gram statistics alone.

  6. Knowl 6 — Counterfactual substitutions do not teach the construction as effectively

    empirical result

    To distinguish learning the AANN pattern from learning arbitrary regularities in its lexical items, the authors replaced every training AANN with either ANAN, such as “a ninety whopping LMs,” or NAAN, such as “ninety whopping a LMs,” and evaluated each model on matched tests for AANN, ANAN, and NAAN. Models exposed to ANAN or NAAN learned those counterfactual patterns less successfully than models learned the grammatical AANN pattern. For example, the NAAN-trained model reached about 43% on unseen AANNs but only about 37% on NAANs, even though NAANs—not AANNs—were present in its training data. Replacing 300 detected AANNs, approximately 12% of them, with ANANs or NAANs produced nearly the same results as complete replacement. This pollution control suggests that the zero-shot AANN preference was not primarily caused by a small number of AANNs missed by the detector.

  7. Knowl 7 — Related phenomena selected for corpus ablation

    data/table

    The second experiment removes frequent phenomena that could provide structural or statistical evidence for AANNs, while retaining the same total token count through up-sampling. The target phenomena and their corpus frequencies were:

    Could not parse LaTeX table

    The DT+ANN condition tests whether phrases such as “the beautiful five days” support analogy from definite determiners. The measure-noun conditions test whether plural measure phrases can behave as singular units, and the frequency-balancing condition removes the strong corpus bias that adjectives are more likely than numerals to follow an indefinite article. The random-removal condition controls for effects caused merely by deleting a large quantity of training data.

  8. Knowl 8 — Frequent related constructions support zero-shot AANN judgments

    empirical result

    When AANNs were also removed, ablating the related phenomena caused substantial additional reductions in models’ normalized scores for unseen AANNs. The largest effects came from removing evidence that measure noun phrases can behave as singular units and from balancing the frequencies of adjectives and numerals after indefinite articles. The authors report that removing the measure-noun evidence made unseen AANNs less likely by 36.5% on average. Removing determiner-plus-adjective-plus-numeral-plus-noun sequences had little additional effect beyond removing AANNs alone, showing that simply deleting similar-looking strings was not sufficient to explain the result. Randomly deleting an equally large number of tokens did not reproduce the reductions, while 4-gram models showed mostly null effects except for the article–adjective/numeral frequency manipulation. When AANNs were retained, related-phenomenon ablations had smaller but still sometimes significant effects, indicating that direct AANN evidence and indirect evidence from frequent related constructions contribute additively.

  9. Knowl 9 — Variability in observed AANN fillers increases productivity

    empirical result

    The authors divided the 2,448 observed AANN utterances into two equal groups of 1,224 using median splits based on the number of unique fillers in the adjective, numeral, and noun slots. They trained models on corpora exposing the models either to low- or high-variability AANNs, considering all three open slots jointly and each slot separately. Models exposed to highly variable adjectives and nouns assigned higher normalized scores to novel AANNs than models exposed to restricted, repetitive fillers. The same direction held when all three slots were considered jointly. The numeral-only comparison was the exception: low- and high-variability numeral conditions produced roughly similar scores. Scores for both variability conditions fell between the NO-AANN and unablated conditions, suggesting that models use variation in observed slot fillers as evidence that the construction’s positions are productive.

  10. Knowl 10 — AANN learning includes lexical-semantic restrictions

    empirical result

    Across all 3,420 human-rated AANN types, the unmodified and NO-AANN models showed lexical preferences that broadly matched human judgments. Both models preferred quantitative adjectives, such as “mere” or “hefty,” and acceptable qualitative adjectives over stubbornly distributive adjectives, such as “blue” in “a blue five pencils.” These patterns persisted even when the models had seen no detected AANNs during training. The result suggests that generalization from related input can reproduce not only the AANN’s broad word-order pattern but also some lexical restrictions on which adjective and noun combinations are acceptable.

  11. Knowl 11 — Scope and methodological limitations

    limitation

    The study examines only one rare construction in English and trains only on English text, so it does not establish that the same learning mechanisms apply across constructions or languages. The controlled-rearing design requires repeatedly training models from scratch and is computationally expensive. The AANN detector had high but imperfect estimated recall, although pollution controls reduced concern that missed examples drove the zero-shot result. The evaluations target formal acceptability and lexical restrictions rather than downstream interpretation or task performance; preliminary analyses suggested that the corpus manipulations left lexical-semantic properties largely unchanged, but those properties were not directly ablated.

Coverage note — Detailed regular-expression listings, dependency-parser patterns, per-category plotting details, licensing information, and the corrigendum were omitted because they are implementation or provenance details rather than additional load-bearing findings.

References

  1. 1.Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, and Stephan Vogel. 2014. The AMARA corpus: Building parallel language resources for the educational domain. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), pages 1856–1862, Reykjavik, Iceland. European Language Resources Association (ELRA).
  2. 2.R Harald Baayen. 2009. Corpus linguistics in morphology: Morphological productivity. Corpus Linguistics. An International Handbook, pages 900–919.
  3. 3.Marco Baroni. 2022. On the proper role of linguistically oriented deep net analysis in linguistic theorising. In Algebraic Structures in Natural Language, pages 1–16. CRC Press.
  4. 4.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.
  5. 5.Joan Bybee. 1995. Regular morphology and the lexicon. Language and Cognitive Processes, 10(5):425–455.
  6. 6.N. Chomsky. 1957. Syntactic Structures. The Hague: Mouton.
  7. 7.N. Chomsky. 1965. Aspects of the Theory of Syntax. MIT Press, Cambridge, MA.
  8. 8.N. Chomsky. 1986. Knowledge of language: Its nature, origin, and use. Praeger Publishers.
  9. 9.Noam Chomsky, Ian Roberts, and Jeffrey Watumull. 2023. Noam Chomsky: The False Promise of ChatGPT. The New York Times.
  10. 10.Mary Dalrymple and Tracy Holloway King. 2019. An amazing four doctoral dissertations. Argumentum, 15(2019). Publisher: Debreceni Egyetemi Kiado.
  11. 11.Ronen Eldan and Yuanzhi Li. 2023. TinyStories: How Small Can Language Models Be and Still Speak Coherent English? arXiv:2305.07759.
  12. 12.Richard Futrell, Ethan Wilcox, Takashi Morita, Peng Qian, Miguel Ballesteros, and Roger Levy. 2019. Neural language models as psycholinguistic subjects: Representations of syntactic state. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 32–42, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Martin Gerlach and Francesc Font-Clos. 2020. A standardized Project Gutenberg corpus for statistical analysis of natural language and quantitative linguistics. Entropy, 22(1):126.
  14. 14.Adele E Goldberg. 1995. Constructions: A Construction Grammar Approach to Argument Structure. University of Chicago Press.
  15. 15.Adele E Goldberg. 2005. Constructions at Work: The Nature of Generalization in Language. Oxford University Press.
  16. 16.Adele E Goldberg. 2019. Explain me this: Creativity, competition, and the partial productivity of constructions. Princeton University Press.
  17. 17.Kenneth Heafield. 2011. KenLM: Faster and smaller language model queries. In Proceedings of the Sixth Workshop on Statistical Machine Translation, pages 187–197, Edinburgh, Scotland. Association for Computational Linguistics.
  18. 18.Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2016. The Goldilocks Principle: Reading Children’s Books with Explicit Memory Representations. In 4th International Conference on Learning Representations, ICLR 2016.
  19. 19.Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020. spaCy: Industrial-strength natural language processing in python.
  20. 20.Philip A. Huebner, Elior Sulem, Fisher Cynthia, and Dan Roth. 2021. BabyBERTa: Learning more grammar with small-scale child-directed language. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 624–646, Online. Association for Computational Linguistics.
  21. 21.Jaap Jumelet, Milica Denic, Jakub Szymanik, Dieuwke Hupkes, and Shane Steinert-Threlkeld. 2021. Language models use monotonicity to assess NPI licensing. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4958–4969, Online. Association for Computational Linguistics.
  22. 22.Julie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald, and Christopher Potts. 2024. Mission: Impossible language models. arXiv:2401.06416.
  23. 23.Richard S Kayne. 2007. On the syntax of quantity in English. Linguistic theory and South Asian languages: Essays in honour of Ka Jayaseelan, 102:73.
  24. 24.Caitlin Keenan. 2013. “A pleasant three days in Philadelphia”: Arguments for a pseudopartitive analysis. University of Pennsylvania Working Papers in Linguistics, 19(1):11.
  25. 25.Najoung Kim, Tal Linzen, and Paul Smolensky. 2022. Uncontrolled Lexical Exposure Leads to Overestimation of Compositional Generalization in Pretrained Models. arXiv:2212.10769.
  26. 26.Jey Han Lau, Alexander Clark, and Shalom Lappin. 2017. Grammaticality, acceptability, and probability: A probabilistic view of linguistic knowledge. Cognitive Science, 41(5):1202–1241.
  27. 27.Cara Su-Yi Leong and Tal Linzen. 2023. Language models can learn exceptions to syntactic rules. In Proceedings of the Society for Computation in Linguistics 2023, pages 133–144, Amherst, MA. Association for Computational Linguistics.
  28. 28.Cara Su-Yi Leong and Tal Linzen. 2024. Testing learning hypotheses using neural networks by manipulating learning data. arXiv:2407.04593.
  29. 29.Bai Li, Zining Zhu, Guillaume Thomas, Frank Rudzicz, and Yang Xu. 2022. Neural reality of argument structure constructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7410–7423, Dublin, Ireland. Association for Computational Linguistics.
  30. 30.Tal Linzen. 2020. How can we accelerate progress towards human-like linguistic generalization? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5210–5217, Online. Association for Computational Linguistics.
  31. 31.Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016. Assessing the ability of LSTMs to learn syntax-sensitive dependencies. Transactions of the Association for Computational Linguistics, 4:521–535.
  32. 32.Pierre Lison and Jörg Tiedemann. 2016. OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 923–929, Portoro!, Slovenia. European Language Resources Association (ELRA).
  33. 33.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692.
  34. 34.Kaiji Lu, Piotr Mardziel, Fangjing Wu, Preetam Amancharla, and Anupam Datta. 2020. Gender bias in neural natural language processing. Logic, Language, and Security: Essays Dedicated to Andre Scedrov on the Occasion of His 65th Birthday, pages 189–202.
  35. 35.B. MacWhinney. 2000. The CHILDES project: Tools for analyzing talk. Lawrence Erlbaum Hillsdale, New Jersey.
  36. 36.Kyle Mahowald. 2023. A discerning several thousand judgments: GPT-3 rates the article + adjective + numeral + noun construction. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 265–273, Dubrovnik, Croatia. Association for Computational Linguistics.
  37. 37.Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. 2024. Dissociating language and thought in large language models. Trends in Cognitive Sciences.
  38. 38.Rowan Hall Maudslay, Hila Gonen, Ryan Cotterell, and Simone Teufel. 2019. It’s all in the name: Mitigating gender bias with name-based counterfactual data substitution. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5267–5275, Hong Kong, China. Association for Computational Linguistics.
  39. 39.R. Thomas McCoy, Paul Smolensky, Tal Linzen, Jianfeng Gao, and Asli Celikyilmaz. 2023. How much do language models copy from their training data? evaluating linguistic novelty in text generation using RAVEN. Transactions of the Association for Computational Linguistics, 11:652–670.
  40. 40.Kanishka Misra. 2022. minicons: Enabling flexible behavioral and representational analyses of transformer language models. arXiv:2203.13112.
  41. 41.Kanishka Misra and Najoung Kim. 2023. Abstraction via exemplars? A representational case study on lexical category inference in BERT. In BUCLD 48: Proceedings of the 48th annual Boston University Conference on Language Development, Boston, USA.
  42. 42.Timothy J O’Donnell. 2015. Productivity and reuse in language: A theory of linguistic computation and storage. MIT Press.
  43. 43.Daniel N Osherson, Edward E Smith, Ormond Wilkie, Alejandro Lopez, and Eldar Shafir. 1990. Category-based Induction. Psychological Review, 97(2):185.
  44. 44.Abhinav Patil, Jaap Jumelet, Yu Ying Chiu, Andy Lapastora, Peter Shen, Lexie Wang, Clevis Willrich, and Shane Steinert-Threlkeld. 2024. Filtered Corpus Training (FiCT) Shows that Language Models can Generalize from Indirect Evidence. arXiv:2405.15750.
  45. 45.Adam Pauls and Dan Klein. 2012. Large-scale syntactic language modeling with treelets. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 959–968, Jeju Island, Korea. Association for Computational Linguistics.
  46. 46.Lisa Pearl. 2022. Poverty of the stimulus without tears. Language Learning and Development, 18(4):415–454.
  47. 47.Christopher Potts. 2023. Characterizing English Preposing in PP constructions. Ms., Stanford University.
  48. 48.Jakob Prange, Nathan Schneider, and Vivek Srikumar. 2021. Supertagging the long tail with tree-structured decoding of complex categories. Transactions of the Association for Computational Linguistics, 9:243–260.
  49. 49.Geoffrey K Pullum. 2017. Theory, data, and the epistemology of syntax. In Grammatische Variation. Empirische Zugänge und theoretische Modellierung, pages 283–298. de Gruyter.
  50. 50.Geoffrey K Pullum and Barbara C Scholz. 2002. Empirical assessment of stimulus poverty arguments. The Linguistic Review, 19(1-2):9–50.
  51. 51.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI.
  52. 52.Roger Schwarzschild. 2011. Stubborn Distributivity, Multiparticipant Nouns and the Count/Mass Distinction. In Proceedings of NELS, volume 39, pages 661–678. Graduate Linguistics Students Association, University of Massachusetts. Issue: 2.
  53. 53.Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021. Masked language modeling and the distributional hypothesis: Order word matters pre-training for little. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2888–2913, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  54. 54.Stephanie Solt. 2007. Two types of modified cardinals. In International Conference on Adjectives. Lille.
  55. 55.Andreas Stolcke, Klaus Ries, Noah Coccaro, Elizabeth Shriberg, Rebecca Bates, Daniel Jurafsky, Paul Taylor, Rachel Martin, Carol Van Ess-Dykema, and Marie Meteer. 2000. Dialogue act modeling for automatic tagging and recognition of conversational speech. Computational Linguistics, 26(3):339–374.
  56. 56.Laura Suttle and Adele E Goldberg. 2011. The partial productivity of constructions as induction. Linguistics, 49(6):1237–1269.
  57. 57.Harish Tayyar Madabushi, Laurence Romain, Dagmar Divjak, and Petar Milin. 2020. CxGBERT: BERT meets construction grammar. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4020–4032, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  58. 58.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288.
  59. 59.Yu-Hsiang Tseng, Cing-Fang Shih, Pin-Er Chen, Hsin-Yu Chou, Mao-Chang Ku, and Shu-Kai Hsieh. 2022. CxLM: A construction and context-aware language model. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 6361–6369, Marseille, France. European Language Resources Association.
  60. 60.Tim Veenboer and Jelke Bloem. 2023. Using collostructional analysis to evaluate BERT’s representation of linguistic constructions. In Findings of the Association for Computational Linguistics: ACL 2023, pages 12937–12951, Toronto, Canada. Association for Computational Linguistics.
  61. 61.Alex Warstadt. 2022. Artificial Neural Networks as Models of Human Language Acquisition. New York University.
  62. 62.Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Juan Ciro, Rafael Mosquera, Bhargavi Paranjabe, Adina Williams, Tal Linzen, and Ryan Cotterell. 2023. Findings of the BabyLM challenge: Sample-efficient pretraining on developmentally plausible corpora. In Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning, pages 1–34, Singapore. Association for Computational Linguistics.
  63. 63.Lucas Weber, Jaap Jumelet, Elia Bruni, and Dieuwke Hupkes. 2021. Language modelling as a multi-task problem. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2049–2060, Online. Association for Computational Linguistics.
  64. 64.Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick. 2021. Frequency effects on syntactic rule learning in transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 932–948, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  65. 65.Leonie Weissweiler, Valentin Hofmann, Abdullatif Köksal, and Hinrich Schütze. 2022. The better your syntax, the better your semantics? probing pretrained language models for the English comparative correlative. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10859–10882, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  66. 66.Leonie Weissweiler, Abdullatif Köksal, and Hinrich Schütze. 2024. Hybrid human-LLM corpus construction and LLM evaluation for rare linguistic phenomena. arXiv:2403.06965.
  67. 67.Ethan Wilcox, Roger Levy, Takashi Morita, and Richard Futrell. 2018. What do RNN language models learn about filler–gap dependencies? In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 211–221, Brussels, Belgium. Association for Computational Linguistics.
  68. 68.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  69. 69.Fei Xu and Joshua B Tenenbaum. 2007. Word learning as Bayesian inference. Psychological Review, 114(2):245.
  70. 70.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. OPT: Open Pre-trained Transformer Language Models. arXiv:2205.01068.

Citation

MLA
Misra, K., and K. Mahowald. “Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 913–29, https://doi.org/10.18653/v1/2024.emnlp-main.53.
APA
Misra, K., & Mahowald, K. (2024). Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 913–929. https://doi.org/10.18653/v1/2024.emnlp-main.53
Chicago
Misra, K., and K. Mahowald. 2024. “Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 913–29. https://doi.org/10.18653/v1/2024.emnlp-main.53.
Harvard
Misra, K. and Mahowald, K. (2024) “Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 913–929. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.53.
Vancouver
1. Misra K, Mahowald K (2024) Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 913–929

BibTeX

@inproceedings{misra-mahowald-2024-language,
    title = "Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing {AANN}s",
    author = "Misra, Kanishka  and
      Mahowald, Kyle",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.53/",
    doi = "10.18653/v1/2024.emnlp-main.53",
    pages = "913--929"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/