The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative

Leonie WeissweilerValentin HofmannAbdullatif KöksalHinrich Schütze

article2022EMNLP55 citations

Reveals a critical disconnect in pretrained language models by showing that while models like BERT and DeBERTa reliably identify the syntactic structure of comparative correlative constructions, they consistently fail to understand and apply their underlying semantic meaning.

Listen

Modern natural language processing frequently evaluates pretrained language models to determine whether their performance reflects true linguistic competence or merely superficial pattern matching. While existing evaluations predominantly rely on generative grammar rules, the article investigates whether language models align with Construction Grammar—a linguistic framework positing that grammar consists of learned pairings of form and meaning. Specifically, the article examines whether prominent models can both recognize the structure of the English comparative correlative (for example, "the more, the merrier") and apply the cause-and-effect meaning it conveys.

The article evaluates BERT, RoBERTa, and DeBERTa across two distinct experimental tracks: syntactic recognition and semantic application. To assess syntax, the authors trained simple classification probes using representations from both synthetic minimal sentence pairs generated via context-free grammars and diverse, naturally occurring sentences extracted from the C4 web corpus. To evaluate semantics, the authors designed a zero-shot masked token prediction task requiring models to infer outcomes based on comparative correlative premises (such as "The stronger you are, the faster you are"). They paired this with calibration techniques and bias controls to isolate the models' true semantic reasoning from confounding factors like vocabulary and recency biases.

The analysis revealed a profound divergence between syntactic recognition and semantic reasoning. First, all evaluated models successfully recognize the syntactic form of the comparative correlative, achieving over 80% classification accuracy from intermediate layers onward and near-perfect accuracy with DeBERTa on artificial datasets. Second, despite this structural mastery, no model demonstrated a functional understanding of the construction's semantic meaning, performing around the 50% chance baseline across zero-shot inference tasks. Third, model predictions were heavily distorted by superficial heuristics: vocabulary bias caused up to 99.66% of predictions to flip based solely on the specific adjectives used, and BERT displayed strong recency bias (flipping up to 30.44% of decisions). Finally, targeted probability calibration reduced some biases but failed to lift semantic reasoning performance meaningfully above chance.

These findings indicate that pretrained language models easily master complex syntactic patterns through standard pretraining but fail to internalize the relational semantics inherent to grammatical constructions. For technical leaders and decision-makers, this highlights a critical operational risk: models cannot be assumed to understand the causal logic of structured text simply because they process or generate grammatically sophisticated sentences. Relying on current language models for zero-shot logical reasoning or automated policy compliance involving comparative conditional statements carries significant error risk.

To address these deficiencies, development teams should not rely on surface fluency as a proxy for language understanding. The article suggests that resolving these semantic limitations may require moving beyond standard text pretraining toward training paradigms that incorporate grounded interaction, real-world observation, or novel model architectures designed to capture non-compositional form-meaning mappings. Further research should extend construction-based evaluations to additional linguistic patterns and multilingual settings to establish clearer benchmarks for genuine machine understanding.

Readers should note that the study's conclusions are bound to the English comparative correlative and masked language model architectures (BERT, RoBERTa, and DeBERTa). While indirect probing and prompting techniques carry inherent noise, the extensive calibration and dual-dataset methodology provide high confidence that current models suffer from genuine semantic deficits in construction-level reasoning.

arXiv: 2210.13181
Cover for The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative

Abstract

Construction Grammar (CxG) is a paradigm from cognitive linguistics emphasising the connection between syntax and semantics. Rather than rules that operate on lexical items, it posits constructions as the central building blocks of language, i.e., linguistic units of different granularity that combine syntax and semantics. As a first step towards assessing the compatibility of CxG with the syntactic and semantic knowledge demonstrated by state-of-the-art pretrained language models (PLMs), we present an investigation of their capability to classify and understand one of the most commonly studied constructions, the English comparative correlative (CC). We conduct experiments examining the classification accuracy of a syntactic probe on the one hand and the models’ behaviour in a semantic application task on the other, with BERT, RoBERTa, and DeBERTa as the example PLMs. Our results show that all three investigated PLMs are able to recognise the structure of the CC but fail to use its meaning. While human-like performance of PLMs on many NLP tasks has been alleged, this indicates that PLMs still suffer from substantial shortcomings in central domains of linguistic knowledge.

Table of Contents

  • 1 Introduction
  • 2 Construction Grammar
  • 2.1 Overview
  • 2.2 Construction Grammar and NLP
  • 2.3 The English Comparative Correlative
  • 3 Syntax
  • 3.1 Probing Methods
  • 3.1.1 Synthetic Data
  • 3.1.2 Corpus-based Minimal Pairs
  • 3.1.3 The Probe
  • 3.2 Probing Results
  • 3.2.1 Artificial Data
  • 3.2.2 Corpus Data
  • 4 Semantics
  • 4.1 Probing Methods
  • 4.1.1 Usage-based Testing
  • 4.1.2 Biases
  • 4.1.3 Calibration
  • 4.2 Results
  • 4.2.1 Problem Analysis
  • 5 Related Work
  • 5.1 Construction Grammar in NLP
  • 5.2 Probing
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Disparity Between Syntactic Recognition and Semantic Understanding of the English Comparative Correlative in PLMs

    empirical result

    Pretrained language models (PLMs) such as BERT, RoBERTa, and DeBERTa demonstrate strong capability in recognizing the syntactic structure of the English comparative correlative (CC)—a construction exemplified by phrases like "the more, the merrier"—achieving probing classification accuracies above 80% to nearly 100% across middle and higher layers. However, when tested in zero-shot masked token prediction tasks requiring deductive inference based on the CC's conditional meaning (e.g., inferring "Therefore, X is [faster] than Y" from "The stronger you are, the faster you are"), the models perform at or near chance level (~50%). This establishes a clear divide where syntactic form recognition in PLMs does not translate into functional semantic understanding of constructional meaning.

  2. Knowl 2 — CFG-Based Synthetic Minimal Pair Generation for Comparative Correlative Probing

    model/method

    To probe whether language models can recognize the syntactic structure of the comparative correlative (CC) without lexical confounds, synthetic minimal pairs are generated using context-free grammars (CFGs).

    Positive CC instances follow the structural pattern: "The [ADV-er] the [NUM] [NOUN] [VERB]"\text{"The [ADV-er] the [NUM] [NOUN] [VERB]"} (e.g., "The harder the two cats fight").

    Negative (non-CC) instances reorder the exact same lexical items into: "The [ADJ-er] [NUM] [VERB] the [NOUN]"\text{"The [ADJ-er] [NUM] [VERB] the [NOUN]"} (e.g., "The harder two fight the cats").

    In the negative pattern, the numeral shifts from a noun modifier to a subject noun head, the verb acts transitively rather than intransitively, and the comparative adverb functions as a nominal modifier. To introduce variable sentence lengths and syntactic complexity, the CFG inserts optional initial phrases (e.g., "Nowadays,", "It follows that"), inter-clause parentheticals (e.g., "and by the way, they mean that this is always true"), and adverbial/prepositional adjuncts (e.g., "under the sun", "during the morning"). Mutually disjoint vocabularies for adverbs, numerals, nouns, and verbs are used for training and test sets.

  3. Knowl 3 — Corpus-Based POS Pattern Filtering for Natural Comparative Correlative Pairs

    model/method

    To evaluate syntactic recognition across natural, diverse variations of the comparative correlative (CC), sentences are extracted from the C4 corpus matching the template: "The" followed by a comparative adjective or adverb (ending in "-er" or preceded by "more"), followed later in the sentence by a second identical comparative pattern. Extracted sentences are grouped by their sequence of Part-of-Speech (POS) tags, and POS patterns are annotated as true CC instances (positive) or non-CC instances (negative). To ensure out-of-distribution syntactic generalization during probing, the set of POS tag patterns in the training split is completely disjoint from the POS tag patterns in the evaluation split.

  4. Knowl 4 — Feature-Balanced Dataset Construction for Syntactic Probing

    experimental setup

    To prevent probing classifiers from exploiting superficial structural correlations (such as sentence length or clause distance), balanced datasets DfD_f are constructed for each target feature f∈{sentence length,CC start position,second clause start position,inter-clause distance}f \in \{\text{sentence length}, \text{CC start position}, \text{second clause start position}, \text{inter-clause distance}\}:

    Df=⋃v∈fv⋃l∗∈{positive,negative}S(D,v,l∗,n∗)D_f = \bigcup_{v \in f_v} \bigcup_{l^* \in \{\text{positive}, \text{negative}\}} S(D, v, l^*, n^*)

    where fvf_v is the set of feature values, and S(D,v,l∗,n∗)S(D, v, l^*, n^*) is a sampling function returning exactly n∗n^* instances from dataset DD with feature value vv and binary label l∗l^*.

    To assess out-of-length generalization, the training set DftrainD^{\text{train}}_f is restricted to the lowest quartile of the feature distribution:

    fvtrain=[vfmin⁡,vfmin⁡+14(vfmax⁡−vfmin⁡)]f_v^{\text{train}} = \left[v^{\min}_f, v^{\min}_f + \frac{1}{4}(v^{\max}_f - v^{\min}_f)\right]

    while the evaluation set DftestD^{\text{test}}_f spans the entire range [vfmin⁡,vfmax⁡][v^{\min}_f, v^{\max}_f]. A logistic regression classifier is trained on mean-pooled PLM sentence embeddings to classify instances as positive or negative CC constructions.

  5. Knowl 5 — Layer-Wise Syntactic Probing Accuracy across PLM Architectures

    empirical result

    Probing logistic regression classifiers trained on mean-pooled layer representations of BERT-Large (340M parameters), RoBERTa-Large (355M parameters), and DeBERTa-Large (1.5B parameters) reveal strong syntactic encoding of the comparative correlative:

    1. Performance Hierarchy: DeBERTa-Large achieves near-perfect classification accuracy (~100%) across nearly all layers on both synthetic and corpus datasets, consistently outperforming RoBERTa-Large, which in turn outperforms BERT-Large.
    2. Layer Emergence: All models achieve above 80% accuracy on natural corpus data from middle layers onwards, demonstrating that models distinguish CC from non-CC structures across unseen POS patterns.
    3. Feature Invariance: For synthetic minimal pairs, performance is stable across variations in sentence length, construction start position, and inter-clause distance, though BERT-Large and RoBERTa-Large exhibit slight degradation on longer sentences in early layers, which reverses in the final semantic layers.
  6. Knowl 6 — Zero-Shot Masked Deduction Task for Constructional Semantics

    model/method

    To test whether masked pretrained language models understand the semantics of the comparative correlative (CC), a zero-shot inference prompt is constructed from four clauses:

    1. Primary CC Rule: "The [ADJ1]-er you are, the [ADJ2]-er you are."\text{"The [ADJ1]-er you are, the [ADJ2]-er you are."}
    2. Contrastive Antonym CC Rule: "The [ANT1]-er you are, the [ANT2]-er you are."\text{"The [ANT1]-er you are, the [ANT2]-er you are."}
    3. Specific Premise: "NAME1 is [ADJ1]-er than NAME2."\text{"NAME1 is [ADJ1]-er than NAME2."}
    4. Target Deduction Sentence: "Therefore, NAME1 is [MASK] than NAME2."\text{"Therefore, NAME1 is [MASK] than NAME2."}

    Here, [ANT1][\text{ANT1}] and [ANT2][\text{ANT2}] are the exact semantic antonyms of [ADJ1][\text{ADJ1}] and [ADJ2][\text{ADJ2}] respectively. Under valid semantic comprehension of the CC, the model should assign a higher probability at the [MASK][\text{MASK}] token to the logically entailed comparative [ADJ2][\text{ADJ2}] than to the distractor [ANT2][\text{ANT2}] (P(ADJ2∣context)>P(ANT2∣context)P(\text{ADJ2} \mid \text{context}) > P(\text{ANT2} \mid \text{context})).

  7. Knowl 7 — Bias Dissection and Decision Flip Metric in Zero-Shot Prompting

    model/method

    In prompt-based zero-shot evaluation of masked language models on comparative correlatives, three potential confounding heuristics are isolated using modified prompt schemata:

    1. Recency Bias (S2S2): The order of the two CC rules is inverted such that the correct adjective ADJ2\text{ADJ2} appears in the second CC sentence, placing it physically closer to the [MASK][\text{MASK}] position than ANT2\text{ANT2}.
    2. Vocabulary / Lexical Identity Bias (S3S3): The positions of ADJ2\text{ADJ2} and ANT2\text{ANT2} in the premise rules are swapped while maintaining the token order, testing whether the model favors specific words regardless of context.
    3. Name Association Bias (S4S4): The names in the premise and target deduction are swapped (NAME1↔NAME2\text{NAME1} \leftrightarrow \text{NAME2}).

    To quantify model susceptibility to these biases, the decision flip metric is computed as the percentage of test instances for which the model's binary decision (arg⁡max⁡a∈{ADJ2,ANT2}P(a∣context)\arg\max_{a \in \{\text{ADJ2}, \text{ANT2}\}} P(a \mid \text{context})) inverts between the baseline prompt (S1S1) and the bias probe (S2,S3,S2, S3, or S4S4).

  8. Knowl 8 — Contextual Prior Calibration for Masked Zero-Shot Inference

    equation

    To correct for token frequency and prompt biases in zero-shot masked language model probing, calibrated prediction probabilities Pc(a∣Sb)P_c(a \mid S_b) for a target adjective a∈{ADJ2,ANT2}a \in \{\text{ADJ2}, \text{ANT2}\} in base sentence SbS_b are computed by dividing by the unconditioned prior over uninformative contexts:

    Pc(a∣Sb)=P(a∣Sb)15∑i=15P(a∣Ci)P_c(a \mid S_b) = \frac{P(a \mid S_b)}{\frac{1}{5}\sum_{i=1}^5 P(a \mid C_i)}

    where {C1,…,C5}\{C_1, \dots, C_5\} are five uninformative context samples generated under one of three calibration schemas:

    • Short Context Calibration (S5S5): Removing the CC rule sentences entirely, keeping only the scenario premise and masked conclusion.
    • Name Replacement Calibration (S6S6): Retaining the CC rules and premise but replacing the names in the conclusion with unrelated names (NAME3,NAME4\text{NAME3}, \text{NAME4}).
    • Adjective Replacement Calibration (S7S7): Retaining the CC rules and premise structure but replacing the premise adjective with an unrelated adjective (ADJ3\text{ADJ3}).
  9. Knowl 9 — Empirical Semantic Probing Accuracy and Bias Susceptibility across PLMs

    data/table

    Evaluation of 8 language model variants across 144,800 test instances reveals that models fail to perform semantic deduction for comparative correlatives, driven heavily by recency and vocabulary biases.

    Model S1 Acc (%) S2 Acc (%) S2 Flip (%) S3 Flip (%) S4 Flip (%)
    BERTBase\text{BERT}_{\text{Base}} 37.65 64.64 26.98 75.69 2.70
    BERTLarge\text{BERT}_{\text{Large}} 36.85 67.21 30.44 73.31 2.32
    RoBERTaBase\text{RoBERTa}_{\text{Base}} 61.60 52.84 9.91 76.18 2.76
    RoBERTaLarge\text{RoBERTa}_{\text{Large}} 55.71 68.00 14.33 79.47 4.33
    DeBERTaBase\text{DeBERTa}_{\text{Base}} 49.72 49.80 0.91 99.66 1.07
    DeBERTaLarge\text{DeBERTa}_{\text{Large}} 50.88 51.40 7.04 94.83 2.23
    DeBERTaXL\text{DeBERTa}_{\text{XL}} 47.73 49.33 5.46 89.28 2.51
    DeBERTaXXL\text{DeBERTa}_{\text{XXL}} 47.34 48.72 3.59 82.09 1.13

    In the baseline setting (S1S1), RoBERTa and DeBERTa variants score close to random chance (47.34%–61.60%), while BERT scores below chance (36.85%–37.65%) due to recency bias favoring the distractor ANT2\text{ANT2}. Swapping rule order (S2S2) raises BERT accuracy to 64.64%–67.21%, confirming strong recency bias (26.98%–30.44% decision flips). DeBERTa exhibits extreme vocabulary bias (S3S3 flip rate of 82.09%–99.66%), choosing tokens based on pretraining surface frequency rather than constructional entailment. Calibration methods (S5,S6,S7S5, S6, S7) fail to consistently overcome these biases, leaving calibrated accuracies near 50%.

  10. Knowl 10 — Methodological Scope and Evaluation Limitations of Comparative Correlative Probing

    limitation

    Probing pretrained language models (PLMs) for construction grammar (CxG) comprehension is subject to three core limitations:

    1. Indirect Probing: Probing evaluates internal linguistic representations only indirectly via downstream classification or zero-shot prediction tasks, which can be influenced by external surface artifacts and token frequency biases.
    2. Single Construction and Language Scope: The experimental analysis is restricted exclusively to the English comparative correlative (CC) and does not cover other grammatical constructions (e.g., argument structure constructions) or other natural languages.
    3. Lack of Interactive Grounding: Pretraining text corpora lack embodied interaction and pragmatic feedback, which may be necessary for neural models to acquire the functional causal semantics associated with complex grammatical constructions.

Coverage note — Detailed CFG production rules (Algorithms 1 and 2 in the appendix) were summarized into general generation principles rather than reproduced verbatim as code algorithms, as they represent specific vocabulary enumerations.

References

  1. 1.Anne Abeillé and Robert D Borsley. 2008. Comparative correlatives and parameters. Lingua, 118(8):1139–1157.
  2. 2.Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017. What do neural machine translation models learn about morphology? In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 861–872, Vancouver, Canada. Association for Computational Linguistics.
  3. 3.Giulia ML Bencini and Adele E Goldberg. 2000. The contribution of argument structure constructions to sentence meaning. Journal of Memory and Language, 43(4):640–651.
  4. 4.Leonard Bloomfield. 1933. Language. Holt, Rinehart & Winston, New York, NY.
  5. 5.Noam Chomsky. 1988. Generative grammar. Studies in English linguistics and literature.
  6. 6.Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2126–2136, Melbourne, Australia. Association for Computational Linguistics.
  7. 7.Peter W Culicover and Ray Jackendoff. 1999. The view from the periphery: The english comparative correlative. Linguistic inquiry, 30(4):543–571.
  8. 8.Dorottya Demszky, Devyani Sharma, Jonathan H. Clark, Vinodkumar Prabhakaran, and Jacob Eisenstein. 2021. Learning to recognize dialect features. In Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HTL) 2021.
  9. 9.Marcel Den Dikken. 2005. Comparative correlatives comparatively. Linguistic Inquiry, 36(4):497–532.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Jesse Dunietz, Lori Levin, and Jaime Carbonell. 2017. Automatically tagging constructions of causation and their slot-fillers. Transactions of the Association for Computational Linguistics, 5:117–133.
  12. 12.Jonathan Dunn. 2019. Frequency vs. association for constraint selection in usage-based construction grammar. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 117–128, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Charles J Fillmore. 1986. Varieties of conditional sentences. In Eastern States Conference on Linguistics, volume 3, pages 163–182.
  14. 14.Charles J. Fillmore. 1989. Grammatical construction: Theory and the familiar dichotomies. In Rainer Dietrich and Carl F. Graumann, editors, Language processing in social context, pages 17–38. North-Holland, Amsterdam.
  15. 15.Charles J. Fillmore, Paul Kay, and Mary C. O’Connor. 1988. Regularity and idiomaticity in grammatical constructions: The case of let alone. Language, 64(3):501–538.
  16. 16.Adele Goldberg. 1995. Constructions: A construction grammar approach to argument structure. University of Chicago Press, Chicago, IL.
  17. 17.Adele Goldberg. 2006. Constructions at work: The nature of generalization in language. Oxford University Press, Oxford, UK.
  18. 18.Adele E Goldberg. 2003. Constructions: A new theoretical approach to language. Trends in cognitive sciences, 7(5):219–224.
  19. 19.Yoav Goldberg. 2019. Assessing bert’s syntactic abilities. arXiv preprint arXiv:1901.05287.
  20. 20.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020. Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations.
  21. 21.Martin Hilpert. 2006. A synchronic perspective on the grammaticalization of Swedish future constructions. Nordic Journal of Linguistics, 29(2):151–173.
  22. 22.Thomas Hoffmann, Jakob Horsch, and Thomas Brunner. 2019. The more data, the better: A usage-based account of the english comparative correlative construction. Cognitive Linguistics, 30(1):1–36.
  23. 23.Thomas Hoffmann and Graeme Trousdale. 2013. The Oxford handbook of construction grammar. Oxford University Press.
  24. 24.Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn’t always right. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  25. 25.Matthew Honnibal and Ines Montani. 2018. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. To appear.
  26. 26.Eiichi Iwasaki and Andrew Radford. 2009. Comparative correlatives in english: A minimalist-cartographic analysis.
  27. 27.Matt A Johnson and Adele E Goldberg. 2013. Evidence for automatic accessing of constructional meaning: Jabberwocky sentences prime associated verbs. Language and Cognitive Processes, 28(10):1439–1452.
  28. 28.Nora Kassner and Hinrich Schütze. 2020. Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7811–7818, Online. Association for Computational Linguistics.
  29. 29.Paul Kay and Charles J. Fillmore. 1999. Grammatical constructions and linguistic generalizations: The What’s X doing Y? construction. Language, 75(1):1–33.
  30. 30.George Lakoff. 1987. Women, fire, and dangerous things: What categories reveal about the mind. University of Chicago Press, Chicago, IL.
  31. 31.Ronald W. Langacker. 1987. Foundations of cognitive grammar: Theoretical prerequisites. Stanford University Press, Stanford, CA.
  32. 32.Bai Li, Zining Zhu, Guillaume Thomas, Frank Rudzicz, and Yang Xu. 2022. Neural reality of argument structure constructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7410–7423, Dublin, Ireland. Association for Computational Linguistics.
  33. 33.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  34. 34.Tânia Marques and Katrien Beuls. 2016. Evaluation strategies for computational construction grammars. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 1137–1146, Osaka, Japan. The COLING 2016 Organizing Committee.
  35. 35.Rebecca Marvin and Tal Linzen. 2018. Targeted syntactic evaluation of language models. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1192–1202, Brussels, Belgium. Association for Computational Linguistics.
  36. 36.Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428–3448, Florence, Italy. Association for Computational Linguistics.
  37. 37.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The LAMBADA dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1525–1534, Berlin, Germany. Association for Computational Linguistics.
  38. 38.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, G. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830.
  39. 39.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HLT) 2018.
  40. 40.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  41. 41.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, et al. Language models are unsupervised multitask learners.
  42. 42.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  43. 43.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online. Association for Computational Linguistics.
  44. 44.Harish Tayyar Madabushi, Laurence Romain, Dagmar Divjak, and Petar Milin. 2020. CxGBERT: BERT meets construction grammar. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4020–4032, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  45. 45.Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019. What do you learn from context? probing for sentence structure in contextualized word representations. In International Conference on Learning Representations (ICLR) 7.
  46. 46.Yu-Hsiang Tseng, Cing-Fang Shih, Pin-Er Chen, Hsin-Yu Chou, Mao-Chang Ku, and Shu-Kai Hsieh. 2022. CxLM: A construction and context-aware language model. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 6361–6369, Marseille, France. European Language Resources Association.
  47. 47.Ivan Vulic, Edoardo M. Ponti, Robert Litschko, Goran Glavaš, and Anna Korhonen. 2020. Probing pretrained language models for lexical semantics. In Conference on Empirical Methods in Natural Language Processing (EMNLP) 2020.
  48. 48.Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohan­aney, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020. BLiMP: The benchmark of linguistic minimal pairs for English. Transactions of the Association for Computational Linguistics, 8:377–392.
  49. 49.Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick. 2021. Frequency effects on syntactic rule learning in transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 932–948, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  50. 50.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771.
  51. 51.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 12697–12706. PMLR.

Citation

MLA
Weissweiler, L., et al. “The Better Your Syntax, the Better Your Semantics? Probing Pretrained Language Models for the English Comparative Correlative”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 10859–82, https://doi.org/10.18653/v1/2022.emnlp-main.746.
APA
Weissweiler, L., Hofmann, V., Köksal, A., & Schütze, H. (2022). The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 10859–10882. https://doi.org/10.18653/v1/2022.emnlp-main.746
Chicago
Weissweiler, L., V. Hofmann, A. Köksal, and H. Schütze. 2022. “The Better Your Syntax, the Better Your Semantics? Probing Pretrained Language Models for the English Comparative Correlative”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 10859–82. https://doi.org/10.18653/v1/2022.emnlp-main.746.
Harvard
Weissweiler, L. et al. (2022) “The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10859–10882. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.746.
Vancouver
1. Weissweiler L, Hofmann V, Köksal A, Schütze H (2022) The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10859–10882

BibTeX

@inproceedings{weissweiler-etal-2022-better,
    title = "The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the {E}nglish Comparative Correlative",
    author = {Weissweiler, Leonie  and
      Hofmann, Valentin  and
      K{\"o}ksal, Abdullatif  and
      Sch{\"u}tze, Hinrich},
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.746/",
    doi = "10.18653/v1/2022.emnlp-main.746",
    pages = "10859--10882"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/