The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative
Leonie WeissweilerValentin HofmannAbdullatif KöksalHinrich Schütze
Reveals a critical disconnect in pretrained language models by showing that while models like BERT and DeBERTa reliably identify the syntactic structure of comparative correlative constructions, they consistently fail to understand and apply their underlying semantic meaning.
Modern natural language processing frequently evaluates pretrained language models to determine whether their performance reflects true linguistic competence or merely superficial pattern matching. While existing evaluations predominantly rely on generative grammar rules, the article investigates whether language models align with Construction Grammar—a linguistic framework positing that grammar consists of learned pairings of form and meaning. Specifically, the article examines whether prominent models can both recognize the structure of the English comparative correlative (for example, "the more, the merrier") and apply the cause-and-effect meaning it conveys.
The article evaluates BERT, RoBERTa, and DeBERTa across two distinct experimental tracks: syntactic recognition and semantic application. To assess syntax, the authors trained simple classification probes using representations from both synthetic minimal sentence pairs generated via context-free grammars and diverse, naturally occurring sentences extracted from the C4 web corpus. To evaluate semantics, the authors designed a zero-shot masked token prediction task requiring models to infer outcomes based on comparative correlative premises (such as "The stronger you are, the faster you are"). They paired this with calibration techniques and bias controls to isolate the models' true semantic reasoning from confounding factors like vocabulary and recency biases.
The analysis revealed a profound divergence between syntactic recognition and semantic reasoning. First, all evaluated models successfully recognize the syntactic form of the comparative correlative, achieving over 80% classification accuracy from intermediate layers onward and near-perfect accuracy with DeBERTa on artificial datasets. Second, despite this structural mastery, no model demonstrated a functional understanding of the construction's semantic meaning, performing around the 50% chance baseline across zero-shot inference tasks. Third, model predictions were heavily distorted by superficial heuristics: vocabulary bias caused up to 99.66% of predictions to flip based solely on the specific adjectives used, and BERT displayed strong recency bias (flipping up to 30.44% of decisions). Finally, targeted probability calibration reduced some biases but failed to lift semantic reasoning performance meaningfully above chance.
These findings indicate that pretrained language models easily master complex syntactic patterns through standard pretraining but fail to internalize the relational semantics inherent to grammatical constructions. For technical leaders and decision-makers, this highlights a critical operational risk: models cannot be assumed to understand the causal logic of structured text simply because they process or generate grammatically sophisticated sentences. Relying on current language models for zero-shot logical reasoning or automated policy compliance involving comparative conditional statements carries significant error risk.
To address these deficiencies, development teams should not rely on surface fluency as a proxy for language understanding. The article suggests that resolving these semantic limitations may require moving beyond standard text pretraining toward training paradigms that incorporate grounded interaction, real-world observation, or novel model architectures designed to capture non-compositional form-meaning mappings. Further research should extend construction-based evaluations to additional linguistic patterns and multilingual settings to establish clearer benchmarks for genuine machine understanding.
Readers should note that the study's conclusions are bound to the English comparative correlative and masked language model architectures (BERT, RoBERTa, and DeBERTa). While indirect probing and prompting techniques carry inherent noise, the extensive calibration and dual-dataset methodology provide high confidence that current models suffer from genuine semantic deficits in construction-level reasoning.
- Paper: A Structural Probe for Finding Syntax in Word Representations, John Hewitt et al. (2019). Introduces structural probing methods to test whether pretrained language models encode syntactic trees, providing the foundational probing framework used in the source paper.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). Establishes how pretrained representations sequentially encode syntax and semantics across layers, establishing the baseline framework of PLM linguistic representations examined by the source.
- Paper: What Does BERT Learn about the Structure of Language?, Ganesh Jawahar et al. (2019). Demonstrates layer-wise probing of syntactic and semantic structures in BERT, directly motivating the probing methodology for specific linguistic constructions.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). Shows that language models can leverage syntactic heuristics rather than genuine linguistic understanding, motivating the source's dissociation between syntactic probing and semantic task performance.
- Paper: Neural Network Acceptability Judgments, Alex Warstadt et al. (2018). Establishes benchmarks for evaluating whether neural language models possess true grammatical acceptability and linguistic competence.
- Paper: What Does BERT Look at? An Analysis of BERT’s Attention, Kevin Clark et al. (2019). Investigates how attention mechanisms in transformer models capture syntactic dependencies and grammatical functions.
- Paper: Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs, Kanishka Misra et al. (2024). Extends the investigation of constructional learning in language models by analyzing how models acquire rare grammatical constructions via structural generalization.
- Paper: When transformers learn "impossible" languages, what do they learn?, Ram Janarthan et al. (2026). Deepens the inquiry into model linguistic competence by assessing whether transformers acquire genuine grammatical sensitivity or superficial patterns when exposed to unnatural grammars.
