Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing
Linlu QiuPeter ShawPanupong PasupatTianze ShiJonathan HerzigEmily PitlerFei ShaKristina Toutanova
Reveals that increasing model size alone fails to resolve out-of-distribution compositional generalization failures during standard fine-tuning for semantic parsing, while establishing prompt tuning as a more effective adaptation strategy across language models scaled up to 540 billion parameters.
Modern language models have achieved significant performance gains across many tasks by scaling up model size. However, deploying these models in real-world systems requires compositional generalization—the ability to correctly process novel combinations of previously observed concepts. In semantic parsing, which translates natural language into formal logical representations, models frequently encounter new compositions not present in training data. Whether simply increasing model size resolves this fundamental bottleneck across different deployment techniques has remained an open and urgent question for AI practitioners.
The article systematically evaluates how model scale impacts compositional generalization in semantic parsing across three distinct adaptation methods: full fine-tuning, prompt tuning, and in-context learning. The study tests encoder-decoder models up to 11 billion parameters and decoder-only models up to 540 billion parameters across four established semantic parsing benchmarks containing both synthetic and real-world natural language queries.
The findings show that scaling models does not uniformly resolve compositional generalization challenges. First, fully fine-tuning language models produces mostly flat or negative scaling curves, meaning larger fine-tuned models frequently perform no better—or even worse—than smaller ones. Second, while in-context learning demonstrates positive gains from scale, its absolute accuracy remains poor; the 540-billion-parameter model using in-context learning is generally outperformed by far smaller fine-tuned models. Third, prompt tuning exhibits positive scaling curves and frequently outperforms standard fine-tuning at larger scales. Finally, detailed error analysis reveals divergent trends: while larger models generate fewer output syntax errors, large fine-tuned models are significantly more prone to memorizing training distributions and failing to recombine known elements.
These results demonstrate that increasing parameter scale alone is an expensive and insufficient strategy for solving compositional generalization when relying on standard fine-tuning. Because full parameter tuning encourages large models to overfit to shallow statistical patterns and prior training outputs, engineering teams risk wasting computational resources without improving generalization reliability. Conversely, parameter-efficient methods like prompt tuning better preserve core representations while benefiting from increased model scale.
Organizations developing semantic parsing systems should reconsider relying purely on larger fine-tuned models to solve compositional failures. Decision-makers should prioritize parameter-efficient adaptation methods, such as prompt tuning, and implement constrained decoding mechanisms to eliminate persistent syntax errors. System designers using in-context learning should focus on building advanced example-retrieval pipelines that maximize structural diversity rather than relying solely on surface-level text similarity.
These conclusions are bounded by specific experimental conditions, including unconstrained decoding, fixed prompting formats, and single-run evaluations on selected model families. Nonetheless, the evidence strongly supports that structural inductive biases, improved retrieval, and parameter-efficient tuning are required alongside model scale to achieve robust compositional understanding.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). Its compositionality-gap framing clarifies the central distinction between knowing familiar pieces and successfully combining them in novel ways.
- Paper: Exploring Length Generalization in Large Language Models, Cem Anil et al. (2022). Its finding that scaling and standard fine-tuning fail on length generalization provides useful prior evidence for interpreting the source’s tests of out-of-distribution composition.
- Paper: Scaling Laws for Neural Language Models, Jared Kaplan et al. (2020). Its scaling-law results establish the expectation that larger models improve performance, which the source tests against the specific challenge of compositional generalization.
- Paper: P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks, Xiao Liu et al. (2022). Its evaluation of prompt tuning across model scales supplies background for the source’s comparison of parameter-efficient prompting with full fine-tuning.
- Paper: Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning, Saibo Geng et al. (2023). It develops grammar-constrained decoding as a practical way to enforce valid structured outputs, extending the source’s proposed remedy for syntax errors in semantic parsing.
- Paper: Unveiling the Generalization Power of Fine-Tuned Large Language Models, Haoran Yang et al. (2024). It extends the source’s concern about fine-tuning-induced generalization failures by testing how those effects carry across domains and tasks.
