Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing

Linlu QiuPeter ShawPanupong PasupatTianze ShiJonathan HerzigEmily PitlerFei ShaKristina Toutanova

article2022EMNLP61 citations

Reveals that increasing model size alone fails to resolve out-of-distribution compositional generalization failures during standard fine-tuning for semantic parsing, while establishing prompt tuning as a more effective adaptation strategy across language models scaled up to 540 billion parameters.

Listen

Modern language models have achieved significant performance gains across many tasks by scaling up model size. However, deploying these models in real-world systems requires compositional generalization—the ability to correctly process novel combinations of previously observed concepts. In semantic parsing, which translates natural language into formal logical representations, models frequently encounter new compositions not present in training data. Whether simply increasing model size resolves this fundamental bottleneck across different deployment techniques has remained an open and urgent question for AI practitioners.

The article systematically evaluates how model scale impacts compositional generalization in semantic parsing across three distinct adaptation methods: full fine-tuning, prompt tuning, and in-context learning. The study tests encoder-decoder models up to 11 billion parameters and decoder-only models up to 540 billion parameters across four established semantic parsing benchmarks containing both synthetic and real-world natural language queries.

The findings show that scaling models does not uniformly resolve compositional generalization challenges. First, fully fine-tuning language models produces mostly flat or negative scaling curves, meaning larger fine-tuned models frequently perform no better—or even worse—than smaller ones. Second, while in-context learning demonstrates positive gains from scale, its absolute accuracy remains poor; the 540-billion-parameter model using in-context learning is generally outperformed by far smaller fine-tuned models. Third, prompt tuning exhibits positive scaling curves and frequently outperforms standard fine-tuning at larger scales. Finally, detailed error analysis reveals divergent trends: while larger models generate fewer output syntax errors, large fine-tuned models are significantly more prone to memorizing training distributions and failing to recombine known elements.

These results demonstrate that increasing parameter scale alone is an expensive and insufficient strategy for solving compositional generalization when relying on standard fine-tuning. Because full parameter tuning encourages large models to overfit to shallow statistical patterns and prior training outputs, engineering teams risk wasting computational resources without improving generalization reliability. Conversely, parameter-efficient methods like prompt tuning better preserve core representations while benefiting from increased model scale.

Organizations developing semantic parsing systems should reconsider relying purely on larger fine-tuned models to solve compositional failures. Decision-makers should prioritize parameter-efficient adaptation methods, such as prompt tuning, and implement constrained decoding mechanisms to eliminate persistent syntax errors. System designers using in-context learning should focus on building advanced example-retrieval pipelines that maximize structural diversity rather than relying solely on surface-level text similarity.

These conclusions are bounded by specific experimental conditions, including unconstrained decoding, fixed prompting formats, and single-run evaluations on selected model families. Nonetheless, the evidence strongly supports that structural inductive biases, improved retrieval, and parameter-efficient tuning are required alongside model scale to achieve robust compositional understanding.

arXiv: 2205.12253
Cover for Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing

Abstract

Despite their strong performance on many tasks, pre-trained language models have been shown to struggle on out-of-distribution compositional generalization. Meanwhile, recent work has shown considerable improvements on many NLP tasks from model scaling. Can scaling up model size also improve compositional generalization in semantic parsing? We evaluate encoder-decoder models up to 11B parameters and decoder-only models up to 540B parameters, and compare model scaling curves for three different methods for applying a pre-trained language model to a new task: fine-tuning all parameters, prompt tuning, and in-context learning. We observe that fine-tuning generally has flat or negative scaling curves on out-of-distribution compositional generalization in semantic parsing evaluations. In-context learning has positive scaling curves, but is generally outperformed by much smaller fine-tuned models. Prompt-tuning can outperform fine-tuning, suggesting further potential improvements from scaling as it exhibits a more positive scaling curve. Additionally, we identify several error trends that vary with model scale. For example, larger models are generally better at modeling the syntax of the output space, but are also more prone to certain types of overfitting. Overall, our study highlights limitations of current techniques for effectively leveraging model scale for compositional generalization, while our analysis also suggests promising directions for future work.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 3 Experimental Setup
  • 3.1 Datasets
  • 3.2 Models
  • 3.3 Retrievers
  • 4 Results & Analysis
  • 4.1 Main Results
  • 4.2 Error Analysis
  • 4.3 Task Analysis
  • 4.3.1 Distribution Shift
  • 4.3.2 Output Space
  • 4.4 Retriever Analysis
  • 5 Conclusion
  • 6 Limitations
  • Acknowledgements
  • References
  • Appendix
  • A Dataset Details
  • A.1 Dataset Sizes
  • A.2 Intermediate Representation
  • B Experiment Details
  • B.1 Experimental Setup
  • B.2 Number of Exemplars
  • C Additional Results
  • C.1 Additional Retriever Analysis
  • C.2 Results on Individual Splits
  • C.3 Example Prediction Errors

Knowls

  1. Knowl 1 — Experimental design for measuring scale and adaptation

    experimental setup

    The study measures exact-match accuracy for semantic parsing, mapping natural-language inputs to structured meaning representations, across four datasets: COGS, CFQ, GeoQuery, and SMCalFlow-CS. It compares five encoder-decoder T5 models from 60M to 11B parameters and three decoder-only PaLM models from 8B to 540B parameters. The adaptation methods span full parameter updates (fine-tuning), a learned prompt with the pretrained model frozen (prompt tuning), and no parameter updates with retrieved demonstrations at inference time (in-context learning, or ICL). T5 is used for fine-tuning and prompt tuning; PaLM is used for fine-tuning and ICL. Evaluation includes in-distribution and compositional splits, including novel linguistic structures, compound/template shifts, sequence-length splits, and cross-domain combinations. COGS and CFQ are evaluated on sampled test subsets; CFQ uses 1,000 test examples per split, while SMCalFlow-CS cross-domain evaluation has 8-, 16-, or 32-example training conditions.

  2. Knowl 2 — Full fine-tuning usually does not turn scale into compositional gains

    empirical result

    Across the compositional semantic-parsing evaluations, increasing model size while fine-tuning all parameters generally produces flat or declining accuracy curves; CFQ is the main dataset-level exception, with some splits showing gains. This pattern appears for both T5 and PaLM, so simply scaling a fully fine-tuned pretrained model is not a reliable way to improve compositional generalization. The flat curves are not explained only by a lack of room to improve: on a downsampled, less-saturated CFQ in-distribution split, T5 fine-tuning still has a negative scaling trend. The paper does not establish a single cause for the weak scaling.

  3. Knowl 3 — Prompt tuning has more favorable scaling than full fine-tuning

    empirical result

    For T5, prompt tuning freezes the pretrained model and learns a prompt of 100 tunable parameters; its compositional-generalization scaling curves are generally more positive than those of full fine-tuning, and prompt tuning sometimes exceeds fine-tuning at the same model size. For example, on GeoQuery TMCD1, T5-11B reaches 83.3% exact match with prompt tuning versus 62.4% with fine-tuning. Prompt tuning also closes the fine-tuning gap on in-distribution evaluations as model size grows. These findings indicate that the adaptation method changes how effectively model scale translates into semantic-parsing accuracy; they do not show that prompt tuning wins on every dataset or split.

  4. Knowl 4 — In-context learning improves with scale but depends strongly on retrieval

    empirical result

    PaLM in-context learning (ICL) shows positive scaling trends, but on most compositional splits even the largest ICL model performs below substantially smaller fine-tuned models. ICL performance also varies markedly with how demonstrations are selected. Non-oracle retrievers use only the input query: BM25 retrieves by lexical similarity, while a BERT-base retriever uses cosine similarity of query [CLS] embeddings. Oracle retrievers additionally use the gold output: one uses BM25 similarity to that output, and another uses Jaccard overlap of output compounds (parent-child symbol pairs), with exemplars required to contain the target's component symbols when available. On the PaLM-540B development set, the target-overlap oracle reaches 43.5% on SMCalFlow-CS cross-domain, compared with 9.5% for input-based BM25 and 1.4% for BERT retrieval; on CFQ MCD, the corresponding scores are 21.7%, 8.0%, and 7.8%. These oracle results are not directly comparable to ordinary ICL because the retrievers use gold outputs. Adding exemplars helps until performance plateaus; the plateau generally occurs with fewer examples for oracle retrieval than for non-oracle retrieval. ICL prompts use as many retrieved examples as fit within 1,920 tokens.

  5. Knowl 5 — Larger models generally produce fewer syntactically malformed outputs

    empirical result

    The study estimates output syntax errors by measuring the fraction of predictions with unbalanced parentheses during unconstrained greedy decoding. Across most GeoQuery and SMCalFlow-CS compositional splits and adaptation methods, this fraction tends to decrease as model size increases, indicating improved modeling of output syntax. Nevertheless, malformed outputs remain common for some conditions, particularly SMCalFlow-CS. The measure is only a proxy for syntax validity, and the results suggest that scaling alone does not eliminate the potential value of constrained decoding.

  6. Knowl 6 — Fine-tuned larger models can fail more often to recombine seen components

    empirical result

    Two measures track failures to compose previously seen output elements: on GeoQuery compositional splits, the fraction of incorrect predictions identical to an output in the training set; on SMCalFlow-CS cross-domain splits, the fraction of predictions containing functions from only one domain. With full fine-tuning, larger models are often more likely to show these errors, consistent with overfitting to the training distribution instead of recombining elements for the test composition. For example, larger fine-tuned models can represent a request to schedule a meeting with a team as a single-domain calendar subject string rather than combining calendar and organization-query functions. The study does not observe the same general trend for prompt tuning or ICL, where fewer or no pretrained-model parameters are updated.

  7. Knowl 7 — Models underproduce long outputs, with length trends varying by adaptation

    empirical result

    On GeoQuery and SMCalFlow-CS length splits, average prediction length is strongly associated with accuracy, and model predictions are generally shorter than the gold test outputs. This indicates difficulty generating sequences longer than those represented in training. As T5 grows under full fine-tuning, its average prediction length declines on GeoQuery's length split and stays roughly flat on SMCalFlow-CS's length split. By contrast, average prediction length grows with PaLM scale in ICL and with T5 scale under prompt tuning. Thus, the length behavior differs by adaptation method as well as model size.

  8. Knowl 8 — Incorrect outputs often favor frequent training-set trigrams

    empirical result

    To examine reliance on shallow output-distribution patterns, the study fits an add-one-smoothed count-based trigram language model to training outputs and compares its average token likelihood for predictions and gold outputs on GeoQuery template and TMCD splits. Incorrect predictions tend to be more likely under this trigram model than the correct outputs. On GeoQuery TMCD1, 72% of errors from fine-tuned T5-11B contain a predicted trigram that occurs more often in the training data than the corresponding correct trigram. Increasing scale does not reliably reduce this tendency under fine-tuning; prompt tuning shows a more favorable trend, and the tendency decreases as PaLM scale grows for ICL. This analysis supports, but does not prove, the hypothesis that some errors reflect overfitting to shallow output statistics.

  9. Knowl 9 — Limited headroom does not fully explain flat scaling curves

    empirical result

    The study estimates that dataset issues—ambiguous or inconsistent annotations and unseen output symbols—account for about 30% of errors on GeoQuery in-distribution and 70% on SMCalFlow-CS single-domain evaluation, compared with about 15% on GeoQuery compositional evaluation and 10% on SMCalFlow-CS cross-domain evaluation. These estimates come from manually inspecting up to 20 examples on which all models are incorrect. The larger share of such errors on in-distribution evaluations suggests less headroom there, but the remaining evidence does not make saturation a complete explanation for weak fine-tuning scaling. To probe less-saturated conditions, the study downsamples CFQ training data from about 95,743 examples to about 1,000, leaving only one or two exceptions in coverage of test symbols. Accuracy drops for every adaptation method, with a particularly large drop for fine-tuning; prompt tuning is less sensitive to this reduction in training data.

  10. Knowl 10 — Output representation affects accuracy more than scaling trends

    empirical result

    The study compares original and alternative output formats on two synthetic datasets: COGS lambda-calculus outputs with variables versus an equivalent variable-free representation, and CFQ's original SPARQL queries versus a reversible intermediate representation. The original, more complex formats yield lower accuracy, while the scaling trends remain similar across formats. Error patterns differ by adaptation method: more than 60% of fine-tuning and prompt-tuning errors on CFQ involve missing conjuncts, a problem largely mitigated by the intermediate representation's grouping of conjuncts. ICL instead often produces extra conjuncts absent from the gold output. On COGS, ICL also struggles with distinctions among entity types, with a 29-percentage-point absolute accuracy loss associated with these errors even when demonstrations include examples of each type.

Coverage note — The paper's remaining limitations—single runs per split, prompt tuning as the only parameter-efficient tuning method tested, and no compute- or energy-normalized scaling comparison—are omitted as standalone knowls because they qualify the study's scope rather than add a separate principal finding; per-split plots and individual error examples are likewise omitted because they instantiate the aggregate findings above.

References

  1. 1.Samira Abnar, Mostafa Dehghani, Behnam Neyshabur, and Hanie Sedghi. 2021. Exploring the limits of large scale pre-training. ArXiv preprint, abs/2110.02095.
  2. 2.Ekin Akyürek, Afra Feyza Akyürek, and Jacob Andreas. 2021. Learning to recombine and resample data for compositional generalization. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  3. 3.Jacob Andreas. 2020. Good-enough compositional data augmentation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7556–7566, Online. Association for Computational Linguistics.
  4. 4.Jacob Andreas, John Bufe, David Burkett, Charles Chen, Josh Clausman, Jean Crawford, Kate Crim, Jordan DeLoach, Leah Dorner, Jason Eisner, Hao Fang, Alan Guo, David Hall, Kristin Hayes, Kellie Hill, Diana Ho, Wendy Iwaszuk, Smriti Jha, Dan Klein, Jayant Krishnamurthy, Theo Lanman, Percy Liang, Christopher H. Lin, Ilya Lintsbakh, Andy McGovern, Aleksandr Nisnevich, Adam Pauls, Dmitrij Petters, Brent Read, Dan Roth, Subhro Roy, Jesse Rusak, Beth Short, Div Slomin, Ben Snyder, Stephon Striplin, Yu Su, Zachary Tellman, Sam Thomson, Andrei Vorobev, Izabela Witoszko, Jason Wolfe, Abby Wray, Yuchen Zhang, and Alexander Zotov. 2020. Task-oriented dialogue as dataflow synthesis. Transactions of the Association for Computational Linguistics, 8:556–571.
  5. 5.Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021. Program synthesis with large language models. ArXiv preprint, abs/2108.07732.
  6. 6.Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma. 2021. Explaining neural scaling laws. ArXiv preprint, abs/2102.06701.
  7. 7.Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al. 2018. Relational inductive biases, deep learning, and graph networks. ArXiv preprint, abs/1806.01261.
  8. 8.Jörg Bornschein, Francesco Visin, and Simon Osindero. 2020. Small data, big decisions: Model selection in the small-data regime. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 1035–1044. PMLR.
  9. 9.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  10. 10.Xinyun Chen, Chen Liang, Adams Wei Yu, Dawn Song, and Denny Zhou. 2020. Compositional generalization via neural-symbolic stack machines. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  11. 11.Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, and Tao Yu. 2022. Binding language models in symbolic languages. ArXiv preprint, abs/2210.02875.
  12. 12.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankar Anarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. ArXiv preprint, abs/2204.02311.
  13. 13.Henry Conklin, Bailin Wang, Kenny Smith, and Ivan Titov. 2021. Meta-learning to compositionally generalize. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3322–3335, Online. Association for Computational Linguistics.
  14. 14.Róbert Csordás, Kazuki Irie, and Juergen Schmidhuber. 2021. The devil is in the detail: Simple tricks improve systematic generalization of transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 619–634, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, and Dragomir Radev. 2018. Improving text-to-SQL evaluation methodology. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 351–360, Melbourne, Australia. Association for Computational Linguistics.
  17. 17.Daniel Furrer, Marc van Zee, Nathan Scales, and Nathanael Schärli. 2020. Compositional generalization in semantic parsing: Pre-training vs. specialized architectures. ArXiv preprint, abs/2007.08970.
  18. 18.Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart. 2019. Scaling description of generalization with number of parameters in deep learning. ArXiv preprint, abs/1901.01608.
  19. 19.Behrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna, Maxim Krikun, Xavier Garcia, Ciprian Chelba, and Colin Cherry. 2021. Scaling laws for neural machine translation. ArXiv preprint, abs/2109.07740.
  20. 20.Jonathan Gordon, David Lopez-Paz, Marco Baroni, and Diane Bouchacourt. 2020. Permutation equivariant models for compositional generalization in language. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  21. 21.Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2022. Towards a unified view of parameter-efficient transfer learning. In International Conference on Learning Representations.
  22. 22.Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B. Brown, Prafulla Dhariwal, Scott Gray, Chris Hallacy, Benjamin Mann, Alec Radford, Aditya Ramesh, Nick Ryder, Daniel M. Ziegler, John Schulman, Dario Amodei, and Sam McCandlish. 2020. Scaling laws for autoregressive generative modeling. ArXiv preprint, abs/2010.14701.
  23. 23.Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish. 2021. Scaling laws for transfer. CoRR.
  24. 24.Jonathan Herzig and Jonathan Berant. 2019. Don’t paraphrase, detect! rapid and effective data collection for semantic parsing. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3810–3820, Hong Kong, China. Association for Computational Linguistics.
  25. 25.Jonathan Herzig and Jonathan Berant. 2021. Span-based semantic parsing for compositional generalization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 908–921, Online. Association for Computational Linguistics.
  26. 26.Jonathan Herzig, Peter Shaw, Ming-Wei Chang, Kelvin Guu, Panupong Pasupat, and Yuan Zhang. 2021. Unlocking compositional generalization in pre-trained models using intermediate representations. ArXiv preprint, abs/2104.07478.
  27. 27.Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory F. Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. 2017. Deep learning scaling is predictable, empirically. ArXiv preprint, abs/1712.00409.
  28. 28.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022. Training compute-optimal large language models. CoRR, abs/2203.15556.
  29. 29.Maor Ivgi, Yair Carmon, and Jonathan Berant. 2022. Scaling laws under the microscope: Predicting transformer performance from small scale experiments. arXiv preprint arXiv:2202.06387.
  30. 30.Robin Jia and Percy Liang. 2016. Data recombination for neural semantic parsing. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12–22, Berlin, Germany. Association for Computational Linguistics.
  31. 31.Yichen Jiang and Mohit Bansal. 2021. Inducing transformer’s compositional generalization ability via auxiliary sequence prediction tasks. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6253–6265, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  32. 32.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray andkaplan2020scaling Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. ArXiv preprint, abs/2001.08361.
  33. 33.Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, Dmitry Tsarkov, Xiao Wang, Marc van Zee, and Olivier Bousquet. 2020. Measuring compositional generalization: A comprehensive method on realistic data. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  34. 34.Najoung Kim and Tal Linzen. 2020. COGS: A compositional generalization challenge based on semantic interpretation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9087–9105, Online. Association for Computational Linguistics.
  35. 35.Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. 2022. Fine-tuning can distort pretrained features and underperform out-of-distribution. ArXiv preprint, abs/2202.10054.
  36. 36.Brenden M. Lake. 2019. Compositional generalization through meta sequence-to-sequence learning. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 9788–9798.
  37. 37.Brenden M. Lake and Marco Baroni. 2018. Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pages 2879–2888. PMLR.
  38. 38.Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. 2017. Building machines that learn and think like people. Behavioral and brain sciences, 40.
  39. 39.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  40. 40.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  41. 41.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  42. 42.Yuanpeng Li, Liang Zhao, Jianyu Wang, and Joel Hestness. 2019. Compositional generalization for primitive substitutions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4293–4302, Hong Kong, China. Association for Computational Linguistics.
  43. 43.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021a. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ArXiv preprint, abs/2107.13586.
  44. 44.Qian Liu, Shengnan An, Jian-Guang Lou, Bei Chen, Zeqi Lin, Yan Gao, Bin Zhou, Nanning Zheng, and Dongmei Zhang. 2020. Compositional generalization by learning analytical expressions. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  45. 45.Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2021b. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. ArXiv preprint, abs/2110.07602.
  46. 46.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. CoRR.
  47. 47.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. ArXiv preprint, abs/2104.08786.
  48. 48.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? ArXiv preprint, abs/2202.12837.
  49. 49.Benjamin Newman, John Hewitt, Percy Liang, and Christopher D. Manning. 2020. The EOS decision and length extrapolation. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 276–291, Online. Association for Computational Linguistics.
  50. 50.Maxwell I. Nye, Armando Solar-Lezama, Josh Tenenbaum, and Brenden M. Lake. 2020. Learning compositional rules via neural program synthesis. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  51. 51.Santiago Ontanón, Joshua Ainslie, Vaclav Cvicek, and Zachary Fisher. 2021. Making transformers solve compositional tasks. ArXiv preprint, abs/2108.04378.
  52. 52.Inbar Oren, Jonathan Herzig, and Jonathan Berant. 2021. Finding needles in a haystack: Sampling structurally-diverse training sets from synthetic data for compositional generalization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10793–10809, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  53. 53.Inbar Oren, Jonathan Herzig, Nitish Gupta, Matt Gardner, and Jonathan Berant. 2020. Improving compositional generalization in semantic parsing. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2482–2495, Online. Association for Computational Linguistics.
  54. 54.Emmanouil Antonios Platanios, Adam Pauls, Subhro Roy, Yuchen Zhang, Alexander Kyte, Alan Guo, Sam Thomson, Jayant Krishnamurthy, Jason Wolfe, Jacob Andreas, and Dan Klein. 2021. Value-agnostic conversational semantic parsing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3666–3681, Online. Association for Computational Linguistics.
  55. 55.Linlu Qiu, Peter Shaw, Panupong Pasupat, Pawel Nowak, Tal Linzen, Fei Sha, and Kristina Toutanova. 2022. Improving compositional generalization with latent structure and data augmentation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4341–4362, Seattle, United States. Association for Computational Linguistics.
  56. 56.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training gopher. CoRR.
  57. 57.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  58. 58.Nitarshan Rajkumar, Raymond Li, and Dzmitry Bahdanau. 2022. Evaluating the text-to-sql capabilities of large language models. ArXiv preprint, abs/2204.00498.
  59. 59.Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. In CHI ’21: CHI Conference on Human Factors in Computing Systems, Virtual Event / Yokohama Japan, May 8-13, 2021, Extended Abstracts, pages 314:1–314:7. ACM.
  60. 60.Stephen E. Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: BM25 and beyond. Found. Trends Inf. Retr., 3(4):333–389.
  61. 61.Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit. 2020. A constructive prediction of the generalization error across scales. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  62. 62.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2021. Learning to retrieve prompts for in-context learning. ArXiv preprint, abs/2112.08633.
  63. 63.Luana Ruiz, Joshua Ainslie, and Santiago Ontañón. 2021. Iterative decoding for compositional generalization in transformers. ArXiv preprint, abs/2110.04169.
  64. 64.Jake Russin, Jason Jo, Randall C O’Reilly, and Yoshua Bengio. 2019. Compositional generalization in a deep seq2seq model by separating syntax and semantics. ArXiv preprint, abs/1904.09708.
  65. 65.Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. 2021. PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9895–9901, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  66. 66.Nathan Schucher, Siva Reddy, and Harm de Vries. 2022. The power of prompt tuning for low-resource semantic parsing. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 148–156, Dublin, Ireland. Association for Computational Linguistics.
  67. 67.Peter Shaw, Ming-Wei Chang, Panupong Pasupat, and Kristina Toutanova. 2021. Compositional generalization and natural language variation: Can a semantic parsing approach handle both? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 922–938, Online. Association for Computational Linguistics.
  68. 68.Richard Shin and Benjamin Van Durme. 2021. Few-shot semantic parsing with language models trained on code. ArXiv preprint, abs/2112.08696.
  69. 69.Richard Shin, Christopher Lin, Sam Thomson, Charles Chen, Subhro Roy, Emmanouil Antonios Platanios, Adam Pauls, Dan Klein, Jason Eisner, and Benjamin Van Durme. 2021. Constrained language models yield few-shot semantic parsers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7699–7715, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  70. 70.Lappoon R Tang and Raymond J Mooney. 2001. Using multiple clause constructors in inductive logic programming for semantic parsing. In European Conference on Machine Learning, pages 466–477. Springer.
  71. 71.Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler. 2021. Scale efficiently: Insights from pre-training and fine-tuning transformers. ArXiv preprint, abs/2109.10686.
  72. 72.Dmitry Tsarkov, Tibor Tihon, Nathan Scales, Nikola Momchev, Danila Sinopalnikov, and Nathanael Schärli. 2021. *-cfq: Analyzing the scalability of machine learning on a compositional task. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pages 9949–9957. AAAI Press.
  73. 73.Bailan Wang, Mirella Lapata, and Ivan Titov. 2021. Structured reordering for modeling latent alignments in sequence transduction. Advances in Neural Information Processing Systems, 34.
  74. 74.Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel. 2022. What language model architecture and pretraining objective work best for zero-shot generalization? CoRR, abs/2204.05832.
  75. 75.Mitchell Wortsman, Gabriel Ilharco, Mike Li, Jong Wook Kim, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt. 2021. Robust fine-tuning of zero-shot models. ArXiv preprint, abs/2109.01903.
  76. 76.Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2022. Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models. ArXiv preprint, abs/2201.05966.
  77. 77.Jingfeng Yang, Le Zhang, and Diyi Yang. 2022. SUBS: subtree substitution for compositional semantic parsing. CoRR, abs/2205.01538.
  78. 78.Pengcheng Yin, Hao Fang, Graham Neubig, Adam Pauls, Emmanouil Antonios Platanios, Yu Su, Sam Thomson, and Jacob Andreas. 2021. Compositional generalization for neural semantic parsing via span-level supervised attention. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2810–2823, Online. Association for Computational Linguistics.
  79. 79.John M Zelle and Raymond J Mooney. 1996. Learning to parse database queries using inductive logic programming. In Proceedings of the thirteenth national conference on Artificial intelligence-Volume 2, pages 1050–1055.
  80. 80.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068.
  81. 81.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 12697–12706. PMLR.
  82. 82.Hao Zheng and Mirella Lapata. 2021. Compositional generalization via semantic tagging. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 1022–1032, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  83. 83.Wang Zhu, Peter Shaw, Tal Linzen, and Fei Sha. 2021. Learning to generalize compositionally by transferring across semantic parsing tasks. ArXiv preprint, abs/2111.05013.

Citation

MLA
Qiu, L., et al. “Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 9157–79, https://doi.org/10.18653/v1/2022.emnlp-main.624.
APA
Qiu, L., Shaw, P., Pasupat, P., Shi, T., Herzig, J., Pitler, E., Sha, F., & Toutanova, K. (2022). Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 9157–9179. https://doi.org/10.18653/v1/2022.emnlp-main.624
Chicago
Qiu, L., P. Shaw, P. Pasupat, et al. 2022. “Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 9157–79. https://doi.org/10.18653/v1/2022.emnlp-main.624.
Harvard
Qiu, L. et al. (2022) “Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 9157–9179. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.624.
Vancouver
1. Qiu L, Shaw P, Pasupat P, Shi T, Herzig J, Pitler E, Sha F, Toutanova K (2022) Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 9157–9179

BibTeX

@inproceedings{qiu-etal-2022-evaluating,
    title = "Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing",
    author = "Qiu, Linlu  and
      Shaw, Peter  and
      Pasupat, Panupong  and
      Shi, Tianze  and
      Herzig, Jonathan  and
      Pitler, Emily  and
      Sha, Fei  and
      Toutanova, Kristina",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.624/",
    doi = "10.18653/v1/2022.emnlp-main.624",
    pages = "9157--9179"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/