Exploring Length Generalization in Large Language Models

Cem AnilYuhuai WuAnders AndreassenAitor LewkowyczVedant MisraVinay V. RamaseshAmbrose SloneGuy Gur-AriEthan DyerBehnam Neyshabur

article2022NeurIPS227 citations

Reveals that while standard finetuning fails to extrapolate reasoning to longer problem instances regardless of model scale, combining in-context learning with scratchpad prompting substantially improves length generalization in large language models.

Listen

Real-world reasoning tasks—such as mathematical problem solving, program execution, and theorem proving—naturally increase in difficulty as problem length increases. In these domains, long training examples are exceedingly rare, making it critical for artificial intelligence systems to learn generalizable rules from short instances and extrapolate them to longer ones. The article evaluates the ability of transformer-based large language models to perform this type of out-of-distribution reasoning, termed length generalization, across model scales and training methodologies.

To systematically analyze this capability, the article investigates two controlled algorithmic benchmarks: a parity task (determining whether the count of ones in a bit-string is odd or even) and a Boolean variable assignment task (tracking sequential program execution). The authors conducted extensive empirical tests using decoder-only language models ranging from 244 million up to 128 billion parameters, comparing standard fine-tuning, intermediate step generation (scratchpad or chain-of-thought methods), and in-context few-shot prompting.

The investigation produced four central findings. First, standard fine-tuning consistently fails to generalize to longer problems regardless of model size; even when models achieve near 100% accuracy on training lengths, their performance rapidly degrades toward random guessing on longer instances. Second, transformer architectures exhibit a strong bias toward learning parallel counting shortcuts rather than sequential step-by-step algorithms, making in-distribution training loss an unreliable predictor of out-of-distribution success. Third, contrary to prior assumptions, fine-tuning models to produce step-by-step scratchpad solutions also fails to generalize to longer inputs due to attention failures across longer token contexts. Fourth, combining pre-trained models with few-shot scratchpad prompting dramatically improves length generalization without any parameter fine-tuning, successfully enabling models to map short step-by-step templates to instances up to five times longer than prompt demonstrations.

These findings indicate that scaling model size, data volume, or fine-tuning compute is insufficient on its own to teach models fundamental algorithmic execution. Organizations deploying language models for complex, multi-step workflows face significant performance and reliability risks if they rely solely on standard fine-tuning. Instead, in-context scratchpad prompting activates template-following capabilities inherent in large pre-trained models, allowing them to handle longer problem sequences more effectively without costly architectural modifications.

Decision-makers and engineering teams should prioritize few-shot scratchpad prompting over standard fine-tuning pipelines when designing systems for sequential, multi-step reasoning. Fine-tuning should be applied cautiously, as combining fine-tuning with scratchpads only provides benefits if the base pre-trained model already possesses strong zero-shot baseline performance on the target domain. Future work should focus on developing attention mechanisms that prevent distractor interference and testing whether hybrid prompting strategies generalize to broader non-synthetic applications.

The article's conclusions are supported by controlled synthetic experiments that isolate algorithmic state tracking. However, readers should note that confidence is highest for structured, deterministic tasks, and additional research is required to determine how cleanly these findings translate to highly unstructured, natural language domains.

arXiv: 2207.04901
Cover for Exploring Length Generalization in Large Language Models

Abstract

The ability to extrapolate from short problem instances to longer ones is an important form of out-of-distribution generalization in reasoning tasks, and is crucial when learning from datasets where longer problem instances are rare. These include theorem proving, solving quantitative mathematics problems, and reading/summarizing novels. In this paper, we run careful empirical studies exploring the length generalization capabilities of transformer-based language models. We first establish that naively finetuning transformers on length generalization tasks shows significant generalization deficiencies independent of model scale. We then show that combining pretrained large language models’ in-context learning abilities with scratchpad prompting (asking the model to output solution steps before producing an answer) results in a dramatic improvement in length generalization. We run careful failure analyses on each of the learning modalities and identify common sources of mistakes that highlight opportunities in equipping language models with the ability to generalize to longer problems.

Table of Contents

  • 1 Introduction
  • 2 Length Generalization
  • 2.1 Tasks
  • 3 Standard Finetuning Fails at Length Generalization
  • 3.1 Scale Doesn't Improve Length Generalization
  • 3.2 Transformers Prefer Parallel Strategies over Sequential Ones
  • 3.3 In-Distribution Generalization Doesn't Predict OOD Generalization on Length Generalization Tasks
  • 4 Scratchpad Finetuning Still Fails at Length Generalization
  • 5 Scratchpad Prompting Significantly Improves Length Generalization
  • 5.1 Few-shot scratchpad
  • 6 Related Works
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Deterministic Markov Process Characterization of Length Generalization

    definition

    Length generalization in multi-step reasoning problems can be formalized as state extrapolation within a deterministic Markov process.

    Let s0∈Ss_0 \in \mathcal{S} denote an initial world state sampled from a state space S\mathcal{S}, and let (f1,f2,…,fT)(f_1, f_2, \dots, f_T) be a sequence of TT deterministic state transformation functions ft:S→Sf_t: \mathcal{S} \to \mathcal{S}. The true final state sTs_T after TT steps is given by the composition: sT=fT(fT−1(…f1(s0)… ))s_T = f_T(f_{T-1}(\dots f_1(s_0)\dots))

    The problem length is defined as the number of sequential state transitions TT. In a length generalization benchmark, a learning agent receives the tuple (s0,f1,…,fT)(s_0, f_1, \dots, f_T) and must predict sTs_T. The training dataset consists of problem instances with lengths drawn from an in-distribution interval T∈[Tmin⁡train,Tmax⁡train]T \in [T_{\min}^{\text{train}}, T_{\max}^{\text{train}}], while evaluation is performed out-of-distribution on instances where T>Tmax⁡trainT > T_{\max}^{\text{train}}.

  2. Knowl 2 — Algorithmic Tasks for Measuring Length Generalization

    experimental setup

    Two deterministic synthetic tasks are designed to evaluate length generalization in decoder-only transformer language models:

    1. Parity Task: Given an input bit-string b=(b1,b2,…,bn)∈{0,1}nb = (b_1, b_2, \dots, b_n) \in \{0, 1\}^n, the target output is the parity ∑i=1nbi(mod2)∈{0,1}\sum_{i=1}^n b_i \pmod 2 \in \{0, 1\} (or labeled "odd" vs "even"). Length is defined either as the total number of bits nn or as the number of state flips (the count of 11s). A natural-language variant formats the bit-string as sequential coin flips where "flips" changes state and "doesn't flip" maintains state.

    2. Boolean Variable Assignment Task: The input is a valid Python snippet containing sequential variable assignments (e.g., x=Truex = \text{True}, y=Falsey = \text{False}, z=x and yz = x \text{ and } y), followed by a query for the value of the variable in the final line. The task is evaluated on two splits:

    • Diverse Split: Contains a broad variety of Boolean operations and branching execution flows.
    • Chain-like Split: Restricts operations strictly to composition over previously defined variables, eliminating redundant operations and maximizing sequential dependency length.
  3. Knowl 3 — Invariance of Finetuning Length Generalization Failures to Model Scale

    empirical result

    Supervised finetuning of standard decoder-only transformer language models (LaMDA models spanning 244M, 422M, 1B, and 64B parameters) fails to extrapolate to longer sequence lengths, and increasing model scale by more than two orders of magnitude does not reduce the failure rate.

    When trained on in-distribution problem lengths using the AdaFactor optimizer until validation convergence (e.g., lengths 10 to 21 for Parity; lengths 3 to 8 for Chain-like Variable Assignment):

    • All model sizes achieve near 100% in-distribution accuracy.
    • For test lengths exceeding the training maximum (lengths 22 to 40 for Parity; lengths 9 to 19 for Variable Assignment), accuracy drops rapidly across all model sizes (244M, 422M, 1B, 64B) along nearly identical degradation curves, quickly approaching random chance (50% on binary parity and ~60% on variable assignment).
  4. Knowl 4 — Transformer Bias Toward Parallel Counting and Computational Graph Depth over Sequential Execution

    empirical result

    When finetuned on sequential algorithmic tasks without scratchpads, transformers learn non-sequential pooling shortcuts rather than step-by-step state tracking:

    1. Counting Shortcut on Parity: When input length is held constant at 30 tokens and the model is trained only on inputs with 10 to 20 ones, the model achieves 100% in-distribution accuracy but drops to chance accuracy on inputs with 1 to 9 or 21 to 30 ones. The model learns a permutation-invariant global pooling and thresholding operation over the count of ones rather than sequential accumulation.

    2. Computational Graph Depth on Variable Assignment: For a program represented as a directed acyclic computational graph where nodes are variables and edges are operations, computational graph depth is defined as the length of the longest dependency chain connecting to the queried variable. Transformer accuracy correlates with computational graph depth rather than the total number of lines in the program: models finetuned on programs up to 16 lines successfully solve programs with out-of-distribution line counts (up to 32 lines) as long as the computational graph depth remains within the in-distribution training range.

  5. Knowl 5 — Decoupling of In-Distribution Loss and Out-of-Distribution Length Generalization

    empirical result

    In-distribution cross-entropy loss does not predict out-of-distribution (OOD) length generalization performance in language models on algorithmic tasks.

    Models with identical transformer architectures trained on identical datasets to nearly identical in-distribution cross-entropy loss and in-distribution validation accuracy (~100%) exhibit starkly diverging out-of-distribution performance based solely on optimization hyperparameter choices. For example, on the Parity task, varying the learning rate between 2×10−52\times 10^{-5} and 2×10−32\times 10^{-3} and the batch size between 32 and 128 results in OOD accuracy on longer sequences ranging from near-zero to over 60%, despite indistinguishable in-distribution performance.

  6. Knowl 6 — Supervised Scratchpad Finetuning Fails to Generalize in Length Due to Distractor Attention Failure

    empirical result

    Training transformer models (125M to 64B parameters) to output intermediate step-by-step reasoning tokens (scratchpad finetuning) does not solve length generalization in the supervised finetuning regime. The model displays rapid accuracy decay on out-of-distribution lengths similar to standard direct-answer finetuning.

    An investigation of potential architectural causes reveals:

    • Positional Biases: Adding padding tokens to align relative positional encoding bins (e.g., T5 relative position biases) between input and scratchpad tokens improves performance slightly but fails to resolve length generalization collapse.
    • End-of-Sequence (EOS) Prediction: Premature EOS emission is not the root cause of failure.
    • Attention Degradation on Distractors: Per-step error analysis demonstrates that when evaluated on out-of-distribution lengths, error rates spike during the initial in-distribution scratchpad steps. The failure stems from the self-attention mechanism failing to attend to the correct input tokens in the presence of additional out-of-distribution input tokens (distractors).
  7. Knowl 7 — Extrapolation via In-Context Scratchpad Prompting

    empirical result

    Prompting frozen, pretrained large language models (such as LaMDA 128B) with a few short scratchpad demonstrations enables length extrapolation to problem lengths far exceeding those demonstrated in the prompt.

    When prompted with 3-shot examples of short instances (e.g., length 3 coin-flip parity), the model performs variable-length template matching, applying the step-by-step reasoning structure to queries up to length 20. The per-step error rate remains roughly constant across scratchpad steps (between 0.5% and 2.5% per step), in contrast to finetuned models where per-step error rates spike catastrophically on longer sequence lengths.

  8. Knowl 8 — Performance Trade-offs of Finetuning, Prompting, and Scratchpads in Length Generalization

    data/table

    The interaction among finetuning, few-shot in-context prompting, and scratchpad reasoning produces distinct in-distribution and out-of-distribution length generalization behaviors in language models:

    Techniques In-distribution Out-of-distribution Improves with scale
    Fine-tune Near-perfect Poor Poor
    Prompting Poor Poor Poor
    Fine-tune + Prompting Near-perfect Poor Poor
    Fine-tune + Scratchpad Near-perfect Poor Poor
    Prompting + Scratchpad Nontrivial Nontrivial Nontrivial / Strong
    Fine-tune + Prompting + Scratchpad Near-perfect Task-dependent Strong

    Supervised finetuning achieves near-perfect in-distribution execution but consistently fails out-of-distribution across all scales, with or without scratchpads. In-context few-shot prompting combined with scratchpad generation is the only single paradigm that achieves length extrapolation and improves with parameter scale.

  9. Knowl 9 — Base Model Task Competence Governs Few-Shot Scratchpad Finetuning Success

    empirical result

    Combining supervised finetuning with few-shot scratchpad prompting yields out-of-distribution length generalization only when the underlying pretrained base model already possesses non-finetuned competence on the task:

    • Pretrained Competence Present (Natural Language Coin-Flip Parity): When the pretrained base model already achieves strong few-shot scratchpad performance prior to finetuning, few-shot finetuning with scratchpad substantially boosts accuracy on both in-distribution (approaching 100%) and out-of-distribution lengths (maintaining ~85% accuracy up to length 20).
    • Pretrained Competence Absent (Raw Boolean Variable Assignment): When the base model cannot perform the task zero-shot/few-shot before finetuning, few-shot scratchpad finetuning fails to extrapolate, exhibiting the same steep accuracy degradation on out-of-distribution lengths as zero-shot scratchpad finetuning.

Coverage note — None was omitted; all primary theoretical definitions, algorithmic tasks, empirical findings across finetuning, scratchpad, and prompting modalities, error analyses, and comparative tables were included.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  2. 2.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. Impact of pretraining term frequencies on few-shot reasoning. arXiv preprint arXiv:2202.07206, 2022.
  3. 3.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114, 2021.
  4. 4.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models, 2022. URL https://arxiv.org/abs/2201.11903.
  5. 5.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. Solving quantitative reasoning problems with language models. arXiv preprint arXiv:2206.14858, 2022.
  6. 6.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.
  7. 7.Haoyu Wang, Mo Yu, Xiaoxiao Guo, Rajarshi Das, Wenhan Xiong, and Tian Gao. Do multi-hop readers dream of reasoning chains? arXiv preprint arXiv:1910.14520, 2019.
  8. 8.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  9. 9.Peter Clark, Oyvind Tafjord, and Kyle Richardson. Transformers as soft reasoners over language. arXiv preprint arXiv:2002.05867, 2020.
  10. 10.Yuhuai Wu, Albert Qiaochu Jiang, Jimmy Ba, and Roger Grosse. Int: An inequality benchmark for evaluating generalization in theorem proving. arXiv preprint arXiv:2007.02924, 2020.
  11. 11.Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR, 2018.
  12. 12.Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. Set transformer: A framework for attention-based permutation-invariant neural networks. In International Conference on Machine Learning, pages 3744–3753. PMLR, 2019.
  13. 13.Mirac Suzgun, Sebastian Gehrmann, Yonatan Belinkov, and Stuart M Shieber. Lstm networks can perform dynamic counting. arXiv preprint arXiv:1906.03648, 2019.
  14. 14.Vaishnavh Nagarajan, Anders Andreassen, and Behnam Neyshabur. Understanding the failure modes of out-of-distribution generalization. arXiv preprint arXiv:2010.15775, 2020.
  15. 15.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114, 2021.
  16. 16.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.
  17. 17.Benjamin Newman, John Hewitt, Percy Liang, and Christopher D Manning. The eos decision and length extrapolation. arXiv preprint arXiv:2010.07174, 2020.
  18. 18.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  19. 19.Yann Dubois, Gautier Dagan, Dieuwke Hupkes, and Elia Bruni. Location attention for extrapolation to longer sequences. arXiv preprint arXiv:1911.03872, 2019.
  20. 20.Kenton Murray and David Chiang. Correcting length bias in neural machine translation. arXiv preprint arXiv:1808.10006, 2018.
  21. 21.Gilad Yehudai, Ethan Fetaya, Eli Meirom, Gal Chechik, and Haggai Maron. From local structures to size generalization in graph neural networks. In International Conference on Machine Learning, pages 11975–11986. PMLR, 2021.
  22. 22.Da Ju, Stephen Roller, Sainbayar Sukhbaatar, and Jason Weston. Staircase attention for recurrent processing of sequences. arXiv preprint arXiv:2106.04279, 2021.
  23. 23.Ofir Press, Noah Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=R8sQPpGCv0.
  24. 24.Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner. Unveiling transformers with lego: a synthetic reasoning task. arXiv preprint arXiv:2206.04301, 2022.
  25. 25.Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, and Tom Goldstein. Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks. Advances in Neural Information Processing Systems, 34, 2021.
  26. 26.Arpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam, Furong Huang, Micah Goldblum, and Tom Goldstein. End-to-end algorithm synthesis with recurrent networks: Logical extrapolation without overthinking. arXiv preprint arXiv:2202.05826, 2022.
  27. 27.Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser. Universal transformers. arXiv preprint arXiv:1807.03819, 2018.
  28. 28.Łukasz Kaiser and Ilya Sutskever. Neural gpus learn algorithms. arXiv preprint arXiv:1511.08228, 2015.
  29. 29.R Thomas McCoy, Robert Frank, and Tal Linzen. Does syntax need to grow on trees? sources of hierarchical inductive bias in sequence-to-sequence networks. Transactions of the Association for Computational Linguistics, 8:125–140, 2020.
  30. 30.Eugene Kharitonov and Rahma Chaabouni. What they do when in doubt: a study of inductive biases in seq2seq learners. arXiv preprint arXiv:2006.14953, 2020.
  31. 31.He He, Sheng Zha, and Haohan Wang. Unlearn dataset bias in natural language inference by fitting the residual. arXiv preprint arXiv:1908.10763, 2019.
  32. 32.R Thomas McCoy, Ellie Pavlick, and Tal Linzen. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. arXiv preprint arXiv:1902.01007, 2019.
  33. 33.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014.

Citation

MLA
Anil, C., et al. “Exploring Length Generalization in Large Language Models”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 38546–56, https://proceedings.neurips.cc/paper_files/paper/2022/file/fb7451e43f9c1c35b774bcfad7a5714b-Paper-Conference.pdf.
APA
Anil, C., Wu, Y., Andreassen, A., Lewkowycz, A., Misra, V., Ramasesh, V., Slone, A., Gur-Ari, G., Dyer, E., & Neyshabur, B. (2022). Exploring Length Generalization in Large Language Models. Advances in Neural Information Processing Systems, 35, 38546–38556. https://proceedings.neurips.cc/paper_files/paper/2022/file/fb7451e43f9c1c35b774bcfad7a5714b-Paper-Conference.pdf
Chicago
Anil, C., Y. Wu, A. Andreassen, et al. 2022. “Exploring Length Generalization in Large Language Models”. Advances in Neural Information Processing Systems 35: 38546–56. https://proceedings.neurips.cc/paper_files/paper/2022/file/fb7451e43f9c1c35b774bcfad7a5714b-Paper-Conference.pdf.
Harvard
Anil, C. et al. (2022) “Exploring Length Generalization in Large Language Models”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 38546–38556. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/fb7451e43f9c1c35b774bcfad7a5714b-Paper-Conference.pdf.
Vancouver
1. Anil C, Wu Y, Andreassen A, Lewkowycz A, Misra V, Ramasesh V, Slone A, Gur-Ari G, Dyer E, Neyshabur B (2022) Exploring Length Generalization in Large Language Models. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 38546–38556

BibTeX

@inproceedings{anil2022exploring,
  title = {Exploring Length Generalization in Large Language Models},
  author = {Anil, Cem and Wu, Yuhuai and Andreassen, Anders and Lewkowycz, Aitor and Misra, Vedant and Ramasesh, Vinay and Slone, Ambrose and Gur-Ari, Guy and Dyer, Ethan and Neyshabur, Behnam},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {38546-38556},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/fb7451e43f9c1c35b774bcfad7a5714b-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission