Exploring Length Generalization in Large Language Models
Cem AnilYuhuai WuAnders AndreassenAitor LewkowyczVedant MisraVinay V. RamaseshAmbrose SloneGuy Gur-AriEthan DyerBehnam Neyshabur
Reveals that while standard finetuning fails to extrapolate reasoning to longer problem instances regardless of model scale, combining in-context learning with scratchpad prompting substantially improves length generalization in large language models.
Real-world reasoning tasks—such as mathematical problem solving, program execution, and theorem proving—naturally increase in difficulty as problem length increases. In these domains, long training examples are exceedingly rare, making it critical for artificial intelligence systems to learn generalizable rules from short instances and extrapolate them to longer ones. The article evaluates the ability of transformer-based large language models to perform this type of out-of-distribution reasoning, termed length generalization, across model scales and training methodologies.
To systematically analyze this capability, the article investigates two controlled algorithmic benchmarks: a parity task (determining whether the count of ones in a bit-string is odd or even) and a Boolean variable assignment task (tracking sequential program execution). The authors conducted extensive empirical tests using decoder-only language models ranging from 244 million up to 128 billion parameters, comparing standard fine-tuning, intermediate step generation (scratchpad or chain-of-thought methods), and in-context few-shot prompting.
The investigation produced four central findings. First, standard fine-tuning consistently fails to generalize to longer problems regardless of model size; even when models achieve near 100% accuracy on training lengths, their performance rapidly degrades toward random guessing on longer instances. Second, transformer architectures exhibit a strong bias toward learning parallel counting shortcuts rather than sequential step-by-step algorithms, making in-distribution training loss an unreliable predictor of out-of-distribution success. Third, contrary to prior assumptions, fine-tuning models to produce step-by-step scratchpad solutions also fails to generalize to longer inputs due to attention failures across longer token contexts. Fourth, combining pre-trained models with few-shot scratchpad prompting dramatically improves length generalization without any parameter fine-tuning, successfully enabling models to map short step-by-step templates to instances up to five times longer than prompt demonstrations.
These findings indicate that scaling model size, data volume, or fine-tuning compute is insufficient on its own to teach models fundamental algorithmic execution. Organizations deploying language models for complex, multi-step workflows face significant performance and reliability risks if they rely solely on standard fine-tuning. Instead, in-context scratchpad prompting activates template-following capabilities inherent in large pre-trained models, allowing them to handle longer problem sequences more effectively without costly architectural modifications.
Decision-makers and engineering teams should prioritize few-shot scratchpad prompting over standard fine-tuning pipelines when designing systems for sequential, multi-step reasoning. Fine-tuning should be applied cautiously, as combining fine-tuning with scratchpads only provides benefits if the base pre-trained model already possesses strong zero-shot baseline performance on the target domain. Future work should focus on developing attention mechanisms that prevent distractor interference and testing whether hybrid prompting strategies generalize to broader non-synthetic applications.
The article's conclusions are supported by controlled synthetic experiments that isolate algorithmic state tracking. However, readers should note that confidence is highest for structured, deterministic tasks, and additional research is required to determine how cleanly these findings translate to highly unstructured, natural language domains.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces chain-of-thought prompting as a mechanism to elicit multi-step reasoning in large language models, providing the baseline prompting paradigm examined and tested for length generalization in the source.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). Presents foundational findings on length extrapolation in transformer architectures and the failure modes of standard positional encodings on out-of-distribution sequence lengths.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Establishes in-context few-shot prompting capabilities in autoregressive large language models, which the source explores as an alternative to fine-tuning for algorithmic generalization.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). Provides early empirical analysis of transformer program execution and algorithmic synthesis across scales, setting the groundwork for evaluating execution tasks on extended sequences.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). Demonstrates how neural models rely on superficial statistical shortcuts rather than robust algorithmic rules, contextualizing why transformers fail out-of-distribution length generalization.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Directly tackles the easy-to-hard and length generalization failures identified in standard chain-of-thought prompting by introducing sequential subproblem decomposition.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). Provides a formal analysis of where step-by-step reasoning breaks down on longer proof lengths due to greedy search and local deduction failures.
- Paper: Why think step by step? Reasoning emerges from the locality of experience, Ben Prystawski et al. (2023). Offers theoretical and empirical explanations for why step-by-step intermediate generation succeeds over direct prediction by analyzing the statistical locality of training data.
- Paper: A Length-Extrapolatable Transformer, Yutao Sun et al. (2023). Develops architectural innovations like extrapolatable position embeddings to directly address the attention degradation on long sequences documented in the source.
- Paper: Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision, Zhiqing Sun et al. (2024). Extends out-of-distribution reasoning from simple synthetic length extrapolation to harder multi-step math and coding problems using easy-to-hard verifier supervision.
- Paper: GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models, Iman Mirzadeh et al. (2025). Investigates the fragility of chain-of-thought reasoning under controlled structural and complexity variations using symbolic math problem modifications.
- Paper: Training Large Language Models to Reason in a Continuous Latent Space, Shibo Hao et al. (2024). Addresses the inefficiency and error accumulation of long autoregressive reasoning chains by executing intermediate steps directly within continuous latent space.
