Auto-Regressive Next-Token Predictors are Universal Learners
Eran Malach
Proves that even simple linear next-token predictors trained on chain-of-thought sequences can approximate any efficiently computable Turing machine, demonstrating that the reasoning power of modern language models stems fundamentally from autoregressive supervision rather than specific transformer architectures.
Modern artificial intelligence heavily relies on massive large language models that perform complex reasoning and problem-solving. A central question facing decision-makers and technical leaders is whether these advanced capabilities stem from complex neural architectures, such as transformers, or from the fundamental training objective of predicting the next token in a sequence. The article addresses this fundamental question to determine what truly drives reasoning performance in machine learning models.
The article demonstrates that the power of modern language models is primarily driven by sequential, step-by-step next-token prediction rather than complex model architectures. It evaluates how basic linear predictors and shallow multi-layer networks perform when trained with intermediate reasoning sequences, commonly known as chain of thought, to show that these simple frameworks can theoretically learn any computation that a standard computer can execute.
To establish these results, the article combines formal mathematical proofs with targeted empirical experiments. The theoretical analysis adapts standard learning theory to the sequential next-token prediction setting, establishing proof that linear predictors can simulate general computer algorithms when provided with intermediate steps. The empirical evaluation tests these principles in practice: a basic linear network is trained on a synthetic short-story dataset across fifty evaluation prompts, and a shallow 775-million-parameter network without attention mechanisms is trained on four-digit multiplication tasks using one hundred million training sequences.
The article presents several key findings. Theoretically, auto-regressive next-token prediction allows simple linear models to approximate any efficiently computable function, provided they receive step-by-step reasoning data. The article also introduces length complexity, showing that the number of intermediate reasoning tokens can be systematically traded against computational and sample complexity. In empirical tests, a shallow four-layer network without attention achieved 96.9% exact accuracy on four-digit multiplication, matching a specialized seven-billion-parameter transformer and substantially outperforming leading commercial models, which achieved approximately 1% to 5% accuracy on direct calculation without intermediate steps. Additionally, a simple linear model trained on short stories produced grammatically sound and coherent text across standard benchmarks.
These findings carry significant strategic implications for machine learning design, computational cost, and resource allocation. They indicate that complex reasoning can be unlocked in smaller, simpler models without relying exclusively on compute-heavy architectures, provided that high-quality intermediate supervision is available. Organizations can potentially achieve high task performance and lower operational latency by investing in rich step-by-step training data rather than solely scaling model parameter counts.
Leaders should focus machine learning strategies on data quality and task decomposition, especially for logical and algorithmic domains. When designing systems for mathematical, symbolic, or multi-step logic tasks, teams should integrate explicit step-by-step supervision into data pipelines rather than assuming larger architectures are necessary. Further work should explore automated generation of intermediate data sequences and investigate the trade-offs of length complexity in production environments.
The primary limitation of this approach is its heavy dependence on detailed intermediate supervision, as generating extensive step-by-step training data can be labor-intensive and costly. Furthermore, the empirical validation focused on controlled domains—short stories and arithmetic—meaning confidence is high for structured algorithmic tasks, but additional validation is required before extending these architectural simplifications to broader, open-ended tasks.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This foundational work introduced chain-of-thought prompting in autoregressive language models, establishing the empirical reasoning paradigm whose theoretical and architectural necessity is scrutinized in the source paper.
- Paper: Why think step by step? Reasoning emerges from the locality of experience, Ben Prystawski et al. (2023). This paper establishes how step-by-step intermediate variables reduce estimation bias in autoregressive models, providing essential theoretical and statistical groundwork for understanding why sequential next-token prediction drives reasoning.
- Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). This study demonstrates that small student models can acquire multi-step reasoning when trained on intermediate teacher traces, directly preceding the source paper's thesis that simple autoregressive learners thrive on step-by-step supervision.
- Paper: Exploring Length Generalization in Large Language Models, Cem Anil et al. (2022). This paper examines how length generalization and sequential scratchpads operate in transformer architectures, directly motivating the source's formal introduction of length complexity in next-token predictors.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). This formal analysis shows that autoregressive models execute local reasoning steps effectively but struggle with global planning, contextualizing the source's findings on the expressive power and limits of sequential next-token predictors.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). This work demonstrates that simple sequential prompting triggers multi-step problem solving in autoregressive models, providing key empirical context for the source's claim that sequential prediction fundamentally unlocks universal computation.
- Paper: CoT-Valve: Length-Compressible Chain-of-Thought Tuning, Xinyin Ma et al. (2025). This work builds on the trade-offs of intermediate reasoning sequences by introducing parameter-space mechanisms to compress chain-of-thought length while maintaining task accuracy.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). This paper advances the study of sequential intermediate supervision by training language models to internalize non-linear trial-and-error search and meta-reasoning traces.
- Paper: s1: Simple test-time scaling, Niklas Muennighoff et al. (2025). This paper operationalizes the test-time length trade-off proven in the source by implementing budget forcing to scale reasoning performance with minimal compute and data.
- Paper: Training Large Language Models to Reason in a Continuous Latent Space, Shibo Hao et al. (2024). This study extends sequential autoregressive reasoning beyond discrete token sequences by demonstrating how intermediate steps can be computed effectively in a continuous latent space.
- Paper: Hierarchical Reasoning Model, Guan Wang et al. (2025). This work explores an alternative architectural paradigm for sequential reasoning, using a compact recurrent framework with deep supervision to solve complex planning tasks without token-by-token explicit traces.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). This study analyzes the internal structure of long autoregressive reasoning traces, showing how complex multi-step generation manifests as simulated multi-perspective dialogue.
- Paper: Understanding R1-Zero-Like Training: A Critical Perspective, Zichen Liu et al. (2025). This paper examines whether reasoning in autoregressive models emerges autonomously from reinforcement learning or relies on latent capabilities instilled during next-token pre-training.
- Paper: Efficient Reasoning on the Edge, Yelysei Bondarenko et al. (2026). This work applies the principles of compact reasoning models and length control to edge deployment, demonstrating how to adapt small models under tight compute budgets.
