Learning to (Learn at Test Time): RNNs with Expressive Hidden States
Yu SunXinhao LiKaran DalalJiarui XuArjun VikramGenghan ZhangYann DuboisXinlei ChenXiaolong WangSanmi Koyejo
Introduces Test-Time Training layers that treat RNN hidden states as internal machine learning models updated via self-supervised gradient steps, achieving linear-time sequence modeling that continues to improve across long contexts where existing architectures plateau.
Modern language models face a fundamental trade-off between computational efficiency and the ability to process long context. Self-attention mechanisms achieve superior language understanding across long contexts, but their processing cost grows quadratically with sequence length. Conversely, recurrent neural networks (RNNs) offer linear time complexity and constant inference cost per token, yet their performance degrades across long sequences because fixed-size hidden states fail to capture complex dependencies among thousands or millions of tokens.
The article introduces and evaluates Test-Time Training (TTT) layers, a novel sequence modeling framework designed to retain linear computational complexity while significantly enhancing hidden state expressiveness. To overcome standard RNN memory limitations, the authors frame the hidden state itself as an internal machine learning model whose parameters are continuously updated on test sequences using self-supervised gradient descent.
The researchers implemented two concrete versions: TTT-Linear, which uses a linear model as its internal state, and TTT-MLP, which employs a two-layer multi-layer perceptron. They evaluated these architectures against competitive Transformer and Mamba (a modern RNN baseline) models at scales ranging from 125 million to 1.3 billion parameters across context windows up to 32,000 tokens using standard text benchmarks. To ensure practical execution on modern hardware, the authors implemented mini-batch test-time updates and a mathematically equivalent dual formulation that accelerates training more than fivefold on hardware accelerators.
The empirical findings demonstrate that TTT layers maintain performance comparable to leading baselines in short contexts while outperforming existing RNN architectures in long contexts. At an 8,000-token context length on the Pile dataset and a 32,000-token context length on the Books benchmark, both TTT variants achieved lower perplexity than Mamba. Furthermore, like Transformers, TTT layers steadily decrease prediction perplexity as context grows through 32,000 tokens, whereas Mamba's ability to utilize additional context plateaus after 16,000 tokens. Theoretical analyses also established that linear attention and standard self-attention represent special parametric and non-parametric cases within this broader TTT formulation.
These results indicate that embedding self-supervised learning directly into recurrent hidden states resolves the compression bottleneck of linear-time sequence models without incurring the prohibitive memory growth of full attention. For engineering deployments, TTT-Linear delivers constant inference latency per token and matches or beats standard training speeds, demonstrating immediate viability for long-sequence tasks. While TTT-MLP demonstrates higher theoretical expressiveness in long contexts, its hardware input/output complexity currently introduces notable wall-clock time overhead during generation.
Organizations developing large-scale language systems should evaluate TTT-Linear as a drop-in recurrent alternative for long-context workloads where full self-attention is cost-prohibitive. Prior to deploying deeper configurations like TTT-MLP in production, engineering teams should conduct targeted hardware and kernel optimization pilots to alleviate memory bandwidth bottlenecks. Future research should focus on extending evaluations to multi-million token sequences, exploring convolutional architectures for the internal learner, and building dedicated multi-device pipeline parallelization.
- Paper: Test-Time Training with Self-Supervision for Generalization under Distribution Shifts, Yu Sun et al. (2019). Introduces the foundational Test-Time Training (TTT) framework that the source adapts directly into recurrent sequence hidden states.
- Paper: Mamba: Linear-Time Sequence Modeling with Selective State Spaces, Albert Gu et al. (2023). Introduces Mamba, the primary state-space modern recurrent baseline that the source benchmarks against and aims to surpass in long-context modeling.
- Paper: Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, Tri Dao et al. (2024). Establishes structured state-space duality connecting linear attention and RNNs, providing the theoretical context for the source's dual-formulation acceleration.
- Paper: Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos et al. (2020). Formulates linear attention as an autoregressive recurrent model, which the source establishes as a special parametric case of Test-Time Training.
- Paper: Simple linear attention language models balance the recall-throughput tradeoff, Simran Arora et al. (2024). Analyzes the fundamental memory-versus-recall tradeoff in linear causal RNNs that the source solves via self-supervised hidden state updates.
- Paper: Hungry Hungry Hippos: Towards Language Modeling with State Space Models, Daniel Y. Fu et al. (2023). Demonstrates the expressivity and associative recall limitations of traditional state-space models that motivated expressive hidden state architectures.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Introduces standard Transformer self-attention, defining the long-context performance benchmark and non-parametric limit analyzed in the source.
- Paper: End-to-End Test-Time Training for Long Context, Arnuv Tandon et al. (2025). Extends the source's layer-level test-time training concept into an end-to-end framework directly trained across entire Transformer backbones.
- Paper: Test-Time Training with KV Binding Is Secretly Linear Attention, Junchen Liu et al. (2026). Critically re-examines and generalizes the source's TTT mechanism by mathematically demonstrating its equivalence to learned linear attention.
- Paper: Test-time regression: a unifying framework for designing sequence models with associative memory, Ke Alexander Wang et al. (2025). Unifies test-time training sequence layers and attention mechanisms into a comprehensive online test-time regression framework.
- Paper: Fast Weight Attention for Continual Learning, Yifan Zhang et al. (2026). Applies test-time regression and fast-weight updates with explicit normalization and write-after-read causality for online continual learning.
- Paper: Memory Caching: RNNs with Growing Memory, Ali Behrouz et al.. Provides an alternative approach to overcoming the fixed hidden-state compression bottleneck of RNNs by caching checkpointed memory states.
- Paper: The Surprising Effectiveness of Test-Time Training for Few-Shot Learning, Ekin Akyrek et al. (2025). Applies gradient-based test-time parameter updates to improve LLM adaptation and in-context reasoning on complex out-of-distribution benchmarks.
- Paper: Dynamic Linear Attention, Xin Wang et al. (2026). Builds on subquadratic linear recurrent backbones like Mamba-2 by introducing dynamic, information-aware state allocation.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). Surveys the broader landscape of efficient LLM architectures, situating linear RNNs and test-time training within modern subquadratic sequence modeling.
