Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
Koyena PalJiuding SunAndrew YuanByron C. WallaceDavid Bau
Demonstrates through linear probing and causal state transplantation that single intermediate hidden states in autoregressive transformers encode information capable of predicting tokens multiple steps into the future.
Large language models are standardly trained to generate text step-by-step by predicting only the immediate next word piece, or token. However, as these systems grow more complex and are deployed in high-stakes environments, understanding their internal mechanics becomes critical for safety, auditing, and computational efficiency. The article addresses whether internal representations inside a transformer model contain information about words several steps ahead, rather than just the single upcoming token.
The main objective of the article is to demonstrate that an individual internal vector—known as a hidden state—encodes sufficient signal to reliably anticipate text multiple positions into the future. The researchers evaluate this hypothesis using the 6-billion-parameter GPT-J model, testing how accurately future tokens can be decoded from a single state.
To conduct this evaluation, the authors used a sample of 100,000 training examples and 1,000 test examples drawn from the Pile dataset. They compared four decoding techniques across all 28 layers of the model: two linear translation models that map intermediate states either directly to vocabulary words or to final-layer states, a causal transplantation method inserting the state into generic fixed phrases, and an optimized soft prompt method that learns a specialized prefix to steer text generation from that transplanted state. Predictions were evaluated up to four steps ahead using accuracy and statistical surprise metrics.
The findings confirm that single hidden states encode information multiple tokens ahead. The learned prompt method achieved the highest performance, anticipating tokens two steps ahead with 48.4% accuracy, three steps ahead with 43.7% accuracy, and four steps ahead with 46.9% accuracy. This substantially outperformed standard word-association baselines (20.1%) as well as linear probes (reaching at most 29.2%). Crucially, future token information peaked in the middle computational layers (around layer 14), in sharp contrast to immediate next-token prediction, which concentrates in the final layers. Furthermore, decoding accuracy strongly correlated with the model's confidence, rising from 26% when the model was uncertain to 95% when it was highly confident.
These results demonstrate that language models plan multi-token sequences internally before emitting them, rather than operating purely as step-by-step predictors. This finding has direct implications for model interpretability, model editing, and inference efficiency. By showing that middle layers already commit to future trajectory concepts, the work provides a foundation for auditing model reasoning chains and building early-exit systems that reduce computational costs and latency. Based on these insights, the authors developed a visualization tool named Future Lens to inspect multi-step planning across layers.
Organizations developing or deploying large language models should consider integrating multi-token probing frameworks into their model inspection, safety auditing, and optimization pipelines. Before deploying these techniques in production, further research is required to evaluate broader model families, languages beyond English, and longer prediction horizons beyond four steps. While the study provides high-confidence evidence within GPT-J-6B, practitioners should treat the results cautiously until validated across different model architectures and diverse dataset distributions.
- Paper: Transformer Feed-Forward Layers Are Key-Value Memories, Mor Geva et al. (2020). This foundational study demonstrates how intermediate transformer layers project concepts and vocabulary predictions into hidden states, providing the core mechanistic foundation for Future Lens's probing across layers.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). Understanding how representation geometry and contextuality evolve across transformer layers is essential context for analyzing how future tokens are encoded within intermediate hidden representations.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). This paper establishes the layer-wise probing methodology that Future Lens extends from traditional linguistic properties to multi-token lookahead planning.
- Paper: Next-Latent Prediction Transformers Learn Compact World Models, Jayden Teoh et al. (2025). Building on the insight that single hidden states encode future trajectory signals, this work explicitly trains transformers with an auxiliary objective to predict future latent states.
- Paper: Training Large Language Models to Reason in a Continuous Latent Space, Shibo Hao et al. (2024). This work directly operationalizes continuous multi-step internal planning by training language models to reason over sequential latent states rather than discrete token emissions.
- Paper: Language Models are Injective and Hence Invertible, Giorgos Nikolaou et al. (2025). This work explores the complementary theoretical and algorithmic invertibility of hidden states across layers, demonstrating how token information can be recovered directly from representations.
- Paper: Learn from your own latents and not from tokens: A sample-complexity theory, Daniel J. Korchinski et al. (2026). This theoretical study extends the investigation of latent representations by demonstrating the sample complexity advantages of learning from internal latent predictions rather than raw tokens.
