Future Lens: Anticipating Subsequent Tokens from a Single Hidden State

Koyena PalJiuding SunAndrew YuanByron C. WallaceDavid Bau

article2023CoNLL108 citations

Demonstrates through linear probing and causal state transplantation that single intermediate hidden states in autoregressive transformers encode information capable of predicting tokens multiple steps into the future.

Listen

Large language models are standardly trained to generate text step-by-step by predicting only the immediate next word piece, or token. However, as these systems grow more complex and are deployed in high-stakes environments, understanding their internal mechanics becomes critical for safety, auditing, and computational efficiency. The article addresses whether internal representations inside a transformer model contain information about words several steps ahead, rather than just the single upcoming token.

The main objective of the article is to demonstrate that an individual internal vector—known as a hidden state—encodes sufficient signal to reliably anticipate text multiple positions into the future. The researchers evaluate this hypothesis using the 6-billion-parameter GPT-J model, testing how accurately future tokens can be decoded from a single state.

To conduct this evaluation, the authors used a sample of 100,000 training examples and 1,000 test examples drawn from the Pile dataset. They compared four decoding techniques across all 28 layers of the model: two linear translation models that map intermediate states either directly to vocabulary words or to final-layer states, a causal transplantation method inserting the state into generic fixed phrases, and an optimized soft prompt method that learns a specialized prefix to steer text generation from that transplanted state. Predictions were evaluated up to four steps ahead using accuracy and statistical surprise metrics.

The findings confirm that single hidden states encode information multiple tokens ahead. The learned prompt method achieved the highest performance, anticipating tokens two steps ahead with 48.4% accuracy, three steps ahead with 43.7% accuracy, and four steps ahead with 46.9% accuracy. This substantially outperformed standard word-association baselines (20.1%) as well as linear probes (reaching at most 29.2%). Crucially, future token information peaked in the middle computational layers (around layer 14), in sharp contrast to immediate next-token prediction, which concentrates in the final layers. Furthermore, decoding accuracy strongly correlated with the model's confidence, rising from 26% when the model was uncertain to 95% when it was highly confident.

These results demonstrate that language models plan multi-token sequences internally before emitting them, rather than operating purely as step-by-step predictors. This finding has direct implications for model interpretability, model editing, and inference efficiency. By showing that middle layers already commit to future trajectory concepts, the work provides a foundation for auditing model reasoning chains and building early-exit systems that reduce computational costs and latency. Based on these insights, the authors developed a visualization tool named Future Lens to inspect multi-step planning across layers.

Organizations developing or deploying large language models should consider integrating multi-token probing frameworks into their model inspection, safety auditing, and optimization pipelines. Before deploying these techniques in production, further research is required to evaluate broader model families, languages beyond English, and longer prediction horizons beyond four steps. While the study provides high-confidence evidence within GPT-J-6B, practitioners should treat the results cautiously until validated across different model architectures and diverse dataset distributions.

Cover for Future Lens: Anticipating Subsequent Tokens from a Single Hidden State

Table of Contents

  • 1 Introduction
  • 2 Methods
  • 2.1 Direct Vocabulary Prediction
  • 2.2 Linear Model Approximation
  • 2.3 Fixed Prompt Causal Intervention
  • 2.4 Learned Prompt Causal Intervention
  • 3 Experiments and Results
  • 3.1 Data
  • 3.2 Evaluation Metrics
  • 3.3 Experimental Setup
  • 3.4 Unveiling Subsequent Tokens
  • 4 Related Work
  • 5 Discussion
  • Acknowledgements
  • References
  • A Appendix
  • Additional Figures
  • Limitations

Knowls

  1. Knowl 1 — Causal Intervention with Learned Soft Prompts for Multi-Token Decoding

    model/method

    To extract future token predictions implicitly encoded in a single hidden state hTl∈Rdhh_T^l \in \mathbb{R}^{d_h} (from layer ll at input prefix token position TT) of an autoregressive transformer GG, a causal transplantation framework is employed with a parameterized prefix (soft prompt).

    Let x=[x1,…,xT]x = [x_1, \dots, x_T] denote the original input sequence yielding subsequent token generations [xT+1,…,xT+N][x_{T+1}, \dots, x_{T+N}], with target next-step output distribution yT+N=G([x1,…,xT+N])y_{T+N} = G([x_1, \dots, x_{T+N}]). A continuous prompt prefix c(l)=[c1(l),…,cM(l)]c^{(l)} = [c_1^{(l)}, \dots, c_M^{(l)}] of length M=10M=10 tokens is introduced for layer ll. The hidden representation at position MM of the prefix context run at layer ll is replaced with the transplanted hidden state hTlh_T^l:

    y^M+N=G([c1(l),…,cM(l),xT+1,…,xT+N]∥hMl:=hTl)\hat{y}_{M+N} = G([c_1^{(l)}, \dots, c_M^{(l)}, x_{T+1}, \dots, x_{T+N}] \parallel h_M^l := h_T^l)

    With all parameters of the underlying transformer GG frozen, the soft prompt vectors c(l)c^{(l)} are optimized by back-propagating gradients directly through the network to minimize the Kullback-Leibler (KL) divergence to the target distribution:

    arg⁡min⁡c(l)DKL(y^M+N∥yT+N)\arg\min_{c^{(l)}} D_{\text{KL}}(\hat{y}_{M+N} \parallel y_{T+N})

    Optimizing gradients directly through the frozen transformer yields a significantly more effective context for unlocking subsequent token predictions than routing gradients through an auxiliary multi-layer perceptron.

  2. Knowl 2 — Quantitative Comparison of Future Token Anticipation Methods

    data/table

    Evaluating multi-token probing across future horizons N∈{0,1,2,3}N \in \{0, 1, 2, 3\} (where N=0N=0 corresponds to the immediate next token xT+1x_{T+1}, and N≥1N \ge 1 corresponds to subsequent future tokens xT+1+Nx_{T+1+N}) on a 1000-token test set from The Pile using GPT-J-6B demonstrates that single hidden states contain substantial signal about tokens beyond the immediate next step. The learned prompt intervention significantly outperforms linear decoders and fixed prompt interventions.

    Method N=0 N=1 N=2 N=3
    Accuracy (%)
    Learned Prompt 97.0 48.4 43.7 46.9
    Fixed Prompt 97.0 20.8 30.0 36.5
    Linear Hidden State (HS) 98.0 29.2 19.0 15.8
    Linear Vocabulary (VOCAB) 85.7 27.5 19.4 14.7
    Surprisal
    Learned Prompt 0.6 4.5 4.4 3.9
    Fixed Prompt 0.6 8.8 6.5 5.7
    Linear Hidden State (HS) 0.8 14.1 13.2 13.1
    Linear Vocabulary (VOCAB) 0.9 15.3 14.4 14.2

    For N=1N=1, the learned prompt achieves 48.4% top-1 accuracy, which is more than double the 20.1% accuracy achieved by an empirical corpus bigram baseline. The learned prompt also achieves the lowest surprisal across all future steps (N=1,2,3N=1, 2, 3), confirming that the underlying transformer layers facilitate reading out multi-token predictions that simple linear mappings cannot fully decode.

  3. Knowl 3 — Middle-Layer Concentration of Future Token Information

    empirical result

    In an autoregressive transformer (GPT-J-6B), the layer-wise distribution of information for predicting future tokens differs fundamentally between immediate next-token prediction (N=0N = 0) and multi-token lookahead (N≥1N \ge 1):

    1. For the immediate next token (N=0N = 0), probing accuracy increases monotonically through the network and reaches its maximum at the final layer (L=28L = 28), achieving over 97% Precision@1.
    2. For subsequent future tokens (N∈{1,2,3}N \in \{1, 2, 3\}), predictive accuracy exhibits an inverted U-shape across the network depth, peaking sharply in the intermediate layers around layer l=14l = 14 before dropping off in the deepest layers.

    This demonstrates that intermediate representations in middle layers temporarily encode contextual multi-token trajectories before the network narrows its computation in the final layers toward outputting the single immediate next token.

  4. Knowl 4 — Linear Probing Methods for Future Token Anticipation

    model/method

    Two linear probing architectures can be trained to predict future tokens xT+N+1x_{T+N+1} and target distributions yT+Ny_{T+N} from an individual intermediate hidden state hTl∈Rdhh_T^l \in \mathbb{R}^{d_h} at layer ll of an autoregressive transformer GG:

    1. Direct Vocabulary Prediction Model: A parameterized linear mapping gθ:Rdh→Rdvg_\theta: \mathbb{R}^{d_h} \to \mathbb{R}^{d_v} directly maps hTlh_T^l to the full vocabulary logit space z^T+N=gθ(hTl)\hat{z}_{T+N} = g_\theta(h_T^l), with probability distribution:

    y^T+N=softmax(gθ(hTl))\hat{y}_{T+N} = \text{softmax}(g_\theta(h_T^l))

    1. Future Hidden State Approximation Model: An extension of the tuned lens where a parameterized linear map fθ:Rdh→Rdhf_\theta: \mathbb{R}^{d_h} \to \mathbb{R}^{d_h} projects hTlh_T^l to the predicted final-layer hidden state h^T+NL\hat{h}_{T+N}^L corresponding to future token position T+NT+N:

    h^T+NL=fθ(hTl)\hat{h}_{T+N}^L = f_\theta(h_T^l)

    The vocabulary distribution is then decoded via the transformer's fixed pretrained decoder head D:Rdh→[0,1]dvD: \mathbb{R}^{d_h} \to [0, 1]^{d_v} as y^T+N=D(h^T+NL)\hat{y}_{T+N} = D(\hat{h}_{T+N}^L). Reusing the pretrained decoder head slightly improves prediction at layers ll close to the final layer LL relative to direct vocabulary regression.

  5. Knowl 5 — Problem Formulation for Subsequent-Token Extraction from Hidden States

    definition

    Let an autoregressive language model G:X→YG: \mathcal{X} \to \mathcal{Y} over vocabulary VV of size ∣V∣=dv|V| = d_v process an input sequence of tokens x=[x1,…,xT]∈Xx = [x_1, \dots, x_T] \in \mathcal{X} with xi∈Vx_i \in V. The model decomposes across LL layers as:

    G(x)=D(bL(…(b2(b1(E(x))))… ))G(x) = D(b_L(\dots(b_2(b_1(E(x))))\dots))

    where E:V→RdhE: V \to \mathbb{R}^{d_h} embeds tokens into H0=(h10,…,hT0)H^0 = (h_1^0, \dots, h_T^0), each layer bl:Rdh×T→Rdh×Tb_l: \mathbb{R}^{d_h \times T} \to \mathbb{R}^{d_h \times T} computes Hl=(h1l,…,hTl)H^l = (h_1^l, \dots, h_T^l), and D:Rdh→[0,1]dvD: \mathbb{R}^{d_h} \to [0, 1]^{d_v} is the pretrained decoder head mapping the final hidden state hTLh_T^L to next-token distribution yT=D(hTL)y_T = D(h_T^L).

    When greedy generation is iteratively unrolled for NN steps, subsequent token distributions are given by:

    yT+i=G([x1,…,xT+i]),xT+i+1=arg⁡max⁡yT+ifor i∈{0,1,…,N−1}y_{T+i} = G([x_1, \dots, x_{T+i}]), \quad x_{T+i+1} = \arg\max y_{T+i} \quad \text{for } i \in \{0, 1, \dots, N-1\}

    The task of multi-token anticipation from a single hidden state consists of predicting the full distribution yT+Ny_{T+N} (or argmax token xT+N+1x_{T+N+1}) for N≥1N \ge 1 using only the information contained in a single internal vector hTlh_T^l at a single layer l≤Ll \le L and position TT.

  6. Knowl 6 — Future Token Decodability by Input Token Category

    data/table

    Performance of N=1N = 1 future token prediction from the hidden state at layer l=14l = 14 across different lexical properties of the context token xTx_T demonstrates that learned prompt causal intervention consistently surpasses linear probes, especially on longer tokens and lowercase words following whitespace.

    Context Token Property Linear: Vocab Linear: Hidden State Fixed Context Learned Context
    Lowercase No Space 21.7% 25.2% 9.2% 32.5%
    Lowercase With Space 26.4% 20.8% 19.2% 51.9%
    Uppercase No Space 29.2% 26.3% 0.0% 23.3%
    Uppercase With Space 26.3% 26.3% 10.5% 31.6%
    Token length <4< 4 26.5% 24.9% 21.8% 46.9%
    Token length ≥4\ge 4 23.9% 24.4% 18.0% 52.1%
    Punctuation 28.7% 28.7% 16.6% 47.8%
    Numerical 12.5% 16.7% 20.8% 33.3%

    The largest performance gains of the learned prompt model over linear probes occur for Lowercase With Space (51.9% vs. 20.8%--26.4%) and Token length >= 4 (52.1% vs. 23.9%--24.4%). This suggests that complex multi-token completions for full words rely on computations distributed across the pretrained network parameters rather than representations that can be extracted via direct linear readouts.

  7. Knowl 7 — Confidence Calibration and Generality of Future Token Encodings

    empirical result

    Decoding future tokens from single hidden states exhibits strong confidence calibration and broad generality across text types:

    1. Calibration with Model Confidence: Probing accuracy for future token prediction (N=1N = 1) via learned prompt intervention scales monotonically with the base language model's output confidence on its immediate prediction: accuracy is 26% for model confidence in [0,30%)[0, 30\%), 57% for [30,60%)[30, 60\%), 77% for [60,90%)[60, 90\%), and 95% for [90,100%][90, 100\%]. Similar trends hold for N=2N = 2 and N=3N = 3.
    2. Generality Beyond Named Entities: Evaluating learned prompt probing specifically on multi-token named entities yields accuracies of 44%, 42%, and 37% for N=1,2,3N = 1, 2, 3, which is comparable to or slightly below performance on general text distributions. Future token encoding is therefore a general phenomenon in the model rather than an artifact restricted to memorized multi-token entity names.
    3. Objective Generalization: Optimizing soft prompts specifically on the N=1N = 1 objective generalizes robustly to decoding farther horizons (N=2,3N = 2, 3), whereas optimizing soft prompts directly on N=2N = 2 or N=3N = 3 fails to transfer back to N=1N = 1.
  8. Knowl 8 — Experimental Setup for Multi-Token Probing in GPT-J-6B

    experimental setup

    Multi-token lookahead experiments are conducted using GPT-J-6B (L=28L = 28 transformer layers, hidden dimension dh=4096d_h = 4096, vocabulary size ∣V∣=50,400|V| = 50,400) on text drawn from The Pile dataset:

    • Training Sets: Linear probing models are trained on 100,000 sampled token positions with an average context length of 518 tokens. Learned continuous soft prompts are trained on a subset of 10,000 token samples from this pool.
    • Test Set: Evaluation is performed on 1,000 sampled token locations with an average context length of 535 tokens where GPT-J-6B correctly predicts the immediate next token (N=0N = 0).
    • Evaluation Metrics:
      • Precision@k\text{Precision@}k: Evaluates whether the probe's top predicted token matches one of the top-kk tokens generated by the base model unrolled to position T+NT+N.
      • Surprisal\text{Surprisal}: Negative log probability −log⁡pG(xprobe)-\log p_{G}(x_{\text{probe}}) assigned by the base model to the top token predicted by the probing method.
  9. Knowl 9 — Future Lens Probing Visualization

    model/method

    The Future Lens is an interpretability visualization tool built upon learned prompt causal intervention. Given an arbitrary input sequence processed by a transformer, the tool extracts the hidden state htlh_t^l at every layer l∈{1,…,L}l \in \{1, \dots, L\} and every token position tt, transplants it into a trained soft prompt, and decodes the anticipated sequence of subsequent tokens (e.g., up to 4 future tokens).

    Visualizing the resulting grid (l,t)(l, t) with cell shading mapped to average output confidence reveals how future concept representations evolve across layers. For example, on the prompt prefix Marty McFly from, early layers initially predict generic geographic tokens (such as Australia or Boston), while middle layers around layer 6 transition to predicting movie-related phrases (Back to the Future), demonstrating the layer-wise semantic refinement of future context.

  10. Knowl 10 — Limitations of Multi-Token Representation Probing

    limitation

    The empirical findings on multi-token hidden state decoding carry several specific limitations:

    1. Model Scope: Experiments are restricted to a single autoregressive architecture, GPT-J-6B; evaluation on other LLM families and scale regimes remains unverified.
    2. Data Scope: Probing training is conducted on a 100,000-token English subset of The Pile, which is small relative to full pretraining datasets.
    3. Lookahead Horizon: Evaluations examine lookahead horizons up to 4 tokens ahead (N∈{0,1,2,3}N \in \{0, 1, 2, 3\}), leaving the maximum horizon of decodable information in a single hidden state uncharacterized.

Coverage note — None was omitted; all primary methods (linear models, fixed and learned prompt interventions), experimental metrics, empirical findings (layer-wise dynamics, token property breakdowns, confidence calibration), visualization tools, and stated limitations are fully represented.

References

  1. 1.Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023. Eliciting latent predictions from transformers with the tuned lens.
  2. 2.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023. Quantifying memorization across neural language models. In The Eleventh International Conference on Learning Representations.
  3. 3.Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019. The secret sharer: Evaluating and testing unintended memorization in neural networks. In Proceedings of the 28th USENIX Conference on Security Symposium, SEC’19, page 267–284, USA. USENIX Association.
  4. 4.Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, and Colin Raffel. 2021. Extracting training data from large language models. In USENIX Security Symposium.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: pre-training of deep bidirectional transformers for language understanding. CoRR, abs/1810.04805.
  6. 6.Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva. 2023. Jump to conclusions: Shortcutting transformers with linear transformations.
  7. 7.Jeffrey L. Elman. 1990. Finding structure in time. Cognitive Science, 14(2):179–211.
  8. 8.Vitaly Feldman and Chiyuan Zhang. 2020. What neural networks memorize and why: Discovering the long tail via influence estimation. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20, Red Hook, NY, USA. Curran Associates Inc.
  9. 9.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The pile: An 800gb dataset of diverse text for language modeling.
  10. 10.Mor Geva, Avi Caciularu, Guy Dar, Paul Roit, Shoval Sadde, Micah Shlain, Bar Tamir, and Yoav Goldberg. 2022. LM-debugger: An interactive tool for inspection and intervention in transformer-based language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 12–21, Abu Dhabi, UAE. Association for Computational Linguistics.
  11. 11.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484–5495, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  12. 12.Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023. Finding neurons in a haystack: Case studies with sparse probing.
  13. 13.Adi Haviv, Ido Cohen, Jacob Gidron, Roei Schuster, Yoav Goldberg, and Mor Geva. 2023. Understanding transformer memorization recall through idioms. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 248–264, Dubrovnik, Croatia. Association for Computational Linguistics.
  14. 14.Daphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A. Choquette-Choo, and Nicholas Carlini. 2023. Preventing verbatim memorization in language models gives a false sense of privacy.
  15. 15.Michael I Jordan. 1997. Serial order: A parallel distributed processing approach. In Advances in psychology, volume 121, pages 471–495. Elsevier.
  16. 16.Shahar Katz and Yonatan Belinkov. 2023. Interpreting transformer’s attention dynamic memory and visualizing the semantic information flow of gpt.
  17. 17.Jun Kong, Jin Wang, Liang-Chih Yu, and Xuejie Zhang. 2022. Accelerating inference for pretrained language models by unified multi-perspective early exiting. In Proceedings of the 29th International Conference on Computational Linguistics, pages 4677–4686, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  18. 18.Eric Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, and Byron C Wallace. 2021. Does bert pretrained on clinical notes reveal sensitive data? arXiv preprint arXiv:2104.07762.
  19. 19.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190.
  20. 20.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022a. Locating and editing factual associations in GPT. Advances in Neural Information Processing Systems, 36.
  21. 21.Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2022b. Mass editing memory in a transformer. arXiv preprint arXiv:2210.07229.
  22. 22.nostalgebraist. 2020. interpreting gpt: the logit lens.
  23. 23.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  24. 24.Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, and Donald Metzler. 2022. Confident adaptive language modeling. In Advances in Neural Information Processing Systems.
  25. 25.Yixuan Su, Deng Cai, Yan Wang, David Vandyke, Simon Baker, Piji Li, and Nigel Collier. 2021. Non-autoregressive text generation with pre-trained language models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 234–243, Online. Association for Computational Linguistics.
  26. 26.Jiuding Sun, Chantal Shaib, and Byron C Wallace. 2023. Evaluating the zero-shot robustness of instruction-tuned language models. arXiv preprint arXiv:2306.11270.
  27. 27.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  28. 28.Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019. Universal adversarial triggers for attacking and analyzing nlp. arXiv preprint arXiv:1908.07125.
  29. 29.Ben Wang and Aran Komatsuzaki. 2021. Gpt-j-6b: A 6 billion parameter autoregressive language model.
  30. 30.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. Emergent abilities of large language models. Transactions on Machine Learning Research. Survey Certification.
  31. 31.Yisheng Xiao, Lijun Wu, Junliang Guo, Juntao Li, Min Zhang, Tao Qin, and Tie yan Liu. 2023. A survey on non-autoregressive generation for neural machine translation and beyond.
  32. 32.Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. 2021. BERxiT: Early exiting for BERT with better fine-tuning and extension to regression. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 91–104, Online. Association for Computational Linguistics.
  33. 33.Minjia Zhang and Yuxiong He. 2020. Accelerating training of transformer-based language models with progressive layer dropping.

Citation

MLA
Pal, K., et al. “Future Lens: Anticipating Subsequent Tokens from a Single Hidden State”. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), 2023, pp. 548–60, https://doi.org/10.18653/v1/2023.conll-1.37.
APA
Pal, K., Sun, J., Yuan, A., Wallace, B. C., & Bau, D. (2023). Future Lens: Anticipating Subsequent Tokens from a Single Hidden State. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), 548–560. https://doi.org/10.18653/v1/2023.conll-1.37
Chicago
Pal, K., J. Sun, A. Yuan, B. C. Wallace, and D. Bau. 2023. “Future Lens: Anticipating Subsequent Tokens from a Single Hidden State”. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), 548–60. https://doi.org/10.18653/v1/2023.conll-1.37.
Harvard
Pal, K. et al. (2023) “Future Lens: Anticipating Subsequent Tokens from a Single Hidden State”, Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). Association for Computational Linguistics, pp. 548–560. Available at: https://doi.org/10.18653/v1/2023.conll-1.37.
Vancouver
1. Pal K, Sun J, Yuan A, Wallace BC, Bau D (2023) Future Lens: Anticipating Subsequent Tokens from a Single Hidden State. In: Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). Association for Computational Linguistics, pp 548–560

BibTeX

@inproceedings{pal-etal-2023-future,
    title = "Future Lens: Anticipating Subsequent Tokens from a Single Hidden State",
    author = "Pal, Koyena  and
      Sun, Jiuding  and
      Yuan, Andrew  and
      Wallace, Byron  and
      Bau, David",
    editor = "Jiang, Jing  and
      Reitter, David  and
      Deng, Shumin",
    booktitle = "Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL)",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.conll-1.37/",
    doi = "10.18653/v1/2023.conll-1.37",
    pages = "548--560"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/