Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers

Sotiris AnagnostidisDario PavlloLuca BiggioLorenzo NociAurélien LucchiThomas Hofmann

article2023NeurIPS79 citations

Proposes a learnable context-pruning mechanism for autoregressive language models that discards up to 80% of uninformative tokens during generation to cut inference memory and double throughput without hurting downstream accuracy.

Listen

Deploying modern autoregressive Large Language Models (LLMs) at scale creates severe memory and computational bottlenecks. Because standard self-attention mechanisms evaluate all pairs of tokens across an entire sequence, resource demands scale quadratically with context length. While existing mitigation strategies compress context windows or enforce static attention patterns, they often sacrifice relevant information regardless of context content. The article evaluates a novel dynamic context pruning technique—termed Adaptively Sparse Attention—that enables transformer models to learn which uninformative tokens to permanently drop during generation, thereby optimizing memory usage and processing speed without sacrificing core task performance.

The researchers integrated a lightweight, learnable pruning mechanism into pre-trained transformer architectures. The method evaluates token relevance at each layer using a sparse sigmoid function and a regularized training objective, discarding unnecessary cached representations permanently from memory. To support practical deployment, the authors developed a specialized batched memory data structure that dynamically recycles the memory slots of pruned tokens into contiguous blocks. The approach was systematically benchmarked on standard GPT-2 models (ranging from 124 million to 1.5 billion parameters) across standard language modeling datasets, zero-shot benchmarks (such as HellaSwag, PIQA, and WinoGrande), and preliminary tests on newer Pythia models.

The findings show that up to 80% of the input context can be pruned without significant degradation in model perplexity or zero-shot downstream task performance. Pruning drastically cuts Key-Value cache memory requirements, enabling up to 2x larger batch sizes or substantially longer input contexts on identical hardware. For long context sequences (1,000 tokens), this memory reduction translates to a 50% decrease in per-step latency and up to a 189% increase in overall token throughput compared to dense baselines. In addition, the pruning behavior offers enhanced interpretability: the models dynamically learned to prune less informative words primarily at sentence boundaries (punctuation) and successfully cleared irrelevant history across abrupt topic switches.

These results demonstrate that inference in autoregressive language models is predominantly memory-bound rather than compute-bound, meaning that context pruning offers high-leverage practical efficiency. By freeing up memory, organizations can serve higher query volumes on existing infrastructure, lowering deployment costs and hardware constraints. Crucially, the pruning framework functions orthogonally to other popular optimization techniques, such as weight quantization and weight pruning, allowing multiple efficiency methods to be stacked together.

Organizations deploying transformer models should consider context pruning as a fine-tuning enhancement to reduce operational inference costs on long-form generation tasks. Future work should pilot this mechanism on larger contemporary foundation models using parameter-efficient fine-tuning methods like LoRA. While empirical confidence in the reported gains is high across the tested architectures, stakeholders should note that the primary evaluations were conducted on GPT-2 models up to context lengths of 1,024 to 2,048 tokens; validation on newer, ultra-large language models with much longer contexts remains an important next step.

arXiv: 2305.15805
Cover for Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers

Abstract

Autoregressive Transformers adopted in Large Language Models (LLMs) are hard to scale to long sequences. Despite several works trying to reduce their computational cost, most of LLMs still adopt attention layers between all pairs of tokens in the sequence, thus incurring a quadratic cost. In this study, we present a novel approach that dynamically prunes contextual information while preserving the model’s expressiveness, resulting in reduced memory and computational requirements during inference. Our method employs a learnable mechanism that determines which uninformative tokens can be dropped from the context at any point across the generation process. By doing so, our approach not only addresses performance concerns but also enhances interpretability, providing valuable insight into the model’s decision-making process. Our technique can be applied to existing pre-trained models through a straightforward fine-tuning process, and the pruning strength can be specified by a sparsity parameter. Notably, our empirical findings demonstrate that we can effectively prune up to 80% of the context without significant performance degradation on downstream tasks, offering a valuable tool for mitigating inference costs. Our reference implementation achieves up to 2× increase in inference throughput and even greater memory savings.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Adaptively Sparse Attention
  • 3.2 Sparse Sigmoid
  • 3.3 Regularized Objective
  • 4 Experiments
  • 4.1 Results
  • 5 Discussion
  • References
  • A Experimental Setup
  • B Training Results and Ablations
  • C Additional Results
  • D Discussion

Knowls

  1. Knowl 1 — Adaptively sparse attention with irreversible, layer-specific context pruning

    model/method

    The method adds a learned pruning mask to each causal self-attention layer of a decoder-only Transformer. Let Xℓ∈Rn×dX^\ell\in\mathbb{R}^{n\times d} be the layer-ℓ\ell token representations for a sequence of length nn, let rr be the interaction dimension, and let WQintℓ,WKintℓ∈Rd×rW_{Q_{\mathrm{int}}}^\ell,W_{K_{\mathrm{int}}}^\ell\in\mathbb{R}^{d\times r} be layer-specific projections. The interaction queries and keys are

    Qintℓ=XℓWQintℓ,Kintℓ=XℓWKintℓ.Q_{\mathrm{int}}^\ell=X^\ell W_{Q_{\mathrm{int}}}^\ell,\qquad K_{\mathrm{int}}^\ell=X^\ell W_{K_{\mathrm{int}}}^\ell.

    For token indices j<kj<k, define the local keep score ak,jℓa_{k,j}^\ell and cumulative interaction mask Ik,jℓI_{k,j}^\ell by

    ak,jℓ=σ ⁣((Qintℓ)k,:⊤(Kintℓ)j,:r+βℓ),Ik,jℓ=∏t=j+1kat,jℓ,a_{k,j}^\ell=\sigma\!\left(\frac{(Q_{\mathrm{int}}^\ell)_{k,:}^{\top}(K_{\mathrm{int}}^\ell)_{j,:}}{\sqrt r}+\beta^\ell\right),\qquad I_{k,j}^\ell=\prod_{t=j+1}^{k}a_{t,j}^\ell,

    where σ\sigma is the sparse sigmoid and βℓ∈R\beta^\ell\in\mathbb{R} is a learned layer bias. The remaining mask entries are Ik,kℓ=1I_{k,k}^\ell=1 and Ik,jℓ=0I_{k,j}^\ell=0 for j>kj>k, preserving causal masking and preventing a token from dropping itself. For a self-attention head with query, key, and value matrices Q,K,VQ,K,V and head dimension pp, the mask modifies attention as

    SA⁡(Q,K,V)=softmax⁡ ⁣(QK⊤p+log⁡Iℓ)V,\operatorname{SA}(Q,K,V)=\operatorname{softmax}\!\left(\frac{QK^{\top}}{\sqrt p}+\log I^\ell\right)V,

    where the logarithm and softmax are applied elementwise across attention targets and log⁡0\log 0 represents complete masking. Values between zero and one partially suppress an attended token, while zero removes it. Because Ik,jℓI_{k,j}^\ell is a cumulative product, once token jj is dropped it remains unavailable to all later tokens at that layer. Different layers make their pruning decisions independently. Computing the interaction projections and pruning logic costs O(ndr+n2r)O(n d r+n^2r), which is below the self-attention cost when r<dr<d.

  2. Knowl 2 — Sparse sigmoid for trainable near-binary pruning

    definition

    The pruning mechanism uses an α\alpha-sigmoid rather than an ordinary sigmoid. For an input score x∈Rx\in\mathbb{R} and p∈[0,1]p\in[0,1], its output is

    σ(x)=α-sigmoid⁡(x)=arg⁡max⁡p∈[0,1](px+Hα(p)),\sigma(x)=\operatorname{\alpha\text{-}sigmoid}(x)=\arg\max_{p\in[0,1]}\left(px+H_\alpha(p)\right),

    with Tsallis entropy

    Hα(p)={p−pα+(1−p)−(1−p)αα(α−1),α≠1,−plog⁡p−(1−p)log⁡(1−p),α=1.H_\alpha(p)= \begin{cases} \dfrac{p-p^\alpha+(1-p)-(1-p)^\alpha}{\alpha(\alpha-1)}, & \alpha\ne 1,\\[6pt] -p\log p-(1-p)\log(1-p), & \alpha=1. \end{cases}

    Small α\alpha values provide smoother outputs and useful gradients early in fine-tuning, whereas increasing α\alpha makes the outputs increasingly sparse and close to binary decisions. In the experiments, α\alpha starts at 11 and is increased with a cosine schedule; values above 88 did not improve the final results. At inference time the learned function is replaced by a step function, corresponding to the limiting case α→∞\alpha\to\infty. The layer biases were initialized to βℓ=2.0\beta^\ell=2.0 so that pruning initially has a bias toward retaining context.

  3. Knowl 3 — Sparsity-regularized fine-tuning objective

    equation

    A pretrained autoregressive language model with parameters θ\theta is fine-tuned on a token sequence T=(t1,…,tn)T=(t_1,\ldots,t_n) using the usual language-modeling cross-entropy together with a penalty on retained context. Let fθ(T)f_\theta(T) denote the model predictions, let shift⁡(T)\operatorname{shift}(T) denote the next-token targets, let LL be the number of Transformer layers, and let Ii,jℓI_{i,j}^\ell be the adaptive mask value for j<ij<i at layer ℓ\ell. The training objective is

    L(θ,T)=Llm(θ,T)+Lsparsity(θ,T),\mathcal{L}(\theta,T)=\mathcal{L}_{\mathrm{lm}}(\theta,T)+\mathcal{L}_{\mathrm{sparsity}}(\theta,T),

    where

    Llm(θ,T)=CE⁡ ⁣(fθ(T),shift⁡(T))\mathcal{L}_{\mathrm{lm}}(\theta,T)=\operatorname{CE}\!\left(f_\theta(T),\operatorname{shift}(T)\right)

    and

    Lsparsity(θ,T)=2γLn(n−1)∑ℓ=1L∑i=1n∑j=1i−1Ii,jℓ,\mathcal{L}_{\mathrm{sparsity}}(\theta,T)=\frac{2\gamma}{L n(n-1)}\sum_{\ell=1}^{L}\sum_{i=1}^{n}\sum_{j=1}^{i-1}I_{i,j}^\ell,

    with regularization strength γ>0\gamma>0. The normalization averages the mask values over the Ln(n−1)/2L n(n-1)/2 potentially prunable causal interactions. Increasing γ\gamma encourages more context removal. For a current position ii, the paper defines sparsity as the fraction of preceding tokens that have been dropped, namely (number of dropped tokens among positions ≤i)/i(\text{number of dropped tokens among positions }\le i)/i.

  4. Knowl 4 — Packed batched key-value cache with recyclable slots

    algorithm

    The inference implementation maintains a separate dynamic cache for each Transformer layer and batch element. Each stored item contains a token identifier together with its cached key KK, value VV, and interaction key KintK_{\mathrm{int}}.

    The cache supports three batched operations:

    • push() inserts a newly generated token and its three cached tensors into the leftmost available memory slot.
    • get() returns the current keys, values, and interaction keys as a view of the underlying memory buffer, together with a binary mask indicating padding and any remaining gaps.
    • remove(mask) erases the entries selected by the binary mask after the current layer has made its pruning decisions.

    Removed tokens are not retained for later autoregressive steps. Their slots are recycled by subsequent tokens, because self-attention is invariant to the physical ordering of the key-value set. The buffer is dynamically resized as sequences grow and is consolidated whenever the effective load factor falls below 0.90.9, where the load factor is the ratio between the longest active sequence and buffer capacity. This keeps the storage sufficiently packed for contiguous GPU processing while supporting unequal prompt lengths, termination times, and pruning patterns across a batch. The implementation has O(n)O(n) push() and remove() operations and O(1)O(1) get(); a tested O(log⁡n)O(\log n) priority-queue alternative was slower in practice on the GPU.

  5. Knowl 5 — Fine-tuning configuration for GPT-2 evaluations

    experimental setup

    The experiments fine-tune pretrained Hugging Face GPT-2 models on subsets of English Wikipedia 20220301.en and BookCorpus. The evaluated models have the following configurations:

    Model Parameters Layers Heads Dimension dd
    GPT-2-small 124M 12 12 768
    GPT-2-medium 350M 24 16 1024
    GPT-2-large 774M 36 20 1280
    GPT-2-xl 1558M 48 25 1600

    All models support a maximum context of 1024 tokens and use vocabulary size 50,257. Fine-tuning runs for 25,000 steps with batch size 6, Adam optimization, no weight decay, and no learning-rate scheduler. The learning rate is 10−410^{-4} for GPT-2-small and GPT-2-medium and 5×10−55\times10^{-5} for GPT-2-large and GPT-2-xl. Adaptive sparse attention uses interaction dimension r=64r=64, the cosine schedule for α\alpha, and PyTorch scaled-dot-product FlashAttention. The dense comparison model is the same GPT-2 architecture fine-tuned without the additional interaction projections. Zero-shot evaluation uses WinoGrande, HellaSwag, PIQA, and LAMBADA, with sparsity averaged over the variable-length prefixes used by those datasets.

  6. Knowl 6 — High context sparsity preserves perplexity

    empirical result

    On GPT-2-small fine-tuned with the described objective, adaptive sparse attention removed up to approximately 80% of the context with little or no perplexity degradation. At context size 1000, a sparsity of 80.35% produced a perplexity that was 0.085 lower than the dense counterpart. Across contexts from 1 to 1024 tokens and at matched sparsity levels, the adaptive method consistently achieved lower perplexity than the local-attention and static sparse-attention baselines. Unlike fixed-window baselines, its achieved sparsity adapts to the current context length even when the regularization strength is fixed.

  7. Knowl 7 — Zero-shot capabilities remain largely intact under pruning

    empirical result

    The pruned GPT-2 models were evaluated without task-specific fine-tuning on WinoGrande, HellaSwag, PIQA, and LAMBADA. Mean zero-shot accuracy generally remained at, or in some cases above, the dense GPT-2 baseline as sparsity increased, including relatively high sparsity levels. The result indicates that the learned context-removal decisions can preserve broad downstream language-model capabilities rather than merely optimizing held-out language-model perplexity. Additional per-task analyses showed that the effect of sparsity is task-dependent, with tasks requiring longer-range interactions being more sensitive in some settings.

  8. Knowl 8 — Inference throughput improves as pruned contexts become memory-efficient

    empirical result

    On an NVIDIA RTX A5000, adaptive pruning initially increased latency for short contexts because of the extra interaction and deletion logic, but surpassed dense GPT-2 as context length grew. The main benefit came from reducing key-value-cache memory and enabling larger batch sizes, rather than from a large reduction in raw FLOPs. For a context size of 1000 tokens, GPT-2-small achieved a 98% additional throughput margin relative to the dense model with only a 0.316 perplexity loss, while GPT-2-medium achieved a 189% additional throughput margin with only a 0.084 perplexity loss. In the latter comparison, GPT-2-medium with γ=1.0\gamma=1.0 was faster than dense GPT-2-small while having 3.769 lower perplexity. The pruning implementation adds cached interaction keys and therefore does not make memory vanish completely, but total cache memory decreases approximately linearly with sparsity.

  9. Knowl 9 — Pruning substantially expands the feasible context window

    data/table

    The following measurements report the maximum sequence length that GPT-2-xl can process on a single NVIDIA RTX A5000 for different batch sizes. The baseline is dense GPT-2-xl; the other columns use adaptive sparse attention at the indicated sparsity. The values demonstrate that memory savings translate directly into substantially longer feasible contexts, especially at 90% sparsity.

    Batch size Baseline 50% sparsity 70% sparsity 90% sparsity
    16 1729 3312 5434 16221
    32 865 1616 2653 7901
    64 433 784 1274 3742
  10. Knowl 10 — Layerwise pruning exposes context-sensitive token importance

    empirical result

    The learned masks provide an interpretable view of which tokens remain available to later generation steps. In a GPT-2-small example trained with γ=0.3\gamma=0.3, pruning was triggered most often by punctuation and stop-word-like boundary tokens, consistent with the model discarding local sentence context after a clause or sentence ends. Token keep probabilities varied substantially by part of speech and by layer, and different layers formed different sparsity patterns rather than following a monotonic depth schedule.

    In a context-switch experiment, the input concatenated three unrelated passages about Lebanon, motorcycles, and Zorba’s Dance. The learned attention masks largely detected the transitions between these passages and suppressed preceding material from unrelated contexts. The resulting attention patterns contained three dense causal triangular submatrices, especially in layers 7–10, showing that the mechanism can isolate locally relevant context without imposing a fixed local-attention window.

Coverage note — Detailed ablations on propagating pruning decisions across depth, freezing the original model parameters, and varying the interaction dimension, together with the supplementary Pythia portability results, were omitted because they are secondary to the ten load-bearing method and main-evaluation knowls.

References

  1. 1.Lalit R Bahl, Frederick Jelinek, and Robert L Mercer. A maximum likelihood approach to continuous speech recognition. IEEE transactions on pattern analysis and machine intelligence, (2):179–190, 1983.
  2. 2.Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
  3. 3.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439, 2020.
  4. 4.Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token merging: Your vit but faster. arXiv preprint arXiv:2210.09461, 2022.
  5. 5.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019.
  6. 6.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, David Belanger, Lucy Colwell, and Adrian Weller. Masked language modeling for proteins via linearly scalable long-context transformers, 2020a.
  7. 7.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794, 2020b.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  9. 9.Zihang Dai, Guokun Lai, Yiming Yang, and Quoc Le. Funnel-transformer: Filtering out sequential redundancy for efficient language processing. Advances in neural information processing systems, 33:4271–4282, 2020.
  10. 10.Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 35:16344–16359, 2022.
  11. 11.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Llm. int8 (): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339, 2022.
  12. 12.Elias Frantar and Dan Alistarh. Massive language models can be accurately pruned in one-shot. arXiv preprint arXiv:2301.00774, 2023a.
  13. 13.Elias Frantar and Dan Alistarh. Sparsegpt: Massive language models can be accurately pruned in one-shot, 2023b.
  14. 14.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022.
  15. 15.Elias Frantar, Sidak Pal Singh, and Dan Alistarh. Optimal brain compression: A framework for accurate post-training quantization and pruning, 2023.
  16. 16.Yaru Hao, Li Dong, Furu Wei, and Ke Xu. Self-attention attribution: Interpreting information interactions inside transformer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 12963–12971, 2021.
  17. 17.Babak Hassibi, David G. Stork, and Gregory J. Wolff. Optimal brain surgeon and general network pruning. IEEE International Conference on Neural Networks, pages 293–299 vol.1, 1993.
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pages 1026–1034, 2015.
  19. 19.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  20. 20.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  21. 21.Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, Shigang Li, and Torsten Hoefler. Data movement is all you need: A case study on optimizing transformers. Proceedings of Machine Learning and Systems, 3:711–732, 2021.
  22. 22.Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. Perceiver: General perception with iterative attention, 2021.
  23. 23.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  24. 24.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning, pages 5156–5165. PMLR, 2020.
  25. 25.Sehoon Kim, Sheng Shen, David Thorsley, Amir Gholami, Woosuk Kwon, Joseph Hassoun, and Kurt Keutzer. Learned token pruning for transformers. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 784–794, 2022.
  26. 26.Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451, 2020.
  27. 27.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, et al. Openassistant conversations–democratizing large language model alignment. arXiv preprint arXiv:2304.07327, 2023.
  28. 28.Woosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami. A fast post-training pruning framework for transformers, 2022.
  29. 29.Heejun Lee, Minki Kang, Youngwan Lee, and Sung Ju Hwang. Sparse token transformer with attention back tracking. In The Eleventh International Conference on Learning Representations, 2023.
  30. 30.Juho Lee, Yoonho Lee, Jungtaek Kim, Adam R. Kosiorek, Seungjin Choi, and Yee Whye Teh. Set transformer: A framework for attention-based permutation-invariant neural networks, 2019.
  31. 31.Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu. A survey of transformers. AI Open, 2022.
  32. 32.André Martins, António Farinhas, Marcos Treviso, Vlad Niculae, Pedro Aguiar, and Mario Figueiredo. Sparse and continuous attention mechanisms. Advances in Neural Information Processing Systems, 33:20989–21001, 2020.
  33. 33.Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Pal Singh, and Aurelien Lucchi. Signal propagation in transformers: Theoretical perspectives and the role of rank collapse. arXiv preprint arXiv:2206.03126, 2022.
  34. 34.OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  35. 35.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. fairseq: A fast, extensible toolkit for sequence modeling. arXiv preprint arXiv:1904.01038, 2019.
  36. 36.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  37. 37.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031, 2016.
  38. 38.Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A. Smith, and Lingpeng Kong. Random feature attention, 2021.
  39. 39.Ben Peters, Vlad Niculae, and André FT Martins. Sparse sequence-to-sequence models. arXiv preprint arXiv:1905.05702, 2019.
  40. 40.Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. Efficiently scaling transformer inference. arXiv preprint arXiv:2211.05102, 2022.
  41. 41.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  42. 42.Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlovic, Geir Kjetil Sandve, et al. Hopfield networks ´ is all you need. arXiv preprint arXiv:2008.02217, 2020.
  43. 43.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  44. 44.Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. Linear transformers are secretly fast weight programmers, 2021.
  45. 45.Noam Shazeer. Fast transformer decoding: One write-head is all you need. arXiv preprint arXiv:1911.02150, 2019.
  46. 46.Han Shi, Jiahui Gao, Xiaozhe Ren, Hang Xu, Xiaodan Liang, Zhenguo Li, and James Tin-Yau Kwok. Sparsebert: Rethinking the importance analysis in self-attention. In International Conference on Machine Learning, pages 9547–9557. PMLR, 2021.
  47. 47.Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and policy considerations for deep learning in nlp. arXiv preprint arXiv:1906.02243, 2019.
  48. 48.Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, and Mohit Iyyer. Do long-range language models actually use long-range context? arXiv preprint arXiv:2109.09115, 2021.
  49. 49.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A survey.(2020). arXiv preprint cs.LG/2009.06732, 2020.
  50. 50.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  51. 51.Constantino Tsallis. Possible generalization of boltzmann-gibbs statistics. Journal of statistical physics, 52:479–487, 1988.
  52. 52.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  53. 53.Ashish Vaswani, Samy Bengio, Eugene Brevdo, Francois Chollet, Aidan N Gomez, Stephan Gouws, Llion Jones, Łukasz Kaiser, Nal Kalchbrenner, Niki Parmar, et al. Tensor2tensor for neural machine translation. arXiv preprint arXiv:1803.07416, 2018.
  54. 54.Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-attention with linear complexity, 2020.
  55. 55.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45, 2020.
  56. 56.Guangxuan Xiao, Ji Lin, Mickael Seznec, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. arXiv preprint arXiv:2211.10438, 2022.
  57. 57.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. On layer normalization in the transformer architecture. In International Conference on Machine Learning, pages 10524–10533. PMLR, 2020.
  58. 58.Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. Advances in Neural Information Processing Systems, 35:27168–27183, 2022.
  59. 59.Chulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar. O (n) connections are expressive enough: Universal approximability of sparse transformers. Advances in Neural Information Processing Systems, 33:13783–13794, 2020.
  60. 60.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. Big bird: Transformers for longer sequences. Advances in neural information processing systems, 33:17283–17297, 2020.
  61. 61.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  62. 62.Zhenhai Zhu and Radu Soricut. H-transformer-1d: Fast one-dimensional hierarchical attention for sequences. arXiv preprint arXiv:2107.11906, 2021.

Citation

MLA
Anagnostidis, S., et al. “Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 65202–23, https://proceedings.neurips.cc/paper_files/paper/2023/file/cdaac2a02c4fdcae77ba083b110efcc3-Paper-Conference.pdf.
APA
Anagnostidis, S., Pavllo, D., Biggio, L., Noci, L., Lucchi, A., & Hofmann, T. (2023). Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers. Advances in Neural Information Processing Systems, 36, 65202–65223. https://proceedings.neurips.cc/paper_files/paper/2023/file/cdaac2a02c4fdcae77ba083b110efcc3-Paper-Conference.pdf
Chicago
Anagnostidis, S., D. Pavllo, L. Biggio, L. Noci, A. Lucchi, and T. Hofmann. 2023. “Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers”. Advances in Neural Information Processing Systems 36: 65202–23. https://proceedings.neurips.cc/paper_files/paper/2023/file/cdaac2a02c4fdcae77ba083b110efcc3-Paper-Conference.pdf.
Harvard
Anagnostidis, S. et al. (2023) “Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 65202–65223. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/cdaac2a02c4fdcae77ba083b110efcc3-Paper-Conference.pdf.
Vancouver
1. Anagnostidis S, Pavllo D, Biggio L, Noci L, Lucchi A, Hofmann T (2023) Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 65202–65223

BibTeX

@inproceedings{anagnostidis2023dynamic,
  title = {Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers},
  author = {Anagnostidis, Sotiris and Pavllo, Dario and Biggio, Luca and Noci, Lorenzo and Lucchi, Aurelien and Hofmann, Thomas},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {65202-65223},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/cdaac2a02c4fdcae77ba083b110efcc3-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors