Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale

Hritik BansalKarthik GopalakrishnanSaket DingliwalSravan BodapatiKatrin KirchhoffDan Roth

article2023ACL77 citations

Reveals that up to 70% of attention heads and 20% of feedforward networks in a 66-billion-parameter language model can be pruned without sacrificing in-context learning performance, demonstrating that in-context capabilities are concentrated in a small, shared subset of task-agnostic induction heads.

Listen

Large language models have rapidly grown in size, delivering impressive capabilities such as in-context learning, where a model solves new tasks from just a few examples without retraining. However, operating these massive architectures demands immense computational power, drives high financial costs, and leaves a substantial carbon footprint. As organizations scale up model sizes, a critical question arises: do these models genuinely require all of their billions of parameters to perform in-context tasks, or are large portions of their underlying structure underutilized?

The article aims to evaluate whether the capability to perform in-context learning is distributed across all components of a large language model or concentrated within a small subset. Specifically, it demonstrates how much of a massive 66-billion-parameter model can be removed without significantly degrading its ability to perform zero-shot and few-shot natural language processing tasks.

To investigate this, the analysis evaluates a 66-billion-parameter Open Pre-trained Transformer across 14 diverse language benchmarks spanning question answering, reading comprehension, and commonsense reasoning. The evaluation measures the importance of two primary structural components: multi-headed attention mechanisms, which manage interactions between words, and feed-forward networks, which process token representations. By systematically ranking these components through sensitivity scores and removing them in 10% increments, the article examines how component pruning affects accuracy across zero-shot, one-shot, and five-shot scenarios. It also analyzes model components using a task-independent mathematical framework that isolates fundamental pattern-matching and copying behaviors, known as induction operations.

The findings reveal that large language models contain substantial structural redundancy for in-context learning tasks. First, approximately 70% of the attention heads—representing nearly 15.7 billion parameters—can be removed with minimal decline in overall task performance. Second, feed-forward networks prove far more sensitive to removal; performance drops sharply after pruning just 10% to 20% of these networks (about 4.3 to 8.5 billion parameters), underscoring their critical role in task execution. Third, when removing both components simultaneously, the model maintains strong performance: pruning 60% of attention heads alongside 20% of feed-forward networks results in only a 4% to 5% absolute drop in average accuracy. Fourth, component importance is highly consistent across diverse tasks and prompt styles, showing statistically significant rank correlations. Finally, the attention heads identified as important overlap with the specialized heads responsible for primitive prefix matching and copying, demonstrating that a compact, universal core drives both basic and sophisticated reasoning behaviors.

These insights carry major operational and strategic implications. They suggest that current large models are substantially undertrained for in-context learning, meaning massive parameter counts are used inefficiently. For enterprise deployments, recognizing that over half of a model's attention capacity is non-essential opens up immediate avenues for aggressive model compression. Pruning redundant components can dramatically lower memory overhead, speed up inference response times, and curtail cloud hosting expenditures without compromising application quality.

Moving forward, technical leaders and practitioners should explore structured pruning and targeted model compression pipelines before deploying massive models into production environments. Additionally, researchers should investigate revised pre-training objectives that directly encourage induction and pattern-matching abilities, potentially yielding smaller, compute-optimal models that achieve emergent capabilities at lower parameter counts. Testing these pruning strategies on modern instruction-tuned model variants is recommended before full-scale adoption.

Decision-makers should interpret these results within certain boundaries. The evaluation was conducted primarily on a single 66-billion-parameter model architecture using English-only benchmarks and short prompt contexts of up to five examples. While the article establishes strong empirical evidence that a shared core of parameters handles in-context learning, it notes that the relationships are correlational rather than strictly causal. Nevertheless, the findings offer high confidence that current large models harbor significant structural redundancy that organizations can safely leverage for efficiency optimizations.

Cover for Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale

Abstract

Language models have been shown to perform better with an increase in scale on a wide variety of tasks via the in-context learning paradigm. In this paper, we investigate the hypothesis that the ability of a large language model to in-context learn-perform a task is not uniformly spread across all of its underlying components. Using a 66 billion parameter language model (OPT-66B) across a diverse set of 14 downstream tasks, we find this is indeed the case: ~70% of the attention heads and ~20% of the feed forward networks can be removed with minimal decline in task performance. We find substantial overlap in the set of attention heads (un)important for in-context learning across tasks and number of in-context examples. We also address our hypothesis through a task-agnostic lens, finding that a small set of attention heads in OPT-66B score highly on their ability to perform primitive induction operations associated with in-context learning, namely, prefix matching and copying. These induction heads overlap with task-specific important heads, reinforcing arguments by Olsson et al. (2022) regarding induction head generality to more sophisticated behaviors associated with in-context learning. Overall, our study provides several insights that indicate large language models may be under-trained for in-context learning and opens up questions on how to pre-train language models to more effectively perform in-context learning.

Table of Contents

  • 1 Introduction
  • 2 Background & Methods
  • 2.1 Open Pre-trained Transformer (OPT)
  • 2.2 In-Context Learning & Induction Heads
  • 2.3 Importance Scores
  • 2.3.1 Oracle
  • 2.3.2 Gradient-based
  • 3 Experimental Setup
  • 4 Importance Scores for OPT-66B
  • 4.1 Attention Heads
  • 4.2 Feed Forward Networks
  • 5 Iterative Pruning
  • 5.1 Removing Attention Heads
  • 5.2 Removing FFNs
  • 5.3 Combined Removal of Heads & FFNs
  • 6 Detailed Analysis of Attention Heads
  • 6.1 Cross-Task Analysis
  • 6.1.1 Spearman's Rank Correlation
  • 6.1.2 Generalization Trends
  • 6.2 Cross-Shot Analysis
  • 6.3 Induction Heads in OPT-66B
  • 6.3.1 Are Induction Heads Important?
  • 7 Related Work
  • 8 Conclusion & Future Work
  • 9 Limitations
  • 10 Impact Statement
  • References
  • A Appendix
  • A.1 Head Importance Scores
  • A.2 FFN Importance Scores
  • A.3 Removing Attention Heads
  • A.4 Removing FFNs
  • A.5 Combined Removal of Heads & FFNs
  • A.6 Cross-Task Analysis: Spearman's Rank Correlation
  • A.7 Cross-Task Analysis: Generalization Trends
  • A.8 Details of Prefix Matching and Copying Scores
  • A.9 Importance of Induction Heads to Each Task

Knowls

  1. Knowl 1 — Joint pruning leaves most in-context performance intact

    empirical result

    In OPT-66B, iterative pruning based on shot-specific, task-aggregated importance rankings showed that attention heads and feed-forward networks (FFNs) can be removed together while retaining most average accuracy across 14 NLP tasks. In the 0-shot setting, removing 70% of attention heads (about 15.7B parameters) and 20% of FFNs (about 8.5B parameters) reduced average accuracy by only 5 absolute percentage points. In the 1-shot setting, removing 70% of heads and 10% of FFNs reduced average accuracy by 6 points; in the 5-shot setting, removing 60% of heads and 20% of FFNs reduced it by 4 points. The pruning rates at which accuracy began to decline differed somewhat from the rates observed when pruning each component type separately.

  2. Knowl 2 — Most attention heads can be pruned before accuracy falls sharply

    empirical result

    Across 0-, 1-, and 5-shot in-context learning on 14 NLP tasks, OPT-66B's average accuracy remained fairly stable as attention heads were iteratively removed in order of increasing task- and shot-specific importance. The decline became pronounced only after roughly 70% of the 4,608 heads had been pruned. Most individual tasks followed a similar trend, although in the 0-shot setting CB and WSC accuracy could increase after 70% of heads were removed. These results indicate that many heads contribute little to the measured in-context performance under the tested conditions.

  3. Knowl 3 — FFNs are less redundant than attention heads for in-context learning

    empirical result

    In OPT-66B, removing the least important FFNs had a more limited tolerance than removing attention heads. In the 0-shot setting, average accuracy across 14 tasks changed little until about 20% of the 64 FFNs had been removed. PIQA, Winogrande, and RTE retained their accuracy even with 30% of FFNs removed, corresponding to about 13B parameters. In the 1- and 5-shot settings, the sharp accuracy decline began at approximately 10% FFN removal. This contrast with the higher head-pruning tolerance supports the paper's conclusion that FFNs play an important role in in-context learning.

  4. Knowl 4 — Attention-head importance rankings are positively related across tasks

    empirical result

    For OPT-66B, attention-head importance rankings showed statistically significant positive Spearman rank correlations (p<0.01p<0.01) between every pair of the 14 evaluated tasks, and between each task ranking and the task-aggregate ranking, in the zero- and few-shot settings. In the 5-shot results, pairwise correlations ranged from 0.15 to 0.47; ReCoRD generally had lower correlations with other tasks than they had with one another. Thus, task rankings are not identical, but heads judged relatively important or unimportant tend to cluster across tasks. When applied to unseen tasks, the aggregate ranking transferred well to MathQA, where its accuracy was within 1–2 points of pruning using MathQA's own ranking, but transferred less well to LAMBADA.

  5. Knowl 5 — Task-important heads overlap with heads scoring for induction operations

    empirical result

    In OPT-66B, the attention heads with high scores for both prefix matching and copying—primitive operations in which a model attends to a previous occurrence of a token and reproduces its following token—formed a sparse subset and overlapped with heads important for downstream in-context learning. When heads were pruned from least to most important according to the task-aggregate rankings, much of the total prefix-matching score remained after 20% were removed; its decline became steeper after about 40% pruning. By contrast, total copying score declined rapidly and consistently as heads were pruned, including in zero- and few-shot settings. The observations support an association between induction capacity and downstream task importance, but do not establish that induction capacity causes the more sophisticated task behavior.

  6. Knowl 6 — Important heads and FFNs occupy different layer regions

    empirical result

    Importance patterns in OPT-66B differed by component type. Attention heads with high task-specific or task-averaged importance scores were concentrated mainly in intermediate layers. For FFNs, removing any network in layers 1–30 usually had comparable or beneficial effects on most tasks in the 0- and 1-shot settings; in the 5-shot setting, important FFNs appeared in both early and later layers. Later-layer FFN importance also varied substantially across tasks: individually removing some FFNs changed WSC or MultiRC accuracy by as much as 20 absolute points in either direction.

  7. Knowl 7 — Head importance rankings overlap across numbers of in-context examples

    empirical result

    Across the 14 evaluated tasks in OPT-66B, the mean Spearman rank correlation between attention-head importance rankings was 0.41 for 1-shot versus 5-shot, 0.39 for 0-shot versus 1-shot, and 0.37 for 0-shot versus 5-shot. The reported variance across tasks was 0.001, and the correlations were statistically significant (p<0.01p<0.01). These results show non-trivial overlap in which heads are relatively important or unimportant across shot counts, with somewhat greater agreement between the two few-shot settings.

  8. Knowl 8 — Component importance scores guide the pruning experiments

    model/method

    The study evaluated OPT-66B, a 64-layer decoder with 72 attention heads per layer (4,608 heads total) and 64 FFNs, on accuracy for 14 tasks: ARC Easy, ARC Challenge, OpenBookQA, HellaSwag, PIQA, Winogrande, BoolQ, CB, COPA, MultiRC, ReCoRD, RTE, WiC, and WSC. Evaluations used 0, 1, or 5 randomly sampled in-context examples. For an attention head hh, the gradient-based importance score on dataset DD was the expected absolute inner product between its activation and the gradient of the autoregressive target loss with respect to that activation. Here xx is the prompt, y=(y_1,ldots,y_{T_y}) is the target sequence, Ah([x;y])A^h([x;y]) is the head's output on their concatenation, and TyT_y is the number of target tokens; the loss is mean negative log-likelihood:

    Ih(D)=E(x,y)∼D[∣Ah([x;y])⊤∂L(y∣x)∂Ah([x;y])∣],L(y∣x)=−1Ty∑j=1Tylog⁡p(yj∣x,y1:j−1).I_h(D)=\mathbb{E}_{(x,y)\sim D}\left[\left|A^h([x;y])^\top\frac{\partial\mathcal{L}(y\mid x)}{\partial A^h([x;y])}\right|\right],\qquad \mathcal{L}(y\mid x)=-\frac{1}{T_y}\sum_{j=1}^{T_y}\log p(y_j\mid x,y_{1:j-1}).

    For a component CC and performance metric PMP_M of model MM, the oracle score was the metric difference before and after removing that component: ISC(D)=PM(D)−PM∖C(D)IS_C(D)=P_M(D)-P_{M\setminus C}(D). The study used gradient scores for heads and oracle scores for FFNs. Components were ranked from least to most important separately by task and shot setting, pruned in 10% increments, and evaluated after each increment. Combined pruning used shot-specific task-aggregate importance rankings.

  9. Knowl 9 — Induction capacity was measured on repeated random-token sequences

    model/method

    To score prefix matching and copying in OPT-66B without relying on downstream task labels, the authors evaluated each attention head on 100 randomly generated sequences of unique tokens with varying lengths. They first excluded the 4% most frequent and 4% least frequent vocabulary items. For prefix matching, each sequence was repeated four times; the score measured attention from a repeated token to the token immediately following an earlier occurrence of that same token, averaged over eligible positions. For copying, each head's contribution to vocabulary predictions was measured by whether it raised the score of the maximally attended preceding token relative to other attendable input tokens, using a normalized positive-score comparison. Prefix-matching sequence lengths before repetition ranged from 25 to 223 tokens; copying sequence lengths ranged from 100 to 892 tokens. Scores were averaged across the 100 sequences.

  10. Knowl 10 — Study conclusions are limited to the tested model and prompting regime

    limitation

    The study does not establish a causal link between a head's ability to perform induction operations and its importance for downstream in-context learning. Its experiments used one model, OPT-66B, at most five in-context examples, randomly selected examples, and monolingual downstream tasks. The authors also report that they do not yet explain why many attention heads appear unimportant or why importance rankings overlap across tasks and shot settings.

Coverage note — Detailed per-task heatmaps and pruning curves are omitted because their aggregate findings and notable exceptions are captured in the knowls above.

References

  1. 1.Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. arXiv preprint arXiv:1905.13319.
  2. 2.Sajid Anwar, Kyuyeon Hwang, and Wonyong Sung. 2017. Structured pruning of deep convolutional neural networks. ACM Journal on Emerging Technologies in Computing Systems (JETC), 13(3):1–18.
  3. 3.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439.
  4. 4.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. 2022. Gpt-neox-20b: An open-source autoregressive language model. arXiv preprint arXiv:2204.06745.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared J Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  6. 6.Stephanie CY Chan, Ishita Dasgupta, Junkyung Kim, Dharshan Kumaran, Andrew K Lampinen, and Felix Hill. 2022. Transformers generalize differently from information stored in context vs in weights. arXiv preprint arXiv:2210.05675.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  8. 8.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  9. 9.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread.
  10. 10.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPoFi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. A framework for few-shot language model evaluation.
  11. 11.Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant. 2022. What can transformers learn in-context? a case study of simple function classes. arXiv preprint arXiv:2208.01066.
  12. 12.Marwa El Halabi, Suraj Srinivas, and Simon Lacoste-Julien. 2022. Data-efficient structured pruning via submodular optimization. arXiv preprint arXiv:2203.04940.
  13. 13.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556.
  14. 14.Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn’t always right. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  15. 15.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  16. 16.Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. 2016. Pruning filters for efficient convnets. arXiv preprint arXiv:1608.08710.
  17. 17.Jiaoda Li, Ryan Cotterell, and Mrinmaya Sachan. 2021. Differentiable subset pruning of transformer heads. Transactions of the Association for Computational Linguistics, 9:1442–1459.
  18. 18.Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham. 2021. Jurassic-1: Technical details and evaluation. White Paper. AI21 Labs.
  19. 19.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, Dublin, Ireland and Online. Association for Computational Linguistics.
  20. 20.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  21. 21.Paul Michel, Omer Levy, and Graham Neubig. 2019. Are sixteen heads really better than one? Advances in neural information processing systems, 32.
  22. 22.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789.
  23. 23.Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022a. Noisy channel language model prompting for few-shot text classification. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5316–5330, Dublin, Ireland. Association for Computational Linguistics.
  24. 24.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022b. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  25. 25.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2022. Reframing instructional prompts to GPTk’s language. In Findings of the Association for Computational Linguistics: ACL 2022, pages 589–612, Dublin, Ireland. Association for Computational Linguistics.
  26. 26.Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. 2016. Pruning convolutional neural networks for resource efficient inference. arXiv preprint arXiv:1611.06440.
  27. 27.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022. In-context learning and induction heads. Transformer Circuits Thread.
  28. 28.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031.
  29. 29.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  30. 30.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022. Impact of pretraining term frequencies on few-shot reasoning. arXiv preprint arXiv:2202.07206.
  31. 31.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2671, Seattle, United States. Association for Computational Linguistics.
  32. 32.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106.
  33. 33.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990.
  34. 34.Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4593–4601, Florence, Italy. Association for Computational Linguistics.
  35. 35.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  36. 36.Jesse Vig and Yonatan Belinkov. 2019. Analyzing the structure of attention in a transformer language model. arXiv preprint arXiv:1906.04284.
  37. 37.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv preprint arXiv:1905.09418.
  38. 38.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems, 32.
  39. 39.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  40. 40.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2021. An explanation of in-context learning as implicit bayesian inference. arXiv preprint arXiv:2111.02080.
  41. 41.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830.
  42. 42.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  43. 43.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pages 12697–12706. PMLR.

Citation

MLA
Bansal, H., et al. “Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 11833–56, https://doi.org/10.18653/v1/2023.acl-long.660.
APA
Bansal, H., Gopalakrishnan, K., Dingliwal, S., Bodapati, S., Kirchhoff, K., & Roth, D. (2023). Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11833–11856. https://doi.org/10.18653/v1/2023.acl-long.660
Chicago
Bansal, H., K. Gopalakrishnan, S. Dingliwal, S. Bodapati, K. Kirchhoff, and D. Roth. 2023. “Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11833–56. https://doi.org/10.18653/v1/2023.acl-long.660.
Harvard
Bansal, H. et al. (2023) “Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11833–11856. Available at: https://doi.org/10.18653/v1/2023.acl-long.660.
Vancouver
1. Bansal H, Gopalakrishnan K, Dingliwal S, Bodapati S, Kirchhoff K, Roth D (2023) Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 11833–11856

BibTeX

@inproceedings{bansal-etal-2023-rethinking,
    title = "Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale",
    author = "Bansal, Hritik  and
      Gopalakrishnan, Karthik  and
      Dingliwal, Saket  and
      Bodapati, Sravan  and
      Kirchhoff, Katrin  and
      Roth, Dan",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.660/",
    doi = "10.18653/v1/2023.acl-long.660",
    pages = "11833--11856"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/