Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale
Hritik BansalKarthik GopalakrishnanSaket DingliwalSravan BodapatiKatrin KirchhoffDan Roth
Reveals that up to 70% of attention heads and 20% of feedforward networks in a 66-billion-parameter language model can be pruned without sacrificing in-context learning performance, demonstrating that in-context capabilities are concentrated in a small, shared subset of task-agnostic induction heads.
Large language models have rapidly grown in size, delivering impressive capabilities such as in-context learning, where a model solves new tasks from just a few examples without retraining. However, operating these massive architectures demands immense computational power, drives high financial costs, and leaves a substantial carbon footprint. As organizations scale up model sizes, a critical question arises: do these models genuinely require all of their billions of parameters to perform in-context tasks, or are large portions of their underlying structure underutilized?
The article aims to evaluate whether the capability to perform in-context learning is distributed across all components of a large language model or concentrated within a small subset. Specifically, it demonstrates how much of a massive 66-billion-parameter model can be removed without significantly degrading its ability to perform zero-shot and few-shot natural language processing tasks.
To investigate this, the analysis evaluates a 66-billion-parameter Open Pre-trained Transformer across 14 diverse language benchmarks spanning question answering, reading comprehension, and commonsense reasoning. The evaluation measures the importance of two primary structural components: multi-headed attention mechanisms, which manage interactions between words, and feed-forward networks, which process token representations. By systematically ranking these components through sensitivity scores and removing them in 10% increments, the article examines how component pruning affects accuracy across zero-shot, one-shot, and five-shot scenarios. It also analyzes model components using a task-independent mathematical framework that isolates fundamental pattern-matching and copying behaviors, known as induction operations.
The findings reveal that large language models contain substantial structural redundancy for in-context learning tasks. First, approximately 70% of the attention heads—representing nearly 15.7 billion parameters—can be removed with minimal decline in overall task performance. Second, feed-forward networks prove far more sensitive to removal; performance drops sharply after pruning just 10% to 20% of these networks (about 4.3 to 8.5 billion parameters), underscoring their critical role in task execution. Third, when removing both components simultaneously, the model maintains strong performance: pruning 60% of attention heads alongside 20% of feed-forward networks results in only a 4% to 5% absolute drop in average accuracy. Fourth, component importance is highly consistent across diverse tasks and prompt styles, showing statistically significant rank correlations. Finally, the attention heads identified as important overlap with the specialized heads responsible for primitive prefix matching and copying, demonstrating that a compact, universal core drives both basic and sophisticated reasoning behaviors.
These insights carry major operational and strategic implications. They suggest that current large models are substantially undertrained for in-context learning, meaning massive parameter counts are used inefficiently. For enterprise deployments, recognizing that over half of a model's attention capacity is non-essential opens up immediate avenues for aggressive model compression. Pruning redundant components can dramatically lower memory overhead, speed up inference response times, and curtail cloud hosting expenditures without compromising application quality.
Moving forward, technical leaders and practitioners should explore structured pruning and targeted model compression pipelines before deploying massive models into production environments. Additionally, researchers should investigate revised pre-training objectives that directly encourage induction and pattern-matching abilities, potentially yielding smaller, compute-optimal models that achieve emergent capabilities at lower parameter counts. Testing these pruning strategies on modern instruction-tuned model variants is recommended before full-scale adoption.
Decision-makers should interpret these results within certain boundaries. The evaluation was conducted primarily on a single 66-billion-parameter model architecture using English-only benchmarks and short prompt contexts of up to five examples. While the article establishes strong empirical evidence that a shared core of parameters handles in-context learning, it notes that the relationships are correlational rather than strictly causal. Nevertheless, the findings offer high confidence that current large models harbor significant structural redundancy that organizations can safely leverage for efficiency optimizations.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). GPT-3 establishes the zero-shot and few-shot in-context learning setup and scaling context that this study uses to investigate which model components support those abilities.
- Paper: Are Sixteen Heads Really Better than One?, Paul Michel et al. (2019). Its attention-head ablation and sensitivity-based pruning provide the direct methodological precedent for assessing how much attention structure the source model can lose.
- Paper: Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned, Elena Voita et al. (2019). This study identifies specialized, important attention heads and prunes the rest, grounding the source’s analysis of head importance and redundancy.
- Paper: A Fast Post-Training Pruning Framework for Transformers, Woosuk Kwon et al. (2022). Its post-training framework jointly evaluates attention heads and feed-forward filters, preparing readers for the source’s comparison of those components’ pruning sensitivity.
- Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). Chinchilla’s compute-optimal scaling results provide essential context for the source’s claim that large models may be undertrained relative to their parameter counts.
- Paper: On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model, Seongjin Shin et al. (2022). Its controlled study of how pretraining data affects few-shot in-context learning helps frame the source’s focus on the model’s internal components as another determinant of that capability.
- Paper: SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks, Jiwon Song et al. (2024). SLEB carries component-level redundancy analysis into training-free removal of whole Transformer blocks in models up to 70B, testing a broader structural route to efficient deployment.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). SparseGPT extends large-model compression from selectively pruning attention heads and feed-forward networks to one-shot pruning of model weights at 175B scale.
- Paper: Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language Models, Peijie Dong et al. (2024). Pruner-Zero continues post-training compression by automating the search for pruning metrics, extending the source’s component-importance analysis into weight-level pruning.
- Paper: ZipLM: Inference-Aware Structured Pruning of Language Models, Eldar Kurtic et al. (2023). ZipLM turns structural pruning into a hardware-aware optimization, extending the source’s evidence of removable model structure toward measured inference speedups.
- Paper: SliceGPT: Compress Large Language Models by Deleting Rows and Columns, Saleh Ashkboos et al. (2024). SliceGPT extends structured model reduction beyond ranking heads and feed-forward networks by shrinking the network’s embedding dimension through row and column deletion.
