In-Context Learning with Long-Context Models: An In-Depth Exploration
Amanda BertschMaor IvgiEmily XiaoUri AlonJonathan BerantMatthew R. GormleyGraham Neubig
Demonstrates that scaling in-context learning to thousands of demonstrations can rival model finetuning while significantly reducing sensitivity to example ordering and selection.
Recent advances in artificial intelligence have expanded the context windows of large language models, allowing them to process tens of thousands of words at once. Traditionally, adapting language models to specific tasks required either providing a handful of examples directly in the prompt—known as in-context learning—or computationally expensive finetuning to update the underlying model weights. With larger context capacities, organizations can now fit thousands of demonstration examples directly into a prompt, presenting a potential alternative to dedicated model training. However, the performance dynamics, operational costs, and core mechanisms of in-context learning at this extreme scale have remained poorly understood.
The article evaluates the effectiveness and underlying properties of long-context in-context learning across multiple open-source and proprietary models, directly comparing this approach against example retrieval methods and model finetuning.
The researchers conducted comprehensive empirical experiments across five diverse classification datasets (spanning 6 to 151 target classes) and one dialogue summarization task. They tested open-source models with context capacities ranging from 4,000 to 128,000 tokens—predominantly variants of Llama 2, Mistral, and Qwen—alongside frontier proprietary models including Claude 3.5 Sonnet and Llama 3.1 405B. The evaluation systematically varied the volume of prompt demonstrations from standard few-shot scales up to several thousand examples, analyzing the effects of random sampling, keyword- and semantic-based retrieval, parameter-efficient finetuning (LoRA), full-model finetuning, prompt ordering, and sparse attention mechanisms.
The investigation produced four central findings. First, performance scales substantially as prompt demonstrations increase; expanding from 10 to 1,000 examples yielded accuracy gains of up to 50.8 percentage points (averaging a 36.8-point gain across datasets), with performance frequently matching or outperforming parameter-efficient finetuning. Second, the necessity of dynamically retrieving relevant examples diminishes in the long-context regime: while retrieval provided up to a 51.5-point advantage over random selection in small prompts, this gap narrowed to under 5 points when over 1,500 demonstrations were included. Third, long-context prompts exhibit greater robustness to random example ordering, reducing label-flipping sensitivity by more than half compared to short prompts; however, grouping demonstrations by label severely degraded accuracy (by up to 25.7 points). Fourth, performance gains stem primarily from the model retrieving relevant patterns at inference time rather than deep cross-referencing across the demonstration pool, meaning sparse, blockwise attention patterns recover up to 95% of full attention accuracy.
These findings suggest that long-context in-context learning offers a viable third deployment paradigm between static few-shot prompting and dedicated model finetuning. Rather than maintaining custom fine-tuned weights for every business task or performing costly per-query example retrieval, organizations can provide a single, extensive set of randomly sampled demonstrations, precompute and cache the demonstration embeddings once, and reuse them across inference queries. While full-model finetuning remains superior when massive labeled training datasets vastly exceed the context window, long-context prompting provides an agile, low-maintenance alternative that eliminates training-time overhead.
Practitioners should evaluate trade-offs based on data availability, task variety, and operational constraints. For organizations managing numerous specialized tasks with moderate amounts of labeled data (hundreds to thousands of examples), caching long demonstration prompts is recommended to minimize maintenance and avoid the serving complexity of task-specific models. When preparing prompt data, teams must ensure examples are shuffled rather than sorted by category to prevent severe performance penalties. If latency or inference compute costs are primary constraints, or if tens of thousands of training instances are available, traditional full-model finetuning remains the preferred choice.
These conclusions are supported by rigorous ablations, but readers should note certain limitations. The experimental results focus primarily on classification tasks and 7-billion to 8-billion parameter open-source architectures; performance benefits plateaued more rapidly on frontier models, and highly generative workflows will fit fewer examples due to output length constraints. In addition, prompt-based methods still require cross-attention compute during inference. Organizations should conduct targeted pilot tests within their specific domain constraints before transitioning away from established finetuning or retrieval pipelines.
- Paper: A Survey on In-context Learning, Qingxiu Dong et al. (2024). This survey maps the main ICL mechanisms, demonstration choices, and evaluation approaches that frame the source’s investigation of ICL at unusually large demonstration counts.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). Its experiments isolate how demonstration labels, inputs, and formats affect ICL, providing useful context for interpreting the source’s findings on shuffling and label grouping.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). Its controlled study shows that long-context models can use evidence unevenly across an input, establishing an important baseline for the source’s tests of long-context ICL.
- Paper: MetaICL: Learning to Learn In Context, Sewon Min et al. (2022). MetaICL establishes how models trained to infer tasks from demonstrations behave, giving a foundation for the source’s study of scaling demonstration counts.
- Paper: Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation, Zichong Li et al. (2026). It extends the study of long-context positional effects by training models to resist them, a direct continuation of the source’s analysis of how long-context ICL responds to input arrangement.
- Paper: ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning, Yanjun Zhao et al. (2026). It carries long-context utilization research into a concrete reasoning method that selects and replays evidence, extending the source’s investigation of what models gain from large contexts.
- Paper: Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL, Xuhan Tong et al. (2026). Its theoretical treatment of demonstration counts and selection extends the source’s empirical findings into a framework for understanding how prompt design shapes ICL generalization.
