In-Context Learning with Long-Context Models: An In-Depth Exploration

Amanda BertschMaor IvgiEmily XiaoUri AlonJonathan BerantMatthew R. GormleyGraham Neubig

article2025NAACL161 citationsSAC Award for Language Modeling

Demonstrates that scaling in-context learning to thousands of demonstrations can rival model finetuning while significantly reducing sensitivity to example ordering and selection.

Listen

Recent advances in artificial intelligence have expanded the context windows of large language models, allowing them to process tens of thousands of words at once. Traditionally, adapting language models to specific tasks required either providing a handful of examples directly in the prompt—known as in-context learning—or computationally expensive finetuning to update the underlying model weights. With larger context capacities, organizations can now fit thousands of demonstration examples directly into a prompt, presenting a potential alternative to dedicated model training. However, the performance dynamics, operational costs, and core mechanisms of in-context learning at this extreme scale have remained poorly understood.

The article evaluates the effectiveness and underlying properties of long-context in-context learning across multiple open-source and proprietary models, directly comparing this approach against example retrieval methods and model finetuning.

The researchers conducted comprehensive empirical experiments across five diverse classification datasets (spanning 6 to 151 target classes) and one dialogue summarization task. They tested open-source models with context capacities ranging from 4,000 to 128,000 tokens—predominantly variants of Llama 2, Mistral, and Qwen—alongside frontier proprietary models including Claude 3.5 Sonnet and Llama 3.1 405B. The evaluation systematically varied the volume of prompt demonstrations from standard few-shot scales up to several thousand examples, analyzing the effects of random sampling, keyword- and semantic-based retrieval, parameter-efficient finetuning (LoRA), full-model finetuning, prompt ordering, and sparse attention mechanisms.

The investigation produced four central findings. First, performance scales substantially as prompt demonstrations increase; expanding from 10 to 1,000 examples yielded accuracy gains of up to 50.8 percentage points (averaging a 36.8-point gain across datasets), with performance frequently matching or outperforming parameter-efficient finetuning. Second, the necessity of dynamically retrieving relevant examples diminishes in the long-context regime: while retrieval provided up to a 51.5-point advantage over random selection in small prompts, this gap narrowed to under 5 points when over 1,500 demonstrations were included. Third, long-context prompts exhibit greater robustness to random example ordering, reducing label-flipping sensitivity by more than half compared to short prompts; however, grouping demonstrations by label severely degraded accuracy (by up to 25.7 points). Fourth, performance gains stem primarily from the model retrieving relevant patterns at inference time rather than deep cross-referencing across the demonstration pool, meaning sparse, blockwise attention patterns recover up to 95% of full attention accuracy.

These findings suggest that long-context in-context learning offers a viable third deployment paradigm between static few-shot prompting and dedicated model finetuning. Rather than maintaining custom fine-tuned weights for every business task or performing costly per-query example retrieval, organizations can provide a single, extensive set of randomly sampled demonstrations, precompute and cache the demonstration embeddings once, and reuse them across inference queries. While full-model finetuning remains superior when massive labeled training datasets vastly exceed the context window, long-context prompting provides an agile, low-maintenance alternative that eliminates training-time overhead.

Practitioners should evaluate trade-offs based on data availability, task variety, and operational constraints. For organizations managing numerous specialized tasks with moderate amounts of labeled data (hundreds to thousands of examples), caching long demonstration prompts is recommended to minimize maintenance and avoid the serving complexity of task-specific models. When preparing prompt data, teams must ensure examples are shuffled rather than sorted by category to prevent severe performance penalties. If latency or inference compute costs are primary constraints, or if tens of thousands of training instances are available, traditional full-model finetuning remains the preferred choice.

These conclusions are supported by rigorous ablations, but readers should note certain limitations. The experimental results focus primarily on classification tasks and 7-billion to 8-billion parameter open-source architectures; performance benefits plateaued more rapidly on frontier models, and highly generative workflows will fit fewer examples due to output length constraints. In addition, prompt-based methods still require cross-attention compute during inference. Organizations should conduct targeted pilot tests within their specific domain constraints before transitioning away from established finetuning or retrieval pipelines.

  • Paper: A Survey on In-context Learning, Qingxiu Dong et al. (2024). This survey maps the main ICL mechanisms, demonstration choices, and evaluation approaches that frame the source’s investigation of ICL at unusually large demonstration counts.
  • Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). Its experiments isolate how demonstration labels, inputs, and formats affect ICL, providing useful context for interpreting the source’s findings on shuffling and label grouping.
  • Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). Its controlled study shows that long-context models can use evidence unevenly across an input, establishing an important baseline for the source’s tests of long-context ICL.
  • Paper: MetaICL: Learning to Learn In Context, Sewon Min et al. (2022). MetaICL establishes how models trained to infer tasks from demonstrations behave, giving a foundation for the source’s study of scaling demonstration counts.
Cover for In-Context Learning with Long-Context Models: An In-Depth Exploration

Abstract

As model context lengths continue to increase, the number of demonstrations that can be provided in-context approaches the size of entire training datasets. We study the behavior of in-context learning (ICL) at this extreme scale on multiple datasets and models. We show that, for many datasets with large label spaces, performance continues to increase with thousands of demonstrations. We contrast this with example retrieval and finetuning: example retrieval shows excellent performance at low context lengths but has diminished gains with more demonstrations; finetuning is more data hungry than ICL but can exceed long-context ICL performance with additional data. We use the ICL setting to study several properties of both in-context learning and long-context models. We show that long-context ICL is less sensitive to random input shuffling than short-context ICL, that grouping of same-label examples negatively impacts performance, and that the performance boosts do not arise from cumulative gain from encoding many examples together. We conclude that long-context ICL can be an effective tool, and may not require long-context for encoding the demonstration set at all.¹

Table of Contents

  • 1 Introduction
  • 2 Experimental setup
  • 3 Long-context ICL
  • 3.1 Compared settings
  • 3.2 In-context results
  • 3.3 Comparison with finetuning
  • 4 Properties of long-context ICL
  • 5 Why does long-context ICL help?
  • 6 Related Work
  • 7 Conclusion
  • 8 Limitations
  • 9 Broader impacts
  • References
  • A Saturation
  • B Using ICL as a testbed for long-context model properties
  • C Full ICL results across datasets
  • C.1 Random selection ICL across all models
  • C.2 Retrieval ICL across all models
  • C.3 Comparing retrieval, random selection, and finetuning
  • D Constrained decoding details
  • E Finetuning
  • F Block attention patterns
  • G Prompt formatting and examples from datasets
  • G.1 TREC
  • G.2 TREC-fine
  • G.3 NLU
  • G.4 Banking-77
  • G.5 Clinic-150
  • G.6 SAMSum
  • H Additional details
  • H.1 Models selected
  • H.2 Computational cost

Knowls

  1. Knowl 1 — Classification accuracy continues to improve with thousands of in-context examples

    empirical result

    Across five classification datasets and models with extended context windows, in-context learning (ICL) generally improves as the demonstration set grows, including beyond 2,000 examples. For Llama-2-7B adapted to an 80K-token context, increasing the set from 10 to 1,000 demonstrations yields gains of up to 50.8 accuracy points and an average gain of 36.8 points across the five classification datasets. The scaling curves plotted on page 19 show this trend across TREC, TREC-fine, NLU, Banking-77, and Clinic-150. The results establish that many-shot ICL can continue to usefully scale well beyond the few-shot regime, though the amount of improvement varies by task and model.

  2. Knowl 2 — Retrieval matters less as the demonstration set grows

    empirical result

    BM25 and BERTScore-Recall retrieval often outperform random demonstration selection at low shot counts, but their advantage and their differences from each other shrink as more examples are included. Which retriever performs better at short contexts depends on the dataset; at long contexts their performance becomes nearly indistinguishable. On Banking-77, where retrieval helps most, BM25 improves accuracy over random selection by 51.5 points at one example but by only 4.9 points at 1,500 examples. At the longest context lengths studied, a single randomly selected set of demonstrations incurs no more than a 5-point penalty relative to retrieval; on TREC at 2,000 examples, the penalty is 1.8 points. This supports reusing and caching one encoded demonstration set when inference efficiency matters.

  3. Knowl 3 — ICL and finetuning trade data efficiency against inference cost

    empirical result

    In comparisons using Llama-2-7B on the classification tasks, ICL generally exceeds LoRA finetuning when relatively few task examples are available. With substantially more training data, full finetuning often reaches the highest performance, whereas LoRA usually does not exceed long-context ICL. Clinic-150, the task with 151 labels, is a notable case: neither full finetuning nor LoRA outperforms ICL at the same number of examples. Finetuning can nevertheless be preferable when additional data are available and inference cost is important, because it reduces the need to process a long prompt at every prediction. These conclusions are specific to the studied tasks, models, and finetuning procedures.

  4. Knowl 4 — Random example-order sensitivity weakens in long-context ICL

    empirical result

    The authors measured the fraction of predictions whose labels change after shuffling the in-context examples, averaging over three shuffles. Across all five classification datasets, the fraction of labels changed with 1,000 demonstrations is less than half the fraction changed with 10 demonstrations. Thus, long-context ICL remains sensitive to order to some degree, but substantially less so than short-context ICL. The plotted measurements on page 6 show the declining trend across datasets.

  5. Knowl 5 — Grouping demonstrations by label harms long-context performance

    empirical result

    Sorting demonstrations so that examples with the same label occur together has little effect when only a few examples are present, but increasingly hurts performance as the demonstration set grows. On Clinic-150 with Llama-2-32K, sorting by label reduces accuracy by 25.7 percentage points at 1,169 demonstrations. The result suggests that exposure to examples with different labels in nearby context is useful; grouping like labels disrupts this contextualization. The plotted comparison on page 6 shows the growing gap between ordinary ordering and label-sorted ordering.

  6. Knowl 6 — Blockwise attention can encode demonstrations with limited long-range attention

    model/method

    The study tests a blockwise attention mask for encoding a long demonstration set. Divide the demonstrations into consecutive blocks of bb examples. Each demonstration block can attend to the first block, which acts as an attention sink, and to the two preceding local blocks; the test example can attend to all demonstrations. This limits long-range attention among demonstrations while retaining global access for the prediction. The mask is a variation on Star Attention; the attention-mask diagrams on page 25 visualize the sink and local-block pattern.

    On Banking-77 with Llama-2-80K, block size b=50b=50, and 500 demonstrations, the tested masks produce the following accuracies:

    Attention pattern Block size Sink blocks Local blocks Accuracy
    Full causal attention 500 – – 80.04
    Current block only 50 0 0 18.72
    One preceding local block 50 0 1 51.16
    Sink block only 50 1 0 62.15
    Sink and one local block 50 1 1 75.12
    Sink and two local blocks 50 1 2 77.51

    Neither a sink block nor local attention alone approaches full attention. Combining both is substantially more effective; the chosen pattern uses two preceding local blocks to reduce the performance gap.

  7. Knowl 7 — Performance depends on both local contextualization and total demonstration count

    empirical result

    Let bb be the number of examples per attention block and kk the total number of demonstrations available to the test example. When the same examples are used and bb is varied, blockwise attention approaches full-attention performance: on Banking-77, a block size of 50 recovers 95% of full-attention performance, with comparable block-size ranges observed on the other studied datasets. Very small blocks (below 10 examples in the reported setting) yield near-zero performance, indicating that some local contextualization is needed.

    When bb is fixed and additional blocks are added to increase kk, the effect depends on block size. With extremely small blocks, adding many poorly contextualized examples performs worse than attending to a single block. Once local contextualization reaches a minimal level—around b=10b=10 in the Banking-77 experiment—adding more blocks substantially improves accuracy. These findings support the authors’ hypothesis that additional accessible examples, rather than extensive cross-attention among all demonstrations during encoding, account for much of the long-context benefit.

  8. Knowl 8 — ICL saturation occurs before the context limit on many tasks

    empirical result

    The study defines the saturation point as the smallest tested number of demonstrations at which accuracy reaches 95% of that model’s maximum accuracy on the dataset. Saturation generally occurs later for datasets with larger label spaces; for example, Banking-77 and Clinic-150 do not saturate within the 4K-token Llama-2 context. In longer-context models, saturation often still occurs before the maximum number of demonstrations that fits. The following values give the saturation point and, in parentheses, the maximum number of examples fitting the context window; “--” means the model’s maximum accuracy was achieved using its full context.

    Dataset Llama-2 Llama-2-32K Llama-2-80K Mistral
    TREC 20 (140) 100 (1129) 75 (2000) 50 (1129)
    TREC-fine 75 (131) 250 (1056) 500 (2000) 500 (1091)
    NLU 100 (162) 500 (1309) 500 (2000) 250 (1309)
    Banking-77 – (100) 500 (838) 750 (1750) 500 (860)
    Clinic-150 – (145) 750 (1169) 1000 (2000) 750 (1212)

    The measurements show both that a model need not always use its full context to reach near-maximum accuracy and that saturation in one model does not imply that additional demonstrations could not help a longer-context model.

  9. Knowl 9 — Long-context ICL also improves summarization scores

    empirical result

    On SAMSum, a dialogue-summarization dataset, the authors observe improving generation performance as the number of demonstrations increases, through at least 250-shot ICL. They evaluate with BERTScore and report similar trends with ROUGE. Because the conversations and summaries are longer than classification examples, fewer demonstrations fit in a given context. Retrieval is less consistently beneficial for this generation task and sometimes performs worse than random selection.

  10. Knowl 10 — Experimental comparison spans five classifiers, one generation task, and multiple context lengths

    experimental setup

    The experiments cover five classification datasets—TREC (6 labels; 5,452 training examples), TREC-fine (50 labels; 5,452), NLU (68 labels; 19,286), Banking-77 (77 labels; 10,003), and Clinic-150 (151 labels; 15,250)—and the SAMSum summarization dataset (14,732 training examples). The evaluated models are Llama-2-7B with a 4K context, its 32K and 80K context-adapted variants, Mistral-7B-v0.2, and Qwen2.5-7B; prompt lengths are restricted to each model’s trained context length. For classification, the study compares random demonstrations, BM25 or BERTScore-Recall retrieval, full finetuning, and LoRA finetuning. Random-selection ICL averages results over 10 random training-set shuffles, taking the first nn examples from each shuffle. Evaluation uses 250 subsampled test examples per dataset; classification is evaluated by accuracy (with macro-F1 also computed), and SAMSum by BERTScore (with ROUGE used to check trends). Classification ICL uses constrained decoding to restrict outputs to valid labels.

Coverage note — The appendix’s detailed finetuning hyperparameter ablations, classification-head initialization comparisons, copying test, and frontier-model measurements are omitted because they are secondary analyses and do not alter the central many-shot scaling, comparison, or attention findings.

References

  1. 1.Shantanu Acharya, Fei Jia, and Boris Ginsburg. Star attention: Efficient llm inference over long sequences, 2024. URL https://arxiv.org/abs/2411.17116.
  2. 2.Rishabh Agarwal, Avi Singh, Lei M. Zhang, Bernd Bohnet, Stephanie Chan, Ankesh Anand, Zaheer Abbas, Azade Nova, John D. Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle. Many-shot in-context learning, 2024.
  3. 3.Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan J. Hubinger, Yuntao Bai, Trenton Bricken, Timothy Maxwell, Nicholas Schiefer, Jamie Sully, and Alex Hernandez. Many-shot jailbreaking, 2024.
  4. 4.Anthropic. Introducing claude 3.5 sonnet, 2024.
  5. 5.Akari Asai, Sneha Kudugunta, Xinyan Velocity Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, and Hannaneh Hajishirzi. Buffet: Benchmarking large language models for few-shot cross-lingual transfer, 2023.
  6. 6.Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew Gormley. Unlimiformer: Long-range transformers with unlimited length input. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (eds.), Advances in Neural Information Processing Systems. Curran Associates, Inc., 2023.
  7. 7.Stella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Gregory Anthony, Shivanshu Purohit, and Edward Raff. Emergent and predictable memorization in large language models. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  8. 8.Necva Bölücü, Maciej Rybinski, and Stephen Wan. impact of sample selection on in-context learning for entity extraction from scientific writing. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.338.
  9. 9.Iñigo Casanueva, Tadas Temcinas, Daniela Gerz, Matthew Henderson, and Ivan Vulic. Efficient intent detection with dual sentence encoders. In Proceedings of the 2nd Workshop on Natural Language Processing for Conversational AI, pp. 38–45, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.nlp4convai-1.5.
  10. 10.Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. Extending context window of large language models via positional interpolation, 2023.
  11. 11.Hyunsoo Cho, Hyuhng Joon Kim, Junyeob Kim, Sang-Woo Lee, Sang goo Lee, Kang Min Yoo, and Taeuk Kim. Prompt-augmented linear probing: Scaling beyond the limit of few-shot in-context learners, 2023.
  12. 12.Google Deepmind. Our next-generation model: Gemini 1.5, 2024.
  13. 13.Gilad Deutch, Nadav Magar, Tomer Bar Natan, and Guy Dar. In-context learning and gradient descent revisited, 2024.
  14. 14.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaoqing Ellen Tan, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aaron Grattafiori, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alex Vaughan, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Franco, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, Danny Wyatt, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Firat Ozgenel, Francesco Caggioni, Francisco Guzmán, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Govind Thattai, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Karthik Prasad, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kun Huang, Kunal Chawla, Kushal Lakhotia, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Maria Tsimpoukelli, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikolay Pavlovich Laptev, Ning Dong, Ning Zhang, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Rohan Maheswari, Russ Howes, Ruty Rinott, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Kohler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vítor Albiero, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaofang Wang, Xiaojian Wu, Xiaolan Wang, Xide Xia, Xilun Wu, Xinbo Gao, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yuchen Hao, Yundi Qian, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, and Zhiwei Zhao. The llama 3 herd of models, 2024.
  15. 15.Yao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue, Hannaneh Hajishirzi, Yoon Kim, and Hao Peng. Data engineering for scaling language models to 128k context, 2024.
  16. 16.Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. Samsum corpus: A human-annotated dialogue dataset for abstractive summarization. In Proceedings of the 2nd Workshop on New Frontiers in Summarization. Association for Computational Linguistics, 2019. doi: 10.18653/v1/d19-5409. URL http://dx.doi.org/10.18653/v1/D19-5409.
  17. 17.Junxian Guo, Haotian Tang, Shang Yang, Zhekai Zhang, Zhijian Liu, and Song Han. Block Sparse Attention. https://github.com/mit-han-lab/Block-Sparse-Attention, 2024.
  18. 18.Shivanshu Gupta, Matt Gardner, and Sameer Singh. Coverage-based example selection for in-context learning, 2023. URL https://arxiv.org/abs/2305.14907.
  19. 19.Chi Han, Qifan Wang, Hao Peng, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang. Lm-infinite: Zero-shot extreme length generalization for large language models, 2024.
  20. 20.Zhixiong Han, Yaru Hao, Li Dong, Yutao Sun, and Furu Wei. Prototypical calibration for few-shot learning of language models, 2022.
  21. 21.Yaru Hao, Yutao Sun, Li Dong, Zhixiong Han, Yuxian Gu, and Furu Wei. Structured prompting: Scaling in-context learning to 1, 000 examples. ArXiv preprint, abs/2212.06713, 2022.
  22. 22.Roee Hendel, Mor Geva, and Amir Globerson. In-context learning creates task vectors. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.624.
  23. 23.Eduard Hovy, Laurie Gerber, Ulf Hermjakob, Chin-Yew Lin, and Deepak Ravichandran. Toward semantics-based answer pinpointing. In Proceedings of the First International Conference on Human Language Technology Research, 2001.
  24. 24.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022.
  25. 25.Maor Ivgi, Uri Shaham, and Jonathan Berant. Efficient long-text understanding with short-text models. Transactions of the Association for Computational Linguistics, 2023. doi: 10.1162/tacl_a_00547.
  26. 26.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7b, 2023.
  27. 27.Damjan Kalajdzievski. A rank stabilization scaling factor for fine-tuning with lora. ArXiv preprint, abs/2312.03732, 2023.
  28. 28.Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, and Jason Mars. An evaluation dataset for intent classification and out-of-scope prediction. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 1311–1316, Hong Kong, China, 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1131.
  29. 29.Itay Levy, Ben Bogin, and Jonathan Berant. Diverse demonstrations improve in-context compositional generalization. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada, 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.78.
  30. 30.Mosh Levy, Alon Jacoby, and Yoav Goldberg. Same task, more tokens: the impact of input length on the reasoning performance of large language models, 2024.
  31. 31.Dacheng Li, Rulin Shao, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang. How long can context length of open-source LLMs truly promise? In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following, 2023a.
  32. 32.Jiaqi Li, Mengmeng Wang, Zilong Zheng, and Muhan Zhang. Loogle: Can long-context language models understand long contexts?, 2023b.
  33. 33.Mukai Li, Shansan Gong, Jiangtao Feng, Yiheng Xu, Jun Zhang, Zhiyong Wu, and Lingpeng Kong. In-context learning with many demonstration examples, 2023c.
  34. 34.Tianle Li, Ge Zhang, Quy Duc Do, Xiang Yue, and Wenhu Chen. Long-context llms struggle with long in-context learning, 2024.
  35. 35.Xin Li and Dan Roth. Learning question classifiers. In COLING 2002: The 19th International Conference on Computational Linguistics, 2002.
  36. 36.Chin-Yew Lin. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pp. 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics. URL https://aclanthology.org/W04-1013/.
  37. 37.Ziqian Lin and Kangwook Lee. Dual operating modes of in-context learning, 2024.
  38. 38.Hao Liu, Matei Zaharia, and Pieter Abbeel. Ring attention with blockwise transformers for near-infinite context, 2023.
  39. 39.Haokun Liu, Derek Tam, Muqeeth Mohammed, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (eds.), Advances in Neural Information Processing Systems, 2022.
  40. 40.Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 2024. ISSN 2307-387X. doi: 10.1162/tacl_a_00638.
  41. 41.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8086–8098, Dublin, Ireland, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.556.
  42. 42.Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft, 2022.
  43. 43.Aristides Milios, Siva Reddy, and Dzmitry Bahdanau. In-context learning for text classification with many labels. In Dieuwke Hupkes, Verna Dankers, Khuyagbaatar Batsuren, Koustuv Sinha, Amirhossein Kazemnejad, Christos Christodoulopoulos, Ryan Cotterell, and Elia Bruni (eds.), Proceedings of the 1st GenBench Workshop on (Benchmarking) Generalisation in NLP, Singapore, 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.genbench-1.14.
  44. 44.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. MetaICL: Learning to learn in context. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2791–2809, Seattle, United States, 2022a. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.201.
  45. 45.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 11048–11064, Abu Dhabi, United Arab Emirates, 2022b. Association for Computational Linguistics.
  46. 46.Marius Mosbach, Tiago Pimentel, Shauli Ravfogel, Dietrich Klakow, and Yanai Elazar. Few-shot fine-tuning vs. in-context learning: A fair comparison and evaluation. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.779.
  47. 47.Jane Pan, Tianyu Gao, Howard Chen, and Danqi Chen. What in-context learning “learns” in-context: Disentangling task recognition and task learning. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.527.
  48. 48.Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. Yarn: Efficient context window extension of large language models, 2023.
  49. 49.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. True few-shot learning with language models. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 11054–11070, 2021.
  50. 50.Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud D. Karpas, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. Parallel context windows for large language models. In Annual Meeting of the Association for Computational Linguistics, 2022.
  51. 51.Stephen Robertson and Hugo Zaragoza. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends in Information Retrieval, 2009. doi: 10.1561/1500000019.
  52. 52.Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. Code llama: Open foundation models for code, 2024.
  53. 53.Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting, 2023.
  54. 54.Damien Sileo, Tim Van De Cruys, Camille Pradel, and Philippe Muller. Mining discourse markers for unsupervised sentence representation learning. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 3477–3486, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/N19-1351.
  55. 55.Woomin Song, Seunghyuk Oh, Sangwoo Mo, Jaehyung Kim, Sukmin Yun, Jung-Woo Ha, and Jinwoo Shin. Hierarchical context merging: Better long context understanding for pre-trained LLMs. In The Twelfth International Conference on Learning Representations, 2024.
  56. 56.Qwen Team. Qwen2.5: A party of foundation models, September 2024. URL https://qwenlm.github.io/blog/qwen2.5/.
  57. 57.TogetherAI. Llama-2-7b-32k-instruct - and fine-tuning for llama-2 models with together api, 2023.
  58. 58.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models, 2023.
  59. 59.Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek, Yuhuai Wu, Henryk Michalewski, and Piotr Miłos. Focused transformer: Contrastive training for context scaling. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (eds.), Advances in Neural Information Processing Systems. Curran Associates, Inc., 2023.
  60. 60.Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. Transformers learn in-context by gradient descent, 2023.
  61. 61.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-demos.6.
  62. 62.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks, 2024.
  63. 63.Pawel Swietojanski Xingkun Liu, Arash Eshghi and Verena Rieser. Benchmarking natural language understanding services for building conversational agents. In Proceedings of the Tenth International Workshop on Spoken Dialogue Systems Technology (IWSDS), Ortigia, Siracusa (SR), Italy, 2019. Springer.
  64. 64.Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, and Hao Ma. Effective long-context scaling of foundation models, 2023.
  65. 65.Paiheng Xu, Fuxiao Liu, Zongxia Li, and Hyemi Song. Towards understanding in-context learning with contrastive demonstrations and saliency maps, 2023.
  66. 66.Howard Yen, Tianyu Gao, and Danqi Chen. Long-context language modeling with parallel context encoding, 2024.
  67. 67.Cecilia Ying and Stephen Thomas. Label errors in BANKING77. In Proceedings of the Third Workshop on Insights from Negative Results in NLP, pp. 139–143, Dublin, Ireland, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.insights-1.19.
  68. 68.LILI YU, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis. MEGABYTE: Predicting million-byte sequences with multiscale transformers. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  69. 69.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert, 2020. URL https://arxiv.org/abs/1904.09675.
  70. 70.Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D. Manning. Position-aware attention and supervised data improve slot filling. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 35–45, Copenhagen, Denmark, 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1004.
  71. 71.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. In Marina Meila and Tong Zhang (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pp. 12697–12706. PMLR, 2021.
  72. 72.Dawei Zhu, Nan Yang, Liang Wang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. Pose: Efficient context window extension of llms via positional skip-wise training, 2024.

Citation

MLA
Bertsch, A., et al. “In-Context Learning with Long-Context Models: An In-Depth Exploration”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 12119–49, https://doi.org/10.18653/v1/2025.naacl-long.605.
APA
Bertsch, A., Ivgi, M., Xiao, E., Alon, U., Berant, J., Gormley, M. R., & Neubig, G. (2025). In-Context Learning with Long-Context Models: An In-Depth Exploration. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 12119–12149. https://doi.org/10.18653/v1/2025.naacl-long.605
Chicago
Bertsch, A., M. Ivgi, E. Xiao, et al. 2025. “In-Context Learning with Long-Context Models: An In-Depth Exploration”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 12119–49. https://doi.org/10.18653/v1/2025.naacl-long.605.
Harvard
Bertsch, A. et al. (2025) “In-Context Learning with Long-Context Models: An In-Depth Exploration”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12119–12149. Available at: https://doi.org/10.18653/v1/2025.naacl-long.605.
Vancouver
1. Bertsch A, Ivgi M, Xiao E, Alon U, Berant J, Gormley MR, Neubig G (2025) In-Context Learning with Long-Context Models: An In-Depth Exploration. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 12119–12149

BibTeX

@inproceedings{bertsch-etal-2025-context,
    title = "In-Context Learning with Long-Context Models: An In-Depth Exploration",
    author = "Bertsch, Amanda  and
      Ivgi, Maor  and
      Xiao, Emily  and
      Alon, Uri  and
      Berant, Jonathan  and
      Gormley, Matthew R.  and
      Neubig, Graham",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.605/",
    doi = "10.18653/v1/2025.naacl-long.605",
    pages = "12119--12149",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/