NoLiMa: Long-Context Evaluation Beyond Literal Matching

Ali ModarressiHanieh DeilamsalehyFranck DernoncourtTrung BuiRyan A. RossiSeunghyun YoonHinrich Schtze

article2025ICML140 citations

Introduces the NoLiMa benchmark to expose how modern long-context language models fail to retrieve information through latent associations when deprived of surface-level lexical matches.

Listen

Modern large language models claim the capability to process massive contexts ranging from 128,000 to over one million tokens. Standard benchmarks evaluate these capabilities using retrieval tasks where a targeted fact is hidden within extensive irrelevant text. However, existing benchmarks predominantly rely on queries that share exact, literal word matches with the target information. This creates a critical blind spot in real-world applications such as search, summarization, and retrieval-augmented generation, where user queries rarely share exact wording with the relevant facts buried in extensive documents.

The article introduces and evaluates NOLIMA, a benchmark designed to assess long-context retrieval and latent reasoning without relying on literal word matches. The primary objective is to measure how well language models identify and retrieve relevant information when they must rely on underlying associative reasoning—such as real-world knowledge or commonsense connections—across growing context lengths.

The authors constructed a controlled dataset of 58 question-and-fact pairs embedded within haystacks of curated book snippets up to 128,000 tokens long. The evaluation rigorously removed distracting words and accidental answers from the background text to prevent confounding factors. The study evaluated 13 widely used commercial and open-weight language models, including GPT-4o, Gemini 1.5 Pro, and Llama 3.3 70B, running over 7,500 tests per context length to measure accuracy, the effect of multi-step associative hops, and the impact of irrelevant literal distractors.

The evaluation revealed several critical findings. First, while nearly all models achieve high accuracy (above 85% to 99%) in short contexts under 1,000 tokens, their performance collapses as context length grows. At 32,000 tokens, 11 of the 13 evaluated models retained less than half of their short-context baseline score, and even leading models like GPT-4o declined from 99.3% to 69.7%. Second, the effective reliable context length for most models was 2,000 tokens or fewer, falling vastly short of their claimed capacities of 128,000 tokens or more. Third, increasing the complexity of associative reasoning from one hop to two hops accelerated the performance decline across all models. Fourth, chain-of-thought prompting and specialized reasoning models improved accuracy modestly but failed to prevent severe degradation in contexts exceeding 16,000 tokens. Finally, introducing an irrelevant sentence with literal overlap to the query severely disrupted model retrieval, cutting GPT-4o's effective length to just 1,000 tokens.

These findings demonstrate that current transformer attention mechanisms rely heavily on surface-level keyword matching rather than robust deep reasoning across extended text. In production environments, relying on advertised context limits introduces significant operational risk, as models may fail to retrieve critical facts or become easily misled by superficial distractors. Organizations cannot assume that expanded context windows solve long-document comprehension without addressing this underlying retrieval vulnerability.

Organizations deploying large language models should not rely solely on vendor-advertised context windows for tasks that require semantic inference. System architects should design retrieval-augmented generation pipelines that minimize context length and filter out superficial keyword distractors before passing text to the model. Researchers and benchmark developers must also adopt evaluation suites that eliminate literal overlap to test genuine comprehension. Limitations of the article include the synthetic nature of the fact-needle templates and cost constraints that limited exhaustive evaluations beyond 32,000 tokens for all models. Nevertheless, the high volume of controlled tests provides strong confidence that current models suffer from severe attention degradation in long contexts when literal cues are absent.

Cover for NoLiMa: Long-Context Evaluation Beyond Literal Matching

Abstract

Recent large language models (LLMs) support long contexts ranging from 128K to 1M tokens. A popular method for evaluating these capabilities is the needle-in-a-haystack (NIAH) test, which involves retrieving a “needle” (relevant information) from a “haystack” (long irrelevant context). Extensions of this approach include increasing distractors, fact chaining, and in-context reasoning. However, in these benchmarks, models can exploit existing literal matches between the needle and haystack to simplify the task. To address this, we introduce NoLiMa, a benchmark extending NIAH with a carefully designed needle set, where questions and needles have minimal lexical overlap, requiring models to infer latent associations to locate the needle within the haystack. We evaluate 13 popular LLMs that claim to support contexts of at least 128K tokens. While they perform well in short contexts (<1K), performance degrades significantly as context length increases. At 32K, for instance, 11 models drop below 50% of their strong short-length baselines. Even GPT-4o, one of the top-performing exceptions, experiences a reduction from an almost-perfect baseline of 99.3% to 69.7%. Our analysis suggests these declines stem from the increased difficulty the attention mechanism faces in longer contexts when literal matches are absent, making it harder to retrieve relevant information. Even models enhanced with reasoning capabilities or CoT prompting struggle to maintain performance in long contexts. We publicly release the dataset and evaluation code at https://github.com/adobe-research/NoLiMa.

Table of Contents

  • Abstract
  • 1. Introduction
  • 2. Related Work
  • 3. NOLIMA
  • 3.1. Haystack Filtering Pipeline
  • 4. Experiments
  • 4.1. Dataset Configuration
  • 4.2. Models
  • 4.3. Evaluation Setup & Metric
  • 4.4. Results
  • 4.4.1. Latent Hops & Inversion
  • 4.4.2. Needle Placement Depth Analysis
  • 4.4.3. CoT Prompting
  • 4.4.4. Ablation Study: Literal Match Effect
  • 5. Conclusion
  • Impact Statement
  • Acknowledgments
  • References
  • A. Needle Set Design & Considerations
  • B. Models
  • C. Task Prompt Templates & Inference Settings
  • D. Distractor Design
  • E. Results Beyond 32K & Recent LLMs
  • F. One- & Two-hop Results
  • G. Raw Needle Placement Depth Plots

Knowls

  1. Knowl 1 — NOLIMA tests associative retrieval while suppressing literal-match cues

    model/method

    NOLIMA is a needle-in-a-haystack benchmark designed to test whether a language model can retrieve a character from a long text when the question and the relevant fact are linked by an association rather than shared keywords. Each needle names a randomly selected character and states a fact containing a needle keyword WnW_n; the question contains an associated query keyword WqW_q and asks which character fits. For example, a needle can say that a character lives next to the Semper Opera House, while the question asks who has been to Dresden. One-hop items connect WqW_q to WnW_n directly through an association; two-hop items require an additional link, such as connecting a city to its state. The set also varies fact order: in the default form the character precedes WnW_n, while in the inverted form the character follows it.

    The 58 question–needle pairs come from five groups of templates, covering country/city and landmark associations and dietary restrictions linked to drinks or foods. The authors selected simple, uniquely associated keyword pairs intended to be answerable in short contexts, used diverse character names not already present in the haystacks, and sought to avoid language-specific cues. The design and its example templates are illustrated in the paper’s page-3 template example and page-13 template inventory.

    As a measure of lexical overlap, the authors report ROUGE precision between each question and its relevant context, in the order R-1/R-2/R-L: NOLIMA scores 0.069/0.002/0.067; vanilla NIAH, 0.905/0.789/0.855; RULER S-NIAH, 0.571/0.461/0.500; and BABILong (0K), 0.553/0.238/0.522. Other comparisons are ∞Bench QA, 0.966/0.545/0.960; ∞Bench MC, 0.946/0.506/0.932; RULER QA, 0.809/0.437/0.693; and HELMET RAG, 0.689/0.304/0.555. Thus the benchmark substantially reduces question-to-evidence lexical overlap relative to these evaluated datasets.

  2. Knowl 2 — Most evaluated models lose accuracy well before their advertised context limits

    empirical result

    Across 13 evaluated language models, short-context base scores were generally high, but accuracy declined as context length increased. The paper defines effective length as the largest tested context length whose accuracy exceeds 85% of that model’s base score. The page-5 results table reports the following accuracies (percent); values in parentheses after each base score are the 85% thresholds.

    Model Claimed Effective Base (threshold) 1K 2K 4K 8K 16K 32K
    GPT-4o 128K 8K 99.3 (84.4) 98.1 98.0 95.7 89.2 81.6 69.7
    Llama 3.3 70B 128K 2K 97.3 (82.7) 94.2 87.4 81.5 72.1 59.5 42.7
    Llama 3.1 405B 128K 2K 94.7 (80.5) 89.0 85.0 74.5 60.1 48.4 38.0
    Llama 3.1 70B 128K 2K 94.5 (80.3) 91.0 81.8 71.2 62.7 51.8 43.2
    Gemini 1.5 Pro 2M 2K 92.6 (78.7) 86.4 82.7 75.4 63.9 55.5 48.2
    Jamba 1.5 Mini 256K <1K 92.4 (78.6) 76.3 74.1 70.8 62.2 52.7 43.6
    Command R+ 128K <1K 90.9 (77.3) 77.0 73.5 66.2 39.5 21.3 7.4
    Gemini 2.0 Flash 1M 4K 89.4 (76.0) 87.7 87.5 77.9 64.7 48.2 41.0
    Mistral Large 2 128K 2K 87.9 (74.7) 86.1 85.5 73.3 51.4 32.6 18.8
    Claude 3.5 Sonnet 200K 4K 87.5 (74.4) 85.4 84.0 77.6 61.7 45.7 29.8
    Gemini 1.5 Flash 1M <1K 84.7 (72.0) 68.6 61.6 51.0 44.4 35.5 28.6
    GPT-4o mini 128K <1K 84.8 (72.1) 67.7 58.2 44.2 32.6 20.6 13.7
    Llama 3.1 8B 128K 1K 76.7 (65.2) 65.7 54.4 44.1 31.9 22.6 14.2

    At 32K, 11 of the 13 models score at or below half their base score. GPT-4o is the strongest exception in this evaluation, but its accuracy still falls from 99.3% at base to 69.7% at 32K, and its effective length is 8K rather than its claimed 128K.

  3. Knowl 3 — Haystacks combine filtered book snippets with repeated needle placements

    experimental setup

    The evaluation haystacks are built from 10 open-license books, each containing at least 50K tokens. After filtering, the authors repeatedly select a book at random, extract a continuous snippet shorter than 250 tokens, and append it until the haystack exceeds 2K lines and 60K tokens. This concatenation of short snippets is intended to reduce the risk that models answer by memorizing publicly available book text.

    For each tested context length, each needle is inserted at 26 evenly spaced positions. Five randomly generated haystacks, 58 question–needle pairs, and 26 placements yield 7,540 tests per context length. Evaluations cover 250, 500, 1K, 2K, 4K, 8K, 16K, and 32K tokens. Accuracy is the proportion of tests in which the returned answer contains the correct character name.

    The base score is computed by first averaging each question–needle pair’s accuracy across the five haystacks at each of 250, 500, and 1K tokens, taking that pair’s highest score across those three lengths, and averaging these maxima over all pairs. Normalized accuracy is a model’s accuracy divided by its base score; effective length is the largest tested length with accuracy above 85% of base.

  4. Knowl 4 — Two-pass filtering removes superficial distractors and plausible false answers

    model/method

    The haystack filtering pipeline removes both words that could provide unintended lexical or semantic cues and content that might be mistaken for an answer. In the first pass, the authors embed the haystack’s words using Contriever, compare word embeddings with the question keywords using dot-product similarity, and manually inspect the top 20 similar words for each query keyword WqW_q. Words judged highly similar by meaning or substring are flagged, and sentences containing flagged words are removed.

    In the second pass, questions are paired with short chunks of the already filtered text and given to an instruction-tuned language model with a short instruction and four few-shot examples. Chunks are 1,000 characters long with an 800-character stride (approximately 250 tokens). The model returns either N/A or an explanation identifying a possible answer. A reviewer manually checks flagged content and removes spans that could yield a plausible but unintended answer; the scan is repeated until it identifies no further content for removal. The page-4 workflow diagram depicts these two successive filtering passes and the manual review loop. Llama 3.3 70B was used for filtering; in a control test with needles inserted into 100 randomly selected chunks per question, it flagged conflicting information with 99.8% accuracy.

  5. Knowl 5 — Two-hop reasoning and inverted fact order make retrieval harder

    empirical result

    NOLIMA’s two-hop questions are harder than its one-hop questions at the same context length, and the performance gap generally grows as contexts lengthen. For example, Llama 3.3 70B’s separately computed effective length is 8K for one-hop items but 1K for two-hop items; GPT-4o’s is 16K for one-hop and 8K for two-hop. Both subsets have strong short-context base scores, so the difference concerns length generalization rather than simply whether the association can be solved in a short context.

    Inverted needles, in which the character name follows the needle keyword rather than preceding it, are also more difficult than default-order needles. The authors propose that causal attention may contribute: in default order, a question’s association to WnW_n can help retrieve the earlier character name, whereas in inverted order the model must rely on weaker signals linking the name after WnW_n to the query. This is an interpretation, not a demonstrated mechanism; the authors state that deeper mechanistic analysis is needed.

  6. Knowl 6 — Aligned needle positions indicate that long-context load dominates in two-hop cases

    empirical result

    Needle-position analyses distinguish position effects from context-length effects. In the full-context sweep, Llama 3.3 70B shows a lost-in-the-middle pattern, particularly for one-hop items. For two-hop items, increasing context length lowers performance across the position sweep, including near the beginning and end of the haystack.

    To compare like-for-like positions, the authors also align needles within the final 2K tokens across different context lengths. They use 51 placements and a moving average with a window of 12 to make trends clearer. In this aligned region, one-hop performance still tends to dip toward the middle, but two-hop curves are comparatively stable across positions while their overall accuracy drops as total context length grows. Because the question-to-needle relative distance remains fixed across aligned positions, the authors interpret the two-hop decline as more consistent with difficulty attending over the larger context than with position encoding alone. The page-7 plots show both the full sweeps and the aligned final-2K comparison.

  7. Knowl 7 — Chain-of-thought and reasoning models improve scores but do not remove long-context failures

    empirical result

    For Llama 3.3 70B, prompting for a short chain-of-thought (CoT) improves accuracy on both one-hop and two-hop items, with larger relative gains on two-hop items. Scores below are percentages at 4K, 8K, 16K, and 32K; the increase row is the reported relative improvement.

    Task Prompt 4K 8K 16K 32K
    One-hop Without CoT 90.3 84.1 73.2 56.2
    One-hop With CoT 95.6 91.1 82.6 60.6
    One-hop Increase rate 5.9% 8.3% 12.8% 7.8%
    Two-hop Without CoT 70.7 57.4 42.7 25.9
    Two-hop With CoT 82.4 70.1 56.7 34.3
    Two-hop Increase rate 16.5% 22.1% 32.7% 32.4%

    The gains do not close the gap: two-hop CoT results remain poor at 16K and 32K. On NOLIMA-Hard, a subset of the 10 most difficult question–needle pairs, reasoning models also fall below half of base accuracy at 32K. Their 32K scores are 31.1 for GPT-o1, 18.9 for GPT-o3 Mini, and 20.7 for DeepSeek R1 Distill-Llama-70B; their respective base scores are 99.9, 98.8, and 99.9. Llama 3.3 70B scores 8.9 at 32K without CoT and 10.1 with CoT on this subset. Thus, strong short-context performance and additional reasoning steps do not ensure long-context generalization when lexical cues are absent.

  8. Knowl 8 — Literal matches sharply simplify NOLIMA questions

    empirical result

    Two ablations with Llama 3.3 70B show that literal overlap between a question and the relevant needle substantially improves retrieval. Direct questions explicitly include the needle keyword WnW_n, making the task resemble ordinary NIAH. Multiple-choice (MC) questions retain the latent association task but offer four character names as options: the correct name from the needle and three names from the haystack.

    Condition 8K 16K 32K
    Direct question 98.3 98.5 98.5
    One-hop, original 84.1 73.2 56.2
    One-hop, MC with literal match 98.7 97.4 93.1
    Two-hop, original 57.4 42.7 25.9
    Two-hop, MC with literal match 96.3 94.6 87.2

    Both direct questions and MC options greatly outperform the original latent-association questions at long lengths. The MC results show that even when the reasoning link is retained, adding names as options and a literal cue can make the search much easier.

  9. Knowl 9 — Irrelevant query-keyword matches distract models from the relevant fact

    empirical result

    The authors test the reverse case: a haystack contains a literal match to the query keyword, but that match is irrelevant to the answer and the relevant needle still has minimal lexical overlap with the question. Each added distractor is a neutral newspaper sentence about WqW_q, using a template such as an article or photo about that keyword. It is placed between 20% and 80% of context length and at least 20% of context length away from the needle.

    Adding these distractors reduces length generalization for both GPT-4o and Llama 3.3 70B. GPT-4o’s effective length falls from 8K without distractors to 1K with them. The base scores also fall, from 99.3 to 93.8 for GPT-4o and from 97.3 to 84.4 for Llama 3.3 70B; normalized performance nevertheless shows further degradation with length. The page-9 performance plot compares the normalized curves with and without distractors.

  10. Knowl 10 — Newer models still fall short of claimed context lengths

    empirical result

    Additional evaluations show that the length-generalization problem persists in newer models and at contexts beyond 32K. GPT-4.1, with a claimed 1M-token context, has a 16K effective length; its accuracy is 97.0% at base, 79.8% at 32K, and 64.7% at 128K. Gemini 2.5 Flash, also claiming 1M tokens, has a 2K effective length and scores 48.4% at 32K. Llama 4 Maverick has a 2K effective length and scores 24.5% at 32K.

    In a separate extension to 64K and 128K, GPT-4o scores 62.4% and 56.0%, respectively, from a 99.3% base; Gemini 2.0 Flash scores 33.0% and 16.4%, respectively, from an 89.4% base. For this extension, the authors used 11 needle placements per context length instead of 26, and limited the haystack to 127,500 tokens for both models to fit GPT-4o’s token limit. These results show that some models retain meaningful accuracy at very long lengths, but performance remains far below short-context accuracy and advertised limits.

Coverage note — The full 58-item keyword/template inventory and complete per-model one-hop and two-hop score matrices are omitted because their design principles and distinct aggregate length-generalization patterns are captured here without duplicating the main model results.

References

  1. 1.Agarwal, R., Singh, A., Zhang, L. M., Bohnet, B., Rosias, L., Chan, S. C., Zhang, B., Faust, A., and Larochelle, H. Many-shot in-context learning. In ICML 2024 Workshop on In-Context Learning, 2024. URL https://openreview.net/forum?id=goi7DFHlqS.
  2. 2.Anthropic, A. Claude 3.5 sonnet model card addendum. Claude-3.5 Model Card, 3, 2024.
  3. 3.Arora, S., Eyuboglu, S., Timalsina, A., Johnson, I., Poli, M., Zou, J., Rudra, A., and Re, C. Zoology: Measuring and improving recall in efficient language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=LY3ukUANko.
  4. 4.Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C. Using fast weights to attend to the recent past. Advances in neural information processing systems, 29, 2016.
  5. 5.Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., Dong, Y., Tang, J., and Li, J. LongBench: A bilingual, multitask benchmark for long context understanding. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3119–3137, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.172. URL https://aclanthology.org/2024.acl-long.172/.
  6. 6.Chang, Y., Lo, K., Goyal, T., and Iyyer, M. Booookscore: A systematic exploration of book-length summarization in the era of LLMs. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=7Ttk3RzDeu.
  7. 7.Chen, S., Wong, S., Chen, L., and Tian, Y. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595, 2023.
  8. 8.Cohere For AI. c4ai-command-r-plus-08-2024, 2024. URL https://huggingface.co/CohereForAI/c4ai-command-r-plus-08-2024.
  9. 9.DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025.
  10. 10.Dong, Z., Tang, T., Li, J., Zhao, W. X., and Wen, J.-R. BAMBOO: A comprehensive benchmark for evaluating long text modeling capacities of large language models. In Calzolari, N., Kan, M.-Y., Hoste, V., Lenci, A., Sakti, S., and Xue, N. (eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pp. 2086–2099, Torino, Italia, May 2024. ELRA and ICCL. URL https://aclanthology.org/2024.lrec-main.188/.
  11. 11.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  12. 12.Gemini Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  13. 13.Gemini Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  14. 14.Goldman, O., Jacovi, A., Slobodkin, A., Maimon, A., Dagan, I., and Tsarfaty, R. Is it really long context if all you need is retrieval? towards genuinely difficult long context NLP. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 16576–16586, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.924. URL https://aclanthology.org/2024.emnlp-main.924/.
  15. 15.Graves, A., Wayne, G., and Danihelka, I. Neural turing machines, 2014. URL https://arxiv.org/abs/1410.5401.
  16. 16.Hengle, A., Bajpai, P., Dan, S., and Chakraborty, T. Multilingual needle in a haystack: Investigating long-context behavior of multilingual large language models. arXiv preprint arXiv:2408.10151, 2024.
  17. 17.Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., and Ginsburg, B. RULER: What’s the real context size of your long-context language models? In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=kIoBbc76Sy.
  18. 18.Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024.
  19. 19.Izacard, G., Caron, M., Hosseini, L., Riedel, S., Bojanowski, P., Joulin, A., and Grave, E. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=jKN1pXi7b0.
  20. 20.Jiang, B., Xie, Y., Hao, Z., Wang, X., Mallick, T., Su, W. J., Taylor, C. J., and Roth, D. A peek into token bias: Large language models are not yet genuine reasoners. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 4722–4756, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.272. URL https://aclanthology.org/2024.emnlp-main.272/.
  21. 21.Kamradt, G. Needle in a haystack-pressure testing llms. Github Repository, pp. 28, 2023.
  22. 22.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=e2TBb5y0yFf.
  23. 23.Kuratov, Y., Bulatov, A., Anokhin, P., Rodkin, I., Sorokin, D. I., Sorokin, A., and Burtsev, M. BABILong: Testing the limits of LLMs with long context reasoning-in-a-haystack. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. URL https://openreview.net/forum?id=u7m2CG84BQ.
  24. 24.Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023.
  25. 25.Lee, J., Chen, A., Dai, Z., Dua, D., Sachan, D. S., Boratko, M., Luan, Y., Arnold, S. M. R., Perot, V., Dalmia, S., Hu, H., Lin, X., Pasupat, P., Amini, A., Cole, J. R., Riedel, S., Naim, I., Chang, M.-W., and Guu, K. Can long-context language models subsume retrieval, rag, sql, and more?, 2024. URL https://arxiv.org/abs/2406.13121.
  26. 26.Levy, M., Jacoby, A., and Goldberg, Y. Same task, more tokens: the impact of input length on the reasoning performance of large language models. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15339–15353, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.818. URL https://aclanthology.org/2024.acl-long.818/.
  27. 27.Li, M., Zhang, S., Liu, Y., and Chen, K. Needlebench: Can llms do retrieval and reasoning in 1 million context window?, 2024. URL https://arxiv.org/abs/2407.11963.
  28. 28.Lin, C.-Y. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pp. 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics. URL https://aclanthology.org/W04-1013/.
  29. 29.Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173, 2024. doi: 10.1162/tacl a 00638. URL https://aclanthology.org/2024.tacl-1.9/.
  30. 30.Maharana, A., Lee, D.-H., Tulyakov, S., Bansal, M., Barbieri, F., and Fang, Y. Evaluating very long-term conversational memory of LLM agents. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 13851–13870, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.747. URL https://aclanthology.org/2024.acl-long.747/.
  31. 31.Meta, A. Llama 3.3 model card. 2024. URL https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md.
  32. 32.Mistral, A. Mistral large 2. Mistral Large 2 Blogpost, 2024. URL https://mistral.ai/news/mistral-large-2407/.
  33. 33.Mohtashami, A. and Jaggi, M. Random-access infinite context length for transformers. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=7eHn64wOVy.
  34. 34.Navigli, R., Conia, S., and Ross, B. Biases in large language models: Origins, inventory, and discussion. J. Data and Information Quality, 15(2), June 2023. ISSN 1936-1955. doi: 10.1145/3597307. URL https://doi.org/10.1145/3597307.
  35. 35.Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. In-context learning and induction heads, 2022. URL https://arxiv.org/abs/2209.11895.
  36. 36.OpenAI. Openai o3-mini system card. 2025. URL https://openai.com/index/o3-mini-system-card/.
  37. 37.OpenAI, : Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024.
  38. 38.Peng, B., Quesnelle, J., Fan, H., and Shippole, E. YaRN: Efficient context window extension of large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=wHBfxhZu1u.
  39. 39.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  40. 40.Team, G. Gemma 3. 2025. URL https://goo.gle/Gemma3Report.
  41. 41.Team, J., Lenz, B., Arazi, A., Bergman, A., Manevich, A., Peleg, B., Aviram, B., Almagor, C., Fridman, C., Padnos, D., et al. Jamba-1.5: Hybrid transformer-mamba models at scale. arXiv preprint arXiv:2408.12570, 2024.
  42. 42.Vodrahalli, K., Ontanon, S., Tripuraneni, N., Xu, K., Jain, S., Shivanna, R., Hui, J., Dikkala, N., Kazemi, M., Fatemi, B., Anil, R., Dyer, E., Shakeri, S., Vij, R., Mehta, H., Ramasesh, V., Le, Q., Chi, E., Lu, Y., Firat, O., Lazaridou, A., Lespiau, J.-B., Attaluri, N., and Olszewska, K. Michelangelo: Long context evaluations beyond haystacks via latent structure queries, 2024. URL https://arxiv.org/abs/2409.12640.
  43. 43.Wang, M., Chen, L., Cheng, F., Liao, S., Zhang, X., Wu, B., Yu, H., Xu, N., Zhang, L., Luo, R., Li, Y., Yang, M., Huang, F., and Li, Y. Leave no document behind: Benchmarking long-context LLMs with extended multi-doc QA. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 5627–5646, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.322. URL https://aclanthology.org/2024.emnlp-main.322/.
  44. 44.Wei, J., Wang, X., Schuurmans, D., Bosma, M., ichter, b., Xia, F., Chi, E., Le, Q. V., and Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, pp. 24824–24837. Curran Associates, Inc., 2022.
  45. 45.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. Transformers: State-of-the-art natural language processing. In Liu, Q. and Schlangen, D. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, October 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-demos.6. URL https://aclanthology.org/2020.emnlp-demos.6/.
  46. 46.Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H. Effective long-context scaling of foundation models. In Duh, K., Gomez, H., and Bethard, S. (eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 4643–4663, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.naacl-long.260. URL https://aclanthology.org/2024.naacl-long.260/.
  47. 47.Yen, H., Gao, T., Hou, M., Ding, K., Fleischer, D., Izsak, P., Wasserblat, M., and Chen, D. Helmet: How to evaluate long-context language models effectively and thoroughly. arXiv preprint arXiv:2410.02694, 2024.
  48. 48.Zhang, X., Chen, Y., Hu, S., Xu, Z., Chen, J., Hao, M. K., Han, X., Thai, Z. L., Wang, S., Liu, Z., and Sun, M. ∞bench: Extending long context evaluation beyond 100k tokens, 2024. URL https://arxiv.org/abs/2402.13718.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/