NoLiMa: Long-Context Evaluation Beyond Literal Matching
Ali ModarressiHanieh DeilamsalehyFranck DernoncourtTrung BuiRyan A. RossiSeunghyun YoonHinrich Schtze
Introduces the NoLiMa benchmark to expose how modern long-context language models fail to retrieve information through latent associations when deprived of surface-level lexical matches.
Modern large language models claim the capability to process massive contexts ranging from 128,000 to over one million tokens. Standard benchmarks evaluate these capabilities using retrieval tasks where a targeted fact is hidden within extensive irrelevant text. However, existing benchmarks predominantly rely on queries that share exact, literal word matches with the target information. This creates a critical blind spot in real-world applications such as search, summarization, and retrieval-augmented generation, where user queries rarely share exact wording with the relevant facts buried in extensive documents.
The article introduces and evaluates NOLIMA, a benchmark designed to assess long-context retrieval and latent reasoning without relying on literal word matches. The primary objective is to measure how well language models identify and retrieve relevant information when they must rely on underlying associative reasoning—such as real-world knowledge or commonsense connections—across growing context lengths.
The authors constructed a controlled dataset of 58 question-and-fact pairs embedded within haystacks of curated book snippets up to 128,000 tokens long. The evaluation rigorously removed distracting words and accidental answers from the background text to prevent confounding factors. The study evaluated 13 widely used commercial and open-weight language models, including GPT-4o, Gemini 1.5 Pro, and Llama 3.3 70B, running over 7,500 tests per context length to measure accuracy, the effect of multi-step associative hops, and the impact of irrelevant literal distractors.
The evaluation revealed several critical findings. First, while nearly all models achieve high accuracy (above 85% to 99%) in short contexts under 1,000 tokens, their performance collapses as context length grows. At 32,000 tokens, 11 of the 13 evaluated models retained less than half of their short-context baseline score, and even leading models like GPT-4o declined from 99.3% to 69.7%. Second, the effective reliable context length for most models was 2,000 tokens or fewer, falling vastly short of their claimed capacities of 128,000 tokens or more. Third, increasing the complexity of associative reasoning from one hop to two hops accelerated the performance decline across all models. Fourth, chain-of-thought prompting and specialized reasoning models improved accuracy modestly but failed to prevent severe degradation in contexts exceeding 16,000 tokens. Finally, introducing an irrelevant sentence with literal overlap to the query severely disrupted model retrieval, cutting GPT-4o's effective length to just 1,000 tokens.
These findings demonstrate that current transformer attention mechanisms rely heavily on surface-level keyword matching rather than robust deep reasoning across extended text. In production environments, relying on advertised context limits introduces significant operational risk, as models may fail to retrieve critical facts or become easily misled by superficial distractors. Organizations cannot assume that expanded context windows solve long-document comprehension without addressing this underlying retrieval vulnerability.
Organizations deploying large language models should not rely solely on vendor-advertised context windows for tasks that require semantic inference. System architects should design retrieval-augmented generation pipelines that minimize context length and filter out superficial keyword distractors before passing text to the model. Researchers and benchmark developers must also adopt evaluation suites that eliminate literal overlap to test genuine comprehension. Limitations of the article include the synthetic nature of the fact-needle templates and cost constraints that limited exhaustive evaluations beyond 32,000 tokens for all models. Nevertheless, the high volume of controlled tests provides strong confidence that current models suffer from severe attention degradation in long contexts when literal cues are absent.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). Its controlled evidence-position experiments establish the long-context retrieval failures that NoLiMa sharpens by removing literal query–answer overlap.
- Paper: Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models, Mosh Levy et al. (2024). Its finding that irrelevant input length degrades multi-step reasoning provides a direct baseline for NoLiMa’s tests of associative retrieval in growing contexts.
- Paper: Large Language Models Can Be Easily Distracted by Irrelevant Context, Freda Shi et al. (2023). Its experiments show how irrelevant and overlapping distractors derail model reasoning, framing NoLiMa’s specific test of literal-overlap interference.
- Paper: LooGLE: Can Long-Context Language Models Understand Long Contexts?, Jiaqi Li et al. (2024). Its benchmark separates short-range success from long-dependency failure, preparing readers for NoLiMa’s more controlled evaluation of retrieval without surface cues.
- Paper: Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA, Minzheng Wang et al. (2024). Its multi-document benchmark establishes the realistic long-context evaluation problem that NoLiMa isolates into associative retrieval and distractor effects.
- Paper: Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context, Keivan Alizadeh et al. (2026). It follows NoLiMa’s diagnosis of long-context failure by testing uncertainty-guided context handling as a strategy for selecting and processing relevant information.
- Paper: Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale, Siddharth Gollapudi et al. (2026). It extends the retrieval question to million-token corpora, measuring where in-context retrieval loses its signal and testing targeted remedies.
- Paper: A Benchmark for Deep Information Synthesis, Debjit Paul et al. (2026). It broadens long-context evaluation beyond embedded fact retrieval to multi-source research tasks requiring planning, evidence gathering, and synthesis.
