Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
Minzheng WangLongze ChenFu ChengShengyi LiaoXinghua ZhangBingli WuHaiyang YuNan XuLei ZhangRun Luo
Introduces Loong, a realistic long-context benchmark that tests large language models across distributed multi-document scenarios where every included text is necessary to derive the correct answer.
Modern large language models frequently claim the ability to process massive amounts of information simultaneously, often advertising context windows spanning hundreds of thousands of words. However, existing industry benchmarks evaluate these models using synthetic shortcuts, such as hiding a single fact inside unrelated filler text or concentrating evidence within a single document. These evaluation methods fail to reflect realistic enterprise workflows—such as aggregating multi-year financial statements, comparing case law, or synthesizing research papers—where missing even one document leads to incorrect conclusions.
The article introduces and evaluates Loong, a realistic benchmark designed to rigorously assess how language models handle extended multi-document question answering. The benchmark tests whether models can synthesize distributed information across realistic contexts where every provided document contains essential evidence.
To construct this benchmark, the authors collected 1,600 verified test cases across financial reports, legal cases, and academic papers, spanning both English and Chinese. The evaluation spans input lengths from 10,000 to over 250,000 tokens and tests four distinct capabilities: locating a specific document among distractors (Spotlight Locating), comparing information across multiple sources (Comparison), grouping distributed evidence (Clustering), and multi-step reasoning over sequential data (Chain of Reasoning). Using this dataset, the authors benchmarked seven advanced models, including proprietary systems like Gemini-1.5-pro and GPT-4o, alongside leading open-source models, while also evaluating the impact of retrieval-augmented generation (RAG).
The analysis reveals that current language models struggle significantly in realistic multi-document scenarios. Even the top-performing model, Gemini-1.5-pro, achieved an overall average score of only 55.37 out of 100, with a perfect completion rate of just 27%, while GPT-4o scored 53.47 with a 26% perfect rate. Model performance drops sharply as context lengths expand; for example, models trained on 128,000-token windows degraded significantly once inputs exceeded 50,000 tokens, revealing a substantial gap between advertised window sizes and effective processing capabilities. Furthermore, while models performed relatively well on simple single-document lookup tasks, they degraded on complex clustering and comparison tasks that require multi-source synthesis. Finally, adding retrieval-augmented generation caused overall performance to decline—lowering GPT-4o's score from 53.47 to between 32.85 and 46.52 depending on retrieval settings—because standard search mechanisms failed to retrieve all required documents from the collection.
These findings indicate that organizations face material operational and compliance risks if they rely on advertised context window sizes for tasks that demand comprehensive multi-document synthesis. Standard retrieval pipelines cannot reliably replace true long-context processing when evidence is dispersed across many sources. For enterprise leaders, deploying models for multi-document auditing, legal discovery, or financial synthesis without strict verification creates a high risk of omission errors and hallucinations.
Organizations should treat advertised context capacities with caution and avoid relying entirely on basic retrieval mechanisms for tasks requiring complete information synthesis. Model developers must train architectures on context lengths that exceed their intended operating windows to establish genuine reliability across the entire input. Moving forward, engineering efforts should prioritize end-to-end long-context training and more sophisticated retrieval architectures capable of ensuring full evidence coverage.
These conclusions are supported by a rigorous evaluation methodology, although the benchmark is limited to three domain categories (financial, legal, and academic) due to high expert annotation costs. Confidence in the relative performance rankings and failure modes of current long-context models remains high under realistic multi-document conditions.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). LongBench established the foundational bilingual benchmark for evaluating multi-task long-context processing that the source builds upon and critiques for relying on synthetic shortcuts.
- Paper: SCROLLS: Standardized CompaRison Over Long Language Sequences, Uri Shaham et al. (2022). SCROLLS provides essential prior work on standardizing evaluations across long-document tasks that require synthesizing dispersed information.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). This paper identifies the 'lost in the middle' phenomenon and positional failure modes in long-context retrieval, providing critical context for the retrieval degradation studied in the source.
- Paper: Large Language Models Can Be Easily Distracted by Irrelevant Context, Freda Shi et al. (2023). This work demonstrates how language models are vulnerable to distractors in contextual inputs, directly informing the distractor-filtering and spotlight tasks evaluated in Loong.
- Paper: Longformer: The Long-Document Transformer, Iz Beltagy et al. (2020). Longformer establishes early architectural mechanisms for processing long sequences, offering foundational background on attention scaling over extended texts.
- Paper: Recursive Language Models, Alex L. Zhang et al. (2025). Recursive Language Models address context degradation by treating long inputs programmatically and recursively executing sub-queries to overcome the multi-document synthesis limits identified in the source.
- Paper: Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale, Siddharth Gollapudi et al. (2026). This study advances past the retrieval failure modes discovered in the source by examining whether models can perform accurate, end-to-end in-context retrieval at a 1-million-token scale.
- Paper: Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context, Keivan Alizadeh et al. (2026). This framework builds upon recursive long-context models by using uncertainty-guided program search to enhance reasoning and prevent omission over extended documents.
- Paper: LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Huiqiang Jiang et al. (2024). LongLLMLingua introduces query-aware prompt compression and reordering to mitigate the context length degradation and multi-document retrieval drop-offs highlighted in the source.
- Paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization, Darren Edge et al. (2024). GraphRAG presents a structured knowledge graph approach to solve the global information aggregation and multi-document synthesis challenges where standard RAG pipelines fail.
- Paper: DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, DeepSeek AI (2026). DeepSeek-V4 develops native million-token context architectures with compressed KV caching to improve model reliability on ultra-long multi-document inputs.
- Paper: World Model on Million-Length Video And Language With Blockwise RingAttention, Hao Liu 0055 et al. (2025). Large World Model scales sequence processing to one million tokens across language and video to address the context capacity bottlenecks detailed in the benchmark.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). This paper investigates methods to make RAG systems resilient to noisy or partially retrieved contexts, targeting the precise retrieval vulnerabilities uncovered in the source.
