Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA

Minzheng WangLongze ChenFu ChengShengyi LiaoXinghua ZhangBingli WuHaiyang YuNan XuLei ZhangRun Luo

article2024EMNLP124 citations

Introduces Loong, a realistic long-context benchmark that tests large language models across distributed multi-document scenarios where every included text is necessary to derive the correct answer.

Listen

Modern large language models frequently claim the ability to process massive amounts of information simultaneously, often advertising context windows spanning hundreds of thousands of words. However, existing industry benchmarks evaluate these models using synthetic shortcuts, such as hiding a single fact inside unrelated filler text or concentrating evidence within a single document. These evaluation methods fail to reflect realistic enterprise workflows—such as aggregating multi-year financial statements, comparing case law, or synthesizing research papers—where missing even one document leads to incorrect conclusions.

The article introduces and evaluates Loong, a realistic benchmark designed to rigorously assess how language models handle extended multi-document question answering. The benchmark tests whether models can synthesize distributed information across realistic contexts where every provided document contains essential evidence.

To construct this benchmark, the authors collected 1,600 verified test cases across financial reports, legal cases, and academic papers, spanning both English and Chinese. The evaluation spans input lengths from 10,000 to over 250,000 tokens and tests four distinct capabilities: locating a specific document among distractors (Spotlight Locating), comparing information across multiple sources (Comparison), grouping distributed evidence (Clustering), and multi-step reasoning over sequential data (Chain of Reasoning). Using this dataset, the authors benchmarked seven advanced models, including proprietary systems like Gemini-1.5-pro and GPT-4o, alongside leading open-source models, while also evaluating the impact of retrieval-augmented generation (RAG).

The analysis reveals that current language models struggle significantly in realistic multi-document scenarios. Even the top-performing model, Gemini-1.5-pro, achieved an overall average score of only 55.37 out of 100, with a perfect completion rate of just 27%, while GPT-4o scored 53.47 with a 26% perfect rate. Model performance drops sharply as context lengths expand; for example, models trained on 128,000-token windows degraded significantly once inputs exceeded 50,000 tokens, revealing a substantial gap between advertised window sizes and effective processing capabilities. Furthermore, while models performed relatively well on simple single-document lookup tasks, they degraded on complex clustering and comparison tasks that require multi-source synthesis. Finally, adding retrieval-augmented generation caused overall performance to decline—lowering GPT-4o's score from 53.47 to between 32.85 and 46.52 depending on retrieval settings—because standard search mechanisms failed to retrieve all required documents from the collection.

These findings indicate that organizations face material operational and compliance risks if they rely on advertised context window sizes for tasks that demand comprehensive multi-document synthesis. Standard retrieval pipelines cannot reliably replace true long-context processing when evidence is dispersed across many sources. For enterprise leaders, deploying models for multi-document auditing, legal discovery, or financial synthesis without strict verification creates a high risk of omission errors and hallucinations.

Organizations should treat advertised context capacities with caution and avoid relying entirely on basic retrieval mechanisms for tasks requiring complete information synthesis. Model developers must train architectures on context lengths that exceed their intended operating windows to establish genuine reliability across the entire input. Moving forward, engineering efforts should prioritize end-to-end long-context training and more sophisticated retrieval architectures capable of ensuring full evidence coverage.

These conclusions are supported by a rigorous evaluation methodology, although the benchmark is limited to three domain categories (financial, legal, and academic) due to high expert annotation costs. Confidence in the relative performance rankings and failure modes of current long-context models remains high under realistic multi-document conditions.

Cover for Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA

Abstract

Long-context modeling capabilities have garnered widespread attention, leading to the emergence of Large Language Models (LLMs) with ultra-context windows. Meanwhile, benchmarks for evaluating long-context LLMs are gradually catching up. However, existing benchmarks employ irrelevant noise texts to artificially extend the length of test cases, diverging from the real-world scenarios of long-context applications. To bridge this gap, we propose a novel long-context benchmark, Loong, aligning with realistic scenarios through extended multi-document question answering (QA). Unlike typical document QA, in Loong's test cases, each document is relevant to the final answer, ignoring any document will lead to the failure of the answer. Furthermore, Loong introduces four types of tasks with a range of context lengths: Spotlight Locating, Comparison, Clustering, and Chain of Reasoning, to facilitate a more realistic and comprehensive evaluation of long-context understanding. Extensive experiments indicate that existing long-context language models still exhibit considerable potential for enhancement. Retrieval augmented generation (RAG) achieves poor performance, demonstrating that Loong can reliably assess the model's long-context modeling capabilities.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Long-Context Language Models
  • 2.2 Long-Context Benchmarks
  • 2.3 Retrieval Augmented Language Models
  • 3 Loong: A Long-Context Benchmark
  • 3.1 Overview
  • 3.2 Evaluation Task
  • 3.2.1 Spotlight Locating
  • 3.2.2 Comparison
  • 3.2.3 Clustering
  • 3.2.4 Chain of Reasoning
  • 3.3 Benchmark Construction
  • 3.3.1 Data Collection
  • 3.3.2 Annotation Process
  • 3.3.3 Quality Control
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Main Results
  • 4.3 Task Analysis
  • 4.4 Scaling Law of Context Window
  • 4.5 RAG or Not
  • 5 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A GPT4-as-the-Judge Prompt
  • B Test Case
  • B.1 Spotlight Locating
  • B.2 Sequential Enumeration
  • B.3 Extremum Acquisition
  • B.4 Range Awareness
  • B.5 Report Integration
  • B.6 Citation&Reference
  • B.7 Case Classification
  • B.8 Temporal Analysis
  • B.9 Citation Chain
  • B.10 Link the Links
  • B.11 Solitaire
  • C Length Distribution
  • D RAG Detailed Results
  • E Comparison of Evidence Distribution
  • F Comparison of Results with Other Benchmarks
  • G Results of Recall Rate by RAG
  • H Detailed URL of Document Source

Knowls

  1. Knowl 1 — Loong Benchmark Overview and Multi-Document Evidence Dispersion Design

    definition

    Loong is a bilingual (English and Chinese) question-answering benchmark designed to evaluate the long-context understanding capabilities of Large Language Models (LLMs) across multi-document inputs. Unlike conventional long-context benchmarks where the evidence is concentrated in a single passage or document with added noise, Loong is designed under the "leave no document behind" principle:

    1. Evidence Dispersion: Every document in an input set contains necessary evidence; omitting or misinterpreting any document leads to failure on the final answer. An average test case comprises approximately 11 documents.
    2. Real-World Domains: Documents are curated from three realistic domains: Financial Reports (700 test instances), Legal Cases (500 test instances), and Academic Papers (400 test instances), primarily sourced from 2024 filings, judgments, and publications.
    3. Length Distribution: The benchmark consists of 1,600 total test instances categorized into four context length intervals:
      • Set 1 (10K–50K tokens; average length 37.8K tokens; 323 instances)
      • Set 2 (50K–100K tokens; average length 75.6K tokens; 564 instances)
      • Set 3 (100K–200K tokens; average length 138.9K tokens; 481 instances)
      • Set 4 (200K–250K tokens; average length 233.9K tokens; 232 instances)
  2. Knowl 2 — Loong Task Taxonomy and Subtask Structure

    definition

    Loong categorizes multi-document evaluation into four high-level task categories comprising 10 distinct subtasks:

    1. Spotlight Locating (250 instances; average 119.3K tokens; English & Chinese): Assesses atomic evidence search across multiple documents by placing the required information within a single target document while other semantically similar documents from the same domain act as noise.
    2. Comparison (300 instances; average 110.6K tokens; English & Chinese): Requires locating dispersed evidence across multiple documents, correlating them, and making comparisons. Subtasks include:
      • Sequential Enumeration (87 instances): Listing specific numerical or textual values of an attribute across documents in a specified order.
      • Extremum Acquisition (143 instances): Identifying the maximum or minimum value of a specific attribute among all documents.
      • Range Awareness (70 instances): Identifying all entities or objects that meet a specified numerical or conceptual condition.
    3. Clustering (641 instances; average 109.8K tokens; English & Chinese): Evaluates multi-source information extraction and grouping based on prescribed criteria. Subtasks include:
      • Report Integration (250 instances): Grouping financial report data into structured categories based on textual or numerical criteria.
      • Citation & Reference (270 instances; English): Identifying mutual citation and reference links between candidate academic papers.
      • Case Classification (121 instances; Chinese): Categorizing legal judgment documents based on given legal causes of action.
    4. Chain of Reasoning (409 instances; average 103.9K tokens; English & Chinese): Evaluates multi-hop sequential reasoning across documents. Subtasks include:
      • Temporal Analysis (100 instances): Analyzing multi-year or multi-quarter temporal trends for target attributes.
      • Citation Chain (130 instances; English): Deducing linear, continuous citation paths (A→B→CA \to B \to C) across papers.
      • Link the Links (113 instances; Chinese): Matching decoupled case factual descriptions with their corresponding court verdicts.
      • Solitaire (66 instances; Chinese): Sorting and matching judgment documents sequentially according to a given sequence of legal action types.
  3. Knowl 3 — Loong Benchmark Construction and Quality Control Pipeline

    model/method

    The construction of the Loong benchmark utilizes a semi-automated pipeline combined with multi-stage verification:

    1. Data Curation Criteria: Documents are collected from official authoritative sources (U.S. SEC and cninf for financial reports; China Judge Online for legal rulings; arXiv and Semantic Scholar for academic papers) based on six criteria: timeliness (overwhelmingly from the year 2024 to minimize pre-training data contamination), public accessibility, long length, parseability, domain categorizability, and source authority. Personal identifiable information is scrubbed during pre-processing.
    2. Annotation Workflows:
      • Attribute Compression (Financial Reports): Key attributes across long documents are extracted via GPT-4o into compressed structured records, allowing subsequent question-answer pair construction without re-reading the entire raw text.
      • Structural Segmentation (Legal Documents): Judgments are separated into fact descriptions and verdicts using rule-based parsing.
      • Citation Graph Traversal (Academic Papers): Semantic Scholar API and LaTeX .bbl files are parsed to extract verified citation edges and linear citation paths.
      • Q&A Generation: Conducted via rule-based templates on structured data and GPT-4o free-form prompt generation.
    3. Quality Control Protocols:
      • Evidence Recall Prompting: Generation prompts require GPT-4o to quote the specific supporting evidence passages alongside the labels.
      • Self-Check: GPT-4o re-evaluates generated candidate labels against source text passages.
      • Human Verification: Domain experts inspect and filter questions for factual accuracy and clarity, selecting 1,600 verified entries from an initial candidate pool of 2,814 entries.
  4. Knowl 4 — LLM Evaluation Protocol and GPT-4-as-a-Judge Metric Formulation

    experimental setup

    To overcome the limitations of lexical overlap metrics (F1F_1 score and ROUGE-L) on long-context QA, Loong employs GPT-4 as an automated evaluator scoring responses against reference gold answers on a 0–100 scale based on three criteria:

    1. Accuracy: Semantic consistency with the gold answer.
    2. Hallucination Absence: Factual correctness without fabricating unsupported facts or incorrect numerical values/orders.
    3. Completeness: Inclusion of all mandatory key points required by the reference answer.

    The evaluation outputs two metrics:

    • Average Score (Avg Score): The arithmetic mean of scores assigned across test queries: Avg Score=1N∑i=1NSi\text{Avg Score} = \frac{1}{N} \sum_{i=1}^N S_i where Si∈[0,100]S_i \in [0, 100] is the GPT-4 assigned score for test sample ii, and NN is the total number of test samples.
    • Perfect Rate: The proportion of test cases that achieve a full score of 100: Perfect Rate=1N∑i=1NI(Si=100)\text{Perfect Rate} = \frac{1}{N} \sum_{i=1}^N \mathbb{I}(S_i = 100) where I(⋅)\mathbb{I}(\cdot) is the indicator function.

    Evaluations are conducted with decoding temperature set to 0. For inputs exceeding model context windows, documents are concatenated in order and truncated when the capacity threshold is reached.

  5. Knowl 5 — Long-Context Model Performance Across Loong Task Categories

    data/table

    Evaluation of frontier closed-source and open-source long-context LLMs across the four task categories of Loong demonstrates that existing models struggle on multi-document reasoning, with even top models failing to achieve passing scores on strict metrics.

    Model Spotlight Loc. Comparison Clustering Chain of Reas. Overall
    Avg Perf Avg Perf Avg Perf Avg Perf Avg Perf
    GPT-4o (128K) 73.95 0.62 50.50 0.28 44.29 0.09 57.95 0.28 53.47 0.26
    Gemini-1.5-pro (1000K) 75.02 0.56 49.94 0.27 44.10 0.09 64.97 0.37 55.37 0.27
    Claude3.5-Sonnet (200K) 58.45 0.49 54.21 0.35 45.77 0.07 43.92 0.25 48.85 0.23
    Qwen2-72B-Instruct (128K) 54.17 0.36 42.38 0.20 36.71 0.04 47.76 0.18 43.29 0.15
    Claude3-Haiku (200K) 68.68 0.59 42.10 0.21 35.04 0.02 47.59 0.17 44.88 0.19
    Kimi-Chat (200K) 60.98 0.50 34.74 0.13 28.76 0.04 38.52 0.15 37.49 0.16
    GLM4-9B-Chat (1000K) 57.35 0.47 40.38 0.20 28.52 0.02 39.94 0.16 38.31 0.16

    Key empirical observations:

    • Spotlight Locating yields the highest scores across all models (e.g., 75.02 for Gemini-1.5-pro and 73.95 for GPT-4o) because evidence resides within a single document.
    • Clustering is the most difficult task category (overall perfect rate ≤0.09\le 0.09 across all models), as it requires exhaustive multi-document extraction and categorization.
    • Gemini-1.5-pro achieves the highest overall Avg Score (55.37) and Perfect Rate (0.27), followed closely by GPT-4o (53.47 Avg, 0.26 Perf). Open-source models trail closed-source models across all task types.
  6. Knowl 6 — Context Length Scaling Degradation and Context Window Ineffective Zones

    empirical result

    Model performance across different context length sets reveals a systematic degradation as input length expands, exposing an "ineffective zone" within nominal context window limits:

    Model Set1 (10K–50K) Set2 (50K–100K) Set3 (100K–200K) Set4 (200K–250K)
    Avg Perf Avg Perf Avg Perf Avg Perf
    GPT-4o (128K) 70.40 0.44 58.38 0.29 46.95 0.19 31.11 0.07
    Gemini-1.5-pro (1000K) 63.36 0.34 55.56 0.26 52.05 0.24 50.70 0.25
    Claude3.5-Sonnet (200K) 63.69 0.37 52.73 0.24 42.06 0.19 30.51 0.08
    Qwen2-72B-Instruct (128K) 60.11 0.29 45.71 0.17 35.94 0.09 28.92 0.06
    Claude3-Haiku (200K) 57.14 0.28 45.45 0.17 41.41 0.18 32.15 0.10
    Kimi-Chat (200K) 55.02 0.24 42.40 0.16 31.37 0.14 13.50 0.05
    GLM4-9B-Chat (1000K) 51.43 0.25 40.19 0.17 37.36 0.16 16.84 0.05

    Key insights on context window scaling:

    1. Ineffective Zone: Models nominalized for 128K context (GPT-4o and Qwen2-72B-Instruct) begin steep degradation as early as the 50K–100K token range (Set 2), indicating that their effective context handling capacity is substantially smaller than their nominal context window.
    2. Context Scaling Law: To reliably support a target context window length LL, models must be pre-trained or fine-tuned on sequences strictly exceeding LL. Gemini-1.5-pro, trained on 1,000K tokens, maintains stable performance across Set 3 (52.05) and Set 4 (50.70), exhibiting minimal degradation compared to models trained on shorter windows.
  7. Knowl 7 — Retrieval-Augmented Generation vs. Full-Context Modeling in Dispersed Evidence QA

    empirical result

    Integrating dense Retrieval-Augmented Generation (RAG) using OpenAI text-embedding-ada-002 or BGE-M3 embeddings across chunk sizes of 1024 tokens and retrieval parameters k∈{5,10,30,50}k \in \{5, 10, 30, 50\} leads to degraded overall performance compared to native long-context modeling:

    • Overall Score Drop: On GPT-4o, baseline native processing achieves an overall Avg Score of 53.47 and Perfect Rate of 0.26. Adding RAG reduces performance across all configurations:
      • Top-k=5k=5: 32.85 Avg / 0.11 Perf (OpenAI), 34.01 Avg / 0.13 Perf (BGE)
      • Top-k=10k=10: 38.80 Avg / 0.14 Perf (OpenAI), 38.71 Avg / 0.15 Perf (BGE)
      • Top-k=30k=30: 44.62 Avg / 0.18 Perf (OpenAI), 45.67 Avg / 0.21 Perf (BGE)
      • Top-k=50k=50: 42.70 Avg / 0.18 Perf (OpenAI), 46.52 Avg / 0.21 Perf (BGE)
    • Qwen2-72B-Instruct: Baseline native achieves 43.29 Avg / 0.15 Perf. With RAG, scores range from 33.22 (Top-k=5k=5) to 42.98 (Top-k=30k=30).
    • Context Fragmentation: In dispersed multi-document QA, chunk-level retrieval breaks cross-document information flow and drops critical intermediate evidence.
    • Length-Dependent Behavior: For inputs within the model's native window (Set 1 and Set 2), RAG substantially underperforms full-context input. RAG only offers minor score preservation in ultra-long contexts exceeding the model's truncation threshold (Set 4, >200K tokens) where truncation would otherwise discard documents completely.
  8. Knowl 8 — Multi-Document Recall Bottleneck of Dense Retrieval in Long Contexts

    data/table

    An analysis of dense retrieval models on financial multi-document test cases reveals that standard retrieval fails to recall passages from all necessary documents when evidence is distributed across all documents.

    Let Recall@n\text{Recall@}n denote the proportion of test queries for which the top-nn retrieved passages contain at least one chunk from every relevant document in the context (11 if all documents are represented, 00 otherwise).

    Retriever Model Recall@10 Recall@20 Recall@30 Recall@40 Recall@50
    OpenAI Embedding (chunk=1024) 0.20 0.34 0.43 0.51 0.58
    OpenAI Embedding (chunk=2048) 0.20 0.36 0.48 0.57 0.62
    BGE Embedding (chunk=1024) 0.19 0.34 0.44 0.53 0.59
    BGE Embedding (chunk=2048) 0.21 0.38 0.49 0.58 0.64

    Even at n=50n=50 retrieved passages, the maximum document coverage rate is only 0.64 (64%). Because retrieving a document chunk does not guarantee capturing the exact evidence passage within that document, the true multi-document evidence recall rate is strictly bounded below this coverage rate, explaining the fundamental performance ceiling of RAG on multi-document reasoning tasks.

  9. Knowl 9 — Comparative Performance Across Long-Context Benchmarks (Loong vs. RULER vs. NOCHA)

    data/table

    Comparing evaluation results across different long-context benchmarks demonstrates distinct evaluation dynamics between real-world multi-document tasks (Loong), synthetic retrieval tasks (RULER), and single-domain narrative comprehension (NOCHA):

    Model Loong RULER NOCHA
    GPT-4o 53.47 – 55.80
    GPT-4-Turbo – 91.60 40.20
    Gemini-1.5-pro 55.37 95.80 48.10
    Claude3.5-Sonnet 48.85 – 41.00
    Qwen2-72B-Instruct 43.29 85.90 43.35
    GLM4-9B-Chat 38.31 89.90 27.07
    Llama3.1-8B 36.31 88.30 16.53
    Phi-3-mini-3.8B 14.54 68.80 9.30

    Key observations:

    1. On synthetic tasks like RULER, smaller models can achieve artificially high scores via synthetic needle-in-a-haystack fine-tuning (e.g., GLM4-9B-Chat scoring 89.90 and Llama3.1-8B scoring 88.30, surpassing Qwen2-72B-Instruct at 85.90). In contrast, on Loong, performance scales strictly with model capacity and parameter size.
    2. Domain variance affects rankings: GPT-4o leads on novel narrative text (55.80 on NOCHA), whereas Gemini-1.5-pro leads on structured multi-document reasoning (55.37 on Loong and 95.80 on RULER).
  10. Knowl 10 — Limitations of the Loong Benchmark

    limitation

    The Loong benchmark has two primary stated limitations:

    1. Domain Coverage: Due to annotation overhead and evaluation efficiency constraints, the benchmark is restricted to three domains: financial reports, legal judgment cases, and academic papers. While representative of long-context multi-document workloads, many other real-world multi-document domains are not covered.
    2. Annotation and Verification Cost: Ensuring ground-truth reliability across multi-document contexts with average sequence lengths exceeding 100K tokens required recruiting domain experts proficient in both Chinese and English to manually review and verify long document sets, restricting the feasibility of scaling the benchmark to larger sample sizes.

Coverage note — No substantial contributed material was omitted. Detailed task prompt templates from Appendix B were synthesized within the task definitions and evaluation setup knowls.

References

  1. 1.Chenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao, Mukai Li, Jun Zhang, Lingpeng Kong, and Xipeng Qiu. 2024. L-eval: Instituting standardized evaluation for long context language models. In Proceedings of ACL, pages 14388–14411.
  2. 2.AI Anthropic. 2024a. The claude 3 model family: Opus, sonnet, haiku. Claude-3 Model Card.
  3. 3.AI Anthropic. 2024b. Claude 3.5 sonnet model card addendum. Claude-3.5 Model Card.
  4. 4.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  5. 5.Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2024. LongBench: A bilingual, multitask benchmark for long context understanding. In Proceedings of ACL, pages 3119–3137.
  6. 6.bloc97. 2023. Ntk-aware scaled rope. https://www.reddit.com/r/LocalLLaMA/comments/14lz7j5/ntkaware_scaled_rope_allows_llama_models_to_have/.
  7. 7.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In Proceedings of ICML, pages 2206–2240.
  8. 8.Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024a. Benchmarking large language models in retrieval-augmented generation. In Proceedings of AAAI, pages 17754–17762.
  9. 9.Longze Chen, Ziqiang Liu, Wanwei He, Yinhe Zheng, Hao Sun, Yunshui Li, Run Luo, and Min Yang. 2024b. Long context is not long at all: A prospector of long-dependency data for large language models. In Proceedings of ACL, pages 8222–8234.
  10. 10.Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595.
  11. 11.Zican Dong, Tianyi Tang, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen. 2024. BAMBOO: A comprehensive benchmark for evaluating long text modeling capacities of large language models. In Proceedings of LREC-COLING, pages 2086–2099.
  12. 12.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of ACL, pages 320–335.
  13. 13.Shahriar Golchin and Mihai Surdeanu. 2023. Time travel in llms: Tracing data contamination in large language models. arXiv preprint arXiv:2308.08493.
  14. 14.Chi Han, Qifan Wang, Hao Peng, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang. 2024. LM-infinite: Zero-shot extreme length generalization for large language models. In Proceedings of NAACL, pages 3991–4008.
  15. 15.Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg. 2024. RULER: What’s the real context size of your long-context language models? In Proceedings of COLM.
  16. 16.Greg Kamradt. 2023. Needle in a haystack - pressure testing llms. https://github.com/gkamradt/LLMTest_NeedleInAHaystack.
  17. 17.Marzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2024. One thousand and one pairs: A "novel" challenge for long-context language models. arXiv preprint arXiv: 2406.16264.
  18. 18.Jiaqi Li, Mengmeng Wang, Zilong Zheng, and Muhan Zhang. 2024. LooGLE: Can long-context language models understand long contexts? In Proceedings of ACL, pages 16304–16333.
  19. 19.Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2024. Calibrating LLM-based evaluator. In Proceedings of LREC-COLING, pages 2638–2656.
  20. 20.OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  21. 21.Arka Pal, Deep Karkhanis, Manley Roberts, Samuel Dooley, Arvind Sundararajan, and Siddartha Naidu. 2023. Giraffe: Adventures in expanding context lengths in llms. arXiv preprint arXiv:2308.10882.
  22. 22.Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2024. YaRN: Efficient context window extension of large language models. In Proceedings of ICLR.
  23. 23.Ofir Press, Noah Smith, and Mike Lewis. 2022. Train short, test long: Attention with linear biases enables input length extrapolation. In Proceedings of ICLR.
  24. 24.Zexuan Qiu, Jingjing Li, Shijue Huang, Wanjun Zhong, and Irwin King. 2024. Clongeval: A chinese benchmark for evaluating long-context large language models. arXiv preprint arXiv:2403.03514.
  25. 25.Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. Parallel context windows for large language models. In Proceedings of ACL, pages 6383–6402.
  26. 26.Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530.
  27. 27.Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950.
  28. 28.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2024. REPLUG: Retrieval-augmented black-box language models. In Proceedings of NAACL, pages 8371–8384.
  29. 29.Mingyang Song, Mao Zheng, and Xuan Luo. 2024. Counting-stars: A simple, efficient, and reasonable strategy for evaluating long-context large language models. arXiv preprint arXiv:2403.11802.
  30. 30.Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063.
  31. 31.Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei. 2023. A length-extrapolatable transformer. In Proceedings of ACL, pages 14590–14604.
  32. 32.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of NeurIPs, page 30.
  33. 33.Cunxiang Wang, Sirui Cheng, Qipeng Guo, Yuanhao Yue, Bowen Ding, Zhikun Xu, Yidong Wang, Xiangkun Hu, Zheng Zhang, and Yue Zhang. 2024. Evaluating open-qa evaluation. In Proceedings of NeurIPs, page 36.
  34. 34.Kevin Wu, Eric Wu, and James Zou. 2024. How faithful are rag models? quantifying the tug-of-war between rag and llms’ internal prior. arXiv preprint arXiv:2404.10198.
  35. 35.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2023. Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453.
  36. 36.Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, and Hao Ma. 2024. Effective long-context scaling of foundation models. In Proceedings of NAACL, pages 4643–4663.
  37. 37.Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, and Bryan Catanzaro. 2024. Retrieval meets long context large language models. In Proceedings of ICLR.
  38. 38.Lei Zhang, Yunshui Li, Ziqiang Liu, Jiaxi Yang, Junhao Liu, Longze Chen, Run Luo, and Min Yang. 2024a. Marathon: A race through the realm of long context with large language models. In Proceedings of ACL, pages 5201–5217.
  39. 39.Xinghua Zhang, Bowen Yu, Haiyang Yu, Yangyu Lv, Tingwen Liu, Fei Huang, Hongbo Xu, and Yongbin Li. 2023. Wider and deeper llm networks are fairer llm evaluators. arXiv preprint arXiv:2308.01862.
  40. 40.Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Hao, Xu Han, Zhen Thai, Shuo Wang, Zhiyuan Liu, and Maosong Sun. 2024b. ∞Bench: Extending long context evaluation beyond 100K tokens. In Proceedings of ACL, pages 15262–15277.
  41. 41.Dawei Zhu, Nan Yang, Liang Wang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. 2024. PoSE: Efficient context window extension of LLMs via positional skip-wise training. In Proceedings of ICLR.

Citation

MLA
Wang, M., et al. “Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 5627–46, https://doi.org/10.18653/v1/2024.emnlp-main.322.
APA
Wang, M., Chen, L., Cheng, F., Liao, S., Zhang, X., Wu, B., Yu, H., Xu, N., Zhang, L., Luo, R., Li, Y., Yang, M., Huang, F., & Li, Y. (2024). Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 5627–5646. https://doi.org/10.18653/v1/2024.emnlp-main.322
Chicago
Wang, M., L. Chen, F. Cheng, et al. 2024. “Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 5627–46. https://doi.org/10.18653/v1/2024.emnlp-main.322.
Harvard
Wang, M. et al. (2024) “Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 5627–5646. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.322.
Vancouver
1. Wang M, Chen L, Cheng F, et al (2024) Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 5627–5646

BibTeX

@inproceedings{wang-etal-2024-leave,
    title = "Leave No Document Behind: Benchmarking Long-Context {LLM}s with Extended Multi-Doc {QA}",
    author = "Wang, Minzheng  and
      Chen, Longze  and
      Cheng, Fu  and
      Liao, Shengyi  and
      Zhang, Xinghua  and
      Wu, Bingli  and
      Yu, Haiyang  and
      Xu, Nan  and
      Zhang, Lei  and
      Luo, Run  and
      Li, Yunshui  and
      Yang, Min  and
      Huang, Fei  and
      Li, Yongbin",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.322/",
    doi = "10.18653/v1/2024.emnlp-main.322",
    pages = "5627--5646"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/