Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems

Philippe LabanAlexander R. FabbriCaiming XiongChien-Sheng Wu

article2024EMNLP118 citations

Introduces SummHay, a synthetic long-context benchmark that exposes critical weaknesses in leading language models and retrieval-augmented generation systems by evaluating their ability to aggregate dispersed insights and accurately cite sources across 100,000-token document collections.

Listen

Modern artificial intelligence systems can now ingest massive amounts of text, either by expanding language model context windows to hundreds of thousands of words or by using retrieval-augmented generation to pull relevant snippets from large document collections. However, evaluating how accurately these systems synthesize and attribute information across extensive corpora remains difficult. Popular evaluation tasks often lack the complexity required to test whether an artificial intelligence system can synthesize recurring insights and accurately track their exact sources.

The article introduces a benchmark task called Summary of a Haystack and evaluates how effectively current long-context language models and retrieval-augmented systems can summarize multi-document collections and cite their sources. The objective is to rigorously measure two essential capabilities: identifying key recurring insights across a large corpus and precisely attributing each insight to the correct source documents.

To conduct this evaluation, the researchers developed a synthetic data pipeline across conversational and news domains. They created ten document collections totaling approximately 100,000 words each, embedding specific, controlled facts across individual documents. The evaluation assessed 14 language models and 50 retrieval-augmented generation configurations across 92 distinct summarization queries. System outputs were scored on insight coverage, citation accuracy, and an overall combined score, then compared against human benchmark performance and automated evaluations.

The investigation revealed that current artificial intelligence systems struggle significantly on this task. Leading long-context models processing the full text directly, such as GPT-4o and Claude 3 Opus, scored below 20% on the combined metric. When provided with an idealized oracle retriever, the best-performing models reached approximately 40% to 58%, still lagging well behind the estimated human baseline of 56.1%. Retrieval-augmented systems generally improved citation accuracy by narrowing the document context, but this often came at the expense of comprehensive insight coverage. Furthermore, testing confirmed significant position bias across models, where performance shifted by 9 to 13 points depending on whether relevant documents were placed at the beginning or end of the input window.

These findings indicate that handling long contexts does not guarantee reliable comprehension or accurate source attribution. In enterprise and high-stakes settings, deploying long-context models without effective retrieval introduces substantial risks of hallucinated citations and missed information. High-quality retrieval components, such as advanced neural rerankers, remain critical for practical deployments, though current architectures still require significant optimization to balance comprehensive coverage with precise attribution.

Organizations should not rely solely on full-context models for critical document synthesis workflows. Instead, practitioners should implement hybrid retrieval-augmented generation pipelines with advanced reranking to optimize attribution accuracy. Future development must focus on improving retrieval precision and mitigating position bias so that automated systems can approach human-level reliability in complex document synthesis.

The findings are constrained by synthetic data assumptions, which feature independently generated documents without cross-references or temporal dynamics found in real-world corpora. Additionally, the human baseline was measured in an assisted setting rather than across the full raw text. Nevertheless, the rigorous multi-domain validation provides high confidence that current models face substantial, measurable gaps when performing long-context synthesis and attribution.

Cover for Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems

Abstract

LLMs and RAG systems are now capable of handling millions of input tokens or more. However, evaluating the output quality of such systems on long-context tasks remains challenging, as tasks like Needle-in-a-Haystack lack complexity. In this work, we argue that summarization can play a central role in such evaluation. We design a procedure to synthesize Haystacks of documents, ensuring that specific insights repeat across documents. The “Summary of a Haystack” (SummHay) task then requires a system to process the Haystack and generate, given a query, a summary that identifies the relevant insights and precisely cites the source documents. Since we have precise knowledge of what insights should appear in a haystack summary and what documents should be cited, we implement a highly reproducible automatic evaluation that can score summaries on two aspects – Coverage and Citation. We generate Haystacks in two domains (conversation, news), and perform a large-scale evaluation of 14 LLMs and corresponding 50 RAG systems. Our findings indicate that SummHay is an open challenge for current systems, as even systems provided with an Oracle signal of document relevance lag our estimate of human performance (56%) by 10+ points on a Joint Score. Without a retriever, long-context LLMs like GPT-4o and Claude 3 Opus score below 20% on SummHay. We show SummHay can also be used to study enterprise RAG systems and position bias in long-context models. We hope future systems can equal and surpass human performance on SummHay.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Summarization Evaluation
  • 2.2 Long-Context LLM Evaluation
  • 2.3 Attribution Evaluation
  • 3 Summary in a Haystack Framework
  • 3.1 Preliminaries
  • 3.2 Haystack Generation
  • 3.3 Haystack Summarization (Figure 1, right)
  • 3.4 SummHay Benchmark
  • 4 Evaluation Protocol
  • 4.1 Evaluation Metrics
  • 4.2 Annotation Reproducibility
  • 4.3 Automatic Metric Validation
  • 5 Results
  • 5.1 Experimental Settings
  • 5.2 Benchmark Results
  • 5.3 Estimating Human Performance
  • 5.4 Position Bias Sensitivity
  • 5.5 No-Document Baseline
  • 6 Conclusion
  • 7 Limitations
  • Ethical Considerations
  • References
  • A Appendix
  • A.1 Haystack Synthesis Details
  • A.1.1 Subtopic Verification
  • A.1.2 Insight Verification
  • A.1.3 Document Verification
  • A.2 Evaluation Prompt
  • A.3 Automatic Results Bias
  • A.4 Details on Establishing SummHay Human Performance
  • A.5 Citation Precision & Recall Analysis
  • A.6 Model Access Details
  • A.7 Additional Output Examples
  • A.8 Additional Discussion

Knowls

  1. Knowl 1 — No knowls extractable from the provided attachment

    limitation

    The paper's contribution could not be decomposed into knowls because no readable content was available from the attached file named 9979b40d-0706-49e0-8391-cf2d7c7c6d67.pdf. To preserve fidelity and avoid inventing methods, results, or limitations not stated in the source, no substantive knowls were produced. A valid extraction would require the actual text of the paper.

Coverage note — No substantive contributed material was extracted: the paper's text was not readable from the provided attachment in this environment, so no methods, results, or analyses could be identified without fabricating content.

Citation

MLA
Laban, P., et al. “Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 9885–903, https://doi.org/10.18653/v1/2024.emnlp-main.552.
APA
Laban, P., Fabbri, A. R., Xiong, C., & Wu, C.-S. (2024). Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 9885–9903. https://doi.org/10.18653/v1/2024.emnlp-main.552
Chicago
Laban, P., A. R. Fabbri, C. Xiong, and C.-S. Wu. 2024. “Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 9885–9903. https://doi.org/10.18653/v1/2024.emnlp-main.552.
Harvard
Laban, P. et al. (2024) “Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 9885–9903. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.552.
Vancouver
1. Laban P, Fabbri AR, Xiong C, Wu C-S (2024) Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 9885–9903

BibTeX

@inproceedings{laban-etal-2024-summary,
    title = "Summary of a Haystack: A Challenge to Long-Context {LLM}s and {RAG} Systems",
    author = "Laban, Philippe  and
      Fabbri, Alexander  and
      Xiong, Caiming  and
      Wu, Chien-Sheng",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.552/",
    doi = "10.18653/v1/2024.emnlp-main.552",
    pages = "9885--9903"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/