Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
Philippe LabanAlexander R. FabbriCaiming XiongChien-Sheng Wu
Introduces SummHay, a synthetic long-context benchmark that exposes critical weaknesses in leading language models and retrieval-augmented generation systems by evaluating their ability to aggregate dispersed insights and accurately cite sources across 100,000-token document collections.
Modern artificial intelligence systems can now ingest massive amounts of text, either by expanding language model context windows to hundreds of thousands of words or by using retrieval-augmented generation to pull relevant snippets from large document collections. However, evaluating how accurately these systems synthesize and attribute information across extensive corpora remains difficult. Popular evaluation tasks often lack the complexity required to test whether an artificial intelligence system can synthesize recurring insights and accurately track their exact sources.
The article introduces a benchmark task called Summary of a Haystack and evaluates how effectively current long-context language models and retrieval-augmented systems can summarize multi-document collections and cite their sources. The objective is to rigorously measure two essential capabilities: identifying key recurring insights across a large corpus and precisely attributing each insight to the correct source documents.
To conduct this evaluation, the researchers developed a synthetic data pipeline across conversational and news domains. They created ten document collections totaling approximately 100,000 words each, embedding specific, controlled facts across individual documents. The evaluation assessed 14 language models and 50 retrieval-augmented generation configurations across 92 distinct summarization queries. System outputs were scored on insight coverage, citation accuracy, and an overall combined score, then compared against human benchmark performance and automated evaluations.
The investigation revealed that current artificial intelligence systems struggle significantly on this task. Leading long-context models processing the full text directly, such as GPT-4o and Claude 3 Opus, scored below 20% on the combined metric. When provided with an idealized oracle retriever, the best-performing models reached approximately 40% to 58%, still lagging well behind the estimated human baseline of 56.1%. Retrieval-augmented systems generally improved citation accuracy by narrowing the document context, but this often came at the expense of comprehensive insight coverage. Furthermore, testing confirmed significant position bias across models, where performance shifted by 9 to 13 points depending on whether relevant documents were placed at the beginning or end of the input window.
These findings indicate that handling long contexts does not guarantee reliable comprehension or accurate source attribution. In enterprise and high-stakes settings, deploying long-context models without effective retrieval introduces substantial risks of hallucinated citations and missed information. High-quality retrieval components, such as advanced neural rerankers, remain critical for practical deployments, though current architectures still require significant optimization to balance comprehensive coverage with precise attribution.
Organizations should not rely solely on full-context models for critical document synthesis workflows. Instead, practitioners should implement hybrid retrieval-augmented generation pipelines with advanced reranking to optimize attribution accuracy. Future development must focus on improving retrieval precision and mitigating position bias so that automated systems can approach human-level reliability in complex document synthesis.
The findings are constrained by synthetic data assumptions, which feature independently generated documents without cross-references or temporal dynamics found in real-world corpora. Additionally, the human baseline was measured in an assisted setting rather than across the full raw text. Nevertheless, the rigorous multi-domain validation provides high confidence that current models face substantial, measurable gaps when performing long-context synthesis and attribution.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). Its controlled tests of how relevant evidence gets used across long contexts establish the position-bias problem that SummHay also evaluates.
- Paper: SCROLLS: Standardized CompaRison Over Long Language Sequences, Uri Shaham et al. (2022). SCROLLS shows how to benchmark synthesis and reasoning over naturally long inputs, providing an important foundation for SummHay’s long-context evaluation.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). ALCE establishes reproducible measures of citation precision and recall, clarifying the citation-quality dimension SummHay scores.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). LongBench’s broad long-context task suite provides useful context for SummHay’s more focused test of multi-document synthesis.
- Paper: Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation, Satyapriya Krishna et al. (2025). FRAMES extends benchmark evaluation to end-to-end multi-document retrieval and reasoning, building on SummHay’s test of synthesis across retrieved sources.
- Paper: A Benchmark for Deep Information Synthesis, Debjit Paul et al. (2026). DEEPSYNTH carries the challenge of evaluating information synthesis into realistic, multi-source analytical tasks.
