SCROLLS: Standardized CompaRison Over Long Language Sequences
Uri ShahamElad SegalMaor IvgiAvia EfratOri YoranAdi HavivAnkit GuptaWenhan XiongMor GevaJonathan Berant
Introduces a standardized benchmark and live leaderboard across seven diverse, naturally long-text tasks to systematically evaluate how effectively language models synthesize and reason over extended sequences.
Most modern natural language processing benchmarks evaluate language models on short inputs such as individual sentences or brief paragraphs. In practical business, legal, and research settings, however, critical information is embedded within lengthy documents such as meeting transcripts, corporate contracts, scientific papers, and government reports. While many recent model architectures claim the ability to process long sequences, previous evaluation methods relied on artificial tasks, narrow academic summarization datasets, or metrics that only test local word prediction. Consequently, stakeholders have lacked a reliable way to assess whether language models can genuinely understand and reason across extended texts.
The article introduces SCROLLS (Standardized CompaRison Over Long Language Sequences), a benchmark designed to evaluate and compare language models on complex reasoning and synthesis across naturally long documents. To accomplish this, the authors handpicked and standardized seven datasets spanning multiple domains, including literature, television scripts, scientific articles, parliamentary meetings, and legal agreements. These datasets cover summarization, open and multiple-choice question answering, and natural language inference. All tasks were converted into a uniform text-to-text format, and the authors established baseline performance levels using standard and length-efficient transformer models across varying sequence lengths.
The analysis revealed that relevant information in the SCROLLS datasets is dispersed across hundreds or thousands of words, making information fusion essential. Baseline experiments demonstrated that providing models with longer context generally improves performance; for example, expanding the input capacity of the Longformer Encoder-Decoder model from 1,024 to 16,384 tokens increased its overall score by 2.1 points. However, a standard BART baseline processing only 1,024 tokens achieved an overall score of 29.01, performing within 0.15 points of Longformer's top score of 29.16 despite processing one-sixteenth of the text. When comparing both models at the same 1,024-token length, BART outperformed Longformer by almost two points. Crucially, all tested models fell drastically short of human performance; on question answering tasks, baseline models scored between 18% and 26%, whereas human agreement ranged between 58% and 93%.
These findings imply that increasing the maximum sequence length of a model does not automatically improve its semantic understanding. Organizations risk incurring high computational costs to process massive context windows without obtaining corresponding improvements in answer accuracy or summary quality. Efficient architectures initialized from standard models without long-sequence pretraining struggle to fully leverage extended context, particularly on smaller datasets where they cannot easily adapt to sparse attention patterns.
To address these shortcomings, developers and researchers should adopt the SCROLLS benchmark to guide the development of specialized pretraining strategies, novel architectures, and hybrid retrieval-augmented methods. The primary limitations of the benchmark include its restriction to English-language materials and the reliance on n-gram overlap metrics such as ROUGE for evaluating long summaries, which may undervalue valid paraphrasing. Nonetheless, the evidence strongly supports the conclusion that current natural language models remain far from solving long-text understanding, highlighting the need for continued architectural and methodological innovation.
- Paper: Longformer: The Long-Document Transformer, Iz Beltagy et al. (2020). Introduces the Longformer and Longformer-Encoder-Decoder (LED) architectures that serve as the primary long-sequence baseline evaluated in the SCROLLS benchmark.
- Paper: Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps, Xanh Ho et al. (2020). Establishes the 2WikiMultiHopQA dataset, which is directly incorporated into the SCROLLS benchmark suite as a key multi-hop reasoning task.
- Paper: Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization, Shashi Narayan et al. (2018). Provides foundational principles and methodology for extreme abstractive summarization across documents, motivating SCROLLS's synthesis-heavy summarization tasks.
- Paper: Hierarchical Neural Story Generation, Angela Fan et al. (2018). Introduces the WritingPrompts dataset and long-form narrative generation paradigms that informed the selection of creative and narrative tasks within long-text suites like SCROLLS.
- Paper: HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering, Zhilin Yang et al. (2018). Pioneers multi-hop question answering requiring cross-paragraph reasoning, establishing the task paradigm adapted by SCROLLS for evaluating long context synthesis.
- Paper: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context, Zihang Dai et al. (2019). Presents Transformer-XL, an early foundational architecture addressing context window limits that motivated subsequent long-sequence modeling benchmarks.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). Extends the standardized evaluation of long-context language models pioneered by SCROLLS into a comprehensive bilingual, multi-task benchmark spanning longer context lengths.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). Directly evaluates and analyzes the empirical limitations of language models on long-context multi-document tasks like those featured in SCROLLS.
- Paper: LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Huiqiang Jiang et al. (2024). Proposes a prompt compression method to mitigate performance degradation and position bias when evaluating large models on long-context benchmarks.
- Paper: Qwen2.5 Technical Report, Qwen et al. (2024). Presents architecture and scaling techniques for models designed to handle sequence lengths well beyond the initial baselines evaluated on SCROLLS.
- Paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization, Darren Edge et al. (2024). Applies hierarchical knowledge graph indexing to address the query-focused long-document summarization challenges highlighted in SCROLLS.
- Paper: Recursive Language Models, Alex L. Zhang et al. (2025). Develops a recursive execution framework that overcomes context window limits on complex aggregation and multi-document reasoning tasks.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). Builds upon long-sequence evaluation paradigms by benchmarking very long-term conversational memory across extended interaction histories.
