SCROLLS: Standardized CompaRison Over Long Language Sequences

Uri ShahamElad SegalMaor IvgiAvia EfratOri YoranAdi HavivAnkit GuptaWenhan XiongMor GevaJonathan Berant

article2022EMNLP195 citations

Introduces a standardized benchmark and live leaderboard across seven diverse, naturally long-text tasks to systematically evaluate how effectively language models synthesize and reason over extended sequences.

Listen

Most modern natural language processing benchmarks evaluate language models on short inputs such as individual sentences or brief paragraphs. In practical business, legal, and research settings, however, critical information is embedded within lengthy documents such as meeting transcripts, corporate contracts, scientific papers, and government reports. While many recent model architectures claim the ability to process long sequences, previous evaluation methods relied on artificial tasks, narrow academic summarization datasets, or metrics that only test local word prediction. Consequently, stakeholders have lacked a reliable way to assess whether language models can genuinely understand and reason across extended texts.

The article introduces SCROLLS (Standardized CompaRison Over Long Language Sequences), a benchmark designed to evaluate and compare language models on complex reasoning and synthesis across naturally long documents. To accomplish this, the authors handpicked and standardized seven datasets spanning multiple domains, including literature, television scripts, scientific articles, parliamentary meetings, and legal agreements. These datasets cover summarization, open and multiple-choice question answering, and natural language inference. All tasks were converted into a uniform text-to-text format, and the authors established baseline performance levels using standard and length-efficient transformer models across varying sequence lengths.

The analysis revealed that relevant information in the SCROLLS datasets is dispersed across hundreds or thousands of words, making information fusion essential. Baseline experiments demonstrated that providing models with longer context generally improves performance; for example, expanding the input capacity of the Longformer Encoder-Decoder model from 1,024 to 16,384 tokens increased its overall score by 2.1 points. However, a standard BART baseline processing only 1,024 tokens achieved an overall score of 29.01, performing within 0.15 points of Longformer's top score of 29.16 despite processing one-sixteenth of the text. When comparing both models at the same 1,024-token length, BART outperformed Longformer by almost two points. Crucially, all tested models fell drastically short of human performance; on question answering tasks, baseline models scored between 18% and 26%, whereas human agreement ranged between 58% and 93%.

These findings imply that increasing the maximum sequence length of a model does not automatically improve its semantic understanding. Organizations risk incurring high computational costs to process massive context windows without obtaining corresponding improvements in answer accuracy or summary quality. Efficient architectures initialized from standard models without long-sequence pretraining struggle to fully leverage extended context, particularly on smaller datasets where they cannot easily adapt to sparse attention patterns.

To address these shortcomings, developers and researchers should adopt the SCROLLS benchmark to guide the development of specialized pretraining strategies, novel architectures, and hybrid retrieval-augmented methods. The primary limitations of the benchmark include its restriction to English-language materials and the reliance on n-gram overlap metrics such as ROUGE for evaluating long summaries, which may undervalue valid paraphrasing. Nonetheless, the evidence strongly supports the conclusion that current natural language models remain far from solving long-text understanding, highlighting the need for continued architectural and methodological innovation.

arXiv: 2201.03533tau-nlp/scrolls
Cover for SCROLLS: Standardized CompaRison Over Long Language Sequences

Abstract

NLP benchmarks have largely focused on short texts, such as sentences and paragraphs, even though long texts comprise a considerable amount of natural language in the wild. We introduce SCROLLS, a suite of tasks that require reasoning over long texts. We examine existing long-text datasets, and handpick ones where the text is naturally long, while prioritizing tasks that involve synthesizing information across the input. SCROLLS contains summarization, question answering, and natural language inference tasks, covering multiple domains, including literature, science, business, and entertainment. Initial baselines, including Longformer Encoder-Decoder, indicate that there is ample room for improvement on SCROLLS. We make all datasets available in a unified text-to-text format and host a live leaderboard to facilitate research on model architecture and pretraining methods.

Table of Contents

  • 1 Introduction
  • 2 Background: Contemporary Evaluation of Long-Text Models
  • 3 The SCROLLS Benchmark
  • 3.1 Datasets
  • 3.2 Preprocessing
  • 3.3 Evaluation
  • 4 Quantitative Analysis
  • 5 Experiments
  • 5.1 Baselines
  • 5.2 Results
  • 6 Conclusion
  • 7 Limitations
  • Acknowledgements
  • References
  • A Dataset Format
  • B Dataset Splits
  • C Original Datasets Results
  • D Hyperparameters
  • E Qualitative Analysis

Knowls

  1. Knowl 1 — SCROLLS Benchmark Suite Composition and Task Specifications

    definition

    SCROLLS (Standardized CompaRison Over Long Language Sequences) is a benchmark for evaluating language models on natural long texts requiring synthesis and reasoning across long discourses. SCROLLS comprises 7 datasets covering summarization, query-based summarization, question answering (both open-ended and multiple-choice), and natural language inference across diverse domains:

    Dataset Task Domain Metric Avg Words Input Avg Words Output Examples (Total)
    GovReport Summarization Government ROUGE 7,886 492.5 19,402
    SummScreenFD Summarization TV ROUGE 5,598 99.6 4,348
    QMSum Query-Based Summarization Meetings ROUGE 9,497 69.7 1,810
    Qasper Question Answering Science F1 3,629 11.4 5,692
    NarrativeQA Question Answering Literature, Film F1 51,653 4.6 71,187
    QuALITY Multiple-Choice QA Literature, Misc Exact Match 4,193 10.3 6,737
    ContractNLI Natural Language Inference Legal Exact Match 1,706 1.4 10,319

    The split distributions across datasets are:

    • GovReport: 17,457 train, 972 validation, 973 test.
    • SummScreenFD: 3,673 train, 338 validation, 337 test (uses the ForeverDreaming transcript subset covering 88 TV shows).
    • QMSum: 1,257 train, 272 validation, 281 test (transcripts from ICSI academic meetings, AMI product meetings, Welsh and Canadian Parliaments).
    • Qasper: 2,567 train, 1,726 validation, 1,399 test (NLP papers from S2ORC with practitioner-annotated questions and answers).
    • NarrativeQA: 55,003 train, 5,878 validation, 10,306 test (books from Project Gutenberg and movie scripts).
    • QuALITY: 2,523 train, 2,086 validation, 2,128 test (Project Gutenberg/OANC stories with questions where 50%50\% are labeled as hard because speed annotators chose wrong answers when skimming).
    • ContractNLI: 7,191 train, 1,037 validation, 2,091 test (non-disclosure agreements from EDGAR/Google evaluated against 17 legal hypotheses for entailment, contradiction, or neutral).
  2. Knowl 2 — Baseline Performance and Context Length Effects on SCROLLS

    empirical result

    Evaluation of naive heuristics, BART-base, and Longformer Encoder-Decoder (LED-base) across multiple input length limits on the SCROLLS benchmark yields the following test performance:

    Model Input Len GovRep SumScr QMSum Qspr Nrtv QALT CNLI Avg
    ROUGE ROUGE ROUGE F1 F1 EM-T/H EM
    Naive - 45.3 / 17.9 / 20.8 19.6 / 1.8 / 11.0 14.2 / 2.0 / 9.3 3.4 1.5 25.2 / 26.1 66.0 19.35
    BART 256 41.9 / 14.2 / 20.3 24.5 / 3.8 / 15.3 29.9 / 8.3 / 20.4 23.3 14.0 26.0 / 25.8 69.8 26.35
    BART 512 45.6 / 16.9 / 21.8 26.3 / 5.1 / 16.2 29.5 / 8.2 / 20.1 24.7 14.5 26.8 / 27.4 71.6 27.58
    BART 1024 47.9 / 18.6 / 22.7 27.2 / 4.9 / 16.7 30.2 / 8.7 / 20.7 26.3 15.4 26.0 / 25.9 77.4 29.01
    LED 1024 40.9 / 16.1 / 23.1 22.7 / 3.6 / 15.1 24.6 / 6.5 / 19.0 24.4 15.2 26.6 / 27.2 73.4 27.06
    LED 4096 52.5 / 23.3 / 26.8 23.0 / 4.1 / 15.1 26.6 / 6.9 / 19.9 25.0 16.3 26.6 / 27.3 71.5 28.30
    LED 16384 56.2 / 26.6 / 28.8 24.2 / 4.5 / 15.4 25.1 / 6.7 / 18.8 26.6 18.5 25.8 / 25.4 71.5 29.16

    Note: GovRep, SumScr, and QMSum report ROUGE-1 / ROUGE-2 / ROUGE-L. For these datasets, the geometric mean of the three ROUGE scores is used to calculate the overall average (Avg). QALT reports EM on Total / Hard test sets; EM-Total is used for Avg.

    Key experimental observations:

    1. Context Expansion Benefit: Extending context increases average performance for BART (+2.66+2.66 points from 256 to 1,024 tokens) and LED (+2.10+2.10 points from 1,024 to 16,384 tokens).
    2. BART vs. LED at Equal Context: When both models are evaluated at 1,024 tokens, standard BART outperforms LED by 1.951.95 points (29.01 vs. 27.06), outperforming LED on 6 out of 7 tasks.
    3. BART 1k vs. LED 16k: Despite LED processing 16×16\times more tokens (16,384 vs. 1,024), BART-1024 achieves an aggregate score of 29.01, trailing LED-16384 (29.16) by only 0.15 points. LED (16k) significantly outperforms BART (1k) on GovReport (56.2 vs 47.9 ROUGE-1) and NarrativeQA (18.5 vs 15.4 F1), which are the two largest datasets by training volume.
  3. Knowl 3 — Evaluation Metrics and Benchmark Score Aggregation in SCROLLS

    equation

    The SCROLLS benchmark evaluates models using private test sets and computes task-specific metrics aggregated into a single SCROLLS score.

    1. Summarization Tasks (GovReport, SummScreenFD, QMSum): Evaluated using ROUGE after converting text to lowercase, replacing non-alphanumeric characters with whitespace, and applying whitespace tokenization. The overall score for each summarization dataset SSummS_{\text{Summ}} is the geometric mean of ROUGE-1\text{ROUGE-1}, ROUGE-2\text{ROUGE-2}, and ROUGE-L\text{ROUGE-L}: SSumm=ROUGE-1×ROUGE-2×ROUGE-L3S_{\text{Summ}} = \sqrt[3]{\text{ROUGE-1} \times \text{ROUGE-2} \times \text{ROUGE-L}}

    2. Question Answering Tasks (Qasper, NarrativeQA): Evaluated using token-level unigram F1F_1 score after lowercasing, stripping punctuation, and removing stopwords. For an input xx with a set of reference answers R(x)R(x) and model prediction y^\hat{y}, the score is: F1(x,y^)=max⁡r∈R(x)F1(y^,r)F_1(x, \hat{y}) = \max_{r \in R(x)} F_1(\hat{y}, r)

    3. Multiple-Choice QA and Classification Tasks (QuALITY, ContractNLI): Evaluated using Exact Match (EM\text{EM}), which normalizes prediction and target by lowercasing, stripping punctuation and stopwords, and collapsing whitespace, scoring 11 if normalized strings match and 00 otherwise. For QuALITY, the score used in the aggregate benchmark is total test set accuracy EM-T\text{EM-T} (while hard subset accuracy EM-H\text{EM-H} is tracked separately).

    4. Aggregate SCROLLS Score: The overall benchmark score SSCROLLSS_{\text{SCROLLS}} is the unweighted arithmetic mean of the individual dataset scores across all 7 tasks: SSCROLLS=17∑i=17SiS_{\text{SCROLLS}} = \frac{1}{7} \sum_{i=1}^7 S_i where SiS_i is the dataset-specific metric score (geometric mean ROUGE for GovReport, SummScreenFD, and QMSum; max reference F1F_1 for Qasper and NarrativeQA; and EM for QuALITY and ContractNLI).

  4. Knowl 4 — Unified Sequence-to-Sequence Format and Data Cleansing in SCROLLS

    model/method

    To evaluate sequence-to-sequence models across all tasks, SCROLLS formats every dataset into a uniform text-to-text representation and applies quality filtering:

    1. Unified Text-to-Text Formatting:

      • For query-driven tasks (QMSum, Qasper, NarrativeQA, QuALITY, ContractNLI), the query (or hypothesis) is prepended to the source document using two newline characters (\n\n) as a separator: Input=Query+"\n\n"+Document\text{Input} = \text{Query} + \text{"\n\n"} + \text{Document}
      • For multiple-choice QA (QuALITY), the four candidate options (labeled (A), (B), (C), (D)) are appended directly to the question query before the double newline separator.
      • For document-only summarization tasks (GovReport, SummScreenFD), the source document alone serves as the input.
      • For instances with multiple valid reference answers (in Qasper and NarrativeQA), each reference is treated as an independent instance during training and validation. During test evaluation, predictions are scored against each reference independently and aggregated via the maximum score per unique input.
    2. Data Cleansing:

      • GovReport: Discards documents where document word count is <2×< 2 \times summary word count or >1000×> 1000 \times summary word count, or where the summary appears verbatim in the document (64 examples removed).
      • Qasper: Discards papers with fewer than 8,192 characters (removes 176 questions over 63 papers).
      • NarrativeQA: Strips metadata, licenses, and HTML headers using start/end book markers.
  5. Knowl 5 — Information Spread Metric for Quantifying Long-Range Dispersion

    equation

    To quantify the extent to which generating a target output requires aggregating information across disparate locations of a long document, the information spread metric measures the spatial dispersion of target bigrams in the input text.

    Let yy be the target output string, represented as a set of unique bigrams By={b1,b2,…,bm}B_y = \{b_1, b_2, \dots, b_m\}. For each bigram bj∈Byb_j \in B_y, let pos(bj,x)∈{1,…,∣x∣}\text{pos}(b_j, x) \in \{1, \dots, |x|\} denote the word index of the first occurrence of bjb_j in the input document xx (if bjb_j appears in xx). Let K⊆ByK \subseteq B_y be the subset of output bigrams present in xx, with ∣K∣=k|K| = k, and let {p1,p2,…,pk}\{p_1, p_2, \dots, p_k\} denote their first-occurrence positions in xx.

    The spread σ(x,y)\sigma(x, y) of the example is defined as the standard deviation of these positions: σ(x,y)=1k∑j=1k(pj−pˉ)2,where pˉ=1k∑j=1kpj\sigma(x, y) = \sqrt{\frac{1}{k} \sum_{j=1}^k (p_j - \bar{p})^2}, \quad \text{where } \bar{p} = \frac{1}{k} \sum_{j=1}^k p_j

    For classification and QA datasets, the query/question is omitted from xx. For ContractNLI, the hypothesis is treated as yy and the non-disclosure agreement as xx.

  6. Knowl 6 — Empirical Distribution of Information Spread in SCROLLS vs Standard Datasets

    empirical result

    Measuring the standard deviation of target bigram positions across input documents reveals substantial differences in information spread between SCROLLS and standard NLP benchmarks:

    1. Summarization Datasets: In canonical short-context summarization (CNN/DM), reference bigrams cluster within a narrow distance (peaking around 10210^2 words). In arXiv summarization, bigrams spread across wider regions (peaking around 10310^3 words). In SCROLLS summarization datasets (GovReport, SummScreenFD, QMSum), target bigrams are dispersed over substantially wider spans, averaging 1.5×1.5\times to 2×2\times greater spread than arXiv.
    2. Question Answering and NLI Datasets: In SQuAD and Natural Questions (long Wikipedia context with short answers), answer bigrams typically cluster within a standard deviation of fewer than 5 words from each other. In contrast, SCROLLS QA datasets (Qasper, QuALITY) and NLI (ContractNLI) exhibit bigram spread values peaking in the range of hundreds to thousands of words (10210^2 to 10310^3 words). In NarrativeQA, while a subset of answers has localized bigrams, a significant proportion is dispersed across long spans (>103>10^3 words).

    This demonstrates that solving SCROLLS tasks requires fusing evidence across widely separated sections of text rather than extracting local token spans.

  7. Knowl 7 — Baseline-to-Human Performance Discrepancy on SCROLLS

    empirical result

    Comparing baseline model performance against human agreement benchmarks reveals a substantial performance deficit across multiple long-text tasks:

    • Qasper (Scientific QA): The best model baseline (LED-16384) achieves 26.6%26.6\% F1F_1, whereas human inter-annotator agreement on multi-annotated questions is 60.9%60.9\% F1F_1 (more than double the best model baseline).
    • QuALITY (Multiple-Choice Story QA): Baseline models achieve at best 26.8%26.8\% Exact Match (BART-512) and 26.6%26.6\% (LED-1024/4096), which is near the random/naive baseline floor (25.2%25.2\%). In contrast, human annotator agreement reaches 93.5%93.5\% EM.
    • NarrativeQA (Book/Script QA): The best baseline model (LED-16384) achieves 18.5%18.5\% F1F_1, whereas human inter-annotator agreement on the test set is 58.7%58.7\% F1F_1 (over 3×3\times higher).
    • Floor Heuristic Comparison: The naive heuristics achieve an average SCROLLS score of 19.3519.35. The model baselines improve on this lower bound by only 7 to 10 points (26.35 to 29.16).
  8. Knowl 8 — Training and Fine-Tuning Setup for SCROLLS Baselines

    experimental setup

    The benchmark establishes baseline evaluation using BART-base (a denoising sequence-to-sequence pretrained transformer) and Longformer Encoder-Decoder base (LED-base, an efficient transformer with sliding-window local attention of 1,024 tokens and global attention on the first token):

    • Optimization: Fine-tuned on each dataset separately using AdamW with parameters β=(0.9,0.98)\beta = (0.9, 0.98), ϵ=10−6\epsilon = 10^{-6}, 16-bit mixed precision (fp16), and gradient checkpointing.
    • Batch Size and Compute: Effective batch size of 131,072131,072 (2172^{17}) tokens achieved via 16,38416,384 tokens per GPU across 8 NVIDIA V100 (32GB) GPUs using data parallelism and gradient accumulation.
    • Training Epochs: 10 epochs for summarization datasets (GovReport, SummScreenFD, QMSum); 20 epochs for Qasper, QuALITY, and ContractNLI; 2 epochs for NarrativeQA.
    • Learning Rate Schedule: Maximum learning rate tuned per dataset validation set among {1e-5,2e-5,5e-5,1e-4,2e-4,5e-4}\{1\text{e-}5, 2\text{e-}5, 5\text{e-}5, 1\text{e-}4, 2\text{e-}4, 5\text{e-}4\}, with a linear warmup for the first 10%10\% of total training steps followed by a linear decay to 0 over the remaining 90%90\%.
    • Regularization and Decoding: Dropout of 0.1 applied throughout all layers. Greedy decoding is used during inference.
  9. Knowl 9 — Limitations of the SCROLLS Benchmark

    limitation

    The SCROLLS benchmark exhibits two primary limitations:

    1. N-gram Overlap Metric Limitations on Summarization: Summarization evaluation relies on ROUGE (nn-gram overlap and longest common subsequence). ROUGE can penalize semantically accurate summaries that paraphrase reference content using different phrasing, which limits the precision of automated evaluation for long-form abstractive outputs until reliable model-based generation metrics are established.
    2. Monolingual English Scope: All datasets and tasks in SCROLLS are currently restricted to the English language, leaving cross-lingual and non-English long-context reasoning for future benchmark extensions.

Coverage note — Qualitative walkthrough examples of individual dataset transcripts and text excerpts from Appendix E and Figures 5-9 were omitted as they serve purely as qualitative illustrations of the datasets and spread metric.

References

  1. 1.Loïc Barrault, Magdalena Biesialska, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljubešić, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Santanu Pal, Matt Post, and Marcos Zampieri. 2020. Findings of the 2020 conference on machine translation (WMT20). In Proceedings of the Fifth Conference on Machine Translation, pages 1–55, Online. Association for Computational Linguistics.
  2. 2.Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. Findings of the 2019 conference on machine translation (WMT19). In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 1–61, Florence, Italy. Association for Computational Linguistics.
  3. 3.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre Wellner. 2005. The ami meeting corpus: A pre-announcement. In International workshop on machine learning for multi-modal interaction, pages 28–39, Berlin, Heidelberg. Springer.
  6. 6.Mingda Chen, Zewei Chu, Sam Wiseman, and Kevin Gimpel. 2021. Summscreen: A dataset for abstractive screenplay summarization.
  7. 7.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. 2020. Rethinking attention with performers.
  8. 8.Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018. A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 615–621, New Orleans, Louisiana. Association for Computational Linguistics.
  9. 9.Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. 2021. A dataset of information-seeking questions and answers anchored in research papers. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4599–4610, Online. Association for Computational Linguistics.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Charles J. Fillmore, Nancy Ide, Dan Jurafsky, and Catherine Macleod. 1998. An american national corpus: a proposal. In LREC.
  12. 12.Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. 2019. MRQA 2019 shared task: Evaluating generalization in reading comprehension. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pages 1–13, Hong Kong, China. Association for Computational Linguistics.
  13. 13.Quentin Fournier, Gaétan Marceau Caron, and Daniel Aloise. 2021. A practical survey on faster and lighter transformers.
  14. 14.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc.
  15. 15.Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang. 2021. Efficient attentions for long document summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1419–1436, Online. Association for Computational Linguistics.
  16. 16.Nancy Ide and Keith Suderman. 2004. The American national corpus first release. In Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC’04), Lisbon, Portugal. European Language Resources Association (ELRA).
  17. 17.A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke, and C. Wooters. 2003. The icsi meeting corpus. In 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP ’03)., volume 1, pages I–I.
  18. 18.Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018. Sharp nearby, fuzzy far away: How neural language models use context. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 284–294, Melbourne, Australia. Association for Computational Linguistics.
  19. 19.Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018. The NarrativeQA reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328.
  20. 20.Yuta Koreeda and Christopher D. Manning. 2021. Contractnli: A dataset for document-level natural language inference for contracts.
  21. 21.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  22. 22.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pretraining for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  23. 23.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  24. 24.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  25. 25.Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020. S2ORC: The semantic scholar open research corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4969–4983, Online. Association for Computational Linguistics.
  26. 26.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations.
  27. 27.Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, and Samuel R. Bowman. 2021. QuALITY: Question answering with long input texts, yes! arXiv preprint arXiv:2112.08608.
  28. 28.Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong. 2021. Random feature attention. In International Conference on Learning Representations.
  29. 29.Ofir Press, Noah A. Smith, and Mike Lewis. 2021a. Shortformer: Better language modeling using shorter inputs. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5493–5505, Online. Association for Computational Linguistics.
  30. 30.Ofir Press, Noah A. Smith, and Mike Lewis. 2021b. Train short, test long: Attention with linear biases enables input length extrapolation.
  31. 31.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789, Melbourne, Australia. Association for Computational Linguistics.
  32. 32.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  33. 33.Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021. Efficient content-based sparse attention with routing transformers. Transactions of the Association for Computational Linguistics, 9:53–68.
  34. 34.Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, and Mohit Iyyer. 2021. Do long-range language models actually use long-range context? In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 807–822, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  35. 35.Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2020a. Synthesizer: Rethinking self-attention in transformer models.
  36. 36.Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. 2021. Long range arena : A benchmark for efficient transformers. In International Conference on Learning Representations.
  37. 37.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020b. Efficient transformers: A survey.
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  39. 39.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  40. 40.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
  41. 41.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  42. 42.Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev. 2021. QMSum: A new benchmark for query-based multi-domain meeting summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5905–5921, Online. Association for Computational Linguistics.

Citation

MLA
Shaham, U., et al. “SCROLLS: Standardized CompaRison Over Long Language Sequences”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 12007–21, https://doi.org/10.18653/v1/2022.emnlp-main.823.
APA
Shaham, U., Segal, E., Ivgi, M., Efrat, A., Yoran, O., Haviv, A., Gupta, A., Xiong, W., Geva, M., Berant, J., & Levy, O. (2022). SCROLLS: Standardized CompaRison Over Long Language Sequences. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 12007–12021. https://doi.org/10.18653/v1/2022.emnlp-main.823
Chicago
Shaham, U., E. Segal, M. Ivgi, et al. 2022. “SCROLLS: Standardized CompaRison Over Long Language Sequences”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 12007–21. https://doi.org/10.18653/v1/2022.emnlp-main.823.
Harvard
Shaham, U. et al. (2022) “SCROLLS: Standardized CompaRison Over Long Language Sequences”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 12007–12021. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.823.
Vancouver
1. Shaham U, Segal E, Ivgi M, et al (2022) SCROLLS: Standardized CompaRison Over Long Language Sequences. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 12007–12021

BibTeX

@inproceedings{shaham-etal-2022-scrolls,
    title = "{SCROLLS}: Standardized {C}ompa{R}ison Over Long Language Sequences",
    author = "Shaham, Uri  and
      Segal, Elad  and
      Ivgi, Maor  and
      Efrat, Avia  and
      Yoran, Ori  and
      Haviv, Adi  and
      Gupta, Ankit  and
      Xiong, Wenhan  and
      Geva, Mor  and
      Berant, Jonathan  and
      Levy, Omer",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.823/",
    doi = "10.18653/v1/2022.emnlp-main.823",
    pages = "12007--12021"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/