LooGLE: Can Long-Context Language Models Understand Long Contexts?

Jiaqi LiMengmeng WangZilong ZhengMuhan Zhang

article2024ACL284 citations

Introduces LooGLE, a benchmark built from recent, long-form documents to evaluate whether language models can genuinely reason over extended dependencies across entire texts rather than relying on short-range retrieval.

Listen

Modern large language models are increasingly engineered to accept vast volumes of text at once. However, existing evaluation benchmarks rely primarily on short texts, outdated public documents that risk data leakage, or simple retrieval tasks that only test isolated sentences. The article introduces a generic evaluation benchmark called LooGLE to rigorously assess whether modern language models truly comprehend long texts rather than merely accommodating them in memory.

The benchmark consists of 776 diverse documents published after 2022—including academic papers, Wikipedia articles, and film scripts—averaging over 19,000 words per document. It incorporates more than 6,400 evaluation instances across seven tasks. Crucially, the authors organized over 1,200 human-hours of cross-validated manual effort to create 1,101 high-quality questions specifically designed to test long-range dependencies, such as multi-source retrieval, timeline reordering, mathematical calculation across distributed facts, and multi-step reasoning.

The findings demonstrate a severe performance gap between superficial processing and genuine comprehension. While top commercial models perform well on short-dependency tasks and summarization (often achieving 70% to 85% accuracy), all evaluated models struggle significantly on long-dependency tasks. Even the leading commercial model, GPT-4 with a 32,000-token capacity, achieves an accuracy of only about 40% to 54% on complex long-range questions. Open-source models exhibit a severe capability drop, with several scoring below 15% accuracy on long-dependency reasoning. Furthermore, standard retrieval-based augmentation methods failed to improve long-range question answering, and prompt-engineering strategies like chain-of-thought yielded mixed or marginal benefits.

These results carry critical implications for enterprise deployment and risk management. Relying on large context windows under the assumption that models accurately analyze entire lengthy reports introduces major operational risks, including high rates of hallucination and incomplete evidence synthesis. Expanding context window size alone does not resolve the inability to model complex temporal relationships, calculations, or interdependencies across long documents.

Organizations and developers should avoid treating expanded context windows as a substitute for true analytical capability. Development efforts must shift toward improving core reasoning, temporal awareness, and multi-hop fact aggregation. Benchmarks must also be continually refreshed with recent texts to prevent memorization artifacts. The study's conclusions are robustly grounded in both automated metrics and aligned human evaluations, though readers should note that the current benchmark is restricted to English documents and constrained by existing baseline prompting frameworks.

arXiv: 2311.04939
Cover for LooGLE: Can Long-Context Language Models Understand Long Contexts?

Abstract

Large language models (LLMs) are typically limited to processing texts within context-window size, which has spurred significant research efforts into enhancing LLMs’ long-context understanding as well as developing high-quality benchmarks to evaluate the ability. However, prior datasets suffer from shortcomings like short length compared to the context window of modern LLMs; outdated documents that might have data leakage problems; and an emphasis on short dependency tasks only. In this paper, we present LooGLE, a Long Context Generic Language Evaluation benchmark. It features documents post-2022, with over 24,000 tokens per document and 6,000 newly generated questions spanning varying dependency ranges in diverse domains. Human annotators meticulously crafted over 1,100 high-quality question-answer (QA) pairs with thorough cross-validation for a most precise assessment of LLMs’ long dependency capabilities. We conduct a comprehensive evaluation of representative LLMs on LooGLE. The results indicate that most LLMs have shockingly bad long context ability and fail to capture long dependencies in the context, even when their context window size is enough to fit the entire document. Our results shed light on enhancing the “true long-context understanding” ability of LLMs instead of merely enlarging their context window.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The LooGLE Benchmark
  • 3.1 Dataset selection and construction
  • 3.2 Long dependency tasks
  • 3.2.1 Long dependency QA construction
  • 4 Benchmark LLMs on LooGLE
  • 4.1 Evaluation methods and metrics
  • 4.2 Results
  • 4.2.1 Analysis on short dependency tasks
  • 4.2.2 Analysis on long dependency tasks
  • 4.2.3 Long context capabilities deep dive
  • 5 Discussion
  • Acknowledgement
  • References
  • A More details of our dataset and experiment settings
  • B Short dependency task definition and generation
  • C Task definition
  • D Timeline reorder evaluation metrics
  • E Further analysis and results
  • F Prompts
  • F.1 Short dependency QA pair generation
  • F.2 Short and long dependency QA
  • F.3 scripts segment summarization for cloze formulation
  • F.4 Cloze
  • F.5 Summarization
  • F.6 Timeline reorder
  • F.7 QA task evaluation by GPT4
  • F.8 Summarization task evaluation by GPT4
  • F.9 Few-Shot CoT for long QA
  • F.10 Zero-Shot CoT for long QA
  • G Examples for long context understanding tasks
  • G.1 Short dependency QA
  • G.2 Cloze
  • G.3 Summarization
  • G.4 Multi-source retrieval
  • G.5 Timeline reorder
  • G.6 Calculation
  • G.7 Comprehension and reasoning
  • H Examples of models' generated outputs
  • H.1 GPT4-32k
  • H.2 GPT4-8k
  • H.3 GPT3.5-turbo-16k
  • H.4 LlamaIndex
  • H.5 ChatGLM2-6B-32k
  • H.6 RWKV-4-14B-raven
  • H.7 LongLLaMa-3B-Instruct
  • H.8 LLaMA2-7B-32K-Instruct

Knowls

  1. Knowl 1 — LooGLE Benchmark Dataset Composition and Statistics

    data/table

    The LooGLE benchmark is designed to evaluate long-context language comprehension on documents published after 2022 to avoid pre-training data leakage. All included documents exceed a minimum length threshold of 10,000 words, with an average length of 19,367 words (over 24,000 tokens). The benchmark encompasses 776 documents across three primary domains: arXiv scientific papers, Wikipedia articles, and movie/TV scripts. These sources provide a total of 6,448 test instances across short-dependency tasks (short question answering and entity cloze) and long-dependency tasks (paper summarization and manually annotated long QA).

    Source # Docs Avg Words Max Words Min Words Avg Tokens Tasks (# Questions)
    arXiv 516 16,988 197,977 10,204 20,887 Summarization (516)
    Wikipedia 105 17,604 46,250 11,285 21,017 Short QA (1,951), Long QA (459)
    Movie TV scripts 155 28,483 62,752 11,089 36,412 Cloze (2,880), Long QA (642)
    Total / Overall 776 19,367 197,977 10,204 24,000+ 6,448 (1,101 Long QA)

    The table details the distribution of source documents, lengths, and associated task instance counts in LooGLE.

  2. Knowl 2 — Long-Dependency Task Taxonomy and Multi-Stage Annotation Protocol in LooGLE

    model/method

    LooGLE categorizes long-dependency evaluation into five core task types requiring inter-dependency modeling across wide textual spans (recommended evidence span >5,000> 5,000 words):

    1. Multi-Source Retrieval (MR) (34.51% of long QA): Requires locating and extracting multiple distinct pieces of evidence distributed across the text and aggregating them into a single response.
    2. Timeline Reordering (TR) (19.53% of long QA): Requires sorting a permuted sequence of events described in the text into chronological order.
    3. Calculation (9.08% of long QA): Requires extracting numerical information across disparate sections of the text and applying arithmetic reasoning.
    4. Comprehension and Reasoning (CR) (36.88% of long QA): Requires multi-hop reasoning, causal inference, and evaluating attributes where answers are not verbatim phrases in the text.
    5. Paper Summarization: Uses entire arXiv research papers as input, with original author abstracts serving as ground-truth references.

    To construct high-quality long QA instances, a three-step cross-validation annotation protocol was executed on 140 documents from Wikipedia and script collections, totaling over 1,260 human-hours:

    • Step 1 (Question & Answer): A student annotator reads the entire document, generates 5–10 deterministic QA pairs spanning varied task types (no more than 4 of the same type per document), and records the exact evidentiary passages spanning the text.
    • Step 2 (Independent Answer & Review): A second annotator, blind to the questioner's answers and evidence, reads the full document, answers the questions independently, and assesses question validity and clarity.
    • Step 3 (Revision & Unification): The initial questioner receives feedback and the independent answer, resolves discrepancies, and unifies the answers into a final ground-truth reference.

    The final set contains 1,101 verified long-dependency QA pairs with an inter-annotator agreement rate of 81.88%.

  3. Knowl 3 — Permutation Sequence Deviation Metrics for Timeline Reordering Evaluation

    equation

    To evaluate model-predicted event sequences against ground-truth chronological sequences in the Timeline Reordering task, candidate answers are parsed via regular expressions to extract ordered index sequences. For two numeric permutation sequences AA and BB of identical length nn, where i[A]i[A] and i[B]i[B] denote the element at index ii in sequences AA and BB respectively, four deviation metrics are defined:

    1. Location Square Deviation (LSDLSD): LSD(A,B)=1n∑i=0n−1(i[A]−i[B])2LSD(A, B) = \frac{1}{n} \sum_{i=0}^{n-1} (i[A] - i[B])^2

    2. Location Mean Deviation (LMDLMD): LMD(A,B)=1n∑i=0n−1∣i[A]−i[B]∣LMD(A, B) = \frac{1}{n} \sum_{i=0}^{n-1} |i[A] - i[B]|

    3. Swap Deviation (SDSD): SD(A,B)=min⁡S∈A→B∑s∈S1SD(A, B) = \min_{S \in A \to B} \sum_{s \in S} 1

    4. Swap Distance Deviation (SDDSDD): SDD(A,B)=min⁡S∈A→B∑s=A(i,j)∈S∣i−j∣SDD(A, B) = \min_{S \in A \to B} \sum_{s=A(i, j) \in S} |i - j|

    where S=A→BS = A \to B denotes a valid sequence of element swaps transforming sequence AA into sequence BB, and s=A(i,j)s = A(i, j) represents an individual swap action between elements at indices ii and jj. Smaller values indicate closer alignment with the ground-truth sequence. Non-standard outputs (empty responses, mismatched lengths, or unparseable formats) are assigned the maximum theoretical deviation.

  4. Knowl 4 — LLM Performance Divergence Between Short-Dependency and Long-Dependency Tasks

    empirical result

    Evaluation across proprietary models (GPT-4-32k, GPT-4-8k, GPT-3.5-turbo-16k, Claude 3 Opus) and open-source long-context LLMs (ChatGLM2-6B-32k, LongLLaMA-3B-Instruct, RWKV-4-14B-raven, LLaMA2-7B-32K-Instruct) reveals a stark performance gap between short-dependency and long-dependency understanding:

    • Short-Dependency Tasks: Models achieve strong performance on localized Short QA (GPT-4-32k reaches 71.52% GPT-4 judgment score, GPT-3.5-turbo-16k reaches 66.82%) and script Cloze (GPT-4-32k achieves 70.50% Exact Match and 80.81% Partial Match).
    • Summarization: Commercial models effectively summarize long research papers, exceeding 80% GPT-4 evaluation accuracy (GPT-3.5-turbo-16k: 86.84%, GPT-4-8k: 85.42%, GPT-4-32k: 82.84%).
    • Long-Dependency QA: All models suffer severe degradation on long QA tasks requiring multi-source interdependency comprehension. Even the highest-performing model, GPT-4-32k, achieves only a 54.09% GPT-4 score, while open-source models score between 2.85% and 21.64%.
    Model Context Bleu1 Bleu4 Rouge1 RougeL Meteor BERTScore GPT-4 Score
    GPT-4-32k 32k 8.55 1.40 25.59 24.04 11.13 80.16 54.09
    GPT-4-8k 8k 8.94 1.01 23.45 21.69 10.18 85.36 42.12
    GPT-3.5-turbo-16k 16k 6.92 1.81 25.02 23.63 10.40 83.79 45.04
    Claude 3 Opus 200k 3.28 0.43 37.95 36.56 9.44 79.58 20.71
    LongLLaMA-3B-Inst 256k 5.64 0.49 17.30 16.29 6.53 84.26 21.64
    RWKV-4-14B-raven 8k 3.88 0.22 20.39 19.20 6.41 81.46 14.32
    ChatGLM2-6B-32k 32k 5.55 0.11 9.41 8.69 4.39 85.78 11.50
    LLaMA2-7B-32K-Inst 32k 0.08 0.00 4.07 4.07 1.06 66.54 2.85

    The table demonstrates that scaling context window length alone (e.g., Claude 3 Opus at 200k context or LongLLaMA at 256k context) does not inherently resolve complex long-range inter-dependency comprehension.

  5. Knowl 5 — Evaluation of LLMs Across Subcategories of Long-Dependency Question Answering

    empirical result

    Disaggregated evaluation across specific long-dependency QA categories highlights pronounced variance in LLM capabilities depending on the underlying reasoning requirement. Evaluating with GPT-4 judgment shows that LLMs perform highest on Comprehension & Reasoning and Multi-Source Retrieval, but consistently fail on Timeline Reordering and Calculation tasks.

    Model Multi-Source Retrieval (%) Timeline Reorder (%) Calculation (%) Comprehension Reasoning (%)
    GPT-4-32k 33.26 26.43 22.30 44.20
    GPT-4-8k 26.59 20.61 16.31 34.42
    GPT-3.5-turbo-16k 24.05 20.88 13.49 32.10
    LlamaIndex 19.38 17.23 11.43 29.53
    ChatGLM2-6B-32k 11.38 10.77 8.45 10.95
    LongLLaMA-3B-Instruct 15.73 8.87 8.87 21.29
    RWKV-4-14B-raven 5.73 4.76 2.08 6.52
    LLaMA2-7B-32K-Instruct 2.23 1.36 1.39 2.67

    The results demonstrate that tracking chronological timelines across tens of thousands of words and identifying dispersed numerical data points for mathematical calculation constitute the most challenging failure modes for long-context language models.

  6. Knowl 6 — Semi-Automated Generation Pipeline for Short-Dependency QA and Cloze Tasks

    algorithm

    Short-dependency QA and cloze tasks in LooGLE are constructed using a semi-automated pipeline that splits documents into localized segments, prompts an LLM for initial question and summary generation, applies entity recognition, and refines items through manual quality inspection.

    Input: Document text DD, target maximum summary length L=500L = 500 words
    Output: Short QA dataset QshortQ_{short}, Cloze dataset QclozeQ_{cloze}
    Initialize Qshort←∅Q_{short} \leftarrow \emptyset, Qcloze←∅Q_{cloze} \leftarrow \emptyset
    Segment document DD into localized chunks S={S1,S2,…,Sm}S = \{S_1, S_2, \dots, S_m\}
    for each segment Si∈SS_i \in S do
        Prompt GPT-3.5-turbo-16k with SiS_i to output QA pairs in JSON format: {(Si,j,qi,j,ai,j)}\{(S_{i,j}, q_{i,j}, a_{i,j})\}
        for each candidate pair (qi,j,ai,j)(q_{i,j}, a_{i,j}) do
            Manually verify that ai,ja_{i,j} is supported by localized segment Si,jS_{i,j}
            Filter non-essential text and eliminate redundant descriptions
            Qshort←Qshort∪{(qi,j,ai,j,Si,j)}Q_{short} \leftarrow Q_{short} \cup \{(q_{i,j}, a_{i,j}, S_{i,j})\}
        end for
    end for
    for each script segment Si∈SS_i \in S do
        Prompt GPT-3.5-turbo-16k to generate objective factual summary YiY_i (word count ≤L\le L)
        Apply BERT-large NER model to extract named entities Ei={e∈Yi∣type(e)∈{Person,Location,Organization}}E_i = \{e \in Y_i \mid \text{type}(e) \in \{\text{Person}, \text{Location}, \text{Organization}\}\}
        Randomly select k≤5k \le 5 distinct entities {e1,…,ek}⊆Ei\{e_1, \dots, e_k\} \subseteq E_i
        Construct cloze question CiC_i by replacing each occurrence of ere_r in YiY_i with placeholder ⟨mask-r⟩\langle\text{mask-}r\rangle
        Qcloze←Qcloze∪{(Ci,{e1,…,ek},Si)}Q_{cloze} \leftarrow Q_{cloze} \cup \{(C_i, \{e_1, \dots, e_k\}, S_i)\}
    end for
    return Qshort,QclozeQ_{short}, Q_{cloze}

    The algorithm formalizes the generation of 1,951 Short QA pairs and 2,880 Cloze instances from segmented Wikipedia and script collections.

  7. Knowl 7 — Inefficacy of Retrieval-Augmented Generation on Long-Dependency Complex Tasks

    empirical result

    Augmenting language models with retrieval mechanisms via LlamaIndex (employing embeddings such as text-embedding-ada-002 or all-mpnet-base-v2 to fetch relevant chunks) provides competitive results on short QA tasks but significantly degrades performance on long-dependency tasks relative to full-document ingestion:

    • On Short QA, LlamaIndex achieves a GPT-4 evaluation score of 59.61% and a BLEU-1 of 33.37%.
    • On Long QA, pairing GPT-4-32k with LlamaIndex retrieval causes performance to drop from 54.09% (full context) to 28.25% (GPT-4 evaluation score). Similarly, GPT-4-8k drops from 42.12% to 26.34% when constrained to retrieved segments.
    • For GPT-3.5-turbo-16k, long QA accuracy drops from 45.04% to 33.24% with LlamaIndex.
    Model Setup Context Bleu1 Rouge1 RougeL Meteor BERTScore GPT-4 Score
    GPT-4-32k (Full Context) 32k 8.55 25.59 24.04 11.13 80.16 54.09
    GPT-4-32k + LlamaIndex 32k 6.08 10.27 9.52 8.54 85.27 28.25
    GPT-4-8k (Full Context) 8k 8.94 23.45 21.69 10.18 85.36 42.12
    GPT-4-8k + LlamaIndex 8k 6.62 11.95 10.99 9.02 85.51 26.34
    GPT-3.5-turbo (Full Context) 16k 6.92 25.02 23.63 10.40 83.79 45.04
    GPT-3.5-turbo + LlamaIndex 16k 6.50 10.93 9.86 8.65 85.63 33.24

    These findings show that standard embedding-based semantic retrieval cannot effectively synthesize information dispersed across multiple distant sections when answering complex multi-hop or global storyline queries.

  8. Knowl 8 — Impact of Chain-of-Thought Prompting Strategies on Long-Dependency Question Answering

    empirical result

    Human evaluation of Chain-of-Thought (CoT) prompting on representative long-context models across long QA task subcategories demonstrates distinct trade-offs between zero-shot and few-shot paradigms:

    • Zero-Shot CoT (prompting with "Let's think step by step"): Yields massive accuracy gains on structured reasoning tasks, increasing Timeline Reordering accuracy from 17.11% to 38.16% (+21.05% absolute) and Calculation accuracy from 16.13% to 27.42% (+11.29% absolute). It produces minimal variation on Comprehension & Reasoning (43.21% vs. 43.83%) and Multi-Source Retrieval (27.34% vs. 26.62%).
    • Few-Shot CoT (providing exemplar demonstrations with detailed rationales): Improves Comprehension & Reasoning to 51.23% (up from 43.21% without CoT) and Multi-Source Retrieval to 28.78%. However, it causes performance in Timeline Reordering (26.32%) and Calculation (20.97%) to degrade substantially compared to zero-shot CoT.

    The degradation under few-shot CoT for temporal and mathematical tasks occurs because specific reasoning rationales and evidentiary sequences in demonstration prompts fail to generalize across distinct document structures, providing misleading search guidance to the model.

  9. Knowl 9 — Impact of Input Length Scaling and Context Window Extension on Task Performance

    empirical result

    Empirical analysis of varying input context lengths and fine-tuning strategies reveals key limitations of naive context window extension:

    1. Input Length Truncation Effects on GPT-4: On paper summarization, expanding the context window of GPT-4-32k from 8k to 32k tokens has negligible effect on output quality (GPT-4 score moves from 82.75% at 8k to 82.84% at 32k) because introductory and concluding sections contain most key factual summaries. In contrast, on Long QA, expanding input length from 8k to 32k tokens steadily improves GPT-4-32k performance (GPT-4 score increases from 38.34% at 8k, to 47.55% at 16k, 50.61% at 24k, and 54.65% at 32k) by reducing information loss from head-tail truncation.
    2. Performance of Position Interpolation vs. Base Models: Fine-tuning models to extend context windows using position interpolation can impair model precision on downstream tasks. Specifically, the base LLaMA2-7B model with a 4k context window achieves higher scores than its extended-context counterpart LLaMA2-7B-32K-Instruct across both short QA (7.06% vs. 2.71% GPT-4 score) and long QA (7.95% vs. 2.85% GPT-4 score; 80.48 vs. 66.54 BERTScore).

    This indicates that current context extension methods can introduce significant noise and lower retrieval precision, compromising baseline instruction-following capabilities.

  10. Knowl 10 — Failure Modes and Error Distribution in Long-Context Large Language Models

    empirical result

    Manual qualitative analysis of failed test cases in long-dependency QA identifies distinct error distributions across task types:

    Task Hallucination (%) Redundant Retr. (%) Insufficient Retr. (%) Irrelevant (%) Refusal (%) Wrong Reasoning (%)
    Calculation 31.11 24.44 15.56 0.00 20.00 0.00
    Multi-Source Retrieval 14.71 31.37 28.43 13.73 13.73 0.00
    Compr. Reasoning 14.29 10.99 21.98 18.68 16.48 10.99

    Key failure modes include:

    • Hallucination: High in Calculation (31.11%), where models generate ungrounded numerical figures.
    • Redundant Retrieval: The primary error in Multi-Source Retrieval (31.37%), where models retrieve irrelevant context alongside valid evidence.
    • Insufficient Retrieval: Occurs in 28.43% of MR and 21.98% of CR tasks, where models fail to locate all required evidence dispersed throughout the text.
    • Refusal / No Relevant Context: Models fail to recall memory and explicitly decline to answer (20.00% in Calculation, 16.48% in CR).
    • Structural Output Failures: Open-source models exhibit extreme format non-adherence on structured tasks like Timeline Reordering, where LongLLaMA-3B produced 100% non-standard outputs, ChatGLM2-6B produced 99.07%, and LLaMA2-7B-32K produced 98.60% (compared to 52.80% for GPT-4-32k), frequently entering degenerate repetition loops or spitting out unrelated source code.

Coverage note — Raw qualitative output transcripts (Appendix H) and exact prompt templates (Appendix F) were omitted from standalone knowls as their methodologies and empirical findings are fully synthesized within the task construction, algorithm, and experimental evaluation knowls.

References

  1. 1.Chen An, Shansan Gong, Ming Zhong, Mukai Li, Jun Zhang, Lingpeng Kong, and Xipeng Qiu. 2023. L-eval: Instituting standardized evaluation for long context language models. ArXiv, abs/2307.11088.
  2. 2.Arian Askari, Suzan Verberne, Amin Abolghasemi, Wessel Kraaij, and Gabriella Pasi. 2024. Retrieval for extremely long queries and documents with rprs: a highly efficient and effective transformer-based reranker. ACM Transactions on Information Systems, 42(5):1–32.
  3. 3.Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. 2023. Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508.
  4. 4.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  5. 5.Arkadii Bessonov, Alexey Staroverov, Huzhenyu Zhang, Alexey K Kovalev, Dmitry Yudin, and Aleksandr I Panov. 2023. Recurrent memory decision transformer. arXiv preprint arXiv:2306.09459.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS), 33:1877–1901.
  7. 7.Aydar Bulatov, Yuri Kuratov, Yermek Kapushev, and Mikhail S Burtsev. 2023. Scaling transformer to 1m tokens and beyond with rmt. arXiv preprint arXiv:2304.11062.
  8. 8.Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev. 2022. Recurrent memory transformer. Advances in Neural Information Processing Systems, 35:11079–11091.
  9. 9.Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023a. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595.
  10. 10.Xuanting Chen, Junjie Ye, Can Zu, Nuo Xu, Rui Zheng, Minlong Peng, Jie Zhou, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023b. How robust is gpt-3.5 to predecessors? a comprehensive study on language understanding tasks. arXiv preprint arXiv:2303.00293.
  11. 11.Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 35:16344–16359.
  12. 12.J Ding, S Ma, L Dong, X Zhang, S Huang, W Wang, N Zheng, and F Wei. 2023. Longnet: Scaling transformers to 1,000,000,000 tokens 2023. arXiv preprint arXiv:2307.02486.
  13. 13.Zican Dong, Tianyi Tang, Lunyi Li, and Wayne Xin Zhao. 2023. A survey on long text modeling with transformers. arXiv preprint arXiv:2302.14502.
  14. 14.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335.
  15. 15.Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang. 2021. Efficient attentions for long document summarization. arXiv preprint arXiv:2104.02112.
  16. 16.Yunpeng Huang, Jingwei Xu, Zixu Jiang, Junyu Lai, Zenan Li, Yuan Yao, Taolue Chen, Lijuan Yang, Zhou Xin, and Xiaoxing Ma. 2023. Advancing transformer architecture in long-context large language models: A comprehensive survey. arXiv preprint arXiv:2311.12351.
  17. 17.Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT, volume 1, page 2.
  18. 18.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213.
  19. 19.Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi. 2023. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267.
  20. 20.Dacheng Li, Rulin Shao, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang. 2023a. How long can context length of open-source llms truly promise? In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following.
  21. 21.Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie, Santiago Ontanon, Manzil Zaheer, Sumit Sanghai, Yiming Yang, Sanjiv Kumar, and Srinadh Bhojanapalli. 2023b. Functional interpolation for relative positions improves long context transformers. arXiv preprint arXiv:2310.04418.
  22. 22.Yucheng Li. 2023. Unlocking context constraints of llms: Enhancing context efficiency of llms with self-information-based content filtering. arXiv preprint arXiv:2304.12102.
  23. 23.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  24. 24.Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023a. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173.
  25. 25.Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2023b. Calibrating llm-based evaluator. arXiv preprint arXiv:2309.13308.
  26. 26.Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, et al. 2023c. Llm360: Towards fully transparent open-source llms. arXiv preprint arXiv:2312.06550.
  27. 27.Clara Meister, Stefan Lazov, Isabelle Augenstein, and Ryan Cotterell. 2021. Is sparse attention more interpretable? arXiv preprint arXiv:2106.01087.
  28. 28.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744.
  29. 29.Arka Pal, Deep Karkhanis, Manley Roberts, Samuel Dooley, Arvind Sundararajan, and Siddartha Naidu. 2023. Giraffe: Adventures in expanding context lengths in llms. arXiv preprint arXiv:2308.10882.
  30. 30.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  31. 31.Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, et al. 2023. Rwkv: Reinventing rnns for the transformer era. arXiv preprint arXiv:2305.13048.
  32. 32.Uri Shaham, Maor Ivgi, Avia Efrat, Jonathan Berant, and Omer Levy. 2023. Zeroscrolls: A zero-shot benchmark for long text understanding. arXiv preprint arXiv:2305.14196.
  33. 33.Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, et al. 2022. Scrolls: Standardized comparison over long language sequences. arXiv preprint arXiv:2201.03533.
  34. 34.Eva Sharma, Chen Li, and Lu Wang. 2019. Bigpatent: A large-scale dataset for abstractive and coherent summarization. arXiv preprint arXiv:1906.03741.
  35. 35.Roshan Sharma, Suyoun Kim, Daniel Lazar, Trang Le, Akshat Shrivastava, Kwanghoon Ahn, Piyush Kansal, Leda Sari, Ozlem Kalinli, and Michael Seltzer. 2023. Augmenting text for spoken language understanding with large language models. arXiv preprint arXiv:2309.09390.
  36. 36.Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. Mpnet: Masked and permuted pre-training for language understanding. Advances in neural information processing systems, 33:16857–16867.
  37. 37.Gaurav Suri, Lily R Slater, Ali Ziaee, and Morgan Nguyen. 2024. Do large language models show decision heuristics similar to humans? a case study using gpt-3.5. Journal of Experimental Psychology: General.
  38. 38.Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. 2020. Long range arena: A benchmark for efficient transformers. arXiv preprint arXiv:2011.04006.
  39. 39.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022. Efficient transformers: A survey. ACM Computing Surveys, 55(6):1–28.
  40. 40.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. ArXiv, abs/2302.13971.
  41. 41.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. musique: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539–554.
  42. 42.Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek, Yuhuai Wu, Henryk Michalewski, and Piotr Miłos. 2024. Focused transformer: Contrastive training for context scaling. Advances in Neural Information Processing Systems, 36.
  43. 43.Alex Wang, Richard Yuanzhe Pang, Angelica Chen, Jason Phang, and Samuel R Bowman. 2022. Squality: Building a long-document summarization dataset the hard way. arXiv preprint arXiv:2205.11465.
  44. 44.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.
  45. 45.Rodrigo Wilkens. 2023. Statistical methods for annotation analysis. Computational Linguistics, pages 763–765.
  46. 46.Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862.
  47. 47.Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et al. 2023. Effective long-context scaling of foundation models. arXiv preprint arXiv:2309.16039.
  48. 48.Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, and Bryan Catanzaro. 2023. Retrieval meets long context large language models. arXiv preprint arXiv:2310.03025.
  49. 49.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600.
  50. 50.Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, and Tao Kong. 2023. What matters in training a gpt4-style language model with multimodal inputs? arXiv preprint arXiv:2307.02469.
  51. 51.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
  52. 52.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36.
  53. 53.Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han. 2023. Don’t make your llm an evaluation benchmark cheater. arXiv preprint arXiv:2311.01964.

Citation

MLA
Li, J., et al. “LooGLE: Can Long-Context Language Models Understand Long Contexts?”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 16304–33, https://doi.org/10.18653/v1/2024.acl-long.859.
APA
Li, J., Wang, M., Zheng, Z., & Zhang, M. (2024). LooGLE: Can Long-Context Language Models Understand Long Contexts?. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16304–16333. https://doi.org/10.18653/v1/2024.acl-long.859
Chicago
Li, J., M. Wang, Z. Zheng, and M. Zhang. 2024. “LooGLE: Can Long-Context Language Models Understand Long Contexts?”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16304–33. https://doi.org/10.18653/v1/2024.acl-long.859.
Harvard
Li, J. et al. (2024) “LooGLE: Can Long-Context Language Models Understand Long Contexts?”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 16304–16333. Available at: https://doi.org/10.18653/v1/2024.acl-long.859.
Vancouver
1. Li J, Wang M, Zheng Z, Zhang M (2024) LooGLE: Can Long-Context Language Models Understand Long Contexts?. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 16304–16333

BibTeX

@inproceedings{li-etal-2024-loogle,
    title = "{L}oo{GLE}: Can Long-Context Language Models Understand Long Contexts?",
    author = "Li, Jiaqi  and
      Wang, Mengmeng  and
      Zheng, Zilong  and
      Zhang, Muhan",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.859/",
    doi = "10.18653/v1/2024.acl-long.859",
    pages = "16304--16333"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/