Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks

Chonghua WangHaodong DuanSongyang ZhangDahua LinKai Chen

article2024NAACL51 citations

Introduces Ada-LEval, a length-adaptable benchmark scaling up to 128k tokens with two rigorous tasks that require whole-document comprehension, exposing severe performance degradation in leading large language models on ultra-long contexts.

Listen

Recent advances in artificial intelligence have led to models claiming context windows capable of processing massive documents containing up to hundreds of thousands of words. However, conventional evaluation benchmarks mainly mix varying text lengths together and focus on tasks like summarization and standard question-answering, which often do not test full-text comprehension or scale to extreme lengths. Consequently, organizations lack reliable measurements to determine whether these models can genuinely reason over extensive, multi-page documents.

The article introduces and evaluates Ada-LEval, a new benchmark designed to rigorously assess the comprehension capabilities of language models across adaptable text lengths ranging from 1,000 to 128,000 tokens.

To conduct this evaluation, the researchers developed two objective tasks that mandate complete document understanding rather than simple surface matching: sorting shuffled book segments into their correct sequence and identifying the best programming answer among numerous distractor candidates. The benchmark was applied to ten major models—including four leading commercial systems and six open-source models—under strictly controlled zero-shot settings with unambiguous accuracy metrics.

The evaluation revealed a steep performance decline across all models as document length increased. Under ultra-long settings exceeding 32,000 tokens, every evaluated model collapsed to near-zero accuracy, matching or falling below random guessing despite vendor claims of ultra-long capabilities. While proprietary commercial models outperformed open-source alternatives at moderate lengths, open-source models deteriorated rapidly to random guess levels once inputs exceeded roughly 4,000 tokens. Detailed error analysis showed that models frequently failed because they could not maintain instruction-following formats in long texts, copied example prompts, or exhibited severe position bias by favoring text placed at the very beginning of the input.

These findings indicate a substantial gap between advertised context window sizes and actual analytical capability. Relying on current models for complex, unassisted reasoning over long documents introduces considerable operational risk and false confidence. While scalable position embedding techniques show promise in extending effective lengths without full model retraining, they do not resolve fundamental reasoning failures over extended inputs.

Organizations should exercise caution before deploying language models for autonomous end-to-end processing of ultra-long documents. Decision-makers should implement multi-stage processing pipelines—such as divide-and-conquer retrieval frameworks—rather than feeding massive documents directly into a single model. Additionally, AI developers must focus future research on improving long-text instruction following and mitigating position bias.

These conclusions are bounded by the high computational costs and API expenses that limited sample sizes for proprietary systems at ultra-long tiers. Furthermore, current open-source models struggled so significantly with basic instruction adherence that measuring their deeper reasoning capabilities at scale remains challenging. Nonetheless, the evidence clearly shows that current systems cannot reliably comprehend ultra-long contexts.

  • Paper: NoLiMa: Long-Context Evaluation Beyond Literal Matching, Ali Modarressi et al. (2025). NoLiMa extends the long-context evaluation challenge by testing retrieval when literal query-to-evidence matches are unavailable, probing a further limit beyond Ada-LEval’s full-document tasks.
Cover for Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks

Abstract

Recently, the large language model (LLM) community has shown increasing interest in enhancing LLMs’ capability to handle extremely long documents. As various long-text techniques and model architectures emerge, the precise and detailed evaluation of models’ long-text capabilities has become increasingly important. Existing long-text evaluation benchmarks, such as L-Eval and LongBench, construct long-text test sets based on open-source datasets, focusing mainly on QA and summarization tasks. These datasets include test samples of varying lengths (from 2k to 32k+) entangled together, making it challenging to assess model capabilities across different length ranges. Moreover, they do not cover the ultra-long settings (100k+ tokens) that the latest LLMs claim to achieve. In this paper, we introduce Ada-LEval, a length-adaptable benchmark for evaluating the long-context understanding of LLMs. Ada-LEval includes two challenging subsets, TSort and BestAnswer, which enable a more reliable evaluation of LLMs’ long context capabilities. These benchmarks support intricate manipulation of the length of test cases, and can easily produce text samples up to 128k tokens. We evaluate 4 state-of-the-art closed-source API models and 6 open-source models with Ada-LEval. The evaluation results demonstrate the limitations of current LLMs, especially in ultra-long-context settings. Our code is available at https://github.com/open-compass/Ada-LEval.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Long-Context Techniques
  • 2.2 Long-Context Language Models
  • 2.3 Long-Context Benchmarks
  • 3 Ada-LEval
  • 3.1 Task Definition
  • 3.2 Source Data Collection
  • 3.3 Test Case Building
  • 4 Evaluation Results
  • 4.1 Experiment Setup
  • 4.2 Long-Context Evaluation Results
  • 4.3 Error Breakdown
  • 4.4 Ultra-Long-Context Evaluation Results
  • 4.5 Ablation Study
  • 4.5.1 Perplexity Evaluation on TSort
  • 4.5.2 Position Bias in BestAnswer
  • 4.5.3 Scalable Position Embeddings
  • 4.5.4 Comparison with Other Long-Context Benchmarks
  • 5 Conclusion
  • 6 Limitations
  • References
  • A Test Case Building Statistics
  • B Evaluation Setups

Knowls

  1. Knowl 1 — Ada-LEval tests long-context understanding with two length-adaptable tasks

    model/method

    Ada-LEval is a benchmark for assessing long-context language-model understanding using test cases whose lengths can be varied and extended to about 128k tokens. It contains two tasks with closed-form accuracy criteria: TSort asks a model to restore the original order of shuffled text segments from a continuous book passage, while BestAnswer asks it to select the questioner-accepted answer from many candidates. TSort length is controlled through the amount of source text; BestAnswer length is controlled largely by adding distractor answers. The benchmark is designed so that solving the tasks requires reasoning over the supplied content, and correct responses can be scored unambiguously.

  2. Knowl 2 — TSort constructs long cases from shuffled contiguous book passages

    model/method

    TSort uses material from BookSum and Project Gutenberg books. Epistolary works and poetry are excluded, and chapter numbers and annotations are removed to reduce superficial ordering cues. A case presents four shuffled segments drawn from contiguous chapters and asks for their exact original permutation; segments retain complete paragraphs rather than cutting through a paragraph. Adjacent paragraphs before and after the selected material are supplied as contextual hints. Construction varies the length limits for segments and neighboring text, as well as the starting-paragraph stride between cases, then filters cases that exceed the token limit after instructions are added. Exact ordering is required for a correct response.

  3. Knowl 3 — BestAnswer uses accepted Stack Overflow answers and distractors

    model/method

    BestAnswer is built from Stack Overflow programming questions. The authors selected 23 tags and collected up to 2,500 popular questions per tag. For each question, the answer explicitly accepted by its asker is the gold answer—not the answer with the most votes—and questions whose accepted answer is not text-only are excluded. Distractors include other answers to the same question that were posted before the accepted answer, plus answers to similar questions with overlapping tags. Long-context versions retain the same questions while adding distractors, making answer choice sets larger as input length grows. At ultra-long lengths, tag-matching constraints are relaxed because the supply of similar questions and answers is limited.

  4. Knowl 4 — Ada-LEval supplies thousands of cases through 16k tokens and ultra-long cases through 128k

    data/table

    Case lengths are measured with the GPT-4 CL100K tokenizer; the reported maximum and average lengths are token counts. The following entries give setting: number of cases built / maximum tokens / average tokens. TSort: 2k: 5,123 / 2,000 / 1,816; 4k: 5,451 / 4,000 / 3,724; 8k: 5,324 / 8,000 / 7,663; 16k: 4,957 / 16,000 / 15,662; 32k: 2,206 / 32,000 / 31,226; 64k: 1,658 / 64,000 / 62,407; 128k: 782 / 127,800 / 121,488. BestAnswer: 1k: 7,526 / 1,128 / 955; 2k: 7,526 / 2,154 / 1,983; 4k: 7,526 / 4,215 / 3,994; 6k: 7,526 / 6,268 / 6,012; 8k: 7,526 / 7,790 / 7,518; 12k: 7,526 / 12,389 / 12,091; 16k: 7,526 / 15,964 / 15,646; 32k: 200 / 32,974 / 32,329; 64k: 200 / 64,216 / 63,274; 128k: 200 / 127,059 / 126,098. The authors use evaluation subsets rather than all built cases.

  5. Knowl 5 — TSort accuracy is near chance for most evaluated models

    empirical result

    In zero-shot long-context evaluation, TSort accuracy requires the exact ordering of four segments. The random-guess baseline is 4.2%, corresponding to one correct permutation among 24. GPT-4-Turbo-0125 scored 15.5%, 16.5%, 8.5%, and 5.5% at 2k, 4k, 8k, and 16k tokens; GPT-4-Turbo-1106 scored 18.5%, 15.5%, 7.5%, and 3.5%, respectively. Thus, the strongest models exceeded chance at shorter lengths, but their accuracy fell to around chance by 16k. Other proprietary and open-source models were generally close to chance even at shorter lengths; for example, GPT-3.5-Turbo-1106 scored 4.0%, 4.5%, 4.5%, and 5.5% over those settings. The evaluation used 200 cases per length for proprietary models and 1,000 for open-source models.

  6. Knowl 6 — BestAnswer accuracy declines sharply as distractors lengthen the input

    empirical result

    On long-context BestAnswer, GPT-4-Turbo-0125 achieved 73.5%, 73.5%, 65.5%, 63.0%, 56.5%, 52.0%, and 44.5% accuracy at 1k, 2k, 4k, 6k, 8k, 12k, and 16k tokens. GPT-4-Turbo-1106 scored 74.0%, 73.5%, 67.5%, 59.5%, 53.5%, 49.5%, and 44.0%. At 16k, Claude-2 scored 11.0%, while GPT-3.5-Turbo-1106 scored 2.5%; open-source models were at 1.0% or below except for no exceptions in that setting (InternLM2-7b scored 0.8%). Random-guess accuracy, averaged over the questions’ individual candidate counts, was 26.7%, 10.1%, 4.5%, 3.0%, 2.3%, 1.4%, and 1.1% across the same settings. The results show a large proprietary-model advantage, but also a substantial accuracy drop with longer inputs, including for GPT-4-Turbo.

  7. Knowl 7 — Evaluated models largely fail on ultra-long TSort and BestAnswer cases

    empirical result

    Ultra-long-context evaluation used 50 cases per setting. On TSort at 32k, 64k, and 128k tokens, GPT-4-Turbo-0125 scored 2.0%, 4.0%, and 2.0%; GPT-4-Turbo-1106 scored 6.0% at all three lengths; Claude-2 scored 0.0% at 32k and 64k and was not evaluated at 128k; Claude-2.1 scored 0.0% at all three lengths. The TSort random baseline was 4.2%. On BestAnswer, GPT-4-Turbo-0125 scored 30.0%, 0.0%, and 0.0%; GPT-4-Turbo-1106 scored 16.0%, 0.0%, and 0.0%; Claude-2 scored 4.0% and 0.0% at 32k and 64k and was not evaluated at 128k; Claude-2.1 scored 4.0%, 0.0%, and 0.0%; and InternLM2-7b scored 0.5%, 0.5%, and 0.0%. BestAnswer random-guess accuracy was 0.6%, 0.3%, and 0.1%. Consequently, no evaluated model exceeded the random baseline on ultra-long TSort, and none produced a correct BestAnswer response beyond 32k.

  8. Knowl 8 — Instruction-following failures and copying account for prominent benchmark errors

    empirical result

    The authors classify many incorrect responses into two observable failure modes: not following the required response format, or returning the answer shown in the in-context example instead of solving the case. GPT-4-Turbo on TSort and all proprietary models on BestAnswer maintained 100% instruction-following rates across the tested long-context settings; instruction following generally worsened with length for other models. Copying was especially prominent for several models: Claude-2’s TSort copy rate was 99.5%, 95.0%, 97.4%, and 96.9% at 2k, 4k, 8k, and 16k, while the corresponding expected rate at which the example answer was actually correct was only 5.0%, 5.0%, 5.0%, and 5.5%. On BestAnswer, Claude-2’s copy rate increased from 21.5% at 1k to 55.0% at 16k, compared with expected rates of 13.0% and 1.5%. The paper reports copying as a predominant error source for Claude-2, LongChat, and Vicuna models, and poor instruction following as a problem for ChatGLM models.

  9. Knowl 9 — TSort perplexity ranking works better than direct instruction following

    empirical result

    For a separate TSort evaluation, the authors scored all 24 permutations of the four segments by perplexity and selected the lowest-perplexity ordering. Open-source models achieved the following accuracy percentages at 2k, 4k, and 8k tokens: LongChat-7b-v1.5-32k, 60.9%, 68.3%, 77.4%; ChatGLM2-6B-32k, 40.5%, 53.5%, 57.5%; ChatGLM3-6B-32k, 50.1%, 57.0%, 59.3%; Vicuna-7b-v1.5-16k, 70.1%, 78.3%, 77.7%; and Vicuna-13b-v1.5-16k, 79.3%, 86.7%, 89.2%. These results contrast with the much lower accuracy when models are asked directly to output the correct ordering, which the authors interpret as evidence of limited instruction following. One possible explanation they note is that the book chapters may have appeared in pretraining; the reported PPL results therefore do not by themselves establish general sorting ability.

  10. Knowl 10 — BestAnswer exposes sensitivity to the position of the gold answer

    empirical result

    To measure position bias, the authors kept each BestAnswer question and its candidates fixed while moving the accepted answer to the front, middle, or rear of the candidate list. All tested models showed substantial position sensitivity, which generally became more evident at longer input lengths. For GPT-4-Turbo-1106 at 16k tokens, accuracy was 82.0% with the gold answer first, 38.5% with it in the middle, and 26.5% with it at the rear. Vicuna-7b-v1.5-16k scored 25.4%, 0.2%, and 0.7% in those positions at 16k. Claude-2 showed a different pattern: at 16k its corresponding accuracies were 6.0%, 5.0%, and 9.5%, and the authors report that it often performed best with the gold answer at the rear. The controlled comparison indicates that accuracy depends not only on context length and answer content, but also on candidate ordering.

  11. Knowl 11 — Scalable positional embeddings improve Vicuna performance beyond its 4k training window

    empirical result

    The authors evaluated ReRoPE, Leaky ReRoPE, and NTK-aware scaled RoPE on Vicuna-v1.5 models originally trained with a 4k context window, using BestAnswer at 1k, 2k, 4k, and 8k. All three methods improved 8k accuracy over the original 4k model while largely retaining short-context performance. For Vicuna-7b, the original 4k model scored 0.0% at 8k; ReRoPE scored 2.3% accuracy (3.2% after excluding instruction-invalid cases), Leaky ReRoPE 1.3% (2.0%), and NTK 3.9% (3.9%). For Vicuna-13b, the original 4k model scored 0.1% (1.0%) at 8k, compared with ReRoPE at 1.5% (2.8%), Leaky ReRoPE at 1.3% (2.6%), and NTK at 2.3% (2.3%). The paired values are overall accuracy / accuracy after excluding cases that failed the response-format instruction. NTK reduced 1k performance relative to the original 4k model: 32.5% versus 39.5% for 7b, and 43.8% versus 49.1% for 13b. The authors report that scalable methods were competitive with the corresponding 16k-trained Vicuna models, while noting that those models also had Flash Attention and further training on 16k conversations.

  12. Knowl 12 — Truncation comparisons support Ada-LEval's emphasis on whole-input comprehension

    empirical result

    The authors compared GPT-4-Turbo-1106 on the 16k BestAnswer set, NarrativeQA, and GovReport after truncating each input to 2k, 4k, or 8k tokens, and on the full input. BestAnswer accuracy was 11.0%, 20.0%, 31.5%, and 44.0%; NarrativeQA F1 was 24.7, 25.6, 29.7, and 33.1; and GovReport ROUGE-L was 30.7, 32.4, 33.6, and 30.9. The average full-input lengths were 15,646 tokens for BestAnswer, 10,276 for NarrativeQA, and 29,872 for GovReport. Because each benchmark uses its own metric, the absolute scores are not directly comparable across tasks; the relevant pattern is that BestAnswer performance dropped more sharply under truncation, while GovReport performance even rose at 4k and 8k relative to the full-input score. The authors take this as evidence that their tasks place greater demands on comprehensive input understanding than these traditional QA and summarization comparisons.

  13. Knowl 13 — Ada-LEval's difficulty limits its diagnostic value for current models

    limitation

    The authors identify two limitations. First, weak instruction following and high rates of copying the in-context example among open-source models can make accuracy a poor measure for distinguishing those models’ underlying long-context capability. Second, Ada-LEval becomes sharply more difficult at ultra-long lengths, where even state-of-the-art proprietary models perform poorly; this limits the benchmark’s applicability for evaluating models that cannot yet solve its tasks reliably.

Coverage note — The only deliberately omitted contributed detail is the exhaustive per-model copy-rate and position-placement breakdown; the knowls retain the main diagnostic patterns and representative exact values. No other substantial contribution is omitted.

References

  1. 1.Chenxin An, Shansan Gong, Ming Zhong, Mukai Li, Jun Zhang, Lingpeng Kong, and Xipeng Qiu. 2023. L-eval: Instituting standardized evaluation for long context language models. arXiv preprint arXiv:2307.11088.
  2. 2.Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. 2023. Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508.
  3. 3.Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiaoyi Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting Huang, Tao Jiang, Penglong Jiao, Zhenjiang Jin, Zhikai Lei, Jiaxing Li, Jingwen Li, Linyang Li, Shuaibin Li, Wei Li, Yining Li, Hongwei Liu, Jiangning Liu, Jiawei Hong, Kaiwen Liu, Kuikun Liu, Xiaoran Liu, Chengqi Lv, Haijun Lv, Kai Lv, Li Ma, Runyuan Ma, Zerun Ma, Wenchang Ning, Linke Ouyang, Jiantao Qiu, Yuan Qu, Fukai Shang, Yunfan Shao, Demin Song, Zifan Song, Zhihao Sui, Peng Sun, Yu Sun, Huanze Tang, Bin Wang, Guoteng Wang, Jiaqi Wang, Jiayu Wang, Rui Wang, Yudong Wang, Ziyi Wang, Xingjian Wei, Qizhen Weng, Fan Wu, Yingtong Xiong, Chao Xu, Ruiliang Xu, Hang Yan, Yirong Yan, Xiaogui Yang, Haochen Ye, Huaiyuan Ying, Jia Yu, Jing Yu, Yuhang Zang, Chuyu Zhang, Li Zhang, Pan Zhang, Peng Zhang, Ruijie Zhang, Shuo Zhang, Songyang Zhang, Wenjian Zhang, Wenwei Zhang, Xingcheng Zhang, Xinyue Zhang, Hui Zhao, Qian Zhao, Xiaomeng Zhao, Fengzhe Zhou, Zaida Zhou, Jingming Zhuo, Yicheng Zou, Xipeng Qiu, Yu Qiao, and Dahua Lin. 2024. Internlm2 technical report.
  4. 4.Howard Chen, Ramakanth Pasunuru, Jason Weston, and Asli Celikyilmaz. 2023a. Walking down the memory maze: Beyond context limit through interactive reading. arXiv preprint arXiv:2310.05029.
  5. 5.Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023b. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595.
  6. 6.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  7. 7.OpenCompass Contributors. 2023. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass.
  8. 8.Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022a. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 35:16344–16359.
  9. 9.Tri Dao, Daniel Y Fu, Khaled K Saab, Armin W Thomas, Atri Rudra, and Christopher Ré. 2022b. Hungry hungry hippos: Towards language modeling with state space models. arXiv preprint arXiv:2212.14052.
  10. 10.Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A Smith, and Matt Gardner. 2021. A dataset of information-seeking questions and answers anchored in research papers. arXiv preprint arXiv:2105.03011.
  11. 11.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. Llm. int8 (): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339.
  12. 12.Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, and Furu Wei. 2023. Longnet: Scaling transformers to 1,000,000,000 tokens. arXiv preprint arXiv:2307.02486.
  13. 13.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323.
  14. 14.Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. 2021. Longt5: Efficient text-to-text transformer for long sequences. arXiv preprint arXiv:2112.07916.
  15. 15.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  16. 16.Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang. 2021. Efficient attentions for long document summarization. arXiv preprint arXiv:2104.02112.
  17. 17.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, et al. 2023. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. arXiv preprint arXiv:2305.08322.
  18. 18.Tomáš Kociskˇy, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018. The narrativeqa reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328.
  19. 19.Wojciech Krysciński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, and Dragomir Radev. 2021. Booksum: A collection of datasets for long-form narrative summarization. arXiv preprint arXiv:2105.08209.
  20. 20.Dacheng Li*, Rulin Shao*, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Ion Stoica, Xuezhe Ma, , and Hao Zhang. 2023. How long can open-source llms truly promise on context length?
  21. 21.Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023a. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172.
  22. 22.Tianyang Liu, Canwen Xu, and Julian McAuley. 2023b. Repobench: Benchmarking repository-level code auto-completion systems. arXiv preprint arXiv:2306.03091.
  23. 23.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  24. 24.OpenAI. 2023. Gpt-4 technical report.
  25. 25.Ofir Press, Noah A Smith, and Mike Lewis. 2021. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409.
  26. 26.Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, et al. 2022. Scrolls: Standardized comparison over long language sequences. arXiv preprint arXiv:2201.03533.
  27. 27.Jianlin Su. 2023. Rectified rotary position embeddings. https://github.com/bojone/rerope.
  28. 28.Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2021. Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864.
  29. 29.Simeng Sun, Yang Liu, Shuohang Wang, Chenguang Zhu, and Mohit Iyyer. 2023. Pearl: Prompting large language models to plan and execute actions over long documents. arXiv preprint arXiv:2305.14564.
  30. 30.Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei. 2022. A length-extrapolatable transformer. arXiv preprint arXiv:2212.10554.
  31. 31.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. 2022. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261.
  32. 32.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  33. 33.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020. Big bird: Transformers for longer sequences. Advances in neural information processing systems, 33:17283–17297.
  34. 34.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. 2022. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414.
  35. 35.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685.

Citation

MLA
Wang, C., et al. “Ada-LEval: Evaluating Long-context LLMs with Length-adaptable Benchmarks”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 3712–24, https://doi.org/10.18653/v1/2024.naacl-long.205.
APA
Wang, C., Duan, H., Zhang, S., Lin, D., & Chen, K. (2024). Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3712–3724. https://doi.org/10.18653/v1/2024.naacl-long.205
Chicago
Wang, C., H. Duan, S. Zhang, D. Lin, and K. Chen. 2024. “Ada-LEval: Evaluating Long-context LLMs with Length-adaptable Benchmarks”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3712–24. https://doi.org/10.18653/v1/2024.naacl-long.205.
Harvard
Wang, C. et al. (2024) “Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3712–3724. Available at: https://doi.org/10.18653/v1/2024.naacl-long.205.
Vancouver
1. Wang C, Duan H, Zhang S, Lin D, Chen K (2024) Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 3712–3724

BibTeX

@inproceedings{wang-etal-2024-ada,
    title = "{A}da-{LE}val: Evaluating long-context {LLM}s with length-adaptable benchmarks",
    author = "Wang, Chonghua  and
      Duan, Haodong  and
      Zhang, Songyang  and
      Lin, Dahua  and
      Chen, Kai",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.205/",
    doi = "10.18653/v1/2024.naacl-long.205",
    pages = "3712--3724"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/