Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
Chonghua WangHaodong DuanSongyang ZhangDahua LinKai Chen
Introduces Ada-LEval, a length-adaptable benchmark scaling up to 128k tokens with two rigorous tasks that require whole-document comprehension, exposing severe performance degradation in leading large language models on ultra-long contexts.
Recent advances in artificial intelligence have led to models claiming context windows capable of processing massive documents containing up to hundreds of thousands of words. However, conventional evaluation benchmarks mainly mix varying text lengths together and focus on tasks like summarization and standard question-answering, which often do not test full-text comprehension or scale to extreme lengths. Consequently, organizations lack reliable measurements to determine whether these models can genuinely reason over extensive, multi-page documents.
The article introduces and evaluates Ada-LEval, a new benchmark designed to rigorously assess the comprehension capabilities of language models across adaptable text lengths ranging from 1,000 to 128,000 tokens.
To conduct this evaluation, the researchers developed two objective tasks that mandate complete document understanding rather than simple surface matching: sorting shuffled book segments into their correct sequence and identifying the best programming answer among numerous distractor candidates. The benchmark was applied to ten major models—including four leading commercial systems and six open-source models—under strictly controlled zero-shot settings with unambiguous accuracy metrics.
The evaluation revealed a steep performance decline across all models as document length increased. Under ultra-long settings exceeding 32,000 tokens, every evaluated model collapsed to near-zero accuracy, matching or falling below random guessing despite vendor claims of ultra-long capabilities. While proprietary commercial models outperformed open-source alternatives at moderate lengths, open-source models deteriorated rapidly to random guess levels once inputs exceeded roughly 4,000 tokens. Detailed error analysis showed that models frequently failed because they could not maintain instruction-following formats in long texts, copied example prompts, or exhibited severe position bias by favoring text placed at the very beginning of the input.
These findings indicate a substantial gap between advertised context window sizes and actual analytical capability. Relying on current models for complex, unassisted reasoning over long documents introduces considerable operational risk and false confidence. While scalable position embedding techniques show promise in extending effective lengths without full model retraining, they do not resolve fundamental reasoning failures over extended inputs.
Organizations should exercise caution before deploying language models for autonomous end-to-end processing of ultra-long documents. Decision-makers should implement multi-stage processing pipelines—such as divide-and-conquer retrieval frameworks—rather than feeding massive documents directly into a single model. Additionally, AI developers must focus future research on improving long-text instruction following and mitigating position bias.
These conclusions are bounded by the high computational costs and API expenses that limited sample sizes for proprietary systems at ultra-long tiers. Furthermore, current open-source models struggled so significantly with basic instruction adherence that measuring their deeper reasoning capabilities at scale remains challenging. Nonetheless, the evidence clearly shows that current systems cannot reliably comprehend ultra-long contexts.
- Paper: SCROLLS: Standardized CompaRison Over Long Language Sequences, Uri Shaham et al. (2022). SCROLLS establishes an earlier standardized benchmark for evaluating comprehension and reasoning across naturally long documents, framing the evaluation gap Ada-LEval addresses.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). Lost in the Middle demonstrates how relevant information’s position within long prompts affects model performance, a key failure mode Ada-LEval also investigates.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2024). LongBench provides an earlier long-context benchmark and length-stratified evaluation framework that helps situate Ada-LEval’s more demanding comprehension tests.
- Paper: NoLiMa: Long-Context Evaluation Beyond Literal Matching, Ali Modarressi et al. (2025). NoLiMa extends the long-context evaluation challenge by testing retrieval when literal query-to-evidence matches are unavailable, probing a further limit beyond Ada-LEval’s full-document tasks.
