General-purpose, long-context autoregressive modeling with Perceiver AR
Curtis HawthorneAndrew JaegleCatalina CangeaSebastian BorgeaudCharlie NashMateusz MalinowskiSander DielemanOriol VinyalsMatthew M. BotvinickIan Simon
Introduces an efficient autoregressive architecture that uses causally masked cross-attention to scale directly to over 100,000 input tokens across text, image, and audio domains without requiring custom sparsity patterns.
Modern artificial intelligence models rely heavily on autoregressive generation, where each new output is predicted based on preceding data. However, real-world data such as high-resolution images, long-form literature, and musical performances contain complex long-range dependencies spanning hundreds of thousands of elements. Standard Transformer models struggle with this scale because their computational and memory costs grow quadratically with sequence length and linearly with model depth. As a result, practitioners are often forced into an undesirable trade-off between truncating the input context to save compute or shallowing the network depth to the detriment of expressive power.
The article introduces and evaluates Perceiver AR, a general-purpose autoregressive model architecture designed to process input contexts exceeding one hundred thousand data points without relying on hand-crafted sparsity patterns or specialized memory systems. The central objective is to demonstrate that decoupling input context length from network depth delivers state-of-the-art density estimation and generation quality across multiple domains, including text, imagery, and audio.
The evaluated approach replaces standard full-sequence self-attention with a single, causally masked cross-attention step that compresses a large input context into a much smaller array of latent representations. These latents are subsequently processed by a deep stack of causally masked self-attention layers where the bulk of the computation takes place. The authors tested this architecture through extensive empirical evaluations across a wide range of tasks, including a 131,072-token synthetic copy benchmark, downsampled 64-by-64 ImageNet image modeling, natural language modeling across public and proprietary corpora of up to 4 million books, and musical audio and symbolic generation on the MAESTRO dataset.
The findings establish that Perceiver AR delivers significant gains in both capability and efficiency. On the synthetic copy task, the model achieved 100% accuracy over a distance of more than 131,000 tokens, proving that training signals propagate reliably through the latent bottleneck. On downsampled 64-by-64 ImageNet images, the architecture achieved a state-of-the-art density estimation of 3.40 bits per dimension, matching or outperforming specialized diffusion and sparse autoregressive models. In language modeling, Perceiver AR set a new benchmark on the PG-19 dataset with a test perplexity of 28.9, substantially outperforming previous architectures such as the Compressive Transformer at 33.6. Furthermore, on large book corpora, compute-matched experiments demonstrated that Perceiver AR consistently outperforms Transformer-XL across various context lengths and model depths. Decoupling the input length also allowed flexible test-time compute scaling: reducing the number of evaluation latents degraded performance gracefully while cutting inference latency by more than half when paired with activation caching.
These results indicate that long-range context can be effectively harnessed across modalities without designing custom, domain-specific attention heuristics. By enabling deep networks to observe extensive history at manageable computational costs, the architecture reduces operational trade-offs between speed, cost, and output fidelity. Additionally, the ability to dynamically adjust the number of latents at test time without retraining allows engineering teams to tailor model deployment to specific hardware, latency, or quality budgets.
Organizations developing models for long-form data should consider Perceiver AR as a unified alternative to complex hybrid memory architectures. Future initiatives should explore incorporating variable latent counts directly into the training process to improve inference flexibility further, as well as testing hybrid configurations with other emerging efficient attention mechanisms. Practitioners must note, however, that on smaller datasets like Wikitext-103, expanding context beyond 2,048 tokens does not yield performance gains and requires careful regularization, such as cross-attention dropout, to prevent overfitting. Overall, confidence in the architecture's scalability and empirical robustness is high across diverse data domains.
- Paper: Perceiver: General Perception with Iterative Attention, Andrew Jaegle et al. (2021). It introduces the cross-attention latent bottleneck architecture that Perceiver AR adapts to autoregressive generation with causal masking.
- Paper: Generating Long Sequences with Sparse Transformers, Rewon Child et al. (2019). It establishes the foundational sparse attention paradigms and long-context benchmarks for autoregressive modeling that Perceiver AR seeks to scale beyond.
- Paper: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context, Zihang Dai et al. (2019). It introduces recurrence and segment-level memory mechanisms to extend autoregressive context lengths, providing a classic baseline against which Perceiver AR's direct attention approach is positioned.
- Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). It provides a comprehensive taxonomy of the computational bottlenecks in standard attention and surveys the landscape of efficient architectures targeting long contexts.
- Paper: Rethinking Attention with Performers, Krzysztof Choromanski et al. (2021). It explores kernel-based linear-complexity approximations to attention, illustrating alternative scaling mechanisms for long sequences.
- Paper: Reformer: The Efficient Transformer, Nikita Kitaev et al. (2020). It demonstrates hashing and reversible layers to tackle quadratic memory growth in long sequence autoregressive generation.
- Paper: Image Transformer, Niki Parmar et al. (2018). It introduces self-attention architectures for autoregressive image density estimation, establishing key likelihood benchmarks on ImageNet evaluated by Perceiver AR.
- Paper: Generative Pretraining From Pixels, Mark Chen et al. (2020). It establishes autoregressive pre-training directly on raw pixels as a scalable paradigm for high-dimensional generative modeling.
- Paper: Ring Attention with Blockwise Transformers for Near-Infinite Context, Hao Liu et al. (2024). It develops distributed blockwise attention across GPU clusters to scale exact long-context autoregressive processing to millions of tokens without latent compression.
- Paper: Titans: Learning to Memorize at Test Time, Ali Behrouz et al. (2024). It designs test-time learning neural memory modules to capture multi-million token sequences as an alternative to fixed latent cross-attention architectures.
- Paper: Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, Tri Dao et al. (2024). It unifies state-space models and attention into structured state space duality, providing linear-time scaling for long contexts.
- Paper: Memorizing Transformers, Yuhuai Wu et al. (2022). It extends long-context autoregressive capabilities on benchmarks like PG-19 by integrating an external approximate nearest-neighbor memory cache.
- Paper: DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression, DeepSeek-AI (2026). It pushes long-sequence inference efficiency further by advancing KV cache compression and sparse attention mechanisms across million-token contexts.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). It builds on unified autoregressive modeling to combine multimodal understanding and visual generation within a single transformer.
