General-purpose, long-context autoregressive modeling with Perceiver AR

Curtis HawthorneAndrew JaegleCatalina CangeaSebastian BorgeaudCharlie NashMateusz MalinowskiSander DielemanOriol VinyalsMatthew M. BotvinickIan Simon

article2022ICML88 citations

Introduces an efficient autoregressive architecture that uses causally masked cross-attention to scale directly to over 100,000 input tokens across text, image, and audio domains without requiring custom sparsity patterns.

Listen

Modern artificial intelligence models rely heavily on autoregressive generation, where each new output is predicted based on preceding data. However, real-world data such as high-resolution images, long-form literature, and musical performances contain complex long-range dependencies spanning hundreds of thousands of elements. Standard Transformer models struggle with this scale because their computational and memory costs grow quadratically with sequence length and linearly with model depth. As a result, practitioners are often forced into an undesirable trade-off between truncating the input context to save compute or shallowing the network depth to the detriment of expressive power.

The article introduces and evaluates Perceiver AR, a general-purpose autoregressive model architecture designed to process input contexts exceeding one hundred thousand data points without relying on hand-crafted sparsity patterns or specialized memory systems. The central objective is to demonstrate that decoupling input context length from network depth delivers state-of-the-art density estimation and generation quality across multiple domains, including text, imagery, and audio.

The evaluated approach replaces standard full-sequence self-attention with a single, causally masked cross-attention step that compresses a large input context into a much smaller array of latent representations. These latents are subsequently processed by a deep stack of causally masked self-attention layers where the bulk of the computation takes place. The authors tested this architecture through extensive empirical evaluations across a wide range of tasks, including a 131,072-token synthetic copy benchmark, downsampled 64-by-64 ImageNet image modeling, natural language modeling across public and proprietary corpora of up to 4 million books, and musical audio and symbolic generation on the MAESTRO dataset.

The findings establish that Perceiver AR delivers significant gains in both capability and efficiency. On the synthetic copy task, the model achieved 100% accuracy over a distance of more than 131,000 tokens, proving that training signals propagate reliably through the latent bottleneck. On downsampled 64-by-64 ImageNet images, the architecture achieved a state-of-the-art density estimation of 3.40 bits per dimension, matching or outperforming specialized diffusion and sparse autoregressive models. In language modeling, Perceiver AR set a new benchmark on the PG-19 dataset with a test perplexity of 28.9, substantially outperforming previous architectures such as the Compressive Transformer at 33.6. Furthermore, on large book corpora, compute-matched experiments demonstrated that Perceiver AR consistently outperforms Transformer-XL across various context lengths and model depths. Decoupling the input length also allowed flexible test-time compute scaling: reducing the number of evaluation latents degraded performance gracefully while cutting inference latency by more than half when paired with activation caching.

These results indicate that long-range context can be effectively harnessed across modalities without designing custom, domain-specific attention heuristics. By enabling deep networks to observe extensive history at manageable computational costs, the architecture reduces operational trade-offs between speed, cost, and output fidelity. Additionally, the ability to dynamically adjust the number of latents at test time without retraining allows engineering teams to tailor model deployment to specific hardware, latency, or quality budgets.

Organizations developing models for long-form data should consider Perceiver AR as a unified alternative to complex hybrid memory architectures. Future initiatives should explore incorporating variable latent counts directly into the training process to improve inference flexibility further, as well as testing hybrid configurations with other emerging efficient attention mechanisms. Practitioners must note, however, that on smaller datasets like Wikitext-103, expanding context beyond 2,048 tokens does not yield performance gains and requires careful regularization, such as cross-attention dropout, to prevent overfitting. Overall, confidence in the architecture's scalability and empirical robustness is high across diverse data domains.

  • Paper: Perceiver: General Perception with Iterative Attention, Andrew Jaegle et al. (2021). It introduces the cross-attention latent bottleneck architecture that Perceiver AR adapts to autoregressive generation with causal masking.
  • Paper: Generating Long Sequences with Sparse Transformers, Rewon Child et al. (2019). It establishes the foundational sparse attention paradigms and long-context benchmarks for autoregressive modeling that Perceiver AR seeks to scale beyond.
  • Paper: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context, Zihang Dai et al. (2019). It introduces recurrence and segment-level memory mechanisms to extend autoregressive context lengths, providing a classic baseline against which Perceiver AR's direct attention approach is positioned.
  • Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). It provides a comprehensive taxonomy of the computational bottlenecks in standard attention and surveys the landscape of efficient architectures targeting long contexts.
  • Paper: Rethinking Attention with Performers, Krzysztof Choromanski et al. (2021). It explores kernel-based linear-complexity approximations to attention, illustrating alternative scaling mechanisms for long sequences.
  • Paper: Reformer: The Efficient Transformer, Nikita Kitaev et al. (2020). It demonstrates hashing and reversible layers to tackle quadratic memory growth in long sequence autoregressive generation.
  • Paper: Image Transformer, Niki Parmar et al. (2018). It introduces self-attention architectures for autoregressive image density estimation, establishing key likelihood benchmarks on ImageNet evaluated by Perceiver AR.
  • Paper: Generative Pretraining From Pixels, Mark Chen et al. (2020). It establishes autoregressive pre-training directly on raw pixels as a scalable paradigm for high-dimensional generative modeling.
Cover for General-purpose, long-context autoregressive modeling with Perceiver AR

Abstract

Real-world data is high-dimensional: a book, image, or musical performance can easily contain hundreds of thousands of elements even after compression. However, the most commonly used autoregressive models, Transformers, are prohibitively expensive to scale to the number of inputs and layers needed to capture this long-range structure. We develop Perceiver AR, an autoregressive, modality-agnostic architecture which uses cross-attention to map long-range inputs to a small number of latents while also maintaining end-to-end causal masking. Perceiver AR can directly attend to over a hundred thousand tokens, enabling practical long-context density estimation without the need for hand-crafted sparsity patterns or memory mechanisms. When trained on images or music, Perceiver AR generates outputs with clear long-term coherence and structure. Our architecture also obtains state-of-the-art likelihood on long-sequence benchmarks, including 64 × 64 ImageNet images and PG-19 books.

Table of Contents

  • 1. Introduction
  • 2. Autoregression and long-context modeling
  • 3. Perceiver AR
  • 4. Related work
  • 4.1. Relationship to Transformer and Transformer-XL
  • 4.2. Relationship to other scalable attention architectures
  • 4.3. Relationship to encoder-decoder architectures
  • 5. Results
  • 5.1. Copy Task
  • 5.1.1. LONG INPUT LENGTH
  • 5.1.2. NUMBER OF TRAINING TARGETS
  • 5.2. ImageNet 64 × 64
  • 5.2.1. VARYING COMPUTE AT TEST TIME
  • 5.2.2. IMPACT OF SEQUENCE ORDERING AND LENGTH
  • 5.3. Project Gutenberg (PG-19)
  • 5.4. Books
  • 5.5. Wikitext-103
  • 5.6. MAESTRO
  • 5.7. Music samples
  • 6. Conclusion
  • 7. Acknowledgements
  • References
  • A. ImageNet Samples
  • B. Books
  • C. A More Detailed Look at Perceiver AR's Internals
  • D. Further Related Work
  • D.1. Efficient attention
  • D.2. Other efficient architectures
  • D.3. Input Tokenization
  • E. Additional Details of the Methods
  • E.1. Memory Usage
  • E.2. Cross-attend Dropout
  • E.3. Activation Caching for Inference
  • E.4. Varying Compute at Test Time
  • F. Training Details
  • F.1. Common
  • F.2. Rotary Position Encodings
  • F.3. Copy Task
  • F.4. ImageNet 64 × 64
  • F.5. Wikitext-103
  • F.6. PG-19
  • F.7. Books
  • F.8. Music Generation Tasks
  • G. Evaluation Details
  • H. Dropout Ablations
  • I. MAESTRO SoundStream

Knowls

  1. Knowl 1 — Perceiver AR Architecture

    model/method

    Perceiver AR is an autoregressive neural architecture that decouples input context length from network depth by utilizing a causally masked cross-attention module followed by a stack of causally masked self-attention layers.

    Given an input sequence X∈RM×CX \in \mathbb{R}^{M \times C}, where MM is the number of input tokens and CC is the channel dimension, the model constructs a smaller latent sequence Z0∈RN×CZ_0 \in \mathbb{R}^{N \times C} with N<MN < M. By default, the query input to the initial cross-attention module is formed from the final NN tokens of the input sequence, XQ=X[−N:,:]X_Q = X[-N:, :].

    The full forward pass is defined as:

    Z0=CrossAttendcm(X,X[−N:,:])Z_0 = \text{CrossAttend}_{cm}(X, X[-N:, :])

    Zl+1=SelfAttendcm(Zl,Zl)for l=0,…,L−1Z_{l+1} = \text{SelfAttend}_{cm}(Z_l, Z_l) \quad \text{for } l = 0, \dots, L-1

    where CrossAttendcm\text{CrossAttend}_{cm} and SelfAttendcm\text{SelfAttend}_{cm} denote causally masked cross-attention and self-attention operations, respectively, each wrapped in pre-layer normalization and followed by residual connections and a two-layer multi-layer perceptron (MLP) with squared-ReLU activations. The final latents ZL∈RN×CZ_L \in \mathbb{R}^{N \times C} are layer-normalized and projected to output vocabulary logits.

    The computational complexity of Perceiver AR is O(MN)+O(LN2)\mathcal{O}(MN) + \mathcal{O}(LN^2), where O(MN)\mathcal{O}(MN) arises from the single initial cross-attention layer and O(LN2)\mathcal{O}(LN^2) is the cost of the LL self-attention layers over NN latents.

  2. Knowl 2 — Causal Masking in Perceiver AR Cross- and Self-Attention

    equation

    To preserve autoregressive dependencies end-to-end when mapping MM input tokens to NN output target latents (N<MN < M), causal masking is enforced in both the initial cross-attention layer and the subsequent latent self-attention layers.

    For the initial cross-attention layer, queries correspond to the trailing NN positions of the input sequence X∈RM×CX \in \mathbb{R}^{M \times C}, indexed by n∈{0,1,…,N−1}n \in \{0, 1, \dots, N-1\}, while keys/values correspond to all positions m∈{0,1,…,M−1}m \in \{0, 1, \dots, M-1\}. Pre-softmax attention logits XQKpre(n,m)X^{\text{pre}}_{QK}(n, m) are causally masked by setting masked entries to −∞-\infty:

    XQKpre(n,m)=−∞if m>n+M−N−1X^{\text{pre}}_{QK}(n, m) = -\infty \quad \text{if } m > n + M - N - 1

    For each latent self-attention layer l∈{0,…,L−1}l \in \{0, \dots, L-1\}, both queries and keys are of length NN. For query index m′∈{0,1,…,N−1}m' \in \{0, 1, \dots, N-1\} and key index m∈{0,1,…,N−1}m \in \{0, 1, \dots, N-1\}, the standard causal self-attention mask is applied:

    XQKpre(m′,m)=−∞if m>m′X^{\text{pre}}_{QK}(m', m) = -\infty \quad \text{if } m > m'

    These masking constraints ensure that the prediction for each position depends strictly on preceding sequence elements.

  3. Knowl 3 — Periodic Memory Resetting for Inference Activation Caching

    algorithm

    Autoregressive generation in Perceiver AR with key-value activation caching requires avoiding the accumulation of latent states across more positions than were present during training (NN). Naive caching would allow long-distance expired latents to influence active latents, introducing out-of-distribution dependencies. Periodic memory resetting maintains a fixed-width activation cache of size NN matching training time.

    Input: Model weights θ\theta, prompt sequence XX, target total length TT, latent capacity NN
    Output: Generated token sequence XX
    Initialize cache buffer with capacity NN
    while length of X<TX < T do
        if cache buffer is full (contains NN latent states) then
            Clear cache buffer
            Execute full uncached forward pass using N/2N/2 latents on the most recent context
            Store the resulting N/2N/2 activation states into cache buffer
        end if
        Compute next token distribution using cached states plus the current step
        Sample next token xnextx_{\text{next}} from distribution
        Append xnextx_{\text{next}} to XX
        Update cache buffer with newly computed activation states
    end while
    return XX
  4. Knowl 4 — Downsampled ImageNet 64x64 Density Estimation

    data/table

    Perceiver AR was evaluated on unconditional density estimation using downsampled 64×6464 \times 64 ImageNet. Images were serialized in raster scan order with RGB subpixel channels (sequence length M=12,289M = 12,289 including a [BOS][\text{BOS}] token). The model used L=60L = 60 self-attention layers with N=1024N = 1024 latents and 770.1M parameters, trained for 750k steps.

    Model Type Bits/Dim
    PixelCNN AR 3.57
    Sparse Transformer AR 3.44
    Routing Transformer AR 3.43
    Combiner AR 3.42
    VDM Diff 3.40
    Perceiver AR (ours) AR 3.40

    Perceiver AR achieved 3.40 bits/dim on the validation set, outperforming prior autoregressive models and matching the Variational Diffusion Model (VDM).

  5. Knowl 5 — Test-Time Compute and Quality Trade-Off via Latent Count

    data/table

    Because Perceiver AR decouples input context length from latent stack width and does not learn position-specific weights in its self-attention layers, the number of latents NN evaluated at inference time can differ from the number used during training (N=1024N=1024). Evaluated on downsampled 64×6464 \times 64 ImageNet (M=12,289M = 12,289 tokens) using a 60-layer model on a single TPUv3 core with activation caching:

    # Latents (NN) Bits/Dim Generation Time (minutes)
    16 3.5664 1.99
    64 3.4946 2.02
    512 3.4139 2.81
    1024 3.4025 3.68
    1536 3.3986 4.69
    2048 3.4018 5.88
    4096 3.5569 12.28

    Increasing the latent count up to N=1536N = 1536 improves likelihood over the train-time configuration (N=1024N = 1024). Reducing NN down to 16 reduces generation time by nearly half while maintaining performance comparable to PixelCNN (3.57 bits/dim).

  6. Knowl 6 — PG-19 Long-Context Language Modeling Results

    data/table

    Perceiver AR was evaluated on the Project Gutenberg (PG-19) benchmark using Subword tokenization. Models were configured with L=60L = 60 self-attention layers, 974.6M parameters, input embeddings of dimension 4096, 128 cross-attention heads, and cross-attention dropout (0.875 for 2048 context, 0.96875 for 4096 context).

    Model Context Length # Layers Val ppl. Test ppl.
    Transformer-XL 512 + 1024 36 45.5 36.3
    Compressive Transformer 512 + 512 + 2×\times512 36 43.4 33.6
    Routing Transformer 8192 22 – 33.2
    Perceiver AR (ours) 2048 60 45.9 28.9
    Perceiver AR (ours) 4096 60 45.9 29.0

    Perceiver AR achieved 28.9 test perplexity, outperforming prior state-of-the-art architectures on the PG-19 benchmark.

  7. Knowl 7 — Long-Context Robustness Under Sequence Ordering Permutations

    empirical result

    On 64×6464 \times 64 ImageNet images (12,288 tokens), autoregressive models were evaluated under two channel orderings: standard raster-scan ordering (where R, G, and B subpixels for each pixel are predicted consecutively) and channel-separated ordering (R→G→BR \to G \to B, where all red subpixels are predicted first, followed by all green, then all blue).

    Input Context Standard Ordering R→G→BR \to G \to B
    1024 3.55 4.63
    12289 3.54 3.53

    For short context models (1024 tokens), separating color channels severely degraded validation performance from 3.55 to 4.63 bits/dim because spatial co-dependencies span long distances. In contrast, the full-context Perceiver AR (12,289 tokens) achieved 3.53 bits/dim on the reordered sequence, demonstrating the capacity of the cross-attention layer to access and utilize long-range antecedents.

  8. Knowl 8 — Compute-Matched Evaluation Against Transformer-XL on Book Corpus

    data/table

    Perceiver AR and Transformer-XL were evaluated on a 4-million book corpus published between 1500 and 2008. Models were trained for 500k steps with batch size 256, matching training throughput (steps per second) across configurations.

    First, 36-layer Perceiver AR models (498.9M parameters, N=1024N=1024 latents) were compared against depth-matched Transformer-XL models (context 1024):

    Model Context Depth Steps/Sec Eval ppl.
    AR 1024 36 2.19 14.006
    T-XL 1024 23 2.17 14.822
    AR 4096 36 2.09 13.806
    T-XL 1024 24 2.06 14.719
    AR 8192 36 1.95 13.791
    T-XL 1024 25 1.97 14.593
    AR 16384 36 1.75 13.749
    T-XL 1024 28 1.76 14.276

    Second, the deepest Transformer-XL that fit in device memory (42 layers, 605.8M parameters) was compared against compute-matched deep Perceiver AR models with varying context lengths:

    Model Context Depth Steps/Sec Eval ppl.
    T-XL 1024 42-layer 1.17 13.253
    AR 1024 62-layer 1.19 12.849
    AR 4096 61-layer 1.21 12.680
    AR 8192 60-layer 1.25 12.660
    AR 16384 56-layer 1.26 12.816

    In all compute-matched configurations, Perceiver AR achieved lower perplexity than Transformer-XL.

  9. Knowl 9 — Fractional Rotary Position Encoding

    model/method

    Perceiver AR uses Rotary Position Encodings (RoPE) applied to only a fraction of the channel dimensions within each query and key attention head (e.g., the first 25% or 50%), leaving the remaining dimensions unrotated.

    Applying RoPE to a subset of dimensions reduces the number of frequency bands used to encode relative distance, allowing the network to simultaneously represent position-dependent and position-agnostic relationships between queries and keys. Additionally, fractional rotation reduces the compute and memory overhead of the rotary matrix operations compared to full-dimension rotation.

  10. Knowl 10 — Synthetic Long-Context Copy Task Performance and Target Count Trade-off

    empirical result

    Perceiver AR was evaluated on a synthetic byte-copying task with sequence length M=217=131,072M = 2^{17} = 131,072 tokens. The input consisted of a [BOS][\text{BOS}] token followed by 65,535 random byte tokens, which were then mirrored for the second half of the sequence and terminated with [EOS][\text{EOS}], yielding a maximum copy distance of 217−22^{17} - 2 tokens.

    A 6-layer Perceiver AR with N=1024N = 1024 latents trained for 25k steps achieved 100% accuracy across 786,432 test tokens on unseen sequences, demonstrating that gradients from 1024 target positions can successfully train the cross-attention bottleneck to retrieve exact tokens over 100k positions away.

    On an 8,192-token context copy task with a 1-layer self-attention stack, model convergence depended jointly on batch size and latent target count:

    Batch Size 1024 Latents 2048 Latents
    64 Failed (<1%<1\%) Converged (100%100\%)
    128 Converged (100%100\%) Converged (100%100\%)

    Doubling the number of latents NN per sequence provides additional target loss signals, compensating for a smaller batch size.

  11. Knowl 11 — Symbolic and Audio Music Modeling with Perceiver AR

    empirical result

    Perceiver AR was evaluated on symbolic MIDI music and neural-codec audio representations using the MAESTRO dataset.

    For symbolic music (MAESTRO v1), a 12-layer Perceiver AR with context length 4096 and N=2048N = 2048 latents achieved a negative log-likelihood (NLL) of 1.82 on the test set, improving upon Music Transformer (1.84 NLL), without using data augmentation. On MAESTRO v3, Perceiver AR achieved 1.91 test NLL.

    For raw audio converted into discrete tokens via the SoundStream neural audio codec at a fixed context length of 65,536 tokens, Perceiver AR achieved the following negative log-likelihoods:

    SoundStream Bitrate Context Duration Test NLL Val NLL
    12 kbps 54.4 s 2.49 2.34
    18 kbps 36.8 s 2.60 2.57
    22 kbps 29.6 s 2.65 2.62

Coverage note — No substantial contributed material was omitted. Wikitext-103 experiments are summarized in the text but omitted as a standalone knowl because the authors note the small benchmark is not bottlenecked by long context and yields results on par with existing baselines.

References

  1. 1.Baevski, A. and Auli, M. Adaptive input representations for neural language modeling. In Proceedings of the International Conference on Learning Representations (ICLR), 2018.
  2. 2.Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C. A neural probabilistic language model. Journal of Machine Learning Research (JMLR), 2003.
  3. 3.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  4. 4.Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I. Generative pretraining from pixels. In Proceedings of the International Conference on Machine Learning (ICML), 2020.
  5. 5.Child, R., Gray, S., Radford, A., and Sutskever, I. Generating long sequences with sparse Transformers. arXiv preprint arXiv:1904.10509, 2019.
  6. 6.Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A. Rethinking attention with Performers. In Proceedings of the International Conference on Learning Representations (ICLR).
  7. 7.Clark, J. H., Garrette, D., Turc, I., and Wieting, J. CANINE: pre-training an efficient tokenization-free encoder for language representation. arXiv preprint arXiv:2103.06874, 2021.
  8. 8.Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R. Transformer-XL: Attentive language models beyond a fixed-length context. In Proceedings of the Annual Meetings of the Association for Computational Linguistics (ACL), 2019.
  9. 9.Dai, Z., Lai, G., Yang, Y., and Le, Q. Funnel-Transformer: Filtering out sequential redundancy for efficient language processing. In Proceedings of Neural Information Processing Systems (NeurIPS), 2020.
  10. 10.Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341, 2020.
  11. 11.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR), 2021.
  12. 12.Graves, A. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850, 2013.
  13. 13.Graves, A., Wayne, G., and Danihelka, I. Neural Turing Machines. arXiv preprint arXiv:1410.5401, 2014.
  14. 14.Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwinska, A., Colmenarejo, S. G., Grefenestette, E., Ramalho, T., Agapiou, J., Badia, A. P., Hermann, K. M., Zwols, Y., Ostrovski, G., Cain, A., King, H., Summerfield, C., Blunsom, P., Kavukcuoglu, K., and Hassabis, D. Hybrid computing using a neural network with dynamic external memory. Nature, 538:471–476, 2016.
  15. 15.Gu, A., Goel, K., and Ré, C. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021.
  16. 16.Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gerard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E. Array programming with NumPy. Nature, 585(7825):357–362, 2020.
  17. 17.Hawthorne, C., Elsen, E., Song, J., Roberts, A., Simon, I., Raffel, C., Engel, J., Oore, S., and Eck, D. Onsets and frames: Dual-objective piano transcription. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2018.
  18. 18.Hawthorne, C., Stasyuk, A., Roberts, A., Simon, I., Huang, C.-Z. A., Dieleman, S., Elsen, E., Engel, J., and Eck, D. Enabling factorized piano music modeling and generation with the MAESTRO dataset. In Proceedings of the International Conference on Learning Representations (ICLR), 2019.
  19. 19.Hendrycks, D. and Gimpel, K. Gaussian error linear units (GELUs). arXiv preprint arXiv:1606.08415, 2016.
  20. 20.Hessel, M., Budden, D., Viola, F., Rosca, M., Sezener, E., and Hennigan, T. Optax: composable gradient transformation and optimisation, in JAX!, 2020. URL http://github.com/deepmind/optax.
  21. 21.Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Shazeer, N., Simon, I., Hawthorne, C., Dai, A. M., Hoffman, M. D., Dinculescu, M., and Eck, D. Music Transformer: Generating music with long-term structure. In Proceedings of the International Conference on Learning Representations (ICLR), 2019.
  22. 22.Jaegle, A., Gimeno, F., Brock, A., Zisserman, A., Vinyals, O., and Carreira, J. Perceiver: General perception with iterative attention. In Proceedings of the International Conference on Machine Learning (ICML), 2021.
  23. 23.Jaegle, A., Borgeaud, S., Alayrac, J.-B., Doersch, C., Ionescu, C., Ding, D., Koppula, S., Zoran, D., Brock, A., Shelhamer, E., Henaff, O., Botvinick, M. M., Zisserman, A., Vinyals, O., and Carreira, J. Perceiver IO: A general architecture for structured inputs & outputs. In Proceedings of the International Conference on Learning Representations (ICLR), 2022.
  24. 24.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Zídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P., and Hassabis, D. Highly accurate protein structure prediction with AlphaFold. Nature, 596:583–589, 2021.
  25. 25.Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F. Transformers are RNNs: Fast autoregressive Transformers with linear attention. In Proceedings of the International Conference on Machine Learning (ICML), 2020.
  26. 26.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations (ICLR), 2015.
  27. 27.Kingma, D. P., Salimans, T., Poole, B., and Ho, J. On density estimation with diffusion models. In Proceedings of Neural Information Processing Systems (NeurIPS), 2021.
  28. 28.Kitaev, N., Kaiser, L., and Levskaya, A. Reformer: The efficient transformer. In Proceedings of the International Conference on Learning Representations (ICLR), 2020.
  29. 29.Kudo, T. and Richardson, J. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the Annual Meetings of the Association for Computational Linguistics (ACL), 2018.
  30. 30.Lakhotia, K., Kharitonov, E., Hsu, W.-N., Adi, Y., Polyak, A., Bolte, B., Nguyen, T.-A., Copet, J., Baevski, A., Mohamed, A., and Dupoux, E. Generative spoken language modeling from raw audio. arXiv preprint arXiv:2102.01192, 2021.
  31. 31.Lee, J., Lee, Y., Kim, J., Kosiorek, A., Choi, S., and Teh, Y. W. Set Transformer: A framework for attention-based permutation-invariant neural networks. In Proceedings of the International Conference on Machine Learning (ICML), 2019.
  32. 32.Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N. Generating Wikipedia by summarizing long sequences. In Proceedings of the International Conference on Learning Representations (ICLR), 2018.
  33. 33.Ma, X., Kong, X., Wang, S., Zhou, C., May, J., Ma, H., and Zettlemoyer, L. LUNA: Linear unified nested attention. In Proceedings of Neural Information Processing Systems (NeurIPS), 2021.
  34. 34.Mehri, S., Kumar, K., Gulrajani, I., Kumar, R., Jain, S., Sotelo, J., Courville, A., and Bengio, Y. SampleRNN: An unconditional end-to-end neural audio generation model. In Proceedings of the International Conference on Learning Representations (ICLR), 2017.
  35. 35.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. In Proceedings of the International Conference on Learning Representations (ICLR), 2017.
  36. 36.Nash, C., Menick, J., Dieleman, S., and Battaglia, P. W. Generating images with sparse representations. In Proceedings of the International Conference on Machine Learning (ICML), 2021.
  37. 37.Nawrot, P., Tworkowski, S., Tyrolski, M., Kaiser, L., Wu, Y., Szegedy, C., and Michalewski, H. Hierarchical Transformers are more efficient language models. arXiv preprint arXiv:2110.13711, 2021.
  38. 38.Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N., and Kong, L. Random feature attention. In Proceedings of the International Conference on Learning Representations (ICLR), 2021.
  39. 39.Polyak, A., Adi, Y., Copet, J., Kharitonov, E., Lakhotia, K., Hsu, W.-N., Mohamed, A., and Dupoux, E. Speech resynthesis from discrete disentangled self-supervised representations. arXiv preprint arXiv:2104.00355, 2021.
  40. 40.Press, O., Smith, N. A., and Lewis, M. Shortformer: Better language modeling using shorter inputs. In Proceedings of the Annual Meetings of the Association for Computational Linguistics (ACL), 2021.
  41. 41.Rabe, M. N. and Staats, C. Self-attention does not need O(n 2) memory. arXiv preprint arXiv:2112.05682, 2021.
  42. 42.Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P. Compressive Transformers for long-range sequence modelling. In Proceedings of the International Conference on Learning Representations (ICLR), 2019.
  43. 43.Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J.-B., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M., Hechtman, B., Weidinger, L., Gabriel, I., Isaac, W., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G. Scaling language models: Methods, analysis & insights from training Gopher. arXiv preprint arXiv:2112.11446, 2021.
  44. 44.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text Transformer. Journal of Machine Learning Research (JMLR), 2020.
  45. 45.Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. ZeRO: Memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC), 2020.
  46. 46.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092, 2021.
  47. 47.Ren, H., Dai, H., Dai, Z., Yang, M., Leskovec, J., Schuurmans, D., and Dai, B. Combiner: Full attention Transformer with sparse computation cost. In Proceedings of Neural Information Processing Systems (NeurIPS), 2021.
  48. 48.Rosenfeld, R. Two decades of statistical language modeling: where do we go from here? Proceedings of the IEEE, 88 (8):1270–1278, 2000.
  49. 49.Roy, A., Saffar, M., Vaswani, A., and Grangier, D. Efficient content-based sparse attention with Routing Transformers. Transactions of the Association for Computational Linguistics (TACL), 9, 2021.
  50. 50.Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M. Image super-resolution via iterative refinement. arXiv preprint arXiv:2104.07636, 2021.
  51. 51.Schmidhuber, J. and Heil, S. Sequential neural text compression. IEEE Transactions on Neural Networks, 7(1): 142–146, 1994.
  52. 52.Shazeer, N. Fast Transformer decoding: one write-head is all you need. arXiv preprint arXiv:1911.02150, 2019.
  53. 53.Simon, I., Huang, C.-Z. A., Engel, J., Hawthorne, C., and Dinculescu, M. Generating piano music with transformer. 2019. URL https://magenta.tensorflow.org/piano-transformer.
  54. 54.So, D. R., Manke, W., Liu, H., Dai, Z., Shazeer, N., and Le, Q. V. Primer: Searching for efficient Transformers for language modeling. arXiv preprint arXiv:2109.08668, 2021.
  55. 55.Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research (JMLR), 2014.
  56. 56.Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y. RoFormer: Enhanced Transformer with rotary position embedding. arXiv preprint arxiv:2104.09864, 2021.
  57. 57.Sun, S., Krishna, K., Mattarella-Micke, A., and Iyyer, M. Do long-range language models actually use long-range context? In Proceedings of the Annual Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021.
  58. 58.Sutskever, I., Vinyals, O., and Le, Q. V. Sequence to sequence learning with neural networks. In Proceedings of Neural Information Processing Systems (NeurIPS), 2014.
  59. 59.Uria, B., Côté, M.-A., Gregor, K., Murray, I., and Larochelle, H. Neural autoregressive distribution estimation. Journal of Machine Learning Research (JMLR), 2016.
  60. 60.van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016a.
  61. 61.van den Oord, A., Kalchbrenner, N., and Kavukcuoglu, K. Pixel recurrent neural networks. In Proceedings of the International Conference on Machine Learning (ICML), 2016b.
  62. 62.van den Oord, A., Vinyals, O., and Kavukcuoglu, K. Neural discrete representation learning. In Proceedings of Neural Information Processing Systems (NeurIPS), 2017.
  63. 63.Vaswani, A., Bengio, S., Brevdo, E., Chollet, F., Gomez, A. N., Gouws, S., Jones, L., Kaiser, L., Kalchbrenner, N., Parmar, N., Sepassi, R., Shazeer, N., and Uszkoreit, J. Tensor2tensor for neural machine translation. arXiv preprint arXiv:1803.07416.
  64. 64.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In Proceedings of Neural Information Processing Systems (NeurIPS), 2017.
  65. 65.Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gulcehre, C., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, Dani Wunsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T., Kavukcuoglu, K., Hassabis, D., Chris, A., and Silver, D. Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature, 575(7782): 350–354, 2019.
  66. 66.Wang, B. Mesh transformer Jax, 2021. URL https://github.com/kingoflolz/mesh-transformer-jax.
  67. 67.Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768, 2020.
  68. 68.Weston, J., Chopra, S., and Bordes, A. Memory networks. In Proceedings of the International Conference on Learning Representations (ICLR), 2015.
  69. 69.Wu, C., Liang, J., Ji, L., Yang, F., Fang, Y., Jiang, D., and Duan, N. NÜWA: Visual synthesis pre-training for neural visual world creation. arXiv preprint arXiv:2111.12417, 2021.
  70. 70.Wu, F., Fan, A., Baevski, A., Dauphin, Y. N., and Auli, M. Pay less attention with lightweight and dynamic convolutions. arXiv preprint arXiv:1901.10430, 2019.
  71. 71.Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C. Memorizing transformers. In Proceedings of the International Conference on Learning Representations (ICLR), 2022. URL https://openreview.net/forum?id=TrjbxzRcnf-.
  72. 72.Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T.-Y. On layer normalization in the Transformer architecture. In Proceedings of the International Conference on Machine Learning (ICML), 2020.
  73. 73.Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al. Big Bird: Transformers for longer sequences. Proceedings of Neural Information Processing Systems (NeurIPS), 33, 2020.
  74. 74.Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M. SoundStream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP), 2021.

Citation

MLA
Hawthorne, C., et al. “General-purpose, Long-context Autoregressive Modeling with Perceiver AR”. International Conference on Machine Learning, vol. 162, 2022, pp. 8535–58, https://proceedings.mlr.press/v162/hawthorne22a.html.
APA
Hawthorne, C., Jaegle, A., Cangea, C., Borgeaud, S., Nash, C., Malinowski, M., Dieleman, S., Vinyals, O., Botvinick, M., Simon, I., Sheahan, H., Zeghidour, N., Alayrac, J.-B., Carreira, J., & Engel, J. (2022). General-purpose, long-context autoregressive modeling with Perceiver AR. International Conference on Machine Learning, 162, 8535–8558. https://proceedings.mlr.press/v162/hawthorne22a.html
Chicago
Hawthorne, C., A. Jaegle, C. Cangea, et al. 2022. “General-purpose, Long-context Autoregressive Modeling with Perceiver AR”. International Conference on Machine Learning 162: 8535–58. https://proceedings.mlr.press/v162/hawthorne22a.html.
Harvard
Hawthorne, C. et al. (2022) “General-purpose, long-context autoregressive modeling with Perceiver AR”, International Conference on Machine Learning. PMLR, pp. 8535–8558. Available at: https://proceedings.mlr.press/v162/hawthorne22a.html.
Vancouver
1. Hawthorne C, Jaegle A, Cangea C, et al (2022) General-purpose, long-context autoregressive modeling with Perceiver AR. In: International Conference on Machine Learning. PMLR, pp 8535–8558

BibTeX

@InProceedings{pmlr-v162-hawthorne22a,
  title = 	 {General-purpose, long-context autoregressive modeling with Perceiver {AR}},
  author =       {Hawthorne, Curtis and Jaegle, Andrew and Cangea, C{\u{a}}t{\u{a}}lina and Borgeaud, Sebastian and Nash, Charlie and Malinowski, Mateusz and Dieleman, Sander and Vinyals, Oriol and Botvinick, Matthew and Simon, Ian and Sheahan, Hannah and Zeghidour, Neil and Alayrac, Jean-Baptiste and Carreira, Joao and Engel, Jesse},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {8535--8558},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/hawthorne22a/hawthorne22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/hawthorne22a.html},
  abstract = 	 {Real-world data is high-dimensional: a book, image, or musical performance can easily contain hundreds of thousands of elements even after compression. However, the most commonly used autoregressive models, Transformers, are prohibitively expensive to scale to the number of inputs and layers needed to capture this long-range structure. We develop Perceiver AR, an autoregressive, modality-agnostic architecture which uses cross-attention to map long-range inputs to a small number of latents while also maintaining end-to-end causal masking. Perceiver AR can directly attend to over a hundred thousand tokens, enabling practical long-context density estimation without the need for hand-crafted sparsity patterns or memory mechanisms. When trained on images or music, Perceiver AR generates outputs with clear long-term coherence and structure. Our architecture also obtains state-of-the-art likelihood on long-sequence benchmarks, including 64x64 ImageNet images and PG-19 books.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/