RandAR: Decoder-only Autoregressive Visual Generation in Random Orders

Ziqi PangTianyuan ZhangFujun LuanYunze ManHao TanKai ZhangWilliam T. FreemanYu-Xiong Wang

article2025CVPR89 citations

Presents a decoder-only visual autoregressive framework that uses position instruction tokens to generate images in arbitrary orders, achieving 2.5x faster parallel decoding alongside zero-shot inpainting, outpainting, and resolution extrapolation without sacrificing visual quality.

Listen

Visual generation models frequently adapt large language model architectures by predicting visual tokens sequentially in a fixed, row-by-row raster order. However, this rigid sequencing forces artificial directional constraints on two-dimensional images, restricting the model from using surrounding visual context and limiting flexibility in downstream editing tasks. Meanwhile, alternative architectures that allow flexible masking often lose compatibility with standard caching optimizations, which hurts operational speed.

The article demonstrates that causal, decoder-only transformers can generate high-quality images in completely arbitrary token sequences without altering their standard architectural foundations. By introducing a framework named RandAR, the evaluation examines whether removing the predefined raster order allows unidirectional generative models to acquire bidirectional contextual reasoning, accelerate generation, and perform zero-shot visual editing.

To achieve random-order generation with minimal architecture changes, the approach inserts a single learnable spatial indicator—termed a position instruction token—prior to each image token to specify its target coordinate. The model was trained across randomly shuffled sequence permutations on standard image datasets such as ImageNet, using established evaluation metrics including Fréchet Inception Distance, Inception Score, and precision-recall measures alongside latency testing on production-grade graphical processing units.

The findings establish that training on random orderings achieves visual fidelity comparable to conventional raster-based models despite the combinatorial complexity of learning across factorial permutations. By decoupling generation from fixed order, the system enables parallel multi-token prediction during inference, delivering an approximate 2.5-fold reduction in generation latency without degrading visual quality. Furthermore, the model demonstrates zero-shot capabilities in image inpainting, horizontal context outpainting via full causal sequence attention, and resolution extrapolation from 256x256 to 512x512 pixels. It also successfully extracts bidirectional semantic representations when fed image sequences in two successive passes, a capability that fixed-order baselines fail to reproduce.

These results show that rigid directional constraints are not necessary for autoregressive visual generation. For technical leaders and product teams, this framework offers a unified path to combine the infrastructure efficiency of plain transformer models—including standard key-value caching—with the editing flexibility previously restricted to specialized masked-image architectures, substantially reducing computational serving costs and system complexity.

Organizations developing generative visual pipelines should consider piloting random-order token strategies to accelerate generation speed and consolidate editing features into single decoder-only models. Immediate engineering next steps include optimizing multi-token scheduling strategies and refining high-frequency detail generation during resolution scaling.

While the core findings are supported by consistent benchmark improvements and controlled ablations, the approach faces limitations in synthesizing intricate high-frequency boundaries during zero-shot resolution expansion. Consequently, readers should exercise caution when deploying resolution extrapolation directly to high-fidelity production assets without dedicated tuning.

  • Paper: Autoregressive Diffusion Models, Emiel Hoogeboom et al. (2022). This work introduces order-agnostic autoregressive and permutation-based generation, establishing the conceptual foundation for training visual models across arbitrary sequence orderings.
  • Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). This paper establishes the discrete visual tokenization framework via VQGAN and subsequent transformer modeling that standard autoregressive visual generators rely upon.
  • Paper: Image Transformer, Niki Parmar et al. (2018). This seminal paper introduces autoregressive modeling of visual pixels with self-attention transformers, providing the baseline sequence-modeling formulation that RandAR generalizes to arbitrary orders.
  • Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). This work demonstrates scaling autoregressive transformers over discrete image tokens, defining the conventional raster-order visual AR paradigm.
  • Paper: Scaling Autoregressive Models for Content-Rich Text-to-Image Generation, Jiahui Yu et al. (2022). This paper explores scaling discrete-token autoregressive image generation models, illustrating the strengths and fixed-order limitations of standard sequence-to-sequence visual models.
Cover for RandAR: Decoder-only Autoregressive Visual Generation in Random Orders

Abstract

We introduce RandAR, a decoder-only visual autoregressive (AR) model capable of generating images in arbitrary token orders. Unlike previous decoder-only AR models that rely on a predefined generation order, RandAR removes this inductive bias, unlocking new capabilities in decoder-only generation. Our essential design enables random order by inserting a “position instruction token” before each image token to be predicted, representing the spatial location of the next image token. Trained on randomly permuted token sequences – a more challenging task than fixed-order generation, RandAR achieves comparable performance to its conventional raster-order counterpart. More importantly, decoder-only transformers trained from random orders acquire new capabilities. For the efficiency of decoder-only AR models, RandAR adopts parallel decoding with KV-Cache at inference time, enjoying 2.5× acceleration without sacrificing generation quality. Additionally, RandAR supports inpainting, outpainting and resolution extrapolation in a zero-shot manner. We hope RandAR inspires new directions for decoder-only visual generation models and broadens their applications across diverse scenarios. Our project

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. RandAR Framework
  • 3.3. RandAR Enables Parallel Decoding
  • 3.4. Zero-shot Applications for RandAR
  • 3.4.1. Inpainting and Class-conditional Image Editing
  • 3.4.2. Outpainting
  • 3.4.3. Resolution Extrapolation
  • 3.4.4. Bi-directional Encoding
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Main Results
  • 4.3. Effects of Parallel Decoding
  • 4.4. Case Studies: Random v.s. Raster Order
  • 4.4.1. Inpainting and Class-conditioned Editing
  • 4.4.2. Outpainting
  • 4.4.3. Resolution Extrapolation
  • 4.4.4. Feature Encoding
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — RandAR Framework for Random-Order Decoder-Only Visual Generation

    model/method

    RandAR is a framework enabling standard GPT-style unidirectional (causal) decoder-only transformers to generate 2D discrete visual tokens in arbitrary generation orders during both training and inference. Unlike standard visual autoregressive models that enforce a fixed raster order (top-left to bottom-right), RandAR eliminates this order bias by inserting a dedicated position instruction token PiP_i immediately before each target visual token xix_i.

    An image is first encoded into a 2D grid of discrete tokens (e.g., 16×16=25616 \times 16 = 256 tokens) using a vector-quantized image tokenizer. For training, the raster-ordered sequence is shuffled according to a random permutation π\pi of length NN, and the final token is omitted to yield [xπ(1),xπ(2),…,xπ(N−1)][x_{\pi(1)}, x_{\pi(2)}, \dots, x_{\pi(N-1)}]. Position instruction tokens corresponding to each spatial coordinate are interleaved with the visual tokens to construct the sequence: [Pπ(1),xπ(1),Pπ(2),xπ(2),…,Pπ(N−1),xπ(N−1),Pπ(N)][P_{\pi(1)}, x_{\pi(1)}, P_{\pi(2)}, x_{\pi(2)}, \dots, P_{\pi(N-1)}, x_{\pi(N-1)}, P_{\pi(N)}]

    This sequence is processed using causal self-attention, where the model is supervised via cross-entropy loss to predict each ground-truth visual token xπ(n)x_{\pi(n)} directly from the preceding position instruction token Pπ(n)P_{\pi(n)} and all prior context. Except for one shared trainable embedding parameter for the position instruction tokens, the transformer backbone follows standard LLaMA designs (RMSNorm, SwiGLU activations, and 2D Rotary Position Embeddings).

  2. Knowl 2 — Autoregressive Image Token Distribution Factorization under Random Permutations

    equation

    Under the RandAR formulation, the joint probability distribution of an image token sequence x=[x1,x2,…,xN]\mathbf{x} = [x_1, x_2, \dots, x_N] given their spatial position instruction sequence P=[P1,P2,…,PN]\mathbf{P} = [P_1, P_2, \dots, P_N] permuted by an arbitrary ordering permutation π\pi is factorized as:

    pθ(x∣P)=∏n=1Npθ(xπ(n)∣Pπ(1),xπ(1),Pπ(2),xπ(2),…,Pπ(n−1),xπ(n−1),Pπ(n))p_\theta(\mathbf{x} \mid \mathbf{P}) = \prod_{n=1}^N p_\theta(x_{\pi(n)} \mid P_{\pi(1)}, x_{\pi(1)}, P_{\pi(2)}, x_{\pi(2)}, \dots, P_{\pi(n-1)}, x_{\pi(n-1)}, P_{\pi(n)})

    where:

    • NN is the total number of visual tokens in the tokenized image grid (e.g., N=256N = 256 for a 16×1616 \times 16 latent representation).
    • π\pi denotes a permutation over the token index set {1,…,N}\{1, \dots, N\}, with π(n)\pi(n) representing the raster index of the nn-th token in the generation order.
    • xπ(n)∈{1,…,V}x_{\pi(n)} \in \{1, \dots, V\} is the discrete visual token at original spatial position π(n)\pi(n), sampled from a codebook vocabulary of size VV.
    • Pπ(n)∈RdP_{\pi(n)} \in \mathbb{R}^d is the position instruction token representing the spatial coordinate (hπ(n),wπ(n))(h_{\pi(n)}, w_{\pi(n)}) of token xπ(n)x_{\pi(n)}.
    • θ\theta represents the trainable parameters of the causal decoder-only transformer.
  3. Knowl 3 — Position Instruction Token Formulation via 2D Rotary Position Embeddings

    model/method

    To condition each next-token prediction step on the spatial coordinate (hi,wi)(h_i, w_i) of the target token without modifying the core transformer architecture, RandAR defines each position instruction token PiP_i using a single shared trainable embedding vector e∈Rde \in \mathbb{R}^d modulated by 2D Rotary Position Embeddings (2D-RoPE):

    Pi=RoPE(e,hi,wi)P_i = \text{RoPE}(e, h_i, w_i)

    where:

    • e∈Rde \in \mathbb{R}^d is a single learnable vector shared across all spatial coordinates in the image.
    • (hi,wi)(h_i, w_i) are the 2D grid coordinates (row and column) of the target image patch.
    • RoPE(⋅,hi,wi)\text{RoPE}(\cdot, h_i, w_i) applies the 2D rotary embedding transformation to ee based on (hi,wi)(h_i, w_i).

    While RoPE is conventionally applied inside self-attention layers to encode relative distances between query and key tokens, applying it directly to a base vector ee serves as an effective global spatial coordinate representation, adding only dd trainable parameters to the overall model.

  4. Knowl 4 — KV-Cache Compatible Parallel Decoding for Decoder-Only Autoregressive Visual Models

    algorithm

    Because RandAR is trained to predict tokens at arbitrary locations conditioned on previously generated tokens, it can decode multiple tokens simultaneously in a single forward pass without fine-tuning while retaining standard key-value (KV) cache compatibility.

    Input: Model parameters θ\theta, total tokens NN, decoding step schedule {Kt}t=1T\{K_t\}_{t=1}^T with ∑t=1TKt=N\sum_{t=1}^T K_t = N, sampling order permutation π\pi
    Output: Generated visual token sequence x\mathbf{x}
    Initialize sequence S=[]S = []
    Initialize KV-cache
    Set cursor c=0c = 0
    for step t=1t = 1 to TT do
        k=Ktk = K_t
        Construct instruction block Binst=[Pπ(c+1),Pπ(c+2),…,Pπ(c+k)]B_{\text{inst}} = [P_{\pi(c+1)}, P_{\pi(c+2)}, \dots, P_{\pi(c+k)}]
        Append BinstB_{\text{inst}} to SS
        Forward BinstB_{\text{inst}} through model θ\theta using KV-cache
        Sample visual tokens [xπ(c+1),…,xπ(c+k)][x_{\pi(c+1)}, \dots, x_{\pi(c+k)}] from the output logits corresponding to each instruction token in BinstB_{\text{inst}}
        Rearrange the sequence tail from [Pπ(c+1),…,Pπ(c+k)][P_{\pi(c+1)}, \dots, P_{\pi(c+k)}] to interleaved form [Pπ(c+1),xπ(c+1),…,Pπ(c+k),xπ(c+k)][P_{\pi(c+1)}, x_{\pi(c+1)}, \dots, P_{\pi(c+k)}, x_{\pi(c+k)}]
        Update c=c+kc = c + k
    end for
    return Reordered token grid x\mathbf{x}

    The decoding schedule KtK_t follows a cosine decay schedule for step sizes. After sampling kk tokens in parallel, the newly added positions and tokens are rearranged into the interleaved pair format [Pj,xj][P_j, x_j] prior to subsequent steps to match the training format.

  5. Knowl 5 — Zero-Shot Hierarchical Resolution Extrapolation and Spatial Contextual Guidance

    model/method

    RandAR enables zero-shot resolution extrapolation (e.g., generating 512×512512 \times 512 images using a model trained strictly on 256×256256 \times 256 images) without fine-tuning by exploiting arbitrary generation orders in a two-stage hierarchical schedule:

    1. Global Layout Generation: Tokens situated at even grid coordinates are sampled first to establish overall image structure, using interpolated 2D-RoPE embeddings across the enlarged spatial grid.
    2. High-Frequency Detail Infilling: The remaining tokens at odd and mixed coordinates are sampled conditioned on the layout tokens. For this stage, the interpolated RoPE is substituted with the top-kk high-frequency components of extrapolated RoPE (adapted from NTK-RoPE context extension) to capture fine textures.

    To further sharpen high-frequency details, Spatial Contextual Guidance (SCG) is applied at inference: alongside the main generation sequence, an auxiliary sequence is maintained where each newly sampled token is randomly dropped with a probability of 25%25\%. The output logits from the full and dropped sequences are linearly combined in an approach analogous to classifier-free guidance.

  6. Knowl 6 — Zero-Shot Bi-Directional Feature Extraction via Two-Pass Forwarding

    model/method

    Although decoder-only causal transformers prevent early tokens from attending to subsequent tokens in a single forward pass, RandAR enables bi-directional representation extraction without architectural changes by passing the interleaved token sequence twice:

    pθ(P1,x1,…,PN,xN⏟Round 1 (Causal Warmup),P1,x1,…,PN,xN⏟Round 2 (Bi-Directional Aggregation))p_\theta\left(\underbrace{P_1, x_1, \dots, P_N, x_N}_{\text{Round 1 (Causal Warmup)}}, \underbrace{P_1, x_1, \dots, P_N, x_N}_{\text{Round 2 (Bi-Directional Aggregation)}}\right)

    The hidden states at the output of the position instruction tokens PiP_i during Round 2 are extracted as the final feature representations. In Round 2, causal attention allows every token to attend to the entire image context from Round 1. While raster-order models degrade when processing doubled sequence lengths out of domain, RandAR generalizes to this two-pass sequence format zero-shot due to its training on randomized permutation orders.

  7. Knowl 7 — Zero-Shot Inpainting and Full-Sequence Attention Outpainting

    model/method

    RandAR supports zero-shot image editing and boundary extension without task-specific training:

    • Inpainting and Class-Conditional Editing: Known context tokens and their corresponding position instruction tokens are placed at the beginning of the input sequence. The position instruction tokens corresponding to masked/editing areas are appended, and the model autoregressively samples the missing visual tokens using the full unmasked context from all spatial directions.
    • Outpainting via Full Sequence Attention: To extend an image beyond its boundaries (e.g., from 256×256256 \times 256 to 256×1024256 \times 1024, a 3×3\times horizontal extension), RoPE frequencies are extrapolated to the target canvas dimension. Unlike raster-order models that rely on sliding-window attention and discard non-local context, RandAR places all original image tokens first and uses full causal attention across the entire expanded sequence length, producing coherent long-range structures.
  8. Knowl 8 — Class-Conditional Generation Benchmarks on ImageNet 256x256

    data/table

    RandAR achieves image generation quality comparable to standard raster-order autoregressive baselines on class-conditional ImageNet 256×256256 \times 256 (50k50\text{k} samples), despite learning over the 256!≈8×10506256! \approx 8 \times 10^{506} permutation space.

    Model #Params FID ↓\downarrow IS ↑\uparrow Precision ↑\uparrow Recall ↑\uparrow Steps
    Raster-order Counterpart 343M 2.20 274.26 0.80 0.59 256
    Raster-order Counterpart 775M 2.16 282.71 0.80 0.61 256
    RandAR-L 343M 2.55 288.82 0.81 0.58 88
    RandAR-XL 775M 2.25 317.77 0.80 0.60 88
    RandAR-XL 775M 2.22 314.21 0.80 0.60 256
    RandAR-XXL 1.4B 2.15 321.97 0.79 0.62 88

    All models use the LLaMAGen VQGAN tokenizer (downsampling factor 16×16\times, vocabulary size 16,384) and were trained for 300 epochs (360k iterations) with batch size 1024 using AdamW without exponential moving average.

  9. Knowl 9 — Latency and Speedup of Parallel Decoding with KV-Cache on A100 GPU

    data/table

    Parallel decoding in RandAR reduces decoding steps and lowers inference latency while maintaining KV-cache compatibility. Generating a batch of 64 images (256×256256 \times 256 resolution, batch size 128 effective with classifier-free guidance) on a single NVIDIA A100 GPU (40GB VRAM) achieves up to 2.5×2.5\times speedup over full sequential decoding:

    Method Latency (s) #Params #Steps Parallel Decoding KV-Cache
    RandAR 16.8 1.4B 256 Supported Supported
    RandAR 6.6 1.4B 88 Supported Supported
    RandAR 4.6 1.4B 48 Supported Supported
    LLaMAGen 15.9 1.4B 256 Noncompatible Supported
    MAR 53.3 943M 64 Supported Noncompatible
    MAR 220.0 943M 256 Supported Noncompatible

    Decreasing decoding steps from 256 to 88 reduces generation time from 16.8s to 6.6s (2.5×2.5\times acceleration) with negligible change in generation fidelity.

  10. Knowl 10 — Representation Learning Evaluation on SPair-71k and ImageNet Linear Probing

    data/table

    Evaluating representations extracted by XL-sized (775M) models on zero-shot semantic correspondence (SPair-71k, measuring Percentage of Correct Keypoints, PCK) and ImageNet linear probing (classification accuracy) demonstrates that RandAR effectively aggregates bi-directional context in a second forward pass, whereas raster-order models fail to generalize:

    Model SPair-71k PCK (Image) ↑\uparrow SPair-71k PCK (Point) ↑\uparrow Top-1 Acc ↑\uparrow Top-5 Acc ↑\uparrow
    RasterAR 24.5 28.6 62.6 83.9
    RasterAR (w/ 2nd Round) 3.6 3.9 58.3 80.7
    RandAR 22.1 25.8 57.3 80.3
    RandAR (w/ 2nd Round) 31.3 36.4 63.1 84.2

    In Round 2, RandAR's semantic keypoint matching accuracy jumps from 22.1%22.1\% to 31.3%31.3\% (per image) and Top-1 classification improves from 57.3%57.3\% to 63.1%63.1\%, outperforming the raster baseline.

  11. Knowl 11 — Ablation of Position Instruction Token Architecture Choices

    data/table

    Comparison of different position instruction token designs evaluated on an XL-sized (775M) RandAR model trained for 100k iterations on ImageNet 256×256256 \times 256:

    Design Choice FID ↓\downarrow Inception Score ↑\uparrow #Steps
    RandAR (Shared embedding + 2D RoPE) 2.82 293.6 88
    Dense (256 unique learnable embeddings) 3.07 290.6 88
    Dense Merge (Added directly to prior visual token) 3.37 307.6 256

    The default design (a single shared learnable embedding rotated by 2D-RoPE) outperforms using 256 separate dense spatial embeddings. Merging the position instruction directly into preceding visual tokens degrades performance under parallel decoding and requires sequential 256-step generation to function.

  12. Knowl 12 — Limitations in High-Frequency Detail Generation for Zero-Shot Resolution Extrapolation

    limitation

    When performing zero-shot resolution extrapolation to synthesize higher-resolution images (such as 512×512512 \times 512 from a 256×256256 \times 256 training setup), the model struggles to accurately generate intricate, high-frequency structures and sharp object boundaries that were unobserved in the lower-resolution training distribution (e.g., complex mechanical structures like space shuttle surfaces). Additionally, the method currently requires manual construction of coordinate schedules and has not been extended to arbitrary non-integer resolution scaling factors.

Coverage note — None was omitted; all key contributions, mathematical equations, zero-shot application mechanisms (parallel decoding, inpainting, outpainting, resolution extrapolation, bi-directional feature extraction), ablation studies, benchmark data tables, and limitations are covered.

References

  1. 1.Andrew Brock. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018.
  2. 2.Tim Brooks, Aleksander Holynski, and Alexei A Efros. InstructPix2Pix: Learning to follow image editing instructions. In CVPR, 2023.
  3. 3.Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  4. 4.Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao. Medusa: Simple llm inference acceleration framework with multiple decoding heads. In ICML, 2024.
  5. 5.Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman. Maskgit: Masked generative image transformer. In CVPR, 2022.
  6. 6.Huiwen Chang, Han Zhang, Jarred Barber, AJ Maschinot, Jose Lezama, Lu Jiang, Ming-Hsuan Yang, Kevin Murphy, William T Freeman, Michael Rubinstein, et al. Muse: Text-to-image generation via masked generative transformers. In ICML, 2023.
  7. 7.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  8. 8.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. In NeurIPS, 2021.
  9. 9.Zheng Ding, Mengqi Zhang, Jiajun Wu, and Zhuowen Tu. Patched denoising diffusion models for high-resolution image synthesis. In ICLR, 2023.
  10. 10.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In CVPR, 2021.
  11. 11.Lijie Fan, Tianhong Li, Siyang Qin, Yuanzhen Li, Chen Sun, Michael Rubinstein, Deqing Sun, Kaiming He, and Yonglong Tian. Fluid: Scaling autoregressive text-to-image generative models with continuous tokens. arXiv preprint arXiv:2410.13863, 2024.
  12. 12.Yao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue, Hannaneh Hajishirzi, Yoon Kim, and Hao Peng. Data engineering for scaling language models to 128k context. In ICML, 2024.
  13. 13.Shanghua Gao, Pan Zhou, Ming-Ming Cheng, and Shuicheng Yan. Masked diffusion transformer is a strong image synthesizer. In ICCV, 2023.
  14. 14.Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. Olmo: Accelerating the science of language models. arXiv preprint arXiv:2402.00838, 2024.
  15. 15.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, 2022.
  16. 16.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, 2017.
  17. 17.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  18. 18.Minguk Kang, Jun-Yan Zhu, Richard Zhang, Jaesik Park, Eli Shechtman, Sylvain Paris, and Taesung Park. Scaling up gans for text-to-image synthesis. In CVPR, 2023.
  19. 19.Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019.
  20. 20.Sehoon Kim, Coleman Richard Charles Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer. SqueezeLLM: Dense-and-sparse quantization. In ICML, 2024.
  21. 21.Diederik P Kingma. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  22. 22.Tuomas Kynkäanniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Improved precision and recall metric for assessing generative models. In NeurIPS, 2019.
  23. 23.Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. Autoregressive image generation using residual quantization. In CVPR, 2022.
  24. 24.Tianhong Li, Huiwen Chang, Shlok Mishra, Han Zhang, Dina Katabi, and Dilip Krishnan. Mage: Masked generative encoder to unify representation learning and image synthesis. In CVPR, 2023.
  25. 25.Tianhong Li, Yonglong Tian, He Li, Mingyang Deng, and Kaiming He. Autoregressive image generation without vector quantization. In NeurIPS, 2024.
  26. 26.Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum. Compositional visual generation with composable diffusion models. In ECCV, 2022.
  27. 27.Wenze Liu, Le Zhuo, Yi Xin, Sheng Xia, Peng Gao, and Xiangyu Yue. Customize your visual autoregressive recipe with set autoregressive modeling. arXiv preprint arXiv:2410.10511, 2024.
  28. 28.Yinhan Liu. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 364, 2019.
  29. 29.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019.
  30. 30.Zhuoyan Luo, Fengyuan Shi, Yixiao Ge, Yujiu Yang, Limin Wang, and Ying Shan. Open-magvit2: An open-source project toward democratizing auto-regressive visual generation. arXiv preprint arXiv:2409.04410, 2024.
  31. 31.Nanye Ma, Mark Goldstein, Michael S Albergo, Nicholas M Boffi, Eric Vanden-Eijnden, and Saining Xie. SIT: Exploring flow and diffusion-based generative models with scalable interpolant transformers. arXiv preprint arXiv:2401.08740, 2024.
  32. 32.Juhong Min, Jongmin Lee, Jean Ponce, and Minsu Cho. Spair-71k: A large-scale benchmark for semantic correspondence. arXiv preprint arXiv:1908.10543, 2019.
  33. 33.Ziqi Pang, Ziyang Xie, Yunze Man, and Yu-Xiong Wang. Frozen transformers in language models are effective visual encoder layers. In ICLR, 2024.
  34. 34.William Peebles and Saining Xie. Scalable diffusion models with transformers. In ICCV, 2023.
  35. 35.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding with unsupervised learning. Technical report, OpenAI, 2018.
  36. 36.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020.
  37. 37.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, 2021.
  38. 38.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022.
  39. 39.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In NeurIPS, 2016.
  40. 40.Axel Sauer, Katja Schwarz, and Andreas Geiger. Stylegan-xl: Scaling stylegan to large diverse datasets. In SIGGRAPH, 2022.
  41. 41.Noam Shazeer. Fast transformer decoding: One write-head is all you need. arXiv preprint arXiv:1911.02150, 2019.
  42. 42.Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  43. 43.Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  44. 44.Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan. Autoregressive model beats diffusion: Llama for scalable image generation. arXiv preprint arXiv:2406.06525, 2024.
  45. 45.Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan. Emergent correspondence from image diffusion. In NeurIPS, 2023.
  46. 46.Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024.
  47. 47.Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang. Visual autoregressive modeling: Scalable image generation via next-scale prediction. In NeurIPS, 2024.
  48. 48.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  49. 49.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  50. 50.Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, and Alex Graves. Conditional image generation with pixelcnn decoders. In NeurIPS, 2016.
  51. 51.Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, Yingli Zhao, Yulong Ao, Xuebin Min, Tao Li, Boya Wu, Bo Zhao, Bowen Zhang, Liangdong Wang, Guang Liu, Zheqi He, Xi Yang, Jingjing Liu, Yonghua Lin, Tiejun Huang, and Zhongyuan Wang. Emu3: Next-token prediction is all you need. arXiv preprint arXiv:2409.18869, 2024.
  52. 52.Yuqing Wang, Shuhuai Ren, Zhijie Lin, Yujin Han, Haoyuan Guo, Zhenheng Yang, Difan Zou, Jiashi Feng, and Xihui Liu. Parallelized autoregressive visual generation. arXiv preprint arXiv:2412.15119, 2024.
  53. 53.Mark Weber, Lijun Yu, Qihang Yu, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen. Maskbit: Embedding-free image generation via bit tokens. arXiv preprint arXiv:2409.16211, 2024.
  54. 54.Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, and Mike Zheng Shou. Show-o: One single transformer to unify multimodal understanding and generation. arXiv preprint arXiv:2408.12528, 2024.
  55. 55.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. Xlnet: Generalized autoregressive pretraining for language understanding. In NeurIPS, 2019.
  56. 56.Lijun Yu, Yong Cheng, Kihyuk Sohn, Jose Lezama, Han Zhang, Huiwen Chang, Alexander G. Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang. Magvit: Masked generative video transformer. In CVPR, 2023.
  57. 57.Lijun Yu, Jose Lezama, Nitesh B. Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G. Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A. Ross, and Lu Jiang. Language model beats diffusion–tokenizer is key to visual generation. In ICLR, 2024.
  58. 58.Qihang Yu, Ju He, Xueqing Deng, Xiaohui Shen, and Liang-Chieh Chen. Randomized autoregressive visual generation. arXiv preprint arXiv:2411.00776, 2024.
  59. 59.Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen. An image is worth 32 tokens for reconstruction and generation. arXiv preprint arXiv:2406.07550, 2024.
  60. 60.Biao Zhang and Rico Sennrich. Root mean square layer normalization. In NeurIPS, 2019.
  61. 61.Qinsheng Zhang, Jiaming Song, Xun Huang, Yongxin Chen, and Ming-Yu Liu. Diffcollage: Parallel generation of large content with diffusion models. In CVPR, 2023.
  62. 62.Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy. Transfusion: Predict the next token and diffuse images with one multi-modal model. arXiv preprint arXiv:2408.11039, 2024.
  63. 63.Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee. Segment everything everywhere all at once. In NeurIPS, 2024.

Citation

MLA
Pang, Z., et al. “RandAR: Decoder-only Autoregressive Visual Generation in Random Orders”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 45–55, https://doi.org/10.1109/CVPR52734.2025.00014.
APA
Pang, Z., Zhang, T., Luan, F., Man, Y., Tan, H., Zhang, K., Freeman, W. T., & Wang, Y.-X. (2025). RandAR: Decoder-only Autoregressive Visual Generation in Random Orders. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 45–55. https://doi.org/10.1109/CVPR52734.2025.00014
Chicago
Pang, Z., T. Zhang, F. Luan, et al. 2025. “RandAR: Decoder-only Autoregressive Visual Generation in Random Orders”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 45–55. https://doi.org/10.1109/CVPR52734.2025.00014.
Harvard
Pang, Z. et al. (2025) “RandAR: Decoder-only Autoregressive Visual Generation in Random Orders”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 45–55. Available at: https://doi.org/10.1109/CVPR52734.2025.00014.
Vancouver
1. Pang Z, Zhang T, Luan F, Man Y, Tan H, Zhang K, Freeman WT, Wang Y-X (2025) RandAR: Decoder-only Autoregressive Visual Generation in Random Orders. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 45–55

BibTeX

@inproceedings{Pang_2025, title={RandAR: Decoder-only Autoregressive Visual Generation in Random Orders}, url={http://dx.doi.org/10.1109/CVPR52734.2025.00014}, DOI={10.1109/cvpr52734.2025.00014}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Pang, Ziqi and Zhang, Tianyuan and Luan, Fujun and Man, Yunze and Tan, Hao and Zhang, Kai and Freeman, William T. and Wang, Yu-Xiong}, year={2025}, month=June, pages={45–55} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE