RandAR: Decoder-only Autoregressive Visual Generation in Random Orders
Ziqi PangTianyuan ZhangFujun LuanYunze ManHao TanKai ZhangWilliam T. FreemanYu-Xiong Wang
Presents a decoder-only visual autoregressive framework that uses position instruction tokens to generate images in arbitrary orders, achieving 2.5x faster parallel decoding alongside zero-shot inpainting, outpainting, and resolution extrapolation without sacrificing visual quality.
Visual generation models frequently adapt large language model architectures by predicting visual tokens sequentially in a fixed, row-by-row raster order. However, this rigid sequencing forces artificial directional constraints on two-dimensional images, restricting the model from using surrounding visual context and limiting flexibility in downstream editing tasks. Meanwhile, alternative architectures that allow flexible masking often lose compatibility with standard caching optimizations, which hurts operational speed.
The article demonstrates that causal, decoder-only transformers can generate high-quality images in completely arbitrary token sequences without altering their standard architectural foundations. By introducing a framework named RandAR, the evaluation examines whether removing the predefined raster order allows unidirectional generative models to acquire bidirectional contextual reasoning, accelerate generation, and perform zero-shot visual editing.
To achieve random-order generation with minimal architecture changes, the approach inserts a single learnable spatial indicator—termed a position instruction token—prior to each image token to specify its target coordinate. The model was trained across randomly shuffled sequence permutations on standard image datasets such as ImageNet, using established evaluation metrics including Fréchet Inception Distance, Inception Score, and precision-recall measures alongside latency testing on production-grade graphical processing units.
The findings establish that training on random orderings achieves visual fidelity comparable to conventional raster-based models despite the combinatorial complexity of learning across factorial permutations. By decoupling generation from fixed order, the system enables parallel multi-token prediction during inference, delivering an approximate 2.5-fold reduction in generation latency without degrading visual quality. Furthermore, the model demonstrates zero-shot capabilities in image inpainting, horizontal context outpainting via full causal sequence attention, and resolution extrapolation from 256x256 to 512x512 pixels. It also successfully extracts bidirectional semantic representations when fed image sequences in two successive passes, a capability that fixed-order baselines fail to reproduce.
These results show that rigid directional constraints are not necessary for autoregressive visual generation. For technical leaders and product teams, this framework offers a unified path to combine the infrastructure efficiency of plain transformer models—including standard key-value caching—with the editing flexibility previously restricted to specialized masked-image architectures, substantially reducing computational serving costs and system complexity.
Organizations developing generative visual pipelines should consider piloting random-order token strategies to accelerate generation speed and consolidate editing features into single decoder-only models. Immediate engineering next steps include optimizing multi-token scheduling strategies and refining high-frequency detail generation during resolution scaling.
While the core findings are supported by consistent benchmark improvements and controlled ablations, the approach faces limitations in synthesizing intricate high-frequency boundaries during zero-shot resolution expansion. Consequently, readers should exercise caution when deploying resolution extrapolation directly to high-fidelity production assets without dedicated tuning.
- Paper: Autoregressive Diffusion Models, Emiel Hoogeboom et al. (2022). This work introduces order-agnostic autoregressive and permutation-based generation, establishing the conceptual foundation for training visual models across arbitrary sequence orderings.
- Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). This paper establishes the discrete visual tokenization framework via VQGAN and subsequent transformer modeling that standard autoregressive visual generators rely upon.
- Paper: Image Transformer, Niki Parmar et al. (2018). This seminal paper introduces autoregressive modeling of visual pixels with self-attention transformers, providing the baseline sequence-modeling formulation that RandAR generalizes to arbitrary orders.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). This work demonstrates scaling autoregressive transformers over discrete image tokens, defining the conventional raster-order visual AR paradigm.
- Paper: Scaling Autoregressive Models for Content-Rich Text-to-Image Generation, Jiahui Yu et al. (2022). This paper explores scaling discrete-token autoregressive image generation models, illustrating the strengths and fixed-order limitations of standard sequence-to-sequence visual models.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). Show-o extends unified visual transformer generation by exploring joint architectures across autoregressive and discrete diffusion formulations for visual understanding and synthesis.
