Muse: Text-To-Image Generation via Masked Generative Transformers

Huiwen ChangHan ZhangJarred BarberAaron MaschinotJosé LezamaLu JiangMing-Hsuan YangKevin Patrick MurphyWilliam T. FreemanMichael Rubinstein

article2023ICML746 citations

Presents Muse, a text-to-image Transformer based on masked discrete token modeling that achieves state-of-the-art generation fidelity significantly faster than diffusion or autoregressive baselines while enabling zero-shot editing without model fine-tuning.

Listen

Recent breakthroughs in text-to-image synthesis have enabled high-quality visual generation from natural language descriptions. However, existing leading systems—primarily pixel-space diffusion models and autoregressive transformers—suffer from extreme computational overhead and slow generation speeds during practical deployment. Pixel diffusion models require numerous sequential denoising passes, while autoregressive approaches generate visual tokens sequentially one by one. This latency creates significant operational bottlenecks for real-time services and scalable interactive products.

The article demonstrates Muse, a masked generative transformer designed to achieve state-of-the-art image quality and alignment while operating significantly faster than diffusion and autoregressive baselines. The study evaluates the architecture’s generation fidelity, alignment accuracy, processing speed, and zero-shot editing capabilities across established benchmark datasets.

To overcome inference bottlenecks, the approach reformulates image generation into a discrete masked token modeling framework. The system combines frozen representations from a pre-trained large language model (T5-XXL) with discrete visual tokens produced by a convolutional image tokenizer. Image synthesis is carried out through a two-stage cascade: a base transformer generates low-resolution visual tokens (256x256 pixel equivalent) by predicting randomly masked elements in parallel, and a super-resolution transformer translates them into high-resolution tokens (512x512 pixel equivalent). The models were trained on 460 million image-text pairs using accelerated hardware.

The findings confirm substantial gains in generation speed and output quality. First, the 3-billion-parameter model generates 512x512 images in approximately 1.3 seconds, achieving more than a 10-fold speedup over comparable diffusion and autoregressive systems like Imagen-3B and Parti-3B, and roughly 3 times the speed of Stable Diffusion v1.4. Second, the architecture reaches superior alignment and image fidelity, obtaining a 6.06 Fréchet Inception Distance score on the CC3M benchmark and an 7.88 zero-shot score on MS-COCO with a high CLIP alignment score of 0.32. Third, in human evaluation studies involving 1,650 diverse prompts, raters preferred the model's text-image alignment over Stable Diffusion v1.4 by a ratio of 2.7 to 1 (70.6% preference versus 25.4%). Finally, the masked token formulation directly supports zero-shot image editing applications—such as inpainting, outpainting, and mask-free modification—without requiring model fine-tuning or inversion.

These results establish that non-diffusion masked transformers offer a superior efficiency-to-quality profile for production environments. Faster inference dramatically lowers compute infrastructure costs, reduces per-query latency, and expands the viability of real-time image generation and interactive design tooling without sacrificing fine-grained language comprehension.

Given the demonstrated latency advantages, organizations developing generative visual systems should explore masked discrete token architectures as cost-effective alternatives to standard diffusion pipelines. However, due to societal risks surrounding misinformation, cultural stereotyping, and lack of consent in web-scale training data, the authors opt not to release the model code or public demos immediately and advise against deploying such systems for generating human faces without rigorous safeguards.

Decision-makers should interpret the findings alongside specific architectural limitations. The model struggles with direct rendering of long, multi-word phrases and frequently miscounts scenes involving high or multiple object cardinalities. While the quantitative benchmark results provide high confidence in the model's efficiency and alignment gains, broader deployment will require further research into dataset bias mitigation and improved multi-object compositional reasoning.

arXiv: 2301.00704
Cover for Muse: Text-To-Image Generation via Masked Generative Transformers

Abstract

We present Muse, a text-to-image Transformer model that achieves state-of-the-art image generation performance while being significantly more efficient than diffusion or autoregressive models. Muse is trained on a masked modeling task in discrete token space: given the text embedding extracted from a pre-trained large language model (LLM), Muse is trained to predict randomly masked image tokens. Compared to pixel-space diffusion models, such as Imagen and DALL-E 2, Muse is significantly more efficient due to the use of discrete tokens and requiring fewer sampling iterations; compared to autoregressive models, such as Parti, Muse is more efficient due to the use of parallel decoding. The use of a pre-trained LLM enables fine-grained language understanding, translating to high-fidelity image generation and the understanding of visual concepts such as objects, their spatial relationships, pose, cardinality etc. Our 900M parameter model achieves a new SOTA on CC3M, with an FID score of 6.06. The Muse 3B parameter model achieves an FID of 7.88 on zero-shot COCO evaluation, along with a CLIP score of 0.32. Muse also directly enables a number of image editing applications without the need to fine-tune or invert the model: inpainting, outpainting, and mask-free editing. More results are available at this https URL

Table of Contents

  • 1 Introduction
  • 2 Model
  • 2.1 Pre-trained Text Encoders
  • 2.2 Semantic Tokenization using VQGAN
  • 2.3 Base Model
  • 2.4 Super-Resolution Model
  • 2.5 Decoder Finetuning
  • 2.6 Variable Masking Rate
  • 2.7 Classifier Free Guidance
  • 2.8 Iterative Parallel Decoding at Inference
  • 3 Results
  • 3.1 Qualitative Performance
  • 3.2 Quantitative Performance
  • 3.2.1 Human evaluation
  • 3.2.2 Inference Speed
  • 3.3 Image Editing
  • 3.3.1 Text-guided Inpainting / outpainting
  • 3.3.2 Zero-shot Mask-free editing
  • 4 Related Work
  • 4.1 Image Generation Models
  • 4.2 Image Tokenizers
  • 4.3 Large Language Models
  • 4.4 Text-Image Models
  • 4.5 Image Editing with Generative Models
  • 5 Discussion and Social Impact
  • References
  • A Appendix.
  • A.1 Base Model Configurations
  • A.2 VQGAN Configurations
  • A.3 Super Resolution Configurations

Citation

MLA
Chang, H., et al. “Muse: Text-To-Image Generation via Masked Generative Transformers”. arXiv, 2023, http://arxiv.org/abs/2301.00704v1.
APA
Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M.-H., Murphy, K., Freeman, W. T., Rubinstein, M., Li, Y., & Krishnan, D. (2023). Muse: Text-To-Image Generation via Masked Generative Transformers. arXiv. http://arxiv.org/abs/2301.00704v1
Chicago
Chang, H., H. Zhang, J. Barber, et al. 2023. “Muse: Text-To-Image Generation via Masked Generative Transformers”. arXiv. http://arxiv.org/abs/2301.00704v1.
Harvard
Chang, H. et al. (2023) “Muse: Text-To-Image Generation via Masked Generative Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.00704v1.
Vancouver
1. Chang H, Zhang H, Barber J, et al (2023) Muse: Text-To-Image Generation via Masked Generative Transformers. arXiv

BibTeX

@article{chang2023muse,
  title = {Muse: Text-To-Image Generation via Masked Generative Transformers},
  author = {Chang, Huiwen and Zhang, Han and Barber, Jarred and Maschinot, AJ and Lezama, Jose and Jiang, Lu and Yang, Ming-Hsuan and Murphy, Kevin and Freeman, William T. and Rubinstein, Michael and Li, Yuanzhen and Krishnan, Dilip},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.00704v1},
  eprint = {2301.00704}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/