Regularized Vector Quantization for Tokenized Image Synthesis
Jiahui ZhangFangneng ZhanChristian TheobaltShijian Lu
Proposes a dual-regularized vector quantization framework with a probabilistic contrastive loss that prevents codebook collapse and aligns training with stochastic sampling for superior image synthesis in autoregressive and diffusion models.
Discrete representation learning, which converts continuous images into compact sequences of discrete tokens, is a foundational technique for modern generative artificial intelligence across autoregressive and diffusion architectures. However, current vector quantization approaches face fundamental trade-offs. Deterministic methods, which select the single best-matching token, suffer from severe codebook collapse where many codebook entries remain unlearned or invalid, while also creating a mismatch with generative inference where tokens are sampled probabilistically. Conversely, purely stochastic methods, which sample tokens from predicted distributions, lead to underutilized codebook capacity and perturbed reconstruction targets that degrade overall image quality.
The article introduces and evaluates a regularized vector quantization framework designed to overcome these limitations. The approach introduces a prior distribution regularization that encourages uniform token usage across the codebook via divergence minimization, preventing codebook collapse and maximizing representational capacity. In addition, it employs a stochastic mask regularization that applies probabilistic sampling to a subset of spatial regions while keeping the rest deterministic, paired with an adaptive probabilistic contrastive loss to enable flexible, artifact-free image reconstruction. The framework was evaluated across standard benchmarks including ADE20K, CelebA-HQ, CUB-200, and MS-COCO on semantic and text-to-image synthesis tasks.
The empirical findings demonstrate that the regularized quantization framework consistently outperforms prevailing deterministic and stochastic baselines in both image reconstruction fidelity and downstream generation quality. Specifically, the method achieved superior generation scores across all benchmarks, reducing the generation Fréchet Inception Distance on ADE20K to 34.47 compared to 38.53 for deterministic VQ-GAN and 37.51 for Gumbel-VQ, while achieving full codebook utilization. Experiments also established that a 40% stochastic masking ratio provides the optimal balance between inference alignment and reconstruction stability, and confirmed that performance scales effectively with larger codebook sizes up to 8,192 entries where conventional models typically stagnate.
These results indicate that generative image systems can achieve higher fidelity and better scaling efficiency without incurring major structural redesigns of underlying generative models. Organizations developing generative vision models should adopt dual-strategy quantization combining prior regularization and hybrid deterministic-stochastic masking to maximize codebook efficiency. While the article provides high confidence across autoregressive and discrete diffusion models on standard 256x256 benchmark resolutions, further testing in larger-scale commercial production pipelines and higher native image resolutions is recommended before enterprise-wide deployment.
- Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). This foundational work establishes the VQGAN framework of discrete tokenization and autoregressive modeling that the source regularizes to eliminate codebook collapse and improve generation.
- Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). This paper introduces vector-quantized representation learning (VQ-VAE), providing the foundational discrete codebook and straight-through estimator mechanisms analyzed and improved in the source.
- Paper: Generating Diverse High-Fidelity Images with VQ-VAE-2, Ali Razavi et al. (2019). This work demonstrates scaling discrete multi-scale vector quantization for high-fidelity image generation, establishing key architectural baselines for downstream tokenized synthesis.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). This work highlights the two-stage visual tokenization paradigm for text-to-image autoregressive generation, demonstrating the importance of codebook efficiency in downstream generative pipelines.
- Paper: Scaling Autoregressive Models for Content-Rich Text-to-Image Generation, Jiahui Yu et al. (2022). This study analyzes the scaling behavior of autoregressive visual token modeling, motivating the source's focus on maximizing codebook capacity without codebook collapse.
- Paper: Masked Autoencoders Are Effective Tokenizers for Diffusion Models, Hao Chen et al. (2025). This work extends tokenization research by demonstrating how masked autoencoding creates structured latent spaces specifically optimized as tokenizers for generative diffusion models.
- Paper: RandAR: Decoder-only Autoregressive Visual Generation in Random Orders, Ziqi Pang et al. (2025). This paper generalizes autoregressive visual token generation by moving beyond fixed raster order to arbitrary token sequences, directly benefiting from robust and collapse-free discrete token representations.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). This work advances downstream generative modeling over continuous and discrete latent spaces using scaled rectified flow transformers for high-resolution image synthesis.
- Paper: FiT: Flexible Vision Transformer for Diffusion Model, Zeyu Lu et al. (2024). This paper explores flexible tokenization and transformer modeling across arbitrary aspect ratios and resolutions, building on modern advances in visual token sequence modeling.
- Paper: Qwen-Image-VAE-2.0 Technical Report, Zekai Zhang et al. (2026). This technical report advances high-compression visual autoencoding architectures to preserve semantic and visual fidelity for downstream generative modeling.
