Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation
Minghui HuYujie WangTat-Jen ChamJianfei YangPonnuthurai N. Suganthan
Proposes a discrete diffusion framework over vector-quantized latent codes that overcomes the sequential bias of autoregressive generation, achieving competitive image synthesis and training-free inpainting with substantially fewer parameters and faster sampling.
Generating high-resolution digital images using deep learning traditionally involves significant trade-offs. Standard compression-based generative models rely on sequential scanning methods that process an image piece by piece, causing them to miss global context and suffer from severe sequential bias. Conversely, advanced diffusion models capture global context effectively but operate directly on high-resolution image space, demanding thousands of iterative steps that result in excessive computing costs and impractically slow generation speeds.
The article introduces and evaluates the Vector Quantised Discrete Diffusion Model (VQ-DDM), a generative framework designed to produce high-fidelity images efficiently. The primary objective is to demonstrate that pairing discrete image compression with a discrete diffusion model can capture full global image context while substantially reducing model size and computational runtimes.
The authors implemented a two-stage approach. First, an autoencoder compresses images into compact grids of discrete visual codes (a visual dictionary or codebook). To resolve the common issue where most dictionary codes go unused, the authors developed a "Re-build and Fine-tune" (ReFiT) clustering technique that maximizes dictionary utilization. Second, a discrete diffusion model learns to generate these visual codes simultaneously rather than sequentially. The authors validated the framework across standard benchmark image datasets, including CelebA-HQ face images, LSUN-Church scenes, and ImageNet, assessing image quality, dictionary utilization, model size, and generation speed against established baselines.
The evaluations yielded several key findings. First, the proposed ReFiT technique increased visual codebook utilization from roughly 32% up to 97–100% on benchmark datasets, improving image reconstruction quality while using up to 32 times fewer codebook entries. Second, VQ-DDM generated competitive image quality with dramatically fewer computational resources; using only 117 million to 120 million parameters, it outperformed older 10-billion-parameter models and matched the visual quality of 800-million-parameter transformer-based systems. Third, by operating within a compressed discrete space, VQ-DDM produced images 10 to 100 times faster than standard continuous diffusion models, reducing 50,000-image generation time from approximately 1,000 hours to around 10 hours on standard hardware. Finally, because the framework evaluates the entire image layout simultaneously, it performed flexible image inpainting and completion on arbitrary missing regions without requiring specialized retraining.
These findings indicate that generative image systems do not need massive parameter scales or slow, full-resolution diffusion processes to achieve top-tier visual fidelity. By significantly lowering computing and energy requirements, this method lowers deployment costs, shortens development cycles, and reduces infrastructure overhead for visual generation tasks. Furthermore, eliminating sequential scanning bias provides greater consistency in downstream tasks such as partial image editing and restoration.
Organizations developing or deploying visual generation technologies should consider adopting discrete diffusion pipelines over massive autoregressive architectures when latency and computing budgets are constraints. Teams should also apply dictionary optimization techniques like ReFiT to existing discrete representation pipelines to maximize code utilization and eliminate wasted model capacity. Further validation across broader multimodal domains, such as audio, video, and cross-modal generation, is recommended before wide-scale deployment.
Readers should note certain boundaries in the presented results. The diffusion process requires a large number of discrete steps (e.g., up to 4,000 steps during training), which can introduce training fluctuations and potential quality limitations on extremely large, complex, and diverse datasets. Nonetheless, for standard resolution image generation benchmarks, the experimental evidence provides high confidence in the model's computational efficiency, parameter reduction, and reconstruction fidelity.
- Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). Introduces Vector Quantised-Variational AutoEncoders (VQ-VAE), providing the foundational discrete latent codebook representation that the source model relies on for image generation.
- Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). Formulates discrete denoising diffusion probabilistic models (D3PMs), establishing the core mathematical mechanisms for performing diffusion processes over discrete categorical state spaces.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Presents the foundational denoising diffusion probabilistic model (DDPM) framework, providing the baseline diffusion principles that the source adapts to discrete visual codebooks.
- Paper: Generating Diverse High-Fidelity Images with VQ-VAE-2, Ali Razavi et al. (2019). Establishes VQ-VAE-2 with autoregressive priors over discrete latent spaces, demonstrating the specific two-stage paradigm and scanning limitations that the source aims to overcome with diffusion.
- Paper: Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions, Emiel Hoogeboom et al. (2021). Develops multinomial diffusion models for categorical distributions, offering essential theoretical background for diffusing over categorical tokens.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Introduces non-Markovian sampling for accelerated diffusion inference, establishing core fast-sampling formulations relevant to diffusion modeling.
- Paper: Regularized Vector Quantization for Tokenized Image Synthesis, Jiahui Zhang et al. (2023). Directly tackles codebook utilization and collapse in tokenized generative synthesis, advancing the quality and stability of the discrete codebooks crucial to models like VQ-DDM.
- Paper: A Continuous Time Framework for Discrete Denoising Models, Andrew Campbell et al. (2022). Extends discrete diffusion modeling into continuous-time Markov chains, generalizing the discrete-state formulation used by discrete token generative models.
- Paper: RandAR: Decoder-only Autoregressive Visual Generation in Random Orders, Ziqi Pang et al. (2025). Explores an alternative approach to overcoming the fixed raster-order limitation in visual autoregressive generation by enabling bidirectional context via random-order generation.
- Paper: Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning, Ting Chen et al. (2023). Presents Bit Diffusion as an alternative technique to bridge continuous diffusion models with discrete visual data representations.
- Paper: Diffusion Models in Vision: A Survey, Florinel-Alin Croitoru et al. (2022). Surveys the broader ecosystem of vision-based diffusion models and taxonomizes various discrete and continuous formulations across visual tasks.
