Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation

Minghui HuYujie WangTat-Jen ChamJianfei YangPonnuthurai N. Suganthan

article2022CVPR53 citations

Proposes a discrete diffusion framework over vector-quantized latent codes that overcomes the sequential bias of autoregressive generation, achieving competitive image synthesis and training-free inpainting with substantially fewer parameters and faster sampling.

Listen

Generating high-resolution digital images using deep learning traditionally involves significant trade-offs. Standard compression-based generative models rely on sequential scanning methods that process an image piece by piece, causing them to miss global context and suffer from severe sequential bias. Conversely, advanced diffusion models capture global context effectively but operate directly on high-resolution image space, demanding thousands of iterative steps that result in excessive computing costs and impractically slow generation speeds.

The article introduces and evaluates the Vector Quantised Discrete Diffusion Model (VQ-DDM), a generative framework designed to produce high-fidelity images efficiently. The primary objective is to demonstrate that pairing discrete image compression with a discrete diffusion model can capture full global image context while substantially reducing model size and computational runtimes.

The authors implemented a two-stage approach. First, an autoencoder compresses images into compact grids of discrete visual codes (a visual dictionary or codebook). To resolve the common issue where most dictionary codes go unused, the authors developed a "Re-build and Fine-tune" (ReFiT) clustering technique that maximizes dictionary utilization. Second, a discrete diffusion model learns to generate these visual codes simultaneously rather than sequentially. The authors validated the framework across standard benchmark image datasets, including CelebA-HQ face images, LSUN-Church scenes, and ImageNet, assessing image quality, dictionary utilization, model size, and generation speed against established baselines.

The evaluations yielded several key findings. First, the proposed ReFiT technique increased visual codebook utilization from roughly 32% up to 97–100% on benchmark datasets, improving image reconstruction quality while using up to 32 times fewer codebook entries. Second, VQ-DDM generated competitive image quality with dramatically fewer computational resources; using only 117 million to 120 million parameters, it outperformed older 10-billion-parameter models and matched the visual quality of 800-million-parameter transformer-based systems. Third, by operating within a compressed discrete space, VQ-DDM produced images 10 to 100 times faster than standard continuous diffusion models, reducing 50,000-image generation time from approximately 1,000 hours to around 10 hours on standard hardware. Finally, because the framework evaluates the entire image layout simultaneously, it performed flexible image inpainting and completion on arbitrary missing regions without requiring specialized retraining.

These findings indicate that generative image systems do not need massive parameter scales or slow, full-resolution diffusion processes to achieve top-tier visual fidelity. By significantly lowering computing and energy requirements, this method lowers deployment costs, shortens development cycles, and reduces infrastructure overhead for visual generation tasks. Furthermore, eliminating sequential scanning bias provides greater consistency in downstream tasks such as partial image editing and restoration.

Organizations developing or deploying visual generation technologies should consider adopting discrete diffusion pipelines over massive autoregressive architectures when latency and computing budgets are constraints. Teams should also apply dictionary optimization techniques like ReFiT to existing discrete representation pipelines to maximize code utilization and eliminate wasted model capacity. Further validation across broader multimodal domains, such as audio, video, and cross-modal generation, is recommended before wide-scale deployment.

Readers should note certain boundaries in the presented results. The diffusion process requires a large number of discrete steps (e.g., up to 4,000 steps during training), which can introduce training fluctuations and potential quality limitations on extremely large, complex, and diverse datasets. Nonetheless, for standard resolution image generation benchmarks, the experimental evidence provides high confidence in the model's computational efficiency, parameter reduction, and reconstruction fidelity.

arXiv: 2112.01799
  • Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). Introduces Vector Quantised-Variational AutoEncoders (VQ-VAE), providing the foundational discrete latent codebook representation that the source model relies on for image generation.
  • Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). Formulates discrete denoising diffusion probabilistic models (D3PMs), establishing the core mathematical mechanisms for performing diffusion processes over discrete categorical state spaces.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Presents the foundational denoising diffusion probabilistic model (DDPM) framework, providing the baseline diffusion principles that the source adapts to discrete visual codebooks.
  • Paper: Generating Diverse High-Fidelity Images with VQ-VAE-2, Ali Razavi et al. (2019). Establishes VQ-VAE-2 with autoregressive priors over discrete latent spaces, demonstrating the specific two-stage paradigm and scanning limitations that the source aims to overcome with diffusion.
  • Paper: Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions, Emiel Hoogeboom et al. (2021). Develops multinomial diffusion models for categorical distributions, offering essential theoretical background for diffusing over categorical tokens.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Introduces non-Markovian sampling for accelerated diffusion inference, establishing core fast-sampling formulations relevant to diffusion modeling.
Cover for Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation

Abstract

The integration of Vector Quantised Variational AutoEncoder (VQ-VAE) with autoregressive models as generation part has yielded high-quality results on image generation. However, the autoregressive models will strictly follow the progressive scanning order during the sampling phase. This leads the existing VQ series models to hardly escape the trap of lacking global information. Denoising Diffusion Probabilistic Models (DDPM) in the continuous domain have shown a capability to capture the global context, while generating high-quality images. In the discrete state space, some works have demonstrated the potential to perform text generation and low resolution image generation. We show that with the help of a content-rich discrete visual codebook from VQ-VAE, the discrete diffusion model can also generate high fidelity images with global context, which compensates for the deficiency of the classical autoregressive model along pixel space. Meanwhile, the integration of the discrete VAE with the diffusion model resolves the drawback of conventional autoregressive models being oversized, and the diffusion model which demands excessive time in the sampling process when generating images. It is found that the quality of the generated images is heavily dependent on the discrete visual codebook. Extensive experiments demonstrate that the proposed Vector Quantised Discrete Diffusion Model (VQ-DDM) is able to achieve comparable performance to top-tier methods with low complexity. It also demonstrates outstanding advantages over other vectors quantised with autoregressive models in terms of image inpainting tasks without additional training.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 2.1. Diffusion Models in continuous state space
  • 2.2. Discrete Representation of Images
  • 3. Methods
  • 3.1. Discrete Diffusion Model
  • 3.2. Re-build and Fine-tune Strategy
  • 4. Experiments and Analysis
  • 4.1. Datasets and Implementation Details
  • 4.2. Codebook Quality
  • 4.3. Generation Quality
  • 4.4. Image Inpainting
  • 5. Related Work
  • 5.1. Vector Quantised Variational Autoencoders
  • 5.2. Diffusion Models
  • 6. Conclusion
  • Limitations
  • References

Knowls

  1. Knowl 1 — Two-stage VQ-DDM image-generation framework

    model/method

    Vector Quantised Discrete Diffusion Model (VQ-DDM) generates images through two stages. First, a discrete variational autoencoder maps an image xx to a low-resolution discrete code map Z0Z_0 using an encoder, a vector-quantisation codebook, and a decoder. Second, a discrete diffusion model learns the prior distribution of complete code maps and generates a code map from noise, after which the decoder converts the generated map into an image.

    Unlike an autoregressive prior that predicts code indices in a fixed spatial order, the VQ-DDM reverse model receives the entire noisy code map at every step and jointly updates the map. Consequently, the reverse process can use spatially global information while operating on a 16×1616\times16 latent map rather than on a full-resolution image.

  2. Knowl 2 — Uniform-resampling discrete diffusion process

    equation

    For a discrete latent variable with KK categories, let zt∈{0,1}Kz_t\in\{0,1\}^K be its one-hot representation at diffusion step tt, let βt∈(0,1]\beta_t\in(0,1] be the noise rate, let IKI_K be the K×KK\times K identity matrix, and let 1K\mathbf{1}_K be the all-ones vector. VQ-DDM uses the transition matrix

    Qt=(1−βt)IK+βtK1K1KT,Q_t=(1-\beta_t)I_K+\frac{\beta_t}{K}\mathbf{1}_K\mathbf{1}_K^{\mathsf T},

    so that the state is retained with probability 1−βt1-\beta_t and is resampled uniformly from all KK categories with probability βt\beta_t:

    q(zt∣zt−1)=Cat⁡(zt; (1−βt)zt−1+βtK1K).q(z_t\mid z_{t-1})=\operatorname{Cat}\left(z_t;\,(1-\beta_t)z_{t-1}+\frac{\beta_t}{K}\mathbf{1}_K\right).

    With αt=1−βt\alpha_t=1-\beta_t and αˉt=∏r=1tαr\bar\alpha_t=\prod_{r=1}^{t}\alpha_r, the noisy state can be sampled directly from the clean state z0z_0 according to

    q(zt∣z0)=Cat⁡(zt; αˉtz0+1−αˉtK1K).q(z_t\mid z_0)=\operatorname{Cat}\left(z_t;\,\bar\alpha_t z_0+\frac{1-\bar\alpha_t}{K}\mathbf{1}_K\right).

    The experiments use a cosine schedule, defined for continuous time t∈[0,T]t\in[0,T] by αˉ(t)=f(t)/f(0)\bar\alpha(t)=f(t)/f(0), where f(t)=cos⁡2 ⁣(t/T+s1+sπ2)f(t)=\cos^2\!\left(\frac{t/T+s}{1+s}\frac{\pi}{2}\right), ss is the schedule offset, and T=4000T=4000 diffusion steps.

  3. Knowl 3 — Global-context reverse process and discrete diffusion objective

    model/method

    The reverse model predicts a distribution over the clean code z^0\hat z_0 from the complete noisy code map Zt∈{0,1}K×h×wZ_t\in\{0,1\}^{K\times h\times w} and the timestep tt. The predictor is a U-Net with self-attention; it estimates a noise representation and converts it into logits for z^0\hat z_0. For a single spatial position, define

    θk(zt,z^0)=[αtzt,k+1−αtK][αˉt−1z^0,k+1−αˉt−1K],\theta_k(z_t,\hat z_0)=\left[\alpha_t z_{t,k}+\frac{1-\alpha_t}{K}\right]\left[\bar\alpha_{t-1}\hat z_{0,k}+\frac{1-\bar\alpha_{t-1}}{K}\right],

    where zt,kz_{t,k} and z^0,k\hat z_{0,k} are the one-hot or predicted probabilities for category kk, k∈{1,…,K}k\in\{1,\ldots,K\}. If N[θ]k=θk/∑j=1Kθj\mathcal N[\theta]_k=\theta_k/\sum_{j=1}^{K}\theta_j denotes normalization, the learned reverse transition is

    pθ(zt−1∣zt)=Cat⁡(zt−1; N[θ(zt,z^0)]),p_\theta(z_{t-1}\mid z_t)=\operatorname{Cat}\left(z_{t-1};\,\mathcal N[\theta(z_t,\hat z_0)]\right),

    with pθ(z0∣z1)=Cat⁡(z0;z^0)p_\theta(z_0\mid z_1)=\operatorname{Cat}(z_0;\hat z_0). Although the terminal code variables can be sampled independently from a uniform categorical distribution, the neural network processes the whole map jointly, allowing each reverse update to capture dependencies among all spatial positions.

    For t>1t>1, the training contribution is the categorical KL divergence between the exact posterior based on the clean code and the model posterior based on the predicted clean code:

    Lt=∑k=1KN[θ(zt,z0)]klog⁡N[θ(zt,z0)]kN[θ(zt,z^0)]k.\mathcal L_t=\sum_{k=1}^{K}\mathcal N[\theta(z_t,z_0)]_k\log\frac{\mathcal N[\theta(z_t,z_0)]_k}{\mathcal N[\theta(z_t,\hat z_0)]_k}.

    The complete objective is the variational lower bound over the discrete forward and reverse chains.

  4. Knowl 4 — Re-build and Fine-tune codebook strategy

    algorithm

    Re-build and Fine-tune (ReFiT) replaces an under-utilised codebook with a smaller codebook whose entries are sampled from valid encoder features, then restores reconstruction quality by fine-tuning.

    Inputs are a trained encoder EsE_s, decoder DsD_s, original codebook, training images, target codebook capacity KtK_t, and feature-sampling count PP. The output is a discrete autoencoder with a rebuilt codebook of capacity KtK_t.

    Input: trained encoder E_s, decoder D_s, training images, target capacity K_t, sample count P
    Output: rebuilt codebook Z_t and fine-tuned discrete autoencoder
    Encode every training image with E_s and collect its latent feature vectors
    Sample P feature vectors uniformly from the collected feature set
    Run k-means with AFK-MC2 seeding on the P sampled vectors using K_t clusters
    Set the K_t cluster centres as the new codebook Z_t
    Replace the original codebook with Z_t
    Fine-tune the discrete autoencoder using the rebuilt codebook
    Freeze the encoder during fine-tuning
    Use decoder learning rate 1e-6 and discriminator learning rate 2e-6 with batch size 8
    Return the rebuilt codebook and fine-tuned decoder

    The paper uses P=20,000P=20{,}000 for CelebA-HQ and LSUN-Church, P=50,000P=50{,}000 for ImageNet, and also tests P=100,000P=100{,}000. ReFiT addresses the fact that straight-through codebook training updates only the entries selected by current encoder features, leaving many entries unused or trapped near poor initial values.

  5. Knowl 5 — Training and implementation configuration

    experimental setup

    The experiments evaluate VQ-DDM on CelebA-HQ and LSUN-Church, and evaluate ReFiT codebook reconstruction on CelebA-HQ and ImageNet. Images are resized to 256×256256\times256 pixels and compressed by a factor of 16 into a latent representation of size 1×16×161\times16\times16. The discrete autoencoder follows the VQ-GAN training strategy.

    The diffusion prior uses a U-Net with self-attention and T=4000T=4000 diffusion steps with the cosine noise schedule. For ReFiT, features are sampled uniformly from training images and then uniformly across the 16×1616\times16 feature map. During fine-tuning, the encoder is frozen; the decoder learning rate is 10−610^{-6}, the discriminator learning rate is 2×10−62\times10^{-6}, and the batch size is 8.

  6. Knowl 6 — ReFiT improves codebook utilisation and reconstruction

    data/table

    This comparison measures how many codebook entries appear in the data and how faithfully the discrete autoencoder reconstructs images. Usage is the fraction of codebook entries observed in the relevant dataset, and FID is computed between reconstructed images and original images; lower FID is better. ReFiT obtains nearly complete codebook utilisation with capacities 512 or 1024, while the baseline VQ-GAN codebook uses only a minority of its entries.

    Model Latent size Capacity Usage CelebA Usage ImageNet FID CelebA FID ImageNet
    VQ-VAE-2 Cascade 512 65% – – 10
    DALL-E 323×2 8192 – – – 32.01
    VQ-GAN 161×6 16384 – 5.96% – 4.98
    VQ-GAN 161×6 1024 31.85% 33.67% 10.18 7.94
    ReFiT (P=100k) 161×6 1024 – 100% – 4.98
    ReFiT (P=20k) 161×6 1024 97.07% 100% 5.59 5.99
    ReFiT (P=20k) 161×6 512 93.06% 100% 5.64 6.95

    At capacity 1024 and P=20,000P=20{,}000, ReFiT raises CelebA-HQ usage from 31.85% to 97.07% and reduces reconstruction FID from 10.18 to 5.59. Reducing capacity from 1024 to 512 causes only a small FID increase on CelebA-HQ and an approximately one-point increase on ImageNet, enabling the diffusion model to use a much smaller category set.

  7. Knowl 7 — Image-generation quality with a compact discrete prior

    data/table

    The following comparison reports FID on unconditional CelebA-HQ images at 256×256256\times256 resolution. Parameter counts and FLOPs refer only to the generation or inference stage for one image; lower FID and lower computation are preferred. VQ-DDM with ReFiT reaches substantially better FID than the same model without ReFiT while using 117 million parameters and roughly 1.04--1.06 billion FLOPs.

    Method FID Parameters FLOPs
    Likelihood-based
    GLOW 60.9 220M 540G
    NVAE 40.3 1.26G 185G
    VQ-DDM (K=1024, without ReFiT) 22.6 117M 1.06G
    VAEBM 20.4 127M 8.22G
    VQ-DDM (K=512, with ReFiT) 18.8 117M 1.04G
    DC-VAE 15.8 – –
    VQ-DDM (K=1024, with ReFiT) 13.2 117M 1.06G
    DDIM (T=100) 10.9 114M 124G
    VQ-GAN + Transformer 10.2 802M 102G
    Likelihood-free
    PG-GAN 8.0 46.1M 14.1G

    For CelebA-HQ, the K=1024K=1024 ReFiT model has training negative log-likelihood 1.258, test negative log-likelihood 1.286, and FID 13.2. On LSUN-Church, the corresponding K=1024K=1024 model has training negative log-likelihood 1.803, test negative log-likelihood 1.756, and FID 16.9. Within the tested range, increasing codebook capacity improves generation after ReFiT, but the authors report that excessively many entries can cause model collapse.

  8. Knowl 8 — Discrete latent diffusion substantially reduces sampling cost

    empirical result

    For generating 50,000 256×256256\times256 images on one NVIDIA 2080Ti GPU, the reported approximate times are 1,000 hours for a DDPM using 1,000 reverse steps, 100 hours for DDIM using 100 steps, and about 10 hours for VQ-DDM using 1,000 discrete reverse steps. VQ-DDM performs the diffusion computation on 16×1616\times16 categorical code maps and decodes the resulting maps afterward, rather than denoising full-resolution images at every step.

    The reported generation-stage comparison therefore shows that the discrete latent formulation can be roughly two orders of magnitude faster than the DDPM baseline and about one order of magnitude faster than the DDIM baseline under these settings, while retaining competitive image quality.

  9. Knowl 9 — Mask diffusion enables context-independent image inpainting

    model/method

    VQ-DDM performs inpainting directly in the discrete latent space without additional training. Let m∈{0,1}h×wm\in\{0,1\}^{h\times w} be a spatial mask, with m=0m=0 denoting masked positions and m=1m=1 denoting retained context. Encode the masked image into a discrete map z0z_0, diffuse it to obtain z~t∼q(zt∣z0)\tilde z_t\sim q(z_t\mid z_0), and use a uniformly sampled categorical code C∼Cat⁡(K,1/K)C\sim\operatorname{Cat}(K,1/K) to initialize the terminal masked state according to

    z~Tm=(1−m)⊙z~T+m⊙C.\tilde z_T^{m}=(1-m)\odot\tilde z_T+m\odot C.

    At the first reverse step, sample zT−1z_{T-1} from pθ(zT−1∣z~Tm)p_\theta(z_{T-1}\mid\tilde z_T^m). At subsequent steps, sample the predicted complete map from pθ(zt−1∣ztm)p_\theta(z_{t-1}\mid z_t^m) and overwrite the retained positions with the corresponding diffused context:

    zt−1m=(1−m)⊙zt−1+m⊙z~t−1.z_{t-1}^{m}=(1-m)\odot z_{t-1}+m\odot\tilde z_{t-1}.

    The decoder converts the final completed latent map into an image. Because every reverse update conditions on the full latent map rather than a fixed scan order, the method produces consistent completions for arbitrary masks. The experiments include masking the upper 62.5% of an image, retaining only a lower-right quarter, and masking the perimeter while retaining the centre. VQ-DDM produced more coherent completions than the compared sliding-window VQ-GAN transformer and required no task-specific retraining.

  10. Knowl 10 — Limitation on large and complex datasets

    limitation

    The complete discrete diffusion process requires many reverse steps. The resulting long chain makes training fluctuate and can limit image-generation quality, so the authors caution that VQ-DDM may underperform on large-scale and complex datasets.

Coverage note — Continuous Gaussian diffusion preliminaries, standard VQ-VAE/VQ-GAN background, related work, and qualitative sample images were omitted because they are not additional contributions beyond the reconstructed VQ-DDM method and its experiments.

References

  1. 1.Jacob Austin, Daniel Johnson, Jonathan Ho, Danny Tarlow, and Rianne van den Berg. Structured denoising diffusion models in discrete state-spaces. arXiv preprint arXiv:2107.03006, 2021. 1, 5, 7
  2. 2.Olivier Bachem, Mario Lucic, Hamed Hassani, and Andreas Krause. Fast and provably good seedings for k-means. Advances in neural information processing systems, 29:55–63, 2016. 5
  3. 3.Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. Scheduled sampling for sequence prediction with recurrent neural networks. arXiv preprint arXiv:1506.03099, 2015. 1
  4. 4.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pre-training from pixels. In International Conference on Machine Learning, pages 1691–1703. PMLR, 2020. 1, 6
  5. 5.Xi Chen, Nikhil Mishra, Mostafa Rohaninejad, and Pieter Abbeel. Pixelsnail: An improved autoregressive generative model. In International Conference on Machine Learning, pages 864–872. PMLR, 2018. 1, 6
  6. 6.Prafulla Dhariwal and Alex Nichol. Diffusion models beat gans on image synthesis. arXiv e-prints, pages arXiv–2105, 2021. 1, 7
  7. 7.Patrick Esser, Robin Rombach, Andreas Blattmann, and Bjorn Ommer. Imagebart: Bidirectional context with multinomial diffusion for autoregressive image synthesis. arXiv preprint arXiv:2108.08827, 2021. 7
  8. 8.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12873–12883, 2021. 3, 5, 6, 7
  9. 9.Vincent Fortuin, Matthias Huser, Francesco Locatello, Heiko Strathmann, and Gunnar Ratsch. Som-vae: Interpretable discrete representation learning on time series. arXiv preprint arXiv:1806.02199, 2018. 6
  10. 10.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. arXiv preprint arxiv:2006.11239, 2020. 1, 2, 3, 5, 6
  11. 11.Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation. arXiv preprint arXiv:2106.15282, 2021. 7
  12. 12.Emiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forre, and Max Welling. Argmax flows and multinomial diffusion: Towards non-autoregressive language models. arXiv preprint arXiv:2102.05379, 2021. 1, 4, 6, 7
  13. 13.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016. 4
  14. 14.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196, 2017. 5, 6, 7
  15. 15.Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. Transformers in vision: A survey. arXiv preprint arXiv:2101.01169, 2021. 1
  16. 16.Diederik P Kingma and Prafulla Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. arXiv preprint arXiv:1807.03039, 2018. 6, 7
  17. 17.Diederik P Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. arXiv preprint arXiv:2107.00630, 2021. 8
  18. 18.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. 2
  19. 19.Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. Diffwave: A versatile diffusion model for audio synthesis. In International Conference on Learning Representations, 2020. 1
  20. 20.Chris J Maddison, Andriy Mnih, and Yee Whye Teh. The concrete distribution: A continuous relaxation of discrete random variables. arXiv preprint arXiv:1611.00712, 2016. 4
  21. 21.Alex Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. arXiv preprint arXiv:2102.09672, 2021. 2, 3, 4, 5
  22. 22.Aaron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu. Conditional image generation with pixelcnn decoders. arXiv preprint arXiv:1606.05328, 2016. 1
  23. 23.Gaurav Parmar, Dacheng Li, Kwonjoon Lee, and Zhuowen Tu. Dual contradistinctive generative autoencoder. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 823–832, 2021. 6, 7
  24. 24.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092, 2021. 1, 5, 6
  25. 25.Ali Razavi, Aaron van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with vq-vae-2. In Advances in neural information processing systems, pages 14866–14876, 2019. 5
  26. 26.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015. 5
  27. 27.Aurko Roy, Ashish Vaswani, Arvind Neelakantan, and Niki Parmar. Theory and experiments on vector quantized autoencoders. arXiv preprint arXiv:1805.11063, 2018. 6
  28. 28.Abhishek Sinha, Jiaming Song, Chenlin Meng, and Stefano Ermon. D2c: Diffusion-denoising models for few-shot conditional generation. arXiv preprint arXiv:2106.06819, 2021. 1
  29. 29.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015. 1, 2, 7
  30. 30.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations, 2020. 2, 6, 7
  31. 31.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. In Proceedings of the 33rd Annual Conference on Neural Information Processing Systems, 2019. 3
  32. 32.Arash Vahdat and Jan Kautz. Nvae: A deep hierarchical variational autoencoder. arXiv preprint arXiv:2007.03898, 2020. 6, 7
  33. 33.Arash Vahdat, Karsten Kreis, and Jan Kautz. Score-based generative modeling in latent space. Advances in Neural Information Processing Systems, 34, 2021. 1
  34. 34.Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pages 6309–6318, 2017. 1, 3, 6
  35. 35.Aaron Van Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. Pixel recurrent neural networks. In International Conference on Machine Learning, pages 1747–1756. PMLR, 2016. 1, 6
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017. 5
  37. 37.Antoine Wehenkel and Gilles Louppe. Diffusion priors in variational autoencoders. In ICML Workshop on Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models, 2021. 8
  38. 38.Zhisheng Xiao, Karsten Kreis, Jan Kautz, and Arash Vahdat. Vaebm: A symbiosis between variational autoencoders and energy-based models. In International Conference on Learning Representations, 2020. 6, 7
  39. 39.Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015. 5
  40. 40.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018. 3
  41. 41.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networkss. In Computer Vision (ICCV), 2017 IEEE International Conference on, 2017. 3

Citation

MLA
Hu, M., et al. “Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation”. arXiv, 2021, http://arxiv.org/abs/2112.01799v1.
APA
Hu, M., Wang, Y., Cham, T.-J., Yang, J., & Suganthan, P. N. (2021). Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation. arXiv. http://arxiv.org/abs/2112.01799v1
Chicago
Hu, M., Y. Wang, T.-J. Cham, J. Yang, and P. N. Suganthan. 2021. “Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation”. arXiv. http://arxiv.org/abs/2112.01799v1.
Harvard
Hu, M. et al. (2021) “Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.01799v1.
Vancouver
1. Hu M, Wang Y, Cham T-J, Yang J, Suganthan PN (2021) Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation. arXiv

BibTeX

@article{hu2021global,
  title = {Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation},
  author = {Hu, Minghui and Wang, Yujie and Cham, Tat-Jen and Yang, Jianfei and Suganthan, P. N.},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.01799v1},
  eprint = {2112.01799}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE