Regularized Vector Quantization for Tokenized Image Synthesis

Jiahui ZhangFangneng ZhanChristian TheobaltShijian Lu

article2023CVPR56 citations

Proposes a dual-regularized vector quantization framework with a probabilistic contrastive loss that prevents codebook collapse and aligns training with stochastic sampling for superior image synthesis in autoregressive and diffusion models.

Listen

Discrete representation learning, which converts continuous images into compact sequences of discrete tokens, is a foundational technique for modern generative artificial intelligence across autoregressive and diffusion architectures. However, current vector quantization approaches face fundamental trade-offs. Deterministic methods, which select the single best-matching token, suffer from severe codebook collapse where many codebook entries remain unlearned or invalid, while also creating a mismatch with generative inference where tokens are sampled probabilistically. Conversely, purely stochastic methods, which sample tokens from predicted distributions, lead to underutilized codebook capacity and perturbed reconstruction targets that degrade overall image quality.

The article introduces and evaluates a regularized vector quantization framework designed to overcome these limitations. The approach introduces a prior distribution regularization that encourages uniform token usage across the codebook via divergence minimization, preventing codebook collapse and maximizing representational capacity. In addition, it employs a stochastic mask regularization that applies probabilistic sampling to a subset of spatial regions while keeping the rest deterministic, paired with an adaptive probabilistic contrastive loss to enable flexible, artifact-free image reconstruction. The framework was evaluated across standard benchmarks including ADE20K, CelebA-HQ, CUB-200, and MS-COCO on semantic and text-to-image synthesis tasks.

The empirical findings demonstrate that the regularized quantization framework consistently outperforms prevailing deterministic and stochastic baselines in both image reconstruction fidelity and downstream generation quality. Specifically, the method achieved superior generation scores across all benchmarks, reducing the generation Fréchet Inception Distance on ADE20K to 34.47 compared to 38.53 for deterministic VQ-GAN and 37.51 for Gumbel-VQ, while achieving full codebook utilization. Experiments also established that a 40% stochastic masking ratio provides the optimal balance between inference alignment and reconstruction stability, and confirmed that performance scales effectively with larger codebook sizes up to 8,192 entries where conventional models typically stagnate.

These results indicate that generative image systems can achieve higher fidelity and better scaling efficiency without incurring major structural redesigns of underlying generative models. Organizations developing generative vision models should adopt dual-strategy quantization combining prior regularization and hybrid deterministic-stochastic masking to maximize codebook efficiency. While the article provides high confidence across autoregressive and discrete diffusion models on standard 256x256 benchmark resolutions, further testing in larger-scale commercial production pipelines and higher native image resolutions is recommended before enterprise-wide deployment.

arXiv: 2303.06424
  • Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). This foundational work establishes the VQGAN framework of discrete tokenization and autoregressive modeling that the source regularizes to eliminate codebook collapse and improve generation.
  • Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). This paper introduces vector-quantized representation learning (VQ-VAE), providing the foundational discrete codebook and straight-through estimator mechanisms analyzed and improved in the source.
  • Paper: Generating Diverse High-Fidelity Images with VQ-VAE-2, Ali Razavi et al. (2019). This work demonstrates scaling discrete multi-scale vector quantization for high-fidelity image generation, establishing key architectural baselines for downstream tokenized synthesis.
  • Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). This work highlights the two-stage visual tokenization paradigm for text-to-image autoregressive generation, demonstrating the importance of codebook efficiency in downstream generative pipelines.
  • Paper: Scaling Autoregressive Models for Content-Rich Text-to-Image Generation, Jiahui Yu et al. (2022). This study analyzes the scaling behavior of autoregressive visual token modeling, motivating the source's focus on maximizing codebook capacity without codebook collapse.
Cover for Regularized Vector Quantization for Tokenized Image Synthesis

Abstract

Quantizing images into discrete representations has been a fundamental problem in unified generative modeling. Predominant approaches learn the discrete representation either in a deterministic manner by selecting the best-matching token or in a stochastic manner by sampling from a predicted distribution. However, deterministic quantization suffers from severe codebook collapse and misalignment with inference stage while stochastic quantization suffers from low codebook utilization and perturbed reconstruction objective. This paper presents a regularized vector quantization framework that allows to mitigate above issues effectively by applying regularization from two perspectives. The first is a prior distribution regularization which measures the discrepancy between a prior token distribution and the predicted token distribution to avoid codebook collapse and low codebook utilization. The second is a stochastic mask regularization that introduces stochasticity during quantization to strike a good balance between inference stage misalignment and unperturbed reconstruction objective. In addition, we design a probabilistic contrastive loss which serves as a calibrated metric to further mitigate the perturbed reconstruction objective. Extensive experiments show that the proposed quantization framework outperforms prevailing vector quantization methods consistently across different generative models including auto-regressive models and diffusion models.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Vector Quantization
  • 2.2. Tokenized Image Synthesis
  • 3. Method
  • 3.1. Prior Distribution Regularization
  • 3.2. Stochastic Mask Regularization
  • 3.3. Probabilistic Contrastive Loss
  • 4. Experiments
  • 4.1. Experimental Settings
  • 4.2. Quantitative Evaluation
  • 4.3. Qualitative Evaluation
  • 4.4. Ablation Study
  • 5. Conclusions
  • 6. Acknowledgments
  • References

Knowls

  1. Knowl 1 — Regularized vector quantization architecture

    model/method

    The proposed regularized vector quantization framework represents an image XX using an encoder EE, decoder GG, and a learned codebook Z={zn}n=1N∈RN×dZ=\{z_n\}_{n=1}^{N}\in\mathbb{R}^{N\times d}, where NN is the number of codebook embeddings and dd is their embedding dimension. The encoder produces a token-probability vector for each position in an H×WH\times W spatial grid, giving P∈RH×W×NP\in\mathbb{R}^{H\times W\times N}. Each spatial position is quantized to a codebook embedding, and the resulting H×WH\times W embeddings are passed to GG to reconstruct the image.

    The framework combines deterministic quantization, which selects the most probable token, with stochastic quantization, which samples a token using Gumbel sampling. A learned image-token representation can subsequently be modeled by an autoregressive or diffusion model; at generation time, sampled token indices are mapped back to codebook embeddings and decoded by GG. The proposed framework adds prior-distribution regularization, stochastic mask regularization, and probabilistic contrastive reconstruction to address codebook collapse, low codebook utilization, inference-stage mismatch, and reconstruction perturbation.

  2. Knowl 2 — Prior distribution regularization for codebook utilization

    model/method

    To prevent codebook collapse and low codebook utilization, the method assumes that codebook tokens should follow a uniform prior distribution. For a codebook of size NN, the prior is Pprior=[1/N,…,1/N]∈RNP_{\mathrm{prior}}=[1/N,\ldots,1/N]\in\mathbb{R}^{N}. For each spatial position i∈{1,…,HW}i\in\{1,\ldots,H W\}, let pi∈{0,1}Np_i\in\{0,1\}^{N} be the one-hot vector representing the selected token. The empirical posterior token distribution is

    Ppost=1HW∑i=1HWpi=[p1,…,pN],P_{\mathrm{post}}=\frac{1}{H W}\sum_{i=1}^{H W}p_i=[p_1,\ldots,p_N],

    where pnp_n is the empirical frequency of codebook token nn. The regularizer minimizes the KL divergence from the empirical posterior to the uniform prior:

    Lkl=KL(Ppost,Pprior)=−∑n=1Npnlog⁡(1/Npn).\mathcal{L}_{\mathrm{kl}}=\mathrm{KL}(P_{\mathrm{post}},P_{\mathrm{prior}})=-\sum_{n=1}^{N}p_n\log\left(\frac{1/N}{p_n}\right).

    Minimizing this term encourages the quantizer to use codebook embeddings more uniformly, thereby reducing collapse of embeddings to a small subset and increasing the effective information capacity of the codebook.

  3. Knowl 3 — Stochastic mask regularization

    model/method

    Stochastic mask regularization combines deterministic and stochastic token selection at different image locations. Let P∈RH×W×NP\in\mathbb{R}^{H\times W\times N} be the predicted token probabilities and let M∈{0,1}H×WM\in\{0,1\}^{H\times W} be a randomly sampled spatial mask. A value Mij=1M_{ij}=1 selects Gumbel-based stochastic quantization at position (i,j)(i,j), whereas Mij=0M_{ij}=0 selects deterministic Argmax quantization. If XargmaxX_{\mathrm{argmax}} and XgumbelX_{\mathrm{gumbel}} are the corresponding quantized embedding grids, the reconstruction loss is

    Lrec=∥X−G(Xargmax⊙(1−M)+Xgumbel⊙M)∥1,\mathcal{L}_{\mathrm{rec}}=\left\|X-G\left(X_{\mathrm{argmax}}\odot(1-M)+X_{\mathrm{gumbel}}\odot M\right)\right\|_{1},

    where XX is the original image, GG is the decoder, and ⊙\odot denotes spatially broadcast elementwise multiplication. Deterministic regions preserve alignment with the reconstruction target, while stochastic regions introduce uncertainty resembling token sampling at the inference stage of a generative model. The paper uses a masking ratio of 40%40\% by default.

  4. Knowl 4 — Hard-forward, differentiable-backward quantization procedure

    algorithm

    The proposed quantization procedure applies different forward and backward computations to deterministic and stochastic spatial regions. It uses straight-through-style reparameterization: hard one-hot selections are used to form the forward quantized embeddings, while soft probability-weighted embeddings are used for gradient backpropagation.

    Input: Encoded vector x_ij, token distribution P_ij = [p_1, ..., p_N], codebook Z, mask value M_ij
    Output: Quantized embedding x_hat_ij
    if M_ij = 0 then
        Forward: index = Argmax(P_ij)
        Forward: index_hard = OneHot(index)
        Forward: x_hat_ij = Matmul(index_hard, Z)
        Backward: index_soft = Softmax(P_ij)
        Backward: x_hat_ij = Matmul(index_soft, Z)
    else
        Sample independent gumbels from Gumbel(0, 1)
        Forward: index = Argmax(P_ij + gumbels)
        Forward: index_hard = OneHot(index)
        Forward: x_hat_ij = Matmul(index_hard, Z)
        Backward: index_soft = Softmax(P_ij + gumbels)
        Backward: x_hat_ij = Matmul(index_soft, Z)
    end if
    return x_hat_ij
  5. Knowl 5 — Probabilistic contrastive loss for elastic reconstruction

    equation

    For spatially corresponding feature patches from an original image and its reconstruction, the method treats patches at the same location as positive pairs and patches at different locations as negative pairs. Let Y=[y1,…,yL]Y=[y_1,\ldots,y_L] and Z=[z1,…,zL]Z=[z_1,\ldots,z_L] be the original-image and reconstructed-image feature patches, respectively; LL is the number of patches and τ\tau is the contrastive temperature. For each patch ii, let zs,iz_{s,i} be the randomly sampled codebook embedding and zq,iz_{q,i} the best-matching embedding selected by deterministic Argmax. The perturbation weight is

    wi=∥zs,i−zq,i∥22,w_i=\left\|z_{s,i}-z_{q,i}\right\|_2^2,

    and its normalized value is wi′=wi/∑k=1Lwkw_i'=w_i/\sum_{k=1}^{L}w_k, so that ∑iwi′=1\sum_i w_i'=1. The probabilistic contrastive loss is

    Lpcl=−∑i=1Llog⁡wi′exp⁡(yi⋅zi/τ)wi′exp⁡(yi⋅zi/τ)+1L∑j=1j≠iLexp⁡(yi⋅zj/τ).\mathcal{L}_{\mathrm{pcl}}=-\sum_{i=1}^{L}\log\frac{w_i'\exp(y_i\cdot z_i/\tau)}{w_i'\exp(y_i\cdot z_i/\tau)+\frac{1}{L}\sum_{\substack{j=1\\j\neq i}}^{L}\exp(y_i\cdot z_j/\tau)}.

    The weight adapts the positive-pair attraction to the discrepancy caused by token sampling: sampled embeddings farther from their best-matching embeddings receive a different contrastive pulling force. The factor 1/L1/L balances the negative term against the positive term. The paper extracts the image features with pretrained VGG-19 layers relu1_2, relu2_2, relu3_3, relu4_3, and relu5_3; the contrastive loss is used together with perceptual loss during training.

  6. Knowl 6 — Evaluation protocol and implementation settings

    experimental setup

    The method is evaluated on ADE20K and CelebA-HQ for semantic image synthesis, and on CUB-200 and MS-COCO for text-to-image synthesis. Autoregressive generative models are used for ADE20K and CelebA-HQ, while a diffusion model is used for CUB-200 and MS-COCO. Reconstruction quality is measured by reconstructed-image FID, denoted FID[R], and PSNR; generated-image quality is measured by FID[G]. Lower FID and higher PSNR are preferred.

    Following the VQ-GAN configuration, all quantizers use a 16×1616\times16 spatial feature grid, a codebook of N=1024N=1024 embeddings, and 256×256256\times256 images. The proposed method uses a 40%40\% stochastic masking ratio by default. Comparisons include VQ-VAE, deterministic VQ-GAN, and Gumbel-VQ, a stochastic VQ-GAN variant using Gumbel-Softmax.

  7. Knowl 7 — Cross-dataset reconstruction and generation results

    data/table

    The main comparison evaluates reconstructed images and generated images across semantic and text-to-image synthesis. Reg-VQ is the proposed regularized quantizer. FID[R] measures reconstructed-image quality, PSNR[R] measures pixel-level reconstruction accuracy, and FID[G] measures generated-image quality; ↓\downarrow means lower is better and ↑\uparrow means higher is better. The proposed method obtains the best FID[R] and FID[G] on all four datasets, although VQ-VAE or VQ-GAN can have higher PSNR in some cases.

    Could not parse LaTeX table
  8. Knowl 8 — Component ablation on ADE20K

    data/table

    An ablation on ADE20K semantic image synthesis starts from deterministic VQ-GAN and adds the proposed components incrementally. Prior distribution regularization improves reconstruction FID, PSNR, and generation FID. Adding stochastic mask regularization further improves FID[R] and FID[G] but lowers PSNR[R], reflecting the trade-off between stochastic inference alignment and exact pixel reconstruction. Vanilla contrastive loss provides an additional improvement, while the proposed probabilistic contrastive loss gives the strongest final results.

    Could not parse LaTeX table

    Relative to the baseline, the complete method reduces FID[R] from 28.1728.17 to 23.6923.69 and FID[G] from 38.5338.53 to 34.4734.47.

  9. Knowl 9 — Effect of the stochastic masking ratio

    empirical result

    On ADE20K, the paper compares deterministic vector quantization (D-VQ), fully stochastic vector quantization (S-VQ), and stochastic mask regularization (MaskReg) across masking ratios from 0.00.0 to 1.01.0. VQ-GAN tokenizers and an autoregressive generator are used; PSNR evaluates reconstruction and FID evaluates generation.

    The experiments identify a 40%40\% masking ratio as the best setting for the proposed MaskReg method when reconstruction and generation are considered jointly, with the best generation quality measured by FID. The result demonstrates that neither fully deterministic quantization nor fully stochastic quantization provides the best overall trade-off. The paper also emphasizes that a higher PSNR does not necessarily imply better generative quality, because exact reconstruction and compatibility with randomly sampled inference tokens are different objectives.

  10. Knowl 10 — Scaling behavior with codebook size

    empirical result

    The paper evaluates codebook sizes N∈{1024,2048,4096,6144,8192}N\in\{1024,2048,4096,6144,8192\} on ADE20K and CUB-200. Autoregressive generation is used for ADE20K and diffusion generation for CUB-200. Visualizations at N=1024N=1024, 40964096, and 81928192 show that deterministic VQ-GAN develops increasingly severe codebook collapse as the codebook grows: many embeddings become invalid or unused. In contrast, regularized quantization maintains high codebook utilization even at N=8192N=8192.

    Quantitative FID curves show that Reg-VQ continues to improve reconstruction and generation quality as the codebook size increases, whereas VQ-GAN tends to plateau at a worse FID level. Thus, the uniform-prior regularization is particularly beneficial when larger codebooks are used, where deterministic quantization suffers from stronger collapse.

Coverage note — The paper's qualitative image examples and supplementary implementation details were not made separate knowls because they illustrate the quantitative comparisons or provide non-load-bearing training details already covered by the method and experimental protocol.

References

  1. 1.Jacob Austin, Daniel Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg. Structured denoising diffusion models in discrete state-spaces. Advances in Neural Information Processing Systems, 34, 2021. 3
  2. 2.Alexei Baevski, Steffen Schneider, and Michael Auli. vq-wav2vec: Self-supervised learning of discrete speech representations. arXiv preprint arXiv:1910.05453, 2019. 1, 2, 3, 6
  3. 3.Tadas Baltrusaitis, Chaitanya Ahuja, and Louis-Philippe Morency. Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2):423–443, 2018. 1
  4. 4.Yoshua Bengio, Nicholas Leonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013. 2
  5. 5.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pre-training from pixels. In International Conference on Machine Learning, pages 1691–1703. PMLR, 2020. 3
  6. 6.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34, 2021. 1, 3
  7. 7.Xiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, and Nenghai Yu. Peco: Perceptual codebook for bert pre-training of vision transformers. arXiv preprint arXiv:2111.12710, 2021. 5
  8. 8.Patrick Esser, Robin Rombach, and Björn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021. 1, 2, 3, 5, 6, 7
  9. 9.Karol Gregor, Ivo Danihelka, Andriy Mnih, Charles Blundell, and Daan Wierstra. Deep autoregressive networks. In International Conference on Machine Learning, pages 1242–1250. PMLR, 2014. 1, 3
  10. 10.Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quantized diffusion model for text-to-image synthesis. arXiv preprint arXiv:2111.14822, 2021. 3, 6
  11. 11.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, pages 6626–6637, 2017. 5
  12. 12.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020. 1, 3
  13. 13.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997. 3
  14. 14.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016. 2, 3
  15. 15.Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In Proceedings of the European Conference on Computer Vision (ECCV), pages 694–711. Springer, 2016. 5
  16. 16.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, 2014. 5, 6
  17. 17.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of the IEEE International Conference on Computer Vision, pages 3730–3738, 2015. 5, 6
  18. 18.Chris J Maddison, Daniel Tarlow, and Tom Minka. A* sampling. In NIPS, 2014. 3
  19. 19.Todd K. Moon. The expectation-maximization algorithm. IEEE Signal processing magazine, 13(6):47–60, 1996. 3
  20. 20.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018. 5
  21. 21.Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. arXiv preprint arXiv:1711.00937, 2017. 1, 2, 3
  22. 22.Taesung Park, Alexei A Efros, Richard Zhang, and Jun-Yan Zhu. Contrastive learning for unpaired image-to-image translation. In European Conference on Computer Vision, pages 319–345. Springer, 2020. 2, 5
  23. 23.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021. 1, 3, 6
  24. 24.Ali Razavi, Aaron Van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems, 32, 2019. 2
  25. 25.Jason Tyler Rolfe. Discrete variational autoencoders. arXiv preprint arXiv:1609.02200, 2016. 1
  26. 26.Aurko Roy, Ashish Vaswani, Arvind Neelakantan, and Niki Parmar. Theory and experiments on vector quantized autoencoders. arXiv preprint arXiv:1805.11063, 2018. 1, 3
  27. 27.Woncheol Shin, Gyubok Lee, Jiyoung Lee, Joonseok Lee, and Edward Choi. Translation-equivariant image quantizer for bi-directional image-text generation. arXiv preprint arXiv:2112.00384, 2021. 2
  28. 28.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 5
  29. 29.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015. 3
  30. 30.Aaron Van Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. Pixel recurrent neural networks. In International Conference on Machine Learning, pages 1747–1756. PMLR, 2016. 3
  31. 31.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 1, 3
  32. 32.Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. High-resolution image synthesis and semantic manipulation with conditional gans. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8798–8807, 2018. 6
  33. 33.Greg CG Wei and Martin A Tanner. A monte carlo implementation of the em algorithm and the poor man’s data augmentation algorithms. Journal of the American statistical Association, 85(411):699–704, 1990. 3
  34. 34.Peter Welinder, Steve Branson, Takeshi Mita, Catherine Wah, Florian Schroff, Serge Belongie, and Pietro Perona. Caltech-ucsd birds 200. California Institute of Technology, 2010. 5, 6
  35. 35.Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with improved vqgan. arXiv preprint arXiv:2110.04627, 2021. 2
  36. 36.Yingchen Yu, Fangneng Zhan, Rongliang Wu, Jianxiong Pan, Kaiwen Cui, Shijian Lu, Feiying Ma, Xuansong Xie, and Chunyan Miao. Diverse image inpainting with bidirectional and autoregressive transformers. arXiv preprint arXiv:2104.12335, 2021. 3
  37. 37.Yingchen Yu, Fangneng Zhan, Rongliang Wu, Jiahui Zhang, Shijian Lu, Miaomiao Cui, Xuansong Xie, Xian-Sheng Hua, and Chunyan Miao. Towards counterfactual image manipulation via clip. In Proceedings of the 30th ACM International Conference on Multimedia, pages 3637–3645, 2022. 1
  38. 38.Fangneng Zhan, Yingchen Yu, Rongliang Wu, Jiahui Zhang, Kaiwen Cui, Changgong Zhang, and Shijian Lu. Autoregressive image synthesis with integrated quantization. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XVI, pages 110–127. Springer, 2022. 2, 3
  39. 39.Fangneng Zhan, Yingchen Yu, Rongliang Wu, Jiahui Zhang, and Shijian Lu. Multimodal image synthesis and editing: A survey. arXiv preprint arXiv:2112.13592, 2021. 1
  40. 40.Fangneng Zhan, Yingchen Yu, Rongliang Wu, Jiahui Zhang, Shijian Lu, and Changgong Zhang. Marginal contrastive correspondence for guided image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10663–10672, 2022. 5
  41. 41.Fangneng Zhan, Jiahui Zhang, Yingchen Yu, Rongliang Wu, and Shijian Lu. Modulated contrast for versatile image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18280–18290, 2022. 2
  42. 42.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 633–641, 2017. 1, 5, 6, 7

Citation

MLA
Zhang, J., et al. “Regularized Vector Quantization for Tokenized Image Synthesis”. arXiv, 2023, http://arxiv.org/abs/2303.06424v2.
APA
Zhang, J., Zhan, F., Theobalt, C., & Lu, S. (2023). Regularized Vector Quantization for Tokenized Image Synthesis. arXiv. http://arxiv.org/abs/2303.06424v2
Chicago
Zhang, J., F. Zhan, C. Theobalt, and S. Lu. 2023. “Regularized Vector Quantization for Tokenized Image Synthesis”. arXiv. http://arxiv.org/abs/2303.06424v2.
Harvard
Zhang, J. et al. (2023) “Regularized Vector Quantization for Tokenized Image Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.06424v2.
Vancouver
1. Zhang J, Zhan F, Theobalt C, Lu S (2023) Regularized Vector Quantization for Tokenized Image Synthesis. arXiv

BibTeX

@article{zhang2023regularized,
  title = {Regularized Vector Quantization for Tokenized Image Synthesis},
  author = {Zhang, Jiahui and Zhan, Fangneng and Theobalt, Christian and Lu, Shijian},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.06424v2},
  eprint = {2303.06424}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE