SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization

Yuhta TakidaTakashi ShibuyaWei-Hsiang LiaoChieh-Hsin LaiJunki OhmuraToshimitsu UesakaNaoki MurataShusuke TakahashiToshiyuki KumakuraYuki Mitsufuji

article2022ICML113 citations

Proposes a principled variational Bayes framework for discrete latent representations that eliminates codebook collapse in VQ-VAEs through self-annealing stochastic quantization without relying on heuristic tricks.

Listen

Generative modeling commonly relies on vector-quantized variational autoencoders to represent complex data into compact, discrete codes. However, these systems frequently suffer from codebook collapse, a failure mode where the model utilizes only a tiny fraction of its available representational capacity. To bypass this issue, standard approaches rely on brittle training heuristics and complex workarounds, including manual codebook resets and extensive hyperparameter tuning. These challenges motivate the need for a principled, self-adjusting framework that reliably maximizes code capacity without manual interventions.

The article demonstrates and evaluates a framework called the stochastically quantized variational autoencoder (SQ-VAE). Its primary objective is to replace heuristic training with a mathematically grounded variational Bayesian formulation that incorporates stochastic quantization and self-annealing to ensure high codebook utilization.

To evaluate this framework, the authors conducted extensive empirical experiments spanning diverse vision and speech tasks. The datasets examined include standard image benchmarks (MNIST, Fashion-MNIST, CIFAR10, CelebA, and CelebAHQ) as well as speech corpora (VCTK and ZeroSpeech 2019). The approach introduces trainable probability distributions—specifically Gaussian distributions for continuous data and von Mises–Fisher distributions for categorical data—that enable the model to start with broad stochastic exploration and naturally transition into deterministic encoding as training converges.

The findings show that SQ-VAE consistently outperforms standard vector-quantized autoencoders across multiple domains. First, SQ-VAE achieves substantially higher codebook utilization (measured by perplexity) and lower reconstruction errors; for instance, on categorical face mask data, it cut pixel error roughly in half from 6.95% down to 3.51% while improving intersection-over-union scores from 59.7% to 74.6%. Second, the architecture scales predictably: increasing codebook size or dimensionality directly improves performance in SQ-VAE, whereas conventional models show little benefit and remain prone to collapse. Third, on continuous image benchmarks, SQ-VAE lowered reconstruction mean squared error from 1.33 to under 0.98 on CelebA and delivered superior image generation metrics. Finally, speech experiments confirmed clear gains in spectrogram reconstruction accuracy over baseline models.

These results demonstrate that generative pipelines can eliminate ad-hoc heuristics and fragile manual tuning without sacrificing fidelity. By stabilizing latent representations, SQ-VAE reduces engineering overhead, streamlines model deployment, and provides a dependable path toward data compression and multi-modal representation learning.

Organizations developing generative models or discrete compression systems should consider SQ-VAE as a drop-in replacement for standard vector-quantized architectures. Before production rollout for specialized downstream tasks, teams should conduct domain-specific validation; while spectrogram reconstruction improved in speech evaluations, discriminative phonetic unit discovery scores remained on par with baselines, indicating that representation quality depends on the end objective.

The article's conclusions are strongly supported across multiple random seeds, image scales, and modalities. However, minor limitations remain regarding specific hyperparameter variants in variable-length sequential data, as certain complex variance parameterizations showed training instability on speech and large-scale vision tasks.

arXiv: 2205.07547sony/sq-vae
  • Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). Read the original VQ-VAE first to understand the codebook, straight-through estimator, and training heuristics that SQ-VAE revisits to address codebook collapse.
  • Paper: An Introduction to Variational Autoencoders, Diederik P. Kingma et al. (2019). Its derivation of the VAE objective and reparameterized training provides the variational framework that SQ-VAE extends with stochastic quantization.
Cover for SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization

Abstract

One noted issue of vector-quantized variational autoencoder (VQ-VAE) is that the learned discrete representation uses only a fraction of the full capacity of the codebook, also known as codebook collapse. We hypothesize that the training scheme of VQ-VAE, which involves some carefully designed heuristics, underlies this issue. In this paper, we propose a new training scheme that extends the standard VAE via novel stochastic dequantization and quantization, called stochastically quantized variational autoencoder (SQ-VAE). In SQ-VAE, we observe a trend that the quantization is stochastic at the initial stage of the training but gradually converges toward a deterministic quantization, which we call self-annealing. Our experiments show that SQ-VAE improves codebook utilization without using common heuristics. Furthermore, we empirically show that SQ-VAE is superior to VAE and VQ-VAE in vision- and speech-related tasks.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 3. Stochastically Quantized VAE
  • 3.1. Overview of SQ-VAE
  • 3.2. Gaussian SQ-VAE
  • 3.3. Self-annealed Quantization
  • 3.4. vMF SQ-VAE for Categorical Distributions
  • 4. Related Work
  • 5. Experiments
  • 5.1. Continuous Data Distribution
  • 5.2. Categorical Distributions
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Notations and Definitions
  • B. Derivation Details
  • B.1. Gaussian SQ-VAE
  • B.2. vMF SQ-VAE
  • C. Training Procedures of SQ-VAEs
  • D. Self-Annealed Quantization
  • D.1. Similarity between SQ-VAE and conventional VAE
  • D.2. Proof of Proposition 1
  • D.3. Details of Section 3.3
  • D.4. Case of vMF SQ-VAE
  • E. Experimental Details
  • E.1. Gaussian SQ-VAE on Image Datasets
  • E.1.1. DATASETS AND PREPROCESSING
  • E.1.2. MODEL DESCRIPTION AND TRAINING
  • E.1.3. RECONSTRUCTED AND GENERATED SAMPLES ON CELEBA 64 × 64
  • E.2. Gaussian SQ-VAE on Speech Dataset
  • E.2.1. DATASETS AND PREPROCESSING
  • E.2.2. MODEL DESCRIPTION AND TRAINING
  • E.2.3. DETAILS OF EXPERIMENTAL RESULTS
  • E.2.4. RECONSTRUCTED SAMPLES
  • E.2.5. ACOUSTIC UNIT DISCOVERY
  • E.3. vMF SQ-VAE on Vision Dataset
  • E.3.1. DATASETS AND PREPROCESSING
  • F. Experiments on CelebA HQ 256 × 256

Knowls

  1. Knowl 1 — Unreadable source: no paper content was accessible for knowledge extraction

    limitation

    No body text, equations, tables, or figures of the paper were provided in a readable form; only the file identifier ff10a519-ecc6-4be7-a4d3-983457f14c4b.pdf was available. Consequently, the paper's methods, models, theoretical results, experimental setups, empirical results, and stated limitations could not be identified, and no knowledge grains can be faithfully extracted from its contribution. Any knowl produced in this situation would be fabrication rather than extraction, so none is reported.

Coverage note — The attached file arrived as a bare filename with no readable text, figures, or tables, so no genuine contributed content of the paper could be extracted; no domain material was deliberately omitted — it was simply never accessible.

Citation

MLA
Takida, Y., et al. “SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization”. International Conference on Machine Learning, vol. 162, 2022, pp. 20987–1012, https://proceedings.mlr.press/v162/takida22a.html.
APA
Takida, Y., Shibuya, T., Liao, W., Lai, C.-H., Ohmura, J., Uesaka, T., Murata, N., Takahashi, S., Kumakura, T., & Mitsufuji, Y. (2022). SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization. International Conference on Machine Learning, 162, 20987–21012. https://proceedings.mlr.press/v162/takida22a.html
Chicago
Takida, Y., T. Shibuya, W. Liao, et al. 2022. “SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization”. International Conference on Machine Learning 162: 20987–21012. https://proceedings.mlr.press/v162/takida22a.html.
Harvard
Takida, Y. et al. (2022) “SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization”, International Conference on Machine Learning. PMLR, pp. 20987–21012. Available at: https://proceedings.mlr.press/v162/takida22a.html.
Vancouver
1. Takida Y, Shibuya T, Liao W, Lai C-H, Ohmura J, Uesaka T, Murata N, Takahashi S, Kumakura T, Mitsufuji Y (2022) SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization. In: International Conference on Machine Learning. PMLR, pp 20987–21012

BibTeX

@InProceedings{pmlr-v162-takida22a,
  title = 	 {{SQ}-{VAE}: Variational {B}ayes on Discrete Representation with Self-annealed Stochastic Quantization},
  author =       {Takida, Yuhta and Shibuya, Takashi and Liao, Weihsiang and Lai, Chieh-Hsin and Ohmura, Junki and Uesaka, Toshimitsu and Murata, Naoki and Takahashi, Shusuke and Kumakura, Toshiyuki and Mitsufuji, Yuki},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {20987--21012},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/takida22a/takida22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/takida22a.html},
  abstract = 	 {One noted issue of vector-quantized variational autoencoder (VQ-VAE) is that the learned discrete representation uses only a fraction of the full capacity of the codebook, also known as codebook collapse. We hypothesize that the training scheme of VQ-VAE, which involves some carefully designed heuristics, underlies this issue. In this paper, we propose a new training scheme that extends the standard VAE via novel stochastic dequantization and quantization, called stochastically quantized variational autoencoder (SQ-VAE). In SQ-VAE, we observe a trend that the quantization is stochastic at the initial stage of the training but gradually converges toward a deterministic quantization, which we call self-annealing. Our experiments show that SQ-VAE improves codebook utilization without using common heuristics. Furthermore, we empirically show that SQ-VAE is superior to VAE and VQ-VAE in vision- and speech-related tasks.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/