Built independently by an author, for readers. Read the story and support ChapterPal

keyword

tokenized image synthesis

Tokenized image synthesis is a generative modeling paradigm in which images are created by representing visual data as a collection of discrete tokens drawn from a learned codebook and using generative models to predict these tokens before reconstructing them into pixels. In this framework, an encoder and a vector quantization mechanism compress continuous visual features into a grid or sequence of discrete categorical indices. A downstream generative architecture, such as an autoregressive transformer, a masked generative model, or a discrete diffusion model, is then trained to learn the probability distribution of these tokens and generate new valid sequences. Finally, a paired decoder maps the predicted token representations back into high-fidelity continuous images, adapting discrete sequence modeling techniques from natural language processing to visual generation tasks.

1 item

Regularized Vector Quantization for Tokenized Image Synthesis

Regularized Vector Quantization for Tokenized Image Synthesis

Jiahui Zhang, Fangneng Zhan, Christian Theobalt, Shijian Lu

OrganizationsMax Planck Institute for InformaticsNanyang Technological University

Why you should read this

Proposes a dual-regularized vector quantization framework with a probabilistic contrastive loss that prevents codebook collapse and aligns training with stochastic sampling for superior image synthesis in autoregressive and diffusion models.

Quantizing images into discrete representations has been a fundamental problem in unified generative modeling. Predominant approaches learn the discrete representation either in a deterministic manner by selecting the best-matching token or in a stochastic manner by sampling from a predicted distribution. However, deterministic quantization suffers from severe codebook collapse and misalignment with inference stage while stochastic quantization suffers from low codebook utilization and perturbed reconstruction objective. This paper presents a regularized vector quantization framework that allows to mitigate above issues effectively by applying regularization from two perspectives. The first is a prior distribution regularization which measures the discrepancy between a prior token distribution and the predicted token distribution to avoid codebook collapse and low codebook utilization. The second is a stochastic mask regularization that introduces stochasticity during quantization to strike a good balance between inference stage misalignment and unperturbed reconstruction objective. In addition, we design a probabilistic contrastive loss which serves as a calibrated metric to further mitigate the perturbed reconstruction objective. Extensive experiments show that the proposed quantization framework outperforms prevailing vector quantization methods consistently across different generative models including auto-regressive models and diffusion models.

Added

2026-09-26