keyword
probabilistic contrastive loss
A probabilistic contrastive loss is a machine learning objective function used in representation learning that distinguishes true matching data points from unrelated negative samples by modeling their relative association as a probability distribution. Typically formulated through noise-contrastive estimation and categorical cross-entropy over similarity scores, this loss optimizes the model to assign high probability to related pairs—such as temporally continuous features, augmented views, or target discrete tokens—while minimizing the probability assigned to randomly sampled negatives. By framing the contrastive discrimination task in terms of density ratios or normalized probability distributions, the objective maximizes a lower bound on the mutual information between shared representations, allowing neural networks to extract compact, high-level features and maintain calibrated latent spaces without requiring explicit reconstruction of high-dimensional inputs.
2 items

Regularized Vector Quantization for Tokenized Image Synthesis
Jiahui Zhang, Fangneng Zhan, Christian Theobalt, Shijian Lu
Why you should read this
Proposes a dual-regularized vector quantization framework with a probabilistic contrastive loss that prevents codebook collapse and aligns training with stochastic sampling for superior image synthesis in autoregressive and diffusion models.
Quantizing images into discrete representations has been a fundamental problem in unified generative modeling. Predominant approaches learn the discrete representation either in a deterministic manner by selecting the best-matching token or in a stochastic manner by sampling from a predicted distribution. However, deterministic quantization suffers from severe codebook collapse and misalignment with inference stage while stochastic quantization suffers from low codebook utilization and perturbed reconstruction objective. This paper presents a regularized vector quantization framework that allows to mitigate above issues effectively by applying regularization from two perspectives. The first is a prior distribution regularization which measures the discrepancy between a prior token distribution and the predicted token distribution to avoid codebook collapse and low codebook utilization. The second is a stochastic mask regularization that introduces stochasticity during quantization to strike a good balance between inference stage misalignment and unperturbed reconstruction objective. In addition, we design a probabilistic contrastive loss which serves as a calibrated metric to further mitigate the perturbed reconstruction objective. Extensive experiments show that the proposed quantization framework outperforms prevailing vector quantization methods consistently across different generative models including auto-regressive models and diffusion models.
Added
2026-09-26

Representation Learning with Contrastive Predictive Coding
Aäron van den Oord, Yazhe Li, Oriol Vinyals
Why you should read this
Demonstrates how a single unsupervised learning method can extract meaningful representations across completely different data types—speech, images, text, and video—by predicting future information in a compressed latent space rather than trying to reconstruct raw data.
While supervised learning has enabled great progress in many applications, unsupervised learning has not seen such widespread adoption, and remains an important and challenging endeavor for artificial intelligence. In this work, we propose a universal unsupervised learning approach to extract useful representations from high-dimensional data, which we call Contrastive Predictive Coding. The key insight of our model is to learn such representations by predicting the future in latent space by using powerful autoregressive models. We use a probabilistic contrastive loss which induces the latent space to capture information that is maximally useful to predict future samples. It also makes the model tractable by using negative sampling. While most prior work has focused on evaluating representations for a particular modality, we demonstrate that our approach is able to learn useful representations achieving strong performance on four distinct domains: speech, images, text and reinforcement learning in 3D environments.
Added
2026-03-25
License
Published with permission
