Built independently by an author, for readers. Read the story and support ChapterPal

keyword

contrastive divergence

Contrastive divergence is an approximate learning algorithm used in machine learning to train energy-based probabilistic models, such as restricted Boltzmann machines and Markov random fields, by estimating the gradient of the log-likelihood function. In standard maximum likelihood estimation, calculating the gradient requires computing an intractable normalization term known as the partition function, which typically demands running Markov chain Monte Carlo simulations until they reach thermal equilibrium. Contrastive divergence accelerates this process by initializing the Markov chain directly at observed data points and running Gibbs sampling for only a small, fixed number of steps rather than until full convergence. The algorithm then updates the model parameters by measuring the difference between the feature correlations in the original data and those in the reconstructed samples, thereby lowering the energy assigned to observed data while raising the energy assigned to competing, model-generated states.

11 items

Noise-contrastive estimation: A new estimation principle for unnormalized statistical models

Noise-contrastive estimation: A new estimation principle for unnormalized statistical models

Michael Gutmann, Aapo Hyvönen

OrganizationsHelsinki Institute for Information TechnologyUniversity of Helsinki

Why you should read this

Introduces a novel estimation principle for unnormalized statistical models that avoids complex integration by treating parameter estimation as a classification task between data and artificial noise.

We present a new estimation principle for parameterized statistical models. The idea is to perform nonlinear logistic regression to discriminate between the observed data and some artificially generated noise, using the model log-density function in the regression nonlinearity. We show that this leads to a consistent (convergent) estimator of the parameters, and analyze the asymptotic variance. In particular, the method is shown to directly work for unnormalized models, i.e. models where the density function does not integrate to one. The normalization constant can be estimated just like any other parameter. For a tractable ICA model, we compare the method with other estimation methods that can be used to learn unnormalized models, including score matching, contrastive divergence, and maximum-likelihood where the normalization constant is estimated with importance sampling. Simulations show that noise-contrastive estimation offers the best trade-off between computational and statistical efficiency. The method is then applied to the modeling of natural images: We show that the method can successfully estimate a large-scale two-layer model and a Markov random field.

Added

2026-10-04

License

Published with permission

Multimodal learning with deep Boltzmann machines

Multimodal learning with deep Boltzmann machines

Nitish Srivastava, Ruslan Salakhutdinov

OrganizationsUniversity of Toronto

Why you should read this

Proposes a Multimodal Deep Boltzmann Machine that learns a joint generative model across disparate modalities like images and text, enabling effective classification, cross-modal retrieval, and the reconstruction of missing inputs.

A Deep Boltzmann Machine is described for learning a generative model of data that consists of multiple and diverse input modalities. The model can be used to extract a unified representation that fuses modalities together. We find that this representation is useful for classification and information retrieval tasks. The model works by learning a probability density over the space of multimodal inputs. It uses states of latent variables as representations of the input. The model can extract this representation even when some modalities are absent by sampling from the conditional distribution over them and filling them in. Our experimental results on bi-modal data consisting of images and text show that the Multimodal DBM can learn a good generative model of the joint space of image and text inputs that is useful for information retrieval from both unimodal and multimodal queries. We further demonstrate that this model significantly outperforms SVMs and LDA on discriminative tasks. Finally, we compare our model to other deep learning methods, including autoencoders and deep belief networks, and show that it achieves noticeable gains.

Added

2026-09-18

Generative Modeling by Estimating Gradients of the Data Distribution

Generative Modeling by Estimating Gradients of the Data Distribution

Yang Song, Stefano Ermon

OrganizationsStanford University

Why you should read this

Proposes the Noise Conditional Score Network (NCSN), demonstrating that training a network to estimate gradients (scores) at multiple noise levels solves the manifold hypothesis problem.

We introduce a new generative model where samples are produced via Langevin dynamics using gradients of the data distribution estimated with score matching. Because gradients can be ill-defined and hard to estimate when the data resides on low-dimensional manifolds, we perturb the data with different levels of Gaussian noise, and jointly estimate the corresponding scores, i.e., the vector fields of gradients of the perturbed data distribution for all noise levels. For sampling, we propose an annealed Langevin dynamics where we use gradients corresponding to gradually decreasing noise levels as the sampling process gets closer to the data manifold. Our framework allows flexible model architectures, requires no sampling during training or the use of adversarial methods, and provides a learning objective that can be used for principled model comparisons. Our models produce samples comparable to GANs on MNIST, CelebA and CIFAR-10 datasets, achieving a new state-of-the-art inception score of 8.87 on CIFAR-10. Additionally, we demonstrate that our models learn effective representations via image inpainting experiments.

Added

2026-02-25

A Fast Learning Algorithm for Deep Belief Nets

A Fast Learning Algorithm for Deep Belief Nets

Geoffrey E. Hinton, Simon Osindero, Yee‐Whye Teh

OrganizationsNational University of SingaporeUniversity of Toronto

Why you should read this

Introduces a layer-wise unsupervised pre-training method that effectively initializes weights for deep architectures preventing early optimization stalls.

We show how to use "complementary priors" to eliminate the explaining-away effects that make inference difficult in densely connected belief nets that have many hidden layers. Using complementary priors, we derive a fast, greedy algorithm that can learn deep, directed belief networks one layer at a time, provided the top two layers form an undirected associative memory. The fast, greedy algorithm is used to initialize a slower learning procedure that fine-tunes the weights using a contrastive version of the wake-sleep algorithm. After fine-tuning, a network with three hidden layers forms a very good generative model of the joint distribution of handwritten digit images and their labels. This generative model gives better digit classification than the best discriminative learning algorithms. The low-dimensional manifolds on which the digits lie are modeled by long ravines in the free-energy landscape of the top-level associative memory, and it is easy to explore these ravines by using the directed connections to display what the associative memory has in mind.

Added

2026-02-21

Rectified Linear Units Improve Restricted Boltzmann Machines

Rectified Linear Units Improve Restricted Boltzmann Machines

Vinod Nair, Geoffrey E. Hinton

OrganizationsUniversity of Toronto

Why you should read this

Demonstrates how replacing binary hidden units with rectified linear units significantly improves feature learning in restricted Boltzmann machines for object recognition and face verification tasks.

Restricted Boltzmann machines were developed using binary stochastic hidden units. These can be generalized by replacing each binary unit by an infinite number of copies that all have the same weights but have progressively more negative biases. The learning and inference rules for these "Stepped Sigmoid Units" are unchanged. They can be approximated efficiently by noisy, rectified linear units. Compared with binary units, these units learn features that are better for object recognition on the NORB dataset and face verification on the Labeled Faces in the Wild dataset. Unlike binary units, rectified linear units preserve information about relative intensities as information travels through multiple layers of feature detectors.

Added

2025-09-24

License

Published with permission