keyword
contrastive divergence
Contrastive divergence is an approximate learning algorithm used in machine learning to train energy-based probabilistic models, such as restricted Boltzmann machines and Markov random fields, by estimating the gradient of the log-likelihood function. In standard maximum likelihood estimation, calculating the gradient requires computing an intractable normalization term known as the partition function, which typically demands running Markov chain Monte Carlo simulations until they reach thermal equilibrium. Contrastive divergence accelerates this process by initializing the Markov chain directly at observed data points and running Gibbs sampling for only a small, fixed number of steps rather than until full convergence. The algorithm then updates the model parameters by measuring the difference between the feature correlations in the original data and those in the reconstructed samples, thereby lowering the energy assigned to observed data while raising the energy assigned to competing, model-generated states.
11 items

Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann, Aapo Hyvönen
Why you should read this
Introduces a novel estimation principle for unnormalized statistical models that avoids complex integration by treating parameter estimation as a classification task between data and artificial noise.
We present a new estimation principle for parameterized statistical models. The idea is to perform nonlinear logistic regression to discriminate between the observed data and some artificially generated noise, using the model log-density function in the regression nonlinearity. We show that this leads to a consistent (convergent) estimator of the parameters, and analyze the asymptotic variance. In particular, the method is shown to directly work for unnormalized models, i.e. models where the density function does not integrate to one. The normalization constant can be estimated just like any other parameter. For a tractable ICA model, we compare the method with other estimation methods that can be used to learn unnormalized models, including score matching, contrastive divergence, and maximum-likelihood where the normalization constant is estimated with importance sampling. Simulations show that noise-contrastive estimation offers the best trade-off between computational and statistical efficiency. The method is then applied to the modeling of natural images: We show that the method can successfully estimate a large-scale two-layer model and a Markov random field.
Added
2026-10-04
License
Published with permission

Restricted Boltzmann machines for collaborative filtering
Ruslan Salakhutdinov, Andriy Mnih, Geoffrey Hinton
Why you should read this
Demonstrates how Restricted Boltzmann Machines can be effectively adapted to massive, sparse recommendation datasets like the Netflix Prize, outperforming standard matrix factorization techniques and significantly boosting accuracy when blended in ensembles.
Most of the existing approaches to collaborative filtering cannot handle very large data sets. In this paper we show how a class of two-layer undirected graphical models, called Restricted Boltzmann Machines (RBM’s), can be used to model tabular data, such as user’s ratings of movies. We present efficient learning and inference procedures for this class of models and demonstrate that RBM’s can be successfully applied to the Netflix data set, containing over 100 million user/movie ratings. We also show that RBM’s slightly outperform carefully-tuned SVD models. When the predictions of multiple RBM models and multiple SVD models are linearly combined, we achieve an error rate that is well over 6% better than the score of Netflix’s own system.
Added
2026-09-27
License
Published with permission

Generative Flow Networks for Discrete Probabilistic Modeling
Dinghuai Zhang, Nikolay Malkin, Zhen Liu, Alexandra Volokhova, Aaron C. Courville, Yoshua Bengio
Why you should read this
Proposes energy-based generative flow networks to overcome the slow mixing of traditional MCMC methods in high-dimensional discrete spaces by jointly training an energy function alongside a generative policy that amortizes mode-hopping exploration.
We present energy-based generative flow networks (EB-GFN), a novel probabilistic modeling algorithm for high-dimensional discrete data. Building upon the theory of generative flow networks (GFlowNets; Bengio et al., 2021b), we model the generation process by a stochastic data construction policy and thus amortize expensive MCMC exploration into a fixed number of actions sampled from a GFlowNet. We show how GFlowNets can approximately perform large-block Gibbs sampling to mix between modes. We propose a framework to jointly train a GFlowNet with an energy function, so that the GFlowNet learns to sample from the energy distribution, while the energy learns with an approximate MLE objective with negative samples from the GFlowNet. We demonstrate EB-GFN's effectiveness on various probabilistic modeling tasks. Code is publicly available at github.com/zdhnarsil/EB_GFN.
Added
2026-09-26

Deep Gaussian Processes
Andreas C. Damianou, Neil D. Lawrence
Why you should read this
Introduces deep Gaussian processes and a variational inference framework that enables fully Bayesian hierarchical modeling and automated architecture selection on small datasets without overfitting.
In this paper we introduce deep Gaussian process (GP) models. Deep GPs are a deep belief network based on Gaussian process mappings. The data is modeled as the output of a multivariate GP. The inputs to that Gaussian process are then governed by another GP. A single layer model is equivalent to a standard GP or the GP latent variable model (GP-LVM). We perform inference in the model by approximate variational marginalization. This results in a strict lower bound on the marginal likelihood of the model which we use for model selection (number of layers and nodes per layer). Deep belief networks are typically applied to relatively large data sets using stochastic gradient descent for optimization. Our fully Bayesian treatment allows for the application of deep models even when data is scarce. Model selection by our variational bound shows that a five layer hierarchy is justified even when modelling a digit data set containing only 150 examples.
Added
2026-09-25

Multimodal learning with deep Boltzmann machines
Nitish Srivastava, Ruslan Salakhutdinov
Why you should read this
Proposes a Multimodal Deep Boltzmann Machine that learns a joint generative model across disparate modalities like images and text, enabling effective classification, cross-modal retrieval, and the reconstruction of missing inputs.
A Deep Boltzmann Machine is described for learning a generative model of data that consists of multiple and diverse input modalities. The model can be used to extract a unified representation that fuses modalities together. We find that this representation is useful for classification and information retrieval tasks. The model works by learning a probability density over the space of multimodal inputs. It uses states of latent variables as representations of the input. The model can extract this representation even when some modalities are absent by sampling from the conditional distribution over them and filling them in. Our experimental results on bi-modal data consisting of images and text show that the Multimodal DBM can learn a good generative model of the joint space of image and text inputs that is useful for information retrieval from both unimodal and multimodal queries. We further demonstrate that this model significantly outperforms SVMs and LDA on discriminative tasks. Finally, we compare our model to other deep learning methods, including autoencoders and deep belief networks, and show that it achieves noticeable gains.
Added
2026-09-18

Deep Canonical Correlation Analysis
Galen Andrew, Raman Arora, Jeff Bilmes, Karen Livescu
We introduce Deep Canonical Correlation Analysis (DCCA), a method to learn complex nonlinear transformations of two views of data such that the resulting representations are highly linearly correlated. Parameters of both transformations are jointly learned to maximize the (regularized) total correlation. It can be viewed as a nonlinear extension of the linear method canonical correlation analysis (CCA). It is an alternative to the nonparametric method kernel canonical correlation analysis (KCCA) for learning correlated nonlinear transformations. Unlike KCCA, DCCA does not require an inner product, and has the advantages of a parametric method: training time scales well with data size and the training data need not be referenced when computing the representations of unseen instances. In experiments on two real-world datasets, we find that DCCA learns representations with significantly higher correlation than those learned by CCA and KCCA. We also introduce a novel non-saturating sigmoid function based on the cube root that may be useful more generally in feedforward neural networks.
Added
2026-09-16

Estimation of Non-Normalized Statistical Models by Score Matching
Aapo Hyvärinen
One often wants to estimate statistical models where the probability density function is known only up to a multiplicative normalization constant. Typically, one then has to resort to Markov Chain Monte Carlo methods, or approximations of the normalization constant. Here, we propose that such models can be estimated by minimizing the expected squared distance between the gradient of the log-density given by the model and the gradient of the log-density of the observed data. While the estimation of the gradient of log-density function is, in principle, a very difficult non-parametric problem, we prove a surprising result that gives a simple formula for this objective function. The density function of the observed data does not appear in this formula, which simplifies to a sample average of a sum of some derivatives of the log-density given by the model. The validity of the method is demonstrated on multivariate Gaussian and independent component analysis models, and by estimating an overcomplete filter set for natural image data.
Added
2026-09-16

Deep Boltzmann Machines
Ruslan Salakhutdinov, Geoffrey Hinton
Why you should read this
Introduces a scalable learning and pre-training framework for Deep Boltzmann Machines that combines variational inference with persistent Markov chains to effectively train multi-layer undirected generative models with millions of parameters.
We present a new learning algorithm for Boltzmann machines that contain many layers of hidden variables. Data-dependent expectations are estimated using a variational approximation that tends to focus on a single mode, and data-independent expectations are approximated using persistent Markov chains. The use of two quite different techniques for estimating the two types of expectation that enter into the gradient of the log-likelihood makes it practical to learn Boltzmann machines with multiple hidden layers and millions of parameters. The learning can be made more efficient by using a layer-by-layer “pre-training” phase that allows variational inference to be initialized with a single bottom-up pass. We present results on the MNIST and NORB datasets showing that deep Boltzmann machines learn good generative models and perform well on handwritten digit and visual object recognition tasks.
Added
2026-09-15
License
Published with permission

Generative Modeling by Estimating Gradients of the Data Distribution
Yang Song, Stefano Ermon
Why you should read this
Proposes the Noise Conditional Score Network (NCSN), demonstrating that training a network to estimate gradients (scores) at multiple noise levels solves the manifold hypothesis problem.
We introduce a new generative model where samples are produced via Langevin dynamics using gradients of the data distribution estimated with score matching. Because gradients can be ill-defined and hard to estimate when the data resides on low-dimensional manifolds, we perturb the data with different levels of Gaussian noise, and jointly estimate the corresponding scores, i.e., the vector fields of gradients of the perturbed data distribution for all noise levels. For sampling, we propose an annealed Langevin dynamics where we use gradients corresponding to gradually decreasing noise levels as the sampling process gets closer to the data manifold. Our framework allows flexible model architectures, requires no sampling during training or the use of adversarial methods, and provides a learning objective that can be used for principled model comparisons. Our models produce samples comparable to GANs on MNIST, CelebA and CIFAR-10 datasets, achieving a new state-of-the-art inception score of 8.87 on CIFAR-10. Additionally, we demonstrate that our models learn effective representations via image inpainting experiments.
Added
2026-02-25

A Fast Learning Algorithm for Deep Belief Nets
Geoffrey E. Hinton, Simon Osindero, Yee‐Whye Teh
Why you should read this
Introduces a layer-wise unsupervised pre-training method that effectively initializes weights for deep architectures preventing early optimization stalls.
We show how to use "complementary priors" to eliminate the explaining-away effects that make inference difficult in densely connected belief nets that have many hidden layers. Using complementary priors, we derive a fast, greedy algorithm that can learn deep, directed belief networks one layer at a time, provided the top two layers form an undirected associative memory. The fast, greedy algorithm is used to initialize a slower learning procedure that fine-tunes the weights using a contrastive version of the wake-sleep algorithm. After fine-tuning, a network with three hidden layers forms a very good generative model of the joint distribution of handwritten digit images and their labels. This generative model gives better digit classification than the best discriminative learning algorithms. The low-dimensional manifolds on which the digits lie are modeled by long ravines in the free-energy landscape of the top-level associative memory, and it is easy to explore these ravines by using the directed connections to display what the associative memory has in mind.
Added
2026-02-21

Rectified Linear Units Improve Restricted Boltzmann Machines
Vinod Nair, Geoffrey E. Hinton
Why you should read this
Demonstrates how replacing binary hidden units with rectified linear units significantly improves feature learning in restricted Boltzmann machines for object recognition and face verification tasks.
Restricted Boltzmann machines were developed using binary stochastic hidden units. These can be generalized by replacing each binary unit by an infinite number of copies that all have the same weights but have progressively more negative biases. The learning and inference rules for these "Stepped Sigmoid Units" are unchanged. They can be approximated efficiently by noisy, rectified linear units. Compared with binary units, these units learn features that are better for object recognition on the NORB dataset and face verification on the Labeled Faces in the Wild dataset. Unlike binary units, rectified linear units preserve information about relative intensities as information travels through multiple layers of feature detectors.
Added
2025-09-24
License
Published with permission
