Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Gibbs sampling

Gibbs sampling is a Markov chain Monte Carlo algorithm used to generate a sequence of random samples from a joint multivariate probability distribution when direct sampling is mathematically or computationally intractable. The algorithm operates iteratively by updating one variable or subset of variables at a time, drawing each new value directly from its conditional distribution given the current values of all remaining variables. Because full conditional distributions are frequently easier to compute and sample from than the complete joint distribution, Gibbs sampling reduces a complex multidimensional sampling task into a sequence of simpler, lower-dimensional steps. Over successive iterations, the resulting sequence of states forms a Markov chain whose stationary distribution converges to the target joint distribution, making it a widely utilized method for approximate statistical inference in Bayesian modeling and machine learning.

15 items

One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, Shi Pu, Yaole Wang, Gang Yue, Yue Cao, Hang Su, Jun Zhu

OrganizationsBeijing Academy of Artificial IntelligencePazhou Laboratory (Huangpu)Renmin University of ChinaShengshu TechnologyTsinghua University

Why you should read this

Proposes UniDiffuser, a single transformer-based framework that captures marginal, conditional, and joint multi-modal distributions to handle diverse generation tasks—including text-to-image, image-to-text, and paired generation—without requiring task-specific models or extra computational overhead.

This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is – learning diffusion models for marginal, conditional, and joint distributions can be unified as predicting the noise in the perturbed data, where the perturbation levels (i.e. timesteps) can be different for different modalities. Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model – perturbs data in all modalities instead of a single modality, inputs individual timesteps in different modalities, and predicts the noise of all modalities instead of a single modality. UniDiffuser is parameterized by a transformer for diffusion models to handle input types of different modalities. Implemented on large-scale paired image-text data, UniDiffuser is able to perform image, text, text-to-image, image-to-text, and image-text pair generation by setting proper timesteps without additional overhead. In particular, UniDiffuser is able to produce perceptually realistic samples in all tasks and its quantitative results (e.g., the FID and CLIP score) are not only superior to existing general-purpose models but also comparable to the bespoke models (e.g., Stable Diffusion and DALL·E 2) in representative tasks (e.g., text-to-image generation). Our code is available at https://github.com/thu-ml/unidiffuser.

Added

2026-09-28

The Author-Topic Model for Authors and Documents

The Author-Topic Model for Authors and Documents

Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers, Padhraic Smyth

OrganizationsStanford UniversityUniversity of California, Irvine

Why you should read this

Extends Latent Dirichlet Allocation by simultaneously modeling document content and authorship distributions, providing a probabilistic framework to discover research interests, measure author similarities, and analyze multi-authored texts.

We introduce the author-topic model, a generative model for documents that extends Latent Dirichlet Allocation (LDA; Blei, Ng, & Jordan, 2003) to include authorship information. Each author is associated with a multinomial distribution over topics and each topic is associated with a multinomial distribution over words. A document with multiple authors is modeled as a distribution over topics that is a mixture of the distributions associated with the authors. We apply the model to a collection of 1,700 NIPS conference papers and 160,000 CiteSeer abstracts. Exact inference is intractable for these datasets and we use Gibbs sampling to estimate the topic and author distributions. We compare the performance with two other generative models for documents, which are special cases of the author-topic model: LDA (a topic model) and a simple author model in which each author is associated with a distribution over words rather than a distribution over topics. We show topics recovered by the author-topic model, and demonstrate applications to computing similarity between authors and entropy of author output.

Added

2026-09-24

A Tutorial on Learning with Bayesian Networks

A Tutorial on Learning with Bayesian Networks

David Heckerman

Why you should read this

Explains the foundational principles and practical algorithms for learning both the structure and parameters of Bayesian networks from complete and incomplete data, bridging statistical inference, prior knowledge integration, and causal discovery.

A Bayesian network is a graphical model that encodes probabilistic relationships among variables of interest. When used in conjunction with statistical techniques, the graphical model has several advantages for data analysis. One, because the model encodes dependencies among all variables, it readily handles situations where some data entries are missing. Two, a Bayesian network can be used to learn causal relationships, and hence can be used to gain understanding about a problem domain and to predict the consequences of intervention. Three, because the model has both a causal and probabilistic semantics, it is an ideal representation for combining prior knowledge (which often comes in causal form) and data. Four, Bayesian statistical methods in conjunction with Bayesian networks offer an efficient and principled approach for avoiding the overfitting of data. In this paper, we discuss methods for constructing Bayesian networks from prior knowledge and summarize Bayesian statistical methods for using data to improve these models. With regard to the latter task, we describe methods for learning both the parameters and structure of a Bayesian network, including techniques for learning with incomplete data. In addition, we relate Bayesian-network methods for learning to techniques for supervised and unsupervised learning. We illustrate the graphical-modeling approach using a real-world case study.

Added

2026-09-13

An Introduction to Variational Methods for Graphical Models

An Introduction to Variational Methods for Graphical Models

MICHAEL I. JORDAN, ZOUBIN GHAHRAMANI, TOMMI S. JAAKKOLA, LAWRENCE K. SAUL

OrganizationsAT&T Labs—ResearchMassachusetts Institute of TechnologyUniversity College LondonUniversity of California Berkeley

Why you should read this

Establishes a foundational framework for approximate inference and learning in graphical models by using convex duality to convert intractable probabilistic calculations into tractable optimization problems.

This paper presents a tutorial introduction to the use of variational methods for inference and learning in graphical models (Bayesian networks and Markov random fields). We present a number of examples of graphical models, including the QMR-DT database, the sigmoid belief network, the Boltzmann machine, and several variants of hidden Markov models, in which it is infeasible to run exact inference algorithms. We then introduce variational methods, which exploit laws of large numbers to transform the original graphical model into a simplified graphical model in which inference is efficient. Inference in the simplified model provides bounds on probabilities of interest in the original model. We describe a general framework for generating variational transformations based on convex duality. Finally we return to the examples and demonstrate how variational algorithms can be formulated in each case.

Added

2026-09-10

The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo

The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo

Matthew D. Hoffman, Andrew Gelman

OrganizationsColumbia University

Why you should read this

Introduces the No-U-Turn Sampler (NUTS), an algorithm that automates path length selection and step-size adaptation in Hamiltonian Monte Carlo to enable efficient, hands-free sampling for high-dimensional Bayesian models.

Hamiltonian Monte Carlo (HMC) is a Markov chain Monte Carlo (MCMC) algorithm that avoids the random walk behavior and sensitivity to correlated parameters that plague many MCMC methods by taking a series of steps informed by first-order gradient information. These features allow it to converge to high-dimensional target distributions much more quickly than simpler methods such as random walk Metropolis or Gibbs sampling. However, HMC's performance is highly sensitive to two user-specified parameters: a step size {\epsilon} and a desired number of steps L. In particular, if L is too small then the algorithm exhibits undesirable random walk behavior, while if L is too large the algorithm wastes computation. We introduce the No-U-Turn Sampler (NUTS), an extension to HMC that eliminates the need to set a number of steps L. NUTS uses a recursive algorithm to build a set of likely candidate points that spans a wide swath of the target distribution, stopping automatically when it starts to double back and retrace its steps. Empirically, NUTS perform at least as efficiently as and sometimes more efficiently than a well tuned standard HMC method, without requiring user intervention or costly tuning runs. We also derive a method for adapting the step size parameter {\epsilon} on the fly based on primal-dual averaging. NUTS can thus be used with no hand-tuning at all. NUTS is also suitable for applications such as BUGS-style automatic inference engines that require efficient "turnkey" sampling algorithms.

Added

2026-09-09

Creative Commons License
A Fast Learning Algorithm for Deep Belief Nets

A Fast Learning Algorithm for Deep Belief Nets

Geoffrey E. Hinton, Simon Osindero, Yee‐Whye Teh

OrganizationsNational University of SingaporeUniversity of Toronto

Why you should read this

Introduces a layer-wise unsupervised pre-training method that effectively initializes weights for deep architectures preventing early optimization stalls.

We show how to use "complementary priors" to eliminate the explaining-away effects that make inference difficult in densely connected belief nets that have many hidden layers. Using complementary priors, we derive a fast, greedy algorithm that can learn deep, directed belief networks one layer at a time, provided the top two layers form an undirected associative memory. The fast, greedy algorithm is used to initialize a slower learning procedure that fine-tunes the weights using a contrastive version of the wake-sleep algorithm. After fine-tuning, a network with three hidden layers forms a very good generative model of the joint distribution of handwritten digit images and their labels. This generative model gives better digit classification than the best discriminative learning algorithms. The low-dimensional manifolds on which the digits lie are modeled by long ravines in the free-energy landscape of the top-level associative memory, and it is easy to explore these ravines by using the directed connections to display what the associative memory has in mind.

Added

2026-02-21