Built independently by an author, for readers. Read the story and support ChapterPal

keyword

variational inference

Variational inference is a computational method in Bayesian statistics and machine learning that approximates complex, intractable probability distributions by transforming the inference problem into an optimization task. Instead of generating random samples from a target posterior distribution, as in Markov Chain Monte Carlo techniques, variational inference selects a family of tractable probability distributions and optimizes their parameters to find the candidate closest to the true distribution. This closeness is typically measured by minimizing the Kullback-Leibler divergence to the target posterior, which is mathematically equivalent to maximizing the evidence lower bound on the marginal likelihood of the observed data. Because it relies on standard continuous optimization techniques such as gradient descent, variational inference scales effectively to large datasets and complex probabilistic models, making it widely used for training latent variable models, deep generative architectures, and Bayesian neural networks.

63 items

Model-Based Imitation Learning for Urban Driving

Model-Based Imitation Learning for Urban Driving

Anthony Hu, Gianluca Corrado, Nicolas Griffiths, Zachary Murez, Corina Gurau, Hudson Yeo, Alex Kendall, Roberto Cipolla, Jamie Shotton

Why you should read this

Proposes MILE, a camera-only model-based imitation learning framework that jointly models world dynamics and driving policies entirely from offline demonstration videos, achieving state-of-the-art closed-loop urban driving in CARLA by planning actions purely within imagined latent rollouts.

An accurate model of the environment and the dynamic agents acting in it offers great potential for improving motion planning. We present MILE: a Model-based Imitation LEarning approach to jointly learn a model of the world and a policy for autonomous driving. Our method leverages 3D geometry as an inductive bias and learns a highly compact latent space directly from high-resolution videos of expert demonstrations. Our model is trained on an offline corpus of urban driving data, without any online interaction with the environment. MILE improves upon prior state-of-the-art by 31% in driving score on the CARLA simulator when deployed in a completely new town and new weather conditions. Our model can predict diverse and plausible states and actions, that can be interpretably decoded to bird’s-eye view semantic segmentation. Further, we demonstrate that it can execute complex driving manoeuvres from plans entirely predicted in imagination. Our approach is the first camera-only method that models static scene, dynamic scene, and ego-behaviour in an urban driving environment. The code and model weights are available at https://github.com/wayveai/mile.

Added

2026-10-05

Temporal-Difference Variational Continual Learning

Temporal-Difference Variational Continual Learning

Luckeciano Carvalho Melo, Alessandro Abate, Yarin Gal

OrganizationsUniversity of Oxford

Why you should read this

Introduces a temporal-difference-inspired variational continual learning objective that regularizes model updates using multiple past posterior estimates to prevent compounding approximation errors and reduce catastrophic forgetting.

Machine Learning models in real-world applications must continuously learn new tasks to adapt to shifts in the data-generating distribution. Yet, for Continual Learning (CL), models often struggle to balance learning new tasks (plasticity) with retaining previous knowledge (memory stability). Consequently, they are susceptible to Catastrophic Forgetting, which degrades performance and undermines the reliability of deployed systems. In the Bayesian CL literature, variational methods tackle this challenge by employing a learning objective that recursively updates the posterior distribution while constraining it to stay close to its previous estimate. Nonetheless, we argue that these methods may be ineffective due to compounding approximation errors over successive recursions. To mitigate this, we propose new learning objectives that integrate the regularization effects of multiple previous posterior estimations, preventing individual errors from dominating future posterior updates and compounding over time. We reveal insightful connections between these objectives and Temporal-Difference methods, a popular learning mechanism in Reinforcement Learning and Neuroscience. Experiments on challenging CL benchmarks show that our approach effectively mitigates Catastrophic Forgetting, outperforming strong Variational CL methods.

Added

2026-10-05

Improved off-policy training of diffusion samplers

Improved off-policy training of diffusion samplers

Marcin Sendera, Minsu Kim, Sarthak Mittal, Pablo Lemos, Luca Scimeca, Jarrid Rector-Brooks, Alexandre Adam, Yoshua Bengio, Nikolay Malkin

OrganizationsCIELA InstituteCIFARDreamfoldJagiellonian UniversityKorea Advanced Institute of Science and TechnologyMilaUniversité de MontréalUniversity of Edinburgh

Why you should read this

Presents a unified benchmark and an effective replay-buffer exploration strategy for training diffusion models to sample from unnormalized energy densities via continuous generative flow networks.

We study the problem of training diffusion models to sample from a distribution with a given unnormalized density or energy function. We benchmark several diffusion-structured inference methods, including simulation-based variational approaches and off-policy methods (continuous generative flow networks). Our results shed light on the relative advantages of existing algorithms while bringing into question some claims from past work. We also propose a novel exploration strategy for off-policy methods, based on local search in the target space with the use of a replay buffer, and show that it improves the quality of samples on a variety of target distributions. Our code for the sampling methods and benchmarks studied is made public at (link) as a base for future work on diffusion models for amortized inference.

Added

2026-10-04

GFlowNet-EM for Learning Compositional Latent Variable Models

GFlowNet-EM for Learning Compositional Latent Variable Models

Edward J. Hu, Nikolay Malkin, Moksh Jain, Katie E. Everett, Alexandros Graikos, Yoshua Bengio

OrganizationsCIFARGoogleMassachusetts Institute of TechnologyMilaStony Brook UniversityUniversité de Montréal

Why you should read this

Proposes GFlowNet-EM, a framework that replaces the intractable expectation-maximization E-step with an amortized GFlowNet sampler to train expressive latent variable models over discrete compositional structures without imposing restrictive independence assumptions.

Latent variable models (LVMs) with discrete compositional latents are an important but challenging setting due to a combinatorially large number of possible configurations of the latents. A key tradeoff in modeling the posteriors over latents is between expressivity and tractable optimization. For algorithms based on expectation-maximization (EM), the E-step is often intractable without restrictive approximations to the posterior. We propose the use of GFlowNets, algorithms for sampling from an unnormalized density by learning a stochastic policy for sequential construction of samples, for this intractable E-step. By training GFlowNets to sample from the posterior over latents, we take advantage of their strengths as amortized variational inference algorithms for complex distributions over discrete structures. Our approach, GFlowNet-EM, enables the training of expressive LVMs with discrete compositional latents, as shown by experiments on non-context-free grammar induction and on images using discrete variational autoencoders (VAEs) without conditional independence enforced in the encoder.

Added

2026-10-04

Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

Aaron J. Havens, Benjamin Kurt Miller, Bing Yan, Carles Domingo-Enrich, Anuroop Sriram, Daniel S. Levine, Brandon M. Wood, Bin Hu, Brandon Amos, Brian Karrer, Xiang Fu, Guan-Horng Liu, Ricky T. Q. Chen

OrganizationsMetaMicrosoftNew York UniversityUniversity of Illinois Urbana-Champaign

Why you should read this

Introduces a scalable stochastic optimal control framework that trains diffusion samplers from unnormalized densities with dramatically fewer expensive energy evaluations, enabling efficient amortized molecular conformer generation.

We introduce Adjoint Sampling, a highly scalable and efficient algorithm for learning diffusion processes that sample from unnormalized densities, or energy functions. It is the first on-policy approach that allows significantly more gradient updates than the number of energy evaluations and model samples, allowing us to scale to much larger problem settings than previously explored by similar methods. Our framework is theoretically grounded in stochastic optimal control and shares the same theoretical guarantees as Adjoint Matching, being able to train without the need for corrective measures that push samples towards the target distribution. We show how to incorporate key symmetries, as well as periodic boundary conditions, for modeling molecules in both cartesian and torsional coordinates. We demonstrate the effectiveness of our approach through extensive experiments on classical energy functions, and further scale up to neural network-based energy models where we perform amortized conformer generation across many molecular systems. To encourage further research in developing highly scalable sampling methods, we plan to open source these challenging benchmarks, where successful methods can directly impact progress in computational chemistry. Code & and Benchmarks provided at github.com/facebookresearch/adjoint_sampling.

Added

2026-10-03

NETS: A Non-equilibrium Transport Sampler

NETS: A Non-equilibrium Transport Sampler

Michael Samuel Albergo, Eric Vanden-Eijnden

OrganizationsCapital Fund ManagementHarvard UniversityNew York UniversitySociety of FellowsThe NSF Institute for Artificial Intelligence and Fundamental Interactions

Why you should read this

Introduces an unbiased sampling framework for unnormalized target distributions that augments non-equilibrium Langevin dynamics with a learned drift field trained without backpropagation to minimize importance weight variance.

We introduce the Non-Equilibrium Transport Sampler (NETS), an algorithm for sampling from unnormalized probability distributions. NETS builds on non-equilibrium sampling strategies that transport a simple base distribution into the target distribution in finite time, as pioneered in Neal’s annealed importance sampling (AIS). In the continuous-time setting, this transport is accomplished by evolving walkers using Langevin dynamics with a time-dependent potential, while simultaneously evolving importance weights to debias their solutions following Jarzynski’s equality. The key innovation of NETS is to add to the dynamics a learned drift term that offsets the need for these corrective weights by minimizing their variance through an objective that can be estimated without backpropagation and provably bounds the Kullback-Leibler divergence between the estimated and target distributions. NETS provides unbiased samples and features a tunable diffusion coefficient that can be adjusted after training to maximize the effective sample size. In experiments on standard benchmarks, high-dimensional Gaussian mixtures, and statistical lattice field theory models, NETS shows compelling performances.

Added

2026-10-03

An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference

An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference

Jeremias Knoblauch, Jack Jewson, Theodoros Damoulas

OrganizationsThe Alan Turing InstituteUniversity of Warwick

Why you should read this

Presents Generalized Variational Inference, a modular optimization framework that extends standard Bayesian updating to handle misspecified priors, misspecified likelihoods, and computational constraints in deep probabilistic models.

We advocate an optimization-centric view of Bayesian inference. Our inspiration is the representation of Bayes’ rule as infinite-dimensional optimization (Csiszár, 1975; Donsker and Varadhan, 1975; Zellner, 1988). Equipped with this perspective, we study Bayesian inference when one does not have access to (1) well-specified priors, (2) well-specified likelihoods, (3) infinite computing power. While these three assumptions underlie the standard Bayesian paradigm, they are typically inappropriate for modern Machine Learning applications. We propose addressing this through an optimization-centric generalization of Bayesian posteriors that we call the Rule of Three (RoT). The RoT can be justified axiomatically and recovers Bayesian, PAC-Bayesian and VI posteriors as special cases. While the RoT is primarily a conceptual and theoretical device, it also encompasses a novel sub-class of tractable posteriors which we call Generalized Variational Inference (GVI) posteriors. Just as the RoT, GVI posteriors are specified by three arguments: a loss, a divergence and a variational family. They also possess a number of desirable properties, including modularity, Frequentist consistency and an interpretation as approximate ELBO. We explore applications of GVI posteriors, and show that they can be used to improve robustness and posterior marginals on Bayesian Neural Networks and Deep Gaussian Processes.

Added

2026-10-01

Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal Reasoning

Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal Reasoning

Wenhao Ding, Haohong Lin, Bo Li, Ding Zhao

OrganizationsCarnegie Mellon UniversityUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes GRADER, a goal-conditioned reinforcement learning framework that treats causal graphs as latent variables to jointly discover environment causal structure and train generalizable, interpretable policies with theoretical performance guarantees.

As a pivotal component to attaining generalizable solutions in human intelligence, reasoning provides great potential for reinforcement learning (RL) agents’ generalization towards varied goals by summarizing part-to-whole arguments and discovering cause-and-effect relations. However, how to discover and represent causalities remains a huge gap that hinders the development of causal RL. In this paper, we augment Goal-Conditioned RL (GCRL) with Causal Graph (CG), a structure built upon the relation between objects and events. We novelly formulate the GCRL problem into variational likelihood maximization with CG as latent variables. To optimize the derived objective, we propose a framework with theoretical performance guarantees that alternates between two steps: using interventional data to estimate the posterior of CG; using CG to learn generalizable models and interpretable policies. Due to the lack of public benchmarks that verify generalization capability under reasoning, we design nine tasks and then empirically show the effectiveness of the proposed method against five baselines on these tasks. Further theoretical analysis shows that our performance improvement is attributed to the virtuous cycle of causal discovery, transition modeling, and policy training, which aligns with the experimental evidence in extensive ablation studies. Code is available on https://github.com/GilgameshD/GRADER.

Added

2026-09-26

Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

Juncheng Li, Junlin Xie, Long Qian, Linchao Zhu, Siliang Tang, Fei Wu, Yi Yang, Yueting Zhuang, Xin Eric Wang

OrganizationsUniversity of California, Santa CruzUniversity of Technology SydneyZhejiang University

Why you should read this

Introduces compositional temporal grounding benchmarks alongside a hierarchical variational cross-graph reasoning framework that aligns multi-level video and language structures to accurately localize video segments described by novel word combinations.

Temporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language descriptions, temporal grounding allows activity grounding beyond pre-defined classes and has received increasing attention in recent years. The semantic diversity is rooted in the principle of compositionality in linguistics, where novel semantics can be systematically described by combining known words in novel ways (compositional generalization). However, current temporal grounding datasets do not specifically test for the compositional generalizability. To systematically measure the compositional generalizability of temporal grounding models, we introduce a new Compositional Temporal Grounding task and construct two new dataset splits, i.e., Charades-CG and ActivityNet-CG. Evaluating the state-of-the-art methods on our new dataset splits, we empirically find that they fail to generalize to queries with novel combinations of seen words. To tackle this challenge, we propose a variational cross-graph reasoning framework that explicitly decomposes video and language into multiple structured hierarchies and learns fine-grained semantic correspondence among them. Experiments illustrate the superior compositional generalizability of our approach. The repository of this work is at https://github.com/YYJMJC/Compositional-Temporal-Grounding.

Added

2026-09-26

Joint Bayesian Inference of Graphical Structure and Parameters with a Single Generative Flow Network

Joint Bayesian Inference of Graphical Structure and Parameters with a Single Generative Flow Network

Tristan Deleu, Mizu Nishikawa-Toomey, Jithendaraa Subramanian, Nikolay Malkin, Laurent Charlin, Yoshua Bengio

OrganizationsHEC MontréalMcGill UniversityMila – Québec Artificial Intelligence InstituteUniversité de Montréal

Why you should read this

Presents JSP-GFN, a single Generative Flow Network that jointly infers both the graph structure and continuous parameters of Bayesian networks through a two-phase sampling process, enabling tractable posterior inference for expressive non-linear models without intractable marginalizations.

Generative Flow Networks (GFlowNets), a class of generative models over discrete and structured sample spaces, have been previously applied to the problem of inferring the marginal posterior distribution over the directed acyclic graph (DAG) of a Bayesian Network, given a dataset of observations. Based on recent advances extending this framework to non-discrete sample spaces, we propose in this paper to approximate the joint posterior over not only the structure of a Bayesian Network, but also the parameters of its conditional probability distributions. We use a single GFlowNet whose sampling policy follows a two-phase process: the DAG is first generated sequentially one edge at a time, and then the corresponding parameters are picked once the full structure is known. Since the parameters are included in the posterior distribution, this leaves more flexibility for the local probability models of the Bayesian Network, making our approach applicable even to non-linear models parametrized by neural networks. We show that our method, called JSP-GFN, offers an accurate approximation of the joint posterior, while comparing favorably against existing methods on both simulated and real data.

Added

2026-09-26

InfoDiffusion: Representation Learning Using Information Maximizing Diffusion Models

InfoDiffusion: Representation Learning Using Information Maximizing Diffusion Models

Yingheng Wang, Yair Schiff, Aaron Gokaslan, Weishen Pan, Fei Wang, Christopher De Sa, Volodymyr Kuleshov

OrganizationsCornell University

Why you should read this

Proposes InfoDiffusion, a framework that integrates mutual information regularization into diffusion models to extract semantically meaningful, disentangled low-dimensional latent representations without sacrificing generative sample quality.

While diffusion models excel at generating high-quality samples, their latent variables typically lack semantic meaning and are not suitable for representation learning. Here, we propose InfoDiffusion, an algorithm that augments diffusion models with low-dimensional latent variables that capture high-level factors of variation in the data. InfoDiffusion relies on a learning objective regularized with the mutual information between observed and hidden variables, which improves latent space quality and prevents the latents from being ignored by expressive diffusion-based decoders. Empirically, we find that InfoDiffusion learns disentangled and human-interpretable latent representations that are competitive with state-of-the-art generative and contrastive methods, while retaining the high sample quality of diffusion models. Our method enables manipulating the attributes of generated images and has the potential to assist tasks that require exploring a learned latent space to generate quality samples, e.g., generative design.

Added

2026-09-26

Bayesian Invariant Risk Minimization

Bayesian Invariant Risk Minimization

Yong Lin, Hanze Dong, Hao Wang, Tong Zhang

OrganizationsRutgers UniversityThe Hong Kong University of Science and Technology

Why you should read this

Reveals that Invariant Risk Minimization fails in deep networks because overfitting causes it to degenerate into empirical risk minimization, and resolves this failure mode by incorporating Bayesian inference over classifier posteriors to consistently improve out-of-distribution generalization.

Generalization under distributional shift is an open challenge for machine learning. Invariant Risk Minimization (IRM) is a promising framework to tackle this issue by extracting invariant features. However, despite the potential and popularity of IRM, recent works have reported negative results of it on deep models. We argue that the failure can be primarily attributed to deep models’ tendency to overfit the data. Specifically, our theoretical analysis shows that IRM degenerates to empirical risk minimization (ERM) when overfitting occurs. Our empirical evidence also provides supports: IRM methods that work well in typical settings significantly deteriorate even if we slightly enlarge the model size or lessen the training data. To alleviate this issue, we propose Bayesian Invariant Risk Minimization (BIRM) by introducing Bayesian inference into the IRM. The key motivation is to estimate the penalty of IRM based on the posterior distribution of classifiers (as opposed to a single classifier), which is much less prone to overfitting. Extensive experimental results on four datasets demonstrate that BIRM consistently outperforms the existing IRM baselines significantly.

Added

2026-09-26

Sampling-Based Accuracy Testing of Posterior Estimators for General Inference

Sampling-Based Accuracy Testing of Posterior Estimators for General Inference

Pablo Lemos, Adam Coogan, Yashar Hezaveh, Laurence Perreault Levasseur

OrganizationsCIELA InstituteFlatiron InstituteMila – Québec Artificial Intelligence InstituteUniversité de Montréal

Why you should read this

Introduces Tests of Accuracy with Random Points (TARP), a sample-only coverage testing framework that provides necessary and sufficient guarantees for validating high-dimensional generative posterior estimators without requiring explicit density evaluations.

Parameter inference, i.e. inferring the posterior distribution of the parameters of a statistical model given some data, is a central problem to many scientific disciplines. Generative models can be used as an alternative to Markov Chain Monte Carlo methods for conducting posterior inference, both in likelihood-based and simulation-based problems. However, assessing the accuracy of posteriors encoded in generative models is not straightforward. In this paper, we introduce ‘Tests of Accuracy with Random Points’ (TARP) coverage testing as a method to estimate coverage probabilities of generative posterior estimators. Our method differs from previously-existing coverage-based methods, which require posterior evaluations. We prove that our approach is necessary and sufficient to show that a posterior estimator is accurate. We demonstrate the method on a variety of synthetic examples, and show that TARP can be used to test the results of posterior inference analyses in high-dimensional spaces. We also show that our method can detect inaccurate inferences in cases where existing methods fail.

Added

2026-09-26

Normalizing Flows are Capable Generative Models

Normalizing Flows are Capable Generative Models

Shuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot, Jiatao Gu, Huangjie Zheng, Tianrong Chen, Miguel ngel Bautista, Navdeep Jaitly, Joshua M. Susskind

OrganizationsApple

Why you should read this

Demonstrates that normalizing flows can rival diffusion models in image generation quality while achieving state-of-the-art exact likelihood estimation through a scalable Transformer-based architecture paired with noise augmentation and guidance.

Normalizing Flows (NFs) are likelihood-based models for continuous inputs. They have demonstrated promising results on both density estimation and generative modeling tasks, but have received relatively little attention in recent years. In this work, we demonstrate that NFs are more powerful than previously believed. We present TARFLOW: a simple and scalable architecture that enables highly performant NF models. TARFLOW can be thought of as a Transformer-based variant of Masked Autoregressive Flows (MAFs): it consists of a stack of autoregressive Transformer blocks on image patches, alternating the autoregression direction between layers. TARFLOW is straightforward to train end-to-end, and capable of directly modeling and generating pixels. We also propose three key techniques to improve sample quality: Gaussian noise augmentation during training, a post training denoising procedure, and an effective guidance method for both class-conditional and unconditional settings. Putting these together, TARFLOW sets new state-of-the-art results on likelihood estimation for images, beating the previous best methods by a large margin, and generates samples with quality and diversity comparable to diffusion models, for the first time with a stand-alone NF model. We make our code available at https://github.com/apple/ml-tarflow.

Added

2026-09-26

Pre-trained Gaussian Processes for Bayesian Optimization

Pre-trained Gaussian Processes for Bayesian Optimization

Zi Wang, George E. Dahl, Kevin Swersky, Chansoo Lee, Zachary Nado, Justin Gilmer, Jasper Snoek, Zoubin Ghahramani

OrganizationsGoogle

Why you should read this

Proposes HyperBO, a transfer learning framework that pre-trains Gaussian process priors across related tuning tasks to achieve near-zero theoretical regret and speed up deep learning hyperparameter optimization by at least threefold over existing methods.

Bayesian optimization (BO) has become a popular strategy for global optimization of expensive real-world functions. Contrary to a common expectation that BO is suited to optimizing black-box functions, it actually requires domain knowledge about those functions to deploy BO successfully. Such domain knowledge often manifests in Gaussian process (GP) priors that specify initial beliefs on functions. However, even with expert knowledge, it is non-trivial to quantitatively define a prior. This is especially true for hyperparameter tuning problems on complex machine learning models, where landscapes of tuning objectives are often difficult to comprehend. We seek an alternative practice for setting these functional priors. In particular, we consider the scenario where we have data from similar functions that allow us to pre-train a tighter distribution a priori. We detail what pre-training entails for GPs using a KL divergence based loss function, and propose a new pre-training based BO framework named HyperBO. Theoretically, we show bounded posterior predictions and near-zero regrets for HyperBO without assuming the “ground truth” GP prior is known. To verify our approach in realistic setups, we collect a large multi-task hyperparameter tuning dataset by training tens of thousands of configurations of near-state-of-the-art deep learning models on popular image and text datasets, as well as a protein sequence dataset. Our results show that on average, HyperBO is able to locate good hyperparameters at least 3 times more efficiently than the best competing methods on both our new tuning dataset and existing multi-task BO benchmarks.

Added

2026-09-26

Pathfinder: Parallel quasi-Newton variational inference

Pathfinder: Parallel quasi-Newton variational inference

Lu Zhang, Bob Carpenter, Andrew Gelman, Aki Vehtari

OrganizationsAalto UniversityColumbia UniversityFlatiron InstituteUniversity of Southern California

Why you should read this

Introduces Pathfinder, a parallel variational inference algorithm that constructs normal approximations along quasi-Newton optimization trajectories to produce high-quality posterior draws with one to two orders of magnitude fewer gradient evaluations than standard methods.

We propose Pathfinder, a variational method for approximately sampling from differentiable probability densities. Starting from a random initialization, Pathfinder locates normal approximations to the target density along a quasi-Newton optimization path, with local covariance estimated using the inverse Hessian estimates produced by the optimizer. Pathfinder returns draws from the approximation with the lowest estimated Kullback-Leibler (KL) divergence to the target distribution. We evaluate Pathfinder on a wide range of posterior distributions, demonstrating that its approximate draws are better than those from automatic differentiation variational inference (ADVI) and comparable to those produced by short chains of dynamic Hamiltonian Monte Carlo (HMC), as measured by 1-Wasserstein distance. Compared to ADVI and short dynamic HMC runs, Pathfinder requires one to two orders of magnitude fewer log density and gradient evaluations, with greater reductions for more challenging posteriors. Importance resampling over multiple runs of Pathfinder improves the diversity of approximate draws, reducing 1-Wasserstein distance further and providing a measure of robustness to optimization failures on plateaus, saddle points, or in minor modes. The Monte Carlo KL divergence estimates are embarrassingly parallelizable in the core Pathfinder algorithm, as are multiple runs in the resampling version, further increasing Pathfinder’s speed advantage with multiple cores.

Added

2026-09-26

Bayesian Model Selection, the Marginal Likelihood, and Generalization

Bayesian Model Selection, the Marginal Likelihood, and Generalization

Sanae Lotfi, Pavel Izmailov, Gregory W. Benton, Micah Goldblum, Andrew Gordon Wilson

OrganizationsNew York University

Why you should read this

Demonstrates how marginal likelihood can negatively correlate with generalization in deep learning and introduces conditional marginal likelihood as an effective remedy for neural architecture search and large-scale hyperparameter learning.

How do we compare between hypotheses that are entirely consistent with observations? The marginal likelihood (aka Bayesian evidence), which represents the probability of generating our observations from a prior, provides a distinctive approach to this foundational question, automatically encoding Occam's razor. Although it has been observed that the marginal likelihood can overfit and is sensitive to prior assumptions, its limitations for hyperparameter learning and discrete model comparison have not been thoroughly investigated. We first revisit the appealing properties of the marginal likelihood for learning constraints and hypothesis testing. We then highlight the conceptual and practical issues in using the marginal likelihood as a proxy for generalization. Namely, we show how marginal likelihood can be negatively correlated with generalization, with implications for neural architecture search, and can lead to both underfitting and overfitting in hyperparameter learning. We provide a partial remedy through a conditional marginal likelihood, which we show is more aligned with generalization, and practically valuable for large-scale hyperparameter learning, such as in deep kernel learning.

Added

2026-09-26

A Variational Perspective on Solving Inverse Problems with Diffusion Models

A Variational Perspective on Solving Inverse Problems with Diffusion Models

Morteza Mardani, Jiaming Song, Jan Kautz, Arash Vahdat

OrganizationsNVIDIA

Why you should read this

Proposes a variational framework that transforms diffusion-based posterior sampling into stochastic optimization, enabling the use of standard solvers to perform image restoration without task-specific retraining.

Diffusion models have emerged as a key pillar of foundation models in visual domains. One of their critical applications is to universally solve different downstream inverse tasks via a single diffusion prior without re-training for each task. Most inverse tasks can be formulated as inferring a posterior distribution over data (e.g., a full image) given a measurement (e.g., a masked image). This is however challenging in diffusion models since the nonlinear and iterative nature of the diffusion process renders the posterior intractable. To cope with this challenge, we propose a variational approach that by design seeks to approximate the true posterior distribution. We show that our approach naturally leads to regularization by denoising diffusion process (RED-Diff) where denoisers at different timesteps concurrently impose different structural constraints over the image. To gauge the contribution of denoisers from different timesteps, we propose a weighting mechanism based on signal-to-noise-ratio (SNR). Our approach provides a new variational perspective for solving inverse problems with diffusion models, allowing us to formulate sampling as stochastic optimization, where one can simply apply off-the-shelf solvers with lightweight iterates. Our experiments for image restoration tasks such as inpainting and superresolution demonstrate the strengths of our method compared with state-of-the-art sampling-based diffusion models.

Added

2026-09-26