Built independently by an author, for readers. Read the story and support ChapterPal

keyword

variational gradient descent

Variational gradient descent is an optimization-based approximate inference method that iteratively transforms a collection of sample particles or a probability distribution to approximate a complex, intractable target distribution. Rather than relying on traditional stochastic sampling chains or restricted parametric families, it frames distributional approximation as a functional gradient descent or gradient flow over a probability space, systematically minimizing a discrepancy measure such as the Kullback-Leibler divergence. In typical implementations, such as Stein variational gradient descent, the update dynamics combine a driving force that pulls particles toward regions of high target density with a repulsive or smoothing interaction that prevents particles from collapsing into a single point, thereby preserving the diversity and spread of the distribution. This framework provides a scalable, deterministic alternative for sampling and uncertainty quantification in applications ranging from Bayesian inference to generative modeling and policy optimization.

3 items

Improving Task-free Continual Learning by Distributionally Robust Memory Evolution

Improving Task-free Continual Learning by Distributionally Robust Memory Evolution

Zhenyi Wang, Li Shen, Le Fang, Qiuling Suo, Tiehang Duan, Mingchen Gao

OrganizationsJD.comMetaUniversity at Buffalo

Why you should read this

Proposes a distributionally robust memory evolution framework that uses Wasserstein gradient flows to dynamically update replay buffers, preventing overfitting to stored samples and mitigating catastrophic forgetting in task-free continual learning.

Task-free continual learning (CL) aims to learn a non-stationary data stream without explicit task definitions and not forget previous knowledge. The widely adopted memory replay approach could gradually become less effective for long data streams, as the model may memorize the stored examples and overfit the memory buffer. Second, existing methods overlook the high uncertainty in the memory data distribution since there is a big gap between the memory data distribution and the distribution of all the previous data examples. To address these problems, for the first time, we propose a principled memory evolution framework to dynamically evolve the memory data distribution by making the memory buffer gradually harder to be memorized with distributionally robust optimization (DRO). We then derive a family of methods to evolve the memory buffer data in the continuous probability measure space with Wasserstein gradient flow (WGF). The proposed DRO is w.r.t the worst-case evolved memory data distribution, thus guarantees the model performance and learns significantly more robust features than existing memory-replay-based methods. Extensive experiments on existing benchmarks demonstrate the effectiveness of the proposed methods for alleviating forgetting. As a by-product of the proposed framework, our method is more robust to adversarial examples than existing task-free CL methods.

Added

2026-09-26

A Variational Perspective on Solving Inverse Problems with Diffusion Models

A Variational Perspective on Solving Inverse Problems with Diffusion Models

Morteza Mardani, Jiaming Song, Jan Kautz, Arash Vahdat

OrganizationsNVIDIA

Why you should read this

Proposes a variational framework that transforms diffusion-based posterior sampling into stochastic optimization, enabling the use of standard solvers to perform image restoration without task-specific retraining.

Diffusion models have emerged as a key pillar of foundation models in visual domains. One of their critical applications is to universally solve different downstream inverse tasks via a single diffusion prior without re-training for each task. Most inverse tasks can be formulated as inferring a posterior distribution over data (e.g., a full image) given a measurement (e.g., a masked image). This is however challenging in diffusion models since the nonlinear and iterative nature of the diffusion process renders the posterior intractable. To cope with this challenge, we propose a variational approach that by design seeks to approximate the true posterior distribution. We show that our approach naturally leads to regularization by denoising diffusion process (RED-Diff) where denoisers at different timesteps concurrently impose different structural constraints over the image. To gauge the contribution of denoisers from different timesteps, we propose a weighting mechanism based on signal-to-noise-ratio (SNR). Our approach provides a new variational perspective for solving inverse problems with diffusion models, allowing us to formulate sampling as stochastic optimization, where one can simply apply off-the-shelf solvers with lightweight iterates. Our experiments for image restoration tasks such as inpainting and superresolution demonstrate the strengths of our method compared with state-of-the-art sampling-based diffusion models.

Added

2026-09-26