Built independently by an author, for readers. Read the story and support ChapterPal

keyword

backward policy

A backward policy is a conditional probability distribution used in sequential generative modeling frameworks, such as Generative Flow Networks and continuous diffusion samplers, that defines how to transition backward from a given state to a preceding state along a trajectory. While a forward policy iteratively constructs a sample toward a target distribution or reward, the backward policy models the reverse process by distributing probability mass over earlier ancestor states that could have led to the current state. In training objectives such as trajectory balance and detailed balance, the backward policy can be fixed or parameterized and learned jointly with the forward policy, serving as an auxiliary reference that enables tractable path evaluation, credit assignment, and off-policy training over complete or partial trajectories.

1 item

Improved off-policy training of diffusion samplers

Improved off-policy training of diffusion samplers

Marcin Sendera, Minsu Kim, Sarthak Mittal, Pablo Lemos, Luca Scimeca, Jarrid Rector-Brooks, Alexandre Adam, Yoshua Bengio, Nikolay Malkin

OrganizationsCIELA InstituteCIFARDreamfoldJagiellonian UniversityKorea Advanced Institute of Science and TechnologyMilaUniversité de MontréalUniversity of Edinburgh

Why you should read this

Presents a unified benchmark and an effective replay-buffer exploration strategy for training diffusion models to sample from unnormalized energy densities via continuous generative flow networks.

We study the problem of training diffusion models to sample from a distribution with a given unnormalized density or energy function. We benchmark several diffusion-structured inference methods, including simulation-based variational approaches and off-policy methods (continuous generative flow networks). Our results shed light on the relative advantages of existing algorithms while bringing into question some claims from past work. We also propose a novel exploration strategy for off-policy methods, based on local search in the target space with the use of a replay buffer, and show that it improves the quality of samples on a variety of target distributions. Our code for the sampling methods and benchmarks studied is made public at (link) as a base for future work on diffusion models for amortized inference.

Added

2026-10-04