keyword
backward policy
A backward policy is a conditional probability distribution used in sequential generative modeling frameworks, such as Generative Flow Networks and continuous diffusion samplers, that defines how to transition backward from a given state to a preceding state along a trajectory. While a forward policy iteratively constructs a sample toward a target distribution or reward, the backward policy models the reverse process by distributing probability mass over earlier ancestor states that could have led to the current state. In training objectives such as trajectory balance and detailed balance, the backward policy can be fixed or parameterized and learned jointly with the forward policy, serving as an auxiliary reference that enables tractable path evaluation, credit assignment, and off-policy training over complete or partial trajectories.
1 item

