Learning GFlowNets From Partial Episodes For Improved Convergence And Stability
Kanika MadanJarrid Rector-BrooksMaksym KorablyovEmmanuel BengioMoksh JainAndrei Cristian NicaTom BoscYoshua BengioNikolay Malkin
Introduces Subtrajectory Balance, a GFlowNet training objective inspired by TD() that learns from partial action sequences to balance gradient bias and variance, accelerating convergence and enabling effective training in long-horizon, sparse-reward environments.
Generating diverse, high-value discrete objects—such as drug molecules, functional proteins, and specialized biological sequences—is a central challenge in computational design. Generative flow networks have emerged as a powerful framework to sample discrete objects proportionally to an unnormalized reward function. However, existing training objectives face severe practical limits. Local methods suffer from slow credit assignment across long action sequences, while trajectory-level methods propagate reward signals across complete episodes at the cost of high gradient variance and instability in sparse-reward environments.
The main objective of the article is to introduce and evaluate subtrajectory balance, a flexible training objective parameterized by a mixing coefficient that enables models to learn from partial action subsequences of varying lengths. The authors demonstrate how this method addresses the core gradient bias-variance tradeoff to improve training stability and convergence.
The authors evaluate subtrajectory balance through systematic computational experiments and gradient dynamics analysis. Testing spans synthetic hypergrid environments of varying dimensions and reward sparsity, synthetic bit sequence generation, and three complex biological design benchmarks: small molecule inhibitor synthesis, antimicrobial peptide generation, and fluorescent protein generation. The empirical evaluations compare the proposed method directly against existing network objectives and standard reinforcement learning baselines across thousands of training trajectories.
The analysis reveals several key findings. First, subtrajectory balance achieves consistently faster convergence and higher stability across all test environments compared to existing trajectory balance and detailed balance methods. Second, the method succeeds in challenging environments where past methods fail completely, such as grid tasks with extreme reward sparsity where trajectory balance fails to discover target modes. Third, on biological sequence tasks, the proposed method substantially outperforms alternative objectives and reinforcement learning baselines; for fluorescent protein generation, it achieves a top-100 mean reward of 1.18 compared to 0.76 for trajectory balance, while maintaining high sequence diversity. Finally, gradient analyses confirm that the weighting parameter smoothly interpolates between high-bias, low-variance local updates and low-bias, high-variance full-trajectory updates, with small-batch subtrajectory gradients providing a more accurate estimate of the large-batch target gradient.
These findings indicate that generative flow networks can now be reliably scaled to tasks with long sequential trajectories and sparse rewards without incurring significant computational overhead. Because the loss can be computed using single forward and backward neural network passes per sampled trajectory, the method provides substantial gains in performance and exploration efficiency without demanding extra network evaluations. This lowers the practical barrier and computational cost of training generative models for high-value scientific discovery pipelines.
Organizations applying generative flow networks to molecular, genetic, or sequential design should adopt subtrajectory balance as a default training objective. Practitioners should tune the mixing parameter near intermediate values (such as 0.8 to 1.9 depending on domain length) to balance credit assignment and gradient variance. Further exploration is recommended to test dynamic, learnable weighting strategies and to apply the objective to settings where rewards are available for incomplete intermediate states.
The empirical findings are validated across multiple random runs and diverse domains, giving high confidence in the method's core advantages. However, users should note that optimal performance still relies on selecting an appropriate mixing parameter and exploration schedule, and results in real-world biological applications remain constrained by the accuracy of the underlying proxy reward models.
- Paper: Generative Flow Networks for Discrete Probabilistic Modeling, Dinghuai Zhang et al. (2022). This paper introduces GFlowNets and trajectory-balance training, the framework and objectives that SubTB(λ) compares and generalizes.
- Paper: High-Dimensional Continuous Control Using Generalized Advantage Estimation, John Schulman et al. (2016). Its Generalized Advantage Estimation explains the λ-controlled bias–variance tradeoff in reinforcement learning that motivates SubTB(λ).
- Paper: GFlowNet Foundations, Yoshua Bengio et al. (2023). Its account of GFlowNet flow conservation and detailed-balance objectives clarifies the local training objectives that SubTB(λ) extends.
- Paper: Local Search GFlowNets, Minsu Kim et al. (2024). After SubTB(λ) introduces flexible GFlowNet training objectives, this work shows how GFlowNets can be further adapted with local refinement and prioritized learning for stronger candidate discovery.
