Learning GFlowNets From Partial Episodes For Improved Convergence And Stability

Kanika MadanJarrid Rector-BrooksMaksym KorablyovEmmanuel BengioMoksh JainAndrei Cristian NicaTom BoscYoshua BengioNikolay Malkin

article2023ICML157 citations

Introduces Subtrajectory Balance, a GFlowNet training objective inspired by TD(λ\lambda) that learns from partial action sequences to balance gradient bias and variance, accelerating convergence and enabling effective training in long-horizon, sparse-reward environments.

Listen

Generating diverse, high-value discrete objects—such as drug molecules, functional proteins, and specialized biological sequences—is a central challenge in computational design. Generative flow networks have emerged as a powerful framework to sample discrete objects proportionally to an unnormalized reward function. However, existing training objectives face severe practical limits. Local methods suffer from slow credit assignment across long action sequences, while trajectory-level methods propagate reward signals across complete episodes at the cost of high gradient variance and instability in sparse-reward environments.

The main objective of the article is to introduce and evaluate subtrajectory balance, a flexible training objective parameterized by a mixing coefficient that enables models to learn from partial action subsequences of varying lengths. The authors demonstrate how this method addresses the core gradient bias-variance tradeoff to improve training stability and convergence.

The authors evaluate subtrajectory balance through systematic computational experiments and gradient dynamics analysis. Testing spans synthetic hypergrid environments of varying dimensions and reward sparsity, synthetic bit sequence generation, and three complex biological design benchmarks: small molecule inhibitor synthesis, antimicrobial peptide generation, and fluorescent protein generation. The empirical evaluations compare the proposed method directly against existing network objectives and standard reinforcement learning baselines across thousands of training trajectories.

The analysis reveals several key findings. First, subtrajectory balance achieves consistently faster convergence and higher stability across all test environments compared to existing trajectory balance and detailed balance methods. Second, the method succeeds in challenging environments where past methods fail completely, such as grid tasks with extreme reward sparsity where trajectory balance fails to discover target modes. Third, on biological sequence tasks, the proposed method substantially outperforms alternative objectives and reinforcement learning baselines; for fluorescent protein generation, it achieves a top-100 mean reward of 1.18 compared to 0.76 for trajectory balance, while maintaining high sequence diversity. Finally, gradient analyses confirm that the weighting parameter smoothly interpolates between high-bias, low-variance local updates and low-bias, high-variance full-trajectory updates, with small-batch subtrajectory gradients providing a more accurate estimate of the large-batch target gradient.

These findings indicate that generative flow networks can now be reliably scaled to tasks with long sequential trajectories and sparse rewards without incurring significant computational overhead. Because the loss can be computed using single forward and backward neural network passes per sampled trajectory, the method provides substantial gains in performance and exploration efficiency without demanding extra network evaluations. This lowers the practical barrier and computational cost of training generative models for high-value scientific discovery pipelines.

Organizations applying generative flow networks to molecular, genetic, or sequential design should adopt subtrajectory balance as a default training objective. Practitioners should tune the mixing parameter near intermediate values (such as 0.8 to 1.9 depending on domain length) to balance credit assignment and gradient variance. Further exploration is recommended to test dynamic, learnable weighting strategies and to apply the objective to settings where rewards are available for incomplete intermediate states.

The empirical findings are validated across multiple random runs and diverse domains, giving high confidence in the method's core advantages. However, users should note that optimal performance still relies on selecting an appropriate mixing parameter and exploration schedule, and results in real-world biological applications remain constrained by the accuracy of the underlying proxy reward models.

arXiv: 2209.12782
  • Paper: Local Search GFlowNets, Minsu Kim et al. (2024). After SubTB(λ) introduces flexible GFlowNet training objectives, this work shows how GFlowNets can be further adapted with local refinement and prioritized learning for stronger candidate discovery.
Cover for Learning GFlowNets From Partial Episodes For Improved Convergence And Stability

Abstract

Generative flow networks (GFlowNets) are a family of algorithms for training a sequential sampler of discrete objects under an unnormalized target density and have been successfully used for various probabilistic modeling tasks. Existing training objectives for GFlowNets are either local to states or transitions, or propagate a reward signal over an entire sampling trajectory. We argue that these alternatives represent opposite ends of a gradient bias-variance tradeoff and propose a way to exploit this tradeoff to mitigate its harmful effects. Inspired by the TD(λ) algorithm in reinforcement learning, we introduce subtrajectory balance or SubTB(λ), a GFlowNet training objective that can learn from partial action subsequences of varying lengths. We show that SubTB(λ) accelerates sampler convergence in previously studied and new environments and enables training GFlowNets in environments with longer action sequences and sparser reward landscapes than what was possible before. We also perform a comparative analysis of stochastic gradient dynamics, shedding light on the bias-variance tradeoff in GFlowNet training and the advantages of subtrajectory balance.

Table of Contents

  • 1. Introduction
  • 2. Method
  • 2.1. Preliminaries
  • 2.2. GFlowNet training objectives
  • 2.3. Subtrajectory balance: Learning from partial episodes
  • 3. Related work
  • 4. Experiments
  • 4.1. Hypergrid: Robustness to sparse rewards
  • 4.1.1. A CLOSER LOOK AT GRADIENT VARIANCE
  • 4.2. Small molecule synthesis
  • 4.3. Sequence generation
  • 4.3.1. BIT SEQUENCES
  • 4.3.2. ANTIMICROBIAL PEPTIDE GENERATION
  • 4.3.3. FLUORESCENT PROTEIN GENERATION
  • 5. Discussion and conclusion
  • Acknowledgments
  • References
  • A. Experiment details: Hypergrid
  • A.1. Additional experiments
  • A.2. More on bias and variance: The effect of learned state flows
  • B. Experiment details: Molecules
  • C. Experiment details: Bit sequences
  • D. Experiment details: Antimicrobial peptide generation
  • E. Experiment details: Fluorescent protein generation
  • F. Inverse protein folding: Non-autoregressive sequence generation

Knowls

  1. Knowl 1 — Subtrajectory balance enforces the reward-proportional terminal distribution

    model/method

    Consider a directed acyclic state graph with forward transition probabilities PF(si+1∣si)P_F(s_{i+1}\mid s_i), backward transition probabilities PB(si∣si+1)P_B(s_i\mid s_{i+1}), and a nonnegative state-flow function F(s)F(s). For any contiguous trajectory segment τ=(sm→⋯→sn)\tau=(s_m\to\cdots\to s_n), subtrajectory balance requires

    F(sm)∏i=mn−1PF(si+1∣si)=F(sn)∏i=mn−1PB(si∣si+1).F(s_m)\prod_{i=m}^{n-1}P_F(s_{i+1}\mid s_i) = F(s_n)\prod_{i=m}^{n-1}P_B(s_i\mid s_{i+1}).

    At terminal states xx, the flow is fixed to the reward, F(x)=R(x)F(x)=R(x). The SubTB loss for a segment is the squared logarithm of the ratio of the two sides. If this loss is zero for every partial trajectory, the induced forward policy samples terminal states with probabilities proportional to R(x)R(x). The one-action case recovers detailed balance; the complete-trajectory case recovers trajectory balance with initial flow F(s0)F(s_0).

  2. Knowl 2 — Length-weighted SubTB interpolates between local and complete-trajectory training

    algorithm

    Given a sampled trajectory (s0→⋯→sn)(s_0\to\cdots\to s_n), evaluate the SubTB loss Li:j\mathcal{L}_{i:j} on every nonempty contiguous segment (si→⋯→sj)(s_i\to\cdots\to s_j), where 0≤i<j≤n0\le i<j\le n. For a terminal endpoint, use R(sj)R(s_j) in place of the estimated flow F(sj)F(s_j). Update the model using the normalized weighted loss

    Lλ=∑0≤i<j≤nλj−iLi:j∑0≤i<j≤nλj−i,λ>0.\mathcal{L}_{\lambda} =\frac{\displaystyle\sum_{0\le i<j\le n}\lambda^{j-i}\mathcal{L}_{i:j}} {\displaystyle\sum_{0\le i<j\le n}\lambda^{j-i}}, \qquad \lambda>0.

    At λ=1\lambda=1, all segments receive equal weight. As λ→0+\lambda\to0^+, the objective becomes the average one-step detailed-balance loss; as λ→+∞\lambda\to+\infty, it becomes the complete-trajectory balance loss. The experiments normalized weights over all segments in a batch. A length-nn trajectory has O(n2)O(n^2) segments, but the quadratic overhead is in operations on already computed log-flows and policy logits: gradient evaluation requires only one forward and one backward pass through the networks, so the paper reports little neural-network computation overhead relative to DB or TB.

  3. Knowl 3 — SubTB is hypothesized to reduce variance by learning intermediate flow targets

    theoretical result

    The paper proposes a bias–variance explanation for SubTB, rather than establishing it as a theorem. In trajectory balance, trajectories that share an initial prefix have common flow and policy terms, while their later transitions contribute a stochastic tail to the squared log-ratio target. Subtrajectory balance introduces learned state flows at intermediate states; its segment losses regress these flows against parts of the trajectory-level expressions. Replacing stochastic tail terms with learned flow estimates is expected to reduce gradient variance, but it introduces bias relative to the trajectory-balance gradient. The authors compare this mechanism to variance reduction from learned estimates in actor–critic methods.

  4. Knowl 4 — Intermediate state flows may speed learning through generalization

    assumption

    The paper also hypothesizes that SubTB can improve convergence because a learned state-flow function may generalize across states more effectively than the often high-dimensional forward- and backward-policy logits. This could matter in wide state graphs, where terminal states are numerous but training visits only a small subset, leaving sparse learning signals near termination. The proposed faster-learning effect is a motivation, not a separately proven result.

  5. Knowl 5 — Tabular gradient measurements support a bias–variance tradeoff

    empirical result

    The authors compared DB, TB, and SubTB gradients in a tabular GFlowNet on the harder 8×88\times8 hypergrid, using SubTB with λ=0.8\lambda=0.8. The tabular representation gave each flow and policy logit an independent parameter. Gradient self-consistency was measured by cosine similarity between gradients from small trajectory batches and a 1024-trajectory batch. At the training batch size of 64, DB had the highest self-consistency, TB the lowest, and SubTB lay between them, consistent with decreasing gradient variance in the order TB, SubTB, DB.

    They also compared small-batch gradients with the 1024-trajectory TB gradient as a reference. At intermediate training iterations, the small-batch SubTB gradient was more similar to that reference than the small-batch TB gradient was, despite SubTB's bias. With full-size batches, SubTB's similarity to the TB reference lay between DB and TB. Thus the observed gradient behavior placed SubTB between DB's more biased, more self-consistent estimates and TB's less biased, less self-consistent estimates.

  6. Knowl 6 — SubTB improves convergence on standard and sparse-reward hypergrids

    empirical result

    On two- and four-dimensional hypergrids with multimodal rewards concentrated near the grid corners, the authors compared SubTB with λ=0.9\lambda=0.9 against TB and DB. The standard reward variant had background reward 10−310^{-3}; the harder variant reduced it to 10−410^{-4}. Models were trained with Adam, batch size 16, and 10610^6 sampled trajectories; results were averaged over three random runs. Convergence was measured by the unnormalized L1L^1 distance between the target distribution and the empirical distribution of the last 2⋅1052\cdot10^5 terminal samples.

    SubTB converged faster than TB across the tested grid sizes and showed less variation between seeds. It continued to work well on the harder, sparser rewards, whereas TB failed to discover all target modes on grids larger than 8×88\times8. In an additional restricted-segment experiment, SubTB using only segments of length at most four remained effective on the harder 32×3232\times32 and 40×4040\times40 grids.

  7. Knowl 7 — SubTB improves reward–likelihood fit on molecular generation

    empirical result

    For sEH binder generation, molecules were assembled sequentially from a fixed fragment library in an estimated state space of size 101210^{12}. A pretrained proxy supplied the reward, modified as R(x)=Re(x)βR(x)=R_e(x)^\beta, where ReR_e is the proxy score and β\beta controls reward sharpness. The evaluation metric was the correlation between log⁡R(x)\log R(x) and the model's marginal log sampling probability log⁡pθ(x)\log p_\theta(x) on held-out terminal molecules.

    Across three runs per setting, SubTB—especially with λ=1\lambda=1—achieved higher correlation than DB and TB when each method's hyperparameters were optimized. SubTB was also more robust to hyperparameter choice: the reported best-over-other-parameters results and averages over other settings favored it over DB and TB. The study varied the reward exponent, learning rate, and SubTB weight.

  8. Knowl 8 — SubTB performs best across bit-sequence lengths and token sizes

    empirical result

    In the synthetic sequence task, the target objects were 120-bit strings and the reward was R(x)=exp⁡[−min⁡y∈Md(x,y)]R(x)=\exp[-\min_{y\in M}d(x,y)], where MM is a set of modes and dd is Hamming distance. Actions appended tokens of kk bits, with k∈{1,2,4,6,8,10}k\in\{1,2,4,6,8,10\}, so changing kk varied action-space size and trajectory length without changing the underlying target. The models were evaluated by Spearman correlation between sampling probability and reward on a test set, and by the number of modes found during training.

    SubTB achieved the highest reported test-set correlation across the tested token sizes and discovered modes faster than the other GFlowNet objectives and the A2C, Soft Actor-Critic, and MCMC baselines. The selected SubTB weight was λ=1.9\lambda=1.9; models used a three-layer Transformer and were trained for 50,000 iterations with minibatches of 16.

  9. Knowl 9 — SubTB yields higher reward on peptide and fluorescent-protein generation

    data/table

    The peptide experiment generated antimicrobial sequences of maximum length 60 from 20 amino acids plus an end token. The fluorescent-protein experiment generated fixed-length sequences of length 237 from a 20-amino-acid vocabulary. For each task, 2048 policy samples were used to calculate mean reward and mean pairwise edit distance among the top 100 sequences. The values below are means and standard errors over three runs; diversity is the top-100 pairwise edit-distance metric.

    SubTB gave the highest mean reward on both tasks. On peptides it also produced substantially greater diversity than the other methods. On fluorescent proteins its diversity was similar to TB, while its reward was higher.

    MethodPeptide rewardPeptide diversityProtein rewardProtein diversity
    SubTB0.96±0.020.96\pm0.0242.23±3.442.23\pm3.41.18±0.101.18\pm0.10204.44±0.45204.44\pm0.45
    TB0.90±0.030.90\pm0.0331.42±2.931.42\pm2.90.76±0.190.76\pm0.19204.31±0.44204.31\pm0.44
    FM/DB0.78±0.050.78\pm0.0512.61±1.3212.61\pm1.320.30±0.080.30\pm0.08190.21±6.78190.21\pm6.78
    SAC0.80±0.010.80\pm0.018.36±1.448.36\pm1.440.23±0.030.23\pm0.03120.32±15.57120.32\pm15.57
    A2C with entropy regularization0.79±0.020.79\pm0.027.32±0.767.32\pm0.760.22±0.020.22\pm0.02113.65±21.31113.65\pm21.31
    MCMC0.75±0.020.75\pm0.0212.56±1.4512.56\pm1.450.28±0.010.28\pm0.01169.17±12.44169.17\pm12.44

    For the protein task, SubTB's reward advantage over TB was larger than in the peptide task, while their top-100 diversity values were nearly equal.

  10. Knowl 10 — Intermediate SubTB weights fit the inverse-folding target best

    empirical result

    In the non-autoregressive inverse protein-folding experiment, the target was a Boltzmann distribution over amino-acid sequences of fixed length L=40L=40, scored by a physics-based energy model for a given protein backbone. The sampler first selected a sequence uniformly at random, then made exactly N=40N=40 replacement actions, each modifying a position in the sequence. The forward policy was conditioned on the number of actions already taken, and the backward policy was uniform over the N⋅LN\cdot L actions. Model fit was assessed by correlation between log reward and marginal log sampling probability on held-out terminal sequences.

    The reported comparison across SubTB weights found that intermediate values of λ\lambda gave the best fit to the target distribution, rather than either endpoint corresponding to DB-like local training or TB-like complete-trajectory training.

Coverage note — Detailed network architectures and per-task hyperparameter search grids are omitted where they do not change the main findings; the paper's background, related work, and acknowledgements are outside the contribution scope.

References

  1. 1.Baird, L. Residual algorithms: Reinforcement learning with function approximation. International Conference on Machine Learning (ICML), 1995.
  2. 2.Bengio, E., Pineau, J., and Precup, D. Interference and generalization in temporal difference learning. International Conference on Machine Learning (ICML), 2020.
  3. 3.Bengio, E., Jain, M., Korablyov, M., Precup, D., and Bengio, Y. Flow network based generative models for non-iterative diverse candidate generation. Neural Information Processing Systems (NeurIPS), 2021a.
  4. 4.Bengio, Y., Lahlou, S., Deleu, T., Hu, E., Tiwari, M., and Bengio, E. GFlowNet foundations. arXiv preprint 2111.09266, 2021b.
  5. 5.Brookes, D., Park, H., and Listgarten, J. Conditioning by adaptive sampling for robust design. International Conference on Machine Learning (ICML), 2019.
  6. 6.Chaudhury, S., Lyskov, S., and Gray, J. J. PyRosetta: a script-based interface for implementing molecular modeling algorithms using Rosetta. Bioinformatics, 26(5):689–691, 2010.
  7. 7.Christodoulou, P. Soft actor-critic for discrete action settings. arXiv preprint 1910.07207, 2019.
  8. 8.Deleu, T., Gois, A., Emezue, C., Rankawat, M., Lacoste-Julien, S., Bauer, S., and Bengio, Y. Bayesian structure learning with generative flow networks. Uncertainty in Artificial Intelligence (UAI), 2022.
  9. 9.Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S. Diversity is all you need: Learning skills without a reward function. International Conference on Learning Representations (ICLR), 2018.
  10. 10.Guo, H., Tan, B., Liu, Z., Xing, E. P., and Hu, Z. Text generation with efficient (soft) Q-learning. arXiv preprint 2106.07704, 2021.
  11. 11.Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. Reinforcement learning with deep energy-based policies. International Conference on Machine Learning (ICML), 2017.
  12. 12.Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. International Conference on Machine Learning (ICML), 2018.
  13. 13.Hazan, E., Kakade, S., Singh, K., and Van Soest, A. Provably efficient maximum entropy exploration. International Conference on Machine Learning (ICML), 2019.
  14. 14.Hu, E. J., Malkin, N., Jain, M., Everett, K., Graikos, A., and Bengio, Y. GFlowNet-EM for learning compositional latent variable models. International Conference on Machine Learning (ICML), 2023.
  15. 15.Ilyas, A., Engstrom, L., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A. A closer look at deep policy gradients. International Conference on Learning Representations (ICLR), 2020.
  16. 16.Islam, R., Ahmed, Z., and Precup, D. Marginalized state distribution entropy regularization in policy optimization. arXiv preprint 1912.05128, 2019.
  17. 17.Jain, M., Bengio, E., Hernandez-Garcia, A., Rector-Brooks, J., Dossou, B. F., Ekbote, C., Fu, J., Zhang, T., Kilgour, M., Zhang, D., Simine, L., Das, P., and Bengio, Y. Biological sequence design with GFlowNets. International Conference on Machine Learning (ICML), 2022.
  18. 18.Jin, W., Barzilay, R., and Jaakkola, T. Chapter 11. junction tree variational autoencoder for molecular graph generation. Drug Discovery, pp. 228–249, 2020. ISSN 2041-3211.
  19. 19.Kearns, M. J. and Singh, S. P. Bias-variance error bounds for temporal difference updates. Conference on Learning Theory (COLT), 2000.
  20. 20.Kumar, A., Voet, A., and Zhang, K. Y. Fragment based drug design: from experimental to computational approaches. Current medicinal chemistry, 19(30):5128–5147, 2012.
  21. 21.Malkin, N., Jain, M., Bengio, E., Sun, C., and Bengio, Y. Trajectory balance: Improved credit assignment in GFlowNets. Neural Information Processing Systems (NeurIPS), 2022.
  22. 22.Malkin, N., Lahlou, S., Deleu, T., Ji, X., Hu, E., Everett, K., Zhang, D., and Bengio, Y. GFlowNets and variational inference. International Conference on Learning Representations (ICLR), 2023.
  23. 23.Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. Asynchronous methods for deep reinforcement learning. Neural Information Processing Systems (NIPS), 2016.
  24. 24.Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. Bridging the gap between value and policy based reinforcement learning. Neural Information Processing Systems (NIPS), 2017.
  25. 25.Pan, L., Malkin, N., Zhang, D., and Bengio, Y. Better training of GFlowNets with local credit and incomplete trajectories. International Conference on Machine Learning (ICML), 2023.
  26. 26.Pirtskhalava, M., Amstrong, A. A., Grigolava, M., Chubinidze, M., Alimbarashvili, E., Vishnepolsky, B., Gabrielian, A., Rosenthal, A., Hurt, D. E., and Tartakovsky, M. Dbaasp v3: Database of antimicrobial/cytotoxic activity and structure of peptides as a resource for development of new therapeutics. Nucleic Acids Research, 49(D1):D288–D297, 2021.
  27. 27.Rohl, C. A., Strauss, C. E., Misura, K. M., and Baker, D. Protein structure prediction using Rosetta. In Methods in enzymology, volume 383, pp. 66–93. Elsevier, 2004.
  28. 28.Sarkisyan, K. S., Bolotin, D. A., Meer, M. V., Usmanova, D. R., Mishin, A. S., Sharonov, G. V., Ivankov, D. N., Bozhanova, N. G., Baranov, M. S., Soylemez, O., et al. Local fitness landscape of the green fluorescent protein. Nature, 533(7603):397–401, 2016.
  29. 29.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint 1707.06347, 2017.
  30. 30.Sinai, S., Wang, R., Whatley, A., Slocum, S., Locane, E., and Kelsic, E. AdaLead: A simple and robust adaptive greedy search algorithm for sequence design. arXiv preprint 2010.02141, 2020.
  31. 31.Sutton, R. S. Learning to predict by the methods of temporal differences. Machine learning, 3(1):9–44, 1988.
  32. 32.Sutton, R. S. and Barto, A. G. Reinforcement learning: An introduction. MIT Press, 2018.
  33. 33.Trabucco, B., Geng, X., Kumar, A., and Levine, S. Design-bench: Benchmarks for data-driven offline model-based optimization. International Conference on Machine Learning (ICML), 2022.
  34. 34.Trott, O. and Olson, A. J. AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. Journal of computational chemistry, 31(2):455–461, 2010.
  35. 35.van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J. Deep reinforcement learning and the deadly triad. arXiv preprint 1812.02648, 2018.
  36. 36.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Neural Information Processing Systems (NIPS), 2017.
  37. 37.Williams, R. and Peng, J. Function optimization using connectionist reinforcement learning algorithms. Connection Science, 3:241–, 09 1991. doi: 10.1080/09540099108946587.
  38. 38.Xie, Y., Shi, C., Zhou, H., Yang, Y., Zhang, W., Yu, Y., and Li, L. MARS: Markov molecular sampling for multi-objective drug discovery. International Conference on Learning Representations (ICLR), 2021.
  39. 39.Zhang, C., Cai, Y., Huang, L., and Li, J. Exploration by maximizing Renyi entropy for reward-free RL framework. Association for the Advancement of Artificial Intelligence (AAAI), 2021.
  40. 40.Zhang, D., Malkin, N., Liu, Z., Volokhova, A., Courville, A., and Bengio, Y. Generative flow networks for discrete probabilistic modeling. International Conference on Machine Learning (ICML), 2022.
  41. 41.Zhang, S., Boehmer, W., and Whiteson, S. Deep residual reinforcement learning. Autonomous Agents and Multi-Agent Systems (AAMAS), 2020.
  42. 42.Ziebart, B. D. Modeling purposeful adaptive behavior with the principle of maximum causal entropy. Carnegie Mellon University, 2010.

Citation

MLA
Madan, K., et al. “Learning GFlowNets From Partial Episodes For Improved Convergence And Stability”. International Conference on Machine Learning, vol. 202, 2023, pp. 23467–83, https://proceedings.mlr.press/v202/madan23a.html.
APA
Madan, K., Rector-Brooks, J., Korablyov, M., Bengio, E., Jain, M., Nica, A. C., Bosc, T., Bengio, Y., & Malkin, N. (2023). Learning GFlowNets From Partial Episodes For Improved Convergence And Stability. International Conference on Machine Learning, 202, 23467–23483. https://proceedings.mlr.press/v202/madan23a.html
Chicago
Madan, K., J. Rector-Brooks, M. Korablyov, et al. 2023. “Learning GFlowNets From Partial Episodes For Improved Convergence And Stability”. International Conference on Machine Learning 202: 23467–83. https://proceedings.mlr.press/v202/madan23a.html.
Harvard
Madan, K. et al. (2023) “Learning GFlowNets From Partial Episodes For Improved Convergence And Stability”, International Conference on Machine Learning. PMLR, pp. 23467–23483. Available at: https://proceedings.mlr.press/v202/madan23a.html.
Vancouver
1. Madan K, Rector-Brooks J, Korablyov M, Bengio E, Jain M, Nica AC, Bosc T, Bengio Y, Malkin N (2023) Learning GFlowNets From Partial Episodes For Improved Convergence And Stability. In: International Conference on Machine Learning. PMLR, pp 23467–23483

BibTeX

@InProceedings{pmlr-v202-madan23a,
  title = 	 {Learning {GF}low{N}ets From Partial Episodes For Improved Convergence And Stability},
  author =       {Madan, Kanika and Rector-Brooks, Jarrid and Korablyov, Maksym and Bengio, Emmanuel and Jain, Moksh and Nica, Andrei Cristian and Bosc, Tom and Bengio, Yoshua and Malkin, Nikolay},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {23467--23483},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/madan23a/madan23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/madan23a.html},
  abstract = 	 {Generative flow networks (GFlowNets) are a family of algorithms for training a sequential sampler of discrete objects under an unnormalized target density and have been successfully used for various probabilistic modeling tasks. Existing training objectives for GFlowNets are either local to states or transitions, or propagate a reward signal over an entire sampling trajectory. We argue that these alternatives represent opposite ends of a gradient bias-variance tradeoff and propose a way to exploit this tradeoff to mitigate its harmful effects. Inspired by the TD($\lambda$) algorithm in reinforcement learning, we introduce subtrajectory balance or SubTB($\lambda$), a GFlowNet training objective that can learn from partial action subsequences of varying lengths. We show that SubTB($\lambda$) accelerates sampler convergence in previously studied and new environments and enables training GFlowNets in environments with longer action sequences and sparser reward landscapes than what was possible before. We also perform a comparative analysis of stochastic gradient dynamics, shedding light on the bias-variance tradeoff in GFlowNet training and the advantages of subtrajectory balance.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/