GFlowRL: Scaling Distribution-Matching RL to Large Language Models

Xiaodong LiuMichael XuJack W. StokesPaul SmolenskyDoug BurgerJianfeng Gao

article2026arXiv1 citations

Introduces GFlowRL, a reinforcement learning algorithm that eliminates the instability of distribution-matching methods by replacing learned partition functions with in-batch Monte Carlo estimates, enabling GFlowNet-style reasoning to scale across dense and mixture-of-experts models up to 235B parameters.

Listen

Post-training reinforcement learning is the primary driver of performance in state-of-the-art reasoning language models. However, standard reward-maximizing algorithms such as Proximal Policy Optimization and Group Relative Policy Optimization tend to cause mode collapse by narrowing the model's focus to a single dominant solution path rather than exploring diverse valid reasoning trajectories. While Generative Flow Network (GFlowNet) methods can encourage diverse reasoning by sampling outputs proportional to rewards, prior attempts to apply them to large models have suffered from severe optimization instability, gradient spikes, and high distributed-systems overhead.

The article demonstrates that the primary bottleneck in scaling GFlowNet-style reinforcement learning is an unnecessary auxiliary component: the learned partition function network. To resolve this, the authors evaluate and propose GFlowRL, a streamlined algorithm designed to stabilize and scale distribution-matching reinforcement learning across dense and mixture-of-experts architectures.

The authors conducted diagnostic and large-scale empirical experiments across mathematical reasoning, competitive programming, and adversarial security red-teaming. They evaluated models ranging from 7 billion to 235 billion parameters on standard benchmarks, including Codeforces, LiveCodeBench, AIME, AdvBench, and HarmBench. The core technical approach replaces the learned partition network entirely with a non-parametric in-batch Monte Carlo estimator computed directly from standard rollout groups, augmented with importance-sampling correction and asymmetric flow-gap clipping to stabilize training.

The investigation produced four central findings. First, diagnostic tests revealed that the learned partition function acts primarily as an uninformative noise source; replacing it with pure random noise yielded comparable performance (36.19% vs. 35.61% accuracy), while prior methods suffered 55 gradient explosions exceeding 10^6 across 421 steps compared to zero for GFlowRL. Second, on 14-billion parameter dense coding models, GFlowRL achieved a 2048 Codeforces Elo rating—outperforming DeepCoder-14B by 112 points, FlowRL by 144 points, and OpenAI o1 by 157 points, coming within 25 points of o3-mini. Third, GFlowRL maintained stability in noisy-reward red-teaming environments where prior GFlowNet approaches diverged, attaining leading attack success rates on AdvBench (82.5%) and HarmBench (79.5%). Fourth, the framework scaled seamlessly to sparse mixture-of-experts models up to 235 billion parameters, whereas prior methods failed to converge.

These findings indicate that generative distribution matching can be achieved efficiently without the computational overhead or instability of auxiliary neural networks. By integrating directly into existing rollout pipelines, GFlowRL reduces training risk, prevents costly optimization divergence, and significantly enhances solution diversity—scoring 3.93 out of 5.0 in diversity evaluations compared to 1.21 for reward-maximizing baselines.

Organizations developing reasoning-focused language models should consider replacing unstable partition networks with in-batch Monte Carlo estimation pipelines for post-training workflows. Because GFlowRL uses the same rollout group structure as Group Relative Policy Optimization, engineering teams can adopt the distribution-matching objective with minimal architectural disruption. Further exploration is recommended to evaluate whether this estimation paradigm extends effectively to broader agentic workflows and multimodal tasks.

The conclusions are supported by extensive evaluations across multiple benchmarks and parameter scales. A primary operational limitation is that the in-batch estimator's variance depends on the rollout group size, though the included clipping mechanisms effectively mitigate instability in practice. Stakeholders can have high confidence in these results for dense and sparse reasoning architectures within the evaluated task domains.

  • Paper: GFlowNet Foundations, Yoshua Bengio et al. (2023). Read this foundation first to understand how GFlowNets match sampling probabilities to rewards—the core objective that GFlowRL streamlines for LLM training.
  • Paper: Generative Flow Networks for Discrete Probabilistic Modeling, Dinghuai Zhang et al. (2022). Its discrete-data GFlowNet formulation clarifies how reward-proportional sampling works in token-like spaces, a useful prerequisite for following the paper’s LLM adaptation.

No sufficiently relevant recommendations were found.

Cover for GFlowRL: Scaling Distribution-Matching RL to Large Language Models

Abstract

Generative Flow Networks (GFlowNets) offer a promising alternative to reward-maximizing reinforcement learning (RL) for large reasoning models, encouraging diverse reasoning paths by matching reward distributions rather than collapsing to dominant modes. Recent work shows promise on math and code, but scaling GFlowNet-style RL to modern post-training pipelines remains difficult: as model size, rollout horizon, reward noise, and distributed-systems complexity grow together, a learned prompt-conditional partition function becomes a source of gradient instability and engineering overhead rather than a useful normalizer. Through systematic analysis, we find that the learned partition function, previously treated as essential, can be replaced by an in-batch Monte Carlo estimate computed from the rollout group already required for training. We propose GFlowRL, a streamlined GFlowNet-style RL algorithm that removes the auxiliary partition network entirely while preserving the reward-distribution-matching objective, completed by two stabilizers: importance-sampling correction for rollout/trainer drift and asymmetric flow-gap clipping for outlier residuals. GFlowRL exceeds all counterparts on math, code, and adversarial red-teaming benchmarks, reaching a Codeforces rating of 2048 at the 14B scale (within 25 Elo of o3-mini) and attaining the highest average ASR@1 on AdvBench and HarmBench, outperforming the previous SOTA multi-turn attacker in a regime where FlowRL, a prior GFlowNet-style method, diverges. The same recipe transfers to all evaluated MoE configurations up to 235B parameters, where FlowRL again fails to converge. To our knowledge, GFlowRL is the first GFlowNet-style RL algorithm to scale stably across both dense and sparse architectures. Code will be at: this https URL

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Methodology
  • 3.1 Why the Learned Partition Function Fails
  • 3.2 GFlowRL
  • 4 Experiment
  • 4.1 Experimental Setup
  • 4.2 Diagnostic Investigation
  • 4.3 Main Results
  • 4.4 Ablation Study
  • 5 Related Work
  • 6 Conclusion
  • 7 Acknowledgment
  • References
  • A Limitations and Broader Impacts
  • B Theoretical Justification for GFlowRL
  • C Extended Related Work
  • D Detailed Experimental Setup
  • D.1 Backbone Models
  • D.2 Datasets
  • D.3 Metrics
  • D.4 Baselines
  • D.5 Implementation
  • D.6 Diagnostic and Synthetic Setup
  • D.7 Per-Experiment Training Configurations
  • E Training Curve
  • F Additional Results
  • F.1 Extended Comparison Against Additional Baselines
  • F.2 Dense 32B Math Results
  • F.3 Full results on adversarial red-teaming
  • G Diversity Judge Prompt
  • H Distribution Matching Comparison
  • I Qualitative Comparison: GFlowRL-14B vs. DeepCoder-14B-Preview

Knowls

  1. Knowl 1 — GFlowRL replaces the learned partition network with a rollout-group estimate

    model/method

    GFlowRL trains a language-model policy to match a reward-tilted response distribution without learning a separate partition-function network. For prompt xx, let y(i)y^{(i)} be response ii in a group of GG responses sampled from the rollout policy πold\pi_{\mathrm{old}}; let πθ\pi_\theta be the trainable policy, πref\pi_{\mathrm{ref}} a frozen reference policy, r(x,y)r(x,y) the scalar reward, and β>0\beta>0 the inverse temperature. GFlowRL estimates the prompt's log partition term from that group as

    log⁡Z^t(x)=1G∑i=1G[βr(x,y(i))+log⁡πref(y(i)∣x)−log⁡πold(y(i)∣x)].\widehat{\log Z}_t(x)=\frac{1}{G}\sum_{i=1}^{G}\left[\beta r(x,y^{(i)})+\log\pi_{\mathrm{ref}}(y^{(i)}\mid x)-\log\pi_{\mathrm{old}}(y^{(i)}\mid x)\right].

    For each response, let LiL_i be its token length. The rollout-policy flow gap is

    gi=sg⁡ ⁣[log⁡Z^t(x)]+1Lilog⁡πold(y(i)∣x)πref(y(i)∣x)−βr(x,y(i)),g_i=\operatorname{sg}\!\left[\widehat{\log Z}_t(x)\right]+\frac{1}{L_i}\log\frac{\pi_{\mathrm{old}}(y^{(i)}\mid x)}{\pi_{\mathrm{ref}}(y^{(i)}\mid x)}-\beta r(x,y^{(i)}),

    where sg⁡\operatorname{sg} means stop-gradient. Clip this gap asymmetrically, g~i=clip⁡(gi,−ϵlow,+ϵhigh)\widetilde g_i=\operatorname{clip}(g_i,-\epsilon_{\mathrm{low}},+\epsilon_{\mathrm{high}}), and define the sequence-level importance weight wi=min⁡ ⁣(πθ(y(i)∣x)/πold(y(i)∣x),1+ϵIS)w_i=\min\!\left(\pi_\theta(y^{(i)}\mid x)/\pi_{\mathrm{old}}(y^{(i)}\mid x),1+\epsilon_{\mathrm{IS}}\right). The loss for prompt xx is

    L(θ;x)=1G∑i=1Gwi[g~i+1Lilog⁡πθ(y(i)∣x)πold(y(i)∣x)]2.\mathcal L(\theta;x)=\frac{1}{G}\sum_{i=1}^{G}w_i\left[\widetilde g_i+\frac{1}{L_i}\log\frac{\pi_\theta(y^{(i)}\mid x)}{\pi_{\mathrm{old}}(y^{(i)}\mid x)}\right]^2.

    The rollout-group estimate is a stop-gradient baseline: it centers the residual but has no trainable parameters or independent optimizer. Length normalization limits the influence of long reasoning sequences; importance weighting corrects for drift between rollout and training policies; and asymmetric clipping limits outlier flow gaps while allowing larger positive than negative corrections. The paper's default settings include G=16G=16, β=8\beta=8, ϵlow=0.2\epsilon_{\mathrm{low}}=0.2, ϵhigh=0.28\epsilon_{\mathrm{high}}=0.28, and sequence-level importance sampling with threshold 2.0.

  2. Knowl 2 — The unclipped, unnormalized GFlowRL objective has the reward-tilted distribution as a fixed point

    theoretical result

    Fix a prompt xx. Let πref(y∣x)\pi_{\mathrm{ref}}(y\mid x) be a reference response distribution, r(x,y)r(x,y) a scalar reward, and β>0\beta>0. Consider the population GFlowRL loss without response-length normalization and with flow-gap clipping inactive. If the policy class can represent the target distribution, and a self-consistent fixed point has πθ=πold\pi_\theta=\pi_{\mathrm{old}} and zero loss, then that fixed-point policy is

    πθ(y∣x)=πref(y∣x)exp⁡(βr(x,y))Z(x),Z(x)=∑y′πref(y′∣x)exp⁡(βr(x,y′)).\pi_\theta(y\mid x)=\frac{\pi_{\mathrm{ref}}(y\mid x)\exp(\beta r(x,y))}{Z(x)},\qquad Z(x)=\sum_{y'}\pi_{\mathrm{ref}}(y'\mid x)\exp(\beta r(x,y')).

    At this fixed point, the rollout-based log-partition estimate equals log⁡Z(x)\log Z(x). The result characterizes a zero-loss fixed point; it does not guarantee convergence of finite-batch, nonconvex training. With a finite rollout group of size GG, the paper states that the Monte Carlo estimate is unbiased for its population counterpart and has variance O(1/G)O(1/G). The theorem does not directly apply to the length-normalized objective: when all responses have common length LL, its fixed point corresponds to inverse temperature LβL\beta; when lengths vary, normalization introduces a length-dependent distortion.

  3. Knowl 3 — The learned partition network adds instability without measurable benefit in the diagnostic

    empirical result

    On Qwen2.5-7B mathematical reasoning, replacing FlowRL's learned log partition term with independent Gaussian noise, log⁡Z∼N(0.5,1)\log Z\sim\mathcal N(0.5,1), did not reduce average accuracy: the random-log-partition variant scored 36.19 versus 35.63 for FlowRL. This tests whether the learned prompt-conditional normalizer contributes useful information; the result suggests little benefit in this setting.

    Across 421 training steps, gradient norms differed sharply. FlowRL had minimum 1.58×1021.58\times10^2, maximum 9.59×10169.59\times10^{16}, mean 3.23×10143.23\times10^{14}, median 1.14×1031.14\times10^3, and standard deviation 4.89×10154.89\times10^{15}, with 55 steps at or above 10610^6. Excluding those explosion steps, its median was 1.1×1031.1\times10^3. GFlowRL's corresponding statistics were 3.11×10−23.11\times10^{-2}, 6.18, 6.98×10−26.98\times10^{-2}, 4.30×10−24.30\times10^{-2}, and 3.08×10−13.08\times10^{-1}, with no explosions; GRPO's were 1.68×10−11.68\times10^{-1}, 5.90, 2.38×10−12.38\times10^{-1}, 1.99×10−11.99\times10^{-1}, and 3.37×10−13.37\times10^{-1}, also with no explosions. The authors attribute FlowRL's instability to the unbounded, poorly calibrated learned log-partition term dominating policy updates. In a separate synthetic three-mode distribution-matching test, FlowRL and its random-log-partition variant behaved similarly and did not capture the multimodal target at 500 steps, whereas GFlowRL began to represent its multiple modes.

  4. Knowl 4 — GFlowRL improves dense-model mathematical reasoning results

    empirical result

    On six mathematical benchmarks, the authors evaluated Avg@16 accuracy for dense Qwen2.5 models. For Qwen2.5-7B, GFlowRL achieved an average of 40.92, compared with 35.63 for FlowRL and 32.48 for GRPO. The per-benchmark scores below are percentages as reported.

    Method AIME24 AIME25 AMC23 MATH500 Minerva Olympiad Avg
    GRPO 13.54 9.79 64.53 57.05 23.06 26.88 32.48
    FlowRL 15.41 10.83 54.53 66.96 31.41 34.61 35.63
    GFlowRL 17.29 9.79 67.66 76.89 33.62 40.25 40.92

    GFlowRL's average exceeded GRPO by 8.44 points and FlowRL by 5.29 points; it scored highest on five of the six listed benchmarks and tied GRPO on AIME25. On Qwen2.5-32B, the reported result table gives GFlowRL an average of 50.42, versus 48.39 for FlowRL, 43.26 for PPO, and 38.34 for GRPO. These results support the method's performance on dense-model math tasks at both tested scales.

  5. Knowl 5 — GFlowRL raises code benchmark performance and reaches 2048 Codeforces rating at 14B

    empirical result

    On the DeepSeek-R1-Distilled-Qwen-7B coding model, GFlowRL had the highest reported result on each code benchmark compared with the listed training baselines. It achieved LiveCodeBench Avg@16 38.62 and Pass@16 58.06; Codeforces rating 1646.21 and percentile 88.0%; and HumanEval+ Avg@16 84.93. FlowRL's corresponding results were 37.43, 56.27, 1549.47, 83.3%, and 83.28; GRPO's were 32.75, 52.32, 1313.82, 67.1%, and 80.13.

    At 14B, GFlowRL reached a Codeforces rating of 2048, compared with 1936 for DeepCoder-14B and 1904 for FlowRL-14B. For context, the reported ratings for OpenAI o1 and o3-mini were 1891 and 2073, respectively: GFlowRL was 157 points above o1 and 25 below o3-mini. The GFlowRL and DeepCoder comparison used the same data and two-stage curriculum, although GFlowRL was trained for 400 steps and DeepCoder for 600.

  6. Knowl 6 — GFlowRL attains the highest reported red-teaming success rates while FlowRL fails to converge

    empirical result

    In the authors' SEMA-based red-teaming experiments, attack success rate at one attempt (ASR@1) was measured against three instruction-tuned victim models: Qwen2.5-3B, Llama-3.1-8B, and GPT-4.1-mini. GFlowRL exceeded SEMA's average ASR@1 on both benchmarks, in the reported setting where FlowRL failed to converge and yielded no usable attacker.

    Benchmark Attacker Qwen2.5-3B Llama-3.1-8B GPT-4.1-mini Avg
    AdvBench SEMA 79.9 77.2 83.3 80.1
    AdvBench GFlowRL 80.2 81.2 86.1 82.5
    HarmBench SEMA 74.5 70.6 79.8 75.0
    HarmBench GFlowRL 79.9 73.0 85.5 79.5

    GFlowRL's average advantage over SEMA was 2.4 percentage points on AdvBench and 4.5 points on HarmBench. The paper presents this noisy- and sparse-reward setting as evidence that removing the learned partition term restores stable training where FlowRL does not converge.

  7. Knowl 7 — The same GFlowRL recipe scales to sparse MoE models up to 235B parameters

    empirical result

    GFlowRL was evaluated on Qwen3 mixture-of-experts models without changing the recipe across scales. On Qwen3-30B-A3B, with 3B active parameters, it reached Codeforces rating 1999 and improved the mathematical Avg@16 score from the backbone's 74.52 and GRPO's 75.78 to 78.32. On Qwen3-235B-A22B, the math average was 83.35, versus 81.29 for the backbone and 82.40 for GRPO. GFlowRL improved over GRPO on every listed math benchmark at both model scales. FlowRL failed to converge on both MoE models.

    The 235B result used the 30B configuration's hyperparameters, data, and schedule without modification. Because of GPU constraints, GFlowRL ran for 30 steps, compared with 100 steps for GRPO; the authors report that training had not yet plateaued when stopped. They describe this as stable GFlowNet-style RL training at 235B scale.

  8. Knowl 8 — Flow-gap clipping improves accuracy and controls gradient size

    empirical result

    In a Qwen2.5-7B math ablation trained for 100 steps, removing GFlowRL's asymmetric flow-gap clipping reduced average accuracy across six benchmarks from 40.92 to 37.02. The unclipped variant also had larger gradient norms: mean 0.601 versus 0.095, maximum 16.57 versus 6.184, and standard deviation 1.3785 versus 0.4301. These results support clipping as a stabilizer for outlier residuals in this experiment.

    A separate inverse-temperature ablation on the same benchmark set reported average accuracies of 37.9, 37.2, 38.2, 37.7, and 36.4 for β=1,5,8,10,15\beta=1,5,8,10,15, respectively. The highest tested score was at β=8\beta=8, which the authors use as the default.

  9. Knowl 9 — GFlowRL receives higher judged reasoning-diversity scores than reward-maximizing baselines

    empirical result

    The authors assessed diversity across 16 mathematical solutions per problem using GPT-o4-mini on a 1–5 scale, with scores averaged over five runs. GRPO scored 1.21, PPO 1.15, FlowRL 2.64, and GFlowRL 3.93. Thus, GFlowRL received the highest judged diversity score in this comparison, 3.2 times GRPO's score and about 1.5 times FlowRL's score. The paper reports variance of approximately 0.53 for the judge and interprets the score gap between GFlowRL and PPO/GRPO as substantially larger than this variation.

  10. Knowl 10 — Evaluation scope and estimator trade-offs remain limitations

    limitation

    The paper notes that an in-batch Monte Carlo estimate of the log partition term can have higher variance in principle, particularly when the rollout group is small. Importance-sampling correction and flow-gap clipping controlled training in the evaluated experiments, but the estimator's bias–variance trade-off and alternative lower-variance group estimators were not fully characterized. Experimental coverage was limited to language-model math reasoning, code reasoning, and red-teaming; whether the same estimator-based design generalizes to broader agentic or multimodal reinforcement-learning settings remains unknown.

Coverage note — The synthetic three-Gaussian plots are summarized with the partition-function diagnostic rather than reproduced in detail; per-problem Codeforces examples and full training-configuration tables are omitted because they corroborate the aggregate results or provide implementation context rather than distinct central contributions.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Emmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup, and Yoshua Bengio. Flow network based generative models for non-iterative diverse candidate generation. Neural Information Processing Systems (NeurIPS), 2021.
  3. 3.Yoshua Bengio, Salem Lahlou, Tristan Deleu, Edward J. Hu, Mo Tiwari, and Emmanuel Bengio. Gflownet foundations. Journal of Machine Learning Research, 24(210):1–55, 2023. URL http://jmlr.org/papers/v24/22-0364.html.
  4. 4.BytedTsinghua-SIA. Dapo-math-17k, 2025. URL https://huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k.
  5. 5.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  6. 6.Mingqian Feng, Xiaodong Liu, Weiwei Yang, Jialin Song, Xuekai Zhu, Chenliang Xu, and Jianfeng Gao. SEMA: Simple yet effective learning for multi-turn jailbreak attacks. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=6eSNG1VNkl.
  7. 7.Will Grathwohl, Dami Choi, Yuhuai Wu, Geoff Roeder, and David Duvenaud. Backpropagation through the void: Optimizing control variates for black-box gradient estimation. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=SyzKd1bCW.
  8. 8.Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, et al. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783.
  9. 9.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. Nature, 645:633–638, 2025a.
  10. 10.Weiyang Guo, Zesheng Shi, Zhuo Li, Yequan Wang, Xuebo Liu, Wenya Wang, Fangming Liu, Min Zhang, and Jing Li. Jailbreak-r1: Exploring the jailbreak capabilities of llms via reinforcement learning, 2025b. URL https://arxiv.org/abs/2506.00782.
  11. 11.Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. Olympiad-bench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems. In Annual Meeting of the Association for Computational Linguistics, 2024. URL https://api.semanticscholar.org/CorpusID:267770504.
  12. 12.Edward J Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar, Guillaume Lajoie, Yoshua Bengio, and Nikolay Malkin. Amortizing intractable inference in large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Ouj6p4ca60.
  13. 13.Jian Hu, Jason Klein Liu, Haotian Xu, and Wei Shen. Reinforce++: Stabilizing critic-free policy optimization with global advantage normalization. arXiv preprint arXiv:2501.03262, 2025.
  14. 14.HugginfaceH4. Math-500. URL https://huggingface.co/datasets/HuggingFaceH4/MATH-500.
  15. 15.Naman Jain, Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. In Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu, editors, International Conference on Learning Representations, volume 2025, pages 58791–58831, 2025. URL https://proceedings.iclr.cc/paper_files/paper/2025/file/94074dd5a072d28ff75a76dabed43767-Paper-Conference.pdf.
  16. 16.Koray Kavukcuoglu. Gemini 2.5: Our most intelligent AI model, 2025. URL https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/. Google Blog (The Keyword), Published Mar. 25, 2025.
  17. 17.Kimi-Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al. Kimi k1.5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599, 2025.
  18. 18.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. Solving quantitative reasoning problems with language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, 2022.
  19. 19.Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In The Twelfth International Conference on Learning Representations, 2023.
  20. 20.Yue Liu, Xiaoxin He, Miao Xiong, Jinlan Fu, Shumin Deng, Yingwei Ma, Jiaheng Zhang, and Bryan Hooi. Flipattack: jailbreak llms via flipping. In Proceedings of the 42nd International Conference on Machine Learning, ICML’25. JMLR.org, 2025.
  21. 21.LiveCodeBench. Livecodebench. URL https://github.com/LiveCodeBench/LiveCodeBench.
  22. 22.Michael Luo, Sijun Tan, Roy Huang, Xiaoxiang Shi, Rachel Xin, Colin Cai, Ameen Patel, Alpay Ariyak, Qingyang Wu, Ce Zhang, Li Erran Li, Raluca Ada Popa, Ion Stoica, and Tianjun Zhang. Deepcoder: A fully open-source 14b coder at o3-mini level. https://www.together.ai/blog/deepcoder, 2025. Notion Blog.
  23. 23.Kanika Madan, Jarrid Rector-Brooks, Maksym Korablyov, Emmanuel Bengio, Moksh Jain, Andrei Cristian Nica, Tom Bosc, Yoshua Bengio, and Nikolay Malkin. Learning gflownets from partial episodes for improved convergence and stability. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 23467–23483. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/v202/madan23a.html.
  24. 24.Nikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun, and Yoshua Bengio. Trajectory balance: Improved credit assignment in gflownets. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 5955–5967. Curran Associates, Inc., 2022. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/27b51baca8377a0cf109f6ecc15a0f70-Paper-Conference.pdf.
  25. 25.Nikolay Malkin, Salem Lahlou, Tristan Deleu, Xu Ji, Edward J Hu, Katie E Everett, Dinghuai Zhang, and Yoshua Bengio. GFlownets and variational inference. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=uKiE0VIluA-.
  26. 26.Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. Harmbench: a standardized evaluation framework for automated red teaming and robust refusal. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024.
  27. 27.Sobhan Mohammadpour, Emmanuel Bengio, Emma Frejinger, and Pierre-Luc Bacon. Maximum entropy GFlowNets with soft Q-learning. In Sanjoy Dasgupta, Stephan Mandt, and Yingzhen Li, editors, Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, volume 238 of Proceedings of Machine Learning Research, pages 2593–2601. PMLR, 02–04 May 2024. URL https://proceedings.mlr.press/v238/mohammadpour24a.html.
  28. 28.Nikolas Nüsken and Lorenz Richter. Solving high-dimensional hamilton–jacobi–bellman pdes using neural networks: Perspectives from the theory of controlled diffusions and measures on path space. Partial Differential Equations and Applications, 2, 2021. URL https://doi.org/10.1007/s42985-021-00102-x.
  29. 29.OpenAI. Introducing openai o1. https://openai.com/o1/, 2024. Reasoning-focused large language model.
  30. 30.OpenAI. Introducing openai o3 and o4-mini. https://openai.com/index/introducing-o3-and-o4-mini/, 2025. Large language model announcement.
  31. 31.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback. In Proceedings of the 36th International Conference on Neural Information Processing Systems, 2022.
  32. 32.Maya Pavlova, Erik Brinkman, Krithika Iyer, Vítor Albiero, Joanna Bitton, Hailey Nguyen, Cristian Canton Ferrer, Ivan Evtimov, and Aaron Grattafiori. Automated red teaming with GOAT: the generative offensive agent tester. In ICLR 2025 Workshop on Building Trust in Language Models and Applications, 2025. URL https://openreview.net/forum?id=6uyczU6S2M.
  33. 33.Guilherme Penedo, Anton Lozhkov, Hynek Kydlícek, Loubna Ben Allal, Edward Beeching, Agustín Piqueres Lajarín, Quentin Gallouédec, Nathan Habib, Lewis Tunstall, and Leandro von Werra. Codeforces. https://huggingface.co/datasets/open-r1/codeforces, 2025.
  34. 34.Qwen-Team. Qwen2.5: A party of foundation models, September 2024. URL https://qwen.ai/blog?id=qwen2.5.
  35. 35.Salman Rahman, Liwei Jiang, James Shiffer, Genglin Liu, Sheriff Issaka, Md Rizwan Parvez, Hamid Palangi, Kai-Wei Chang, Yejin Choi, and Saadia Gabriel. X-teaming: Multi-turn jailbreaks and defenses with adaptive multi-agents. In Second Conference on Language Modeling, 2025. URL https://openreview.net/forum?id=gKfj7Jb1kj.
  36. 36.Qibing Ren, Hao Li, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao, Lei Sha, Junchi Yan, Lizhuang Ma, and Jing Shao. LLMs know their vulnerabilities: Uncover safety gaps through natural distribution shifts. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 24763–24785, Vienna, Austria, July 2025. Association for Computational Linguistics. URL https://aclanthology.org/2025.acl-long.1207/.
  37. 37.Lorenz Richter, Ayman Boustati, Nikolas Nüsken, Francisco Ruiz, and Omer Deniz Akyildiz. Vargrad: A low-variance gradient estimator for variational inference. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 13481–13492, 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/9c22c0b51b3202246463e986c7e205df-Paper.pdf.
  38. 38.Mark Russinovich, Ahmed Salem, and Ronen Eldan. Great, now write an article about that: the crescendo multi-turn llm jailbreak attack. In Proceedings of the 34th USENIX Conference on Security Symposium, SEC ’25, USA, 2025. USENIX Association.
  39. 39.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  40. 40.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
  41. 41.Han Shen. On entropy control in LLM-RL algorithms. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=LqazVN5epT.
  42. 42.Chung-En Sun, Xiaodong Liu, Weiwei Yang, Tsui-Wei Weng, Hao Cheng, Aidan San, Michel Galley, and Jianfeng Gao. Advllm: Iterative self-tuning llms for enhanced jailbreaking capabilities, 2025. URL https://arxiv.org/abs/2410.18469.
  43. 43.Thinking Machines. Defeating nondeterminism in llm inference. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/, 2025.
  44. 44.Daniil Tiapkin, Nikita Morozov, Alexey Naumov, and Dmitry P Vetrov. Generative flow networks as entropy-regularized RL. In Sanjoy Dasgupta, Stephan Mandt, and Yingzhen Li, editors, Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, volume 238 of Proceedings of Machine Learning Research, pages 4213–4221. PMLR, 02–04 May 2024. URL https://proceedings.mlr.press/v238/tiapkin24a.html.
  45. 45.George Tucker, Andriy Mnih, Chris Maddison, John Lawson, and Jascha Sohl-Dickstein. Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/ebd6d2f5d60ff9afaeda1a81fc53e2d0-Paper.pdf.
  46. 46.Siddarth Venkatraman, Minsu Kim, Luke Rowe, Moksh Jain, Marcin Sendera, Sarthak Mittal, Luca Scimeca, Mohsin Hasan, Pablo Lemos, Emmanuel Bengio, Alexandre Adam, Jarrid Rector-Brooks, Yoshua Bengio, Glen Berseth, and Nikolay Malkin. Amortizing intractable inference in diffusion models for vision, language, and control. In Proceedings of the 38th International Conference on Neural Information Processing Systems. Curran Associates Inc., 2024.
  47. 47.Lex Weaver and Nigel Tao. The optimal reward baseline for gradient-based reinforcement learning. In Proceedings of the Seventeenth Conference on Uncertainty in Artificial Intelligence, UAI’01, page 538–545. Morgan Kaufmann Publishers Inc., 2001.
  48. 48.Zixuan Weng, Xiaolong Jin, Jinyuan Jia, and Xiangyu Zhang. Foot-in-the-door: A multi-turn jailbreak for LLMs. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 1939–1950. Association for Computational Linguistics, 2025. URL https://aclanthology.org/2025.emnlp-main.100/.
  49. 49.Ronald J Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3):229–256, 1992.
  50. 50.Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M Bayen, Sham Kakade, Igor Mordatch, and Pieter Abbeel. Variance reduction for policy gradient with action-dependent factorized baselines. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=H1tSsb-AW.
  51. 51.Weijia Xu, Nebojsa Jojic, Sudha Rao, Chris Brockett, and Bill Dolan. Echoes in ai: Quantifying lack of plot diversity in llm outputs. Proceedings of the National Academy of Sciences, 122(35):e2504966122, 2025. URL https://www.pnas.org/doi/abs/10.1073/pnas.2504966122.
  52. 52.Hao Yang, Lizhen Qu, Ehsan Shareghi, and Gholamreza Haffari. Jigsaw puzzles: Splitting harmful questions to jailbreak large language models, 2024a. URL https://arxiv.org/abs/2410.11459.
  53. 53.Xikang Yang, Xuehai Tang, Songlin Hu, and Jizhong Han. Chain of attack: a semantic-driven contextual multi-turn attacker for llm, 2024b. URL https://arxiv.org/abs/2405.05610.
  54. 54.Feng Yao, Liyuan Liu, Dinghuai Zhang, Chengyu Dong, Jingbo Shang, and Jianfeng Gao. Your efficient rl framework secretly brings you off-policy rl training. https://fengyao.notion.site/off-policy-rl, 2025. Work in progress.
  55. 55.Fangxu Yu, Lai Jiang, Haoqiang Kang, Shibo Hao, and Lianhui Qin. Flow of reasoning: Training LLMs for divergent reasoning with minimal examples. In Forty-second International Conference on Machine Learning, 2025a. URL https://openreview.net/forum?id=qyMxunrR2j.
  56. 56.Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, juncai liu, LingJun Liu, Xin Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, editors, Advances in Neural Information Processing Systems, volume 38, pages 113222–113244. Curran Associates, Inc., 2025b. URL https://proceedings.neurips.cc/paper_files/paper/2025/file/a4277440d50f1f15d2cb4c14f7e0c0d2-Paper-Conference.pdf.
  57. 57.David W Zhang, Corrado Rainone, Markus Peschl, and Roberto Bondesan. Robust scheduling with GFlownets. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=ZBUthI6wK9h.
  58. 58.Dinghuai Zhang, Nikolay Malkin, Zhen Liu, Alexandra Volokhova, Aaron Courville, and Yoshua Bengio. Generative flow networks for discrete probabilistic modeling. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, editors, Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 26412–26428. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/zhang22v.html.
  59. 59.Yifan Zhang and Team Math-AI. Minerva math. URL https://huggingface.co/datasets/math-ai/minervamath.
  60. 60.Yifan Zhang and Team Math-AI. American mathematics competitions (amc) 2023, 2023. URL https://huggingface.co/datasets/math-ai/amc23.
  61. 61.Yifan Zhang and Team Math-AI. American invitational mathematics examination (aime) 2024, 2024.
  62. 62.Yifan Zhang and Team Math-AI. American invitational mathematics examination (aime) 2025, 2025.
  63. 63.Xuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li, Kaiyan Zhang, Che Jiang, Youbang Sun, Ermo Hua, Yuxin Zuo, Xingtai Lv, Qizheng Zhang, Lin Chen, Fanghao Shao, Bo Xue, Yunchong Song, Zhenjie Yang, Ganqu Cui, Ning Ding, Jianfeng Gao, Xiaodong Liu, Bowen Zhou, Hongyuan Mei, and Zhouhan Lin. FlowRL: Matching reward distributions for LLM reasoning. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=lObnTKbm9U.
  64. 64.Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models, 2023.

Citation

MLA
Liu, X., et al. “GFlowRL: Scaling Distribution-Matching RL to Large Language Models”. arXiv, 2026, http://arxiv.org/abs/2607.13394v1.
APA
Liu, X., Xu, M., Stokes, J. W., Smolensky, P., Burger, D., & Gao, J. (2026). GFlowRL: Scaling Distribution-Matching RL to Large Language Models. arXiv. http://arxiv.org/abs/2607.13394v1
Chicago
Liu, X., M. Xu, J. W. Stokes, P. Smolensky, D. Burger, and J. Gao. 2026. “GFlowRL: Scaling Distribution-Matching RL to Large Language Models”. arXiv. http://arxiv.org/abs/2607.13394v1.
Harvard
Liu, X. et al. (2026) “GFlowRL: Scaling Distribution-Matching RL to Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2607.13394v1.
Vancouver
1. Liu X, Xu M, Stokes JW, Smolensky P, Burger D, Gao J (2026) GFlowRL: Scaling Distribution-Matching RL to Large Language Models. arXiv

BibTeX

@article{liu2026gflowrl,
  title = {GFlowRL: Scaling Distribution-Matching RL to Large Language Models},
  author = {Liu, Xiaodong and Xu, Michael and Stokes, Jack W. and Smolensky, Paul and Burger, Doug and Gao, Jianfeng},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2607.13394v1},
  eprint = {2607.13394}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/