GFlowRL: Scaling Distribution-Matching RL to Large Language Models
Xiaodong LiuMichael XuJack W. StokesPaul SmolenskyDoug BurgerJianfeng Gao
Introduces GFlowRL, a reinforcement learning algorithm that eliminates the instability of distribution-matching methods by replacing learned partition functions with in-batch Monte Carlo estimates, enabling GFlowNet-style reasoning to scale across dense and mixture-of-experts models up to 235B parameters.
Post-training reinforcement learning is the primary driver of performance in state-of-the-art reasoning language models. However, standard reward-maximizing algorithms such as Proximal Policy Optimization and Group Relative Policy Optimization tend to cause mode collapse by narrowing the model's focus to a single dominant solution path rather than exploring diverse valid reasoning trajectories. While Generative Flow Network (GFlowNet) methods can encourage diverse reasoning by sampling outputs proportional to rewards, prior attempts to apply them to large models have suffered from severe optimization instability, gradient spikes, and high distributed-systems overhead.
The article demonstrates that the primary bottleneck in scaling GFlowNet-style reinforcement learning is an unnecessary auxiliary component: the learned partition function network. To resolve this, the authors evaluate and propose GFlowRL, a streamlined algorithm designed to stabilize and scale distribution-matching reinforcement learning across dense and mixture-of-experts architectures.
The authors conducted diagnostic and large-scale empirical experiments across mathematical reasoning, competitive programming, and adversarial security red-teaming. They evaluated models ranging from 7 billion to 235 billion parameters on standard benchmarks, including Codeforces, LiveCodeBench, AIME, AdvBench, and HarmBench. The core technical approach replaces the learned partition network entirely with a non-parametric in-batch Monte Carlo estimator computed directly from standard rollout groups, augmented with importance-sampling correction and asymmetric flow-gap clipping to stabilize training.
The investigation produced four central findings. First, diagnostic tests revealed that the learned partition function acts primarily as an uninformative noise source; replacing it with pure random noise yielded comparable performance (36.19% vs. 35.61% accuracy), while prior methods suffered 55 gradient explosions exceeding 10^6 across 421 steps compared to zero for GFlowRL. Second, on 14-billion parameter dense coding models, GFlowRL achieved a 2048 Codeforces Elo rating—outperforming DeepCoder-14B by 112 points, FlowRL by 144 points, and OpenAI o1 by 157 points, coming within 25 points of o3-mini. Third, GFlowRL maintained stability in noisy-reward red-teaming environments where prior GFlowNet approaches diverged, attaining leading attack success rates on AdvBench (82.5%) and HarmBench (79.5%). Fourth, the framework scaled seamlessly to sparse mixture-of-experts models up to 235 billion parameters, whereas prior methods failed to converge.
These findings indicate that generative distribution matching can be achieved efficiently without the computational overhead or instability of auxiliary neural networks. By integrating directly into existing rollout pipelines, GFlowRL reduces training risk, prevents costly optimization divergence, and significantly enhances solution diversity—scoring 3.93 out of 5.0 in diversity evaluations compared to 1.21 for reward-maximizing baselines.
Organizations developing reasoning-focused language models should consider replacing unstable partition networks with in-batch Monte Carlo estimation pipelines for post-training workflows. Because GFlowRL uses the same rollout group structure as Group Relative Policy Optimization, engineering teams can adopt the distribution-matching objective with minimal architectural disruption. Further exploration is recommended to evaluate whether this estimation paradigm extends effectively to broader agentic workflows and multimodal tasks.
The conclusions are supported by extensive evaluations across multiple benchmarks and parameter scales. A primary operational limitation is that the in-batch estimator's variance depends on the rollout group size, though the included clipping mechanisms effectively mitigate instability in practice. Stakeholders can have high confidence in these results for dense and sparse reasoning architectures within the evaluated task domains.
- Paper: GFlowNet Foundations, Yoshua Bengio et al. (2023). Read this foundation first to understand how GFlowNets match sampling probabilities to rewards—the core objective that GFlowRL streamlines for LLM training.
- Paper: Generative Flow Networks for Discrete Probabilistic Modeling, Dinghuai Zhang et al. (2022). Its discrete-data GFlowNet formulation clarifies how reward-proportional sampling works in token-like spaces, a useful prerequisite for following the paper’s LLM adaptation.
No sufficiently relevant recommendations were found.
