Flow Q-Learning
Seohong ParkQiyang LiSergey Levine
Proposes Flow Q-learning, a method that distills a multi-step flow-matching behavioral model into an expressive one-step actor to bypass unstable backpropagation through time and achieve fast, high-performing offline reinforcement learning.
Data-driven or offline reinforcement learning enables organizations to train autonomous decision-making agents directly from static historical datasets without the cost, delay, or safety risks of live trial-and-error exploration. As historical datasets expand in scale and complexity, the distribution of recorded behaviors becomes increasingly complex and varied. Traditional approaches using simple Gaussian assumptions often fail to capture these intricate patterns, while advanced generative approaches like diffusion models and flow matching are computationally expensive and unstable to optimize using standard value-maximization techniques.
The article demonstrates an offline reinforcement learning framework called Flow Q-Learning (FQL), which integrates expressive generative flow policies into an efficient decision-making pipeline. The core objective is to achieve the high representational capacity of multi-step flow models while maintaining the speed, stability, and simplicity of single-step policy extraction.
To achieve this, the approach decouples behavioral modeling from value optimization. Instead of forcing a multi-step iterative flow model to maximize reward values directly—a process that requires unstable and computationally expensive recursive calculations—FQL trains the flow model purely to clone behaviors from the dataset. Simultaneously, it trains a separate, single-step policy to maximize expected rewards while using distillation to anchor its actions to the flow model's learned distribution. The resulting system produces a fast, single-step policy for operational deployment, avoiding iterative generation at test time and bypassing unstable backpropagation during training.
The empirical evaluation across 73 benchmark locomotion and manipulation tasks yielded several key findings:
- FQL achieved the highest or near-highest performance across benchmark suites, notably scoring an 84% success rate on the challenging D4RL large navigation benchmark, where previous distillation baselines achieved 0% to 54%.
- On complex robotic manipulation tasks characterized by varied, multi-option behaviors, FQL outperformed leading Gaussian-based methods (e.g., reaching 96% versus 91% on single-cube manipulation and 29% versus 12% on double-cube manipulation).
- Across 50 state-based benchmark tasks, FQL's one-step extraction mechanism achieved an aggregate score of 44%, outperforming alternative flow policy extraction strategies such as rejection sampling (30%), recursive backpropagation (29%), and weighted regression (16%).
- When fine-tuned with additional live interactions, FQL adapted seamlessly without structural modifications, matching or outperforming specialized online fine-tuning methods.
- In computational efficiency tests, FQL was significantly faster at inference than multi-step diffusion or rejection-sampling baselines, executing at speeds nearly identical to simple Gaussian policies.
These findings indicate that teams deploying autonomous decision-making systems no longer need to compromise between policy expressiveness and computational speed. Because FQL produces a single-step policy, it significantly reduces test-time latency and hardware computing costs in time-critical operational settings while minimizing training instability.
For practical implementation, engineering teams can adopt FQL as an effective, drop-in framework for offline policy optimization. Practitioners should prioritize tuning the single behavioral regularization coefficient, which controls the trade-off between conservatism and reward maximization based on dataset quality, while keeping default flow parameters such as uniform time distributions.
Confidence in these findings is high for simulated control and navigation environments across state and visual inputs. However, stakeholders should note that FQL has not yet been evaluated on physical hardware in real-world environments, relies on numerical solvers during training distillation, and lacks an intrinsic exploration mechanism to escape local optima during online fine-tuning on certain combinatorial tasks. Initial real-world pilot deployments are recommended to validate performance transfer outside simulation.
- Paper: Flow Matching for Generative Modeling, Yaron Lipman et al. (2023). Introduces continuous normalizing flow matching for generative modeling, providing the fundamental generative framework adapted by Flow Q-Learning for action representations.
- Paper: Offline Reinforcement Learning with Implicit Q-Learning, Ilya Kostrikov et al. (2021). Establishes in-sample dynamic programming via expectile regression, which underpins the offline Q-learning framework utilized in Flow Q-Learning.
- Paper: Efficient Diffusion Policies For Offline Reinforcement Learning, Bingyi Kang et al. (2023). Addresses the slow inference and recursive backpropagation bottlenecks of iterative generative policies in offline RL, motivating FQL's one-step policy formulation.
- Paper: Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow, Xingchao Liu et al. (2023). Develops rectified flow methods that straighten probability paths between distributions to enable highly accurate one-step sampling.
- Paper: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, Sergey Levine et al. (2020). Provides a comprehensive foundation on out-of-distribution action extrapolation and distributional shift challenges in offline reinforcement learning.
- Paper: Conservative Q-Learning for Offline Reinforcement Learning, Aviral Kumar et al. (2020). Formulates standard value-regularization techniques for offline Q-learning benchmarks evaluated in the source paper.
- Paper: D4RL: Datasets for Deep Data-Driven Reinforcement Learning, Justin Fu et al. (2020). Introduces the D4RL benchmark datasets used to empirically validate and compare the performance of Flow Q-Learning.
- Paper: Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control, Carles Domingo-Enrich et al. (2025). Extends the optimization of continuous flow and diffusion policies by introducing memoryless stochastic optimal control for reward fine-tuning without backpropagation bias.
- Paper: Mean Flows for One-step Generative Modeling, Zhengyang Geng et al. (2025). Generalizes one-step flow generation by modeling average velocity fields directly to accelerate generative sampling from scratch.
