Flow Q-Learning

Seohong ParkQiyang LiSergey Levine

article2025ICML180 citations

Proposes Flow Q-learning, a method that distills a multi-step flow-matching behavioral model into an expressive one-step actor to bypass unstable backpropagation through time and achieve fast, high-performing offline reinforcement learning.

Listen

Data-driven or offline reinforcement learning enables organizations to train autonomous decision-making agents directly from static historical datasets without the cost, delay, or safety risks of live trial-and-error exploration. As historical datasets expand in scale and complexity, the distribution of recorded behaviors becomes increasingly complex and varied. Traditional approaches using simple Gaussian assumptions often fail to capture these intricate patterns, while advanced generative approaches like diffusion models and flow matching are computationally expensive and unstable to optimize using standard value-maximization techniques.

The article demonstrates an offline reinforcement learning framework called Flow Q-Learning (FQL), which integrates expressive generative flow policies into an efficient decision-making pipeline. The core objective is to achieve the high representational capacity of multi-step flow models while maintaining the speed, stability, and simplicity of single-step policy extraction.

To achieve this, the approach decouples behavioral modeling from value optimization. Instead of forcing a multi-step iterative flow model to maximize reward values directly—a process that requires unstable and computationally expensive recursive calculations—FQL trains the flow model purely to clone behaviors from the dataset. Simultaneously, it trains a separate, single-step policy to maximize expected rewards while using distillation to anchor its actions to the flow model's learned distribution. The resulting system produces a fast, single-step policy for operational deployment, avoiding iterative generation at test time and bypassing unstable backpropagation during training.

The empirical evaluation across 73 benchmark locomotion and manipulation tasks yielded several key findings:

  • FQL achieved the highest or near-highest performance across benchmark suites, notably scoring an 84% success rate on the challenging D4RL large navigation benchmark, where previous distillation baselines achieved 0% to 54%.
  • On complex robotic manipulation tasks characterized by varied, multi-option behaviors, FQL outperformed leading Gaussian-based methods (e.g., reaching 96% versus 91% on single-cube manipulation and 29% versus 12% on double-cube manipulation).
  • Across 50 state-based benchmark tasks, FQL's one-step extraction mechanism achieved an aggregate score of 44%, outperforming alternative flow policy extraction strategies such as rejection sampling (30%), recursive backpropagation (29%), and weighted regression (16%).
  • When fine-tuned with additional live interactions, FQL adapted seamlessly without structural modifications, matching or outperforming specialized online fine-tuning methods.
  • In computational efficiency tests, FQL was significantly faster at inference than multi-step diffusion or rejection-sampling baselines, executing at speeds nearly identical to simple Gaussian policies.

These findings indicate that teams deploying autonomous decision-making systems no longer need to compromise between policy expressiveness and computational speed. Because FQL produces a single-step policy, it significantly reduces test-time latency and hardware computing costs in time-critical operational settings while minimizing training instability.

For practical implementation, engineering teams can adopt FQL as an effective, drop-in framework for offline policy optimization. Practitioners should prioritize tuning the single behavioral regularization coefficient, which controls the trade-off between conservatism and reward maximization based on dataset quality, while keeping default flow parameters such as uniform time distributions.

Confidence in these findings is high for simulated control and navigation environments across state and visual inputs. However, stakeholders should note that FQL has not yet been evaluated on physical hardware in real-world environments, relies on numerical solvers during training distillation, and lacks an intrinsic exploration mechanism to escape local optima during online fine-tuning on certain combinatorial tasks. Initial real-world pilot deployments are recommended to validate performance transfer outside simulation.

arXiv: 2502.02538
Cover for Flow Q-Learning

Abstract

We present flow Q-learning (FQL), a simple and performant offline reinforcement learning (RL) method that leverages an expressive flow-matching policy to model arbitrarily complex action distributions in data. Training a flow policy with RL is a tricky problem, due to the iterative nature of the action generation process. We address this challenge by training an expressive one-step policy with RL, rather than directly guiding an iterative flow policy to maximize values. This way, we can completely avoid unstable recursive backpropagation, eliminate costly iterative action generation at test time, yet still mostly maintain expressivity. We experimentally show that FQL leads to strong performance across 73 challenging state- and pixel-based OGBench and D4RL tasks in offline RL and offline-to-online RL.

https://seohong.me/projects/fql/

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 3. Flow Q-Learning
  • 4. Prior Work
  • 4.1. How Have Previous Works Trained Diffusion and Flow Policies with RL?
  • 5. Experiments
  • 5.1. Experimental Setup
  • 5.2. Results and Q&As
  • 6. Closing Remarks
  • Acknowledgments
  • Impact Statement
  • References
  • A. Limitations
  • B. Implementation Details
  • C. Ablation Study
  • D. Additional Results
  • E. Experimental Details
  • E.1. Environments, Tasks, and Datasets
  • E.2. Methods and Hyperparameters

Knowls

  1. Knowl 1 — Flow Q-Learning (FQL) with One-Step Guidance

    model/method

    Flow Q-Learning (FQL) is an offline reinforcement learning method for continuous action spaces A=Rd\mathcal{A} = \mathbb{R}^d and state spaces S\mathcal{S} using static transition datasets D={(s,a,r,s′)}\mathcal{D} = \{(s, a, r, s')\}. Rather than optimizing an iterative flow-matching policy directly with reinforcement learning (which requires computationally expensive and numerically unstable backpropagation through time across ODE integration steps), FQL decouples behavioral modeling from value optimization via one-step guidance.

    FQL consists of three primary parametric components:

    1. Behavioral Cloning (BC) Flow Policy: A state- and time-dependent velocity field vθ(t,s,x):[0,1]×S×Rd→Rdv_\theta(t, s, x): [0, 1] \times \mathcal{S} \times \mathbb{R}^d \to \mathbb{R}^d parameterized by θ\theta. It is trained solely via the conditional flow matching objective on dataset actions: LFlow(θ)=Es,a=x1∼D, x0∼N(0,Id), t∼Unif([0,1])[∥vθ(t,s,xt)−(x1−x0)∥22]\mathcal{L}_{\text{Flow}}(\theta) = \mathbb{E}_{s, a=x_1 \sim \mathcal{D},\, x_0 \sim \mathcal{N}(0, I_d),\, t \sim \text{Unif}([0, 1])} \left[ \|v_\theta(t, s, x_t) - (x_1 - x_0)\|^2_2 \right] where xt=(1−t)x0+tx1x_t = (1-t)x_0 + tx_1. Let μθ(s,z)=ψθ(1,s,z)\mu_\theta(s, z) = \psi_\theta(1, s, z) denote the deterministic terminal action generated by numerically integrating the ODE ddtψθ(t,s,z)=vθ(t,s,ψθ(t,s,z))\frac{d}{dt}\psi_\theta(t, s, z) = v_\theta(t, s, \psi_\theta(t, s, z)) from initial noise ψθ(0,s,z)=z∼N(0,Id)\psi_\theta(0, s, z) = z \sim \mathcal{N}(0, I_d) over t∈[0,1]t \in [0, 1].

    2. Action-Value Critic: A state-action value network Qϕ(s,a):S×A→RQ_\phi(s, a): \mathcal{S} \times \mathcal{A} \to \mathbb{R} with target network QϕˉQ_{\bar{\phi}}, trained via the Bellman error using the one-step policy μω\mu_\omega for next-action evaluation: LQ(ϕ)=Es,a,r,s′∼D, z∼N(0,Id), a′=μω(s′,z)[(Qϕ(s,a)−r−γQϕˉ(s′,a′))2]\mathcal{L}_Q(\phi) = \mathbb{E}_{s, a, r, s' \sim \mathcal{D},\, z \sim \mathcal{N}(0, I_d),\, a' = \mu_\omega(s', z)} \left[ \left( Q_\phi(s, a) - r - \gamma Q_{\bar{\phi}}(s', a') \right)^2 \right] where γ∈[0,1)\gamma \in [0, 1) is the discount factor.

    3. One-Step Policy: A direct regression policy network μω(s,z):S×Rd→A\mu_\omega(s, z): \mathcal{S} \times \mathbb{R}^d \to \mathcal{A} parameterized by ω\omega, mapping noise z∼N(0,Id)z \sim \mathcal{N}(0, I_d) in one forward pass to an action aπ=μω(s,z)a^\pi = \mu_\omega(s, z). The actor is trained to maximize critic values while being regularized by distilling from the BC flow policy output μθ(s,z)\mu_\theta(s, z): Lπ(ω)=Es∼D, z∼N(0,Id), aπ=μω(s,z)[−Qϕ(s,aπ)+α∥aπ−μθ(s,z)∥22]\mathcal{L}_\pi(\omega) = \mathbb{E}_{s \sim \mathcal{D},\, z \sim \mathcal{N}(0, I_d),\, a^\pi = \mu_\omega(s, z)} \left[ -Q_\phi(s, a^\pi) + \alpha \|a^\pi - \mu_\theta(s, z)\|^2_2 \right] where α>0\alpha > 0 is a scalar hyperparameter balancing value maximization and behavioral distillation.

    At test time, actions are generated in a single step via a=μω(s,z)a = \mu_\omega(s, z) with z∼N(0,Id)z \sim \mathcal{N}(0, I_d), eliminating iterative numerical ODE integration.

  2. Knowl 2 — Distillation Loss as an Upper Bound on Squared 2-Wasserstein Behavioral Regularization

    theoretical result

    In offline reinforcement learning, behavioral regularization prevents the actor policy from generating out-of-distribution actions. In Flow Q-Learning (FQL), the policy distillation loss serves as a metric-aware behavioral regularizer bounded by the squared 2-Wasserstein distance.

    Let ξ∼N(0,Id)\xi \sim \mathcal{N}(0, I_d) be a dd-dimensional standard normal random variable. For a given state s∈Ss \in \mathcal{S}, let πθ(⋅∣s)∈Δ(A)\pi_\theta(\cdot \mid s) \in \Delta(\mathcal{A}) and πω(⋅∣s)∈Δ(A)\pi_\omega(\cdot \mid s) \in \Delta(\mathcal{A}) be the push-forward probability distributions of ξ\xi under the behavioral flow generator μθ(s,⋅)\mu_\theta(s, \cdot) and the one-step generator μω(s,⋅)\mu_\omega(s, \cdot), respectively. Let Λ(πω,πθ)\Lambda(\pi_\omega, \pi_\theta) denote the set of all joint coupling distributions on A×A\mathcal{A} \times \mathcal{A} with marginals πω\pi_\omega and πθ\pi_\theta, and let W2\mathcal{W}_2 denote the 2-Wasserstein distance under the standard Euclidean metric on A=Rd\mathcal{A} = \mathbb{R}^d.

    The policy distillation loss LDistill(ω)\mathcal{L}_{\text{Distill}}(\omega) is an upper bound on the expected squared 2-Wasserstein distance: LDistill(ω)=Es∼D, z∼N(0,Id)[∥μω(s,z)−μθ(s,z)∥22]≥Es∼D[inf⁡λ∈Λ(πω,πθ)E(x,y)∼λ[∥x−y∥22]]=Es∼D[W2(πω(⋅∣s),πθ(⋅∣s))2]\mathcal{L}_{\text{Distill}}(\omega) = \mathbb{E}_{s \sim \mathcal{D},\, z \sim \mathcal{N}(0, I_d)} \left[ \|\mu_\omega(s, z) - \mu_\theta(s, z)\|^2_2 \right] \ge \mathbb{E}_{s \sim \mathcal{D}} \left[ \inf_{\lambda \in \Lambda(\pi_\omega, \pi_\theta)} \mathbb{E}_{(x, y) \sim \lambda} \left[ \|x - y\|^2_2 \right] \right] = \mathbb{E}_{s \sim \mathcal{D}} \left[ \mathcal{W}_2(\pi_\omega(\cdot \mid s), \pi_\theta(\cdot \mid s))^2 \right]

    Unlike Kullback-Leibler divergence DKLD_{\text{KL}} (used in TD3+BC and AWAC) and χ2\chi^2-divergence (used in Conservative Q-Learning), which are coordinate- and metric-agnostic divergences over probability spaces, the 2-Wasserstein regularizer incorporates geometric distance information in action space, penalizing deviations proportional to their squared Euclidean displacement.

  3. Knowl 3 — Flow Q-Learning (FQL) Algorithm

    algorithm

    Flow Q-Learning (FQL) trains an action-value critic QϕQ_\phi, a velocity field vθv_\theta for a behavioral flow policy πθ\pi_\theta, and an expressive one-step policy μω\mu_\omega.

    function μθ(s,z)\mu_\theta(s, z)
        for t=0,1,…,M−1t = 0, 1, \dots, M - 1 do
            z←z+vθ(t/M,s,z)/Mz \leftarrow z + v_\theta(t / M, s, z) / M
        return zz
    Input: Offline dataset D={(s,a,r,s′)}\mathcal{D} = \{(s, a, r, s')\}, discount γ\gamma, regularization weight α\alpha, Euler step count MM
    Output: Trained one-step actor policy μω\mu_\omega
    Initialize critic parameters ϕ\phi, target critic ϕˉ←ϕ\bar{\phi} \leftarrow \phi, flow parameters θ\theta, one-step policy parameters ω\omega
    while not converged do
        Sample minibatch of transitions {(s,a,r,s′)}∼D\{(s, a, r, s')\} \sim \mathcal{D}
        
        // Critic Update
        Sample z∼N(0,Id)z \sim \mathcal{N}(0, I_d)
        a′←μω(s′,z)a' \leftarrow \mu_\omega(s', z)
        Update ϕ\phi to minimize E[(Qϕ(s,a)−r−γQϕˉ(s′,a′))2]\mathbb{E}\left[\left(Q_\phi(s, a) - r - \gamma Q_{\bar{\phi}}(s', a')\right)^2\right]
        Update target network: ϕˉ←τϕ+(1−τ)ϕˉ\bar{\phi} \leftarrow \tau \phi + (1 - \tau) \bar{\phi}
        
        // Flow BC Update
        Sample x0∼N(0,Id)x_0 \sim \mathcal{N}(0, I_d)
        x1←ax_1 \leftarrow a
        Sample t∼Unif([0,1])t \sim \text{Unif}([0, 1])
        xt←(1−t)x0+tx1x_t \leftarrow (1 - t)x_0 + t x_1
        Update θ\theta to minimize E[∥vθ(t,s,xt)−(x1−x0)∥22]\mathbb{E}\left[\|v_\theta(t, s, x_t) - (x_1 - x_0)\|^2_2\right]
        
        // One-Step Actor Update
        Sample z∼N(0,Id)z \sim \mathcal{N}(0, I_d)
        aπ←μω(s,z)a^\pi \leftarrow \mu_\omega(s, z)
        Update ω\omega to minimize E[−Qϕ(s,aπ)+α∥aπ−μθ(s,z)∥22]\mathbb{E}\left[-Q_\phi(s, a^\pi) + \alpha \|a^\pi - \mu_\theta(s, z)\|^2_2\right]
    return One-step policy μω\mu_\omega

    In standard practice, two critic networks Qϕ1Q_{\phi_1} and Qϕ2Q_{\phi_2} are maintained. The Q loss term in the actor loss uses their mean 12(Qϕ1(s,aπ)+Qϕ2(s,aπ))\frac{1}{2}(Q_{\phi_1}(s, a^\pi) + Q_{\phi_2}(s, a^\pi)). For the critic Bellman target, the mean target 12(Qϕˉ1(s′,a′)+Qϕˉ2(s′,a′))\frac{1}{2}(Q_{\bar{\phi}_1}(s', a') + Q_{\bar{\phi}_2}(s', a')) is used by default, while the minimum min⁡(Qϕˉ1(s′,a′),Qϕˉ2(s′,a′))\min(Q_{\bar{\phi}_1}(s', a'), Q_{\bar{\phi}_2}(s', a')) (clipped double Q-learning) is selected for adroit and large antmaze tasks. Default Euler integration steps are set to M=10M = 10.

  4. Knowl 4 — Offline Reinforcement Learning Performance on OGBench and D4RL Benchmarks

    data/table

    Across 73 continuous control tasks—including 50 state-based OGBench tasks, 5 pixel-based OGBench visual manipulation tasks, 6 D4RL antmaze tasks, and 12 D4RL adroit tasks—Flow Q-Learning (FQL) achieves the highest or near-highest performance across task categories, particularly in multimodal manipulation.

    Task Category Gaussian Policies Diffusion Policies Flow Policies
    BC IQL ReBRAC IDQL SRPO CAC FAWAC FBRAC IFQL FQL
    OGBench antmaze-large (5 tasks) 11±111 \pm 1 53±353 \pm 3 81±581 \pm 5 21±521 \pm 5 11±411 \pm 4 33±433 \pm 4 6±16 \pm 1 60±660 \pm 6 28±528 \pm 5 79±379 \pm 3
    OGBench antmaze-giant (5 tasks) 0±00 \pm 0 4±14 \pm 1 26±826 \pm 8 0±00 \pm 0 0±00 \pm 0 0±00 \pm 0 0±00 \pm 0 4±44 \pm 4 3±23 \pm 2 9±69 \pm 6
    OGBench humanoidmaze-med (5 tasks) 2±12 \pm 1 33±233 \pm 2 22±822 \pm 8 1±01 \pm 0 1±11 \pm 1 53±853 \pm 8 19±119 \pm 1 38±538 \pm 5 60±1460 \pm 14 58±558 \pm 5
    OGBench humanoidmaze-large (5 tasks) 1±01 \pm 0 2±12 \pm 1 2±12 \pm 1 1±01 \pm 0 0±00 \pm 0 0±00 \pm 0 0±00 \pm 0 2±02 \pm 0 11±211 \pm 2 4±24 \pm 2
    OGBench antsoccer-arena (5 tasks) 1±01 \pm 0 8±28 \pm 2 0±00 \pm 0 12±412 \pm 4 1±01 \pm 0 2±42 \pm 4 12±012 \pm 0 16±116 \pm 1 33±633 \pm 6 60±260 \pm 2
    OGBench cube-single (5 tasks) 5±15 \pm 1 83±383 \pm 3 91±291 \pm 2 95±295 \pm 2 80±580 \pm 5 85±985 \pm 9 81±481 \pm 4 79±779 \pm 7 79±279 \pm 2 96±196 \pm 1
    OGBench cube-double (5 tasks) 2±12 \pm 1 7±17 \pm 1 12±112 \pm 1 15±615 \pm 6 2±12 \pm 1 6±26 \pm 2 5±25 \pm 2 15±315 \pm 3 14±314 \pm 3 29±229 \pm 2
    OGBench scene (5 tasks) 5±15 \pm 1 28±128 \pm 1 41±341 \pm 3 46±346 \pm 3 20±120 \pm 1 40±740 \pm 7 30±330 \pm 3 45±545 \pm 5 30±330 \pm 3 56±256 \pm 2
    OGBench puzzle-3x3 (5 tasks) 2±02 \pm 0 9±19 \pm 1 21±121 \pm 1 10±210 \pm 2 18±118 \pm 1 19±019 \pm 0 6±26 \pm 2 14±414 \pm 4 19±119 \pm 1 30±130 \pm 1
    OGBench puzzle-4x4 (5 tasks) 0±00 \pm 0 7±17 \pm 1 14±114 \pm 1 29±329 \pm 3 10±310 \pm 3 15±315 \pm 3 1±01 \pm 0 13±113 \pm 1 25±525 \pm 5 17±217 \pm 2
    D4RL antmaze (6 tasks) 17 57 78 79 74 30±330 \pm 3 44±344 \pm 3 64±764 \pm 7 65±765 \pm 7 84±384 \pm 3
    D4RL adroit (12 tasks) 48 53 59 52±152 \pm 1 51±151 \pm 1 43±243 \pm 2 48±148 \pm 1 50±250 \pm 2 52±152 \pm 1 52±152 \pm 1
    Visual manipulation (5 tasks) - 42±442 \pm 4 60±260 \pm 2 - - - - 22±222 \pm 2 50±550 \pm 5 65±265 \pm 2

    Results are averaged across 8 seeds for state-based tasks and 4 seeds for pixel-based tasks. Scores reflect normalized success rates / returns. FQL outperforms its closest diffusion distillation baseline, Consistency-AC (CAC), and its closest Gaussian regularized actor-critic baseline, ReBRAC, especially in complex robotic manipulation tasks (e.g., cube-double, scene, puzzle) where policy multimodality is essential.

  5. Knowl 5 — Efficacy of Policy Extraction Mechanisms for Generative Offline RL

    empirical result

    When holding the underlying network architecture, base flow model, and codebase identical across flow-based offline RL methods, the choice of policy extraction scheme substantially impacts benchmark performance.

    Evaluating across the 50 state-based tasks in the OGBench benchmark suite yields the following average performance comparison:

    1. One-Step Guidance (FQL): Achieves an aggregate mean score of 44%44\%.
    2. Rejection Sampling (IFQL / Implicit Flow Q-Learning): Achieves an aggregate mean score of 30%30\%. It selects arg⁡max⁡a1,…,aN∼πβQ(s,ai)\arg\max_{a_1, \dots, a_N \sim \pi^\beta} Q(s, a_i) over NN sampled flow actions (N∈{32,64,128}N \in \{32, 64, 128\}), which is computationally expensive at inference and constrained by sample coverage.
    3. Reparameterization with BPTT (FBRAC / Flow Behavior-Regularized Actor-Critic): Achieves an aggregate mean score of 29%29\%. It directly propagates Q-gradients through the numerical ODE integration steps of the flow generator, which causes optimization instability.
    4. Advantage-Weighted Regression (FAWAC / Flow Advantage-Weighted Actor-Critic): Achieves an aggregate mean score of 16%16\%. It uses value-weighted regression max⁡θE[exp⁡(α(Q(s,a)−V(s)))LFlow(θ)]\max_\theta \mathbb{E}[\exp(\alpha(Q(s, a) - V(s))) \mathcal{L}_{\text{Flow}}(\theta)], suffering from effective sample drop and weak policy steering.

    One-step guidance achieves superior policy extraction by leveraging direct reparameterized policy gradients without the optimization pathologies of backpropagation through time.

  6. Knowl 6 — Offline-to-Online Reinforcement Learning Fine-Tuning Performance

    empirical result

    Flow Q-Learning (FQL) supports offline-to-online fine-tuning directly without algorithmic changes, balanced replay sampling, or additional exploration loss terms: transitions collected during online interactions are simply appended to the dataset D\mathcal{D}, and training continues using the exact same actor-critic and flow objectives.

    Across 15 fine-tuning benchmarks (5 representative OGBench tasks, 6 D4RL antmaze tasks, and 4 D4RL adroit tasks) fine-tuned from 1M to 2M steps:

    • On OGBench tasks, FQL improves significantly: antsoccer-arena from 28±828 \pm 8 to 86±586 \pm 5, cube-double from 40±1140 \pm 11 to 92±392 \pm 3, and scene from 82±1182 \pm 11 to 100±1100 \pm 1.
    • On D4RL antmaze tasks, FQL improves to near-perfect scores: antmaze-umaze (97→9997 \to 99), antmaze-umaze-diverse (79→10079 \to 100), antmaze-medium-play (77→9777 \to 97), antmaze-medium-diverse (55→9755 \to 97), antmaze-large-play (66→8466 \to 84), and antmaze-large-diverse (75→9475 \to 94).
    • On D4RL cloned adroit tasks, FQL increases normalized return from 53→14953 \to 149 on pen-cloned, 0→1020 \to 102 on door-cloned, 0→1270 \to 127 on hammer-cloned, and 0→620 \to 62 on relocate-cloned.

    FQL outperforms prior offline RL fine-tuning methods (IQL, ReBRAC, IFQL) as well as algorithms explicitly engineered for offline-to-online transfer (Cal-QL and RLPD).

  7. Knowl 7 — Computational and Inference Runtime Efficiency of FQL

    empirical result

    Training and inference runtimes measured on an NVIDIA RTX A5000 GPU demonstrate that Flow Q-Learning (FQL) is among the most computationally efficient generative policy methods:

    1. Training Speed: On state-based benchmarks (cube-double), FQL requires approximately 1.1 ms1.1\text{ ms} per gradient update step. This is substantially faster than Flow Behavior-Regularized Actor-Critic (FBRAC, ≈1.7 ms\approx 1.7\text{ ms}), which must backpropagate through 10 Euler ODE solver steps, and only slightly slower than Gaussian policy methods (ReBRAC at ≈0.8 ms\approx 0.8\text{ ms}, IQL at ≈0.6 ms\approx 0.6\text{ ms}).
    2. Inference Speed: In state-based environments, FQL generates actions in ≈0.05 ms\approx 0.05\text{ ms} via its one-step direct mapping. Rejection sampling flow policies (IFQL) require ≈1.2 ms\approx 1.2\text{ ms} due to sampling NN action candidates and querying the critic, and multi-step flow policies (FBRAC, FAWAC) require ≈0.8 ms\approx 0.8\text{ ms} for iterative Euler integration.
    3. Pixel Observation Setting: On visual-cube-double, FQL maintains comparable training step time (≈6.0 ms\approx 6.0\text{ ms}) and one-step inference time (≈0.8 ms\approx 0.8\text{ ms}), whereas IFQL requires ≈2.2 ms\approx 2.2\text{ ms} for test-time inference.
  8. Knowl 8 — Ablations on Hyperparameter and Architectural Choices in FQL

    empirical result

    Empirical ablations across navigation and manipulation environments reveal the sensitivity of FQL to its design parameters:

    1. Behavioral Regularization Weight α\alpha: α\alpha is the primary hyperparameter requiring per-environment tuning based on dataset suboptimality. Performance varies substantially across orders of magnitude; standard tasks perform optimally in α∈{3,10,30,100,300,1000}\alpha \in \{3, 10, 30, 100, 300, 1000\}, while high-reward-scale tasks (Adroit) require α∈{1000,3000,10000,30000}\alpha \in \{1000, 3000, 10000, 30000\}.
    2. Critic Target Aggregation: Utilizing the mean target 12(Qϕˉ1(s′,a′)+Qϕˉ2(s′,a′))\frac{1}{2}(Q_{\bar{\phi}_1}(s', a') + Q_{\bar{\phi}_2}(s', a')) outperforms clipped double Q-learning min⁡(Qϕˉ1,Qϕˉ2)\min(Q_{\bar{\phi}_1}, Q_{\bar{\phi}_2}) on the majority of tasks (e.g., cube-double, antmaze-large-diverse). However, the minimum operator is beneficial on specific environments such as adroit dexterous manipulation.
    3. Flow Integration Steps MM: Varying the number of Euler steps M∈{1,4,10,20}M \in \{1, 4, 10, 20\} during flow policy generation reveals that performance is robust for all M≥4M \ge 4, with M=10M=10 providing optimal stability across tasks.
    4. Flow Time Step Sampling Distribution: Comparing uniform time sampling t∼Unif([0,1])t \sim \text{Unif}([0, 1]), beta distribution t∼Beta(1,1.5)t \sim \text{Beta}(1, 1.5), and logit-normal distribution t=sigmoid(t~)t = \text{sigmoid}(\tilde{t}) with t~∼N(0,1)\tilde{t} \sim \mathcal{N}(0, 1) shows minimal variance in task success, demonstrating that the simplest uniform distribution is sufficient.
  9. Knowl 9 — Limitations of Flow Q-Learning

    limitation

    Flow Q-Learning (FQL) possesses three principal limitations:

    1. ODE Integration Overhead during Training: Training the one-step policy requires evaluating the full numerical ODE flow μθ(s,z)\mu_\theta(s, z) via the Euler method to compute the distillation target in Lπ(ω)\mathcal{L}_\pi(\omega), incurring computational overhead relative to purely analytical Gaussian actors.
    2. Absence of Dedicated Online Exploration: FQL lacks explicit optimism bonuses or exploration mechanisms during offline-to-online fine-tuning. In combinatorial search tasks requiring extensive exploration to escape sub-optimal offline behaviors (e.g., puzzle-4x4), FQL can converge to local optima.
    3. Lack of Real-Robot Hardware Validation: FQL has been evaluated exclusively across simulated benchmarks (OGBench and D4RL) and has not been validated on real-world robotic systems or integrated with pre-trained vision-language-action flow foundation models.

Coverage note — None was omitted. All principal contributions—including method formulation, Wasserstein theoretical connection, complete training algorithm, empirical offline RL benchmarks across 73 tasks, policy extraction ablation, offline-to-online fine-tuning, runtime analysis, hyperparameter ablations, and limitations—are fully represented.

References

  1. 1.Ada, S. E., Oztop, E., and Ugur, E. Diffusion policies for out-of-distribution generalization in offline reinforcement learning. IEEE Robotics and Automation Letters (RA-L), 9:3116–3123, 2024.
  2. 2.Ajay, A., Du, Y., Gupta, A., Tenenbaum, J., Jaakkola, T., and Agrawal, P. Is conditional generative modeling all you need for decision-making? In International Conference on Learning Representations (ICLR), 2023.
  3. 3.Albergo, M. S. and Vanden-Eijnden, E. Building normalizing flows with stochastic interpolants. In International Conference on Learning Representations (ICLR), 2023.
  4. 4.Alonso, E., Jelley, A., Micheli, V., Kanervisto, A., Storkey, A., Pearce, T., and Fleuret, F. Diffusion for world modeling: Visual details matter in atari. In Neural Information Processing Systems (NeurIPS), 2024.
  5. 5.An, G., Moon, S., Kim, J.-H., and Song, H. O. Uncertainty-based offline reinforcement learning with diversified q-ensemble. In Neural Information Processing Systems (NeurIPS), 2021.
  6. 6.Arjovsky, M., Chintala, S., and Bottou, L. Wasserstein generative adversarial networks. In International Conference on Machine Learning (ICML), 2017.
  7. 7.Ba, J., Kiros, J. R., and Hinton, G. E. Layer normalization. ArXiv, abs/1607.06450, 2016.
  8. 8.Ball, P. J., Smith, L., Kostrikov, I., and Levine, S. Efficient online reinforcement learning with offline data. In International Conference on Machine Learning (ICML), 2023.
  9. 9.Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al. π0\pi_0: A vision-language-action flow model for general robot control. ArXiv, abs/2410.24164, 2024.
  10. 10.Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q. JAX: composable transformations of Python+NumPy programs, 2018. URL http://github.com/jax-ml/jax.
  11. 11.Chen, C., Deng, F., Kawaguchi, K., Gulcehre, C., and Ahn, S. Simple hierarchical planning with diffusion. In International Conference on Learning Representations (ICLR), 2024a.
  12. 12.Chen, H., Lu, C., Ying, C., Su, H., and Zhu, J. Offline reinforcement learning via high-fidelity generative behavior modeling. In International Conference on Learning Representations (ICLR), 2023.
  13. 13.Chen, H., Lu, C., Wang, Z., Su, H., and Zhu, J. Score regularized policy optimization through diffusion behavior. In International Conference on Learning Representations (ICLR), 2024b.
  14. 14.Chen, H., Zheng, K., Su, H., and Zhu, J. Aligning diffusion behaviors with q-functions for efficient continuous control. In Neural Information Processing Systems (NeurIPS), 2024c.
  15. 15.Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. Decision transformer: Reinforcement learning via sequence modeling. In Neural Information Processing Systems (NeurIPS), 2021.
  16. 16.Chen, T., Wang, Z., and Zhou, M. Diffusion policies creating a trust region for offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2024d.
  17. 17.Collaboration, O. X.-E., O’Neill, A., Rehman, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., Jain, A., et al. Open x-embodiment: Robotic learning datasets and rt-x models. In IEEE International Conference on Robotics and Automation (ICRA), 2024.
  18. 18.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. In Neural Information Processing Systems (NeurIPS), 2021.
  19. 19.Ding, S., Hu, K., Zhang, Z., Ren, K., Zhang, W., Yu, J., Wang, J., and Shi, Y. Diffusion-based reinforcement learning via q-weighted variational policy optimization. In Neural Information Processing Systems (NeurIPS), 2024a.
  20. 20.Ding, Z. and Jin, C. Consistency models as a rich and efficient policy class for reinforcement learning. In International Conference on Learning Representations (ICLR), 2024.
  21. 21.Ding, Z., Jin, C., Liu, D., Zheng, H., Singh, K. K., Zhang, Q., Kang, Y., Lin, Z., and Liu, Y. Dollar: Few-step video generation via distillation and latent reward optimization. ArXiv, abs/2412.15689, 2024b.
  22. 22.Ding, Z., Zhang, A., Tian, Y., and Zheng, Q. Diffusion world model. ArXiv, abs/2402.03570, 2024c.
  23. 23.Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K. Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures. In International Conference on Machine Learning (ICML), 2018.
  24. 24.Esser, P., Kulal, S., Blattmann, A., Entezari, R., M"uller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al. Scaling rectified flow transformers for high-resolution image synthesis. In International Conference on Machine Learning (ICML), 2024.
  25. 25.Fang, L., Liu, R., Zhang, J., Wang, W., and Jing, B. Diffusion actor-critic: Formulating constrained policy iteration as diffusion noise regression for offline reinforcement learning. In International Conference on Learning Representations (ICLR), 2025.
  26. 26.Frans, K., Hafner, D., Levine, S., and Abbeel, P. One step diffusion via shortcut models. In International Conference on Learning Representations (ICLR), 2025.
  27. 27.Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. D4rl: Datasets for deep data-driven reinforcement learning. ArXiv, abs/2004.07219, 2020.
  28. 28.Fu, Y., Wu, D., and Boulet, B. A closer look at offline rl agents. In Neural Information Processing Systems (NeurIPS), 2022.
  29. 29.Fujimoto, S. and Gu, S. S. A minimalist approach to offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2021.
  30. 30.Fujimoto, S., van Hoof, H., and Meger, D. Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning (ICML), 2018.
  31. 31.Gao, R., Hoogeboom, E., Heek, J., Bortoli, V. D., Murphy, K. P., and Salimans, T. Diffusion meets flow matching: Two sides of the same coin, 2024. URL https://diffusionflow.github.io/.
  32. 32.Garg, D., Hejna, J., Geist, M., and Ermon, S. Extreme q-learning: Maxent rl without entropy. In International Conference on Learning Representations (ICLR), 2023.
  33. 33.Hansen-Estruch, P., Kostrikov, I., Janner, M., Kuba, J. G., and Levine, S. Idql: Implicit q-learning as an actor-critic method with diffusion policies. ArXiv, abs/2304.10573, 2023.
  34. 34.He, L., Shen, L., Zhang, L., Tan, J., and Wang, X. Diffcps: Diffusion model based constrained policy search for offline reinforcement learning. ArXiv, abs/2310.05333, 2023.
  35. 35.He, L., Shen, L., Tan, J., and Wang, X. Aligniql: Policy alignment in implicit q-learning through constrained optimization. ArXiv, abs/2405.18187, 2024.
  36. 36.Hendrycks, D. and Gimpel, K. Gaussian error linear units (gelus). ArXiv, abs/1606.08415, 2016.
  37. 37.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Neural Information Processing Systems (NeurIPS), 2020.
  38. 38.Jackson, M. T., Matthews, M. T., Lu, C., Ellis, B., Whiteson, S., and Foerster, J. Policy-guided diffusion. In Reinforcement Learning Conference (RLC), 2024.
  39. 39.Janner, M., Li, Q., and Levine, S. Reinforcement learning as one big sequence modeling problem. In Neural Information Processing Systems (NeurIPS), 2021.
  40. 40.Janner, M., Du, Y., Tenenbaum, J. B., and Levine, S. Planning with diffusion for flexible behavior synthesis. In International Conference on Machine Learning (ICML), 2022.
  41. 41.Kang, B., Ma, X., Du, C., Pang, T., and Yan, S. Efficient diffusion policies for offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2023.
  42. 42.Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. Morel : Model-based offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2020.
  43. 43.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015.
  44. 44.Kostrikov, I., Nair, A., and Levine, S. Offline reinforcement learning with implicit q-learning. In International Conference on Learning Representations (ICLR), 2022.
  45. 45.Kumar, A., Zhou, A., Tucker, G., and Levine, S. Conservative q-learning for offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2020.
  46. 46.Lange, S., Gabel, T., and Riedmiller, M. Batch reinforcement learning. In Reinforcement learning: State-of-the-art, pp. 45–73. Springer, 2012.
  47. 47.Lee, H., Hwang, D., Kim, D., Kim, H., Tai, J. J., Subramanian, K., Wurman, P. R., Choo, J., Stone, P., and Seno, T. Simba: Simplicity bias for scaling up parameters in deep reinforcement learning. In International Conference on Learning Representations (ICLR), 2025.
  48. 48.Lee, J., Jeon, W., Lee, B.-J., Pineau, J., and Kim, K.-E. Optidice: Offline policy optimization via stationary distribution correction estimation. In International Conference on Machine Learning (ICML), 2021a.
  49. 49.Lee, J. M. Introduction to Smooth Manifolds. Springer, 2012.
  50. 50.Lee, S., Seo, Y., Lee, K., Abbeel, P., and Shin, J. Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble. In Conference on Robot Learning (CoRL), 2021b.
  51. 51.Levine, S., Kumar, A., Tucker, G., and Fu, J. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. ArXiv, abs/2005.01643, 2020.
  52. 52.Li, J., Feng, W., Chen, W., and Wang, W. Y. Reward guided latent consistency distillation. Transactions on Machine Learning Research (TMLR), 2024a.
  53. 53.Li, W., Wang, X., Jin, B., and Zha, H. Hierarchical diffusion for offline decision making. In International Conference on Machine Learning (ICML), 2023.
  54. 54.Li, Z., Krohn, R., Chen, T., Ajay, A., Agrawal, P., and Chalvatzaki, G. Learning multimodal behaviors from scratch with diffusion policy gradient. In Neural Information Processing Systems (NeurIPS), 2024b.
  55. 55.Liang, Z., Mu, Y., Ding, M., Ni, F., Tomizuka, M., and Luo, P. Adaptdiffuser: Diffusion models as adaptive self-evolving planners. In International Conference on Machine Learning (ICML), 2023.
  56. 56.Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In International Conference on Learning Representations (ICLR), 2023.
  57. 57.Lipman, Y., Havasi, M., Holderrieth, P., Shaul, N., Le, M., Karrer, B., Chen, R. T. Q., Lopez-Paz, D., Ben-Hamu, H., and Gat, I. Flow matching guide and code. ArXiv, abs/2412.06264, 2024.
  58. 58.Liu, X., Gong, C., and Liu, Q. Flow straight and fast: Learning to generate and transfer data with rectified flow. In International Conference on Learning Representations (ICLR), 2023.
  59. 59.Liu, X., Zhang, X., Ma, J., Peng, J., et al. Instaflow: One step is enough for high-quality diffusion-based text-to-image generation. In International Conference on Learning Representations (ICLR), 2024.
  60. 60.Lu, C., Ball, P., Teh, Y. W., and Parker-Holder, J. Synthetic experience replay. In Neural Information Processing Systems (NeurIPS), 2023a.
  61. 61.Lu, C., Chen, H., Chen, J., Su, H., Li, C., and Zhu, J. Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning. In International Conference on Machine Learning (ICML), 2023b.
  62. 62.Mandlekar, A., Xu, D., Wong, J., Nasiriany, S., Wang, C., Kulkarni, R., Fei-Fei, L., Savarese, S., Zhu, Y., and Mart’in-Mart’in, R. What matters in learning from offline human demonstrations for robot manipulation. In Conference on Robot Learning (CoRL), 2021.
  63. 63.Mao, L., Xu, H., Zhan, X., Zhang, W., and Zhang, A. Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2024.
  64. 64.Mark, M. S., Gao, T., Sampaio, G. G., Srirama, M. K., Sharma, A., Finn, C., and Kumar, A. Policy agnostic rl: Offline rl and online rl fine-tuning of any class and backbone. ArXiv, abs/2412.06685, 2024.
  65. 65.Mazoure, B., Doan, T., Durand, A., Pineau, J., and Hjelm, R. D. Leveraging exploration in off-policy algorithms via normalizing flows. In Conference on Robot Learning (CoRL), 2019.
  66. 66.Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A. Playing atari with deep reinforcement learning. ArXiv, abs/1312.5602, 2013.
  67. 67.Nair, A., Dalal, M., Gupta, A., and Levine, S. Accelerating online reinforcement learning with offline datasets. ArXiv, abs/2006.09359, 2020.
  68. 68.Nakamoto, M., Zhai, Y., Singh, A., Mark, M. S., Ma, Y., Finn, C., Kumar, A., and Levine, S. Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning. In Neural Information Processing Systems (NeurIPS), 2023.
  69. 69.Nauman, M., Ostaszewski, M., Jankowski, K., Mi{\l}o's, P., and Cygan, M. Bigger, regularized, optimistic: scaling for compute and sample-efficient continuous control. In Neural Information Processing Systems (NeurIPS), 2024.
  70. 70.Nikulin, A., Kurenkov, V., Tarasov, D., and Kolesnikov, S. Anti-exploration by random network distillation. In International Conference on Machine Learning (ICML), 2023.
  71. 71.Park, S., Frans, K., Levine, S., and Kumar, A. Is value learning really the main bottleneck in offline rl? In Neural Information Processing Systems (NeurIPS), 2024a.
  72. 72.Park, S., Rybkin, O., and Levine, S. Metra: Scalable unsupervised rl with metric-aware abstraction. In International Conference on Learning Representations (ICLR), 2024b.
  73. 73.Park, S., Frans, K., Eysenbach, B., and Levine, S. Ogbench: Benchmarking offline goal-conditioned rl. In International Conference on Learning Representations (ICLR), 2025.
  74. 74.Peng, X. B., Kumar, A., Zhang, G., and Levine, S. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. ArXiv, abs/1910.00177, 2019.
  75. 75.Peters, J. and Schaal, S. Reinforcement learning by reward-weighted regression for operational space control. In International Conference on Machine Learning (ICML), 2007.
  76. 76.Psenka, M., Escontrela, A., Abbeel, P., and Ma, Y. Learning a diffusion model policy from rewards via q-score matching. In International Conference on Machine Learning (ICML), 2024.
  77. 77.Rafailov, R., Hatch, K. B., Singh, A., Kumar, A., Smith, L., Kostrikov, I., Hansen-Estruch, P., Kolev, V., Ball, P. J., Wu, J., et al. D5rl: Diverse datasets for data-driven deep reinforcement learning. In Reinforcement Learning Conference (RLC), 2024.
  78. 78.Ren, A. Z., Lidard, J., Ankile, L. L., Simeonov, A., Agrawal, P., Majumdar, A., Burchfiel, B., Dai, H., and Simchowitz, M. Diffusion policy policy optimization. In International Conference on Learning Representations (ICLR), 2025.
  79. 79.Sikchi, H. S., Zheng, Q., Zhang, A., and Niekum, S. Dual rl: Unification and new methods for reinforcement and imitation learning. In International Conference on Learning Representations (ICLR), 2024.
  80. 80.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning (ICML), 2015.
  81. 81.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations (ICLR), 2021.
  82. 82.Song, Y., Zhou, Y., Sekhari, A., Bagnell, J. A., Krishnamurthy, A., and Sun, W. Hybrid rl: Using both offline and online data can make rl efficient. In International Conference on Learning Representations (ICLR), 2023.
  83. 83.Suh, H. J. T., Chou, G., Dai, H., Yang, L., Gupta, A., and Tedrake, R. Fighting uncertainty with gradients: Offline reinforcement learning via diffusion score matching. In Conference on Robot Learning (CoRL), 2023.
  84. 84.Sutton, R. S. and Barto, A. G. Reinforcement learning: An introduction. IEEE Transactions on Neural Networks, 16:285–286, 2005.
  85. 85.Tarasov, D., Kurenkov, V., Nikulin, A., and Kolesnikov, S. Revisiting the minimalist approach to offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2023a.
  86. 86.Tarasov, D., Nikulin, A., Akimov, D., Kurenkov, V., and Kolesnikov, S. Corl: Research-oriented deep offline reinforcement learning library. In Neural Information Processing Systems (NeurIPS), 2023b.
  87. 87.Venkatraman, S., Khaitan, S., Akella, R. T., Dolan, J., Schneider, J., and Berseth, G. Reasoning with latent diffusion in offline reinforcement learning. In International Conference on Learning Representations (ICLR), 2024.
  88. 88.Wang, Z., Hunt, J. J., and Zhou, M. Diffusion policies as an expressive policy class for offline reinforcement learning. In International Conference on Learning Representations (ICLR), 2023.
  89. 89.Wu, Y., Tucker, G., and Nachum, O. Behavior regularized offline reinforcement learning. ArXiv, abs/1911.11361, 2019.
  90. 90.Xu, H., Jiang, L., Li, J., Yang, Z., Wang, Z., Chan, V., and Zhan, X. Offline rl with no ood actions: In-sample learning via implicit value regularization. In International Conference on Learning Representations (ICLR), 2023.
  91. 91.Yang, L., Huang, Z., Lei, F., Zhong, Y., Yang, Y., Fang, C., Wen, S., Zhou, B., and Lin, Z. Policy representation via diffusion probability model for reinforcement learning. ArXiv, abs/2305.13122, 2023.
  92. 92.Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T. Mopo: Model-based offline policy optimization. In Neural Information Processing Systems (NeurIPS), 2020.
  93. 93.Yu, Z. and Zhang, X. Actor-critic alignment for offline-to-online reinforcement learning. In International Conference on Machine Learning (ICML), 2023.
  94. 94.Zhang, R., Luo, Z., Sj"olund, J., Sch"on, T. B., and Mattsson, P. Entropy-regularized diffusion policy with q-ensembles for offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2024.
  95. 95.Zhang, S., Zhang, W., and Gu, Q. Energy-weighted flow matching for offline reinforcement learning. In International Conference on Learning Representations (ICLR), 2025.
  96. 96.Zheng, Q., Le, M., Shaul, N., Lipman, Y., Grover, A., and Chen, R. T. Guided flows for generative modeling and decision making. ArXiv, abs/2311.13443, 2023.

Citation

MLA
Park, S., et al. “Flow Q-Learning”. arXiv, 2025, http://arxiv.org/abs/2502.02538v2.
APA
Park, S., Li, Q., & Levine, S. (2025). Flow Q-Learning. arXiv. http://arxiv.org/abs/2502.02538v2
Chicago
Park, S., Q. Li, and S. Levine. 2025. “Flow Q-Learning”. arXiv. http://arxiv.org/abs/2502.02538v2.
Harvard
Park, S., Li, Q. and Levine, S. (2025) “Flow Q-Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.02538v2.
Vancouver
1. Park S, Li Q, Levine S (2025) Flow Q-Learning. arXiv

BibTeX

@article{park2025flow,
  title = {Flow Q-Learning},
  author = {Park, Seohong and Li, Qiyang and Levine, Sergey},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.02538v2},
  eprint = {2502.02538}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/