WarmPrior: Straightening Flow-Matching Policies with Temporal Priors

Sinjae KangChanyoung KimKaixin WangLi ZhaoKimin Lee

article2026arXiv3 citations

Introduces WarmPrior, a temporal source distribution derived from recent action history that straightens probability paths in flow-matching policies to boost robotic manipulation performance across imitation and reinforcement learning.

Listen

Modern robot control increasingly relies on generative models, such as diffusion and flow-matching policies, to translate camera images and sensor data into precise physical movements. Conventionally, these systems generate actions by transforming a completely random, standard normal noise distribution into a sequence of robot actions. However, treating the starting point as pure, stateless noise ignores the continuous and predictable flow of physical motion. This design flaw forces the robot's control policy to reconstruct every motion trajectory from scratch at each step, resulting in curved generation paths, slower computation, and erratic behavioral switching between different task strategies.

The main objective of the article is to demonstrate that replacing standard random noise with a temporally informed starting distribution, termed WarmPrior, significantly enhances the execution performance, computational speed, and learning efficiency of generative robot policies without modifying the underlying neural networks or training losses.

The approach introduces two simple mechanisms to anchor the starting noise distribution to the robot's own recent movement history, combined with a controlled amount of residual noise. The first variant, WarmPrior-Past, anchors the initial noise to the robot's previously executed action. The second, WarmPrior-Preview, trains the model to forecast twice the required action steps and uses its previous future forecast as the starting point for the current step. The authors evaluated this framework across rigorous simulation suites (Robomimic and MimicGen) and physical robot hardware using a Franka robotic arm across diverse manipulation tasks, testing across varying computational budgets and combining the method with reinforcement learning.

The key findings show consistent, broad-based improvements across all tested domains. First, WarmPrior substantially increases robotic task success rates compared to standard noise baselines, with the most pronounced gains occurring on complex, multi-human demonstration tasks and when compute budgets are constrained to single-step inference (for example, improving success rates from 65.9% to 77.8% on a challenging peg-insertion task). Second, WarmPrior-Preview consistently outperforms WarmPrior-Past because future forecasts provide a closer approximation to the target action than simple past-action persistence. Third, the method straightens the learned mathematical transport trajectories by up to 44%, effectively removing the directional ambiguity that normally curves generative flows. Fourth, in reinforcement learning fine-tuning, WarmPrior restricts the agent's exploration space to a 3-times tighter, meaningful region, enabling it to surpass a 95% success rate on the difficult multi-robot transport task for the first time.

These results demonstrate that the starting noise distribution is a highly practical lever for improving robotic systems. Because WarmPrior straightens generation trajectories, it allows robots to operate accurately at minimal inference steps, directly reducing computational latency, hardware costs, and power requirements. It also resolves dangerous oscillations between conflicting strategies during execution by providing smooth temporal continuity. Furthermore, WarmPrior complements existing inference-time smoothing techniques like Real-Time Chunking, as stacking both approaches yields higher success rates than either method alone.

Based on these findings, engineering and product teams building generative robot policies should replace default random Gaussian starting distributions with temporally grounded priors. Teams should prioritize the preview-based variant when planning horizons allow, or use the past-action variant when simplicity is paramount, maintaining moderate noise levels to avoid collapsing into brittle deterministic behaviors. In addition, organizations fine-tuning diffusion-based policies via reinforcement learning should adopt bounded, anchor-centered action spaces to dramatically reduce sample training time.

A key boundary condition of the study is that the noise variance must remain balanced; reducing noise completely to zero degrades the model into a standard regression approach that fails entirely on multimodal tasks. While the findings provide high confidence across multiple benchmark architectures and tabletop manipulation settings, further validation is recommended on highly agile, long-horizon mobile manipulation platforms in unstructured real-world environments before enterprise-wide deployment.

arXiv: 2605.13959sinnnj/warmprior-bc
Cover for WarmPrior: Straightening Flow-Matching Policies with Temporal Priors

Abstract

Generative policies based on diffusion and flow matching have become a dominant paradigm for visuomotor robotic control. We show that replacing the standard Gaussian source distribution with WarmPrior, a simple temporally grounded prior constructed from readily available recent action history, consistently improves success rates on robotic manipulation tasks. We trace this gain to markedly straighter probability paths, echoing the effect of optimal-transport couplings in Rectified Flow. Beyond standard behavior cloning, WarmPrior also reshapes the exploration distribution in prior-space reinforcement learning, improving both sample efficiency and final performance. Collectively, these results identify the source distribution as an important and underexplored design axis in generative robot control.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 3 WarmPrior
  • 4 Main Results
  • 4.1 Simulation
  • 4.2 Real-Robot Experiments
  • 5 Understanding and Extending WarmPrior
  • 5.1 WarmPrior Improves SR by Straightening Flow Trajectories
  • 5.2 WarmPrior as a Tunable Source of Temporal Consistency
  • 5.3 WarmPrior Improves Prior-Space RL Efficiency
  • 6 Conclusion
  • References
  • A Visualizing Success-Rate Uncertainty with Beta Posteriors
  • B Why WarmPrior Straightens Flows: A Branching-Cost Analysis
  • B.1 The branching cost as endpoint ambiguity
  • B.2 A detailed derivation of the WarmPrior bound
  • C Training Details
  • D σ\sigma Ablation
  • E Comparing WarmPrior with Real-Time Chunking

Knowls

  1. Knowl 1 — WarmPrior Source Distribution Formulation

    model/method

    In flow-matching policies for robot manipulation, action trajectories are generated by integrating a velocity field that transports samples from a prior distribution p0p_0 to the data distribution pdata(⋅∣o)p_{\text{data}}(\cdot \mid o) given an observation oo. Rather than using an isotropic Gaussian distribution N(0,I)\mathcal{N}(0, I), WarmPrior constructs a temporally grounded source distribution centered on recent action history.

    Let a0a_0 denote the prior sample of dimension L×daL \times d_a, where LL is the action prediction horizon and dad_a is the per-step action dimension. Given an index set of warm coordinates W⊆{0,…,L−1}\mathcal{W} \subseteq \{0, \dots, L-1\}, a cold complement C={0,…,L−1}∖W\mathcal{C} = \{0, \dots, L-1\} \setminus \mathcal{W}, a prior mean vector μ\mu defined on W\mathcal{W}, and a residual standard deviation scalar σ>0\sigma > 0, the WarmPrior sample is defined coordinate-wise across timestep indices τ∈{0,…,L−1}\tau \in \{0, \dots, L-1\} as:

    a0[τ]={μτ+σετ,τ∈W,ετ,τ∈C,ε∼N(0,I).a_0[\tau] = \begin{cases} \mu_\tau + \sigma \varepsilon_\tau, & \tau \in \mathcal{W}, \\ \varepsilon_\tau, & \tau \in \mathcal{C}, \end{cases} \quad \varepsilon \sim \mathcal{N}(0, I).

    WarmPrior is instantiated in two primary variants:

    1. WarmPrior-Past (WP-Past): Predicts a single chunk of HH actions (L=HL = H) and sets W={0,…,H−1}\mathcal{W} = \{0, \dots, H-1\}. During training, for a sample at index ii, μτPast=adata[i−H+τ]\mu_\tau^{\text{Past}} = a^{\text{data}}[i - H + \tau] using the preceding HH actions from the demonstration buffer (falling back to W=∅\mathcal{W} = \emptyset at episode boundaries). At inference, μτPast=a^prev[τ]\mu_\tau^{\text{Past}} = \hat{a}^{\text{prev}}[\tau], the previously executed action chunk (falling back to W=∅\mathcal{W} = \emptyset at the initial step). The default noise scale is σ=0.5\sigma = 0.5.

    2. WarmPrior-Preview (WP-Preview): Predicts a horizon of 2H2H actions (L=2HL = 2H) but executes only the first HH. The warm set is the first chunk W={0,…,H−1}\mathcal{W} = \{0, \dots, H-1\}, while the preview horizon C={H,…,2H−1}\mathcal{C} = \{H, \dots, 2H-1\} remains cold noise. During training, the ground-truth target acts as the ideal preview proxy: μτPreview=a1[τ]\mu_\tau^{\text{Preview}} = a_1[\tau] for τ∈{0,…,H−1}\tau \in \{0, \dots, H-1\}, where a1∈R2H×daa_1 \in \mathbb{R}^{2H \times d_a}. At inference, the prior mean is the second half of the previous prediction: μτPreview=a^prev[H+τ]\mu_\tau^{\text{Preview}} = \hat{a}^{\text{prev}}[H + \tau] for τ∈{0,…,H−1}\tau \in \{0, \dots, H-1\} (falling back to W=∅\mathcal{W} = \emptyset at the first step). The default noise scale is σ=1.0\sigma = 1.0.

  2. Knowl 2 — Training and Inference Algorithm for Flow-Matching Policies with WarmPrior

    algorithm

    The algorithm specifies training and inference routines for flow-matching robot policies under WarmPrior. The velocity network vθ(t,at,o)v_\theta(t, a_t, o) is trained with standard linear flow matching while only altering the initial sample a0a_0.

    Input: Dataset DD, linear interpolant coefficients α(t)=1−t\alpha(t) = 1-t and β(t)=t\beta(t) = t, noise scale σ\sigma, chunk length HH, prediction horizon LL (L=HL=H for Past, L=2HL=2H for Preview)
    Parameters: Learnable velocity network vθv_\theta
    TRAINING:
    for each training iteration do
        Sample observation, target action, and buffer index (o,a1,i)∼D(o, a_1, i) \sim D, with a1∈RL×daa_1 \in \mathbb{R}^{L \times d_a}
        Draw standard Gaussian noise ε∼N(0,I)\varepsilon \sim \mathcal{N}(0, I) of shape L×daL \times d_a
        Set a0←εa_0 \leftarrow \varepsilon
        if PAST and i≥Hi \ge H (within same episode) then
            a0←adata[i−H:i]+σεa_0 \leftarrow a^{\text{data}}[i-H:i] + \sigma \varepsilon
        else if PREVIEW then
            a0[0:H]←a1[0:H]+σε[0:H]a_0[0:H] \leftarrow a_1[0:H] + \sigma \varepsilon[0:H]
            a0[H:2H]←ε[H:2H]a_0[H:2H] \leftarrow \varepsilon[H:2H]
        end if
        Sample time t∼U(0,1)t \sim \mathcal{U}(0, 1)
        Compute interpolated point at←α(t)a0+β(t)a1a_t \leftarrow \alpha(t) a_0 + \beta(t) a_1
        Compute target velocity a˙t←α˙(t)a0+β˙(t)a1=a1−a0\dot{a}_t \leftarrow \dot{\alpha}(t) a_0 + \dot{\beta}(t) a_1 = a_1 - a_0
        Compute loss L(θ)←∥vθ(t,at,o)−a˙t∥22\mathcal{L}(\theta) \leftarrow \|v_\theta(t, a_t, o) - \dot{a}_t\|_2^2
        Update θ\theta via gradient step on L(θ)\mathcal{L}(\theta)
    end for
    INFERENCE:
    Initialize previous prediction record a^prev←∅\hat{a}^{\text{prev}} \leftarrow \emptyset
    Reset environment and observe initial observation oo
    while episode not done do
        Draw noise ε∼N(0,I)\varepsilon \sim \mathcal{N}(0, I) of shape L×daL \times d_a
        Set a0←εa_0 \leftarrow \varepsilon
        if PAST and a^prev≠∅\hat{a}^{\text{prev}} \ne \emptyset then
            a0←a^prev+σεa_0 \leftarrow \hat{a}^{\text{prev}} + \sigma \varepsilon
        else if PREVIEW and a^prev≠∅\hat{a}^{\text{prev}} \ne \emptyset then
            a0[0:H]←a^prev[H:2H]+σε[0:H]a_0[0:H] \leftarrow \hat{a}^{\text{prev}}[H:2H] + \sigma \varepsilon[0:H]
            a0[H:2H]←ε[H:2H]a_0[H:2H] \leftarrow \varepsilon[H:2H]
        end if
        Generate action chunk a^←FMSAMPLE(vθ,a0,o)\hat{a} \leftarrow \text{FMSAMPLE}(v_\theta, a_0, o) using an ODE integrator
        Execute action chunk a^[0:H]\hat{a}[0:H] on the robot and observe new observation oo
        Update previous prediction a^prev←a^\hat{a}^{\text{prev}} \leftarrow \hat{a}
    end while

    In both variants, when no preceding history is available at episode onset or across trajectory boundaries, the method defaults to standard Gaussian initialization (a0=εa_0 = \varepsilon).

  3. Knowl 3 — Exact Characterization of Flow-Matching Branching Cost

    theoretical result

    Let (A0,A1)∼Πo(A_0, A_1) \sim \Pi_o denote the conditional joint distribution over prior samples A0∈RdA_0 \in \mathbb{R}^d and target actions A1∈RdA_1 \in \mathbb{R}^d given policy observation oo, where d=Hdad = H d_a. Under the linear interpolant At=(1−t)A0+tA1A_t = (1-t)A_0 + tA_1 for t∈[0,1]t \in [0, 1], the population flow-matching loss for a velocity field vt(⋅,o)v_t(\cdot, o) is:

    Lo(v)=∫01EΠo[∥vt(At,o)−(A1−A0)∥22  |  o]dt.\mathcal{L}_o(v) = \int_0^1 \mathbb{E}_{\Pi_o}\left[ \|v_t(A_t, o) - (A_1 - A_0)\|^2_2 \;\middle|\; o \right] dt.

    For each t∈[0,1)t \in [0, 1), the unique L2L^2-minimizer of Lo(v)\mathcal{L}_o(v) is given by the conditional expectation:

    vt⋆(x,o)=E[A1−A0∣At=x,o]=E[A1∣At=x,o]−x1−t.v_t^\star(x, o) = \mathbb{E}[A_1 - A_0 \mid A_t = x, o] = \frac{\mathbb{E}[A_1 \mid A_t = x, o] - x}{1 - t}.

    The irreducible residual flow-matching error, defined as the branching cost B(o):=Lo(v⋆)\mathcal{B}(o) := \mathcal{L}_o(v^\star), is exactly the time-integrated conditional variance of the endpoint A1A_1 given the intermediate state AtA_t:

    B(o)=∫011(1−t)2E[∥A1−E[A1∣At,o]∥22  |  o]dt.\mathcal{B}(o) = \int_0^1 \frac{1}{(1-t)^2} \mathbb{E}\left[ \|A_1 - \mathbb{E}[A_1 \mid A_t, o]\|^2_2 \;\middle|\; o \right] dt.

    This quantity measures the directional ambiguity remaining in A1A_1 after observing AtA_t. If AtA_t uniquely determines A1A_1 almost surely, B(o)=0\mathcal{B}(o) = 0, which corresponds to non-branching optimal transport paths.

  4. Knowl 4 — Upper Bound on Warm-Coordinate Branching Cost

    theoretical result

    Let PWP_{\mathcal{W}} denote the orthogonal projection matrix onto the subspace of warm coordinates W\mathcal{W} of dimension dW=tr⁡(PW)d_{\mathcal{W}} = \operatorname{tr}(P_{\mathcal{W}}), and let PC=I−PWP_{\mathcal{C}} = I - P_{\mathcal{W}} be the projection onto the cold coordinates. Suppose the source distribution is structured as:

    A0=PW(μ+σΞ)+PCΞ,Ξ∼N(0,Id),A_0 = P_{\mathcal{W}}(\mu + \sigma \Xi) + P_{\mathcal{C}} \Xi, \quad \Xi \sim \mathcal{N}(0, I_d),

    where Ξ\Xi is conditionally independent of (A1,μ)(A_1, \mu) given observation oo. The warm-coordinate branching cost is defined as:

    BW(o):=∫011(1−t)2E[∥PWA1−E[PWA1∣At,o]∥22  |  o]dt.\mathcal{B}_{\mathcal{W}}(o) := \int_0^1 \frac{1}{(1-t)^2} \mathbb{E}\left[ \|P_{\mathcal{W}} A_1 - \mathbb{E}[P_{\mathcal{W}} A_1 \mid A_t, o]\|^2_2 \;\middle|\; o \right] dt.

    Under this source structure and linear interpolation At=(1−t)A0+tA1A_t = (1-t)A_0 + tA_1, the warm-coordinate branching cost satisfies the upper bound:

    BW(o)≤E[∥PW(A1−μ)∥22  |  o]+σ2dW.\mathcal{B}_{\mathcal{W}}(o) \le \mathbb{E}\left[ \|P_{\mathcal{W}}(A_1 - \mu)\|^2_2 \;\middle|\; o \right] + \sigma^2 d_{\mathcal{W}}.

    The bound decomposes the irreducible curvature into two terms:

    1. Mean Mismatch: E[∥PW(A1−μ)∥22∣o]\mathbb{E}[\|P_{\mathcal{W}}(A_1 - \mu)\|^2_2 \mid o], which measures how well the prior mean μ\mu predicts the target action chunk A1A_1. For WarmPrior-Preview, this corresponds to forecast error E[∥E∥2∣o]\mathbb{E}[\|E\|^2 \mid o]; for WarmPrior-Past, it corresponds to persistence residual E[∥R∥2∣o]\mathbb{E}[\|R\|^2 \mid o] between consecutive chunks.
    2. Gaussian Dispersion: σ2dW\sigma^2 d_{\mathcal{W}}, which scales with the injected residual variance.
  5. Knowl 5 — Simulation Benchmark Results on Robomimic and MimicGen

    data/table

    Simulation success rates (%) evaluating the standard Gaussian prior N(0,I)\mathcal{N}(0, I) baseline against WarmPrior-Past (WP-Past) and WarmPrior-Preview (WP-Preview) using the Diffusion Policy (ChiTransformer) backbone with action-chunk length H=8H = 8 across numbers of function evaluations NFE∈{9,3,1}\text{NFE} \in \{9, 3, 1\}. Results are averaged over 3 seeds and 200 episodes per seed (top-3 checkpoint average).

    NFE = 9 NFE = 3 NFE = 1
    Task Base WP-Past WP-Preview Base WP-Past WP-Preview Base WP-Past WP-Preview
    Robomimic — state observation
    Square-PH 86.7 88.1 88.1 86.2 88.0 87.9 83.6 86.6 87.3
    Square-MH 65.9 69.2 72.7 65.4 73.2 72.9 65.9 70.1 77.8
    Transport-PH 34.1 36.2 43.3 39.0 44.0 49.1 36.8 39.8 47.6
    Transport-MH 16.3 20.7 24.3 21.3 30.7 30.4 23.3 30.2 34.5
    Tool-Hang-PH 79.4 80.6 82.8 72.3 75.1 75.8 77.7 78.2 81.9
    Robomimic — image observation
    Square-PH 86.9 88.2 88.7 87.7 89.2 89.6 88.7 89.3 89.1
    Square-MH 76.1 78.0 77.8 73.8 77.9 77.1 72.4 77.6 75.1
    Transport-PH 92.8 94.5 94.3 92.1 93.9 94.9 91.3 93.4 93.7
    Transport-MH 74.8 79.7 79.8 73.8 80.0 80.7 74.3 78.6 79.7
    Tool-Hang-PH 43.7 45.8 56.3 36.9 38.4 50.7 41.3 38.9 54.0
    MimicGen — image observation
    Stack 21.4 22.8 31.6 21.3 23.7 30.7 21.3 22.4 28.7
    Coffee 26.8 29.6 34.7 23.3 24.1 33.4 16.2 20.4 29.4
    Threading 13.8 15.5 20.9 16.3 16.6 22.0 12.5 15.6 18.0

    WP-Preview delivers the highest success rates across most domains, with especially pronounced absolute improvements on challenging tasks (e.g., +12.7% on Tool-Hang-PH image at NFE=1\text{NFE}=1, +11.2% on Transport-MH state at NFE=1\text{NFE}=1, and +13.2% on MimicGen Coffee at NFE=1\text{NFE}=1). Improvements over the baseline are largest under the one-step generation regime (NFE=1\text{NFE}=1).

  6. Knowl 6 — Pathwise Curvature Reduction and Flow Straightening

    empirical result

    The pathwise curvature of a learned flow policy path a:[0,1]→RH×daa: [0, 1] \to \mathbb{R}^{H \times d_a} parameterized by velocity field a˙t=vθ(t,at,o)\dot{a}_t = v_\theta(t, a_t, o) is measured via the velocity-variance surrogate:

    κ(o)=∫01∥a˙t−vˉ∥22 dt,where vˉ=∫01a˙t dt.\kappa(o) = \int_0^1 \|\dot{a}_t - \bar{v}\|^2_2 \, dt, \quad \text{where } \bar{v} = \int_0^1 \dot{a}_t \, dt.

    Evaluating κ(o)\kappa(o) with an Euler sampler (N=100N = 100 steps) over 2,000 validation observations on Robomimic state-observation tasks yields the following normalized curvature values (where the standard Gaussian prior N(0,I)\mathcal{N}(0, I) baseline is set to 1.000):

    • Square-PH: Baseline = 1.000, WP-Past = 0.823, WP-Preview = 0.803
    • Square-MH: Baseline = 1.000, WP-Past = 0.705, WP-Preview = 0.559
    • Transport-PH: Baseline = 1.000, WP-Past = 0.720, WP-Preview = 0.692
    • Transport-MH: Baseline = 1.000, WP-Past = 0.695, WP-Preview = 0.637
    • Tool-Hang-PH: Baseline = 1.000, WP-Past = 0.806, WP-Preview = 0.807

    WarmPrior substantially reduces the pathwise curvature across all tasks (up to a 44.1% reduction for WP-Preview on Square-MH). The relative magnitude of curvature reduction closely mirrors the task success rate improvements, supporting the premise that straightening probability trajectories improves generative policy performance.

  7. Knowl 7 — Conditioned-Residual Prior-Space Reinforcement Learning

    model/method

    In prior-space reinforcement learning (such as Diffusion Steering RL / DSRL), an RL policy generates the initial prior sample a0a_0 for a frozen, pretrained generative policy instead of directly predicting action chunk a1a_1. Applying WarmPrior modifies this setup into a conditioned-residual search problem.

    Rather than exploring the entire uninformative prior space RH×da\mathbb{R}^{H \times d_a} with action bound δ=1.5\delta = 1.5 centered at the origin, the RL action space is restricted to a tight residual Δ\Delta around the WarmPrior mean μ\mu:

    a0=μ+Δ,Δ=πRL(o~)∈[−δ,δ]H×da,a_0 = \mu + \Delta, \quad \Delta = \pi_{\text{RL}}(\tilde{o}) \in [-\delta, \delta]^{H \times d_a},

    where the maximum residual magnitude is set to δ=0.5\delta = 0.5 (a 3×3\times reduction in action range), and the observation is augmented to o~=[o,μ]\tilde{o} = [o, \mu] to preserve the Markov property.

    When fine-tuning DSRL-NA on Robomimic tasks using a frozen flow-matching backbone pretrained for 3,000 epochs:

    • On Square, WP-Past and WP-Preview converge faster and exceed a 0.99 success rate, whereas vanilla DSRL-SAC and DSRL-NA baseline plateaus near 0.90.
    • On Transport (the most challenging task), WP-Past and WP-Preview achieve ∼0.97\sim 0.97 success rate, whereas baseline DSRL-SAC and DSRL-NA fail to exceed 0.90.
  8. Knowl 8 — Temporal Consistency and Single-Step Horizon Performance

    empirical result

    Generative behavior cloning policies trained with context-free Gaussian priors N(0,I)\mathcal{N}(0, I) suffer from mode switching in multimodal environments: independent sampling at consecutive inferences causes the robot to oscillate erratically between divergent demonstrated trajectories at decision boundaries. WarmPrior resolves mode switching by implicitly biasing each generation toward the local mode basin established by the previous action.

    When action chunking is completely removed by setting the chunk length to H=1H = 1 (running fresh policy inference at every single environment timestep with NFE=1\text{NFE} = 1):

    • On Transport-MH (State): The standard N(0,I)\mathcal{N}(0, I) baseline collapses from 23.3% (H=8H=8) down to 1.3%, whereas WarmPrior retains high performance (WP-Past reaches ≈28%\approx 28\%, WP-Preview reaches ≈33%\approx 33\%).
    • On MimicGen Coffee (Image): Baseline drops to 16.2%, while WarmPrior recovers performance with gains up to +14.8%+14.8\%.

    These results demonstrate that WarmPrior supplies an implicit temporal consistency mechanism that functions even in the absence of explicit action chunking (H=1H = 1), allowing per-step re-planning without mode oscillation.

  9. Knowl 9 — Ablation of Residual Noise Scale and Failure of Deterministic Limits

    empirical result

    Sweeping the prior standard deviation σ∈{1.5,1.0,0.5,0.3,0.1,0.05,0}\sigma \in \{1.5, 1.0, 0.5, 0.3, 0.1, 0.05, 0\} on Robomimic Square-MH (H=8,NFE=1H = 8, \text{NFE} = 1) reveals a distinct non-monotonic trade-off between trajectory straightness and prior expressiveness:

    • Optimal Values: WP-Preview peaks at σ=1.0\sigma = 1.0 (success rate 0.778), maintaining performance within 0.06 of its optimum across σ∈[0.3,1.5]\sigma \in [0.3, 1.5]. WP-Past peaks at a tighter scale σ=0.5\sigma = 0.5 (success rate 0.701), staying within 0.08 of its peak across σ∈[0.3,1.0]\sigma \in [0.3, 1.0].
    • Deterministic Collapse (σ=0\sigma = 0): In the limit σ→0\sigma \to 0, the source becomes a deterministic point mass a0=μa_0 = \mu, reducing the generative flow model to a deterministic regression mapping μ↦a1\mu \mapsto a_1. In this limit, success rates sharply collapse to ≈0.31\approx 0.31 for WP-Preview and ≈0.05\approx 0.05 for WP-Past.

    Maintaining σ>0\sigma > 0 provides the residual capacity necessary to accommodate prior prediction errors while preserving multimodal generative modeling.

  10. Knowl 10 — Real-Robot Evaluation and Synergy with Real-Time Chunking

    empirical result

    WarmPrior was evaluated on physical Franka Research 3 robot setups across two model families:

    1. GR00T N1.5 VLA Backbone (NFE=4\text{NFE} = 4, 30 teleoperated demonstrations per task) across four manipulation tasks (mean success rate over 3 seeds, 50 trials per seed):

      • Food Waste Disposal: Baseline N(0,I)=0.76\mathcal{N}(0, I) = 0.76, WP-Past = 0.90, WP-Preview = 0.87
      • Cup Stacking: Baseline = 0.33, WP-Past = 0.36, WP-Preview = 0.45
      • Block Stacking: Baseline = 0.43, WP-Past = 0.60, WP-Preview = 0.67
      • Cable Insertion: Baseline = 0.15, WP-Past = 0.36, WP-Preview = 0.41
    2. Orthogonality to Real-Time Chunking (RTC) on π0.5\pi_{0.5} Backbone: On dynamic whole-arm manipulation tasks (Block Throwing and Towel Folding), WarmPrior (training-time prior shaping) was compared against and stacked with RTC (inference-time prefix clamping):

      • Block Throwing: Baseline = 0.32, RTC = 0.57, WarmPrior = 0.48, RTC + WarmPrior = 0.62
      • Towel Folding: Baseline = 0.50, RTC = 0.67, WarmPrior = 0.68, RTC + WarmPrior = 0.82

    The combined configuration outperforms both standalone methods, demonstrating that straightening velocity fields at training time and enforcing continuity via prefix clamping at inference time operate through complementary mechanisms.

Coverage note — None was omitted; all key theoretical derivations, algorithmic variants, empirical simulation/real-robot results, ablations, and reinforcement learning extensions from the paper are covered.

References

  1. 1.Michael S. Albergo, Nicholas M. Boffi, and Eric Vanden-Eijnden. Stochastic interpolants: A unifying framework for flows and diffusions. Journal of Machine Learning Research, 26(209):1–80, 2025.
  2. 2.Jean-David Benamou and Yann Brenier. A computational fluid mechanics solution to the Monge–Kantorovich mass transfer problem. Numerische Mathematik, 84(3):375–393, 2000.
  3. 3.Johan Bjorck, Valts Blukis, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Xiaowei Jiang, Jan Kautz, Kaushil Kundalia, Zhiqi Li, Kevin Lin, Zongyu Lin, Loic Magne, Yunze Man, Ajay Mandlekar, Avnish Narayan, Soroush Nasiriany, Scott Reed, You Liang Tan, Guanzhi Wang, Jing Wang, Qi Wang, Shihao Wang, Jiannan Xiang, Yuqi Xie, Yinzhen Xu, Seonghyeon Ye, Zhiding Yu, Yizhou Zhao, Zhe Zhang, Ruijie Zheng, and Yuke Zhu. GR00T N1.5: An open foundation model for generalist humanoid robots. NVIDIA Isaac GR00T technical report, 2025a.
  4. 4.Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. GR00T N1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734, 2025b.
  5. 5.Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky. π0\pi_0: A vision-language-action flow model for general robot control. In Robotics: Science and Systems, 2025a.
  6. 6.Kevin Black, Manuel Y. Galliker, and Sergey Levine. Real-time execution of action chunking flow policies. arXiv preprint arXiv:2506.07339, 2025b.
  7. 7.Max Braun, Noémie Jaquier, Leonel Rozo, and Tamim Asfour. Riemannian flow matching policy for robot motion learning. In IEEE/RSJ International Conference on Intelligent Robots and Systems, 2024.
  8. 8.Guo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang, Yicheng Liu, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Tuomas Rintamaki, Tyler Poon, Max Ehrlich, Tong Lu, Limin Wang, Bryan Catanzaro, Jan Kautz, Andrew Tao, Zhiding Yu, and Guilin Liu. Eagle 2.5: Boosting long-context post-training for frontier vision-language models. arXiv preprint arXiv:2504.15271, 2025.
  9. 9.Kaiqi Chen, Eugene Lim, Kelvin Lin, Yiyang Chen, and Harold Soh. Don’t start from scratch: Behavioral refinement via interpolant-based policy diffusion. arXiv preprint arXiv:2402.16075, 2024.
  10. 10.Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. In Robotics: Science and Systems, 2023.
  11. 11.Eugenio Chisari, Nick Heppert, Max Argus, Tim Welschehold, Thomas Brox, and Abhinav Valada. Learning robotic manipulation policies from point clouds with conditional flow matching. In Conference on Robot Learning, 2024.
  12. 12.Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning, 2018.
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  14. 14.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, 2020.
  15. 15.Xixi Hu, Bo Liu, Xingchao Liu, and Qiang Liu. AdaFlow: Imitation learning with variance-adaptive flow-based policies. In Advances in Neural Information Processing Systems, 2024.
  16. 16.Michael Janner, Yilun Du, Joshua B. Tenenbaum, and Sergey Levine. Planning with diffusion for flexible behavior synthesis. In International Conference on Machine Learning, 2022.
  17. 17.Jindou Jia, Gen Li, Xiangyu Chen, Tuo An, Yuxuan Hu, Jingliang Li, Xinying Guo, and Jianfei Yang. Action-to-action flow matching. arXiv preprint arXiv:2602.07322, 2026.
  18. 18.Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, et al. DROID: A large-scale in-the-wild robot manipulation dataset. In Robotics: Science and Systems, 2024.
  19. 19.Myungkyu Koo, Daewon Choi, Taeyoung Kim, Kyungmin Lee, Changyeon Kim, Younggyo Seo, and Jinwoo Shin. HAMLET: Switch your vision-language-action model into a history-aware policy. arXiv preprint arXiv:2510.00695, 2025.
  20. 20.Jinhao Li, Yuxuan Cong, Yingqiao Wang, Hao Xia, Shan Huang, Yijia Zhang, Ningyi Xu, and Guohao Dai. STEP: Warm-started visuomotor policies with spatiotemporal consistency prediction. arXiv preprint arXiv:2602.08245, 2026.
  21. 21.Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. Flow matching for generative modeling. In International Conference on Learning Representations, 2023.
  22. 22.Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. In International Conference on Learning Representations, 2023.
  23. 23.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019.
  24. 24.Guanxing Lu, Zifeng Gao, Tianxing Chen, Wenxun Dai, Ziwei Wang, Wenbo Ding, and Yansong Tang. ManiCM: Real-time 3D diffusion policy via consistency model for robotic manipulation. arXiv preprint arXiv:2406.01586, 2024.
  25. 25.Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, and Roberto Martín-Martín. What matters in learning from offline human demonstrations for robot manipulation. In Conference on Robot Learning, 2021.
  26. 26.Ajay Mandlekar, Soroush Nasiriany, Bowen Wen, Iretiayo Akinola, Yashraj Narang, Linxi Fan, Yuke Zhu, and Dieter Fox. MimicGen: A data generation system for scalable robot learning using human demonstrations. In Conference on Robot Learning, 2023.
  27. 27.Robert J. McCann. A convexity principle for interacting gases. Advances in Mathematics, 128(1):153–179, 1997.
  28. 28.William Peebles and Saining Xie. Scalable diffusion models with transformers. In IEEE/CVF International Conference on Computer Vision, 2023.
  29. 29.Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, James Tanner, Quan Vuong, Homer Walke, Anna Walling, Haohuan Wang, Lili Yu, and Ury Zhilinsky. π0.5\pi_{0.5}: A vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054, 2025.
  30. 30.Aram-Alexandre Pooladian, Heli Ben-Hamu, Carles Domingo-Enrich, Brandon Amos, Yaron Lipman, and Ricky T. Q. Chen. Multisample flow matching: Straightening flows with minibatch couplings. In International Conference on Machine Learning, 2023.
  31. 31.Aaditya Prasad, Kevin Lin, Jimmy Wu, Linqi Zhou, and Jeannette Bohg. Consistency policy: Accelerated visuomotor policies via consistency distillation. In Robotics: Science and Systems, 2024.
  32. 32.Qwen Team. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.
  33. 33.Yuyang Shi, Valentin De Bortoli, Andrew Campbell, and Arnaud Doucet. Diffusion Schrödinger bridge matching. In Advances in Neural Information Processing Systems, 2023.
  34. 34.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021.
  35. 35.Alexander Tong, Kilian Fatras, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Guy Wolf, and Yoshua Bengio. Improving and generalizing flow-based generative models with minibatch optimal transport. Transactions on Machine Learning Research, 2024a.
  36. 36.Alexander Tong, Nikolay Malkin, Kilian Fatras, Lazar Atanackovic, Yanlei Zhang, Guillaume Huguet, Guy Wolf, and Yoshua Bengio. Simulation-free Schrödinger bridges via score and flow matching. In International Conference on Artificial Intelligence and Statistics, 2024b.
  37. 37.TRI LBM Team. A careful examination of large behavior models for multitask dexterous manipulation. arXiv preprint arXiv:2507.05331, 2025.
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, 2017.
  39. 39.Andrew Wagenmaker, Mitsuhiko Nakamoto, Yunchu Zhang, Seohong Park, Waleed Yagoub, Anusha Nagabandi, Abhishek Gupta, and Sergey Levine. Steering your diffusion policy with latent space reinforcement learning. arXiv preprint arXiv:2506.15799, 2025.
  40. 40.Zhendong Wang, Zhaoshuo Li, Ajay Mandlekar, Zhenjia Xu, Jiaojiao Fan, Yashraj Narang, Linxi Fan, Yuke Zhu, Yogesh Balaji, Mingyuan Zhou, Ming-Yu Liu, and Yu Zeng. One-step diffusion policy: Fast visuomotor policies via diffusion distillation. In International Conference on Machine Learning, 2025.
  41. 41.Yuxin Wu and Kaiming He. Group normalization. In European Conference on Computer Vision, 2018.
  42. 42.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In IEEE/CVF International Conference on Computer Vision, 2023.
  43. 43.Tony Z. Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. Learning fine-grained bimanual manipulation with low-cost hardware. In Robotics: Science and Systems, 2023.

Citation

MLA
Kang, S., et al. “WarmPrior: Straightening Flow-Matching Policies with Temporal Priors”. arXiv, 2026, https://doi.org/10.48550/arxiv.2605.13959.
APA
Kang, S., Kim, C., Wang, K., Zhao, L., & Lee, K. (2026). WarmPrior: Straightening Flow-Matching Policies with Temporal Priors. arXiv. https://doi.org/10.48550/arxiv.2605.13959
Chicago
Kang, S., C. Kim, K. Wang, L. Zhao, and K. Lee. 2026. “WarmPrior: Straightening Flow-Matching Policies with Temporal Priors”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2605.13959.
Harvard
Kang, S. et al. (2026) “WarmPrior: Straightening Flow-Matching Policies with Temporal Priors”. arXiv. Available at: https://doi.org/10.48550/arxiv.2605.13959.
Vancouver
1. Kang S, Kim C, Wang K, Zhao L, Lee K (2026) WarmPrior: Straightening Flow-Matching Policies with Temporal Priors. https://doi.org/10.48550/arxiv.2605.13959

BibTeX

@misc{https://doi.org/10.48550/arxiv.2605.13959,
  doi = {10.48550/ARXIV.2605.13959},
  url = {https://arxiv.org/abs/2605.13959},
  author = {Kang, Sinjae and Kim, Chanyoung and Wang, Kaixin and Zhao, Li and Lee, Kimin},
  keywords = {Machine Learning (cs.LG), Artificial Intelligence (cs.AI), Robotics (cs.RO), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {WarmPrior: Straightening Flow-Matching Policies with Temporal Priors},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/