WarmPrior: Straightening Flow-Matching Policies with Temporal Priors
Sinjae KangChanyoung KimKaixin WangLi ZhaoKimin Lee
Introduces WarmPrior, a temporal source distribution derived from recent action history that straightens probability paths in flow-matching policies to boost robotic manipulation performance across imitation and reinforcement learning.
Modern robot control increasingly relies on generative models, such as diffusion and flow-matching policies, to translate camera images and sensor data into precise physical movements. Conventionally, these systems generate actions by transforming a completely random, standard normal noise distribution into a sequence of robot actions. However, treating the starting point as pure, stateless noise ignores the continuous and predictable flow of physical motion. This design flaw forces the robot's control policy to reconstruct every motion trajectory from scratch at each step, resulting in curved generation paths, slower computation, and erratic behavioral switching between different task strategies.
The main objective of the article is to demonstrate that replacing standard random noise with a temporally informed starting distribution, termed WarmPrior, significantly enhances the execution performance, computational speed, and learning efficiency of generative robot policies without modifying the underlying neural networks or training losses.
The approach introduces two simple mechanisms to anchor the starting noise distribution to the robot's own recent movement history, combined with a controlled amount of residual noise. The first variant, WarmPrior-Past, anchors the initial noise to the robot's previously executed action. The second, WarmPrior-Preview, trains the model to forecast twice the required action steps and uses its previous future forecast as the starting point for the current step. The authors evaluated this framework across rigorous simulation suites (Robomimic and MimicGen) and physical robot hardware using a Franka robotic arm across diverse manipulation tasks, testing across varying computational budgets and combining the method with reinforcement learning.
The key findings show consistent, broad-based improvements across all tested domains. First, WarmPrior substantially increases robotic task success rates compared to standard noise baselines, with the most pronounced gains occurring on complex, multi-human demonstration tasks and when compute budgets are constrained to single-step inference (for example, improving success rates from 65.9% to 77.8% on a challenging peg-insertion task). Second, WarmPrior-Preview consistently outperforms WarmPrior-Past because future forecasts provide a closer approximation to the target action than simple past-action persistence. Third, the method straightens the learned mathematical transport trajectories by up to 44%, effectively removing the directional ambiguity that normally curves generative flows. Fourth, in reinforcement learning fine-tuning, WarmPrior restricts the agent's exploration space to a 3-times tighter, meaningful region, enabling it to surpass a 95% success rate on the difficult multi-robot transport task for the first time.
These results demonstrate that the starting noise distribution is a highly practical lever for improving robotic systems. Because WarmPrior straightens generation trajectories, it allows robots to operate accurately at minimal inference steps, directly reducing computational latency, hardware costs, and power requirements. It also resolves dangerous oscillations between conflicting strategies during execution by providing smooth temporal continuity. Furthermore, WarmPrior complements existing inference-time smoothing techniques like Real-Time Chunking, as stacking both approaches yields higher success rates than either method alone.
Based on these findings, engineering and product teams building generative robot policies should replace default random Gaussian starting distributions with temporally grounded priors. Teams should prioritize the preview-based variant when planning horizons allow, or use the past-action variant when simplicity is paramount, maintaining moderate noise levels to avoid collapsing into brittle deterministic behaviors. In addition, organizations fine-tuning diffusion-based policies via reinforcement learning should adopt bounded, anchor-centered action spaces to dramatically reduce sample training time.
A key boundary condition of the study is that the noise variance must remain balanced; reducing noise completely to zero degrades the model into a standard regression approach that fails entirely on multimodal tasks. While the findings provide high confidence across multiple benchmark architectures and tabletop manipulation settings, further validation is recommended on highly agile, long-horizon mobile manipulation platforms in unstructured real-world environments before enterprise-wide deployment.
- Paper: Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow, Xingchao Liu et al. (2023). Introduces the Rectified Flow framework and the concept of straightening probability trajectories, which WarmPrior directly builds upon to motivate its temporal prior construction.
- Paper: Diffusion policy: Visuomotor policy learning via action diffusion, Cheng Chi et al. (2023). Establishes the foundational paradigm of diffusion- and generative-based visuomotor policies for robotic manipulation that WarmPrior aims to improve.
- Paper: π0: A Vision-Language-Action Flow Model for General Robot Control, Kevin Black et al. (2024). Demonstrates flow-matching architectures for continuous action generation in generalist robot control, representing the core policy class adapted by WarmPrior.
- Paper: Flow Q-Learning, Seohong Park et al. (2025). Explores generative flow-based policies and behavioral modeling within reinforcement learning pipelines, providing important background for prior-space exploration.
- Paper: Matching Normalizing Flows and Probability Paths on Manifolds, Heli Ben-Hamu et al. (2022). Provides fundamental formulations for continuous flow matching and probability path construction essential to understanding trajectory straightening.
- Paper: What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models, Yu Peng et al. (2026). Extends reinforcement learning fine-tuning for vision-language-action policies by introducing action-level invariance and sensitivity objectives to handle out-of-distribution visual shifts.
- Paper: TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers, Bin Yu et al. (2026). Applies flow-matching action generation within an asymmetric dual-pathway architecture to mitigate catastrophic forgetting in vision-language-action models.
- Paper: Generative Modeling via Drifting, Mingyang Deng et al. (2026). Proposes an alternative single-step generative modeling paradigm using drifting fields that advances real-time sampling efficiency for control and vision tasks.
