FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies

Moritz ReussHongyi ZhouMarcel RuhleOmer Erdincc YaugmurluFabian OttoRudolf Lioutikov

article2025arXiv71 citations

Introduces FLOWER, an efficient 950-million-parameter vision-language-action policy that matches multi-billion-parameter generalist models across 190 tasks and achieves a new state of the art on CALVIN with only 200 GPU hours of pretraining.

Listen

Developing generalist artificial intelligence policies for robotic manipulation is essential for practical automation, yet existing models face severe computational barriers. Leading systems rely on multi-billion-parameter foundation architectures that require tens of thousands of computing hours to train and demand high-end computing memory for real-time deployment. This high cost restricts research access and limits practical deployment on physical hardware.

The article demonstrates an efficient robotic manipulation architecture that matches or surpasses existing state-of-the-art generalist models across diverse robot platforms while containing fewer than one billion parameters. It also evaluates how architectural modifications affect training efficiency, runtime latency, and real-world task generalization.

The authors develop two structural innovations and integrate them into a 947-million-parameter model called FLOWER. First, intermediate-modality fusion extracts contextual features from the middle of a vision-language backbone—pruning between 30% and 50% of its layers—to retain semantic understanding while reallocating capacity to an action-generation module. Second, an action-specific normalization controller shares core modulation weights across layers to reduce parameter overhead by over 20%. The resulting system was pretrained in approximately 200 graphic processing unit (GPU) hours on roughly 250,000 trajectories and evaluated across 190 tasks in ten simulation and real-world benchmarks involving single-arm and dual-arm robots.

The findings show that FLOWER achieves state-of-the-art or highly competitive performance across diverse benchmarks while dramatically reducing computational demands. Pretraining the model required less than 1% of the compute used by leading multi-billion-parameter baselines. In physical kitchen experiments, FLOWER achieved a 61.0% overall task success rate compared to 31.0% for the next-best baseline, and maintained a 51.0% average success rate under severe environmental disruptions such as novel objects, dim lighting, and clutter. Furthermore, the architecture achieved an inference throughput of 311 Hz and required only 1.85 gigabytes of memory, operating over 50 times faster than comparable models while using a fraction of the hardware memory.

These results demonstrate that robotics developers do not need massive multi-billion-parameter models to achieve versatile control. By balancing semantic reasoning and action generation, organizations can significantly lower training expenses, accelerate development cycles, and deploy capable generalist policies directly on cost-effective, commodity hardware without sacrificing task reliability.

Organizations developing robotic systems should adopt intermediate fusion and parameter-efficient normalization techniques to optimize edge deployment. Before widespread physical rollout, teams should conduct pilot studies in target operating environments, particularly for fine-precision manipulation tasks and zero-shot transfers.

While the empirical results are strong, eight of the ten evaluated benchmark suites were conducted in simulated environments, and physical testing was confined to stationary manipulation rather than mobile navigation or locomotion. In addition, the policy relies on iterative trajectory sampling rather than single-pass generation and demonstrated lower zero-shot success on one complex simulation benchmark, indicating that further testing is necessary before deploying the system in critical low-tolerance environments.

Cover for FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies

Abstract

Developing efficient Vision-Language-Action (VLA) policies is crucial for practical robotics deployment, yet current approaches face prohibitive computational costs and resource requirements. Existing diffusion-based VLA policies require multi-billion-parameter models and massive datasets to achieve strong performance. We tackle this efficiency challenge with two contributions: intermediate-modality fusion, which reallocates capacity to the diffusion head by pruning up to 50%50\% of LLM layers, and action-specific Global-AdaLN conditioning, which cuts parameters by 20%20\% through modular adaptation. We integrate these advances into a novel 950 M-parameter VLA called FLOWER. Pretrained in just 200 H100 GPU hours, FLOWER delivers competitive performance with bigger VLAs across 190190 tasks spanning ten simulation and real-world benchmarks and demonstrates robustness across diverse robotic embodiments. In addition, FLOWER achieves a new SoTA of 4.53 on the CALVIN ABC benchmark. Demos, code and pretrained weights are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Intermediate Modality Fusion Vision-Language-Action-Models
  • 3.2 Cross-Action Space Flow Transformer
  • 3.3 Rectified Flow for Action Generation
  • 3.4 FLOWER: Efficient Flow-based Vision-Language-Action Models
  • 4 Evaluation
  • 4.1 Evaluation of Critical Design Decisions for Efficient Flow VLAs
  • 4.2 Simulation Experiments
  • 4.3 Real-World Evaluation and Generalization
  • 5 Conclusion
  • 6 Limitations
  • 7 Acknowledgments
  • References
  • A Pretraining Details
  • A.1 Pretraining Ablation Experiences
  • A.2 Language Prompt for the VLA
  • A.3 Custom Learning Rate Scheduler
  • A.4 Details for Cross-Action Space Flow Transformer
  • B Detailed Experiments
  • B.1 SIMPLER Benchmark Tasks
  • B.1.1 LIBERO Benchmark.
  • B.1.2 Aloha Benchmark.
  • B.2 Real Kitchen Play Dataset.
  • B.2.1 Pretraining Details for Baselines for the Real World Kitchen
  • B.2.2 Failure Cases for different Policies
  • B.3 Generalization Experiments
  • B.3.1 Novel Task Compositions
  • C Extended Related Work

Knowls

  1. Knowl 1 — FLOWER Model Architecture and Parameter Breakdown

    model/method

    Florence With Embodied Flow (FLOWER) is a 947M-parameter Vision-Language-Action (VLA) policy designed for efficient cross-embodiment robot manipulation. It integrates a pruned Vision-Language Model (VLM) backbone with a continuous Flow Transformer head using cross-attention and action-space adaptive normalization.

    The parameter breakdown across the components of FLOWER is:

    • Vision Transformer (ViT): 360M parameters
    • VLM Encoder: 205M parameters (retaining the encoder of Florence-2-L while pruning the full decoder)
    • Flow Transformer: 339M parameters (18 layers, latent dimension d=1024d = 1024, 16 attention heads, 1D Rotary Position Embeddings [RoPE], QK-norm attention, and SwiGLU feedforward blocks with RMSNorm)
    • Global-AdaLN: 28.3M parameters
    • Action Encoders: 3.2M parameters (2-layered MLPs for mapping heterogeneous action spaces into the transformer latent space)
    • Action Heads / Decoders: 31.8K parameters (linear attention decoders)
    • Timestep Embedder: 1.3M parameters
    • Frequency Embedder: 1.3M parameters
    • Conditioning Linear Projection: 1.0M parameters
    • Conditioning Normalization: 1.0K parameters

    Total parameter count is 947M, requiring 1.85 GB of VRAM during inference.

  2. Knowl 2 — Intermediate-Modality Fusion Strategy for Vision-Language-Action Models

    model/method

    Intermediate-modality fusion extracts multimodal representations from an intermediate layer of a pretrained VLM backbone rather than using raw tokens at the input (early fusion) or the final output layer (late fusion). This design is motivated by the fact that intermediate transformer layers capture broad semantic grounding and spatial relationships, whereas the final layers overspecialize in next-token language prediction.

    Depending on the VLM architecture, pruning is applied as follows:

    • Encoder-Decoder VLMs (e.g., Florence-2): The entire language decoder is discarded, retaining only the visual encoder and the text encoder layers. This cuts the number of LLM layers by 50% while improving action-generation efficiency and performance.
    • Decoder-Only VLMs (e.g., SmolFlow2-Video): The final 30% of transformer layers are dropped, removing 20% to 35% of total parameters and reducing per-step inference latency.

    The intermediate hidden states from the VLM are projected through a linear layer, normalized with RMSNorm, and integrated into every layer of the Flow Transformer via cross-attention with Query-Key (QK) normalization.

  3. Knowl 3 — Action-Space Global-AdaLN-Zero Conditioning

    model/method

    In standard Diffusion Transformers (DiTs), Adaptive Layer Normalization (AdaLN-Zero) allocates separate scale and shift projection parameters for each transformer layer, which can introduce up to 30% extra parameters. Action-Space Global-AdaLN-Zero replaces this per-layer overhead by sharing a single modulation generator across all transformer layers while producing action-category-specific modulation signals (e.g., delta end-effector position vs. joint angles).

    To compensate for the reduced expressiveness of shared global modulation weights across layers, lightweight Low-Rank Adaptation (LoRA) adapters are injected into each individual transformer block. The global scale-and-shift modulation vectors are initialized to zero to ensure training stability. This mechanism reduces the parameter count of the normalization and conditioning module by 20% relative to naive AdaLN-Zero without degradation in expressiveness or trajectory prediction accuracy.

  4. Knowl 4 — Rectified Flow Objective for Continuous Action Generation

    equation

    FLOWER formulates conditional action sequence prediction using Rectified Flow, learning straight-line probability paths between standard Gaussian noise and continuous robot action trajectories.

    Given ground-truth action trajectory aˉn,k∈Rda\bar{a}_{n,k} \in \mathbb{R}^{d_a} of chunk length kk for embodiment ee, current state sˉn\bar{s}_n, and language goal gg, the interpolated sample ztz_t at flow time t∈[0,1]t \in [0, 1] is constructed as: zt=(1−t)aˉn,k+tz1,z1∼N(0,I)z_t = (1 - t)\bar{a}_{n,k} + t z_1, \quad z_1 \sim \mathcal{N}(0, I) where tt is sampled from t∼σ(N(0,1))t \sim \sigma(\mathcal{N}(0, 1)) with sigmoid function σ(⋅)\sigma(\cdot), and z1z_1 is standard Gaussian noise.

    The velocity prediction network vθ(zt,t,sˉn,g,e)v_\theta(z_t, t, \bar{s}_n, g, e) is trained by minimizing the mean squared error loss: L(θ)=Et,z1[∥z1−aˉn,k−vθ(zt,t,sˉn,g,e)∥2]\mathcal{L}(\theta) = \mathbb{E}_{t, z_1} \left[ \left\| z_1 - \bar{a}_{n,k} - v_\theta(z_t, t, \bar{s}_n, g, e) \right\|^2 \right]

    At inference time, trajectory generation is performed via numerical integration across N=4N = 4 denoising steps for single-arm manipulation and N=8N = 8 denoising steps for high-frequency (50 Hz) dual-arm manipulation.

  5. Knowl 5 — Pretraining Recipe and OXE-Soup Dataset Distribution

    experimental setup

    FLOWER is pretrained on an 8-dataset mixture (~250,000 trajectories total) sampled from the Open X-Embodiment (OXE) repository, prioritizing diverse scenes, distractors, and robot embodiments (Franka Panda, WidowX, Google Robot, XARM):

    Dataset Proportion (%)
    bridge_dataset 28.62%
    fractal_data 24.68%
    droid 23.50%
    cmu_play_fusion 6.15%
    dobbe 5.94%
    libero_10_no_noops 4.41%
    libero_goal_no_noops 4.07%
    real_kitchen_lang 2.64%

    The action representation mix is 74% delta End-Effector (delta-EEF) and 26% single-arm joint-state data. Pretraining uses a fixed action chunk length of 20, conditioned on a single static image. Optimization uses AdamW with BF16 precision, an initial warmup followed by constant and cosine decay schedules (peak learning rate 1×10−41\times 10^{-4} for the Flow Transformer, 1×10−51\times 10^{-5} for the VLM), a batch size of 256 with 4 gradient accumulation steps (effective batch size 1024), and completes 350,000–360,000 steps in 48 hours using 4 NVIDIA H100 GPUs (~192–200 GPU hours).

  6. Knowl 6 — Ablation of Fusion Strategy, Layer Pruning, and VLA Components

    empirical result

    Ablation experiments conducted on CALVIN ABC (34 free-form language manipulation tasks evaluated on unseen split D) and LIBERO-Long (long-horizon tasks) evaluate fusion types, pruning ratios, and architectural components.

    Fusion Strategy Florence-VLM Success (%) Smol-VLM Success (%)
    CALVIN-ABC LIBERO-Long CALVIN-ABC LIBERO-Long
    Early 57.1 ±\pm 5.3 33.4 ±\pm 6.0 25.8 ±\pm 3.9 44.5 ±\pm 2.7
    Intermediate (ours) 89.5 ±\pm 1.0 93.4 ±\pm 2.0 72.1 ±\pm 5.0 70.7 ±\pm 2.3
    Late 71.2 ±\pm 2.2 61.8 ±\pm 2.5 66.3 ±\pm 2.0 69.2 ±\pm 1.9

    For SmolFlow-500M layer pruning, dropping 30% of transformer layers yields optimal performance on CALVIN ABC (72.1%±5.0%72.1\% \pm 5.0\%) compared to full model (66.3%±2.0%66.3\% \pm 2.0\%), 20% pruning (68.6%±3.2%68.6\% \pm 3.2\%), and 50% pruning (66.4%±6.4%66.4\% \pm 6.4\%).

    Ablations of individual modules on CALVIN ABC measured by average completed sequence length (out of 5 consecutive tasks) show:

    • Full FLOWER: 4.44±0.044.44 \pm 0.04
    • With standard AdaLN: 4.43±0.034.43 \pm 0.03
    • With smaller Flow head (d=384d=384, 6 layers): 2.60±0.092.60 \pm 0.09
    • Without VLM fine-tuning (frozen VLM): 2.65±0.362.65 \pm 0.36
    • With smaller Florence backbone: 4.26±0.044.26 \pm 0.04
    • With discrete action tokenization head: 1.12±0.121.12 \pm 0.12
    • Without VLM (Flow Transformer only): 3.42±0.073.42 \pm 0.07
  7. Knowl 7 — Simulation Benchmark Performance on CALVIN, LIBERO, Aloha, and SIMPLER

    empirical result

    FLOWER matches or exceeds state-of-the-art specialist and generalist models across four major simulation benchmarks:

    1. CALVIN (Average Completed Sequence Length out of 5):

      • ABC →\rightarrow D (Zero-Shot Generalization): FLOWER reaches 4.53±0.044.53 \pm 0.04 (with pretraining) and 4.44±0.044.44 \pm 0.04 (from scratch), outperforming VPP (4.294.29), Seer (4.284.28), MoDE (4.014.01), RoboDual (3.663.66), 3DDA (3.353.35), and OpenVLA (3.273.27).
      • ABCD →\rightarrow D (Data Scaling): FLOWER achieves 4.67±0.044.67 \pm 0.04, outperforming MDT (4.524.52), MoDE (4.394.39), and GR-1 (4.214.21).
      • D →\rightarrow D: FLOWER achieves 4.35±0.024.35 \pm 0.02 vs. RoboUniView (3.853.85) and MDT (3.723.72).
    2. LIBERO (Task Success Rates):

      • Spatial: 97.5%±0.8%97.5\% \pm 0.8\%
      • Object: 99.1%±0.4%99.1\% \pm 0.4\%
      • Goal: 96.1%±0.6%96.1\% \pm 0.6\%
      • Long: 94.9%±1.2%94.9\% \pm 1.2\% (the only generalist policy exceeding 90%90\%, where π0\pi_0 achieves 85.2%85.2\% and OpenVLA achieves 53.7%53.7\%)
      • LIBERO-90: 94.7%±1.0%94.7\% \pm 1.0\%
      • 4-Suite Average: 96.9%±0.7%96.9\% \pm 0.7\% vs. π0\pi_0 (94.2%94.2\%) and OpenVLA (76.5%76.5\%).
    3. Aloha Simulation (500 episodes per task, 50 Hz control):

      • Cube Transfer: 82%82\% (joint-state pretrained) vs. ACT (87%87\%) and Diffusion Policy (6%6\%).
      • Peg Insertion: 26%26\% (joint-state pretrained) vs. ACT (21%21\%) and Diffusion Policy (0%0\%).
    4. SIMPLER (Real2Sim Evaluation):

      • WidowX + Bridge: Average success of 45.0%45.0\% (Put Carrot 13%13\%, Spoon on Towel 71%71\%, Stack Blocks 8%8\%, Eggplant in Basket 88%88\%) vs. CrossFormer (30.0%30.0\%), Octo (16.0%16.0\%), and OpenVLA (1.0%1.0\%).
      • Google Robot: Average success of 31.9%31.9\% vs. RT-1-X (42.4%42.4\%) and OpenVLA (24.5%24.5\%).
  8. Knowl 8 — Real-World Robotic Manipulation and Generalization Performance

    empirical result

    FLOWER was evaluated on a physical Franka Panda robot setup across 20 kitchen manipulation tasks (e.g., interacting with a stove, microwave, oven, toaster, and sink) operating at 6 Hz in joint position space, fine-tuned on 417 demonstrations (45 minutes total teleoperated data):

    • Multi-Task In-Domain Performance: Across all 20 tasks (5 trials each from randomized start positions), FLOWER achieves an overall average success rate of 61.0%61.0\% (average rank 1.25), outperforming OpenVLA (31.0%31.0\%, rank 2.10), CrossFormer (22.0%22.0\%, rank 2.20), and Octo (10.0%10.0\%, rank 2.70).

    • Zero-Shot Out-of-Distribution Generalization:

      • Novel Objects (e.g., black donut, tennis ball, glove, carrot, red cup): FLOWER 33.3%33.3\% vs. OpenVLA 10.0%10.0\%
      • Flashlight Lighting: FLOWER 50.0%50.0\% vs. OpenVLA 25.0%25.0\%
      • Background Distractors / Clutter: FLOWER 69.5%69.5\% vs. OpenVLA 41.7%41.7\%
      • New Multi-Step Task Compositions (e.g., 'Open and Close All Appliances', 'Move Items Between Stovetop and Sink'): FLOWER 51.1%51.1\% (average sequence length 2.56) vs. OpenVLA 16.7%16.7\% (average sequence length 0.55)
      • Generalization Average: FLOWER 51.0%51.0\% vs. OpenVLA 23.4%23.4\% (a 118.0% relative improvement).
  9. Knowl 9 — Inference Latency, Throughput, and VRAM Efficiency Comparison

    empirical result

    Inference metrics measured on a single consumer NVIDIA RTX 4090 GPU (24GB VRAM) using BF16 precision across 1000 generation steps:

    Method Parameter Count Throughput (Hz) ↑\uparrow Latency (s) ↓\downarrow VRAM (MB) ↓\downarrow
    Diffusion Policy (CNN) 0.26B 130.67 0.341 517
    OpenVLA 7.00B 6.09 0.164 14574
    π0\pi_0 3.30B 288.11 0.104 6692
    FLOWER + Late Fusion 1.15B 287.36 0.055 2235
    FLOWER (ours) 0.95B 311.04 0.052 1848

    FLOWER achieves 311.04 Hz control frequency, running 8% faster than π0\pi_0 and ~50x faster than OpenVLA. Its memory footprint (1848 MB) consumes 27.6% of π0\pi_0's VRAM and 12.7% of OpenVLA's VRAM, allowing deployment on commodity edge GPUs.

  10. Knowl 10 — Stated Limitations of FLOWER

    limitation

    The authors identify five specific limitations of FLOWER:

    1. Iterative Sampling Latency: Rectified Flow requires multi-step numerical integration (N=4N = 4 to 88 steps), which remains computationally slower than single-step forward passes from deterministic policies.
    2. Embodiment Scope: Validation is restricted to three robotic manipulation action formats (single-arm delta-EEF, single-arm joint space, and dual-arm bimanual joint space); generalization to mobile manipulation, navigation, or humanoid locomotion is not demonstrated.
    3. Performance Gap on SIMPLER Google Robot: FLOWER underperforms RT-1-X on the SIMPLER Google Robot benchmark (31.9%31.9\% vs. 42.4%42.4\%), suggesting that zero-shot real2sim transfer on certain domains benefits from larger model capacity or wider pretraining data.
    4. Model Size for Edge Devices: Although reduced to 947M parameters, the model size may still be too large for low-power microcontrollers or ultra-constrained real-time robotic systems.
    5. Simulation Evaluation Bias: Eight out of the ten evaluated benchmark suites are conducted in simulation, which may not fully reflect real-world physical dynamics and noise.

Coverage note — No substantial contributed material was omitted. All primary contributions—intermediate modality fusion, action-space Global-AdaLN, Rectified Flow training, pretraining data recipe, simulation and real-world evaluation, inference efficiency, and limitations—are fully covered.

References

  1. 1.M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024.
  2. 2.A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich. Rt-1: Robotics transformer for real-world control at scale. In arXiv preprint arXiv:2212.06817, 2022.
  3. 3.Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy. https://octo-models.github.io, 2023.
  4. 4.X. Li, P. Li, M. Liu, D. Wang, J. Liu, B. Kang, X. Ma, T. Kong, H. Zhang, and H. Liu. Towards generalist robot policies: What matters in building vision-language-action models. arXiv preprint arXiv:2412.14058, 2024.
  5. 5.D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023.
  6. 6.O. X.-E. Collaboration. Open X-Embodiment: Robotic learning datasets and RT-X models. https://arxiv.org/abs/2310.08864, 2023.
  7. 7.K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. pi 0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024.
  8. 8.S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu. Rdt-1b: a diffusion foundation model for bimanual manipulation. arXiv preprint arXiv:2410.07864, 2024.
  9. 9.J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734, 2025.
  10. 10.M. Reuss, J. Pari, P. Agrawal, and R. Lioutikov. Efficient diffusion transformer policies with mixture of expert denoisers for multitask learning, 2024.
  11. 11.L. Pinto and A. Gupta. Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours. In 2016 IEEE international conference on robotics and automation (ICRA), pages 3406–3413. IEEE, 2016.
  12. 12.X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao. Evaluating real-world robot manipulation policies in simulation. arXiv preprint arXiv:2405.05941, 2024.
  13. 13.A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023.
  14. 14.K. Bousmalis, G. Vezzani, D. Rao, C. Devin, A. X. Lee, M. Bauza, T. Davchev, Y. Zhou, A. Gupta, A. Raju, et al. Robocat: A self-improving foundation agent for robotic manipulation. arXiv preprint arXiv:2306.11706, 2023.
  15. 15.N. M. M. Shafiullah, A. Rai, H. Etukuru, Y. Liu, I. Misra, S. Chintala, and L. Pinto. On bringing robots home. arXiv preprint arXiv:2311.16098, 2023.
  16. 16.H. Etukuru, N. Naka, Z. Hu, S. Lee, J. Mehu, A. Edsinger, C. Paxton, S. Chintala, L. Pinto, and N. M. M. Shafiullah. Robot utility models: General policies for zero-shot deployment in new environments. arXiv preprint arXiv:2409.05865, 2024.
  17. 17.Q. Bu, H. Li, L. Chen, J. Cai, J. Zeng, H. Cui, M. Yao, and Y. Qiao. Towards synergistic, generalized, and efficient dual-system for robotic manipulation. arXiv preprint arXiv:2410.08001, 2024.
  18. 18.Y. Shentu, P. Wu, A. Rajeswaran, and P. Abbeel. From llms to actions: Latent codes as bridges in hierarchical robot control. arXiv preprint arXiv:2405.04798, 2024.
  19. 19.J. Wen, Y. Zhu, J. Li, M. Zhu, K. Wu, Z. Xu, N. Liu, R. Cheng, C. Shen, Y. Peng, et al. Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation. arXiv preprint arXiv:2409.12514, 2024.
  20. 20.Y. Hu, Y. Guo, P. Wang, X. Chen, Y.-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen. Video prediction policy: A generalist robot policy with predictive visual representations, 2024. URL https://arxiv.org/abs/2412.14803.
  21. 21.J. Li, Y. Zhu, Z. Tang, J. Wen, M. Zhu, X. Liu, C. Li, R. Cheng, Y. Peng, and F. Feng. Improving vision-language-action models via chain-of-affordance. arXiv preprint arXiv:2412.20451, 2024.
  22. 22.Y. Yue, Y. Wang, B. Kang, Y. Han, S. Wang, S. Song, J. Feng, and G. Huang. Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution. Advances in Neural Information Processing Systems, 37:56619–56643, 2024.
  23. 23.L. Gao, T. D. la Tour, H. Tillman, G. Goh, R. Troll, A. Radford, I. Sutskever, J. Leike, and J. Wu. Scaling and evaluating sparse autoencoders. arXiv preprint arXiv:2406.04093, 2024.
  24. 24.B. Xiao, H. Wu, W. Xu, X. Dai, H. Hu, Y. Lu, M. Zeng, C. Liu, and L. Yuan. Florence-2: Advancing a unified representation for a variety of vision tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4818–4829, 2024.
  25. 25.A. Marafioti, O. Zohar, M. Farré, M. Noyan, E. Bakouch, P. Cuenca, C. Zakka, L. B. Allal, A. Lozhkov, N. Tazi, V. Srivastav, J. Lochner, H. Larcher, M. Morlon, L. Tunstall, L. von Werra, and T. Wolf. Smolvlm: Redefining small and efficient multimodal models. arXiv preprint arXiv:2504.05299, 2025.
  26. 26.B. Zhang and R. Sennrich. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32, 2019.
  27. 27.W. Peebles and S. Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4195–4205, 2023.
  28. 28.Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022.
  29. 29.X. Liu, C. Gong, and Q. Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022.
  30. 30.F. Lin, Y. Hu, P. Sheng, C. Wen, J. You, and Y. Gao. Data scaling laws in imitation learning for robotic manipulation. arXiv preprint arXiv:2410.18647, 2024.
  31. 31.A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y. J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y. Lee, M. Memmel, S. Park, I. Radosavovic, K. Wang, A. Zhan, K. Black, C. Chi, K. B. Hatch, S. Lin, J. Lu, J. Mercat, A. Rehman, P. R. Sanketi, A. Sharma, C. Simpson, Q. Vuong, H. R. Walke, B. Wulfe, T. Xiao, J. H. Yang, A. Yavary, T. Z. Zhao, C. Agia, R. Baijal, M. G. Castro, D. Chen, Q. Chen, T. Chung, J. Drake, E. P. Foster, J. Gao, D. A. Herrera, M. Heo, K. Hsu, J. Hu, D. Jackson, C. Le, Y. Li, K. Lin, R. Lin, Z. Ma, A. Maddukuri, S. Mirchandani, D. Morton, T. Nguyen, A. O’Neill, R. Scalise, D. Seale, V. Son, S. Tian, E. Tran, A. E. Wang, Y. Wu, A. Xie, J. Yang, P. Yin, Y. Zhang, O. Bastani, G. Berseth, J. Bohg, K. Goldberg, A. Gupta, A. Gupta, D. Jayaraman, J. J. Lim, J. Malik, R. Martín-Martín, S. Ramamoorthy, D. Sadigh, S. Song, J. Wu, M. C. Yip, Y. Zhu, T. Kollar, S. Levine, and C. Finn. Droid: A large-scale in-the-wild robot manipulation dataset. 2024.
  32. 32.H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, et al. Bridgedata v2: A dataset for robot learning at scale. In Conference on Robot Learning, pages 1723–1736. PMLR, 2023.
  33. 33.P. Liu, Y. Orru, J. Vakil, C. Paxton, N. M. M. Shafiullah, and L. Pinto. Ok-robot: What really matters in integrating open-knowledge models for robotics. arXiv preprint arXiv:2401.12202, 2024.
  34. 34.O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard. Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022.
  35. 35.B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36, 2024.
  36. 36.T. Z. Zhao, V. Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023.
  37. 37.M. Reuss, Ö. E. Yağmurlu, F. Wenzel, and R. Lioutikov. Multimodal diffusion transformer: Learning versatile behavior from multimodal goals. In Robotics: Science and Systems, 2024.
  38. 38.T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki. 3d diffuser actor: Policy diffusion with 3d scene representations. arXiv preprint arXiv:2402.10885, 2024.
  39. 39.C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion. In Proceedings of Robotics: Science and Systems (RSS), 2023.
  40. 40.S. Haldar, Z. Peng, and L. Pinto. Baku: An efficient transformer for multi-task policy learning. arXiv preprint arXiv:2406.07539, 2024.
  41. 41.R. Doshi, H. R. Walke, O. Mees, S. Dasari, and S. Levine. Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation. In 8th Annual Conference on Robot Learning, 2024.
  42. 42.S. Dasari, O. Mees, S. Zhao, M. K. Srirama, and S. Levine. The ingredients for robotic diffusion transformers. arXiv preprint arXiv:2410.10088, 2024.
  43. 43.N. Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  44. 44.A. Henry, P. R. Dachapally, S. Pawar, and Y. Chen. Query-key normalization for transformers. arXiv preprint arXiv:2010.04245, 2020.
  45. 45.P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first International Conference on Machine Learning, 2024.
  46. 46.X. Li, M. Liu, H. Zhang, C. Yu, J. Xu, H. Wu, C. Cheang, Y. Jing, W. Zhang, H. Liu, H. Li, and T. Kong. Vision-language foundation models as effective robot imitators. arXiv preprint arXiv:2311.01378, 2023.
  47. 47.K. Black, M. Nakamoto, P. Atreya, H. Walke, C. Finn, A. Kumar, and S. Levine. Zero-shot robotic manipulation with pretrained image-editing diffusion models. arXiv preprint arXiv:2310.10639, 2023.
  48. 48.H. Wu, Y. Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong. Unleashing large-scale video generative pre-training for visual robot manipulation. In International Conference on Learning Representations, 2024.
  49. 49.T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki. 3d diffuser actor: Policy diffusion with 3d scene representations. In 8th Annual Conference on Robot Learning, 2024. URL https://openreview.net/forum?id=gqCQxObVz2.
  50. 50.Y. Tian, S. Yang, J. Zeng, P. Wang, D. Lin, H. Dong, and J. Pang. Predictive inverse dynamics models are scalable learners for robotic manipulation. https://arxiv.org/abs/2412.15109, 2024.
  51. 51.F. Liu, F. Yan, L. Zheng, Y. Huang, C. Feng, and L. Ma. Robouniview: Visual-language model with unified view representation for robotic manipulation. arXiv preprint 2406.18977, 2024.
  52. 52.K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. Fast: Efficient action tokenization for vision-language-action models. arXiv preprint arXiv:2501.09747, 2025.
  53. 53.D. Driess, J. T. Springenberg, B. Ichter, L. Yu, A. Li-Bell, K. Pertsch, A. Z. Ren, H. Walke, Q. Vuong, L. X. Shi, et al. Knowledge insulating vision-language-action models: Train fast, run fast, generalize better. arXiv preprint arXiv:2505.23705, 2025.
  54. 54.M. J. Kim, C. Finn, and P. Liang. Fine-tuning vision-language-action models: Optimizing speed and success. arXiv preprint arXiv:2502.19645, 2025.
  55. 55.R. Cadene, S. Alibert, A. Soare, Q. Gallouedec, A. Zouitine, and T. Wolf. Lerobot: State-of-the-art machine learning for real-world robotics in pytorch. https://github.com/huggingface/lerobot, 2024.
  56. 56.X. Jiang, P. Mattes, X. Jia, N. Schreiber, G. Neumann, and R. Lioutikov. A comprehensive user study on augmented reality-based data collection interfaces for robot learning. In Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, pages 333–342, 2024.
  57. 57.W. Huang, I. Mordatch, and D. Pathak. One policy to control them all: Shared modular policies for agent-agnostic control. In International Conference on Machine Learning, pages 4455–4464. PMLR, 2020.
  58. 58.T. Chen, A. Murali, and A. Gupta. Hardware conditioned policies for multi-robot transfer learning. Advances in Neural Information Processing Systems, 31, 2018.
  59. 59.A. Patel and S. Song. Get-zero: Graph embodiment transformer for zero-shot embodiment generalization. arXiv preprint arXiv:2407.15002, 2024.
  60. 60.J. H. Yang, D. Sadigh, and C. Finn. Polybot: Training one policy across robots while embracing variability. In Conference on Robot Learning, pages 2955–2974. PMLR, 2023.
  61. 61.J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  62. 62.Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2020.
  63. 63.J. Song, C. Meng, and S. Ermon. Denoising diffusion implicit models. In ICLR, 2021.
  64. 64.M. Reuss, M. Li, X. Jia, and R. Lioutikov. Goal conditioned imitation learning using score-based diffusion policies. In Proceedings of Robotics: Science and Systems (RSS), 2023.
  65. 65.X. Jia, D. Blessing, X. Jiang, M. Reuss, A. Donat, R. Lioutikov, and G. Neumann. Towards diverse behaviors: A benchmark for imitation learning with human demonstrations. arXiv preprint arXiv:2402.14606, 2024.
  66. 66.M. S. Albergo and E. Vanden-Eijnden. Building normalizing flows with stochastic interpolants. arXiv preprint arXiv:2209.15571, 2022.
  67. 67.N. Funk, J. Urain, J. Carvalho, V. Prasad, G. Chalvatzaki, and J. Peters. Actionflow: Equivariant, accurate, and efficient policies with spatially symmetric flow matching. arXiv preprint arXiv:2409.04576, 2024.
  68. 68.F. Zhang and M. Gienger. Affordance-based robot manipulation with flow matching. arXiv preprint arXiv:2409.01083, 2024.
  69. 69.M. Braun, N. Jaquier, L. Rozo, and T. Asfour. Riemannian flow matching policy for robot motion learning. arXiv preprint arXiv:2403.10672, 2024.
  70. 70.Z. Fu, T. Z. Zhao, and C. Finn. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation. arXiv preprint arXiv:2401.02117, 2024.

Citation

MLA
Reuss, M., et al. “FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies”. arXiv, 2025, http://arxiv.org/abs/2509.04996v1.
APA
Reuss, M., Zhou, H., Rühle, M., Yağmurlu, Ö. E., Otto, F., & Lioutikov, R. (2025). FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies. arXiv. http://arxiv.org/abs/2509.04996v1
Chicago
Reuss, M., H. Zhou, M. Rühle, Ö. E. Yağmurlu, F. Otto, and R. Lioutikov. 2025. “FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies”. arXiv. http://arxiv.org/abs/2509.04996v1.
Harvard
Reuss, M. et al. (2025) “FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2509.04996v1.
Vancouver
1. Reuss M, Zhou H, Rühle M, Yağmurlu ÖE, Otto F, Lioutikov R (2025) FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies. arXiv

BibTeX

@article{reuss2025flower,
  title = {FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies},
  author = {Reuss, Moritz and Zhou, Hongyi and Rühle, Marcel and Yağmurlu, Ömer Erdinç and Otto, Fabian and Lioutikov, Rudolf},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2509.04996v1},
  eprint = {2509.04996}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/