FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
Moritz ReussHongyi ZhouMarcel RuhleOmer Erdincc YaugmurluFabian OttoRudolf Lioutikov
Introduces FLOWER, an efficient 950-million-parameter vision-language-action policy that matches multi-billion-parameter generalist models across 190 tasks and achieves a new state of the art on CALVIN with only 200 GPU hours of pretraining.
Developing generalist artificial intelligence policies for robotic manipulation is essential for practical automation, yet existing models face severe computational barriers. Leading systems rely on multi-billion-parameter foundation architectures that require tens of thousands of computing hours to train and demand high-end computing memory for real-time deployment. This high cost restricts research access and limits practical deployment on physical hardware.
The article demonstrates an efficient robotic manipulation architecture that matches or surpasses existing state-of-the-art generalist models across diverse robot platforms while containing fewer than one billion parameters. It also evaluates how architectural modifications affect training efficiency, runtime latency, and real-world task generalization.
The authors develop two structural innovations and integrate them into a 947-million-parameter model called FLOWER. First, intermediate-modality fusion extracts contextual features from the middle of a vision-language backbone—pruning between 30% and 50% of its layers—to retain semantic understanding while reallocating capacity to an action-generation module. Second, an action-specific normalization controller shares core modulation weights across layers to reduce parameter overhead by over 20%. The resulting system was pretrained in approximately 200 graphic processing unit (GPU) hours on roughly 250,000 trajectories and evaluated across 190 tasks in ten simulation and real-world benchmarks involving single-arm and dual-arm robots.
The findings show that FLOWER achieves state-of-the-art or highly competitive performance across diverse benchmarks while dramatically reducing computational demands. Pretraining the model required less than 1% of the compute used by leading multi-billion-parameter baselines. In physical kitchen experiments, FLOWER achieved a 61.0% overall task success rate compared to 31.0% for the next-best baseline, and maintained a 51.0% average success rate under severe environmental disruptions such as novel objects, dim lighting, and clutter. Furthermore, the architecture achieved an inference throughput of 311 Hz and required only 1.85 gigabytes of memory, operating over 50 times faster than comparable models while using a fraction of the hardware memory.
These results demonstrate that robotics developers do not need massive multi-billion-parameter models to achieve versatile control. By balancing semantic reasoning and action generation, organizations can significantly lower training expenses, accelerate development cycles, and deploy capable generalist policies directly on cost-effective, commodity hardware without sacrificing task reliability.
Organizations developing robotic systems should adopt intermediate fusion and parameter-efficient normalization techniques to optimize edge deployment. Before widespread physical rollout, teams should conduct pilot studies in target operating environments, particularly for fine-precision manipulation tasks and zero-shot transfers.
While the empirical results are strong, eight of the ten evaluated benchmark suites were conducted in simulated environments, and physical testing was confined to stationary manipulation rather than mobile navigation or locomotion. In addition, the policy relies on iterative trajectory sampling rather than single-pass generation and demonstrated lower zero-shot success on one complex simulation benchmark, indicating that further testing is necessary before deploying the system in critical low-tolerance environments.
- Paper: π0: A Vision-Language-Action Flow Model for General Robot Control, Kevin Black et al. (2024). FLOWER directly builds upon and aims to make more efficient the flow-matching VLA paradigm introduced by $\pi_0$ for continuous robot control.
- Paper: Octo: An Open-Source Generalist Robot Policy, O. Team et al. (2024). Octo establishes open-source transformer policies with diffusion action decoding across diverse embodiments, serving as a foundational reference point for FLOWER's efficient design.
- Paper: Diffusion policy: Visuomotor policy learning via action diffusion, Cheng Chi et al. (2023). Diffusion Policy introduces visuomotor control via conditional generative denoising, which provides the underlying formulation adapted and accelerated by FLOWER.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). OpenVLA provides a foundational open-source vision-language-action architecture against which FLOWER optimizes parameter efficiency and training compute.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). RT-2 pioneered co-fine-tuning vision-language backbones for direct robotic control, establishing the VLA paradigm that FLOWER streamlines.
- Paper: π0.5: a Vision-Language-Action Model with Open-World Generalization, Physical Intelligence et al. (2025). $\pi_{0.5}$ scales flow-matching VLA models to open-world mobile manipulation across diverse homes, extending the flow-based control principles optimized in FLOWER.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). DynamicVLA extends lightweight, real-time VLA execution concepts to fast-moving dynamic object manipulation.
- Paper: What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models, Yu Peng et al. (2026). PAIR-VLA presents an RL fine-tuning approach for visual robustness in VLA models, continuing the pursuit of robust deployment for generalist policies.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). CoA-VLA enhances VLA policies with structured visual-textual chain-of-affordance reasoning, offering an advanced reasoning layer atop action diffusion policies.
- Paper: TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers, Bin Yu et al. (2026). TwinBrainVLA explores dual-pathway architectures with flow matching to prevent catastrophic forgetting when adapting generalist VLMs to robotics.
- Paper: RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation, Songming Liu et al. (2025). RDT-1B investigates scaling diffusion foundation models specifically for bimanual manipulation with a unified action space.
- Paper: GR00T N1: An Open Foundation Model for Generalist Humanoid Robots, NVIDIA et al. (2025). GR00T N1 applies dual-system flow-matching architectures to generalist humanoid robotics across large-scale physical embodiments.
