π0.5: a Vision-Language-Action Model with Open-World Generalization

Physical IntelligenceKevin BlackNoah BrownJames DarpinianKaran DhabaliaDanny DriessAdnan EsmailMichael EquiChelsea FinnNiccolo Fusai

article2025arXiv1,891 citations

Presents π0.5\pi_{0.5}, an end-to-end vision-language-action model that leverages heterogeneous co-training across multi-robot and web data to perform dexterous, long-horizon manipulation tasks in previously unseen real-world homes.

Listen

Deploying autonomous robots to perform complex household tasks in novel real-world environments remains a fundamental hurdle in robotics. Traditional vision-language-action models often fail when encountering unseen layouts, unfamiliar objects, or long-horizon tasks, largely because scaling real-world demonstration collection across every possible domestic environment is prohibitively expensive and logistically impractical.

The article demonstrates and evaluates a vision-language-action model, named π0.5, designed to achieve open-world generalization. The central objective is to evaluate whether co-training a unified model on heterogeneous data sources allows mobile manipulators to perform intricate, multi-stage domestic tasks out-of-the-box in entirely new homes.

To address this challenge, the authors implemented a two-stage hierarchical training approach. The model combines a core vision-language transformer backbone with a specialized continuous action expert. During the initial pre-training phase, the architecture learns over discrete tokens representing a broad mixture of data, where 97.6% of the examples come from non-target sources, including web data (visual question answering, image captioning, and object detection), multi-robot laboratory data, and static arm demonstrations across diverse environments. In the post-training phase, the model is specialized using approximately 400 hours of mobile manipulator demonstrations across roughly 100 homes, combined with verbal supervisory instructions and flow matching for real-time continuous control. The authors evaluated the system through extensive multi-trial experiments across 12 testing sites, including six mock rooms and six rooms in three completely unseen real homes, measuring performance on extended tasks lasting up to 10 to 15 minutes.

The findings show that π0.5 achieves high task success in previously unseen environments, matching the performance of baseline models trained directly within the test homes. Transfer across heterogeneous robotic embodiments proved critical; omitting non-mobile or cross-embodiment laboratory data significantly degraded manipulation success. Incorporating multimodal web data substantially enhanced language following and physical reasoning, especially for novel, out-of-distribution object categories. Furthermore, explicit high-level subtask prediction outperformed flat action architectures and external planning models such as zero-shot GPT-4, and language following accuracy improved consistently as the number of distinct training environments scaled up to 104 locations.

These results indicate that robotic foundation models do not require massive fleets of target hardware collecting data in every deployment setting. Instead, broad generalization can be effectively achieved by transferring knowledge from cheaper, heterogeneous data streams, such as static robot arms, web data, and human verbal instructions. This design reduces the cost, time, and hardware wear associated with scaling physical data collection while significantly improving the robustness of general-purpose robots.

Organizations developing or deploying robotic manipulation systems should prioritize heterogeneous co-training frameworks over narrow, single-platform data collection pipelines. Teams should implement structured, two-level inference pipelines that link high-level semantic reasoning to fast continuous action generation, while also capturing verbal supervisor corrections during operation. Future technical roadmaps should focus on expanding model context to handle partial observability—such as room-to-room navigation—and refining grasp execution on complex physical fixtures like non-standard drawer handles.

While the reported results demonstrate high confidence across structured domestic tasks, users should maintain caution. The model still experiences occasional failures due to visual occlusions and repetitive subtask loops, and it relies on relatively short textual instructions rather than complex, highly nuanced user prompts.

Cover for π0.5: a Vision-Language-Action Model with Open-World Generalization

Abstract

In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated impressive results for end-to-end robot control, it remains an open question how far such models can generalize in the wild. We describe π0.5\pi_{0.5}, a new model based on π0\pi_{0} that uses co-training on heterogeneous tasks to enable broad generalization. π0.5\pi_{0.5}\ uses data from multiple robots, high-level semantic prediction, web data, and other sources to enable broadly generalizable real-world robotic manipulation. Our system uses a combination of co-training and hybrid multi-modal examples that combine image observations, language commands, object detections, semantic subtask prediction, and low-level actions. Our experiments show that this kind of knowledge transfer is essential for effective generalization, and we demonstrate for the first time that an end-to-end learning-enabled robotic system can perform long-horizon and dexterous manipulation skills, such as cleaning a kitchen or bedroom, in entirely new homes.

Table of Contents

  • I Introduction
  • II Related Work
  • III Preliminaries
  • IV The π0.5\pi_{0.5} Model and Training Recipe
  • IV-A The π0.5\pi_{0.5} architecture
  • IV-B Combining discrete & continuous action representations
  • IV-C Pre-training
  • IV-D Post-training
  • IV-E Robot system details
  • V Experimental Evaluation
  • V-A Can π0.5\pi_{0.5} generalize to real homes?
  • V-B How does generalization scale with the number of scenes?
  • V-C How important is each part of our co-training recipe?
  • V-D How does π0.5\pi_{0.5} compare to other VLAs?
  • V-E How important is high-level inference?
  • VI Discussion and Future Work
  • References
  • A
  • A-A Contributions
  • A-B Task evaluation rubric
  • A-C Language following experiment setup
  • A-D Per-task performance breakdown
  • A-E Model technical details

Knowls

  1. Knowl 1 — pi0.5 Model Architecture and Attention Topology

    model/method

    The π0.5\pi_{0.5} architecture is a vision-language-action model that combines a 2-billion parameter vision-language model (VLM) backbone with a 300-million parameter continuous action expert transformer.

    1. VLM Backbone: Initialized from pre-trained PaliGemma weights, with width 20482048, depth 1818, MLP dimension 16,38416{,}384, 1818 attention heads, 11 key-value head, and head dimension 256256. It processes a sequence of NN multimodal tokens, including image patches xiI∈Rp×p×3x_i^I \in \mathbb{R}^{p \times p \times 3}, textual task prompts xiw∈Nx_i^w \in \mathbb{N}, tokenized robot proprioceptive state qtq_t, and discrete action tokens encoded using the FAST tokenizer.

    2. Action Expert Transformer: A transformer with width 10241024, depth 1818, MLP dimension 40964096, 1818 attention heads, 11 key-value head, and head dimension 256256. It takes continuous noisy action chunks at:t+Hτ,ωa_{t:t+H}^{\tau, \omega} across an action chunk horizon of H=50H = 50 (where H−1=49H-1 = 49 future steps are predicted) with dimension dd, linearly projected to the 10241024-dimensional embedding space, and outputs continuous action tokens y1:Hay_{1:H}^a that are decoded via a linear projection to predict the flow-matching vector field.

    3. Attention Topology: Embeddings from the VLM and action expert interact via self-attention governed by a structured mask:

    • A bidirectional prefix mask covers all image tokens, prompt text tokens, and proprioceptive state tokens.
    • FAST discrete action tokens attend bidirectionally to the full observation prefix and autoregressively to preceding FAST action tokens.
    • Action expert tokens attend to the observation prefix and bidirectionally to one another, but are masked from attending to FAST action tokens to prevent information leakage.
    • VLM tokens never attend to the action expert tokens, establishing a strictly unidirectional flow of information from the VLM backbone to the action expert.
  2. Knowl 2 — Joint Discrete and Continuous Flow Matching Objective

    equation

    The π0.5\pi_{0.5} model is optimized using a hybrid loss function that unifies autoregressive cross-entropy loss over discrete tokens with a continuous flow-matching regression loss:

    ED,τ,ω[H(x1:M,fθℓ(ot,ℓ))+α∥ω−at:t+H−fθa(at:t+Hτ,ω,ot,ℓ)∥2]\mathbb{E}_{\mathcal{D}, \tau, \omega} \left[ \mathcal{H}\left(x_{1:M}, f_\theta^\ell(o_t, \ell)\right) + \alpha \left\| \omega - a_{t:t+H} - f_\theta^a\left(a_{t:t+H}^{\tau, \omega}, o_t, \ell\right) \right\|^2 \right]

    where:

    • D\mathcal{D} denotes the training dataset.
    • ot=[It1,…,Itn,qt]o_t = [I_t^1, \dots, I_t^n, q_t] represents the multi-camera image observations and the proprioceptive robot configuration at time tt.
    • ℓ\ell is the natural language instruction or prompt.
    • x1:Mx_{1:M} is the target sequence of MM discrete tokens, encompassing text outputs, bounding boxes, and FAST-encoded discrete action chunks.
    • fθℓ(ot,ℓ)f_\theta^\ell(o_t, \ell) represents the predicted text/discrete token logits from the vision-language backbone.
    • H(⋅,⋅)\mathcal{H}(\cdot, \cdot) is the standard cross-entropy loss.
    • at:t+H∈RH×da_{t:t+H} \in \mathbb{R}^{H \times d} is the ground-truth continuous action chunk over horizon H=50H=50 with action dimension dd.
    • ω∼N(0,I)\omega \sim \mathcal{N}(0, I) is Gaussian noise sampled in action space.
    • τ∈[0,1]\tau \in [0, 1] is the flow matching time index.
    • at:t+Hτ,ω=τat:t+H+(1−τ)ωa_{t:t+H}^{\tau, \omega} = \tau a_{t:t+H} + (1 - \tau)\omega is the interpolated noisy action chunk.
    • ω−at:t+H\omega - a_{t:t+H} is the target velocity vector field.
    • fθa(at:t+Hτ,ω,ot,ℓ)f_\theta^a(a_{t:t+H}^{\tau, \omega}, o_t, \ell) is the vector field predicted by the action expert.
    • α≥0\alpha \ge 0 is a scalar trade-off parameter, set to α=0\alpha = 0 during pre-training and α=10.0\alpha = 10.0 during post-training.
  3. Knowl 3 — Hierarchical High-Level Subtask and Low-Level Action Inference

    model/method

    The π0.5\pi_{0.5} model decomposes policy control into a two-level hierarchical inference process executed by a single, unified neural network:

    πθ(at:t+H,ℓ^∣ot,ℓ)=πθ(at:t+H∣ot,ℓ^) πθ(ℓ^∣ot,ℓ)\pi_\theta(a_{t:t+H}, \hat{\ell} \mid o_t, \ell) = \pi_\theta(a_{t:t+H} \mid o_t, \hat{\ell}) \, \pi_\theta(\hat{\ell} \mid o_t, \ell)

    1. High-Level Semantic Inference (πθ(ℓ^∣ot,ℓ)\pi_\theta(\hat{\ell} \mid o_t, \ell)): Given an abstract language command ℓ\ell (such as "clean the kitchen") and the full multi-camera observation oto_t (utilizing forward, backward, and wrist cameras), the model autoregressively generates 2D bounding boxes of relevant objects followed by a textual subtask instruction ℓ^\hat{\ell} (such as "pick up the cutting board"). High-level inference operates at a low temporal frequency.

    2. Low-Level Motor Inference (πθ(at:t+H∣ot,ℓ^)\pi_\theta(a_{t:t+H} \mid o_t, \hat{\ell})): Conditioned on the predicted subtask ℓ^\hat{\ell} and local observations (forward and wrist cameras), the model generates continuous action chunks at:t+Ha_{t:t+H} via 10 iterative flow-matching denoising steps using the action expert. The low-level action distribution depends strictly on the predicted subtask ℓ^\hat{\ell} rather than the global prompt ℓ\ell.

  4. Knowl 4 — Flow Matching Timestep Conditioning and Sampling Distribution

    model/method

    In π0.5\pi_{0.5}, continuous action generation is parameterized by flow matching with custom timestep conditioning and sampling:

    1. Timestep Conditioning: The flow-matching timestep index τ∈[0,1]\tau \in [0, 1] is mapped to a sinusoidal positional encoding ϕ(τ)∈Rw\phi(\tau) \in \mathbb{R}^w and processed by a dedicated multilayer perceptron: swish(W2⋅swish(W1⋅ϕ(τ)))\text{swish}(W_2 \cdot \text{swish}(W_1 \cdot \phi(\tau))), with weight matrices W1,W2∈Rw×wW_1, W_2 \in \mathbb{R}^{w \times w}. This representation modulates each action expert layer via adaptive RMSNorm, decoupling the temporal conditioning from the noisy action chunk projection.

    2. Low-Timestep Sampling: Instead of uniform sampling τ∼U(0,1)\tau \sim \mathcal{U}(0, 1), timesteps during training are drawn from a Beta distribution biased towards lower noise levels:

    p(τ)=Beta(s−τs;α=1.5,β=1)p(\tau) = \text{Beta}\left( \frac{s - \tau}{s}; \alpha = 1.5, \beta = 1 \right)

    where s=0.999s = 0.999 is a cutoff threshold. Values τ>s\tau > s are discarded, which guarantees compatibility with numerical integration step sizes δ>1−s=0.001\delta > 1 - s = 0.001. At test time, continuous action trajectories are integrated in 10 Euler denoising steps.

  5. Knowl 5 — Two-Stage Co-Training Recipe across Heterogeneous Data Sources

    model/method

    The π0.5\pi_{0.5} training procedure consists of two sequential phases combining diverse robotic and non-robotic data sources:

    1. Pre-Training Phase (280k gradient steps, α=0\alpha = 0): The full model is trained purely via next-token cross-entropy prediction using FAST discrete action tokens across five data streams:
    • Diverse Mobile Manipulator data (MM): ~400 hours of dual-arm mobile manipulation across ~100 distinct home environments.
    • Multi-Environment non-mobile robot data (ME): Single- and dual-arm static manipulators deployed across diverse residential kitchens and rooms.
    • Laboratory Cross-Embodiment data (CE): Tabletop single- and dual-arm robotic tasks collected under controlled laboratory setups, including an extended version of the Open X-Embodiment (OXE) dataset.
    • High-Level Subtask prediction data (HL): Demonstrations from MM, ME, and CE annotated with 2D bounding boxes and natural language subtask descriptions.
    • Multimodal Web Data (WD): Image captioning (CapsFusion, COCO), visual question answering (Cambrian-7M, PixMo, VQAv2), and indoor object bounding box localization datasets.
    1. Post-Training Phase (80k gradient steps, α=10.0\alpha = 10.0): The continuous action expert is initialized from scratch and jointly trained with the VLM backbone. The training mix excludes laboratory CE data to specialize the model on home environments, retaining MM, ME, WD, multi-environment HL data, and adding:
    • Verbal Instructions (VI): Real-time language teleoperation demonstrations where human supervisors interactively guide the trained low-level policy with sequential language subtask commands (accounting for ~11% of high-level mobile manipulation examples).
  6. Knowl 6 — Mobile Manipulator Hardware Architecture and Control Loop

    experimental setup

    The robotic evaluation and data collection systems in π0.5\pi_{0.5} utilize two physical mobile manipulator configurations:

    1. Kinematic Platform: Each robot features two 6 DoF robotic arms equipped with 1 DoF parallel-jaw grippers, a 3 DoF holonomic wheeled base (2D linear velocity and 1D angular velocity), and an actuated torso lift mechanism (1D vertical motion or 2D vertical and horizontal motion), yielding an 18 or 19 DoF continuous state and action space.

    2. Perception Setup: Each robot carries four monocular RGB cameras: two wrist-mounted cameras, one front-facing torso camera, and one rear-facing camera. High-level subtask inference consumes all four camera feeds, whereas low-level action generation processes only the two wrist cameras and the front-facing camera.

    3. Control Execution: Policy outputs are normalized to [−1,1][-1, 1] based on the 1st and 99th percentiles of demonstration data and padded with zeros to a uniform maximum dimensionality. The policy directly outputs target joint configurations, gripper state, torso lift positions, and base velocities at 50 Hz in chunks of 50 steps (H=50H=50), which are tracked via low-level proportional-derivative (PD) controllers without intermediate motion planners or collision checkers.

  7. Knowl 7 — Zero-Shot Multi-Stage Manipulation in Unseen Real Homes

    empirical result

    When evaluated in three real residential homes that were completely absent from training data, π0.5\pi_{0.5} executed autonomous long-horizon manipulation routines lasting 10 to 15 minutes given only high-level natural language instructions.

    Task performance was measured across four standardized multi-stage cleanup tasks (10 trials per task per home):

    • Dishes in Sink (max score 8: 4 dishes, +1 for grasp, +1 for placement into sink): achieved 80% to 100% average progress in real kitchens.
    • Items in Drawer (max score 4: grasp object, open drawer, insert object, close drawer): achieved 65% to 85% average progress.
    • Laundry in Basket (max score 3: navigate and pick up garment, place on basket, insert fully): achieved 85% to 100% average progress in real bedrooms.
    • Make the Bed (max score 5: straighten blanket, place first pillow, place second pillow, blanket neatness bonus, pillow neatness bonus): achieved 70% to 85% average progress.

    Performance in mock evaluation environments was statistically consistent with zero-shot performance in real homes.

  8. Knowl 8 — Generalization Scaling with Training Environment Diversity

    empirical result

    Evaluating π0.5\pi_{0.5} post-trained on mobile manipulation datasets containing 3, 12, 22, 53, 82, and 104 distinct home environments (fixed at 40k gradient steps to control for sample count) demonstrated monotonic performance scaling with scene diversity:

    1. Multi-Stage Task Progress: Average progress across four mock home tasks (Dishes in Sink, Items in Drawer, Laundry in Basket, Make Bed) improved from ~40% with 3 training environments to ~80% with 104 environments. A model trained on 104 unseen environments matched the ~82% success rate of an in-domain baseline trained directly in the target test environments.

    2. Language Following & Placement Success: In pick-and-place tests with 5 objects containing distractors:

    • For in-distribution object classes (new instances of seen categories), language following accuracy scaled from ~42% (3 locations) to ~85% (104 locations), with placement success rising from ~20% to ~60%.
    • For out-of-distribution object classes (unseen categories: funnel, pill bottle, grill lighter, lighter, safety goggles), language following improved from ~20% (chance level) to ~62%, with placement success rising from ~10% to ~52%.
  9. Knowl 9 — Co-Training Recipe Ablation on Cross-Embodiment and Web Data

    empirical result

    Ablation experiments on four mock home cleanup tasks (10 trials per policy per task) isolated the contribution of each co-training data source in π0.5\pi_{0.5}:

    1. Excluding Multi-Environment Static Data (no ME): Dropped overall task progress from ~80% to ~54% (p<0.001p < 0.001), with severe degradation on Items in Drawer (~40% progress).

    2. Excluding Laboratory Cross-Embodiment Data (no CE): Reduced task progress from ~80% to ~52% (p<0.001p < 0.001), severely affecting Dishes in Sink (~45% progress).

    3. Excluding Both ME and CE (no ME or CE): Training exclusively on target mobile manipulator data and web data degraded overall task progress to ~40% (p<0.001p < 0.001).

    4. Excluding Multimodal Web Data (no WD): Had negligible effect on overall progress for familiar household tasks (~78% vs. ~80%, p=0.385p=0.385), but caused a catastrophic drop in out-of-distribution language following success (dropping OOD object placement success from ~75% to ~30%).

  10. Knowl 10 — Hierarchical Subtask Inference Ablation and Comparison to Language Planners

    empirical result

    Systematic evaluation of high-level subtask inference architectures across mock home tasks revealed the following performance hierarchy:

    1. Full Hierarchical π0.5\pi_{0.5}: Achieved the highest performance (~80% task progress), outperforming an oracle human supervisor baseline providing ground-truth subtask instructions in real time (~68%, p=0.016p = 0.016).

    2. Implicit High-Level Inference (implicit HL): A baseline that performs single-stage flat control at test time but retains high-level subtask prediction in its training loss achieved ~72% progress (p=0.144p = 0.144), outperforming a model trained completely without subtask data (no HL: ~62%, p=0.011p = 0.011).

    3. Verbal Instructions Ablation (no VI): Removing human language demonstrations during post-training dropped overall progress to ~60% (p=0.009p = 0.009).

    4. Pre-trained LLM Planner (GPT-4 HL): Zero-shot prompting of GPT-4 to output discrete subtask steps for the low-level policy achieved the lowest performance (~58%, p=0.002p = 0.002), highlighting the necessity of domain-specific policy fine-tuning for embodied subtask sequencing.

  11. Knowl 11 — Performance Comparison of pi0.5 against pi0 Baselines

    empirical result

    Across standardized mock home environments, π0.5\pi_{0.5} was benchmarked against the original π0\pi_0 model and an enhanced baseline, π0-FAST+Flow\pi_0\text{-FAST+Flow} (which uses the same hybrid discrete/continuous architecture but trains solely on robot action trajectories without web data or subtask prediction):

    1. Mock Home Cleanup Tasks: π0.5\pi_{0.5} achieved ~80% average progress, significantly outperforming π0-FAST+Flow\pi_0\text{-FAST+Flow} (~60-70%), standard π0\pi_0 trained for 80k steps (~30-40%), and standard π0\pi_0 trained for 300k steps (~40-50%).

    2. Language Following: π0.5\pi_{0.5} achieved an ~85% language following rate and ~70% placement success rate, outperforming π0-FAST+Flow\pi_0\text{-FAST+Flow} (~75% follow, ~65% success) and outperforming standard π0\pi_0 (~18% follow, ~15% success) by a wide margin. This confirms that pre-training on discrete action tokens is substantially more sample- and compute-efficient for language grounding than pure diffusion/flow matching training.

  12. Knowl 12 — Physical, Observational, and High-Level Failure Modes of pi0.5

    limitation

    The authors reported several failure modes and limitations of π0.5\pi_{0.5} during open-world testing:

    1. Hardware Affordance Limits: The system struggles with physical mechanisms requiring high force or specialized mechanics not well-represented in the training data, such as stiff cabinet latches or unusual drawer handles.

    2. Partial Observability: The policy is susceptible to self-occlusions, where the robot's own arm or gripper obstructs visual feedback of the target workspace (e.g., blocking the camera view of a countertop liquid spill while wiping).

    3. High-Level Policy Loops: The autoregressive subtask predictor occasionally gets trapped in cyclic or distracted states, such as repeatedly opening and closing a cabinet or drawer instead of placing an object inside.

    4. Context and Language Constraints: The model operates over a short temporal context without persistent spatial memory, preventing multi-room navigation and long-term object search. Furthermore, language conditioning is constrained to simple prompt structures matching the demonstration distribution.

Coverage note — None was omitted; all primary architectural components, loss formulations, training pipelines, robotic hardware parameters, empirical evaluations across real/mock homes and ablations, and stated limitations are fully covered.

References

  1. 1.AgiBot-World-Contributors, Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xuan Hu, Xu Huang, Shu Jiang, Yuxin Jiang, Cheng Jing, Hongyang Li, Jialu Li, Chiming Liu, Yi Liu, Yuxiang Lu, Jianlan Luo, Ping Luo, Yao Mu, Yuehan Niu, Yixuan Pan, Jiangmiao Pang, Yu Qiao, Guanghui Ren, Cheng Ruan, Jiaqi Shan, Yongjian Shen, Chengshi Shi, Mingkang Shi, Modi Shi, Chonghao Sima, Jianheng Song, Huijie Wang, Wenhao Wang, Dafeng Wei, Chengen Xie, Guo Xu, Junchi Yan, Cunbiao Yang, Lei Yang, Shukai Yang, Maoqing Yao, Jia Zeng, Chi Zhang, Qinglin Zhang, Bin Zhao, Chengyue Zhao, Jiaqi Zhao, and Jianchao Zhu. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. arXiv preprint arXiv:2503.06669, 2025.
  2. 2.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, and Andy Zeng. Do as i can and not as i say: Grounding language in robotic affordances. In arXiv preprint arXiv:2204.01691, 2022.
  3. 3.Suneel Belkhale and Dorsa Sadigh. Minivla: A better vla with a smaller footprint, 2024. URL https://github.com/Stanford-ILIAD/openvla-mini.
  4. 4.Suneel Belkhale, Tianli Ding, Ted Xiao, Pierre Sermanet, Quon Vuong, Jonathan Tompson, Yevgen Chebotar, Debidatta Dwibedi, and Dorsa Sadigh. Rt-h: Action hierarchies using language, 2024. URL https://arxiv.org/abs/2403.01823.
  5. 5.Lucas Beyer, Andreas Steiner, Andre Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, et al. Paligemma: A versatile 3b vlm for transfer. arXiv preprint arXiv:2407.07726, 2024.
  6. 6.Homanga Bharadhwaj, Jay Vakil, Mohit Sharma, Abhinav Gupta, Shubham Tulsiani, and Vikash Kumar. Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 4788–4795. IEEE, 2024.
  7. 7.Johan Bjorck, Fernando Castaneda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734, 2025.
  8. 8.Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky. π0\pi_0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024.
  9. 9.Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Kuang-Huei Lee, Sergey Levine, Yao Lu, Utsav Malla, Deeksha Manjunath, Igor Mordatch, Ofir Nachum, Carolina Parada, Jodilyn Peralta, Emily Perez, Karl Pertsch, Jornell Quiambao, Kanishka Rao, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Kevin Sayed, Jaspiar Singh, Sumedh Sontakke, Austin Stone, Clayton Tan, Huong Tran, Vincent Vanhoucke, Steve Vega, Quan Vuong, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich. Rt-1: Robotics transformer for real-world control at scale. In arXiv preprint arXiv:2212.06817, 2022.
  10. 10.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Advances in Neural Information Processing Systems, 2020.
  11. 11.Hongyi Chen, Yunchao Yao, Ruixuan Liu, Changliu Liu, and Jeffrey Ichnowski. Automating robot failure recovery using vision-language models with optimized prompts. arXiv preprint arXiv:2409.03966, 2024.
  12. 12.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015.
  13. 13.An-Chieh Cheng, Yandong Ji, Zhaojing Yang, Zaitian Gongye, Xueyan Zou, Jan Kautz, Erdem Bıyık, Hongxu Yin, Sifei Liu, and Xiaolong Wang. Navila: Legged robot vision-language-action model for navigation. arXiv preprint arXiv:2412.04453, 2024.
  14. 14.Cheng Chi, Zhenjia Xu, Chuer Pan, Eric Cousineau, Benjamin Burchfiel, Siyuan Feng, Russ Tedrake, and Shuran Song. Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots. In Proceedings of Robotics: Science and Systems (RSS), 2024.
  15. 15.OX-Embodiment Collaboration, A Padalkar, A Pooley, A Jain, A Bewley, A Herzog, A Irpan, A Khazatsky, A Rai, A Singh, et al. Open X-Embodiment: Robotic learning datasets and RT-X models. arXiv preprint arXiv:2310.08864, 1(2), 2023.
  16. 16.Yinpei Dai, Jayjun Lee, Nima Fazeli, and Joyce Chai. Racer: Rich language-guided failure recovery policies for imitation learning. International Conference on Robotics and Automation (ICRA), 2025.
  17. 17.Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn. Robonet: Large-scale multi-robot learning. CoRL, 2019.
  18. 18.Sudeep Dasari, Mohan Kumar Srirama, Unnat Jain, and Abhinav Gupta. An unbiased look at datasets for visuo-motor pre-training. In Conference on Robot Learning, pages 1183–1198. PMLR, 2023.
  19. 19.Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, et al. Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models. arXiv preprint arXiv:2409.17146, 2024.
  20. 20.Dempsey. Reviews-consumer technology. the teardown-amazon astro consumer robot. Engineering & Technology, 18(2):70–71, 2023.
  21. 21.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2019.
  22. 22.Ria Doshi, Homer Walke, Oier Mees, Sudeep Dasari, and Sergey Levine. Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation. In Conference on Robot Learning, 2024.
  23. 23.Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023.
  24. 24.Jiafei Duan, Wentao Yuan, Wilbert Pumacay, Yi Ru Wang, Kiana Ehsani, Dieter Fox, and Ranjay Krishna. Manipulate-anything: Automating real-world robots using vision-language models. arXiv preprint arXiv:2406.18915, 2024.
  25. 25.Frederik Ebert, Yanlai Yang, Karl Schmeckpeper, Bernadette Bucher, Georgios Georgakis, Kostas Daniilidis, Chelsea Finn, and Sergey Levine. Bridge data: Boosting generalization of robotic skills with cross-domain datasets. arXiv preprint arXiv:2109.13396, 2021.
  26. 26.Kiana Ehsani, Tanmay Gupta, Rose Hendrix, Jordi Salvador, Luca Weihs, Kuo-Hao Zeng, Kunal Pratap Singh, Yejin Kim, Winson Han, Alvaro Herrasti, et al. Spoc: Imitating shortest paths in simulation enables effective navigation and manipulation in the real world. arXiv preprint arXiv:2312.02976, 2023.
  27. 27.Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first International Conference on Machine Learning, 2024.
  28. 28.Haritheja Etukuru, Norihito Naka, Zijin Hu, Seungjae Lee, Julian Mehu, Aaron Edsinger, Chris Paxton, Soumith Chintala, Lerrel Pinto, and Nur Muhammad Mahi Shafiullah. Robot utility models: General policies for zero-shot deployment in new environments. arXiv preprint arXiv:2409.05865, 2024.
  29. 29.Hao-Shu Fang, Chenxi Wang, Hongjie Fang, Minghao Gou, Jirong Liu, Hengxu Yan, Wenhai Liu, Yichen Xie, and Cewu Lu. Anygrasp: Robust and efficient grasp perception in spatial and temporal domains. IEEE Transactions on Robotics, 39(5):3929–3945, 2023.
  30. 30.Hao-Shu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Chenxi Wang, Junbo Wang, Haoyi Zhu, and Cewu Lu. Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 653–660. IEEE, 2024.
  31. 31.Theophile Gervet, Soumith Chintala, Dhruv Batra, Jitendra Malik, and Devendra Singh Chaplot. Navigating to objects in the real world. Science Robotics, 8(79):eadf6991, 2023.
  32. 32.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In Computer Vision and Pattern Recognition (CVPR), 2017.
  33. 33.Abhinav Gupta, Adithyavairavan Murali, Dhiraj Prakashchand Gandhi, and Lerrel Pinto. Robot learning in homes: Improving generalization and reducing dataset bias. Advances in neural information processing systems, 31, 2018.
  34. 34.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  35. 35.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15979–15988, 2022.
  36. 36.Yingdong Hu, Fanqi Lin, Tong Zhang, Li Yi, and Yang Gao. Look before you leap: Unveiling the power of gpt-4v in robotic vision-language planning. arXiv preprint arXiv:2311.17842, 2023.
  37. 37.Huang Huang, Fangchen Liu, Letian Fu, Tingfan Wu, Mustafa Mukadam, Jitendra Malik, Ken Goldberg, and Pieter Abbeel. Otter: A vision-language-action model with text-aware visual feature extraction. arXiv preprint arXiv:2503.03734, 2025.
  38. 38.Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International conference on machine learning, pages 9118–9147. PMLR, 2022.
  39. 39.Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024.
  40. 40.Joseph L Jones. Robots at the tipping point: the road to irobot roomba. IEEE Robotics & Automation Magazine, 13(1):76–78, 2006.
  41. 41.Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, Peter David Fagan, Joey Hejna, Masha Itkina, Marion Lepert, Yecheng Jason Ma, Patrick Tree Miller, Jimmy Wu, Suneel Belkhale, Shivin Dass, Huy Ha, Arhan Jain, Abraham Lee, Youngwoon Lee, Marius Memmel, Sungjae Park, Ilija Radosavovic, Kaiyuan Wang, Albert Zhan, Kevin Black, Cheng Chi, Kyle Beltran Hatch, Shan Lin, Jingpei Lu, Jean Mercat, Abdul Rehman, Pannag R Sanketi, Archit Sharma, Cody Simpson, Quan Vuong, Homer Rich Walke, Blake Wulfe, Ted Xiao, Jonathan Heewon Yang, Arefeh Yavary, Tony Z. Zhao, Christopher Agia, Rohan Baijal, Mateo Guaman Castro, Daphne Chen, Qiuyu Chen, Trinity Chung, Jaimyn Drake, Ethan Paul Foster, Jensen Gao, David Antonio Herrera, Minho Heo, Kyle Hsu, Jiaheng Hu, Donovon Jackson, Charlotte Le, Yunshuang Li, Kevin Lin, Roy Lin, Zehan Ma, Abhiram Maddukuri, Suvir Mirchandani, Daniel Morton, Tony Nguyen, Abigail O’Neill, Rosario Scalise, Derick Seale, Victor Son, Stephen Tian, Emi Tran, Andrew E. Wang, Yilin Wu, Annie Xie, Jingyun Yang, Patrick Yin, Yunchu Zhang, Osbert Bastani, Glen Berseth, Jeannette Bohg, Ken Goldberg, Abhinav Gupta, Abhishek Gupta, Dinesh Jayaraman, Joseph J Lim, Jitendra Malik, Roberto Martín-Martín, Subramanian Ramamoorthy, Dorsa Sadigh, Shuran Song, Jiajun Wu, Michael C. Yip, Yuke Zhu, Thomas Kollar, Sergey Levine, and Chelsea Finn. Droid: A large-scale in-the-wild robot manipulation dataset. In Proceedings of Robotics: Science and Systems, 2024.
  42. 42.Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024.
  43. 43.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment anything. arXiv preprint arXiv:2304.02643, 2023.
  44. 44.Boyi Li, Philipp Wu, Pieter Abbeel, and Jitendra Malik. Interactive task planning with language models, 2023.
  45. 45.Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizhong Zhang, et al. Cogact: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation. arXiv preprint arXiv:2411.19650, 2024.
  46. 46.Xiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya, Yoo Sung Jang, Jinghuan Shang, Kanchana Ranasinghe, Ryan Burgert, Mu Cai, Yong Jae Lee, et al. Llara: Supercharging robot learning data for vision-language policy. arXiv preprint arXiv:2406.20095, 2024.
  47. 47.Yi Li, Yuquan Deng, Jesse Zhang, Joel Jang, Marius Memmel, Raymond Yu, Caelan Reed Garrett, Fabio Ramos, Dieter Fox, Anqi Li, et al. Hamster: Hierarchical action models for open-world robot manipulation. arXiv preprint arXiv:2502.05485, 2025.
  48. 48.Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 9493–9500. IEEE, 2023.
  49. 49.Fanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen, Jiacheng You, and Yang Gao. Data scaling laws in imitation learning for robotic manipulation. arXiv preprint arXiv:2410.18647, 2024.
  50. 50.Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022.
  51. 51.Fangchen Liu, Kuan Fang, Pieter Abbeel, and Sergey Levine. Moka: Open-vocabulary robotic manipulation through mark-based visual prompting. In First Workshop on Vision-Language Models for Navigation and Manipulation at ICRA 2024, 2024.
  52. 52.Jiaming Liu, Hao Chen, Pengju An, Zhuoyang Liu, Renrui Zhang, Chenyang Gu, Xiaoqi Li, Ziyu Guo, Sixiang Chen, Mengzhen Liu, et al. Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model. arXiv preprint arXiv:2503.10631, 2025.
  53. 53.Peiqi Liu, Yaswanth Orru, Jay Vakil, Chris Paxton, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto. Ok-robot: What really matters in integrating open-knowledge models for robotics. arXiv preprint arXiv:2401.12202, 2024.
  54. 54.Qiang Liu. Rectified flow: A marginal preserving approach to optimal transport. arXiv preprint arXiv:2209.14577, 2022.
  55. 55.Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and Jun Zhu. Rdt-1b: a diffusion foundation model for bimanual manipulation. arXiv preprint arXiv:2410.07864, 2024.
  56. 56.Jeffrey Mahler, Jacky Liang, Sherdil Niyaz, Michael Laskey, Richard Doan, Xinyu Liu, Juan Aparicio Ojea, and Ken Goldberg. Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics. arXiv preprint arXiv:1703.09312, 2017.
  57. 57.Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Jason Ma, Claire Chen, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Tingfan Wu, Jay Vakil, et al. Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems, 36:655–677, 2023.
  58. 58.Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta. R3m: A universal visual representation for robot manipulation. In CoRL, 2022.
  59. 59.Soroush Nasiriany, Fei Xia, Wenhao Yu, Ted Xiao, Jacky Liang, Ishita Dasgupta, Annie Xie, Danny Driess, Ayzaan Wahid, Zhuo Xu, et al. Pivot: Iterative visual prompting elicits actionable knowledge for vlms. arXiv preprint arXiv:2402.07872, 2024.
  60. 60.Hai Nguyen and Charles C Kemp. Autonomously learning to visually detect where manipulation will succeed. Autonomous Robots, 36:137–152, 2014.
  61. 61.Dantong Niu, Yuvan Sharma, Giscard Biamby, Jerome Quenum, Yutong Bai, Baifeng Shi, Trevor Darrell, and Roei Herzig. Llarva: Vision-action instruction tuning enhances robot learning. arXiv preprint arXiv:2406.11815, 2024.
  62. 62.Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Charles Xu, Jianlan Luo, Tobias Kreiman, You Liang Tan, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An open-source generalist robot policy. In Proceedings of Robotics: Science and Systems, Delft, Netherlands, 2024.
  63. 63.Open X-Embodiment Collaboration, Abhishek Padalkar, Acorn Pooley, Ajinkya Jain, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anikait Singh, Anthony Brohan, Antonin Raffin, Ayzaan Wahid, Ben Burgess-Limerick, Beomjoon Kim, Bernhard Schölkopf, Brian Ichter, Cewu Lu, Charles Xu, Chelsea Finn, Chenfeng Xu, Cheng Chi, Chenguang Huang, Christine Chan, Chuer Pan, Chuyuan Fu, Coline Devin, Danny Driess, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Federico Ceola, Fei Xia, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Ge Yan, Giulio Schiavi, Hao Su, Hao-Shu Fang, Haochen Shi, Heni Ben Amor, Henrik I Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jaehyung Kim, Jan Schneider, Jasmine Hsu, Jeannette Bohg, Jeffrey Bingham, Jiajun Wu, Jialin Wu, Jianlan Luo, Jiayuan Gu, Jie Tan, Jihoon Oh, Jitendra Malik, Jonathan Tompson, Jonathan Yang, Joseph J. Lim, Joao Silvério, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Ken Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Zhang, Keyvan Majd, Krishan Rana, Krishnan Srinivasan, Lawrence Yunliang Chen, Lerrel Pinto, Liam Tan, Lionel Ott, Lisa Lee, Masayoshi Tomizuka, Maximilian Du, Michael Ahn, Mingtong Zhang, Mingyu Ding, Mohan Kumar Srirama, Mohit Sharma, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen, Nicolas Heess, Nikhil J Joshi, Niko Suenderhauf, Norman Di Palo, Nur Muhammad Mahi Shafiullah, Oier Mees, Oliver Kroemer, Pannag R Sanketi, Paul Wohlhart, Peng Xu, Pierre Sermanet, Priya Sundaresan, Quan Vuong, Rafael Rafailov, Ran Tian, Ria Doshi, Roberto Martín-Martín, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shuran Song, Sichun Xu, Siddhant Haldar, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Sudeep Dasari, Suneel Belkhale, Takayuki Osa, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Vidhi Jain, Vincent Vanhoucke, Wei Zhan, Wenxuan Zhou, Wolfram Burgard, Xi Chen, Xiaolong Wang, Xinghao Zhu, Xuanlin Li, Yao Lu, Yevgen Chebotar, Yifan Zhou, Yifeng Zhu, Ying Xu, Yixuan Wang, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh hua Wu, Yujin Tang, Yuke Zhu, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zhuo Xu, and Zichen Jeff Cui. Open X-Embodiment: Robotic learning datasets and RT-X models. https://arxiv.org/abs/2310.08864, 2023.
  64. 64.Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. FAST: Efficient action tokenization for vision-language-action models. Robotics: Science and Systems, 2025.
  65. 65.Dicong Qiu, Wenzong Ma, Zhenfu Pan, Hui Xiong, and Junwei Liang. Open-vocabulary mobile manipulation in unseen dynamic environments with 3d semantic maps. arXiv preprint arXiv:2406.18115, 2024.
  66. 66.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  67. 67.Nur Muhammad Mahi Shafiullah, Anant Rai, Haritheja Etukuru, Yiqian Liu, Ishan Misra, Soumith Chintala, and Lerrel Pinto. On bringing robots home. arXiv preprint arXiv:2311.16098, 2023.
  68. 68.Dhruv Shah, Ajay Sridhar, Arjun Bhorkar, Noriaki Hirose, and Sergey Levine. Gnm: A general navigation model to drive any robot. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 7226–7233. IEEE, 2023.
  69. 69.Dhruv Shah, Ajay Sridhar, Nitish Dashora, Kyle Stachowicz, Kevin Black, Noriaki Hirose, and Sergey Levine. ViNT: A foundation model for visual navigation. In 7th Annual Conference on Robot Learning, 2023. URL https://arxiv.org/abs/2306.14846.
  70. 70.Rutav Shah, Albert Yu, Yifeng Zhu, Yuke Zhu, and Roberto Martín-Martín. Bumble: Unifying reasoning and acting with vision-language models for building-wide mobile manipulation. arXiv preprint arXiv:2410.06237, 2024.
  71. 71.Lucy Xiaoyang Shi, Zheyuan Hu, Tony Z Zhao, Archit Sharma, Karl Pertsch, Jianlan Luo, Sergey Levine, and Chelsea Finn. Yell at your robot: Improving on-the-fly from language corrections. arXiv preprint arXiv:2403.12910, 2024.
  72. 72.Lucy Xiaoyang Shi, Brian Ichter, Michael Equi, Liyiming Ke, Karl Pertsch, Quan Vuong, James Tanner, Anna Walling, Haohuan Wang, Niccolo Fusai, et al. Hi robot: Open-ended instruction following with hierarchical vision-language-action models. arXiv preprint arXiv:2502.19417, 2025.
  73. 73.Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. Progprompt: Generating situated robot task plans using large language models. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 11523–11530. IEEE, 2023.
  74. 74.Austin Stone, Ted Xiao, Yao Lu, Keerthana Gopalakrishnan, Kuang-Huei Lee, Quan Vuong, Paul Wohlhart, Brianna Zitkovich, Fei Xia, Chelsea Finn, et al. Open-world object manipulation using pre-trained vision-language models. arXiv preprint arXiv:2303.00905, 2023.
  75. 75.Andrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, Devon Hjelm, Zhe Gan, Zsolt Kira, and Alexander Toshev. From multimodal llms to generalist embodied agents: Methods and lessons. arXiv preprint arXiv:2412.08442, 2024.
  76. 76.Gemini Robotics Team, Saminda Abeyruwan, Joshua Ainslie, Jean-Baptiste Alayrac, Montserrat Gonzalez Arenas, Travis Armstrong, Ashwin Balakrishna, Robert Baruch, Maria Bauza, Michiel Blokzijl, et al. Gemini robotics: Bringing ai into the physical world. arXiv preprint arXiv:2503.20020, 2025.
  77. 77.Peter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Adithya Jairam Vedagiri IYER, Sai Charitha Akula, Shusheng Yang, Jihan Yang, Manoj Middepogu, Ziteng Wang, et al. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. Advances in Neural Information Processing Systems, 37:87310–87356, 2024.
  78. 78.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  79. 79.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, 2017.
  80. 80.Homer Rich Walke, Kevin Black, Tony Z Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen-Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, et al. BridgeData v2: A dataset for robot learning at scale. In Conference on Robot Learning, pages 1723–1736. PMLR, 2023.
  81. 81.Shu Wang, Muzhi Han, Ziyuan Jiao, Zeyu Zhang, Ying Nian Wu, Song-Chun Zhu, and Hangxin Liu. Llmˆ3: Large language model-based task and motion planning with motion failure reasoning. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 12086–12092. IEEE, 2024.
  82. 82.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  83. 83.Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu, Kun Wu, Zhiyuan Xu, Ning Liu, Ran Cheng, Chaomin Shen, Yaxin Peng, Feifei Feng, and Jian Tang. Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation. arXiv preprint arXiv:2409.12514, 2024.
  84. 84.Junjie Wen, Yichen Zhu, Jinming Li, Zhibin Tang, Chaomin Shen, and Feifei Feng. Dexvla: Vision-language model with plug-in diffusion expert for general robot control. arXiv preprint arXiv:2502.05855, 2025.
  85. 85.Tete Xiao, Ilija Radosavovic, Trevor Darrell, and Jitendra Malik. Masked visual pre-training for motor control. arXiv preprint arXiv:2203.06173, 2022.
  86. 86.Jianwei Yang, Reuben Tan, Qianhui Wu, Ruijie Zheng, Baolin Peng, Yongyuan Liang, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, et al. Magma: A foundation model for multimodal ai agents. arXiv preprint arXiv:2502.13130, 2025.
  87. 87.Qiying Yu, Quan Sun, Xiaosong Zhang, Yufeng Cui, Fan Zhang, Yue Cao, Xinlong Wang, and Jingjing Liu. Capsfusion: Rethinking image-text data at scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14022–14032, 2024.
  88. 88.Michał Zawalski, William Chen, Karl Pertsch, Oier Mees, Chelsea Finn, and Sergey Levine. Robotic control via embodied chain-of-thought reasoning. In Conference on Robot Learning, 2024.
  89. 89.Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Qianli Ma, Song Han, Chelsea Finn, et al. Cot-vla: Visual chain-of-thought reasoning for vision-language-action models. Computer Vision and Pattern Recognition (CVPR), 2025.
  90. 90.Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan. 3d-vla: 3d vision-language-action generative world model. arXiv preprint arXiv:2403.09631, 2024.
  91. 91.Peiyuan Zhi, Zhiyuan Zhang, Yu Zhao, Muzhi Han, Zeyu Zhang, Zhitian Li, Ziyuan Jiao, Baoxiong Jia, and Siyuan Huang. Closed-loop open-vocabulary mobile manipulation with gpt-4v. arXiv preprint arXiv:2404.10220, 2024.
  92. 92.Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In Conference on Robot Learning, pages 2165–2183. PMLR, 2023.

Citation

MLA
Intelligence, P., et al. “$π_{0.5}$: A Vision-Language-Action Model with Open-World Generalization”. arXiv, 2025, http://arxiv.org/abs/2504.16054v1.
APA
Intelligence, P., Black, K., Brown, N., Darpinian, J., Dhabalia, K., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Galliker, M. Y., Ghosh, D., Groom, L., Hausman, K., Ichter, B., Jakubczak, S., Jones, T., Ke, L., LeBlanc, D., … Zhilinsky, U. (2025). $π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization. arXiv. http://arxiv.org/abs/2504.16054v1
Chicago
Intelligence, P., K. Black, N. Brown, et al. 2025. “$π_{0.5}$: A Vision-Language-Action Model with Open-World Generalization”. arXiv. http://arxiv.org/abs/2504.16054v1.
Harvard
Intelligence, P. et al. (2025) “$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2504.16054v1.
Vancouver
1. Intelligence P, Black K, Brown N, et al (2025) $π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization. arXiv

BibTeX

@article{intelligence2025vision,
  title = {$π_\{0.5\}$: a Vision-Language-Action Model with Open-World Generalization},
  author = {Intelligence, Physical and Black, Kevin and Brown, Noah and Darpinian, James and Dhabalia, Karan and Driess, Danny and Esmail, Adnan and Equi, Michael and Finn, Chelsea and Fusai, Niccolo and Galliker, Manuel Y. and Ghosh, Dibya and Groom, Lachy and Hausman, Karol and Ichter, Brian and Jakubczak, Szymon and Jones, Tim and Ke, Liyiming and LeBlanc, Devin and Levine, Sergey and Li-Bell, Adrian and Mothukuri, Mohith and Nair, Suraj and Pertsch, Karl and Ren, Allen Z. and Shi, Lucy Xiaoyang and Smith, Laura and Springenberg, Jost Tobias and Stachowicz, Kyle and Tanner, James and Vuong, Quan and Walke, Homer and Walling, Anna and Wang, Haohuan and Yu, Lili and Zhilinsky, Ury},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2504.16054v1},
  eprint = {2504.16054}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors