Octo: An Open-Source Generalist Robot Policy
O. TeamDibya GhoshH. WalkeKarl PertschKevin BlackOier MeesS. DasariJoey HejnaTobias KreimanCharles Xu
Presents Octo, an open-source generalist robot manipulation policy trained on 800,000 trajectories that supports language and goal-image conditioning and efficiently adapts to new observation and action spaces across diverse robotic platforms using consumer GPUs.
Training specialized robot control policies from scratch for every new task and hardware configuration requires massive data collection and yields narrow operational generalization. While large foundation models have transformed natural language and computer vision, developing equivalent generalist robot policies has remained difficult due to wide variations in robot hardware, camera setups, control interfaces, and compute constraints.
The article introduces and evaluates Octo, an open-source, transformer-based generalist robot policy designed to control diverse robotic arms directly from visual and task inputs and to serve as an adaptable base model for new setups. Specifically, the article demonstrates how a unified model can handle varied camera streams, language instructions, and goal images, while enabling fast, data-efficient adaptation to unseen sensors, action spaces, and robot embodiments.
To build Octo, the authors trained a flexible transformer architecture with a conditional diffusion action-decoding head on 800,000 demonstration trajectories curated from 25 datasets within the Open X-Embodiment repository. The model converts multimodal inputs into tokens using modular encoders and outputs sequences of continuous actions. The authors evaluated Octo across nine real-world robotic platforms at four institutions, measuring its direct out-of-the-box control capabilities, its fine-tuning performance when provided with roughly 100 domain-specific demonstrations on standard consumer hardware, and the impact of core architectural and data choices.
The evaluation produced four primary findings. First, when operating out of the box across multiple robots, Octo achieved a 29% higher average success rate than the leading open-source alternative (RT-1-X) and matched the performance of much larger closed models (such as the 55-billion-parameter RT-2-X). Second, when fine-tuned on new tasks and unseen robot setups, Octo achieved an average success rate of 72%, outperforming standard from-scratch models (20%) and state-of-the-art visual representations (15%) by 52 percentage points. Third, Octo successfully adapted to entirely new observation types, such as force-torque sensing, and new control interfaces, including joint-position control and dual-arm coordination. Fourth, architectural ablations confirmed that pairing a scalable vision-transformer backbone with a continuous diffusion action head and a diverse multi-robot dataset significantly outperformed traditional convolutional networks, mean-squared-error objectives, and discretized action heads.
These findings indicate that generalist robot policies can substantially lower the cost, time, and data collection burdens associated with deploying robotic manipulation systems. By providing an open-source model that fine-tunes within five hours on a single consumer graphics processing unit using only about 100 demonstrations, Octo offers a practical, standardized foundation that eliminates the need to train complex policies from scratch.
Organizations developing robotic manipulation workflows should adopt Octo as an initialization baseline for single- and dual-arm tasks rather than relying on from-scratch policy learning. Future development should prioritize expanding the pretraining corpus with richer multi-view data, autonomous interaction datasets, and broader robotic form factors, including mobile manipulation.
Decision-makers should note certain limitations: Octo currently exhibits degraded zero-shot performance when faced with entirely novel manipulation skills or altered environmental scenes, and it struggles to effectively leverage wrist-camera inputs and complex language nuances due to underrepresentation in the pretraining data. Nonetheless, for standard single- and dual-arm manipulation tasks, confidence in Octo's capabilities as a robust, data-efficient base model remains high.
- Paper: RT-1: Robotics Transformer for Real-World Control at Scale, Anthony Brohan et al. (2023). RT-1 establishes the foundation of training multi-task transformer models on diverse robotic demonstrations, which directly motivates Octo's open-source generalist architecture.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). RT-2 introduces vision-language-action paradigms by transferring broad web pre-training to robot control, providing essential conceptual framing for generalist manipulation policies like Octo.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). PaLM-E demonstrates how multi-modal sensor and language inputs can be unified in large embodied transformer models, establishing key groundwork for Octo's multimodal conditioning.
- Paper: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware, Tony Z. Zhao et al. (2023). This paper presents Action Chunking with Transformers (ACT) for robot imitation learning, offering foundational methods for handling continuous multi-step robotic action prediction.
- Paper: Decision Transformer: Reinforcement Learning via Sequence Modeling, Lili Chen et al. (2021). Decision Transformer establishes the framing of sequential robotic decision-making and control as conditional sequence modeling with transformers.
- Paper: DINOv2: Learning Robust Visual Features without Supervision, Maxime Oquab et al. (2023). DINOv2 provides the robust self-supervised visual feature extractors commonly leveraged as pretrained visual backbones in large transformer-based robot policies.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). OpenVLA extends the open-source generalist VLA paradigm initialized by models like Octo by integrating autoregressive 7B language backbones trained on the Open X-Embodiment dataset.
- Paper: π0: A Vision-Language-Action Flow Model for General Robot Control, Kevin Black et al. (2024). π0 advances generalist robot control beyond Octo's transformer and diffusion structure by pairing pretrained vision-language models with flow matching continuous action experts.
- Paper: RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation, Songming Liu et al. (2025). RDT-1B builds on multi-robot foundation models like Octo by scaling diffusion transformers to 1.2 billion parameters specifically for bimanual manipulation.
- Paper: π0.5: a Vision-Language-Action Model with Open-World Generalization, Physical Intelligence et al. (2025). π0.5 builds directly upon vision-language-action architectures to extend generalist robotic policies into long-horizon, open-world household tasks.
- Paper: TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers, Bin Yu et al. (2026). TwinBrainVLA builds on VLA policies trained on Open X-Embodiment by introducing an asymmetric dual-pathway transformer architecture to prevent catastrophic forgetting.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). DynamicVLA extends generalist VLA manipulation frameworks like Octo to handle dynamic, moving objects through continuous inference and action streaming.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). CoA-VLA enhances generalist robot foundation policies by integrating visual-textual step-by-step affordance reasoning directly into action prediction.
