Octo: An Open-Source Generalist Robot Policy

O. TeamDibya GhoshH. WalkeKarl PertschKevin BlackOier MeesS. DasariJoey HejnaTobias KreimanCharles Xu

article2024Robotics: Science and Systems Conference1,830 citations

Presents Octo, an open-source generalist robot manipulation policy trained on 800,000 trajectories that supports language and goal-image conditioning and efficiently adapts to new observation and action spaces across diverse robotic platforms using consumer GPUs.

Listen

Training specialized robot control policies from scratch for every new task and hardware configuration requires massive data collection and yields narrow operational generalization. While large foundation models have transformed natural language and computer vision, developing equivalent generalist robot policies has remained difficult due to wide variations in robot hardware, camera setups, control interfaces, and compute constraints.

The article introduces and evaluates Octo, an open-source, transformer-based generalist robot policy designed to control diverse robotic arms directly from visual and task inputs and to serve as an adaptable base model for new setups. Specifically, the article demonstrates how a unified model can handle varied camera streams, language instructions, and goal images, while enabling fast, data-efficient adaptation to unseen sensors, action spaces, and robot embodiments.

To build Octo, the authors trained a flexible transformer architecture with a conditional diffusion action-decoding head on 800,000 demonstration trajectories curated from 25 datasets within the Open X-Embodiment repository. The model converts multimodal inputs into tokens using modular encoders and outputs sequences of continuous actions. The authors evaluated Octo across nine real-world robotic platforms at four institutions, measuring its direct out-of-the-box control capabilities, its fine-tuning performance when provided with roughly 100 domain-specific demonstrations on standard consumer hardware, and the impact of core architectural and data choices.

The evaluation produced four primary findings. First, when operating out of the box across multiple robots, Octo achieved a 29% higher average success rate than the leading open-source alternative (RT-1-X) and matched the performance of much larger closed models (such as the 55-billion-parameter RT-2-X). Second, when fine-tuned on new tasks and unseen robot setups, Octo achieved an average success rate of 72%, outperforming standard from-scratch models (20%) and state-of-the-art visual representations (15%) by 52 percentage points. Third, Octo successfully adapted to entirely new observation types, such as force-torque sensing, and new control interfaces, including joint-position control and dual-arm coordination. Fourth, architectural ablations confirmed that pairing a scalable vision-transformer backbone with a continuous diffusion action head and a diverse multi-robot dataset significantly outperformed traditional convolutional networks, mean-squared-error objectives, and discretized action heads.

These findings indicate that generalist robot policies can substantially lower the cost, time, and data collection burdens associated with deploying robotic manipulation systems. By providing an open-source model that fine-tunes within five hours on a single consumer graphics processing unit using only about 100 demonstrations, Octo offers a practical, standardized foundation that eliminates the need to train complex policies from scratch.

Organizations developing robotic manipulation workflows should adopt Octo as an initialization baseline for single- and dual-arm tasks rather than relying on from-scratch policy learning. Future development should prioritize expanding the pretraining corpus with richer multi-view data, autonomous interaction datasets, and broader robotic form factors, including mobile manipulation.

Decision-makers should note certain limitations: Octo currently exhibits degraded zero-shot performance when faced with entirely novel manipulation skills or altered environmental scenes, and it struggles to effectively leverage wrist-camera inputs and complex language nuances due to underrepresentation in the pretraining data. Nonetheless, for standard single- and dual-arm manipulation tasks, confidence in Octo's capabilities as a robust, data-efficient base model remains high.

Cover for Octo: An Open-Source Generalist Robot Policy

Abstract

Large policies pretrained on diverse robot datasets have the potential to transform robotic learning: instead of training new policies from scratch, such generalist robot policies may be finetuned with only a little in-domain data, yet generalize broadly. However, to be widely applicable across a range of robotic learning scenarios, environments, and tasks, such policies need to handle diverse sensors and action spaces, accommodate a variety of commonly used robotic platforms, and finetune readily and efficiently to new domains. In this work, we aim to lay the groundwork for developing open-source, widely applicable, generalist policies for robotic manipulation. As a first step, we introduce Octo, a large transformer-based policy trained on 800k trajectories from the Open X-Embodiment dataset, the largest robot manipulation dataset to date. It can be instructed via language commands or goal images and can be effectively finetuned to robot setups with new sensory inputs and action spaces within a few hours on standard consumer GPUs. In experiments across 9 robotic platforms, we demonstrate that Octo serves as a versatile policy initialization that can be effectively finetuned to new observation and action spaces. We also perform detailed ablations of design decisions for the Octo model, from architecture to training data, to guide future research on building generalist robot models.

Table of Contents

  • I Introduction
  • II Related Work
  • III The Octo Model
  • III-A Architecture
  • III-B Training data
  • III-C Training objective
  • III-D Training Details
  • III-E Model Checkpoints & Code
  • IV Experiments
  • IV-A Octo Controls Multiple Robots Out-of-the-Box
  • IV-B Octo Enables Data-Efficient Learning in New Domains
  • IV-C Design Decisions for Generalist Robot Policy Training
  • V Discussion and Future Work
  • References
  • A Contributions
  • B Octo Code Example
  • C Data mixture
  • D Training Hyperparameters
  • E Things that Worked and Did Not Work (Yet)
  • F Experimental Setups
  • F-A Zero-Shot Evaluations
  • F-B Model Ablations
  • F-C Finetuning Evaluations

Knowls

  1. Knowl 1 — Octo Transformer Architecture with Block-Wise Masked Attention and Readout Tokens

    model/method

    Octo is a generalist robot manipulation policy based on an encoder-decoder style transformer backbone that operates over sequences of task tokens and observation tokens.

    Input Tokenizers

    • Language Inputs: Natural language instructions ℓ\ell are tokenized and processed by a frozen T5-Base (111M parameter) transformer model to yield a fixed sequence of TT=16T_T = 16 language embedding tokens.
    • Image Inputs: Third-person RGB camera observations (256×256256 \times 256 resolution) and wrist camera observations (128×128128 \times 128 resolution), as well as visual goal images gg, are passed through shallow convolutional patch encoders that partition inputs into non-overlapping 16×1616 \times 16 pixel patches. This generates 256 tokens for third-person images and 64 tokens for wrist camera images per observation timestep.
    • Position Embeddings: Learnable positional embeddings pp are added to task and observation tokens, arranged sequentially as [TT,To,1,To,2,… ][T_T, T_{o,1}, T_{o,2}, \dots] over an observation history window (H=2H=2 frames).

    Backbone and Block-Wise Masking

    The unified token sequence is fed into a transformer backbone. Attention is block-wise masked:

    • Observation tokens at timestep tt attend causally to all task tokens TTT_T and observation tokens from timesteps ≤t\le t.
    • Missing modalities (e.g., absent wrist camera streams or unannotated language instructions) are completely masked out.

    Readout Tokens and Action Prediction

    Learned readout tokens TR,tT_{R,t} are placed at each observation timestep tt. A readout token attends to all preceding task and observation tokens up to time tt, but no observation or task token attends to TR,tT_{R,t}. The readout token acts as a passive sequence summary (analogous to a [CLS][\text{CLS}] token). Its output embedding ee is routed to a lightweight action head that predicts a chunk of continuous future actions.

    Modular Downstream Adaptation

    During downstream finetuning, new sensor modalities (e.g., force-torque proprioception) or new output heads (e.g., joint-space control or multi-arm action spaces) are added by introducing new lightweight patch/token encoders, positional embeddings, or output heads without modifying or re-initializing any pretrained transformer backbone weights.

  2. Knowl 2 — Continuous Action Chunk Prediction via Conditional Diffusion Decoding

    model/method

    Octo generates continuous, multi-modal robot action chunks using a conditional Denoising Diffusion Probabilistic Model (DDPM) action decoding head.

    Action Head Architecture

    The action head is a 3-layer multilayer perceptron (MLP) with a hidden dimension of 256, residual connections, and layer normalization. It operates on the transformer backbone's action readout embedding ee.

    Iterative Reverse Diffusion

    Starting from a standard Gaussian noise vector xK∼N(0,I)x^K \sim \mathcal{N}(0, I), the action head computes the action sequence chunk over K=20K = 20 discrete denoising steps using a cosine noise schedule:

    xk−1=α(xk−γϵθ(xk,e,k)+N(0,σ2I))x^{k-1} = \alpha \left(x^k - \gamma \epsilon_\theta(x^k, e, k) + \mathcal{N}(0, \sigma^2 I)\right)

    where:

    • k∈{K,K−1,…,1}k \in \{K, K-1, \dots, 1\} is the diffusion timestep index,
    • xkx^k is the noisy action candidate at step kk,
    • ee is the conditioning embedding from the transformer backbone readout token,
    • ϵθ(xk,e,k)\epsilon_\theta(x^k, e, k) is the learned denoising MLP,
    • α,γ,σ\alpha, \gamma, \sigma are schedule parameters determined by the cosine noise schedule.

    Computational Efficiency

    The transformer backbone is evaluated only once per decision step to compute ee. The entire multi-step iterative reverse diffusion loop (K=20K=20 iterations) is performed exclusively inside the lightweight 3-layer MLP head, eliminating repeated backbone forward passes during inference.

  3. Knowl 3 — Open X-Embodiment Pretraining Dataset Mixture and Curation Protocol

    data/table

    Octo is pretrained on a curated collection of 800,000 robot demonstration trajectories selected from 25 individual datasets within the Open X-Embodiment (OXE) repository.

    Curation and Standardization Rules

    1. Modality Filtering: Datasets lacking RGB visual camera streams are removed.
    2. Action Space Filtering: Only datasets using delta end-effector Cartesian control are retained for pretraining.
    3. Quality and Repetition Filtering: Overly repetitive datasets, datasets with low image resolutions, and datasets comprising overly narrow niche tasks are excluded.
    4. Gripper Normalization: Gripper action dimensions across all datasets are standardized such that +1.0+1.0 corresponds to a fully open gripper and 0.00.0 corresponds to a fully closed gripper. Missing camera channels are zero-padded.
    5. Diversity-Based Sampling Weighting: Datasets are categorized into "more diverse" and "less diverse" based on environment and task variability; diverse datasets receive double sampling weights during training, while repetitive datasets are down-weighted.
    Dataset Pretraining Sampling Weight (%)
    Fractal 17.0%
    Kuka 17.0%
    Bridge 17.0%
    BC-Z 9.1%
    Stanford Hydra Dataset 6.0%
    Language Table 5.9%
    Taco Play 3.6%
    Furniture Bench Dataset 3.3%
    UTAustin Mutex 3.0%
    Austin Sailor Dataset 2.9%
    Roboturk 2.8%
    Toto 2.4%
    Austin Sirius Dataset 2.3%
    Berkeley Autolab UR5 1.5%
    IAMLab CMU Pickup Insert 1.2%
    Viola 1.2%
    Berkeley Fanuc Manipulation 1.0%
    NYU Franka Play Dataset 0.9%
    Jaco Play 0.6%
    Berkeley Cable Routing 0.3%
    Austin Buds Dataset 0.3%
    CMU Stretch 0.2%
    NYU Door Opening 0.1%
    DLR EDAN Shared Control 0.1%
    UCSD Kitchen Dataset <0.1<0.1%

    Fractal, Kuka, and Bridge constitute the three largest components, each contributing 17.0% of the training batches.

  4. Knowl 4 — Octo Model Configurations and Pretraining Hyperparameters

    experimental setup

    Octo is released in two primary architectural scales mirroring Vision Transformer standards:

    • Octo-Small (27M parameters): 12 transformer layers, hidden dimension D=384D = 384, MLP feed-forward dimension 1536, 6 attention heads.
    • Octo-Base (93M parameters): 12 transformer layers, hidden dimension D=768D = 768, MLP feed-forward dimension 3072, 12 attention heads.
    • Octo-Tiny (10M parameters): Used for architecture scaling ablations.

    Pretraining Hyperparameters

    • Optimizer: AdamW with weight decay 0.10.1 and gradient clipping threshold 1.01.0.
    • Learning Rate Schedule: Peak learning rate 3×10−43 \times 10^{-4}, linear warmup for 2,000 steps, followed by reciprocal square-root learning rate decay.
    • Batch Size: 2048.
    • Compute: Octo-Base is trained for 300,000 steps on a TPU v4-128 pod (approx. 14 hours).
    • Observation History: H=2H = 2 frames (current and one previous frame).
    • Task Conditioning Strategy: Hindsight goal relabeling assigns a uniformly sampled future frame in the trajectory as the goal image. Language instructions and goal images are randomly dropped out during training; datasets lacking language annotations exclusively use goal image conditioning.
    • Data Augmentation: Third-person images undergo random cropping, resizing to 256×256256 \times 256, and color jitter, with pixel values normalized to [−1.0,1.0][-1.0, 1.0]. Wrist camera images are resized to 128×128128 \times 128 with color jitter and normalization but without random cropping.
  5. Knowl 5 — Cross-Embodiment Finetuning to New Observations, Actions, and Setups

    data/table

    Octo can be finetuned to new robotic embodiments, observation spaces, and action spaces using ∼100\sim 100 target domain demonstrations. The finetuning protocol updates all model parameters for 50,000 steps using cosine learning rate decay with linear warmup on a single NVIDIA A5000 GPU (<5< 5 hours runtime).

    Finetuning performance is evaluated across six real-world domains and compared against:

    1. ResNet+Transformer Scratch: A 28M parameter policy combining a ResNet visual encoder, FiLM conditioning, and a transformer diffusion decoder trained from scratch on target demonstrations.
    2. VC-1: A ViT-B visual encoder pretrained on 4,000 hours of egocentric video and ImageNet, paired with an MLP action decoder trained via MSE loss.
    Method Berkeley Insertion∗^* Stanford Coffee CMU Baking Berkeley Pick-Up†^\dagger Berkeley Coke Berkeley Bimanual†^\dagger Average
    ResNet+Transformer Scratch 10% 45% 25% 0% 20% 20% 20%
    VC-1 5% 0% 30% 0% 10% 50% 15%
    Octo (Ours) 70% 75% 50% 60% 100% 80% 72%
    • Berkeley Insertion evaluates a novel observation modality (∗^*force-torque sensor inputs).
    • Berkeley Pick-Up and Berkeley Bimanual evaluate novel action spaces (†^\daggerjoint position control, including a 14-dimensional action space on an ALOHA bimanual platform).
    • Berkeley Coke evaluates an unseen robot embodiment (ViperX).
    • Octo achieves an average success rate of 72%, outperforming training from scratch (20%) and pretrained visual representations (15%) by absolute margins of 52% and 57%, respectively.
  6. Knowl 6 — Out-of-the-Box Zero-Shot Multi-Robot Manipulation

    empirical result

    Octo-Base (93M parameters) controls multiple robot platforms out-of-the-box ("zero-shot") without target-domain finetuning, using language instructions or visual goal images across three distinct robot embodiments:

    • WidowX BridgeV2 (6-DOF arm, tabletop manipulation)
    • UR5 Tabletop (wiping tables, bowl-to-bowl sorting)
    • Google RT-1 Robot (picking items, drawer manipulation)

    Comparative Results

    • Outperforming RT-1-X: Octo achieves an average success rate 29% higher than RT-1-X (35M parameters), the leading open-source generalist robot manipulation policy.
    • Competitiveness with Large Vision-Language-Action Models: On the evaluated WidowX and RT-1 Robot benchmarks, Octo performs competitively with RT-2-X (a 55-billion parameter model), despite being more than 500×\times smaller in parameter count.
    • Goal Image vs. Language Conditioning: On WidowX tasks, conditioning Octo on goal images yields a 25% absolute improvement in task success over language conditioning (achieving 80% average success on novel objects with goal images versus 85% on in-distribution tasks with language), demonstrating that visual goal tokens provide richer geometric target specifications.
  7. Knowl 7 — Model Ablations on Policy Architecture, Objective, and Data Mixture

    data/table

    Ablations conducted on the WidowX platform (evaluated over 40 trials across two language-conditioned tasks and two goal-conditioned tasks) measure the impact of architecture, training data diversity, and action prediction objectives on Octo-Small (27M parameters):

    Ablation Category Model Configuration Aggregate Success Rate (%)
    Full Octo Policy Octo-Small (ViT backbone + Diffusion Head + 25-Dataset Mix) 83%
    Training Data RT-X dataset mix (11 datasets) 60%
    Single-robot data (Bridge dataset only) 43%
    Action Policy Head Discretized Action Head (256 bins, Cross-Entropy) 18%
    Continuous Action Head (MSE Loss) 35%
    Architecture ResNet-50 + Transformer 70%

    Key Findings

    1. Architecture: The "transformer-first" ViT patch-based backbone (83%) outperforms combining a ResNet-50 encoder with a small transformer (70%) when scaling to diverse multi-embodiment data.
    2. Action Objective: The conditional diffusion head (83%) substantially outperforms discrete action heads (18%) and continuous MSE loss (35%). Diffusion models multi-modal action distributions while preserving fine continuous control precision.
    3. Data Diversity: Expanding pretraining data from single-robot data (43%) to 11 datasets (60%) and ultimately to 25 datasets (83%) produces consistent, monotonic gains in downstream control capabilities.
  8. Knowl 8 — Model Scaling and Zero-Shot Generalization Boundaries

    empirical result

    Evaluation of Octo across parameter scales and task distributions reveals clear scaling trends and generalization boundaries:

    Parameter Scaling

    Zero-shot manipulation success increases monotonically with model scale from Octo-Tiny (10M) to Octo-Small (27M) and Octo-Base (93M) on both UR5 and WidowX platforms. Larger models exhibit improved visual perception robustness to initial object clutter and significantly fewer premature grasping attempts.

    Generalization Axes (WidowX Platform with Octo-Small)

    • In-Distribution Tasks ("Put carrot on plate", "Put eggplant in pot"): 85% average success (80% and 90%).
    • Novel Objects ("Put bread on plate", "Put spoon on glove"): 80% average success (70% and 90%).
    • Novel Environments ("Put mushroom in pot", "Put spoon on cloth"): 40% average success (20% and 60%).
    • Novel Motor Skills ("Flip cup on its side", "Put block in slot"): 5% average success (10% and 0%).

    While Octo generalizes effectively to novel objects within known environments, its zero-shot transfer degrades in unseen scenes and fails almost entirely when required to execute motor primitives/skills not present in the pretraining demonstrations for that embodiment.

  9. Knowl 9 — Empirical Design Insights and Negative Results in Robot Pretraining

    empirical result

    Extensive empirical ablations during Octo development identified several techniques that critically influenced policy performance:

    What Worked

    • Frame-Level Shuffling and Large Replay Buffer: Loading 25 heterogeneous datasets requires shuffling and interleaving frames before image decoding to support a 500k-frame buffer. Small shuffle buffers (20k) with episode-level interleaving severely degraded zero-shot control. Subsampling ≤100\le 100 random steps per trajectory prevented long episodes from dominating training batches.
    • Action Chunking: Predicting multi-step action trajectories (chunks) generated significantly smoother, more coherent robot trajectories than single-step action predictions. Temporal ensembling of overlapping chunks provided no measurable benefit over simple receding-horizon execution.
    • Patch Token Size: Using 16×1616 \times 16 pixel patches outperformed 32×3232 \times 32 patches, providing higher spatial resolution necessary for precise grasping.

    What Failed

    • Proprioceptive Observation Inputs: Including proprioceptive states during pretraining degraded performance. This failure is attributed to causal confusion, where the policy over-relies on current proprioception to predict near-future movements rather than attending to visual scene context.
    • Relative Gripper Action Parametrization: Parametrizing gripper commands as relative state transitions (+1/0+1/0 only during opening/closing events) produced higher initial grasp success but eliminated recovery and retry behaviors after a failed grasp. Absolute binary gripper commands (+1+1 open, 00 closed) proved superior.
    • Scaling Language Encoders: Scaling the language model from T5-Base (111M) to T5-Large (386M) or finetuning the upper layers of T5 yielded no performance gains, likely due to the lack of rich, diverse, free-form language annotations across current robotics datasets.
  10. Knowl 10 — Limitations in Sensory Modality Balancing and Demonstration-Only Pretraining

    limitation

    The authors identify four principal limitations in the Octo generalist policy:

    1. Wrist Camera Ingestion: Octo struggles to effectively utilize wrist-mounted camera streams. In finetuning experiments, relying solely on a third-person camera often outperformed combining third-person and wrist camera feeds. This is attributed to data imbalance, as only 27% of the pretraining demonstrations contain wrist camera streams.
    2. Language vs. Goal Conditioning Performance Disparity: Language-conditioned control performs notably worse than goal-image-conditioned control, caused by sparse annotations where only 56% of pretraining trajectories include language descriptions.
    3. Imitation Learning from Demonstrations: Octo is trained via supervised behavioral cloning on expert demonstrations from Open X-Embodiment. It lacks mechanisms to incorporate sub-optimal trajectories, offline reinforcement learning objectives, or autonomous online interaction data.
    4. Embodiment Scope: Pretraining and evaluation are restricted to stationary single-arm and dual-arm manipulation platforms, without coverage of mobile manipulators, quadrupeds, or navigation robots.

Coverage note — None was omitted.

References

  1. 1.Scale AI. Introducing scale’s automotive foundation model, 2023. URL https://scale.com/blog/afm1.
  2. 2.Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba. Hindsight experience replay. In NeurIPS, 2017.
  3. 3.Shikhar Bahl, Abhinav Gupta, and Deepak Pathak. Human-to-robot imitation in the wild. arXiv preprint arXiv:2207.09450, 2022.
  4. 4.Shikhar Bahl, Russell Mendonca, Lili Chen, Unnat Jain, and Deepak Pathak. Affordances from human videos as a versatile representation for robotics. In CVPR, 2023.
  5. 5.Suneel Belkhale, Yuchen Cui, and Dorsa Sadigh. Hydra: Hybrid robot actions for imitation learning. arxiv, 2023.
  6. 6.Homanga Bharadhwaj, Jay Vakil, Mohit Sharma, Abhinav Gupta, Shubham Tulsiani, and Vikash Kumar. Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking. arXiv preprint arXiv:2309.01918, 2023.
  7. 7.Kevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Walke, Chelsea Finn, Aviral Kumar, and Sergey Levine. Zero-shot robotic manipulation with pretrained image-editing diffusion models. arXiv preprint arXiv:2310.10639, 2023.
  8. 8.Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, et al. End to end learning for self-driving cars. arXiv preprint arXiv:1604.07316, 2016.
  9. 9.Konstantinos Bousmalis, Giulia Vezzani, Dushyant Rao, Coline Devin, Alex X Lee, Maria Bauza, Todor Davchev, Yuxiang Zhou, Agrim Gupta, Akhil Raju, et al. Robocat: A self-improving foundation agent for robotic manipulation. arXiv preprint arXiv:2306.11706, 2023.
  10. 10.Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022.
  11. 11.Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, et al. Do as i can, not as i say: Grounding language in robotic affordances. In Conference on Robot Learning, pages 287–318. PMLR, 2023.
  12. 12.Serkan Cabi, Sergio Gómez Colmenarejo, Alexander Novikov, Ksenia Konyushkova, Scott Reed, Rae Jeong, Konrad Zolna, Yusuf Aytar, David Budden, Mel Vecerik, Oleg Sushkov, David Barker, Jonathan Scholz, Misha Denil, Nando de Freitas, and Ziyu Wang. Scaling data-driven robotics with reward sketching and batch reinforcement learning. arXiv preprint arXiv:1909.12200, 2019.
  13. 13.Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11621–11631, 2020.
  14. 14.Lawrence Yunliang Chen, Simeon Adebola, and Ken Goldberg. Berkeley UR5 demonstration dataset. https://sites.google.com/view/berkeley-ur5/home.
  15. 15.William Chen, Oier Mees, Aviral Kumar, and Sergey Levine. Vision-language models provide promptable representations for reinforcement learning. arXiv preprint arXiv:2402.02651, 2024.
  16. 16.Zoey Chen, Sho Kiami, Abhishek Gupta, and Vikash Kumar. Genaug: Retargeting behaviors to unseen situations via generative augmentation. arXiv preprint arXiv:2302.06671, 2023.
  17. 17.Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. In Proceedings of Robotics: Science and Systems (RSS), 2023.
  18. 18.Zichen Jeff Cui, Yibin Wang, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto. From play to policy: Conditional behavior generation from uncurated robot data. In The Eleventh International Conference on Learning Representations, 2022.
  19. 19.Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn. Robonet: Large-scale multi-robot learning. In Conference on Robot Learning, pages 885–897. PMLR, 2020.
  20. 20.Shivin Dass, Jullian Yapeter, Jesse Zhang, Jiahui Zhang, Karl Pertsch, Stefanos Nikolaidis, and Joseph J. Lim. CLVR jaco play dataset, 2023. URL https://github.com/clvrai/clvr_jaco_play_dataset.
  21. 21.Pim de Haan, Dinesh Jayaraman, and Sergey Levine. Causal confusion in imitation learning. NeurIPS, 2019.
  22. 22.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  23. 23.Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023.
  24. 24.Maximilian Du, Suraj Nair, Dorsa Sadigh, and Chelsea Finn. Behavior retrieval: Few-shot imitation learning by querying unlabeled datasets. ArXiv, abs/2304.08742, 2023. URL https://api.semanticscholar.org/CorpusID:258186973.
  25. 25.Frederik Ebert, Yanlai Yang, Karl Schmeckpeper, Bernadette Bucher, Georgios Georgakis, Kostas Daniilidis, Chelsea Finn, and Sergey Levine. Bridge data: Boosting generalization of robotic skills with cross-domain datasets. arXiv preprint arXiv:2109.13396, 2021.
  26. 26.Hao-Shu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Chenxi Wang, Junbo Wang, Haoyi Zhu, and Cewu Lu. Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot. Towards Generalist Robots: Learning Paradigms for Scalable Skill Acquisition@ CoRL2023, 3:5, 2023.
  27. 27.Chelsea Finn and Sergey Levine. Deep visual foresight for planning robot motion. In 2017 IEEE International Conference on Robotics and Automation (ICRA), pages 2786–2793. IEEE, 2017.
  28. 28.Zipeng Fu, Tony Z Zhao, and Chelsea Finn. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation. arXiv preprint arXiv:2401.02117, 2024.
  29. 29.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE conference on computer vision and pattern recognition, pages 3354–3361. IEEE, 2012.
  30. 30.Abhinav Gupta, Adithyavairavan Murali, Dhiraj Prakashchand Gandhi, and Lerrel Pinto. Robot learning in homes: Improving generalization and reducing dataset bias. Advances in neural information processing systems, 31, 2018.
  31. 31.Huy Ha, Pete Florence, and Shuran Song. Scaling up and distilling down: Language-guided robot skill acquisition. In Conference on Robot Learning, pages 3766–3777. PMLR, 2023.
  32. 32.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  33. 33.Minho Heo, Youngwoon Lee, Doohyun Lee, and Joseph J. Lim. Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation. In Robotics: Science and Systems, 2023.
  34. 34.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  35. 35.Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gianluca Corrado. Gaia-1: A generative world model for autonomous driving, 2023.
  36. 36.Chenguang Huang, Oier Mees, Andy Zeng, and Wolfram Burgard. Visual language maps for robot navigation. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 10608–10615. IEEE, 2023.
  37. 37.Chenguang Huang, Oier Mees, Andy Zeng, and Wolfram Burgard. Audio visual language maps for robot navigation. In Proceedings of the International Symposium on Experimental Robotics (ISER), Chiang Mai, Thailand, 2023.
  38. 38.Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei. Voxposer: Composable 3d value maps for robotic manipulation with language models. arXiv preprint arXiv:2307.05973, 2023.
  39. 39.Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn. Bc-z: Zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning, pages 991–1002. PMLR, 2022.
  40. 40.Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan. VIMA: Robot manipulation with multimodal prompts. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 14975–15022. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/v202/jiang23b.html.
  41. 41.Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al. QT-Opt: Scalable deep reinforcement learning for vision-based robotic manipulation. arXiv preprint arXiv:1806.10293, 2018.
  42. 42.Dmitry Kalashnikov, Jake Varley, Yevgen Chebotar, Benjamin Swanson, Rico Jonschkowski, Chelsea Finn, Sergey Levine, and Karol Hausman. Scaling up multi-task robotic reinforcement learning. In 5th Annual Conference on Robot Learning, 2021.
  43. 43.Haresh Karnan, Anirudh Nair, Xuesu Xiao, Garrett Warnell, Sören Pirk, Alexander Toshev, Justin Hart, Joydeep Biswas, and Peter Stone. Socially compliant navigation dataset (scand): A large-scale dataset of demonstrations for social navigation. IEEE Robotics and Automation Letters, 7(4):11807–11814, 2022.
  44. 44.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, et al. Segment Anything, April 2023.
  45. 45.Vikash Kumar, Rutav Shah, Gaoyue Zhou, Vincent Moens, Vittorio Caggiano, Abhishek Gupta, and Aravind Rajeswaran. Robohive: A unified framework for robot learning. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. URL https://openreview.net/forum?id=0H5fRQcpQ7.
  46. 46.Teyun Kwon, Norman Di Palo, and Edward Johns. Language models as zero-shot trajectory generators. arXiv preprint arXiv:2310.11604, 2023.
  47. 47.Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. End-to-end training of deep visuomotor policies. The Journal of Machine Learning Research, 17(1):1334–1373, 2016.
  48. 48.Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen. Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection. The International journal of robotics research, 37(4-5):421–436, 2018.
  49. 49.Yixin Lin, Austin S. Wang, Giovanni Sutanto, Akshara Rai, and Franziska Meier. Polymetis. https://facebookresearch.github.io/fairo/polymetis/, 2021.
  50. 50.Huihan Liu, Soroush Nasiriany, Lance Zhang, Zhiyao Bao, and Yuke Zhu. Robot learning on the job: Human-in-the-loop autonomy and learning during deployment. In Robotics: Science and Systems (RSS), 2023.
  51. 51.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2018.
  52. 52.Jianlan Luo, Charles Xu, Xinyang Geng, Gilbert Feng, Kuan Fang, Liam Tan, Stefan Schaal, and Sergey Levine. Multi-stage cable routing through hierarchical imitation learning. arXiv preprint arXiv:2307.08927, 2023.
  53. 53.Jianlan Luo, Charles Xu, Fangchen Liu, Liam Tan, Zipeng Lin, Jeffrey Wu, Pieter Abbeel, and Sergey Levine. Fmb: a functional manipulation benchmark for generalizable robotic learning. arXiv preprint arXiv:2401.08553, 2024.
  54. 54.Corey Lynch and Pierre Sermanet. Language conditioned imitation learning over unstructured data. In RSS, 2021.
  55. 55.Corey Lynch, Ayzaan Wahid, Jonathan Tompson, Tianli Ding, James Betker, Robert Baruch, Travis Armstrong, and Pete Florence. Interactive language: Talking to robots in real time. IEEE Robotics and Automation Letters, 2023.
  56. 56.Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma, Claire Chen, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Pieter Abbeel, Jitendra Malik, Dhruv Batra, Yixin Lin, Oleksandr Maksymets, Aravind Rajeswaran, and Franziska Meier. Where are we in the search for an artificial visual cortex for embodied intelligence? 2023.
  57. 57.Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma, Claire Chen, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Pieter Abbeel, Jitendra Malik, et al. Where are we in the search for an artificial visual cortex for embodied intelligence? arXiv preprint arXiv:2303.18240, 2023.
  58. 58.Ajay Mandlekar, Yuke Zhu, Animesh Garg, Jonathan Booher, Max Spero, Albert Tung, Julian Gao, John Emmons, Anchit Gupta, Emre Orbay, Silvio Savarese, and Li Fei-Fei. RoboTurk: A crowdsourcing platform for robotic skill learning through imitation. CoRR, abs/1811.02790, 2018. URL http://arxiv.org/abs/1811.02790.
  59. 59.Ajay Mandlekar, Yuke Zhu, Animesh Garg, Jonathan Booher, Max Spero, Albert Tung, Julian Gao, John Emmons, Anchit Gupta, Emre Orbay, et al. Roboturk: A crowdsourcing platform for robotic skill learning through imitation. In Conference on Robot Learning, pages 879–893. PMLR, 2018.
  60. 60.Ajay Mandlekar, Soroush Nasiriany, Bowen Wen, Iretiayo Akinola, Yashraj Narang, Linxi Fan, Yuke Zhu, and Dieter Fox. Mimicgen: A data generation system for scalable robot learning using human demonstrations. In 7th Annual Conference on Robot Learning, 2023.
  61. 61.Oier Mees, Lukas Hermann, and Wolfram Burgard. What matters in language conditioned robotic imitation learning over unstructured data. IEEE Robotics and Automation Letters, 7(4):11205–11212, 2022.
  62. 62.Oier Mees, Lukas Hermann, Erick Rosete-Beas, and Wolfram Burgard. Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022.
  63. 63.Oier Mees, Jessica Borja-Diaz, and Wolfram Burgard. Grounding language with visual affordances over unstructured data. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, UK, 2023.
  64. 64.Russell Mendonca, Shikhar Bahl, and Deepak Pathak. Structured world models from human videos. CoRL, 2023.
  65. 65.Soroush Nasiriany, Tian Gao, Ajay Mandlekar, and Yuke Zhu. Learning and retrieval from prior data for skill-based imitation learning. In Conference on Robot Learning (CoRL), 2022.
  66. 66.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pages 8162–8171. PMLR, 2021.
  67. 67.Open X-Embodiment Collaboration, Abhishek Padalkar, Acorn Pooley, Ajinkya Jain, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anikait Singh, Anthony Brohan, Antonin Raffin, Ayzaan Wahid, Ben Burgess-Limerick, Beomjoon Kim, Bernhard Schölkopf, Brian Ichter, Cewu Lu, Charles Xu, Chelsea Finn, Chenfeng Xu, Cheng Chi, Chenguang Huang, Christine Chan, Chuer Pan, Chuyuan Fu, Coline Devin, Danny Driess, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Federico Ceola, Fei Xia, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Ge Yan, Giulio Schiavi, Hao Su, Hao-Shu Fang, Haochen Shi, Heni Ben Amor, Henrik I Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jaehyung Kim, Jan Schneider, Jasmine Hsu, Jeannette Bohg, Jeffrey Bingham, Jiajun Wu, Jialin Wu, Jianlan Luo, Jiayuan Gu, Jie Tan, Jihoon Oh, Jitendra Malik, Jonathan Tompson, Jonathan Yang, Joseph J. Lim, João Silvério, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Ken Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Zhang, Keyvan Majd, Krishan Rana, Krishnan Srinivasan, Lawrence Yunliang Chen, Lerrel Pinto, Liam Tan, Lionel Ott, Lisa Lee, Masayoshi Tomizuka, Maximilian Du, Michael Ahn, Mingtong Zhang, Mingyu Ding, Mohan Kumar Srirama, Mohit Sharma, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen, Nicolas Heess, Nikhil J Joshi, Niko Suenderhauf, Norman Di Palo, Nur Muhammad Mahi Shafiullah, Oier Mees, Oliver Kroemer, Pannag R Sanketi, Paul Wohlhart, Peng Xu, Pierre Sermanet, Priya Sundaresan, Quan Vuong, Rafael Rafailov, Ran Tian, Ria Doshi, Roberto Martín-Martín, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shuran Song, Sichun Xu, Siddhant Haldar, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Sudeep Dasari, Suneel Belkhale, Takayuki Osa, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Vidhi Jain, Vincent Vanhoucke, Wei Zhan, Wenxuan Zhou, Wolfram Burgard, Xi Chen, Xiaolong Wang, Xinghao Zhu, Xuanlin Li, Yao Lu, Yevgen Chebotar, Yifan Zhou, Yifeng Zhu, Ying Xu, Yixuan Wang, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh hua Wu, Yujin Tang, Yuke Zhu, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zhuo Xu, and Zichen Jeff Cui. Open X-Embodiment: Robotic learning datasets and RT-X models. https://arxiv.org/abs/2310.08864, 2023.
  68. 68.OpenAI. GPT-4 Technical Report, March 2023.
  69. 69.Jyothish Pari, Nur Muhammad Shafiullah, Sridhar Pandian Arunachalam, and Lerrel Pinto. The surprising effectiveness of representation learning for visual imitation, 2021.
  70. 70.Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018.
  71. 71.Lerrel Pinto and Abhinav Gupta. Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours. In 2016 IEEE international conference on robotics and automation (ICRA), pages 3406–3413. IEEE, 2016.
  72. 72.Gabriel Quere, Annette Hagengruber, Maged Iskandar, Samuel Bustamante, Daniel Leidner, Freek Stulp, and Joern Vogel. Shared Control Templates for Assistive Robotics. In 2020 IEEE International Conference on Robotics and Automation (ICRA), page 7, Paris, France, 2020.
  73. 73.Ilija Radosavovic, Baifeng Shi, Letian Fu, Ken Goldberg, Trevor Darrell, and Jitendra Malik. Robot learning with sensorimotor pre-training. Conference on Robot Learning, 2023.
  74. 74.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  75. 75.Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gómez Colmenarejo, Alexander Novikov, Gabriel Barth-maron, Mai Giménez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al. A generalist agent. Transactions on Machine Learning Research, 2022.
  76. 76.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models, April 2022.
  77. 77.Erick Rosete-Beas, Oier Mees, Gabriel Kalweit, Joschka Boedecker, and Wolfram Burgard. Latent plans for task agnostic offline reinforcement learning. In Proceedings of the 6th Conference on Robot Learning (CoRL), 2022.
  78. 78.Saumya Saxena, Mohit Sharma, and Oliver Kroemer. Multi-resolution sensing for real-time control with vision-language models. In 7th Annual Conference on Robot Learning, 2023. URL https://openreview.net/forum?id=WuBv9-IGDUA.
  79. 79.Nur Muhammad Mahi Shafiullah, Anant Rai, Haritheja Etukuru, Yiqian Liu, Ishan Misra, Soumith Chintala, and Lerrel Pinto. On bringing robots home, 2023.
  80. 80.Dhruv Shah, Ajay Sridhar, Arjun Bhorkar, Noriaki Hirose, and Sergey Levine. Gnm: A general navigation model to drive any robot. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 7226–7233. IEEE, 2023.
  81. 81.Dhruv Shah, Ajay Sridhar, Nitish Dashora, Kyle Stachowicz, Kevin Black, Noriaki Hirose, and Sergey Levine. ViNT: A foundation model for visual navigation. In 7th Annual Conference on Robot Learning, 2023. URL https://arxiv.org/abs/2306.14846.
  82. 82.Rutav Shah, Roberto Martín-Martín, and Yuke Zhu. MUTEX: Learning unified policies from multimodal task specifications. In 7th Annual Conference on Robot Learning, 2023. URL https://openreview.net/forum?id=PwqiqaaEzJ.
  83. 83.Mohit Shridhar, Lucas Manuelli, and Dieter Fox. Perceiver-actor: A multi-task transformer for robotic manipulation. In Conference on Robot Learning, pages 785–799. PMLR, 2023.
  84. 84.Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. Progprompt: Generating situated robot task plans using large language models. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 11523–11530. IEEE, 2023.
  85. 85.Ajay Sridhar, Dhruv Shah, Catherine Glossop, and Sergey Levine. Nomad: Goal masked diffusion policies for navigation and exploration. arXiv preprint arXiv:2310.07896, 2023.
  86. 86.Austin Stone, Ted Xiao, Yao Lu, Keerthana Gopalakrishnan, Kuang-Huei Lee, Quan Vuong, Paul Wohlhart, Sean Kirmani, Brianna Zitkovich, Fei Xia, et al. Open-world object manipulation using pre-trained vision-language models. In 7th Annual Conference on Robot Learning, 2023.
  87. 87.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2446–2454, 2020.
  88. 88.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. LLaMA: Open and Efficient Foundation Language Models, February 2023.
  89. 89.Samuel Triest, Matthew Sivaprakasam, Sean J Wang, Wenshan Wang, Aaron M Johnson, and Sebastian Scherer. Tartandrive: A large-scale dataset for learning off-road dynamics models. In 2022 International Conference on Robotics and Automation (ICRA), pages 2546–2552. IEEE, 2022.
  90. 90.Homer Walke, Kevin Black, Abraham Lee, Moo Jin Kim, Max Du, Chongyi Zheng, Tony Zhao, Philippe Hansen-Estruch, Quan Vuong, Andre He, Vivek Myers, Kuan Fang, Chelsea Finn, and Sergey Levine. Bridgedata v2: A dataset for robot learning at scale, 2023.
  91. 91.Wayve. Lingo: Natural language for autonomous driving, 2023. URL https://wayve.ai/thinking/lingo-natural-language-autonomous-driving/.
  92. 92.Philipp Wu, Arjun Majumdar, Kevin Stone, Yixin Lin, Igor Mordatch, Pieter Abbeel, and Aravind Rajeswaran. Masked trajectory models for prediction, representation, and control. International Conference on Machine Learning, 2023.
  93. 93.Ge Yan, Kris Wu, and Xiaolong Wang. ucsd kitchens Dataset. August 2023.
  94. 94.Jonathan Heewon Yang, Dorsa Sadigh, and Chelsea Finn. Polybot: Training one policy across robots while embracing variability. In 7th Annual Conference on Robot Learning, 2023. URL https://openreview.net/forum?id=HEIRj51lcS.
  95. 95.Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2636–2645, 2020.
  96. 96.Tianhe Yu, Ted Xiao, Austin Stone, Jonathan Tompson, Anthony Brohan, Su Wang, Jaspiar Singh, Clayton Tan, Jodilyn Peralta, Brian Ichter, et al. Scaling robot learning with semantically imagined experience. arXiv preprint arXiv:2302.11550, 2023.
  97. 97.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12104–12113, 2022.
  98. 98.Tony Z Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023.
  99. 99.Gaoyue Zhou, Victoria Dean, Mohan Kumar Srirama, Aravind Rajeswaran, Jyothish Pari, Kyle Hatch, Aryan Jain, Tianhe Yu, Pieter Abbeel, Lerrel Pinto, Chelsea Finn, and Abhinav Gupta. Train offline, test online: A real robot learning benchmark, 2023.
  100. 100.Xinghao Zhu, Ran Tian, Chenfeng Xu, Mingyu Ding, Wei Zhan, and Masayoshi Tomizuka. Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot. 2023.
  101. 101.Yifeng Zhu, Peter Stone, and Yuke Zhu. Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation. IEEE Robotics and Automation Letters, 7(2):4126–4133, 2022.
  102. 102.Yifeng Zhu, Abhishek Joshi, Peter Stone, and Yuke Zhu. Viola: Imitation learning for vision-based manipulation with object proposal priors, 2023.
  103. 103.Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In 7th Annual Conference on Robot Learning, 2023.

Citation

MLA
Team, O. M., et al. “Octo: An Open-Source Generalist Robot Policy”. arXiv, 2024, http://arxiv.org/abs/2405.12213v2.
APA
Team, O. M., Ghosh, D., Walke, H., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Kreiman, T., Xu, C., Luo, J., Tan, Y. L., Chen, L. Y., Sanketi, P., Vuong, Q., Xiao, T., Sadigh, D., Finn, C., & Levine, S. (2024). Octo: An Open-Source Generalist Robot Policy. arXiv. http://arxiv.org/abs/2405.12213v2
Chicago
Team, O. M., D. Ghosh, H. Walke, et al. 2024. “Octo: An Open-Source Generalist Robot Policy”. arXiv. http://arxiv.org/abs/2405.12213v2.
Harvard
Team, O.M. et al. (2024) “Octo: An Open-Source Generalist Robot Policy”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.12213v2.
Vancouver
1. Team OM, Ghosh D, Walke H, et al (2024) Octo: An Open-Source Generalist Robot Policy. arXiv

BibTeX

@article{team2024octo,
  title = {Octo: An Open-Source Generalist Robot Policy},
  author = {Team, Octo Model and Ghosh, Dibya and Walke, Homer and Pertsch, Karl and Black, Kevin and Mees, Oier and Dasari, Sudeep and Hejna, Joey and Kreiman, Tobias and Xu, Charles and Luo, Jianlan and Tan, You Liang and Chen, Lawrence Yunliang and Sanketi, Pannag and Vuong, Quan and Xiao, Ted and Sadigh, Dorsa and Finn, Chelsea and Levine, Sergey},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.12213v2},
  eprint = {2405.12213}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/