BAKU: An Efficient Transformer for Multi-Task Policy Learning

Siddhant HaldarZhuoran PengLerrel Pinto

article2024NeurIPS87 citations

Presents BAKU, an efficient transformer architecture for multi-task robot imitation learning that combines multi-sensory observation trunks with action chunking to achieve a 91% success rate on 30 real-world tasks using only 17 demonstrations per task.

Listen

Developing generalist robots capable of mastering diverse manipulation tasks remains a major challenge due to the high cost and time required to collect real-world physical demonstration data. Prior approaches often compensate for poor multi-task learning efficiency by gathering massive amounts of teleoperated data, yet the resulting models frequently lag behind specialized, single-task systems. The article addresses this bottleneck by evaluating whether a streamlined model architecture can achieve superior multi-task performance using substantially less demonstration data.

The main objective of the article is to introduce and demonstrate BAKU, a compact, modular multi-task policy learning architecture. BAKU systematically integrates task-conditioned sensory encoders, a causal transformer observation trunk, and interchangeable action prediction heads to generate smooth, multi-step actions from multi-sensory inputs such as images, robot state, and text instructions.

The authors evaluated the framework across both simulated and physical benchmarks. The testing scope included 129 simulated tasks across three benchmark suites (90 manipulation tasks in LIBERO-90, 30 manipulation tasks in Meta-World, and 9 locomotion tasks in DeepMind Control) and 30 real-world robotic manipulation tasks on an xArm manipulator in a kitchen setup. The evaluations compared BAKU against established multi-task architectures, such as RT-1 and MT-ACT, under limited data regimes, including as few as 17 demonstrations per task on the physical robot.

The key findings demonstrate significant performance gains and data efficiency across all testing domains. Across 129 simulated tasks, BAKU achieved an overall 18% absolute improvement over leading baseline methods, reaching a state-of-the-art 90% success rate on the challenging LIBERO-90 benchmark—a 36% absolute improvement. In real-world physical evaluations across 30 manipulation tasks, BAKU achieved an 86% success rate with an MLP action head and a 91% success rate when paired with a vector-quantized multimodal action head, outperforming the strongest baseline by 35%. On multi-step, long-horizon tasks, the framework exceeded baseline performance by 19% on average. Furthermore, ablation analyses revealed that a compact 10-million-parameter model outperformed a larger 114-million-parameter variant, and that combining action chunking with temporal smoothing was critical for high-precision manipulation.

These results demonstrate that architecture design—specifically modular encoding, action chunking, and multimodal action heads—can drastically lower data collection requirements and hardware costs. By achieving high reliability from fewer than 20 demonstrations per task, the framework reduces the risk and timeline associated with deploying generalist robots in real-world operational environments. The findings also suggest that scaling model parameter size indiscriminately can degrade performance through overfitting when training on limited demonstration datasets.

Based on these findings, teams developing robot control systems should prioritize modular transformer architectures that decouple observation encoding from action generation and leverage action chunking for smooth physical execution. For physical deployments with varied human demonstrations, practitioners should opt for multimodal action heads such as vector-quantized transformers over basic unimodal regressors. Further research and pilot testing should focus on skill chaining to execute extended sequences and developing specialized techniques for high-precision sub-skills, such as fine-tolerance door opening and narrow object insertions.

While the findings demonstrate high confidence and consistent reproducibility across random seeds, certain limitations remain. The evaluations primarily tested short single skills or two-step sequential chains within controlled lab environments, and the article did not assess zero-shot generalization to unseen task distributions. Stakeholders should therefore exercise caution when deploying the framework in unconstrained or safety-critical operational settings without task-specific validation.

Cover for BAKU: An Efficient Transformer for Multi-Task Policy Learning

Abstract

Training generalist agents capable of solving diverse tasks is challenging, often requiring large datasets of expert demonstrations. This is particularly problematic in robotics, where each data point requires physical execution of actions in the real world. Thus, there is a pressing need for architectures that can effectively leverage the available training data. In this work, we present BAKU, a simple transformer architecture that enables efficient learning of multi-task robot policies. BAKU builds upon recent advancements in offline imitation learning and meticulously combines observation trunks, action chunking, multi-sensory observations, and action heads to substantially improve upon prior work. Our experiments on 129 simulated tasks across LIBERO, Meta-World suite, and the Deepmind Control suite exhibit an overall 18% absolute improvement over RT-1 and MT-ACT, with a 36% improvement on the harder LIBERO benchmark. On 30 real-world manipulation tasks, given an average of just 17 demonstrations per task, BAKU achieves a 91% success rate. Videos of the robot are best viewed at baku-robot.github.io.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 BAKU
  • 3.1 Sensory Encoders
  • 3.2 Observation Trunk
  • 3.3 Action Head
  • 3.4 Putting it all together
  • 4 Experiments
  • 4.1 How well does BAKU work for multi-task learning?
  • 4.2 How does BAKU perform on real-world tasks?
  • 4.3 How does BAKU perform on long-horizon tasks?
  • 4.4 What design decisions affect multi-task policy learning?
  • 5 Related Work
  • 6 Conclusion and Limitations
  • Acknowledgments and Disclosure of Funding
  • References
  • A Algorithmic Details
  • A.1 FiLM Conditioning
  • A.2 Action Heads
  • A.3 Temporal smoothing over action chunking
  • A.4 Hyperparameters
  • B Simulation Tasks
  • C Robot Tasks
  • D Additional Results and Analysis
  • D.1 Real-World Task-wise Results
  • D.2 Additional Analysis
  • E Broader Impacts
  • NeurIPS Paper Checklist

Knowls

  1. Knowl 1 — BAKU Policy Architecture for Multi-Task Imitation Learning

    model/method

    BAKU is a multi-task robot policy learning architecture that structures imitation learning into three decoupled components:

    1. Sensory Encoders: Modality-specific encoders process multi-sensory inputs—including multiple RGB camera viewpoints, robot proprioceptive state vectors, and task goals (specified via natural language instructions, goal images, or intermediate goal images)—and project all representations to a common embedding dimension d=256d = 256.
    2. Observation Trunk: A causal transformer decoder (an 8-layer, 4-head decoder network based on minGPT) that receives the sensory embeddings as individual observation tokens along with an appended learnable action token. In the base configuration, the trunk processes only the current timestep's observation tokens (tt) to generate an action feature representation at the action token position.
    3. Action Prediction Head: A decoupled action prediction module that takes the action feature representation from the observation trunk and predicts a multi-step continuous action chunk a^t:t+H−1\hat{a}_{t:t+H-1}.

    The total model contains approximately 10M10\text{M} parameters (sensory encoders ≈2.1M\approx 2.1\text{M}, observation trunk ≈6.5M\approx 6.5\text{M}, and action head ≈1.4M\approx 1.4\text{M}). By separating the observation trunk from the action prediction head, the architecture can swap between unimodal regression heads and multimodal generative heads without altering the observation representation.

  2. Knowl 2 — Sensory Encoding and FiLM Conditioning in BAKU

    model/method

    In BAKU, raw sensory inputs from different modalities are encoded into unified vector spaces before being passed to the observation trunk:

    • Visual Streams: RGB images are processed using a ResNet-18 convolutional backbone shared across all camera viewpoints. When conditioned on text instructions, the visual backbone integrates Feature-wise Linear Modulation (FiLM) layers.
    • Language Instructions: Natural language task descriptions are encoded into a conditioning vector zz using a pre-trained 6-layer MiniLM sentence transformer. For an intermediate visual activation tensor xx and text conditioning vector zz, FiLM computes scaling γ(z)\gamma(z) and shifting β(z)\beta(z) vectors via learned linear layers and applies an affine transformation:
    FiLM(x)=γ(z)⊙x+β(z)\text{FiLM}(x) = \gamma(z) \odot x + \beta(z)

    where ⊙\odot represents element-wise multiplication.

    • Proprioceptive State: Robot proprioceptive state sts_t is encoded via a 2-layer multi-layer perceptron (MLP).
    • Projection Layers: Linear MLP projection layers map visual features, proprioceptive embeddings, and instruction vectors into a shared hidden dimension (d=256d = 256), producing the set of observation tokens for the transformer trunk.
  3. Knowl 3 — Concatenated Action Chunking with Exponential Temporal Ensembling

    algorithm

    Rather than auto-regressively predicting actions step-by-step or executing predicted action chunks open-loop, BAKU predicts a chunk of HH future robot control actions as a single concatenated vector a^t:t+H−1∈RH×da\hat{a}_{t:t+H-1} \in \mathbb{R}^{H \times d_a} and applies closed-loop exponential temporal ensembling across overlapping predictions at every timestep.

    Input: Current observation oto_t, goal specification gg, action buffer AA, chunk horizon HH, decay rate mm
    Output: Executable robot action at∗a_t^*
    Query policy π(ot,g)\pi(o_t, g) to predict action chunk a^t:t+H−1=(a^t(t),a^t+1(t),…,a^t+H−1(t))\hat{a}_{t:t+H-1} = (\hat{a}_t^{(t)}, \hat{a}_{t+1}^{(t)}, \dots, \hat{a}_{t+H-1}^{(t)})
    Store a^t:t+H−1\hat{a}_{t:t+H-1} in buffer AA
    Initialize asum←0a_{\text{sum}} \leftarrow 0
    Initialize W←0W \leftarrow 0
    for i=0i = 0 to H−1H - 1 do
        if action predicted for time tt from query at time t−it - i exists in AA as a^t(t−i)\hat{a}_t^{(t-i)} then
            Compute exponential weight wi=exp⁡(−m⋅i)w_i = \exp(-m \cdot i)
            asum←asum+wi⋅a^t(t−i)a_{\text{sum}} \leftarrow a_{\text{sum}} + w_i \cdot \hat{a}_t^{(t-i)}
            W←W+wiW \leftarrow W + w_i
        end if
    end for
    at∗←asum/Wa_t^* \leftarrow a_{\text{sum}} / W
    return at∗a_t^*

    Here, w0=1w_0 = 1 is the weight assigned to the newest prediction for time tt, and predictions generated in older timesteps (t−it - i) receive decaying weights governed by rate parameter mm. This ensembling smooths trajectory execution and counteracts covariate shift without requiring retraining.

  4. Knowl 4 — Modular Action Prediction Head Formulations in BAKU

    model/method

    BAKU's decoupled action head receives the action token embedding from the observation trunk to generate the action chunk a^t:t+H−1\hat{a}_{t:t+H-1}. Five distinct action heads can be integrated into the framework:

    1. Multi-Layer Perceptron (MLP): A 2-layer dense network mapping the action embedding directly to the concatenated action chunk, trained using mean squared error (L2) loss.
    2. Gaussian Mixture Model (GMM): A 2-layer network parameterizing a mixture of 5 Gaussian distributions over continuous action chunks using a Softplus activation on variance terms.
    3. Behavior Transformer (BeT): Partitions demonstration action chunks into 64 clusters using kk-means; a discrete classification head predicts cluster probabilities using focal loss, while an offset head predicts residual vectors trained with L2 loss.
    4. Vector-Quantized Behavior Transformer (VQ-BeT): Employs a residual Vector-Quantized Variational Autoencoder (VQ-VAE) with 2 residual quantization layers (codebook size 16, latent dimension 256) to discretize the action chunk space.
    5. Diffusion Action Head: A 2-layer transformer-based diffusion module that generates action chunks by iteratively denoising Gaussian noise conditioned on the trunk observation representation.
  5. Knowl 5 — Multi-Task Simulation Performance on LIBERO-90, Meta-World, and DeepMind Control

    empirical result

    BAKU was evaluated across 129 simulated tasks from three benchmarks: LIBERO-90 (90 robotic manipulation tasks, 50 demonstrations per task), Meta-World (30 manipulation tasks, 35 demonstrations per task), and DeepMind Control (DMC) suite (9 continuous locomotion tasks, 500 demonstrations per task), with 10 evaluation rollouts per task. Baselines include Robotics Transformer 1 (RT-1) and Multi-Task Action-Chunking Transformer (MT-ACT).

    Method LIBERO-90 (90 tasks) Meta-World (30 tasks) DMC (9 tasks)
    RT-1 0.16 0.65 0.66
    MT-ACT 0.54 0.13 0.59
    BAKU (Base MLP Head) 0.90 0.79 0.70
    BAKU w/ VQ-BeT Head 0.90 0.78 0.70

    On the complex LIBERO-90 benchmark, BAKU achieves a 90% average success rate, outperforming MT-ACT by 36% absolute and RT-1 by 74% absolute. On Meta-World, BAKU achieves a 79% success rate, improving upon RT-1 by 14% absolute. Multi-seed evaluation across 3 seeds on LIBERO-90 confirms robust convergence: BAKU achieves 0.89±0.010.89 \pm 0.01, compared to 0.55±0.010.55 \pm 0.01 for MT-ACT and 0.14±0.020.14 \pm 0.02 for RT-1.

  6. Knowl 6 — Real-World Multi-Task Robotic Manipulation Performance

    empirical result

    Real-world evaluations were conducted on a physical Ufactory xArm 7 manipulator equipped with an xArm Gripper in a multi-task tabletop kitchen environment comprising 30 distinct manipulation tasks (e.g., opening oven doors, picking bottles and bowls, lifting plates, wiping boards). Policies received 4 RGB camera views (128×128128 \times 128) at 10 Hz high-level control, trained on a total of 520 teleoperated demonstrations (an average of 17 demonstrations per task). Each method was evaluated over 5 rollouts per task (150 trials total per method).

    Method Mean Successes (out of 5) Success Rate
    RT-1 1.83 0.37
    MT-ACT 2.80 0.56
    BAKU (Base MLP Head) 4.30 0.86
    BAKU w/ VQ-BeT Head 4.53 0.91

    The base BAKU policy achieves an 86% overall success rate, outperforming MT-ACT by 30% absolute and RT-1 by 49% absolute. Upgrading the MLP head to a multimodal VQ-BeT head further boosts real-world performance to 91% (a 35% absolute gain over MT-ACT).

  7. Knowl 7 — Long-Horizon Multi-Task Policy Learning Performance

    empirical result

    BAKU was evaluated on extended long-horizon sequential manipulation tasks in both simulation (LIBERO-10, comprising 10 multi-stage tasks with 50 demonstrations each) and physical robot execution (5 composite real-world tasks chaining two sub-goals sequentially, with an average of 19 demonstrations per task).

    Method LIBERO-10 (10 tasks) Real Robot (5 tasks)
    MT-ACT 0.68 0.64
    BAKU 0.86 0.84

    BAKU achieves an 86% success rate on LIBERO-10 (versus 68% for MT-ACT) and an 84% success rate on chained real-robot tasks (versus 64% for MT-ACT), demonstrating an average 19% absolute improvement on long-horizon robotic execution.

  8. Knowl 8 — Ablation Analysis of Core Architectural Components in BAKU

    empirical result

    Ablation experiments across LIBERO-90, Meta-World, and DeepMind Control (DMC) systematically evaluate the effect of varying individual components of the BAKU architecture while holding other parameters fixed:

    Category Variant LIBERO-90 Meta-World DMC
    Observation Trunk MLP 0.81 0.78 0.68
    Transformer 0.90 0.79 0.70
    Model Size 4.4M 0.85 0.78 0.68
    10M 0.90 0.79 0.70
    31M 0.87 0.81 0.70
    114M 0.19 0.81 0.68
    Action Head MLP 0.90 0.79 0.70
    GMM 0.84 0.65 0.67
    BeT 0.89 0.78 0.60
    VQ-BeT 0.90 0.78 0.70
    Diffusion 0.89 0.45 0.61
    Action Chunking Without (×\times) 0.76 0.78 0.74
    With ()✓ 0.90 0.79 0.70
    Observation History None 0.90 0.79 0.70
    With last-step loss 0.54 0.08 0.37
    With multi-step loss 0.90 0.82 0.68
    Goal Modality Text 0.90 0.79 N/A
    Goal Image 0.88 0.81 N/A
    Intermediate Image 0.91 0.80 N/A
    FiLM Conditioning Without (×\times) 0.87 0.79 N/A
    With ()✓ 0.90 0.79 N/A

    Key architectural conclusions include:

    • Action Chunking: Crucial for multi-task manipulation, providing a 14% absolute gain on LIBERO-90 (0.90 vs. 0.76).
    • Observation Trunk: Causal transformer decoder outperforms an MLP trunk by 9% on LIBERO-90.
    • Model Capacity: The 10M parameter model performs best overall; scaling to 114M parameters severely degrades LIBERO-90 performance (0.19) due to overfitting on low demonstration data.
    • Observation History: Naive history with last-step loss severely degrades performance (0.54 on LIBERO-90, 0.08 on Meta-World); multi-step supervision recovers performance, but using no observation history achieves equivalent success while reducing input complexity.
    • FiLM Conditioning: Improves language-guided manipulation accuracy on LIBERO-90 from 0.87 to 0.90.
  9. Knowl 9 — Data Efficiency of Multi-Task Policy Learning

    empirical result

    The sample efficiency of BAKU relative to RT-1 and MT-ACT was analyzed across varying quantities of demonstration trajectories per task on LIBERO-90 (5, 10, 25, 50 demonstrations per task) and Meta-World (5, 10, 25, 35 demonstrations per task):

    Benchmark Demonstrations / Task RT-1 MT-ACT BAKU
    LIBERO-90 5 0.00 0.31 0.58
    10 0.01 0.48 0.71
    25 0.04 0.49 0.83
    50 0.16 0.54 0.90
    Meta-World 5 0.40 0.07 0.59
    10 0.49 0.10 0.67
    25 0.62 0.11 0.76
    35 0.65 0.13 0.79

    BAKU maintains superior data efficiency across all demonstration regimes. With only 5 demonstrations per task on LIBERO-90, BAKU achieves a 58% success rate, exceeding MT-ACT trained on the full 50-demonstration dataset (54%) and RT-1 (0%).

  10. Knowl 10 — Vision Encoder Parameter Sharing and Sensory Tokenization Strategies

    empirical result

    Analysis of multi-camera encoding architectures and observation tokenization strategies reveals:

    Category Variant LIBERO-90 Meta-World DMC
    Vision Encoders Shared (Common) 0.90 – –
    View-Specific (Separate) 0.92 – –
    Observation Trunk Input Individual Tokens 0.90 0.79 0.70
    Concatenated Vector 0.87 0.79 0.70
    • Encoder Sharing: Using separate ResNet-18 vision backbones for each camera view yields a 2% improvement on LIBERO-90 (0.92 vs. 0.90) but incurs a 15% parameter overhead per additional camera view (1.5M parameters per encoder in a 10M parameter model). A shared vision encoder is chosen for computational compactness and inference speed, especially for real-world setups with 4 camera views.
    • Tokenization Structure: Passing encoded visual views and proprioception as distinct individual tokens into the transformer observation trunk outperforms concatenating them into a single joint vector by 3% on LIBERO-90 (0.90 vs. 0.87).
  11. Knowl 11 — Limitations of BAKU Multi-Task Policy Learning

    limitation

    The authors identify three principal limitations in the BAKU multi-task policy learning framework:

    1. Negative Transfer on High-Precision Manipulation: Although BAKU achieves high average success across multi-task benchmarks, it exhibits performance degradation on fine-grained, high-precision sub-skills (e.g., opening oven doors or inserting narrow tea bottles into refrigerator door slots). Joint training across diverse tasks of varying difficulty can cause shared representations to suboptimal for tasks requiring high precision.
    2. Absence of Autonomous Dynamic Skill Chaining: The system executes individual skills or pre-determined sequential chains; it does not incorporate autonomous high-level task planning or dynamic skill composition for uncurated long-horizon objectives.
    3. Evaluation of Unseen Task Generalization: The study focuses on multi-task imitation efficiency and in-distribution policy performance rather than zero-shot generalization capabilities to entirely new task distributions as task diversity scales.

Coverage note — Detailed per-task names and individual trial outcome tables for all 30 real-world tasks and 129 simulation tasks were omitted in favor of comprehensive domain-level summaries, aggregate performance metrics, and explicit ablation tables.

References

  1. 1.P. Abbeel and A. Y. Ng. Apprenticeship learning via inverse reinforcement learning. In Proceedings of the twenty-first international conference on Machine learning, page 1, 2004. 10
  2. 2.J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 1
  3. 3.M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691, 2022. 10
  4. 4.J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023. 1
  5. 5.H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V. Kumar. Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking. arXiv preprint arXiv:2309.01918, 2023. 1, 3, 4, 5, 6, 10, 16, 17
  6. 6.A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022. 1, 3, 6, 10, 17
  7. 7.L. Chen, S. Bahl, and D. Pathak. Playfusion: Skill acquisition via diffusion from language-annotated play. In Conference on Robot Learning, pages 2012–2029. PMLR, 2023. 10
  8. 8.X. Chen, S. Xie, and K. He. An empirical study of training self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9640–9649, 2021. 10
  9. 9.Y. Chen, C. Wang, L. Fei-Fei, and C. K. Liu. Sequential dexterity: Chaining dexterous policies for long-horizon manipulation. arXiv preprint arXiv:2309.00987, 2023. 10
  10. 10.C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion. In Proceedings of Robotics: Science and Systems (RSS), 2023. 2, 3, 4, 8, 10, 16
  11. 11.Z. J. Cui, Y. Wang, N. M. M. Shafiullah, and L. Pinto. From play to policy: Conditional behavior generation from uncurated robot data. arXiv preprint arXiv:2210.10047, 2022. 3, 10
  12. 12.H.-S. Fang, C. Wang, H. Fang, M. Gou, J. Liu, H. Yan, W. Liu, Y. Xie, and C. Lu. Anygrasp: Robust and efficient grasp perception in spatial and temporal domains. IEEE Transactions on Robotics, 2023. 10
  13. 13.P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson. Implicit behavioral cloning. In Conference on Robot Learning, pages 158–168. PMLR, 2022. 10
  14. 14.A. Gupta, A. Murali, D. P. Gandhi, and L. Pinto. Robot learning in homes: Improving generalization and reducing dataset bias. Advances in neural information processing systems, 31, 2018. 10
  15. 15.A. Gupta, L. Fan, S. Ganguli, and L. Fei-Fei. Metamorph: Learning universal controllers with transformers. arXiv preprint arXiv:2203.11931, 2022. 10
  16. 16.S. Haldar and L. Pinto. Polytask: Learning unified policies through behavior distillation. arXiv preprint arXiv:2310.08573, 2023. 10
  17. 17.S. Haldar, V. Mathur, D. Yarats, and L. Pinto. Watch and match: Supercharging imitation with regularized optimal transport. In Conference on Robot Learning, pages 32–43. PMLR, 2023. 5, 10
  18. 18.S. Haldar, J. Pari, A. Rai, and L. Pinto. Teach a robot to fish: Versatile imitation from one minute of demonstrations. arXiv preprint arXiv:2303.01497, 2023. 5, 10
  19. 19.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 3
  20. 20.K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000–16009, 2022. 10
  21. 21.M. Heo, Y. Lee, D. Lee, and J. J. Lim. Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation. arXiv preprint arXiv:2305.12821, 2023. 10
  22. 22.D.-A. Huang, Y.-W. Chao, C. Paxton, X. Deng, L. Fei-Fei, J. C. Niebles, A. Garg, and D. Fox. Motion reasoning for goal-based imitation learning. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 4878–4884. IEEE, 2020. 10
  23. 23.A. Hussein, M. M. Gaber, E. Elyan, and C. Jayne. Imitation learning: A survey of learning methods. ACM Computing Surveys (CSUR), 50(2):1–35, 2017. 10
  24. 24.A. Iyer, Z. Peng, Y. Dai, I. Guzey, S. Haldar, S. Chintala, and L. Pinto. Open teach: A versatile teleoperation system for robotic manipulation. arXiv preprint arXiv:2403.07870, 2024. 3, 5
  25. 25.V. Jain, M. Attarian, N. J. Joshi, A. Wahid, D. Driess, Q. Vuong, P. R. Sanketi, P. Sermanet, S. Welker, C. Chan, I. Gilitschenski, Y. Bisk, and D. Dwibedi. Vid2robot: End-to-end video-conditioned policy learning with cross-attention transformers, 2024. 10
  26. 26.E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn. Bc-z: Zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning, pages 991–1002. PMLR, 2022. 10
  27. 27.M. Janner, Q. Li, and S. Levine. Offline reinforcement learning as one big sequence modeling problem. Advances in neural information processing systems, 34:1273–1286, 2021. 10
  28. 28.T. Jurgenson, O. Avner, E. Groshev, and A. Tamar. Sub-goal trees a framework for goal-based reinforcement learning. In International conference on machine learning, pages 5020–5030. PMLR, 2020. 10
  29. 29.A. Karpathy. mingpt: A minimal pytorch re-implementation of the openai gpt. https://github.com/karpathy/minGPT, 2021. 4, 18
  30. 30.A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. arXiv preprint arXiv:2403.12945, 2024. 10
  31. 31.S. Lee, Y. Wang, H. Etukuru, H. J. Kim, N. M. M. Shafiullah, and L. Pinto. Behavior generation with latent actions. arXiv preprint arXiv:2403.03181, 2024. 2, 3, 4, 6, 8, 10, 16
  32. 32.I. Lenz, H. Lee, and A. Saxena. Deep learning for detecting robotic grasps. The International Journal of Robotics Research, 34(4-5):705–724, 2015. 10
  33. 33.T. Lin, Y. Zhang, Q. Li, H. Qi, B. Yi, S. Levine, and J. Malik. Learning visuotactile skills with two multifingered hands. arXiv:2404.16823, 2024. 10
  34. 34.B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36, 2024. 2, 4, 5, 10, 17
  35. 35.C. Lynch and P. Sermanet. Language conditioned imitation learning over unstructured data. arXiv preprint arXiv:2005.07648, 2020. 10
  36. 36.C. Lynch, M. Khansari, T. Xiao, V. Kumar, J. Tompson, S. Levine, and P. Sermanet. Learning latent plans from play. In Conference on robot learning, pages 1113–1132. PMLR, 2020. 4, 8, 10, 16
  37. 37.A. Mandlekar, Y. Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, et al. Roboturk: A crowdsourcing platform for robotic skill learning through imitation. In Conference on Robot Learning, pages 879–893. PMLR, 2018. 1
  38. 38.A. Mandlekar, D. Xu, R. Martín-Martín, S. Savarese, and L. Fei-Fei. Learning to generalize across long-horizon tasks from human demonstrations. arXiv preprint arXiv:2003.06085, 2020. 10
  39. 39.A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín. What matters in learning from offline human demonstrations for robot manipulation. arXiv preprint arXiv:2108.03298, 2021. 2, 10
  40. 40.H. Mei, M. Bansal, and M. Walter. Listen, attend, and walk: Neural mapping of navigational instructions to action sequences. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 30, 2016. 10
  41. 41.A. Y. Ng, S. J. Russell, et al. Algorithms for inverse reinforcement learning. In Icml, volume 1, page 2, 2000. 10
  42. 42.M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 10
  43. 43.A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023. 10
  44. 44.E. Parisotto, J. L. Ba, and R. Salakhutdinov. Actor-mimic: Deep multitask and transfer reinforcement learning. arXiv preprint arXiv:1511.06342, 2015. 1, 10
  45. 45.T. Pearce, T. Rashid, A. Kanervisto, D. Bignell, M. Sun, R. Georgescu, S. V. Macua, S. Z. Tan, I. Momennejad, K. Hofmann, et al. Imitating human behaviour with diffusion models. arXiv preprint arXiv:2301.10677, 2023. 4, 10, 16
  46. 46.E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018. 1, 4, 8, 16
  47. 47.L. Pinto and A. Gupta. Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours. In 2016 IEEE international conference on robotics and automation (ICRA), pages 3406–3413. IEEE, 2016. 10
  48. 48.D. Pomerleau. An autonomous land vehicle in a neural network. Advances in Neural Information Processing Systems, 1, 1998. 10
  49. 49.I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath. Learning humanoid locomotion with transformers. arXiv e-prints, pages arXiv–2303, 2023. 10
  50. 50.I. Radosavovic, B. Zhang, B. Shi, J. Rajasegaran, S. Kamat, T. Darrell, K. Sreenath, and J. Malik. Humanoid locomotion as next token prediction. arXiv preprint arXiv:2402.19469, 2024. 10
  51. 51.A. Raffin, A. Hill, R. Traoré, T. Lesort, N. Díaz-Rodríguez, and D. Filliat. Decoupling feature extraction from policy learning: assessing benefits of state representation learning in goal based robotics. arXiv preprint arXiv:1901.08651, 2019. 10
  52. 52.S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al. A generalist agent. arXiv preprint arXiv:2205.06175, 2022. 10
  53. 53.M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024. 1
  54. 54.N. Reimers and I. Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 11 2019. URL https://arxiv.org/abs/1908.10084. 4
  55. 55.M. Reuss, M. Li, X. Jia, and R. Lioutikov. Goal-conditioned imitation learning using score-based diffusion policies. arXiv preprint arXiv:2304.02532, 2023. 4, 10, 16
  56. 56.S. Ross, G. Gordon, and D. Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pages 627–635. JMLR Workshop and Conference Proceedings, 2011. 10
  57. 57.A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell. Policy distillation. arXiv preprint arXiv:1511.06295, 2015. 1, 10
  58. 58.M. Ryoo, A. Piergiovanni, A. Arnab, M. Dehghani, and A. Angelova. Tokenlearner: Adaptive space-time tokenization for videos. Advances in neural information processing systems, 34: 12786–12797, 2021. 6
  59. 59.C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems, 35: 36479–36494, 2022. 1
  60. 60.N. M. Shafiullah, Z. Cui, A. A. Altanzaya, and L. Pinto. Behavior transformers: Cloning k modes with one stone. Advances in neural information processing systems, 35:22955–22968, 2022. 2, 4, 8, 10, 16
  61. 61.N. M. M. Shafiullah, A. Rai, H. Etukuru, Y. Liu, I. Misra, S. Chintala, and L. Pinto. On bringing robots home. arXiv preprint arXiv:2311.16098, 2023. 1, 10
  62. 62.D. Shah, A. Sridhar, N. Dashora, K. Stachowicz, K. Black, N. Hirose, and S. Levine. Vint: A foundation model for visual navigation. arXiv preprint arXiv:2306.14846, 2023. 10
  63. 63.M. Shridhar, L. Manuelli, and D. Fox. Perceiver-actor: A multi-task transformer for robotic manipulation. In Conference on Robot Learning, pages 785–799. PMLR, 2023. 3
  64. 64.A. Sridhar, D. Shah, C. Glossop, and S. Levine. Nomad: Goal masked diffusion policies for navigation and exploration. arXiv preprint arXiv:2310.07896, 2023. 10
  65. 65.K. Sridhar, S. Dutta, D. Jayaraman, J. Weimer, and I. Lee. Memory-consistent neural networks for imitation learning. arXiv preprint arXiv:2310.06171, 2023. 10
  66. 66.S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor. Language-conditioned imitation learning for robot manipulation tasks. Advances in Neural Information Processing Systems, 33:13139–13150, 2020. 10
  67. 67.Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al. Deepmind control suite. arXiv preprint arXiv:1801.00690, 2018. 2, 5, 17
  68. 68.F. Torabi, G. Warnell, and P. Stone. Recent advances in imitation learning from observation. arXiv preprint arXiv:1905.13566, 2019. 2, 10
  69. 69.H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 1
  70. 70.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 4
  71. 71.U. Viereck, A. Pas, K. Saenko, and R. Platt. Learning a visuomotor controller for real world robotic grasping using simulated depth images. In Conference on robot learning, pages 291–300. PMLR, 2017. 10
  72. 72.C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y. Zhu, and A. Anandkumar. Mimicplay: Long-horizon imitation learning by watching human play. In 7th Annual Conference on Robot Learning, 2023. 3, 9, 10
  73. 73.W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers, 2020. 4
  74. 74.D. Yarats, R. Fergus, A. Lazaric, and L. Pinto. Mastering visual continuous control: Improved data-augmented reinforcement learning. arXiv preprint arXiv:2107.09645, 2021. 5
  75. 75.T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn. Gradient surgery for multi-task learning. Advances in Neural Information Processing Systems, 33:5824–5836, 2020. 1, 10
  76. 76.T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on robot learning, pages 1094–1100. PMLR, 2020. 2, 5, 17
  77. 77.T. Z. Zhao, V. Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023. 2, 3, 4, 5, 10, 16, 17
  78. 78.B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In Conference on Robot Learning, pages 2165–2183. PMLR, 2023. 10

Citation

MLA
Haldar, S., et al. “BAKU: An Efficient Transformer for Multi-Task Policy Learning”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 141208–39, https://proceedings.neurips.cc/paper_files/paper/2024/file/ff887781480973bd3cb6026feb378d1e-Paper-Conference.pdf.
APA
Haldar, S., Peng, Z., & Pinto, L. (2024). BAKU: An Efficient Transformer for Multi-Task Policy Learning. Advances in Neural Information Processing Systems, 37, 141208–141239. https://proceedings.neurips.cc/paper_files/paper/2024/file/ff887781480973bd3cb6026feb378d1e-Paper-Conference.pdf
Chicago
Haldar, S., Z. Peng, and L. Pinto. 2024. “BAKU: An Efficient Transformer for Multi-Task Policy Learning”. Advances in Neural Information Processing Systems 37: 141208–39. https://proceedings.neurips.cc/paper_files/paper/2024/file/ff887781480973bd3cb6026feb378d1e-Paper-Conference.pdf.
Harvard
Haldar, S., Peng, Z. and Pinto, L. (2024) “BAKU: An Efficient Transformer for Multi-Task Policy Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 141208–141239. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/ff887781480973bd3cb6026feb378d1e-Paper-Conference.pdf.
Vancouver
1. Haldar S, Peng Z, Pinto L (2024) BAKU: An Efficient Transformer for Multi-Task Policy Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 141208–141239

BibTeX

@inproceedings{haldar2024baku,
  title = {BAKU: An Efficient Transformer for Multi-Task Policy Learning},
  author = {Haldar, Siddhant and Peng, Zhuoran and Pinto, Lerrel},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {141208-141239},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/ff887781480973bd3cb6026feb378d1e-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors