BAKU: An Efficient Transformer for Multi-Task Policy Learning
Siddhant HaldarZhuoran PengLerrel Pinto
Presents BAKU, an efficient transformer architecture for multi-task robot imitation learning that combines multi-sensory observation trunks with action chunking to achieve a 91% success rate on 30 real-world tasks using only 17 demonstrations per task.
Developing generalist robots capable of mastering diverse manipulation tasks remains a major challenge due to the high cost and time required to collect real-world physical demonstration data. Prior approaches often compensate for poor multi-task learning efficiency by gathering massive amounts of teleoperated data, yet the resulting models frequently lag behind specialized, single-task systems. The article addresses this bottleneck by evaluating whether a streamlined model architecture can achieve superior multi-task performance using substantially less demonstration data.
The main objective of the article is to introduce and demonstrate BAKU, a compact, modular multi-task policy learning architecture. BAKU systematically integrates task-conditioned sensory encoders, a causal transformer observation trunk, and interchangeable action prediction heads to generate smooth, multi-step actions from multi-sensory inputs such as images, robot state, and text instructions.
The authors evaluated the framework across both simulated and physical benchmarks. The testing scope included 129 simulated tasks across three benchmark suites (90 manipulation tasks in LIBERO-90, 30 manipulation tasks in Meta-World, and 9 locomotion tasks in DeepMind Control) and 30 real-world robotic manipulation tasks on an xArm manipulator in a kitchen setup. The evaluations compared BAKU against established multi-task architectures, such as RT-1 and MT-ACT, under limited data regimes, including as few as 17 demonstrations per task on the physical robot.
The key findings demonstrate significant performance gains and data efficiency across all testing domains. Across 129 simulated tasks, BAKU achieved an overall 18% absolute improvement over leading baseline methods, reaching a state-of-the-art 90% success rate on the challenging LIBERO-90 benchmark—a 36% absolute improvement. In real-world physical evaluations across 30 manipulation tasks, BAKU achieved an 86% success rate with an MLP action head and a 91% success rate when paired with a vector-quantized multimodal action head, outperforming the strongest baseline by 35%. On multi-step, long-horizon tasks, the framework exceeded baseline performance by 19% on average. Furthermore, ablation analyses revealed that a compact 10-million-parameter model outperformed a larger 114-million-parameter variant, and that combining action chunking with temporal smoothing was critical for high-precision manipulation.
These results demonstrate that architecture design—specifically modular encoding, action chunking, and multimodal action heads—can drastically lower data collection requirements and hardware costs. By achieving high reliability from fewer than 20 demonstrations per task, the framework reduces the risk and timeline associated with deploying generalist robots in real-world operational environments. The findings also suggest that scaling model parameter size indiscriminately can degrade performance through overfitting when training on limited demonstration datasets.
Based on these findings, teams developing robot control systems should prioritize modular transformer architectures that decouple observation encoding from action generation and leverage action chunking for smooth physical execution. For physical deployments with varied human demonstrations, practitioners should opt for multimodal action heads such as vector-quantized transformers over basic unimodal regressors. Further research and pilot testing should focus on skill chaining to execute extended sequences and developing specialized techniques for high-precision sub-skills, such as fine-tolerance door opening and narrow object insertions.
While the findings demonstrate high confidence and consistent reproducibility across random seeds, certain limitations remain. The evaluations primarily tested short single skills or two-step sequential chains within controlled lab environments, and the article did not assess zero-shot generalization to unseen task distributions. Stakeholders should therefore exercise caution when deploying the framework in unconstrained or safety-critical operational settings without task-specific validation.
- Paper: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware, Tony Z. Zhao et al. (2023). ACT introduces the action-chunking and temporal-ensembling formulation that BAKU explicitly incorporates into its efficient transformer policy design.
- Paper: Octo: An Open-Source Generalist Robot Policy, O. Team et al. (2024). Octo establishes the multimodal, transformer-based generalist-policy framework whose modular observations and action decoding BAKU adapts and simplifies.
- Paper: RT-1: Robotics Transformer for Real-World Control at Scale, Anthony Brohan et al. (2023). RT-1 provides the central large-scale transformer baseline and multi-task robot-policy setting against which BAKU measures its improvements.
- Paper: Decision Transformer: Reinforcement Learning via Sequence Modeling, Lili Chen et al. (2021). Decision Transformer supplies the offline sequence-modeling perspective that motivates treating demonstrated robot trajectories as transformer-predicted action sequences.
- Paper: π0.5: a Vision-Language-Action Model with Open-World Generalization, Physical Intelligence et al. (2025). π0.5 extends BAKU’s efficient generalist-policy agenda toward open-world household manipulation and longer-horizon transfer across unseen environments.
- Paper: RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation, Songming Liu et al. (2025). RDT-1B continues BAKU’s multi-task transformer direction by scaling diffusion-based policy learning to large heterogeneous corpora and bimanual manipulation.
- Paper: GR00T N1: An Open Foundation Model for Generalist Humanoid Robots, NVIDIA et al. (2025). GR00T N1 generalizes BAKU’s sample-efficient multi-task policy concept into a foundation model for humanoid robots and multiple embodiments.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). DynamicVLA extends BAKU-style vision-language-action control to moving objects by addressing latency, anticipation, and continuously updated action execution.
- Paper: TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers, Bin Yu et al. (2026). TwinBrainVLA continues the generalist-policy problem by preserving broad vision-language knowledge while adapting a trainable pathway for robot action generation.
