SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
Kiana EhsaniTanmay GuptaRose HendrixJordi SalvadorLuca WeihsKuo-Hao ZengKunal Pratap SinghYejin KimWinson HanAlvaro Herrasti
Demonstrates that imitating simulated shortest-path heuristic planners across thousands of procedurally generated environments trains end-to-end RGB-only transformer agents that transfer directly to real-world mobile manipulation tasks without requiring reinforcement learning or costly human demonstrations.
Building robotic systems capable of navigating, exploring, and manipulating objects inside everyday home environments remains a major challenge. Traditional approaches rely heavily on reinforcement learning, which requires extensive reward shaping and is computationally slow for complex tasks, or imitation learning using human-collected demonstrations, which is prohibitively expensive to scale. Furthermore, many current methods depend on specialized sensors such as depth maps and GPS coordinates, or assume pre-existing maps of the physical environment.
The article demonstrates that training an embodied robotic agent to clone automated, shortest-path heuristic planners within diverse simulated environments produces effective navigation and manipulation capabilities in both simulation and the physical world using only standard camera images (RGB sensors). The primary objective is to evaluate whether massive procedural data scale and modern transformer architectures can overcome the historical performance limitations of simulated imitation learning without requiring human demonstrations, reinforcement learning, or explicit mapping modules.
To test this approach, the authors developed SPOC (Shortest Path Oracle Clone), an end-to-end transformer-based system embodied in a mobile robot. The model was trained across roughly 200,000 procedurally generated simulated houses populated with over 41,000 unique 3D household objects across 863 categories. The training utilized automated shortest-path navigation and heuristic manipulation planners acting on privileged simulation data. The resulting agents were evaluated on CHORES, a new multi-task evaluation suite covering object navigation, room visitation, and object fetching, as well as CHORESNAV for open-vocabulary language following, followed by real-world physical robot deployments across 88 trials.
The analysis yielded several critical findings. First, SPOC achieved a 49.9% multi-task success rate in unseen simulated environments, outperforming standard reinforcement learning baselines by roughly 30 percentage points while training at twenty times the computational speed. Second, despite being trained solely on shortest-path trajectories, the agent naturally exhibited complex behaviors such as exploring unknown rooms, backtracking, and obstacle avoidance. Third, dataset diversity proved essential: training across 10,000 unique simulated houses improved navigation success by 13.5 percentage points compared to training on 100 houses with identical total episodes. Fourth, modern vision backbones (specifically SigLIP) and longer transformer context windows substantially improved success rates over older CLIP and recurrent architectures. Finally, the simulated model transferred directly to physical robots without fine-tuning, achieving a 56.1% average success rate when paired with an off-the-shelf object detector.
These findings indicate that scaling synthetic training data in procedural simulators offers a practical, highly cost-effective path to developing autonomous robots. The results challenge the assumption that expensive real-world human demonstrations or complex reward engineering are necessary for robust exploration and mobile manipulation. Furthermore, error analysis showed that robot failures stemmed primarily from visual object detection errors rather than navigational planning failures, as providing ground-truth object detection raised navigation success to roughly 85%.
Organizations developing mobile robotics should prioritize scaling simulated procedural diversity and integrating high-capacity vision encoders rather than investing disproportionately in real-world demonstration collection. Further work should focus on strengthening zero-shot object detection and fine-grained physical grasping, which represent the main performance bottlenecks during real-world execution. While the current results are highly promising for indoor navigation and basic pick-and-place tasks, readers should exercise caution when applying these findings to highly cluttered or dynamic environments where physical grasping and contact mechanics require higher precision.
- Paper: Habitat: A Platform for Embodied AI Research, Manolis Savva et al. (2019). Introduces the Habitat simulation framework and benchmarks for scalable embodied AI and shortest-path navigation that directly underpin SPOC's simulation platform and methodologies.
- Paper: Target-driven visual navigation in indoor scenes using deep reinforcement learning, Yuke Zhu et al. (2016). Pioneers target-driven visual navigation in photo-realistic simulated 3D environments (AI2-THOR) and establishes the foundations of training visual navigation agents for real-world deployment.
- Paper: Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments, Peter Anderson et al. (2017). Establishes the standard paradigm and benchmark for vision-and-language navigation in realistic indoor environments, defining the core problem that SPOC addresses through imitation.
- Paper: RT-1: Robotics Transformer for Real-World Control at Scale, Anthony Brohan et al. (2023). Demonstrates how large-scale transformer-based sequence modeling over diverse robot trajectories enables cross-task generalization in mobile manipulation.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). Extends transformer-based robot control by mapping vision and language inputs to action tokens, providing the architectural foundation for language-conditioned embodied transformers.
- Paper: Domain randomization for transferring deep neural networks from simulation to the real world, Josh Tobin et al. (2017). Introduces domain randomization as a fundamental technique for bridging the sim-to-real gap when training visual policies purely in simulation.
- Paper: End-to-End Driving Via Conditional Imitation Learning, Felipe Codevilla et al. (2017). Presents conditional imitation learning for translating high-level directional commands and raw RGB images into low-level control policies.
- Paper: π0: A Vision-Language-Action Flow Model for General Robot Control, Kevin Black et al. (2024). Advances large-scale generalist robot policies by pairing vision-language backbones with continuous flow matching action experts across multi-task mobile manipulation domains.
- Paper: π0.5: a Vision-Language-Action Model with Open-World Generalization, Physical Intelligence et al. (2025). Builds upon large-scale generalist robot imitation models by co-training vision-language-action architectures across heterogeneous web, simulation, and real-home mobile manipulation data.
- Paper: GR00T N1: An Open Foundation Model for Generalist Humanoid Robots, NVIDIA et al. (2025). Extends simulation-to-real visual imitation by integrating large-scale synthetic physics simulations, human video, and diffusion action transformers for generalist humanoid control.
- Paper: Octo: An Open-Source Generalist Robot Policy, O. Team et al. (2024). Provides an open-source transformer-based generalist policy architecture that operates across diverse camera streams and action spaces, extending generalist robot imitation methodologies.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). Enhances language-conditioned vision-action policies by introducing visual chain-of-thought subgoal generation before predicting continuous actions.
- Paper: UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, Jianke Zhang et al. (2025). Unifies semantic reasoning and low-level physical dynamics in vision-language-action agents by incorporating autoregressive visual future prediction into policy learning.
- Paper: RoboDreamer: Learning Compositional World Models for Robot Imagination, Siyuan Zhou et al. (2024). Applies compositional video world models to break down language instructions and generate visual plans for closed-loop downstream robot control.
