SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World

Kiana EhsaniTanmay GuptaRose HendrixJordi SalvadorLuca WeihsKuo-Hao ZengKunal Pratap SinghYejin KimWinson HanAlvaro Herrasti

article2024CVPR70 citations

Demonstrates that imitating simulated shortest-path heuristic planners across thousands of procedurally generated environments trains end-to-end RGB-only transformer agents that transfer directly to real-world mobile manipulation tasks without requiring reinforcement learning or costly human demonstrations.

Listen

Building robotic systems capable of navigating, exploring, and manipulating objects inside everyday home environments remains a major challenge. Traditional approaches rely heavily on reinforcement learning, which requires extensive reward shaping and is computationally slow for complex tasks, or imitation learning using human-collected demonstrations, which is prohibitively expensive to scale. Furthermore, many current methods depend on specialized sensors such as depth maps and GPS coordinates, or assume pre-existing maps of the physical environment.

The article demonstrates that training an embodied robotic agent to clone automated, shortest-path heuristic planners within diverse simulated environments produces effective navigation and manipulation capabilities in both simulation and the physical world using only standard camera images (RGB sensors). The primary objective is to evaluate whether massive procedural data scale and modern transformer architectures can overcome the historical performance limitations of simulated imitation learning without requiring human demonstrations, reinforcement learning, or explicit mapping modules.

To test this approach, the authors developed SPOC (Shortest Path Oracle Clone), an end-to-end transformer-based system embodied in a mobile robot. The model was trained across roughly 200,000 procedurally generated simulated houses populated with over 41,000 unique 3D household objects across 863 categories. The training utilized automated shortest-path navigation and heuristic manipulation planners acting on privileged simulation data. The resulting agents were evaluated on CHORES, a new multi-task evaluation suite covering object navigation, room visitation, and object fetching, as well as CHORESNAV for open-vocabulary language following, followed by real-world physical robot deployments across 88 trials.

The analysis yielded several critical findings. First, SPOC achieved a 49.9% multi-task success rate in unseen simulated environments, outperforming standard reinforcement learning baselines by roughly 30 percentage points while training at twenty times the computational speed. Second, despite being trained solely on shortest-path trajectories, the agent naturally exhibited complex behaviors such as exploring unknown rooms, backtracking, and obstacle avoidance. Third, dataset diversity proved essential: training across 10,000 unique simulated houses improved navigation success by 13.5 percentage points compared to training on 100 houses with identical total episodes. Fourth, modern vision backbones (specifically SigLIP) and longer transformer context windows substantially improved success rates over older CLIP and recurrent architectures. Finally, the simulated model transferred directly to physical robots without fine-tuning, achieving a 56.1% average success rate when paired with an off-the-shelf object detector.

These findings indicate that scaling synthetic training data in procedural simulators offers a practical, highly cost-effective path to developing autonomous robots. The results challenge the assumption that expensive real-world human demonstrations or complex reward engineering are necessary for robust exploration and mobile manipulation. Furthermore, error analysis showed that robot failures stemmed primarily from visual object detection errors rather than navigational planning failures, as providing ground-truth object detection raised navigation success to roughly 85%.

Organizations developing mobile robotics should prioritize scaling simulated procedural diversity and integrating high-capacity vision encoders rather than investing disproportionately in real-world demonstration collection. Further work should focus on strengthening zero-shot object detection and fine-grained physical grasping, which represent the main performance bottlenecks during real-world execution. While the current results are highly promising for indoor navigation and basic pick-and-place tasks, readers should exercise caution when applying these findings to highly cluttered or dynamic environments where physical grasping and contact mechanics require higher precision.

Cover for SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World

Abstract

Reinforcement learning (RL) with dense rewards and imitation learning (IL) with human-generated trajectories are the most widely used approaches for training modern embodied agents. RL requires extensive reward shaping and auxiliary losses and is often too slow and ineffective for long-horizon tasks. While IL with human supervision is effective, collecting human trajectories at scale is extremely expensive. In this work, we show that imitating shortest-path planners in simulation produces agents that, given a language instruction, can proficiently navigate, explore, and manipulate objects in both simulation and in the real world using only RGB sensors (no depth map or GPS coordinates). This surprising result is enabled by our end-to-end, transformer-based, SPOC architecture, powerful visual encoders paired with extensive image augmentation, and the dramatic scale and diversity of our training data: millions of frames of shortest-path-expert trajectories collected inside approximately 200,000 procedurally generated houses containing 40,000 unique 3D assets. Our models, data, training code, and newly proposed 10-task benchmarking suite CHORES are available in spoc-robot.github.io.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Imitation in Procedural Houses
  • 4. The Shortest Path Oracle Clone (SPOC)
  • 5. Procedural Data
  • 5.1. Environments
  • 5.2. Expert Trajectories
  • 6. Benchmark
  • 7. Experiments
  • 7.1. Quantitative Analysis
  • 7.2. Agent Behavior
  • 8. Conclusion
  • References

Knowls

  1. Knowl 1 — SPOC learns navigation and manipulation by cloning shortest-path experts

    model/method

    SPOC (Shortest Path Oracle Clone) is an end-to-end embodied agent for the Stretch RE-1 mobile manipulator. At each time step, it receives an open-vocabulary language instruction and RGB images from two perpendicular cameras: one for navigation and one for arm manipulation. It predicts the next discrete robot action using behavior cloning on trajectories generated by heuristic shortest-path planners in simulation.

    SPOC does not use depth, GPS coordinates, an explicit map, a mapping module, a large language model, reinforcement learning, or human demonstrations. The expert planners use privileged simulator state only while generating training trajectories; that state is unavailable to SPOC during inference. The same learned weights transfer from simulation to physical environments without real-world fine-tuning or visual adaptation.

    The jointly trained task suite covers object-goal navigation (OBJNAV), picking up an object already in view (PICKUP), finding and then picking up an object (FETCH), and visiting every room in a house (ROOMVISIT).

  2. Knowl 2 — Goal-conditioned visual encoding and causal action prediction

    model/method

    SPOC encodes the language goal, two RGB observations, and action history with a transformer architecture. Let GG be the instruction, FnavtF^t_{nav} and FmaniptF^t_{manip} be the navigation and manipulation RGB frames at time tt, and ata^t be the discrete action at that time. A pretrained text encoder produces token representations g=Egoal(G)g = E_{goal}(G). A pretrained image encoder produces patch-feature sequences for both frames; multilayer perceptrons with ReLU and LayerNorm map the image and goal features into a shared transformer dimension. Learned camera-type embeddings distinguish navigation-camera patches from manipulation-camera patches.

    The patch features, goal-token features, and a learned [CLS] token are concatenated and processed by a transformer encoder. The [CLS] output is the goal-conditioned visual vector vt=Evisual(Fnavt,Fmanipt,G)v_t = E_{visual}(F^t_{nav}, F^t_{manip}, G). A causal transformer decoder DD receives the sequence of current and past visual vectors, sinusoidal temporal encodings, and learned embeddings of previous actions. It cross-attends to the goal representation and produces the action distribution

    πt=softmax(linear(D(v0:t,a0:t−1;g)[t])).π_t = softmax(linear(D(v_{0:t}, a_{0:t-1}; g)[t])).

    Here, v0:tv_{0:t} denotes visual vectors from time 00 through tt, a0:t−1a_{0:t-1} denotes previous actions, gg is the encoded instruction, and [t][t] selects the decoder output for the current step. Training uses teacher-forced cross-entropy against the expert action; during deployment, an action is sampled from πtπ_t and fed back at the next step.

    The action space contains 20 discretized actions, including base translation of ±20 cm, base rotations of ±6° and ±30°, arm translations in the xx and zz directions of ±2 cm and ±10 cm, grasper rotations of ±10°, pickup, dropoff, subtask completion, and episode termination. The implementation uses three-layer transformer encoder and decoder stacks, a context window of 100 during training, batch size 224, AdamW with learning rate 0.0002, and 20,000 iterations for single-task models or 50,000 for multitask models. Random temporal-index shifts during training allow inference to use the full past history rather than only the training window.

  3. Knowl 3 — Large-scale procedural training worlds and privileged expert planners

    experimental setup

    Training environments are generated with AI2-THOR and ProcTHOR. The asset collection combines approximately 40,000 household-relevant Objaverse assets with approximately 2,000 existing AI2-THOR instances, yielding 41,133 unique 3D assets across 863 WordNet object types. ProcTHOR generates approximately 200,000 procedurally varied houses containing between 1 and 8 rooms; object instances are partitioned between training and evaluation so that some object categories are evaluated zero-shot.

    The imitation trajectories are generated with simulator-only privileged information. For navigation, a planner computes a shortest path on the environment navigation mesh to a target object or coordinate and then approximately follows it with the robot's discrete movement actions; when the target is an object, the planner rotates until the object is centered in the relevant camera. For manipulation, the planner first moves to a privileged location from which the target is reachable, then uses the known agent and object poses with iterative distance minimization to bring the arm to the object and grasp it. For room visitation, the planner visits room centers in depth-first-search order and emits a subtask-completion signal after each room. The resulting data contains millions of RGB frames and averages approximately 90,000 training episodes per task.

  4. Knowl 4 — CHORES and CHORESNAV benchmark task definitions

    definition

    CHORES (Core HOusehold Robot Evaluations) evaluates joint navigation, recognition, manipulation, and exploration with four tasks:

    • OBJNAV: locate an object category, such as finding a mug.
    • PICKUP: pick up a specified object that is already in the agent's line of sight.
    • FETCH: find an object and then pick it up.
    • ROOMVISIT: traverse the house, signal when a new room has been seen, and signal when all rooms have been visited.

    CHORESNAV extends object navigation to seven open-vocabulary instruction types: OBJNAV for an object category, OBJNAVROOM for an object constrained to a room type, OBJNAVRELATTR for a comparative attribute such as the smallest vase, OBJNAVAFFORD for an intended use such as a container for flowers, OBJNAVLOCALREF for nearby-object references, OBJNAVDESC for a detailed instance description, and ROOMNAV for locating a room type. Instructions may use WordNet hypernyms, so a target such as container can be satisfied by objects such as a vase or mug.

    The benchmark reports Success, episode-length-weighted Success (SEL), and the percentage of rooms visited. CHORES-S uses 15 object categories, while CHORES-L uses the full set of 863 categories. Each evaluation task contains approximately 195 episodes on average.

  5. Knowl 5 — Shortest-path imitation substantially outperforms reinforcement learning in simulation

    empirical result

    On unseen CHORES-S environments, multitask SPOC trained by imitation learning achieves 49.9% average success, essentially matching the 50.0% achieved by a separately trained single-task SPOC model. The single-task reinforcement-learning baseline achieves only 31.2% average success despite reward shaping, a SIGLIP visual backbone, and twice the training time; it obtains 0.0% success on FETCH.

    The task-level success rates are:

    • Single-task RL: OBJNAV 36.5%, PICKUP 71.9%, FETCH 0.0%, ROOMVISIT 16.5%.
    • Single-task SPOC imitation learning: OBJNAV 57.0%, PICKUP 84.2%, FETCH 15.1%, ROOMVISIT 43.7%.
    • Multitask SPOC imitation learning: OBJNAV 55.0%, PICKUP 90.1%, FETCH 14.0%, ROOMVISIT 40.5%.
    • Multitask SPOC with simulator ground-truth target detection: OBJNAV 85.0%, PICKUP 91.2%, FETCH 47.3%, ROOMVISIT 36.7%.

    On CHORES-L, where the full 863-category vocabulary is used, ordinary multitask SPOC obtains 38.6% average success, while the ground-truth-detection variant obtains 63.1%. Thus, multitask imitation does not show the task-competition degradation that is often observed in multitask training, whereas access to perfect target detection produces a large additional improvement.

  6. Knowl 6 — Transformer history, image representation, and context length are decisive design factors

    empirical result

    Ablations on CHORES-S show that both transformer components of SPOC matter. Replacing the transformer action decoder with a GRU while retaining the transformer visual encoder reduces average success from 49.9% to 43.6%. Replacing the transformer visual encoder with a non-transformer goal-conditioned encoder while retaining the transformer decoder reduces average success to 45.1%. The largest degradation occurs on long-horizon tasks such as FETCH and ROOMVISIT.

    The image encoder also strongly affects performance. With the same SPOC architecture, average success is 26.6% using CLIP-RN50, 44.8% using DINOV2-ViT-S/14, and 49.9% using SIGLIP-ViT-B/16. The corresponding OBJNAV success rates are 19.6%, 47.5%, and 55.0%.

    Longer temporal context is especially important. With context windows of 10, 50, and 100 frames, average success is 27.8%, 39.9%, and 49.9%, respectively. FETCH success rises from 2.3% to 4.1% to 14.0%, and ROOMVISIT success rises from 18.0% to 28.0% to 40.5%. SPOC is trained on a limited window for computational efficiency but can attend to all past observations during inference.

  7. Knowl 7 — Training scale and house diversity matter more than exploratory expert trajectories

    empirical result

    OBJNAV performance improves with both the number and diversity of imitation examples. When training on 1,000, 10,000, and 100,000 episodes, OBJNAV success is 19.0%, 39.0%, and 57.0%, respectively. With 100,000 total episodes, distributing them across 10,000 houses with 10 episodes per house gives 57.0% success, compared with 43.5% when using only 100 houses with 1,000 episodes per house, an absolute improvement of 13.5 percentage points.

    Adding explicit exploration to the expert trajectories does not improve the learned policy. A model trained with an exploration-based OBJNAV planner obtains 46.5% success, compared with 57.0% for the standard shortest-path expert. This result indicates that, in the reported setting, broad procedural house diversity is more beneficial than training on expert trajectories that deliberately explore until the target becomes visible.

  8. Knowl 8 — SPOC follows open-vocabulary navigation instructions

    empirical result

    On CHORESNAV-S, SPOC achieves an average Success of 53.6% across seven instruction types. The Success and percentage of rooms visited for each type are: OBJNAV 57.5% and 55.7%; OBJNAVROOM 50.3% and 54.6%; OBJNAVRELATTR 54.6% and 62.2%; OBJNAVAFFORD 62.4% and 53.0%; OBJNAVLOCALREF 45.1% and 51.5%; OBJNAVDESC 30.6% and 49.9%; and ROOMNAV 74.5% and 48.1%.

    On the larger CHORESNAV-L vocabulary, the corresponding Success and room-coverage values are: OBJNAV 38.7% and 53.4%; OBJNAVROOM 54.2% and 55.7%; OBJNAVRELATTR 38.5% and 56.0%; OBJNAVAFFORD 43.5% and 48.0%; OBJNAVLOCALREF 44.5% and 58.7%; OBJNAVDESC 30.5% and 56.8%; and ROOMNAV 67.5% and 49.9%. The average Success is 45.3%. These results cover object recognition, room identification, affordance interpretation, comparative attributes, local references, and detailed descriptions.

  9. Knowl 9 — Simulation-trained SPOC transfers to physical environments without fine-tuning

    empirical result

    The authors evaluate two simulation-trained SPOC models in 88 physical-robot trials across two real-world settings, without changing model weights or applying real-world visual adaptation. RGB-only SPOC achieves 50.0% OBJNAV success, 46.7% PICKUP success, 11.1% FETCH success, and 50.0% ROOMVISIT success, for an average of 39.5%.

    A model trained with simulator ground-truth detection and evaluated with the DETIC object detector achieves 83.3% OBJNAV, 46.7% PICKUP, 44.4% FETCH, and 50.0% ROOMVISIT success, for an average of 56.1%. For the manipulation tasks, the parenthesized Soft Success values—66.7% for PICKUP and 33.3% for FETCH with RGB-only SPOC, and 86.7% and 44.4% with DETIC—count an attempt as successful when the gripper reaches within 6 cm of the object, regardless of whether the heuristic grasp succeeds. The similarity between these Soft Success values and simulation results indicates effective transfer of the learned control policy, while the gap between ordinary and detector-assisted results identifies visual perception as a major real-world bottleneck.

  10. Knowl 10 — Shortest-path imitation produces emergent exploration and backtracking, but perception remains the main failure source

    empirical result

    Although SPOC is trained only to imitate shortest paths to targets, it exhibits behaviors not explicitly present in those demonstrations. In qualitative trajectories, the agent searches multiple rooms, peeks into rooms, backtracks after passing candidate locations, scans multiple sofas before stopping at a laptop, skips irrelevant chairs to reach a chair in another room, and repositions around a table to reach a headset that was initially outside the arm's reachable region. In another example, it visits several fruits and then returns to the highest fruit specified by the instruction.

    The paper reports that failures appear to arise primarily from object perception rather than an inability to explore. Replacing learned target detection with simulator ground-truth detection raises average CHORES-S success from 49.9% to 65.0%, an approximately 15-percentage-point gain, and raises CHORES-L success from 38.6% to 63.1%, a 24.5-percentage-point gain. Ground-truth detection yields 85.0% OBJNAV success on CHORES-S. These results support the paper's qualified conclusion that exploration is largely learned despite shortest-path supervision, whereas imperfect object detection limits performance.

Coverage note — No substantial contributed material was omitted; only supplementary-only low-level planner implementation details, the complete asset-category visualization, and the real-world grasping heuristic's internal design were not expanded because they are not specified in the provided paper.

References

  1. 1.Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sunderhauf, Ian D. Reid, Stephen Gould, and ¨ Anton van den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 3674–3683. Computer Vision Foundation / IEEE Computer Society, 2018. 2
  2. 2.Dhruv Batra, Aaron Gokaslan, Aniruddha Kembhavi, Oleksandr Maksymets, Roozbeh Mottaghi, Manolis Savva, Alexander Toshev, and Erik Wijmans. Objectnav revisited: On evaluation of embodied agents navigating to objects. CoRR, abs/2006.13171, 2020. 2
  3. 3.Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil J. Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Lisa Lee, Tsang-Wei Edward Lee, Sergey Levine, Yao Lu, Henryk Michalewski, Igor Mordatch, Karl Pertsch, Kanishka Rao, Krista Reymann, Michael S. Ryoo, Grecia Salazar, Pannag Sanketi, Pierre Sermanet, Jaspiar Singh, Anikait Singh, Radu Soricut, Huong T. Tran, Vincent Vanhoucke, Quan Vuong, Ayzaan Wahid, Stefan Welker, Paul Wohlhart, Jialin Wu, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich. RT-2: vision-language-action models transfer web knowledge to robotic control. CoRR, abs/2307.15818, 2023. 3
  4. 4.Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alexander Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil J. Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Kuang-Huei Lee, Sergey Levine, Yao Lu, Utsav Malla, Deeksha Manjunath, Igor Mordatch, Ofir Nachum, Carolina Parada, Jodilyn Peralta, Emily Perez, Karol Pertsch, Jornell Quiambao, Kanishka Rao, Michael S. Ryoo, Grecia Salazar, Pannag R. Sanketi, Kevin Sayed, Jaspiar Singh, Sumedh Sontakke, Austin Stone, Clayton Tan, Huong T. Tran, Vincent Vanhoucke, Steve Vega, Quan Vuong, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich. RT-1: robotics transformer for real-world control at scale. In Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 2023, 2023. 2, 3, 7
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakan­tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. 5, 3
  6. 6.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. International Conference on 3D Vision (3DV), 2017. 4
  7. 7.Matthew Chang, Theophile Gervet, Mukul Khanna, Sriram Yenamandra, Dhruv Shah, So Yeon Min, Kavit Shah, Chris Paxton, Saurabh Gupta, Dhruv Batra, et al. Goat: Go to any thing. arXiv preprint arXiv:2311.06430, 2023. 2
  8. 8.Devendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta, and Ruslan Salakhutdinov. Learning to explore using active neural slam. ICLR, 2020.
  9. 9.Devendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta, and Saurabh Gupta. Neural topological slam for visual navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. 2
  10. 10.Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip W. Robinson, and Kristen Grauman. SoundSpaces: Audio-Visual Navigation in 3D Environments. In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part VI, pages 17–36. Springer, 2020. 2
  11. 11.Changan Chen, Carl Schissler, Sanchit Garg, Philip Kobernik, Alexander Clegg, Paul Calamia, Dhruv Batra, Philip W. Robinson, and Kristen Grauman. SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning. In NeurIPS, 2022. 2
  12. 12.Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 15084–15097, 2021. 3
  13. 13.Xiaoyu Chen, Jiachen Hu, Chi Jin, Lihong Li, and Liwei Wang. Understanding Domain Randomization for Sim-to-real Transfer. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. 3
  14. 14.Felipe Codevilla, Matthias Muller, Alexey Dosovitskiy, An­tonio M. Lopez, and Vladlen Koltun. End-to-end driving via conditional imitation learning. 2018 IEEE International Conference on Robotics and Automation (ICRA), pages 1–9, 2017. 3
  15. 15.Ishita Dasgupta, Christine Kaeser-Chen, Kenneth Marino, Arun Ahuja, Sheila Babayan, Felix Hill, and Rob Fergus. Collaborating with language models for embodied reasoning. ArXiv, abs/2302.00763, 2023. 3
  16. 16.Matt Deitke, Winson Han, Alvaro Herrasti, Aniruddha Kembhavi, Eric Kolve, Roozbeh Mottaghi, Jordi Salvador, Dustin Schwenk, Eli VanderBilt, Matthew Wallingford, Luca Weihs, Mark Yatskar, and Ali Farhadi. RoboTHOR: An Open Simulation-to-Real Embodied AI Platform. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 3161–3171. Computer Vision Foundation / IEEE, 2020. 2, 9
  17. 17.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A Universe of Annotated 3D Objects. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13142–13153, 2022. 4, 5, 1, 2
  18. 18.Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Procthor: Large-scale embodied AI using procedural generation. In NeurIPS, 2022. 2, 4, 5, 1
  19. 19.Matt Deitke, Rose Hendrix, Ali Farhadi, Kiana Ehsani, and Aniruddha Kembhavi. Phone2Proc: Bringing Robust Robots into Our Chaotic World. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 9665–9675. IEEE, 2023. 3, 9
  20. 20.Ainaz Eftekhar, Kuo-Hao Zeng, Jiafei Duan, Ali Farhadi, Ani Kembhavi, and Ranjay Krishna. Selective visual representations improve convergence and generalization for embodied ai. arXiv preprint arXiv:2311.04193, 2023. 6
  21. 21.Kiana Ehsani, Winson Han, Alvaro Herrasti, Eli VanderBilt, Luca Weihs, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Manipulathor: A framework for visual object manipulation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 4497–4506. Computer Vision Foundation / IEEE, 2021. 2
  22. 22.Kiana Ehsani, Ali Farhadi, Aniruddha Kembhavi, and Roozbeh Mottaghi. Object manipulation via visual target localization. In Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part XXXIX, pages 321–337. Springer, 2022. 2
  23. 23.Christiane Fellbaum. WordNet: An Electronic Lexical Database. Bradford Books, 1998. 2, 4
  24. 24.Chuang Gan, Jeremy Schwartz, Seth Alter, Damian Mrowca, Martin Schrimpf, James Traer, Julian De Freitas, Jonas Kubilius, Abhishek Bhandwaldar, Nick Haber, Megumi Sano, Kuno Kim, Elias Wang, Michael Lingelbach, Aidan Curtis, Kevin T. Feigelis, Daniel Bear, Dan Gutfreund, David D. Cox, Antonio Torralba, James J. DiCarlo, Josh Tenenbaum, Josh H. McDermott, and Dan Yamins. ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual, 2021. 2
  25. 25.Chuang Gan, Siyuan Zhou, Jeremy Schwartz, Seth Alter, Abhishek Bhandwaldar, Dan Gutfreund, Daniel L. K. Yamins, James J. DiCarlo, Josh H. McDermott, Antonio Torralba, and Joshua B. Tenenbaum. The ThreeDWorld Transport Challenge: A Visually Guided Task-and-Motion Planning Benchmark for Physically Realistic Embodied AI. CoRR, abs/2103.14025, 2021. 2
  26. 26.Xiaofeng Gao, Qiaozi Gao, Ran Gong, Kaixiang Lin, Govind Thattai, and Gaurav S. Sukhatme. DialFRED: Dialogue-Enabled Agents for Embodied Instruction Following. IEEE Robotics Autom. Lett., 7(4):10049–10056, 2022. 2
  27. 27.Theophile Gervet, Soumith Chintala, Dhruv Batra, Jitendra Malik, and Devendra Singh Chaplot. Navigating to objects in the real world. Science Robotics, 2023. 2
  28. 28.Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Bernardo A Pires, and Remi Munos. Neural predictive belief representations. ICLR, 2019. 2
  29. 29.Zhaohan Daniel Guo, Bernardo Avila Pires, Bilal Piot, Jean-Bastien Grill, Florent Altche, R´emi Munos, and Mohammad Gheshlaghi Azar. Bootstrap latent-predictive representations for multitask reinforcement learning. In International Conference on Machine Learning, 2020. 2
  30. 30.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 770–778. IEEE Computer Society, 2016. 3
  31. 31.Daniel Ho, Kanishka Rao, Zhuo Xu, Eric Jang, Mohi Khansari, and Yunfei Bai. RetinaGAN: An Object-aware Approach to Sim-to-Real Transfer. In IEEE International Conference on Robotics and Automation, ICRA 2021, Xi’an, China, May 30 - June 5, 2021, pages 10920–10926. IEEE, 2021. 3
  32. 32.Wenlong Huang, P. Abbeel, Deepak Pathak, and Igor Mordatch. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. ICML, 2022. 2
  33. 33.Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al. Inner monologue: Embodied reasoning through planning with language models. arXiv preprint arXiv:2207.05608, 2022. 3
  34. 34.Brian Ichter, Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, Dmitry Kalashnikov, Sergey Levine, Yao Lu, Carolina Parada, Kanishka Rao, Pierre Sermanet, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Mengyuan Yan, Noah Brown, Michael Ahn, Omar Cortes, Nicolas Sievers, Clayton Tan, Sichun Xu, Diego Reyes, Jarek Rettinghouse, Jornell Quiambao, Peter Pastor, Linda Luu, Kuang-Huei Lee, Yuheng Kuang, Sally Jesmonth, Nikhil J. Joshi, Kyle Jeffrey, Rosario Jauregui Ruano, Jasmine Hsu, Keerthana Gopalakrishnan, Byron David, Andy Zeng, and Chuyuan Kelly Fu. Do as I can, not as I say: Grounding language in robotic affordances. In Conference on Robot Learning, CoRL 2022, 14-18 December 2022, Auckland, New Zealand, pages 287–318. PMLR, 2022. 2, 3
  35. 35.Michael Janner, Qiyang Li, and Sergey Levine. Offline Reinforcement Learning as One Big Sequence Modeling Problem. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 1273–1286, 2021. 3
  36. 36.Parham Mohsenzadeh Kebria, Abbas Khosravi, Syed Moshfeq Salaken, and Saeid Nahavandi. Deep imitation learning for autonomous vehicles based on convolutional neural networks. IEEE/CAA Journal of Automatica Sinica, 7:82–95, 2020. 3
  37. 37.Charles C. Kemp, Aaron Edsinger, Henry M. Clever, and Blaine Matulevich. The Design of Stretch: A Compact, Lightweight Mobile Manipulator for Indoor Human Environments. In 2022 International Conference on Robotics and Automation, ICRA 2022, Philadelphia, PA, USA, May 23-27, 2022, pages 3150–3157. IEEE, 2022. 2, 4, 7
  38. 38.Apoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, and Aniruddha Kembhavi. Simple but effective: CLIP embeddings for embodied AI. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 14809–14818. IEEE, 2022. 2, 3, 6, 7
  39. 39.Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, Aniruddha Kembhavi, Abhinav Kumar Gupta, and Ali Farhadi. AI2-THOR: An Interactive 3D Environment for Visual AI. ArXiv, abs/1712.05474, 2017. 2, 5, 1
  40. 40.Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 4392–4412. Association for Computational Linguistics, 2020. 2
  41. 41.Chengshu Li, Fei Xia, Roberto Martín-Martín, Michael Lingelbach, Sanjana Srivastava, Bokui Shen, Kent Elliott Vainio, Cem Gokmen, Gokul Dharan, Tanish Jain, Andrey Kurenkov, C. Karen Liu, Hyowon Gweon, Jiajun Wu, Li Fei-Fei, and Silvio Savarese. igibson 2.0: Object-centric simulation for robot learning of everyday household tasks. In Conference on Robot Learning, 8-11 November 2021, London, UK, pages 455–465. PMLR, 2021. 2
  42. 42.Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Martín-Martín, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, Mona Anvari, Minjune Hwang, Manasi Sharma, Arman Aydin, Dhruva Bansal, Samuel Hunter, Kyu-Young Kim, Alan Lou, Caleb R. Matthews, Ivan Villa-Renteria, Jerry Huayang Tang, Claire Tang, Fei Xia, Silvio Savarese, Hyowon Gweon, Karen Liu, Jiajun Wu, and Li Fei-Fei. BEHAVIOR-1K: A Benchmark for Embodied AI with 1, 000 Everyday Activities and Realistic Simulation. In Conference on Robot Learning, CoRL 2022, 14-18 December 2022, Auckland, New Zealand, pages 80–93. PMLR, 2022. 2
  43. 43.G. Li, Matthias Muller, Vincent Casser, Neil G. Smith, Dominik Ludewig Michels, and Bernard Ghanem. Oil: Observational imitation learning. Robotics: Science and Systems XV, 2018. 3
  44. 44.Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Liu, Shiqi Zhang, Joydeep Biswas, and Peter Stone. Llm+ p: Empowering large language models with optimal planning proficiency. arXiv preprint arXiv:2304.11477, 2023. 2
  45. 45.Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma, Claire Chen, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Pieter Abbeel, Jitendra Malik, Dhruv Batra, Yixin Lin, Oleksandr Maksymets, Aravind Rajeswaran, and Franziska Meier. Where are we in the search for an artificial visual cortex for embodied intelligence? CoRR, abs/2303.18240, 2023. 3
  46. 46.Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma, Claire Chen, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Pieter Abbeel, Jitendra Malik, et al. Where are we in the search for an artificial visual cortex for embodied intelligence? ICLR, 2023. 2
  47. 47.John P. McCrae, Alexandre Rademaker, Francis Bond, Ewa Rudnicka, and Christiane Fellbaum. English WordNet 2019 - An Open-Source WordNet for English. In Proceedings of the 10th Global Wordnet Conference, GWC 2019, Wroclaw, Poland, July 23-27, 2019, pages 245–252. Global Wordnet Association, 2019. 5, 2, 4
  48. 48.Luc Le Mero, Dewei Yi, Mehrdad Dianati, and Alexandros Mouzakitis. A survey on imitation learning techniques for end-to-end autonomous vehicles. IEEE Transactions on Intelligent Transportation Systems, 23:14128–14147, 2022. 3
  49. 49.Volodymyr Mnih, Adria Puigdom`enech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous Methods for Deep Reinforcement Learning. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, pages 1928–1937. JMLR.org, 2016. 3
  50. 50.Maxime Oquab, Timothee Darcet, Theo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nicolas Ballas, Gabriel Synnaeve, Ishan Misra, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski. Dinov2: Learning robust visual features without supervision, 2023. 2, 7
  51. 51.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback. In NeurIPS, 2022. 5
  52. 52.Yunpeng Pan, Ching-An Cheng, Kamil Saigol, Keuntaek Lee, Xinyan Yan, Evangelos A. Theodorou, and Byron Boots. Imitation learning for agile autonomous driving. The International Journal of Robotics Research, 39:286 – 302, 2019. 3
  53. 53.Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba. VirtualHome: Simulating Household Activities via Programs. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 8494–8502. Computer Vision Foundation / IEEE Computer Society, 2018. 2
  54. 54.Xavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote, Tsung-Yen Yang, Ruslan Partsey, Ruta Desai, Alexander William Clegg, Michal Hlavac, So Yeon Min, Vladimir Vondrus, Theophile Gervet, Vincent-Pierre Berges, John M. Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara Rai, and Roozbeh Mottaghi. Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots. CoRR, abs/2310.13724, 2023. 2
  55. 55.Henry Pulver, Francisco Eiras, Ludovico Carozza, Majd Hawasly, Stefano V Albrecht, and Subramanian Ramamoorthy. Pilot: Efficient planning by imitation learning and optimisation for safe autonomous driving. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1442–1449. IEEE, 2021. 3
  56. 56.Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 9979–9988. Computer Vision Foundation / IEEE, 2020. 2
  57. 57.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, pages 8748–8763. PMLR, 2021. 2, 3
  58. 58.Ram Ramrakhya, Eric Undersander, Dhruv Batra, and Abhishek Das. Habitat-web: Learning embodied object-search strategies from human demonstrations at scale. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 5163–5173. IEEE, 2022. 2, 3, 4
  59. 59.Ram Ramrakhya, Dhruv Batra, Erik Wijmans, and Abhishek Das. Pirlnav: Pretraining with imitation and rl finetuning for objectnav. In Workshop on IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 2, 8, 7
  60. 60.Scott E. Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas. A Generalist Agent. Trans. Mach. Learn. Res., 2022, 2022. 3
  61. 61.Manolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, and Vladlen Koltun. Habitat: A Platform for Embodied AI Research. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 9338–9346. IEEE, 2019. 2
  62. 62.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal Policy Optimization Algorithms. CoRR, abs/1707.06347, 2017. 3
  63. 63.Bokui Shen, Fei Xia, Chengshu Li, Roberto Martín-Martín, Linxi Fan, Guanzhi Wang, Claudia Perez-D’Arpino, Shyamal Buch, Sanjana Srivastava, Lyne Tchapmi, Micael Tchapmi, Kent Vainio, Josiah Wong, Li Fei-Fei, and Silvio Savarese. iGibson 1.0: A Simulation Environment for Interactive Tasks in Large Realistic Scenes. In IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2021, Prague, Czech Republic, September 27 - Oct. 1, 2021, pages 7520–7527. IEEE, 2021. 2
  64. 64.Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. ALFRED: A benchmark for interpreting grounded instructions for everyday tasks. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 10737–10746. Computer Vision Foundation / IEEE, 2020. 2
  65. 65.Kunal Pratap Singh, Jordi Salvador, Luca Weihs, and Aniruddha Kembhavi. Scene Graph Contrastive Learning for Embodied Navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 10884–10894, 2023. 2, 3
  66. 66.Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M Sadler, Wei-Lun Chao, and Yu Su. Llm-planner: Few-shot grounded planning for embodied agents with large language models. Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023. 2
  67. 67.Sanjana Srivastava, Chengshu Li, Michael Lingelbach, Roberto Martín-Martín, Fei Xia, Kent Vainio, Zheng Lian, Cem Gokmen, S. Buch, C. Karen Liu, Silvio Savarese, Hyowon Gweon, Jiajun Wu, and Li Fei-Fei. BEHAVIOR: Benchmark for Everyday Household Activities in Virtual, Interactive, and Ecological Environments. In Conference on Robot Learning, 2021. 2
  68. 68.Andrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John M. Turner, Noah Maestre, Mustafa Mukadam, Devendra Singh Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel X. Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra. Habitat 2.0: Training Home Assistants to Rearrange their Habitat. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 251–266, 2021. 2
  69. 69.Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2017, Vancouver, BC, Canada, September 24-28, 2017, pages 23–30. IEEE, 2017. 3
  70. 70.Sai Vemprala, Rogerio Bonatti, Arthur Bucker, and Ashish Kapoor. Chatgpt for robotics: Design principles and model abilities. Technical Report MSR-TR-2023-8, Microsoft, 2023. 2
  71. 71.Zihao Wang, Shaofei Cai, Anji Liu, Xiaojian Ma, and Yitao Liang. Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents. arXiv preprint arXiv:2302.01560, 2023. 2
  72. 72.Yao Wei, Yanchao Sun, Ruijie Zheng, Sai Vemprala, Rogerio Bonatti, Shuhang Chen, Ratnesh Madaan, Zhongjie Ba, Ashish Kapoor, and Shuang Ma. Is Imitation All You Need? Generalized Decision-Making with Dual-Phase Training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16221–16231, 2023. 3
  73. 73.Luca Weihs, Jordi Salvador, Klemen Kotar, Unnat Jain, Kuo-Hao Zeng, Roozbeh Mottaghi, and Aniruddha Kembhavi. Allenact: A framework for embodied ai research. arXiv preprint arXiv:2008.12760, 2020. 2, 6, 1
  74. 74.Luca Weihs, Matt Deitke, Aniruddha Kembhavi, and Roozbeh Mottaghi. Visual room rearrangement. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  75. 75.Luca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador, Svetlana Lazebnik, Aniruddha Kembhavi, and Alexander G. Schwing. Bridging the imitation gap by adaptive insubordination. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 19134–19146, 2021. 2, 4
  76. 76.Erik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee, Irfan Essa, Devi Parikh, Manolis Savva, and Dhruv Batra. DDPPO: learning near-perfect pointgoal navigators from 2.5 billion frames. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. 2, 3
  77. 77.Fei Xia, Amir R. Zamir, Zhi-Yang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. Gibson Env: Real-World Perception for Embodied Agents. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 9068–9079. Computer Vision Foundation / IEEE Computer Society, 2018. 2
  78. 78.Fanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia, Hao Zhu, Fangchen Liu, Minghua Liu, Hanxiao Jiang, Yifu Yuan, He Wang, Li Yi, Angel X. Chang, Leonidas J. Guibas, and Hao Su. SAPIEN: A SimulAted Part-based Interactive ENvironment. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2
  79. 79.Karmesh Yadav, Ram Ramrakhya, Arjun Majumdar, Vincent-Pierre Berges, Sachit Kuhar, Dhruv Batra, Alexei Baevski, and Oleksandr Maksymets. Offline visual representation learning for embodied navigation. In Workshop on Reincarnating Reinforcement Learning at ICLR 2023, 2023. 2
  80. 80.Brian Yamauchi. A frontier-based approach for autonomous exploration. In Proceedings 1997 IEEE International Symposium on Computational Intelligence in Robotics and Automation CIRA’97.’Towards New Computational Principles for Robotics and Automation’, pages 146–151. IEEE, 1997. 7
  81. 81.Joel Ye, Dhruv Batra, Abhishek Das, and Erik Wijmans. Auxiliary tasks and exploration enable objectnav. CoRR, abs/2104.04112, 2021. 3
  82. 82.Sriram Yenamandra, Arun Ramachandran, Karmesh Yadav, Austin Wang, Mukul Khanna, Theophile Gervet, Tsung-Yen Yang, Vidhi Jain, Alexander William Clegg, John M. Turner, Zsolt Kira, Manolis Savva, Angel X. Chang, Devendra Singh Chaplot, Dhruv Batra, Roozbeh Mottaghi, Yonatan Bisk, and Chris Paxton. Homerobot: Open-vocabulary mobile manipulation. CoRR, abs/2306.11565, 2023. 2
  83. 83.Naoki Yokoyama, Alexander Clegg, Eric Undersander, Sehoon Ha, Dhruv Batra, and Akshara Rai. Adaptive skill coordination for robotic mobile manipulation. ArXiv, abs/2304.00410, 2023. 2
  84. 84.Kuo-Hao Zeng, Luca Weihs, Ali Farhadi, and Roozbeh Mottaghi. Pushing it out of the way: Interactive visual navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021. 2
  85. 85.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. ICCV, abs/2303.15343, 2023. 2, 7
  86. 86.Xu Zhao, Wenchao Ding, Yongqi An, Yinglong Du, Tao Yu, Min Li, Ming Tang, and Jinqiao Wang. Fast segment anything. arXiv preprint arXiv:2306.12156, 2023. 9
  87. 87.Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krahenbühl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. In Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part IX, pages 350–368. Springer, 2022. 8, 10
  88. 88.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 2242–2251. IEEE Computer Society, 2017. 3
  89. 89.Yuke Zhu, Roozbeh Mottaghi, Eric Kolve, Joseph J. Lim, Abhinav Gupta, Li Fei-Fei, and Ali Farhadi. Target-driven visual navigation in indoor scenes using deep reinforcement learning. In 2017 IEEE International Conference on Robotics and Automation, ICRA 2017, Singapore, Singapore, May 29 - June 3, 2017, pages 3357–3364. IEEE, 2017. 2

Citation

MLA
Ehsani, K., et al. “SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 16238–50, https://doi.org/10.1109/CVPR52733.2024.01537.
APA
Ehsani, K., Gupta, T., Hendrix, R., Salvador, J., Weihs, L., Zeng, K.-H., Singh, K. P., Kim, Y., Han, W., Herrasti, A., Krishna, R., Schwenk, D., VanderBilt, E., & Kembhavi, A. (2024). SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16238–16250. https://doi.org/10.1109/CVPR52733.2024.01537
Chicago
Ehsani, K., T. Gupta, R. Hendrix, et al. 2024. “SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16238–50. https://doi.org/10.1109/CVPR52733.2024.01537.
Harvard
Ehsani, K. et al. (2024) “SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World”, 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 16238–16250. Available at: https://doi.org/10.1109/CVPR52733.2024.01537.
Vancouver
1. Ehsani K, Gupta T, Hendrix R, et al (2024) SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 16238–16250

BibTeX

@inproceedings{Ehsani_2024, title={SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World}, url={http://dx.doi.org/10.1109/CVPR52733.2024.01537}, DOI={10.1109/cvpr52733.2024.01537}, booktitle={2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Ehsani, Kiana and Gupta, Tanmay and Hendrix, Rose and Salvador, Jordi and Weihs, Luca and Zeng, Kuo-Hao and Singh, Kunal Pratap and Kim, Yejin and Han, Winson and Herrasti, Alvaro and Krishna, Ranjay and Schwenk, Dustin and VanderBilt, Eli and Kembhavi, Aniruddha}, year={2024}, month=June, pages={16238–16250} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE