AI2-THOR: An Interactive 3D Environment for Visual AI

Eric KolveRoozbeh MottaghiWinson HanEli VanderBiltLuca WeihsAlvaro HerrastiMatt DeitkeKiana EhsaniDaniel GordonYuke Zhu

article2017arXiv1,545 citations

Introduces a near-photorealistic 3D indoor simulation platform that enables embodied AI agents to physically interact with objects and explore complex environments for visual reinforcement learning and task planning.

Listen

Developing artificial intelligence agents capable of complex visual understanding requires training through rich physical interactions rather than static images or video feeds. However, conducting physical robot experiments in the real world is slow, costly, potentially hazardous, and difficult to scale across varied environments. The article presents and evaluates AI2-THOR (The House Of inteRactions), an interactive, near photo-realistic 3D simulation framework designed to train embodied AI agents across diverse indoor tasks, including navigation, manipulation, and instruction following.

The system pairs a front-end Python interface with the Unity 3D game engine, enabling researchers to control multiple agent embodiments such as mobile bases, drones, and multi-joint robotic arms. It incorporates several extensive scene datasets, ranging from 120 artist-crafted rooms to 10,000 procedurally generated houses, along with an interactive object database containing 3,578 assets capable of dynamic state changes like slicing, cooking, breaking, and filling with liquids.

The article demonstrates several significant findings regarding the platform’s performance and research utility. First, AI2-THOR provides an unmatched scale of interaction, outperforming alternative simulators by supporting physical state changes, arm manipulation, audio cues, virtual reality integration, and procedural generation within a single system. Second, pre-training agents on the procedurally generated dataset achieved state-of-the-art visual navigation performance across multiple separate benchmarks without requiring extra domain-specific training data, effectively overcoming severe training overfitting. Third, in performance benchmarks, the platform demonstrated competitive training throughput, averaging 167.7 frames per second on a two-GPU machine compared to 230.5 frames per second for a less interactive alternative.

These findings indicate that highly interactive simulations can effectively serve as safe, rapid, and low-cost proxies for physical robotics research while substantially improving generalization to unseen environments. By allowing simulated models to transfer more reliably to real-world tasks, this approach reduces hardware testing expenses and accelerates development timelines across embodied AI, language grounding, and multi-agent systems.

Organizations advancing robotic AI should leverage procedural simulation environments to pre-train models at scale before physical deployment. Continued development should focus on expanding the variety of interactive physical assets, optimizing simulation speed during complex object manipulations, and further validating simulation-to-reality transfer across diverse physical robot platforms.

While the framework offers high fidelity, readers should note that simulated interactions may not capture every real-world physical nuance. Computational throughput remains constrained by complex multi-object arm collisions and environment resets during reinforcement learning. Nevertheless, extensive cross-benchmark evaluations and broad community adoption support high confidence in the platform's core conclusions.

  • Paper: Target-driven visual navigation in indoor scenes using deep reinforcement learning, Yuke Zhu et al. (2016). This paper introduced the initial prototype of the AI2-THOR simulation environment to enable target-driven visual navigation with deep reinforcement learning.
  • Paper: OpenAI Gym, Greg Brockman et al. (2016). This work established the standardized agent-environment interaction interface that foundationally shaped interactive 3D simulation platforms for reinforcement learning.
  • Paper: Habitat: A Platform for Embodied AI Research, Manolis Savva et al. (2019). Habitat expands on the interactive 3D simulation paradigm established by AI2-THOR by delivering a high-throughput platform for training embodied agents across photorealistic environments.
  • Paper: Objaverse: A Universe of Annotated 3D Objects, Matt Deitke et al. (2022). Objaverse massively scales the 3D asset repository and interactive object diversity needed to build richer, open-vocabulary embodied AI environments.
  • Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). PaLM-E applies embodied sensory grounding to multimodal language models, enabling high-level planning and interaction in complex visual environments.
  • Paper: Octo: An Open-Source Generalist Robot Policy, O. Team et al. (2024). Octo demonstrates generalist robot policy learning by mapping visual inputs and goal instructions directly to actions across diverse embodied manipulation settings.
Cover for AI2-THOR: An Interactive 3D Environment for Visual AI

Abstract

We introduce The House Of inteRactions (THOR), a framework for visual AI research, available at this http URL. AI2-THOR consists of near photo-realistic 3D indoor scenes, where AI agents can navigate in the scenes and interact with objects to perform tasks. AI2-THOR enables research in many different domains including but not limited to deep reinforcement learning, imitation learning, learning by interaction, planning, visual question answering, unsupervised representation learning, object detection and segmentation, and learning models of cognition. The goal of AI2-THOR is to facilitate building visually intelligent models and push the research forward in this domain.

Table of Contents

  • 1 What is AI2-THOR?
  • 2 What does AI2-THOR feature?
  • 2.1 API
  • 2.2 Scene Datasets
  • 2.3 Agents
  • 2.4 Actions
  • 2.5 Image Modalities
  • 2.6 Objects
  • 2.7 Environment Metadata
  • 3 What has AI2-THOR been used for?
  • 4 Why use AI2-THOR?
  • 5 Conclusion
  • References
  • A Contributions
  • B Performance Comparison

Knowls

  1. Knowl 1 — AI2-THOR Architecture and Agent-Simulator Loop

    model/method

    AI2-THOR is an embodied AI simulation platform operating on a client-server architecture that interfaces a Python front-end API with a Unity 3D game engine back-end.

    The interaction cycle follows a synchronous agent-simulator loop:

    1. An embodied agent or Python script issues an action command through the Python API.
    2. The action is transmitted via a local server communication layer to the Unity simulation runtime.
    3. The Unity engine computes physical dynamics, executes the requested action (such as locomotion, continuous arm actuation, or object state changes), evaluates visibility, and renders requested camera perspectives and shader-based sensor buffers.
    4. Unity packages the simulation state into an Event object and returns it to Python. The Event object contains visual frames from all active scene cameras alongside detailed environment metadata, including agent poses, object coordinate bounding frames, object interaction states, and action execution status flags.
  2. Knowl 2 — Scene Datasets in AI2-THOR

    definition

    AI2-THOR includes four primary indoor scene datasets, each sharing a common API and compatibility across supported agent embodiments:

    • iTHOR: A collection of 120 manually designed, room-sized indoor environments covering four room categories: kitchens, living rooms, bedrooms, and bathrooms.
    • RoboTHOR: A set of 89 modular, maze-styled apartment scenes created by professional 3D artists. A subset of these scenes is physically replicated in the real world to evaluate simulation-to-real (sim2real) transfer.
    • ProcTHOR-10K: A dataset of 10,000 procedurally generated, semantically plausible multi-room houses designed to scale training diversity and mitigate agent overfitting to fixed training layouts.
    • ArchitecTHOR: A set of 10 large, artist-created single-story evaluation houses (partitioned into 5 validation and 5 test scenes) used as a benchmark to evaluate whether policies trained on procedurally generated layouts generalize to human-authored architectural designs.
  3. Knowl 3 — Embodied Agent Architectures and Control Paradigms

    model/method

    AI2-THOR supports multiple robotic agent embodiments with varying kinematics and interaction granularities:

    • ManipulaTHOR: An embodied mobile manipulator equipped with a 6-Degree-of-Freedom (6 DoF) robotic arm capable of continuous motion planning, grasping, and manipulating objects.
    • StretchRE1: A mobile manipulator model based on the Hello Robot Stretch RE1 platform with arm-based continuous manipulation capabilities.
    • LoCoBot: A wheeled mobile robot platform configured for navigation and abstracted object interactions.
    • Abstract Agent: A standard first-person virtual agent for discrete navigation and high-level interaction primitives.
    • Drone Agent: An aerial agent designed for navigation and reaction tasks in 3D space.

    Control paradigms in AI2-THOR operate at two levels:

    • Continuous Arm-Based Control: Direct actuation of multi-joint robotic arms to grasp, move, or articulate objects along continuous trajectories.
    • High-Level Abstracted Control: Symbolic action execution (such as OPEN or PICKUP) that triggers automatically when the target object is visible in the agent's field of view and lies within a predefined distance threshold.
  4. Knowl 4 — AI2-THOR Action Space and Interaction Types

    model/method

    AI2-THOR categorizes actions into four distinct types:

    1. Navigation Actions: Locomotion commands configured as discrete steps or continuous displacements. These include translations (e.g., moving ahead by 0.25 m0.25\,\text{m}), rotations (e.g., rotating right by 30∘30^\circ), camera tilt adjustments (e.g., looking up by 30∘30^\circ), and coordinate teleportation.
    2. Interactive Actions:
      • Abstracted Interactions: High-level operations executed directly on visible objects within reaching distance, including pickup, drop, open, close, push, throw, and place.
      • Object State Manipulations: Explicit state transformations such as slicing, cooking, breaking, toggling power state, filling with liquids, or consuming/depleting objects.
      • Continuous Arm Manipulation: Direct joint and end-effector control to grasp and continuously articulate objects (e.g., pulling a drawer open incrementally).
      • Causal Physical Interactions: Secondary physical consequences governed by the physics engine (e.g., objects falling off a tipped table, cups filling when placed under an active faucet, or breakable items shattering upon impact).
    3. Environment Queries: On-demand computations executed via Unity, including shortest navigable paths to target objects, pixel-to-object semantic raycasting, and object 3D convex hull extraction.
    4. Environment State Modifications: Global scene updates such as randomizing material textures, altering ambient and point light sources, adjusting camera image resolution, modifying skybox textures, and toggling rendering quality settings.
  5. Knowl 5 — Visual and Sensor Modalities in AI2-THOR

    definition

    AI2-THOR supports multiple visual and sensory modalities rendered on demand through Unity shaders for any camera attached to an agent or placed in the environment:

    • RGB: High-fidelity color imagery rendered from first-person or third-person (e.g., top-down) camera perspectives.
    • Depth: Metric depth maps representing distance from the camera plane to scene surfaces.
    • Semantic Segmentation: Per-pixel class labeling where every pixel is assigned an integer category identifier corresponding to the semantic class of the visible object or surface.
    • Instance Segmentation: Per-pixel instance labeling where distinct objects of the same class receive unique color or numerical ID encodings.
    • Surface Normals: Per-pixel surface normal orientation vectors rendered in camera or world coordinate space.
  6. Knowl 6 — Interactive Object Database and State Representations

    definition

    AI2-THOR includes a database of 3,578 interactive 3D object models spanning household categories such as furniture, appliances, kitchenware, and tools. Every object asset supports physical interactions alongside discrete and continuous state attributes:

    • Openness: Articulation degree ranging continuously from fully closed (0.00.0) to fully open (1.01.0) for cabinets, drawers, microwaves, refrigerators, and laptops.
    • Toggle State: Binary operational state (on/off) for electronics, lamps, stoves, and faucets.
    • Thermal and Cooking State: State flags tracking whether food items are uncooked, cooked, or burned based on exposure to heat sources.
    • Structural Integrity and Slicing: Discrete structural transitions, including whole versus sliced (e.g., bread, fruits) and intact versus shattered/broken.
    • Cleanliness: Clean versus dirty state flags.
    • Containment and Fill Level: Capacity to hold other objects or fluids.
  7. Knowl 7 — Environment Metadata and Ground-Truth Feedback

    definition

    After executing each action, the AI2-THOR simulator returns an environment metadata structure within the Event payload containing ground-truth simulation variables:

    • Agent Metadata: 3D position (x,y,z)(x, y, z), camera rotation (pitch,yaw,roll)(\text{pitch}, \text{yaw}, \text{roll}), and joint configurations of manipulators.
    • Object Metadata: 3D spatial bounding boxes, precise positions and rotations, visibility flags relative to the agent's current frustum, occlusion metrics, distance from agent, and discrete/continuous state flags (e.g., temperature, toggle state, cleanliness, open percentage).
    • Scene-Level Data: 3D geometric scene bounds, navmesh configuration, and a discrete grid of reachable navigable positions.
    • Action Execution Status: A boolean flag indicating whether the command executed successfully or failed (e.g., due to physical collisions, reachability constraints, or occlusion).

    While typically concealed from agents during inference to prevent heuristic shortcutting, this metadata is used to construct reward functions in reinforcement learning, generate ground-truth demonstration trajectories for imitation learning, and build evaluation benchmarks.

  8. Knowl 8 — Comparison of Embodied AI Simulation Platforms

    data/table

    Embodied AI simulation platforms differ across scene scale, asset counts, interaction capabilities, multi-agent support, sensory feedback, and underlying simulation engines:

    Simulator # of Scenes # of Objects Object States Arm Manipulation Multi-Agent Sound VR Engine
    AI2-THOR ∞\infty 3578 ✓ ✓ ✓ ✓ ✓ Unity
    iGibson 2.0 15 1217 ✓ ✓ ✓ ✓ PyBullet
    Habitat 1.0 1000 – ✓ Magnum
    Habitat 2.0 105 92 ✓ ✓ Magnum
    ThreeDWorld 15 200 ✓ ✓ ✓ Unity
    SAPIEN 0 2346 ✓ PhysX

    AI2-THOR provides procedurally generated scenes (yielding theoretically unbounded scene count ∞\infty via ProcTHOR) alongside 3,578 interactive objects, full support for object state changes, arm-based manipulation, multi-agent interaction, audio rendering, VR interfacing, an interactive editor, and a Unity-based engine.

  9. Knowl 9 — Reinforcement Learning Training Throughput Benchmark

    empirical result

    An ObjectNav agent was trained for 1,000,0001,000,000 steps using reinforcement learning to benchmark training throughput across parallel simulator instances:

    • Hardware and Architecture: Dual GeForce RTX 2080 GPUs, where GPU-0 evaluates the actor-critic policy network and performs model updates, while GPU-1 hosts 28 parallel simulator processes.
    • Agent and Policy: LoCoBot embodiment with identical action spaces and actor-critic architecture across platforms.
    • Protocol: Rollout length of 128 steps, maximum episode length fixed at 500 steps (with the End action disabled to prevent policy-dependent length variance), and synchronized scene advancement occurring every 10 rollouts (10×128×28=35,84010 \times 128 \times 28 = 35,840 steps).
    • Performance Results:
      • AI2-THOR (ProcTHOR-10K): Training frame rate ranged between 145.5145.5 and 179.4179.4 frames per second (FPS), with an average throughput of 167.7167.7 FPS.
      • Habitat 1.0 (HM3D): Training frame rate ranged between 119.7119.7 and 264.3264.3 FPS, with an average throughput of 230.5230.5 FPS.

    The training bottleneck in synchronous reinforcement learning is driven by model forward/backward passes and scene resetting overheads rather than single-process raw rendering speed alone.

Coverage note — Surveys and literature summaries of third-party publications utilizing AI2-THOR across embodied AI subfields were omitted as they represent external downstream applications rather than core framework contributions.

References

  1. 1.Matt Deitke, Winson Han, Alvaro Herrasti, Aniruddha Kembhavi, Eric Kolve, Roozbeh Mottaghi, Jordi Salvador, Dustin Schwenk, Eli VanderBilt, Matthew Wallingford, Luca Weihs, Mark Yatskar, and Ali Farhadi. Robothor: An open simulation-to-real embodied ai platform. In CVPR, 2020. 2, 3, 6, 7
  2. 2.Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Jordi Salvador, Kiana Ehsani, Winson Han, Eric Kolve, Ali Farhadi, Aniruddha Kembhavi, and Roozbeh Mottaghi. Procthor: Large-scale embodied ai using procedural generation. arXiv, 2022. 2, 3, 6, 7, 8, 12
  3. 3.Heming Du, Xin Yu, and Liang Zheng. Learning object relation graph and tentative policy for visual navigation. In ECCV, 2020. 6
  4. 4.Kshitij Dwivedi, Gemma Roig, Aniruddha Kembhavi, and Roozbeh Mottaghi. What do navigation agents learn about their environment? In CVPR, 2022. 7, 8
  5. 5.Kiana Ehsani, Winson Han, Alvaro Herrasti, Eli VanderBilt, Luca Weihs, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Manipulathor: A framework for visual object manipulation. In CVPR, 2021. 3, 8
  6. 6.Kiana Ehsani, Roozbeh Mottaghi, and Ali Farhadi. Segan: Segmenting and generating the invisible. In CVPR, 2018. 8
  7. 7.Chuang Gan, Jeremy Schwartz, Seth Alter, Martin Schrimpf, James Traer, Julian De Freitas, Jonas Kubilius, Abhishek Bhandwaldar, Nick Haber, Megumi Sano, et al. Threedworld: A platform for interactive multi-modal physical simulation. In Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), 2020. 2, 8
  8. 8.Chuang Gan, Yiwei Zhang, Jiajun Wu, Boqing Gong, and Joshua B Tenenbaum. Look, listen, and act: Towards audio-visual embodied navigation. In ICRA, 2020. 6, 7
  9. 9.Xiaofeng Gao, Qiaozi Gao, Ran Gong, Kaixiang Lin, Govind Thattai, and Gaurav S Sukhatme. Dialfred: Dialogue-enabled agents for embodied instruction following. IEEE Robotics and Automation Letters, 2022. 6
  10. 10.Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi. Iqa: Visual question answering in interactive environments. In CVPR, 2018. 6
  11. 11.Unnat Jain, Luca Weihs, Eric Kolve, Ali Farhadi, Svetlana Lazebnik, Aniruddha Kembhavi, and Alexander G. Schwing. A cordial sync: Going beyond marginal policies for multi-agent embodied tasks. In ECCV, 2020. 6
  12. 12.Unnat Jain, Luca Weihs, Eric Kolve, Mohammad Rastegari, Svetlana Lazebnik, Ali Farhadi, Alexander G. Schwing, and Aniruddha Kembhavi. Two body problem: Collaborative visual task completion. In CVPR, 2019. 6, 7
  13. 13.Siddharth Karamcheti, Dorsa Sadigh, and Percy Liang. Learning adaptive language interfaces through decomposition. arXiv, 2020. 6
  14. 14.Charles C Kemp, Aaron Edsinger, Henry M Clever, and Blaine Matulevich. The design of stretch: A compact, lightweight mobile manipulator for indoor human environments. In ICRA, 2022. 3
  15. 15.Apoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, and Aniruddha Kembhavi. Simple but effective: Clip embeddings for embodied ai. In CVPR, 2022. 6
  16. 16.Klemen Kotar, Gabriel Ilharco, Ludwig Schmidt, Kiana Ehsani, and Roozbeh Mottaghi. Contrasting contrastive self-supervised representation learning pipelines. In ICCV, 2021. 8
  17. 17.Klemen Kotar and Roozbeh Mottaghi. Interactron: Embodied adaptive object detection. In CVPR, 2022. 7, 8
  18. 18.Chengshu Li, Fei Xia, Roberto Martín-Martín, Michael Lingelbach, Sanjana Srivastava, Bokui Shen, Kent Elliott Vainio, Cem Gokmen, Gokul Dharan, Tanish Jain, Andrey Kurenkov, Karen Liu, Hyowon Gweon, Jiajun Wu, Li Fei-Fei, and Silvio Savarese. igibson 2.0: Object-centric simulation for robot learning of everyday household tasks. In CoRL, 2021. 2, 8
  19. 19.Qi Li, Kaichun Mo, Yanchao Yang, Hang Zhao, and Leonidas Guibas. Ifr-explore: Learning inter-object functional relationships in 3d indoor scenes. In ICLR, 2022. 7
  20. 20.Xinzhu Liu, Di Guo, Huaping Liu, and Fuchun Sun. Multi-agent embodied visual semantic navigation with scene prior knowledge. IEEE Robotics and Automation Letters, 2022. 6
  21. 21.Martin Lohmann, Jordi Salvador, Aniruddha Kembhavi, and Roozbeh Mottaghi. Learning about objects by learning to interact with them. In NeurIPS, 2020. 8
  22. 22.Yi Lu, Yaran Chen, Dongbin Zhao, and Dong Li. Mgrl: Graph neural network based inference in a markov network with reinforcement learning for visual navigation. Neurocomputing, 2021. 6
  23. 23.So Yeon Min, Devendra Singh Chaplot, Pradeep Ravikumar, Yonatan Bisk, and Ruslan Salakhutdinov. Film: Following instructions in language with modular methods. In ICLR, 2022. 6
  24. 24.Adithyavairavan Murali, Tao Chen, Kalyan Vasudev Alwala, Dhiraj Gandhi, Lerrel Pinto, Saurabh Gupta, and Abhinav Gupta. Pyrobot: An open-source robotics framework for research and benchmarking. arXiv, 2019. 3
  25. 25.Tushar Nagarajan and Kristen Grauman. Learning affordance landscapes for interaction exploration in 3d environments. In NeurIPS, 2020. 7
  26. 26.Tushar Nagarajan and Kristen Grauman. Shaping embodied agent behavior with activity-context priors from egocentric video. In NeurIPS, 2021. 7
  27. 27.Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gokhan Tur, and Dilek Hakkani-Tur. Teach: Task-driven embodied agents that chat. In AAAI, 2022. 6, 7
  28. 28.Alexander Pashevich, Cordelia Schmid, and Chen Sun. Episodic transformer for vision-and-language navigation. In ICCV, 2021. 6
  29. 29.Santhosh Kumar Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alexander Clegg, John M Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X Chang, Manolis Savva, Yili Zhao, and Dhruv Batra. Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI. In Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. 11
  30. 30.Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A platform for embodied ai research. In ICCV, 2019. 8
  31. 31.Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In CVPR, 2020. 6, 7
  32. 32.Andrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Singh Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel X. Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra. Habitat 2.0: Training home assistants to rearrange their habitat. In NeurIPS, 2021. 2, 8
  33. 33.Sinan Tan, Weilai Xiang, Huaping Liu, Di Guo, and Fuchun Sun. Multi-agent embodied question answering in interactive environments. In ECCV, 2020. 6
  34. 34.Luca Weihs, Matt Deitke, Aniruddha Kembhavi, and Roozbeh Mottaghi. Visual room rearrangement. In CVPR, 2021. 7, 8
  35. 35.Luca Weihs, Aniruddha Kembhavi, Kiana Ehsani, Sarah M Pratt, Winson Han, Alvaro Herrasti, Eric Kolve, Dustin Schwenk, Roozbeh Mottaghi, and Ali Farhadi. Learning generalizable visual representations via interactive gameplay. In ICLR, 2021. 6, 8
  36. 36.Luca Weihs, Jordi Salvador, Klemen Kotar, Unnat Jain, Kuo-Hao Zeng, Roozbeh Mottaghi, and Aniruddha Kembhavi. Allenact: A framework for embodied AI research. arXiv, 2020. 11
  37. 37.Mitchell Wortsman, Kiana Ehsani, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Learning to learn how to learn: Self-adaptive visual navigation using meta-learning. In CVPR, 2019. 6
  38. 38.Qi Wu, Cheng-Ju Wu, Yixin Zhu, and Jungseock Joo. Communicative learning with natural gestures for embodied navigation agents with human-in-the-scene. In IROS, 2021. 6, 7
  39. 39.Fanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia, Hao Zhu, Fangchen Liu, Minghua Liu, Hanxiao Jiang, Yifu Yuan, He Wang, Li Yi, Angel X. Chang, Leonidas J. Guibas, and Hao Su. Sapien: A simulated part-based interactive environment. In CVPR, 2020. 8
  40. 40.Wei Yang, Xiaolong Wang, Ali Farhadi, Abhinav Gupta, and Roozbeh Mottaghi. Visual semantic navigation using scene priors. In ICLR, 2019. 6
  41. 41.Rowan Zellers, Ari Holtzman, Matthew E. Peters, Roozbeh Mottaghi, Aniruddha Kembhavi, Ali Farhadi, and Yejin Choi. Piglet: Language grounding through neuro-symbolic interaction in a 3d world. In ACL, 2021. 6
  42. 42.Kuo-Hao Zeng, Roozbeh Mottaghi, Luca Weihs, and Ali Farhadi. Visual reaction: Learning to play catch with your drone. In CVPR, 2020. 3
  43. 43.Yizhou Zhao, Kaixiang Lin, Zhiwei Jia, Qiaozi Gao, Govind Thattai, Jesse Thomason, and Gaurav S Sukhatme. Luminous: Indoor scene generation for embodied ai challenges. arXiv, 2021. 8
  44. 44.Kaiyu Zheng, Rohan Chitnis, Yoonchang Sung, George Konidaris, and Stefanie Tellex. Towards optimal correlational object search. In ICRA, 2022. 6
  45. 45.Yuke Zhu, Roozbeh Mottaghi, Eric Kolve, Joseph J Lim, Abhinav Gupta, Li Fei-Fei, and Ali Farhadi. Target-driven visual navigation in indoor scenes using deep reinforcement learning. In ICRA, 2017. 6, 7, 11

Citation

MLA
Kolve, E., et al. “AI2-THOR: An Interactive 3D Environment for Visual AI”. arXiv, 2017, http://arxiv.org/abs/1712.05474v4.
APA
Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Deitke, M., Ehsani, K., Gordon, D., Zhu, Y., Kembhavi, A., Gupta, A., & Farhadi, A. (2017). AI2-THOR: An Interactive 3D Environment for Visual AI. arXiv. http://arxiv.org/abs/1712.05474v4
Chicago
Kolve, E., R. Mottaghi, W. Han, et al. 2017. “AI2-THOR: An Interactive 3D Environment for Visual AI”. arXiv. http://arxiv.org/abs/1712.05474v4.
Harvard
Kolve, E. et al. (2017) “AI2-THOR: An Interactive 3D Environment for Visual AI”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1712.05474v4.
Vancouver
1. Kolve E, Mottaghi R, Han W, et al (2017) AI2-THOR: An Interactive 3D Environment for Visual AI. arXiv

BibTeX

@article{kolve2017ai2,
  title = {AI2-THOR: An Interactive 3D Environment for Visual AI},
  author = {Kolve, Eric and Mottaghi, Roozbeh and Han, Winson and VanderBilt, Eli and Weihs, Luca and Herrasti, Alvaro and Deitke, Matt and Ehsani, Kiana and Gordon, Daniel and Zhu, Yuke and Kembhavi, Aniruddha and Gupta, Abhinav and Farhadi, Ali},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1712.05474v4},
  eprint = {1712.05474}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors