Target-driven visual navigation in indoor scenes using deep reinforcement learning

Yuke ZhuRoozbeh MottaghiEric KolveJoseph J. LimAbhinav GuptaLi Fei-FeiAli Farhadi

article2016IEEE International Conference on Robotics and Automation1,743 citations

Introduces a goal-conditioned deep reinforcement learning framework and the AI2-THOR simulation platform, allowing agents to find visual targets across unseen indoor scenes and transfer learned policies to physical robots without explicit 3D mapping.

Listen

Autonomous visual navigation in indoor environments is a critical capability for service and mobile robotics. However, standard deep reinforcement learning techniques suffer from severe data inefficiency and an inability to generalize to new goals without costly, time-consuming retraining from scratch. These limitations have historically restricted reinforcement learning models to constrained game environments and made real-world robotic deployment impractical.

The article addresses this challenge by developing and evaluating an end-to-end framework for target-driven visual navigation. The objective is to enable an agent to navigate to a designated visual target using only camera images as input, without pre-existing maps, explicit three-dimensional reconstruction, or target-specific retraining.

To achieve this, the authors introduced a deep siamese actor-critic network that takes both the agent's current visual observation and a picture of the target goal as inputs. By separating generic visual layers from scene-specific layers, the model shares navigational knowledge across different targets and scenes. To bypass the physical risks and slow pace of training on physical hardware, the authors created the AI2-THOR simulation framework, which comprises 32 photo-realistic indoor 3D environments with accurate physics. The system was trained in parallel using 100 threads across simulated scenes and subsequently validated in both continuous simulation space and on a real SCITOS mobile robot.

The evaluation yielded several key findings regarding efficiency, generalization, and practical deployment. First, the proposed model achieved an average trajectory length of 210.7 steps after 100 million frames, outperforming standard reinforcement learning baselines such as four-thread A3C (723.5 steps) and Q-learning (2,539.2 steps). Second, the model successfully generalized to unseen targets within known environments, showing particularly high success rates for targets located close to previously trained areas. Third, knowledge transferred effectively to entirely new scenes: freezing generic layers and training only scene-specific layers accelerated adaptation as the number of prior training scenes grew. Finally, in physical robot trials, transferring simulation-trained weights and fine-tuning with real images produced optimal navigation policies 44% faster than training from scratch.

These results demonstrate that target-driven reinforcement learning can bridge the gap between simulation and real-world indoor navigation. By embedding goals directly into the input space, organizations can eliminate the high computational and operational costs of retraining models for every new task or floor plan. Furthermore, pre-training in realistic 3D simulation platforms significantly mitigates the hardware wear, safety risks, and data collection bottlenecks associated with real-world robot training.

Based on these findings, decision-makers should consider utilizing high-fidelity simulation frameworks as a low-cost, scalable pipeline for robotic navigation training before deploying models onto physical devices. Future work supported by the article includes expanding the diversity and quantity of high-quality 3D training scenes, as well as extending the learning architecture to handle complex physical object manipulations such as opening doors and grasping items.

Readers should note certain limitations when evaluating this approach. While the model performed robustly across lighting and layout variations, real-world evaluations were conducted on a single robot in a relatively small room with discrete grid-based waypoints. In addition, navigating continuous physical space required substantially more training frames (approximately 50 million) compared to discrete actions. Nonetheless, the high consistency between simulated and physical performance provides strong confidence in the viability of target-driven visual navigation frameworks.

  • Paper: Trust Region Policy Optimization, John Schulman et al. (2015). Trust Region Policy Optimization provides the theoretical and algorithmic foundations for deep actor-critic policy optimization that the target-driven navigation model builds upon.
  • Paper: High-Dimensional Continuous Control Using Generalized Advantage Estimation, John Schulman et al. (2016). Generalized Advantage Estimation develops variance-reduction techniques for actor-critic architectures essential to stabilizing reinforcement learning policies trained directly from sensory inputs.
  • Paper: Continuous control with deep reinforcement learning, T. Lillicrap et al. (2015). Deep Deterministic Policy Gradient demonstrates how deep neural networks can serve as end-to-end actor-critic controllers directly from pixel inputs, laying key groundwork for visual agent policies.
  • Paper: End-to-End Training of Deep Visuomotor Policies, Sergey Levine et al. (2015). This work establishes the paradigm of end-to-end training of deep visuomotor policies from raw camera images, serving as a direct conceptual precursor to end-to-end visual navigation.
  • Paper: Deterministic Policy Gradient Algorithms, David Silver et al. (2014). This foundational text establishes the deterministic policy gradient framework that underpins continuous actor-critic reinforcement learning algorithms.
  • Paper: Benchmarking Deep Reinforcement Learning for Continuous Control, Yan Duan et al. (2016). This paper establishes standardized benchmarks and comparative performance baselines for deep reinforcement learning continuous control algorithms used in robotic simulations.
Cover for Target-driven visual navigation in indoor scenes using deep reinforcement learning

Abstract

Two less addressed issues of deep reinforcement learning are (1) lack of generalization capability to new target goals, and (2) data inefficiency i.e., the model requires several (and often costly) episodes of trial and error to converge, which makes it impractical to be applied to real-world scenarios. In this paper, we address these two issues and apply our model to the task of target-driven visual navigation. To address the first issue, we propose an actor-critic model whose policy is a function of the goal as well as the current state, which allows to better generalize. To address the second issue, we propose AI2-THOR framework, which provides an environment with high-quality 3D scenes and physics engine. Our framework enables agents to take actions and interact with objects. Hence, we can collect a huge number of training samples efficiently.

We show that our proposed method (1) converges faster than the state-of-the-art deep reinforcement learning methods, (2) generalizes across targets and across scenes, (3) generalizes to a real robot scenario with a small amount of fine-tuning (although the model is trained in simulation), (4) is end-to-end trainable and does not need feature engineering, feature matching between frames or 3D reconstruction of the environment.

The supplementary video can be accessed at the following link: this https URL.

Table of Contents

  • I Introduction
  • II RELATED WORK
  • III AI2-THOR FRAMEWORK
  • IV TARGET-DRIVEN NAVIGATION MODEL
  • IV-A Problem Statement
  • IV-B Problem Formulation
  • IV-C Learning Setup
  • IV-C1 Action space
  • IV-C2 Observations and Goals
  • IV-C3 Reward design
  • IV-D Model
  • IV-E Training Protocol
  • IV-F Network Architectures
  • V EXPERIMENTS
  • V-A Navigation Results
  • V-B Generalization Across Targets
  • V-C Generalization Across Scenes
  • V-D Continuous Space
  • V-E Robot Experiment
  • VI CONCLUSIONS
  • References

Knowls

  1. Knowl 1 — Target-Driven Visual Navigation Formulation

    definition

    Target-driven visual navigation is framed as finding the minimum-length sequence of actions to steer an agent from its current position and orientation to a goal location specified purely by an RGB image gg taken from the target viewpoint. Formally, the objective is to learn a stochastic policy function π(at∣st,g;u)\pi(a_t \mid s_t, g; \mathbf{u}), parameterized by weights u\mathbf{u}, that defines a probability distribution over discrete actions ata_t conditioned on both the current state observation representation sts_t (an RGB image or sequence of images) and the target image representation gg:

    at∼π(st,g∣u)a_t \sim \pi(s_t, g \mid \mathbf{u})

    Conditioning the policy explicitly on gg allows the agent to generalize to new navigation targets without altering or re-training the network parameters u\mathbf{u}.

  2. Knowl 2 — Deep Siamese Actor-Critic Architecture for Visual Navigation

    model/method

    The target-driven visual navigation model is a deep siamese actor-critic network composed of two main stages:

    1. Generic Siamese Feature Embedding: The current state observation sts_t (consisting of 4 stacked history RGB frames of dimension 224×224×3224 \times 224 \times 3) and the target image gg (224×224×3224 \times 224 \times 3) are independently processed by two weight-shared branches. Each branch uses an ImageNet-pretrained ResNet-50 network (with the classification softmax layer removed and parameters frozen during training) to produce a 2048-dimensional feature vector. Each 2048-dimensional vector is passed through a fully-connected layer to map both inputs into a shared 512-dimensional embedding space.

    2. Embedding Fusion and Scene-Specific Layers: The two 512-dimensional embeddings are concatenated into a 1024-dimensional vector and projected through an embedding fusion layer to form a 512-dimensional joint state-goal representation. This joint representation is routed into scene-specific branches consisting of two fully-connected layers (512 units each) tailored to individual indoor scenes. The scene-specific network terminates in two separate output heads: a policy branch producing a 4-dimensional softmax probability distribution over actions, and a value branch producing a 1-dimensional scalar baseline estimate for advantage actor-critic updates.

  3. Knowl 3 — AI2-THOR Simulation Framework for Visual Interaction and Navigation

    experimental setup

    The House Of inteRactions (AI2-THOR) is a simulation framework that couples a Unity 3D physics and rendering engine directly with deep learning frameworks (e.g., TensorFlow) via Python APIs for online decision-making. The environment includes 32 photorealistic 3D indoor scenes spanning 4 residential scene types: kitchens, living rooms, bedrooms, and bathrooms. Each scene contains an average of 68 interactable object instances with realistic textures, lighting, physical collisions, and state changes (e.g., opening cabinets or sliding objects).

  4. Knowl 4 — Asynchronous Multi-Target Policy Gradient Training Protocol

    algorithm

    The navigation policy is trained using an asynchronous advantage actor-critic (A3C) protocol modified for multi-task multi-target learning:

    Input: Generic siamese layer parameters u_generic, Scene-specific parameters {u_scene_i}
    Input: Global shared RMSProp optimizer with learning rate 7e-4
    Initialize 100 parallel worker threads, each assigned an individual navigation target g in a designated scene i
    for each parallel thread do
        while global training frames < 100 million do
            Sample target g and random starting position in scene i
            Accumulate trajectory states, actions, and rewards under current policy pi(a | s_t, g; u)
            Compute policy gradients and value loss gradients via n-step advantage estimation
            Compute parameter updates delta_generic and delta_scene_i
            Asynchronously update generic siamese parameters u_generic using delta_generic
            Asynchronously update scene-specific parameters u_scene_i using delta_scene_i
        end while
    end for

    Because the generic siamese layers are updated by gradients from all targets across all scenes, they learn domain-general visual representations, while scene-specific layers capture layout-specific topographies.

  5. Knowl 5 — Action Space, Motion Noise, and Reward Function Specification

    experimental setup

    The reinforcement learning environment is configured with the following parameters:

    • Action Space: 4 discrete command-level actions: move forward (0.5 meters), move backward (0.5 meters), turn left (90∘90^\circ), and turn right (90∘90^\circ).
    • System Uncertainty: Actuation noise is modeled by adding zero-mean Gaussian noise to step translations with variance 0.01 (i.e., N(0,0.01)\mathcal{N}(0, 0.01)) and to turns with variance 1.0 (i.e., N(0,1.0)\mathcal{N}(0, 1.0) in degrees).
    • Reward Structure: A terminal reward of +10.0+10.0 is provided upon reaching the target viewpoint, combined with an immediate step penalty of −0.01-0.01 at every action step to penalize trajectory duration.
  6. Knowl 6 — Navigation Trajectory Length Comparison with Baselines

    data/table

    The navigation performance of the target-driven model was evaluated against heuristic and reinforcement learning baselines across 100 random targets in 20 simulated indoor scenes. Each model was evaluated over 10 test episodes per target from random initial locations, capped at a maximum of 10,000 steps per episode. All learning models were trained for 100 million total frames across worker threads.

    Type Method Avg. Trajectory Length
    Heuristic Random walk 2744.3
    Heuristic Shortest path (Upper Bound) 17.6
    Purpose-built RL One-step Q 2539.2
    Purpose-built RL A3C (1 thread) 1241.3
    Purpose-built RL A3C (4 threads) 723.5
    Target-driven RL Single branch (No scene-specific heads) 581.6
    Target-driven RL (Ours) Final (Siamese + Scene-specific heads) 210.7

    The target-driven architecture with scene-specific heads achieves an average trajectory length of 210.7 steps, outperforming standard single-target A3C and a single-branch target-driven model.

  7. Knowl 7 — Preservation of Spatial Layout in Generic Siamese Embeddings

    empirical result

    The 512-dimensional representations generated by the generic siamese embedding layers preserve topological and spatial layouts of the environment without requiring explicit coordinate supervision. A 2D t-SNE projection of the embedding vectors corresponding to first-person observations across varied locations and four camera orientations (0∘,90∘,180∘,270∘0^\circ, 90^\circ, 180^\circ, 270^\circ) aligns with the actual physical 2D floorplan. The Pearson correlation coefficient between pairwise Euclidean distances of the embedding representations and the pairwise Euclidean distances of their true 2D scene coordinates is r=0.62r = 0.62 (p<0.001p < 0.001).

  8. Knowl 8 — Target and Scene Transfer Generalization Dynamics

    empirical result

    The target-driven model exhibits zero-shot and few-shot transfer across targets and scenes:

    1. Target Generalization: When evaluated on unseen targets in known scenes (measured by the success rate of reaching the target within 500 steps across 100 test episodes), the success rate increases monotonically as the number of pre-trained targets in the scene grows from 1 to 8. Furthermore, success rates are strictly higher for new targets situated within smaller topological step distances (1 step vs. 2, 4, or 8 steps) from pre-trained targets.
    2. Scene Generalization: When transferring to 4 previously unseen scenes, adapting only the scene-specific layers while freezing pre-trained generic siamese layers accelerates policy convergence. The convergence rate increases monotonically with the number of scenes (from 1 up to 16) used to pre-train the generic siamese layers.
  9. Knowl 9 — Continuous Space Navigation and Sim-to-Real Transfer

    empirical result

    The target-driven reinforcement learning framework generalizes to continuous physics engines and real-world mobile robot platforms:

    1. Continuous Physics Space: When tested in a continuous 3D simulation with dynamic forces, collisions, and sliding friction for a balcony door-finding task, the model learns a policy reaching the goal in an average of 15 steps (after ~50M training frames), compared to 719 steps for a random-walk agent.
    2. Sim-to-Real Transfer (SCITOS Mobile Robot): In a real-world indoor room discretized into 28 physical locations (with 4 orientations each, totaling 112 viewpoints), transferring generic siamese parameters pre-trained on 20 simulated AI2-THOR scenes and fine-tuning only scene-specific parameters on real robot images achieves convergence to optimal navigation policies 44% faster than training the entire network from scratch on real images.

Coverage note — None was omitted; all key architectural components, training procedures, baseline comparisons, generalization experiments, and sim-to-real robotic transfers are fully covered.

References

  1. 1.S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” JMLR, 2016.
  2. 2.V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al., “Human-level control through deep reinforcement learning,” Nature, 2015.
  3. 3.V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in ICML, 2016.
  4. 4.J. Borenstein and Y. Koren, “Real-time obstacle avoidance for fast mobile robots,” IEEE Trans. on Systems, Man and Cybernetics, 1989.
  5. 5.——, “The vector field histogram-fast obstacle avoidance for mobile robots,” IEEE Trans. on Robotics and Automation, 1991.
  6. 6.D. Kim and R. Nevatia, “Simbolic navigation with a generic map,” in IEEE Workshop on Vision for Robots, 1995.
  7. 7.G. U. G. Oriolo and M. Vendittelli, “On-line map building and navigation for autonomous mobile robots,” in ICRA, 1995.
  8. 8.R. Sim and J. J. Little, “Autonomous vision-based exploration and mapping using hybrid maps and rao-blackwellised particle filters,” in IROS, 2006.
  9. 9.D. Wooden, “A guide to vision-based map building,” IEEE Robotics and Automation Magazine, 2006.
  10. 10.A. J. Davison, “Real time simultaneous localisation and mapping with a single camera,” in ICCV, 2003.
  11. 11.M. Tomono, “3-d object map building using dense object models with sift-based recognition features,” in IROS, 2006.
  12. 12.K. Kidono, J. Miura, and Y. Shirai, “Autonomous visual navigation of a mobile robot using a human guided experience,” Robotics and Autonomous Systems, 2002.
  13. 13.E. Royer, J. Bom, M. Dhome, B. Thuillot, M. Lhuillier, and F. Marmoiton, “Outdoor autonomous navigation using monocular vision,” in IROS, 2005.
  14. 14.H. Haddad, M. Khatib, S. Lacroix, and R. Chatila, “Reactive navigation in outdoor environments using potential fields,” in ICRA, 1998.
  15. 15.S. Lenser and M. Veloso, “Visual sonar: Fast obstacle avoidance using monocular vision,” in IROS, 2003.
  16. 16.A. Remazeilles, F. Chaumette, and P. Gros, “Robot motion control from a visual memory,” in ICRA, 2004.
  17. 17.P. Saeedi, P. D. Lawrence, and D. G. Lowe, “Vision-based 3-d trajectory tracking for unknown environments,” IEEE Trans. on Robotics, 2006.
  18. 18.F. Bonin-Font, A. Ortiz, and G. Oliver, “Visual navigation for mobile robots: A survey,” J. of Intelligent and Robotic Systems, 2008.
  19. 19.K. Konolige, J. Bowman, J. Chen, P. Mihelich, M. Calonder, V. Lepetit, and P. Fua, “View-based maps,” Intl. J. of Robotics Research, 2010.
  20. 20.S. Phillips, A. Jaegle, and K. Daniilidis, “Fast, robust, continuous monocular egomotion computation,” in ICRA, 2016.
  21. 21.C. McManus, B. Upcroft, and P. Newman, “Scene signatures: Localised and point-less features for localisation,” in RSS, 2014.
  22. 22.C. Linegar, W. Churchill, and P. Newman, “Made to measure: Bespoke landmarks for 24-hour, all-weather localisation with a camera,” in ICRA, 2016.
  23. 23.N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” in ICRA, 2004.
  24. 24.J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,” Neural Networks, 2008.
  25. 25.J. Michels, A. Saxena, and A. Ng, “High speed obstacle avoidance using monocular vision and reinforcement learning,” in ICML, 2005.
  26. 26.H. J. Kim, M. I. Jordan, S. Sastry, and A. Y. Ng, “Autonomous helicopter flight via reinforcement learning,” in NIPS, 2004.
  27. 27.T. Kollar and N. Roy, “Trajectory optimization using reinforcement learning for map exploration,” Intl. J. of Robotics Research, 2008.
  28. 28.A. M. S. Barreto, D. Precup, and J. Pineau, “Practical kernel-based reinforcement learning,” JMLR, 2016.
  29. 29.Y. Liang, M. C. Machado, E. Talvitie, and M. Bowling, “State of the art control of atari games using shallow reinforcement learning,” in AAMAS, 2016.
  30. 30.D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al., “Mastering the game of go with deep neural networks and tree search,” Nature, 2016.
  31. 31.A. A. Rusu, S. G. Colmenarejo, Ç. Gülçehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell, “Policy distillation,” in ICLR, 2016.
  32. 32.E. Parisotto, L. J. Ba, and R. Salakhutdinov, “Actor-mimic: Deep multitask and transfer reinforcement learning,” in ICLR, 2016.
  33. 33.J. Wu, I. Yildirim, J. J. Lim, W. T. Freeman, and J. B. Tenenbaum, “Galileo: Perceiving physical object properties by integrating a physics engine with deep learning,” in NIPS, 2015.
  34. 34.R. Mottaghi, H. Bagherinezhad, M. Rastegari, and A. Farhadi, “Newtonian image understanding: Unfolding the dynamics of objects in static images,” in CVPR, 2016.
  35. 35.R. Mottaghi, M. Rastegari, A. Gupta, and A. Farhadi, ““what happens if...” learning to predict the effect of forces in images,” in ECCV, 2016.
  36. 36.M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, “The arcade learning environment: An evaluation platform for general agents,” J. of Artificial Intelligence Research, vol. 47, pp. 253–279, 2013.
  37. 37.M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jakowski, “Vizdoom: A doom-based ai research platform for visual reinforcement learning,” in IEEE Conference on Computational Intelligence and Games, 2016.
  38. 38.A. Lerer, S. Gross, and R. Fergus, “Learning physical intuition of block towers by example,” in ICML, 2016.
  39. 39.M. Johnson, K. Hofmann, T. Hutton, and D. Bignell, “The malmo platform for artificial intelligence experimentation,” in Intl. Joint Conference on Artificial Intelligence, 2016.
  40. 40.A. Handa, V. Patraucean, S. Stent, and R. Cipolla, “Scenenet: An annotated model generator for indoor scene understanding,” in ICRA, 2016.
  41. 41.B. Wymann, E. Espie, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner, “TORCS, the open racing car simulator, v1.3.5,” http://www.torcs.org, 2013.
  42. 42.G. Ros, L. Sellart, J. Materzynska, D. Vazquez, and A. Lopez, “The SYNTHIA Dataset: A large collection of synthetic images for semantic segmentation of urban scenes,” in CVPR, 2016.
  43. 43.A. Gaidon, Q. Wang, Y. Cabon, and E. Vig, “Virtual worlds as proxy for multi-object tracking analysis,” in CVPR, 2016.
  44. 44.M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al., “Tensorflow: Large-scale machine learning on heterogeneous distributed systems,” arXiv preprint arXiv:1603.04467, 2016.
  45. 45.J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The Intl. J. of Robotics Research, 2013.
  46. 46.B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman, “Building machines that learn and think like people,” arXiv preprint arXiv:1604.00289, 2016.
  47. 47.S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in CVPR, 2005.
  48. 48.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016.
  49. 49.L. van der Maaten and G. E. Hinton, “Visualizing data using t-sne,” JMLR, 2008.
  50. 50.M. J.-Y. Chung, A. Pronobis, M. Cakmak, D. Fox, and R. P. N. Rao, “Autonomous question answering with mobile robots in human-populated environments,” in IROS, 2016.

Citation

MLA
Zhu, Y., et al. “Target-driven Visual Navigation in Indoor Scenes Using Deep Reinforcement Learning”. arXiv, 2016, http://arxiv.org/abs/1609.05143v1.
APA
Zhu, Y., Mottaghi, R., Kolve, E., Lim, J. J., Gupta, A., Fei-Fei, L., & Farhadi, A. (2016). Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning. arXiv. http://arxiv.org/abs/1609.05143v1
Chicago
Zhu, Y., R. Mottaghi, E. Kolve, et al. 2016. “Target-driven Visual Navigation in Indoor Scenes Using Deep Reinforcement Learning”. arXiv. http://arxiv.org/abs/1609.05143v1.
Harvard
Zhu, Y. et al. (2016) “Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1609.05143v1.
Vancouver
1. Zhu Y, Mottaghi R, Kolve E, Lim JJ, Gupta A, Fei-Fei L, Farhadi A (2016) Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning. arXiv

BibTeX

@article{zhu2016target,
  title = {Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning},
  author = {Zhu, Yuke and Mottaghi, Roozbeh and Kolve, Eric and Lim, Joseph J. and Gupta, Abhinav and Fei-Fei, Li and Farhadi, Ali},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1609.05143v1},
  eprint = {1609.05143}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF