Target-driven visual navigation in indoor scenes using deep reinforcement learning
Yuke ZhuRoozbeh MottaghiEric KolveJoseph J. LimAbhinav GuptaLi Fei-FeiAli Farhadi
Introduces a goal-conditioned deep reinforcement learning framework and the AI2-THOR simulation platform, allowing agents to find visual targets across unseen indoor scenes and transfer learned policies to physical robots without explicit 3D mapping.
Autonomous visual navigation in indoor environments is a critical capability for service and mobile robotics. However, standard deep reinforcement learning techniques suffer from severe data inefficiency and an inability to generalize to new goals without costly, time-consuming retraining from scratch. These limitations have historically restricted reinforcement learning models to constrained game environments and made real-world robotic deployment impractical.
The article addresses this challenge by developing and evaluating an end-to-end framework for target-driven visual navigation. The objective is to enable an agent to navigate to a designated visual target using only camera images as input, without pre-existing maps, explicit three-dimensional reconstruction, or target-specific retraining.
To achieve this, the authors introduced a deep siamese actor-critic network that takes both the agent's current visual observation and a picture of the target goal as inputs. By separating generic visual layers from scene-specific layers, the model shares navigational knowledge across different targets and scenes. To bypass the physical risks and slow pace of training on physical hardware, the authors created the AI2-THOR simulation framework, which comprises 32 photo-realistic indoor 3D environments with accurate physics. The system was trained in parallel using 100 threads across simulated scenes and subsequently validated in both continuous simulation space and on a real SCITOS mobile robot.
The evaluation yielded several key findings regarding efficiency, generalization, and practical deployment. First, the proposed model achieved an average trajectory length of 210.7 steps after 100 million frames, outperforming standard reinforcement learning baselines such as four-thread A3C (723.5 steps) and Q-learning (2,539.2 steps). Second, the model successfully generalized to unseen targets within known environments, showing particularly high success rates for targets located close to previously trained areas. Third, knowledge transferred effectively to entirely new scenes: freezing generic layers and training only scene-specific layers accelerated adaptation as the number of prior training scenes grew. Finally, in physical robot trials, transferring simulation-trained weights and fine-tuning with real images produced optimal navigation policies 44% faster than training from scratch.
These results demonstrate that target-driven reinforcement learning can bridge the gap between simulation and real-world indoor navigation. By embedding goals directly into the input space, organizations can eliminate the high computational and operational costs of retraining models for every new task or floor plan. Furthermore, pre-training in realistic 3D simulation platforms significantly mitigates the hardware wear, safety risks, and data collection bottlenecks associated with real-world robot training.
Based on these findings, decision-makers should consider utilizing high-fidelity simulation frameworks as a low-cost, scalable pipeline for robotic navigation training before deploying models onto physical devices. Future work supported by the article includes expanding the diversity and quantity of high-quality 3D training scenes, as well as extending the learning architecture to handle complex physical object manipulations such as opening doors and grasping items.
Readers should note certain limitations when evaluating this approach. While the model performed robustly across lighting and layout variations, real-world evaluations were conducted on a single robot in a relatively small room with discrete grid-based waypoints. In addition, navigating continuous physical space required substantially more training frames (approximately 50 million) compared to discrete actions. Nonetheless, the high consistency between simulated and physical performance provides strong confidence in the viability of target-driven visual navigation frameworks.
- Paper: Trust Region Policy Optimization, John Schulman et al. (2015). Trust Region Policy Optimization provides the theoretical and algorithmic foundations for deep actor-critic policy optimization that the target-driven navigation model builds upon.
- Paper: High-Dimensional Continuous Control Using Generalized Advantage Estimation, John Schulman et al. (2016). Generalized Advantage Estimation develops variance-reduction techniques for actor-critic architectures essential to stabilizing reinforcement learning policies trained directly from sensory inputs.
- Paper: Continuous control with deep reinforcement learning, T. Lillicrap et al. (2015). Deep Deterministic Policy Gradient demonstrates how deep neural networks can serve as end-to-end actor-critic controllers directly from pixel inputs, laying key groundwork for visual agent policies.
- Paper: End-to-End Training of Deep Visuomotor Policies, Sergey Levine et al. (2015). This work establishes the paradigm of end-to-end training of deep visuomotor policies from raw camera images, serving as a direct conceptual precursor to end-to-end visual navigation.
- Paper: Deterministic Policy Gradient Algorithms, David Silver et al. (2014). This foundational text establishes the deterministic policy gradient framework that underpins continuous actor-critic reinforcement learning algorithms.
- Paper: Benchmarking Deep Reinforcement Learning for Continuous Control, Yan Duan et al. (2016). This paper establishes standardized benchmarks and comparative performance baselines for deep reinforcement learning continuous control algorithms used in robotic simulations.
- Paper: Learning to Navigate in Complex Environments, Piotr Mirowski et al. (2017). This paper extends end-to-end visual navigation by incorporating auxiliary depth and loop-closure prediction tasks to enhance spatial awareness and sample efficiency in complex 3D environments.
- Paper: Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments, Peter Anderson et al. (2017). This work generalizes visual indoor navigation to follow natural language instructions in realistic multi-room photorealistic environments.
- Paper: Habitat: A Platform for Embodied AI Research, Manolis Savva et al. (2019). Habitat advances the 3D embodied simulation paradigm introduced by AI2-THOR to massively scalable, high-frame-rate photorealistic environments for benchmarking embodied agents.
- Paper: Reinforcement Learning with Unsupervised Auxiliary Tasks, Max Jaderberg et al. (2017). This work develops the UNREAL agent architecture to improve learning efficiency and representation learning in 3D visual environments through unsupervised auxiliary tasks.
- Paper: Domain randomization for transferring deep neural networks from simulation to the real world, Josh Tobin et al. (2017). Domain randomization expands the simulation-to-real transfer methodology used in simulated 3D environments to zero-shot real-world physical deployment.
- Paper: Sim-to-Real Transfer of Robotic Control with Dynamics Randomization, Xue Bin Peng et al. (2017). This paper builds on sim-to-real transfer principles by randomizing physical dynamics during simulation training to bridge the reality gap without real-world training data.
- Paper: Hindsight Experience Replay, Marcin Andrychowicz et al. (2017). Hindsight Experience Replay improves goal-conditioned reinforcement learning sample efficiency by replaying past trajectories with substituted goals.
- Paper: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor, Tuomas Haarnoja et al. (2018). Soft Actor-Critic develops a maximum entropy off-policy framework that substantially enhances sample efficiency and exploration stability in continuous control.
- Paper: Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning, Viktor Makoviychuk et al. (2021). Isaac Gym takes simulation-based robot learning to GPU-native execution, drastically scaling up the data generation efficiency for embodied reinforcement learning.
