Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation
Yicong HongZun WangQi WuStephen Gould
Proposes a candidate waypoint predictor that enables discrete Vision-and-Language Navigation models to operate directly in continuous environments, achieving state-of-the-art performance on the R2R-CE and RxR-CE benchmarks.
Vision-and-language navigation tasks require autonomous agents to interpret human language instructions and navigate through unfamiliar environments. While training agents in discrete spaces—where the environment is represented as a pre-mapped graph of locations—has led to significant algorithmic advances, these agents fail to operate in realistic, continuous spaces where an agent must execute low-level movements without a predefined map. Navigating in continuous spaces typically leaves agents struggling to identify obstacles and accessible paths, resulting in a substantial performance drop of approximately 20% compared to discrete navigation.
The article evaluates whether a visual candidate waypoints predictor can bridge this domain gap by generating accessible local destinations on the fly in continuous environments. This approach aims to demonstrate that agents designed for discrete navigation can be directly transferred to continuous settings using high-level directional actions instead of tedious low-level motor controls.
The authors developed a predictor using visual encoders and a spatial attention network trained on refined 3D indoor spatial graphs across 90 simulated building environments. This module predicts obstacle-free local candidate waypoints within a three-meter radius around the agent. The authors evaluated two distinct navigation architectures—a cross-modal matching agent and a visiolinguistic transformer agent—on continuous navigation benchmarks. They also applied a waypoint augmentation technique during training to diversify the paths and visual perspectives the agents experienced.
The key findings show that supplying predicted waypoints significantly improves continuous navigation. Incorporating waypoint prediction reduced the discrete-to-continuous performance gap by 11.76% in path efficiency for the cross-modal matching agent and by 18.24% for the transformer-based agent. Training with waypoint augmentation further enhanced generalization, allowing the simpler agent to reach a 40.80% success rate in unseen test spaces, matching or exceeding graph-based performance. Additionally, the approach reduced obstacle collision rates from 15% to 7% and established new state-of-the-art benchmarks on standard continuous navigation datasets (R2R-CE and RxR-CE), achieving a 38% to 42% success rate on unseen test paths.
These results demonstrate that the core challenge of continuous navigation is identifying navigable space rather than executing complex motor actions. Decoupling waypoint prediction from instruction following enables standard imitation learning to succeed without complex reinforcement learning pipelines. This dramatically lowers compute requirements, reducing training overhead from 64 graphics processors over five days to a single processor in roughly three and a half days while achieving superior navigation accuracy.
Teams developing embodied artificial intelligence and autonomous service robots should adopt local waypoint prediction modules to transition high-level planning models into continuous real-world environments. Future development should focus on extending waypoint predictors to support language-conditioned navigation and testing the framework across broader embodied tasks, such as object-goal and audio-visual navigation.
A current limitation is that the predictor may occasionally fail to identify valid pathways across challenging architectural features, such as stairs, leading to failed trajectories. Nevertheless, confidence in the primary findings remains high, as the methodology demonstrated strong, consistent improvements across multiple model architectures and standardized evaluation splits.
- Paper: Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments, Peter Anderson et al. (2017). This seminal paper establishes the Room-to-Room benchmark and baseline frameworks for instruction-following in indoor environments, creating the core vision-and-language navigation foundation upon which the source paper builds.
- Paper: Habitat: A Platform for Embodied AI Research, Manolis Savva et al. (2019). It introduces the Habitat simulation platform that provides the underlying 3D photorealistic environments and continuous physics engine essential for evaluating continuous navigation.
- Paper: Target-driven visual navigation in indoor scenes using deep reinforcement learning, Yuke Zhu et al. (2016). It pioneers target-driven visual indoor navigation using deep reinforcement learning across continuous and simulated spaces, introducing foundational concepts for navigating without predefined maps.
- Paper: Learning to Navigate in Complex Environments, Piotr Mirowski et al. (2017). It demonstrates how auxiliary predictive tasks—such as depth and loop closure—allow navigation agents to capture spatial structure directly from sensory inputs.
- Paper: Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation, Shizhe Chen et al. (2022). It establishes a dual-scale graph transformer for vision-and-language navigation that informs how agents balance coarse topological reasoning with local action selection.
- Paper: Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation, Hanqing Wang et al. (2022). It introduces cycle-consistent learning and data augmentation strategies between instruction followers and speakers that underpin the multi-modal agent baselines adapted in the source.
- Paper: PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation, Jialu Li et al. (2023). It extends vision-and-language navigation research by generating text-conditioned panoramic environments to enhance agent generalization across unseen layouts.
- Paper: GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation, Mukul Khanna et al. (2024). It broadens embodied navigation benchmarks from single-episode instruction following to multimodal, lifelong continuous navigation across diverse goal formats.
- Paper: SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World, Kiana Ehsani et al. (2024). It applies scalable imitation learning over shortest-path simulated heuristics to execute continuous navigation and manipulation tasks on physical robots without complex reinforcement learning.
- Paper: Renderable Neural Radiance Map for Visual Navigation, Obin Kwon et al. (2023). It advances visual continuous navigation by creating 2D renderable neural radiance spatial memories for real-time localization and path planning in novel environments.
- Paper: Kosmos-2: Grounding Multimodal Large Language Models to the World, Zhiliang Peng et al. (2023). It scales multimodal grounding by integrating discrete spatial coordinate tokens into large language models to bridge the gap between high-level language and low-level visual grounding.
