Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation
Shizhe ChenPierre-Louis GuhurMakarand TapaswiCordelia SchmidIvan Laptev
Introduces DUET, a dual-scale graph transformer architecture that dynamically fuses coarse topological maps with fine-grained local visual observations to solve long-term action planning and instruction-following challenges in embodied AI.
Autonomous navigation guided by natural language is a foundational capability for mobile robots and virtual assistants operating in homes and workplaces. However, current systems struggle in previously unseen environments when users give high-level, goal-oriented commands such as finding and interacting with a remote object. Existing methods typically restrict an agent to making myopic step-by-step local moves or rely on condensed memory formats, which makes exploring large spaces inefficient, complicates backtracking, and lacks the fine-grained visual detail required to identify specific target objects.
The article develops and evaluates a new navigation framework called the Dual-scale graph Transformer (DUET). The primary objective is to demonstrate that dynamically combining coarse-scale global spatial reasoning with fine-scale local visual grounding substantially improves an autonomous agent's ability to plan long-term routes and locate target objects in unfamiliar environments.
To accomplish this, the authors construct an online topological map that tracks visited and unvisited navigable locations as the agent moves. DUET uses a dual-scale architecture powered by graph transformers: a coarse-scale encoder reasons over the global map and its structural connectivity to select long-range navigation targets, while a fine-scale encoder processes detailed panoramic and object features at the current location to evaluate immediate actions and pinpoint target objects. The system dynamically fuses predictions from both scales. Training combines pretraining on auxiliary multimodal tasks with policy learning guided by an interactive pseudo-demonstrator to correct errors during simulated exploration. The approach was evaluated on three established benchmarks: REVERIE and SOON for goal-oriented navigation, and R2R for detailed step-by-step navigation.
DUET achieves substantial performance gains across all benchmarks. On the unseen test split of the REVERIE benchmark, DUET increased the navigation success rate by 22.11 percentage points over the prior state of the art, rising from 30.40% to 52.51%, while improving target object grounding success penalized by path length from 13.08% to 22.06%. On the SOON benchmark, DUET improved the unseen test success rate from 12.90% to 33.44%, representing a gain of over 20 percentage points. On the step-by-step R2R dataset, DUET established a new state of the art by lifting the unseen test success rate from 65% to 69%. Ablation experiments confirmed that both the coarse-scale global map reasoning and fine-scale local object representations are essential, and that incorporating graph topology into self-attention directly improves navigation path efficiency.
These findings indicate that autonomous agents can overcome exploration bottlenecks and costly backtracking by decoupling long-term route planning from immediate visual grounding. For operational applications, this capability reduces navigation failures, shortens execution timelines in complex layouts, and enhances human-robot interaction by allowing agents to follow natural, high-level commands without requiring tedious step-by-step guidance.
Organizations developing embodied artificial intelligence and autonomous mobile systems should consider adopting dual-scale topological architectures and interactive demonstrator training strategies. Next steps include evaluating DUET in continuous, non-discrete physical environments and validating its performance on physical robotic platforms operating under real-world sensor noise and dynamic obstacles.
Readers should note that the evaluation is conducted within discrete, graph-based simulation environments with access to reliable orientation and location coordinates. While confidence in the benchmark improvements is very high due to consistent gains across multiple standardized datasets, stakeholders should exercise caution when translating these results directly to real-world hardware where mapping errors, camera blur, and moving obstacles may affect performance.
- Paper: Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments, Peter Anderson et al. (2017). Introduces the foundational Vision-and-Language Navigation task and Room-to-Room benchmark that DUET directly builds upon and targets.
- Paper: Target-driven visual navigation in indoor scenes using deep reinforcement learning, Yuke Zhu et al. (2016). Establishes target-driven visual navigation in photo-realistic indoor simulations, providing the core framework for goal-directed embodied agents.
- Paper: Habitat: A Platform for Embodied AI Research, Manolis Savva et al. (2019). Introduces the standardized simulation platform and evaluation protocols widely adopted for training and evaluating embodied navigation agents.
- Paper: Learning to Navigate in Complex Environments, Piotr Mirowski et al. (2017). Demonstrates how auxiliary predictive tasks and memory structures enhance spatial awareness and path planning in 3D navigation.
- Paper: EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought, Yao Mu et al. (2023). Extends embodied multimodal planning by incorporating vision-language foundation models and embodied chain-of-thought reasoning into physical control.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). Advances multi-step embodied action generation by introducing visual chain-of-thought subgoals to guide low-level robotic execution.
- Paper: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces, Jihan Yang et al. (2025). Benchmarks and investigates the 3D spatial reasoning, mapping, and recall capabilities of multimodal large language models in complex indoor scenes.
- Paper: π0.5: a Vision-Language-Action Model with Open-World Generalization, Physical Intelligence et al. (2025). Scales vision-language-action architectures to execute long-horizon domestic tasks across diverse, unseen physical environments.
