Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions
Jing GuEliana StefaniQi WuJesse ThomasonXin Wang
Presents a comprehensive taxonomy of vision-and-language navigation benchmarks organized by communication complexity and task objectives, while systematically evaluating core modeling strategies, datasets, and open research challenges.
Building autonomous systems that understand human language, perceive physical environments, and carry out complex real-world tasks is a critical objective for artificial intelligence and robotics. Vision-and-language navigation addresses this need by investigating how embodied agents interpret natural instructions, navigate physical or simulated spaces, and interact with surroundings. The article systematically reviews the field by categorizing existing tasks, evaluation standards, and methodologies, while highlighting key challenges and strategic directions for future development.
The article establishes a taxonomy that classifies benchmarks across communication complexity and task objectives, ranging from single initial commands to continuous human dialogue, and from simple pathfinding to complex object manipulation. It also organizes technical solutions into four pillars: cross-modal representation learning, action strategy optimization, data-centric techniques, and prior environmental exploration. Across these domains, the analysis reveals that high-level semantic alignment, memory structures, and task-specific pretraining significantly outperform raw sensory models, improving success rates in unseen environments from around 20% in baseline models to upwards of 60–67% in modern systems. Furthermore, data augmentation and curriculum learning mitigate severe data scarcity, while prior exploration bridges the performance gap between known and unknown environments.
These findings indicate that while visual and linguistic reasoning algorithms are maturing rapidly, significant operational risks remain. Most current models rely on unrealistic assumptions—such as idealized teleportation between discrete nodes, perfect localization, static environments, and limited indoor scenes dominated by Western residential layouts. Transferring these simulated systems to continuous, real-world robotics will lead to substantial performance drops and safety hazards unless underlying dynamics, lighting variations, and sensor noise are resolved.
Organizations developing embodied AI should prioritize transitioning from discrete navigation graphs to continuous physical environments and invest in physical object manipulation beyond simple waypoint traversal. Additionally, development roadmaps must incorporate multi-agent and human-robot collaboration, rigorous data privacy safeguards, and geographically diverse testing environments. Because existing research is largely constrained to synthetic or residential simulations, stakeholders should maintain caution regarding immediate physical deployment and fund pilot validations in dynamic, real-world settings such as warehouses and healthcare facilities.
- Paper: Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments, Peter Anderson et al. (2017). This foundational VLN paper introduces the realistic instruction-following task and benchmark that the survey uses to frame the field’s subsequent methods and evaluation.
- Paper: GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation, Mukul Khanna et al. (2024). GOAT-Bench carries VLN beyond single-target episodes by testing lifelong navigation across open-vocabulary, multimodal instructions and persistent environmental memory.
- Paper: PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation, Jialu Li et al. (2023). PanoGen extends VLN training with generated panoramic environments, directly addressing the survey’s concerns about limited data and generalization to unseen places.
- Paper: AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation, Xin Ding et al. (2025). AdaNav continues VLN method development by adapting when an agent reasons, aiming to improve navigation while reducing the cost of unnecessary reasoning.
