Core Challenges in Embodied Vision-Language Planning
Jonathan FrancisNariaki KitamuraFelix LabelleXiaopeng LuIngrid NavarroJean Oh
Presents a unifying taxonomy and comparative review of embodied vision-language planning methods, benchmarks, and simulation environments to clarify critical open challenges for deploying interactive agents in the physical world.
The field of embodied artificial intelligence is rapidly advancing toward building autonomous agents that can collaborate with humans in physical spaces. To operate effectively, these systems must combine visual perception, natural language understanding, and sequential decision-making to complete physical objectives such as household chores, search and rescue, and autonomous delivery. However, existing research has largely studied vision, language, and planning in isolation or through disparate, fragmented subtasks, leaving the overall landscape and real-world deployment challenges poorly understood.
The article establishes a unified taxonomy for Embodied Vision-Language Planning tasks and systematically surveys current algorithmic approaches, simulation platforms, datasets, and evaluation metrics. Through this comprehensive analysis, the article evaluates the state of the art and demonstrates the critical technical gaps that must be overcome to transition these artificial intelligence agents from virtual simulations to reliable real-world operations.
To conduct this evaluation, the authors categorize the field into five primary task families: vision-language navigation, embodied question answering, embodied object referral, vision and dialogue navigation, and embodied goal-directed manipulation. The article reviews widely used modeling paradigms—including supervised imitation learning, reinforcement learning, and multimodal transformer architectures—across standard simulation environments such as Matterport3D, AI Habitat, and AI2-THOR. It also analyzes evaluation metrics measuring success rates, path length, trajectory fidelity, and object interaction accuracy.
The review yields five key findings regarding current capabilities. First, while agents achieve high success rates in training environments, their performance drops substantially in unseen test environments, exposing severe overfitting and poor generalization. Second, recent diagnostic studies show that masking visual or textual inputs results in negligible performance drops, indicating that models often rely on dataset biases and statistical shortcuts rather than genuine cross-modal understanding. Third, hybrid training that combines supervised pre-training with reinforcement learning or reward shaping consistently outperforms single-paradigm methods by enabling error recovery. Fourth, standard success metrics are often insufficient; for example, simple destination-based metrics fail to reflect whether an agent followed safe or instruction-faithful paths, making path-similarity metrics like dynamic time warping more effective for evaluation. Fifth, most current implementations rely on static, turn-based dialogue and stationary worlds, completely ignoring the dynamic changes and interactive communication required in physical deployments.
These findings imply that high benchmark scores on public leaderboards do not translate to operational readiness. In safety-critical or cost-sensitive deployments, relying on current models introduces significant operational risks because agents struggle to adapt to unmapped obstacles, unexpected physical interventions, or ambiguous instructions. Furthermore, the disconnect between simulated high-level actions (such as teleporting between viewpoints) and low-level physical control poses severe sim-to-real transfer bottlenecks that delay commercial adoption.
To address these limitations, the article recommends prioritizing the development of dynamic simulation environments where external events occur independently of the agent's actions. Stakeholders and researchers should adopt standardized, cross-task evaluation suites to benchmark general skills—such as spatial reasoning and object grounding—rather than isolated single-task metrics. Additionally, future efforts must integrate structured commonsense knowledge bases and dynamic, multi-turn dialogue systems to ensure agents can clarify ambiguous instructions and reason about real-world physical constraints before deploying them into human environments.
The conclusions of the article are constrained by the current scope of the literature, which predominantly focuses on single-agent operations in simulated, indoor settings while excluding complex multi-agent dynamics and low-level physical robot hardware control. Nevertheless, the findings provide a highly credible, evidence-based assessment that serves as a vital strategic roadmap for transitioning embodied vision-language planning into robust, real-world robotic systems.
- Paper: Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments, Peter Anderson et al. (2017). This foundational vision-and-language navigation benchmark establishes the instruction-following task and realistic evaluation setting that the survey later organizes within its broader taxonomy.
- Paper: AI2-THOR: An Interactive 3D Environment for Visual AI, Eric Kolve et al. (2017). Reading the original AI2-THOR paper first clarifies the interactive simulation platform and household tasks that underpin the survey’s discussion of embodied environments.
- Paper: Habitat: A Platform for Embodied AI Research, Manolis Savva et al. (2019). Habitat introduces a major embodied-AI simulation platform whose design and navigation experiments provide essential context for the survey’s benchmark comparisons.
- Paper: Target-driven visual navigation in indoor scenes using deep reinforcement learning, Yuke Zhu et al. (2016). This early target-driven navigation study grounds the survey’s account of visual navigation methods in a concrete deep-reinforcement-learning approach and benchmark.
- Paper: EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents, Rui Yang 0010 et al. (2025). Building on the survey’s call for cross-task evaluation, EmbodiedBench tests multimodal agents across planning and low-level control to expose where simulated competence still fails.
- Paper: GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation, Mukul Khanna et al. (2024). GOAT-Bench extends the survey’s navigation landscape from isolated episodes to lifelong, multimodal goal sequences that require agents to reuse environmental memory.
- Paper: SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World, Kiana Ehsani et al. (2024). SPOC directly advances the survey’s sim-to-real challenge by showing how large-scale simulated imitation can support navigation and manipulation in physical homes.
- Paper: An Embodied Generalist Agent in 3D World, Jiangyong Huang et al. (2024). LEO carries the survey’s vision-language planning agenda toward a unified 3D agent that grounds understanding, dialogue, planning, and action in shared representations.
