EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Rui YangHanyang ChenJunyu ZhangMark ZhaoCheng QianKangrui WangQineng WangTeja Venkat KoripellaMarziyeh MovahediManling Li
Presents EmbodiedBench, a multi-environment benchmark spanning high-level semantic planning to low-level physical manipulation across 1,128 tasks, revealing that leading multimodal models still fail at fine-grained control despite strong high-level reasoning.
Developing autonomous, vision-driven embodied agents—such as robotic systems that navigate physical environments and manipulate objects based on human instructions—has become a central focus in artificial intelligence. While large foundation models show strong general reasoning capabilities, systematic evaluation of Multi-modal Large Language Models acting directly as vision-driven physical agents remains limited. The article addresses this gap by presenting EMBODIEDBENCH, a comprehensive evaluation suite designed to assess off-the-shelf multimodal models across hierarchical action tiers and diverse agent competencies.
The main objective of the article is to establish a standardized, multifaceted benchmark that measures how effectively 24 leading proprietary and open-source multimodal models perform both high-level semantic planning and low-level physical control in simulated embodied environments. The benchmark spans 1,128 test instances across four distinct virtual domains: two high-level household planning environments and two low-level environments requiring continuous or discretized navigation and robotic arm manipulation. The evaluation systematically tests models across six core agent capabilities: base execution, commonsense reasoning, complex instruction comprehension, spatial awareness, visual appearance identification, and long-horizon planning.
The investigation yields several critical findings. First, while current multimodal models perform relatively well on high-level semantic planning tasks (such as household organization), they struggle significantly with low-level physical control; even the top-performing model, GPT-4o, achieved an average task success rate of only 28.9% on manipulation tasks. Second, visual input is indispensable for low-level control—removing vision leads to a 40% to 70% drop in navigation success—whereas high-level planning tasks rely almost entirely on text-based semantic cues. Third, long-horizon planning emerges as the single greatest bottleneck across all models, with performance degrading steeply as action sequences lengthen. In addition, proprietary models generally outperform open-source alternatives, though open-source models demonstrate consistent improvements as model scale increases.
These findings have direct operational and strategic implications for organizations investing in robotics and embodied automation. High-level reasoning models cannot be deployed directly into real-world robotic control loops without specialized grounding, as they frequently commit planning errors, miss execution steps, and fail to estimate physical 3D spatial coordinates accurately. Organizations facing safety-critical and cost-sensitive deployment timelines should decouple high-level task planners from low-level execution controllers rather than relying on end-to-end foundation models. Furthermore, ablation experiments show that moderate visual resolutions and visual in-context demonstrations significantly enhance performance, whereas multi-step or multi-view image feeds tend to confuse current models.
To advance toward viable embodied systems, the article recommends prioritizing research in 3D spatial reasoning, visual in-context learning, and hierarchical planning architectures that can reliably manage long-horizon execution. Decision-makers should validate models within tailored simulation pipelines before deploying them into physical hardware. A primary limitation of this study is that all experiments were conducted within virtual simulators rather than physical real-world environments. While this ensures reproducible, safe, and cost-effective benchmarking, stakeholders should exercise caution and conduct physical pilot testing before translating these simulation findings into real-world operational deployments.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). PaLM-E establishes how multimodal language models can connect visual and robot-state inputs to embodied planning, clarifying the agent capabilities EmbodiedBench evaluates.
- Paper: EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought, Yao Mu et al. (2023). EmbodiedGPT shows how vision-language reasoning can be translated into low-level robot actions, providing a direct model-building precursor to EmbodiedBench’s manipulation tests.
- Paper: An Embodied Generalist Agent in 3D World, Jiangyong Huang et al. (2024). LEO demonstrates a unified agent that perceives, plans, and acts in 3D environments, helping frame EmbodiedBench’s evaluation across embodied capabilities.
- Paper: Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning, Tianhe Yu et al. (2019). Meta-World defines standardized simulated manipulation tasks and success measures that help explain the benchmark conventions behind EmbodiedBench’s low-level action evaluation.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). This work establishes language-model task decomposition and executable household plans as an approach to embodied agents, clarifying the high-level planning abilities EmbodiedBench tests.
No sufficiently relevant recommendations were found.
