EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Rui YangHanyang ChenJunyu ZhangMark ZhaoCheng QianKangrui WangQineng WangTeja Venkat KoripellaMarziyeh MovahediManling Li

article2025ICML245 citations

Presents EmbodiedBench, a multi-environment benchmark spanning high-level semantic planning to low-level physical manipulation across 1,128 tasks, revealing that leading multimodal models still fail at fine-grained control despite strong high-level reasoning.

Listen

Developing autonomous, vision-driven embodied agents—such as robotic systems that navigate physical environments and manipulate objects based on human instructions—has become a central focus in artificial intelligence. While large foundation models show strong general reasoning capabilities, systematic evaluation of Multi-modal Large Language Models acting directly as vision-driven physical agents remains limited. The article addresses this gap by presenting EMBODIEDBENCH, a comprehensive evaluation suite designed to assess off-the-shelf multimodal models across hierarchical action tiers and diverse agent competencies.

The main objective of the article is to establish a standardized, multifaceted benchmark that measures how effectively 24 leading proprietary and open-source multimodal models perform both high-level semantic planning and low-level physical control in simulated embodied environments. The benchmark spans 1,128 test instances across four distinct virtual domains: two high-level household planning environments and two low-level environments requiring continuous or discretized navigation and robotic arm manipulation. The evaluation systematically tests models across six core agent capabilities: base execution, commonsense reasoning, complex instruction comprehension, spatial awareness, visual appearance identification, and long-horizon planning.

The investigation yields several critical findings. First, while current multimodal models perform relatively well on high-level semantic planning tasks (such as household organization), they struggle significantly with low-level physical control; even the top-performing model, GPT-4o, achieved an average task success rate of only 28.9% on manipulation tasks. Second, visual input is indispensable for low-level control—removing vision leads to a 40% to 70% drop in navigation success—whereas high-level planning tasks rely almost entirely on text-based semantic cues. Third, long-horizon planning emerges as the single greatest bottleneck across all models, with performance degrading steeply as action sequences lengthen. In addition, proprietary models generally outperform open-source alternatives, though open-source models demonstrate consistent improvements as model scale increases.

These findings have direct operational and strategic implications for organizations investing in robotics and embodied automation. High-level reasoning models cannot be deployed directly into real-world robotic control loops without specialized grounding, as they frequently commit planning errors, miss execution steps, and fail to estimate physical 3D spatial coordinates accurately. Organizations facing safety-critical and cost-sensitive deployment timelines should decouple high-level task planners from low-level execution controllers rather than relying on end-to-end foundation models. Furthermore, ablation experiments show that moderate visual resolutions and visual in-context demonstrations significantly enhance performance, whereas multi-step or multi-view image feeds tend to confuse current models.

To advance toward viable embodied systems, the article recommends prioritizing research in 3D spatial reasoning, visual in-context learning, and hierarchical planning architectures that can reliably manage long-horizon execution. Decision-makers should validate models within tailored simulation pipelines before deploying them into physical hardware. A primary limitation of this study is that all experiments were conducted within virtual simulators rather than physical real-world environments. While this ensures reproducible, safe, and cost-effective benchmarking, stakeholders should exercise caution and conduct physical pilot testing before translating these simulation findings into real-world operational deployments.

arXiv: 2502.09560

No sufficiently relevant recommendations were found.

Cover for EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Abstract

Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents remain underexplored due to the lack of comprehensive evaluation frameworks. To bridge this gap, we introduce EmbodiedBench, an extensive benchmark designed to evaluate vision-driven embodied agents. EmbodiedBench features: (1) a diverse set of 1,128 testing tasks across four environments, ranging from high-level semantic tasks (e.g., household) to low-level tasks involving atomic actions (e.g., navigation and manipulation); and (2) six meticulously curated subsets evaluating essential agent capabilities like commonsense reasoning, complex instruction understanding, spatial awareness, visual perception, and long-term planning. Through extensive experiments, we evaluated 24 leading proprietary and open-source MLLMs within EmbodiedBench. Our findings reveal that: MLLMs excel at high-level tasks but struggle with low-level manipulation, with the best model, GPT-4o, scoring only 28.9% on average. EmbodiedBench provides a multifaceted standardized evaluation platform that not only highlights existing challenges but also offers valuable insights to advance MLLM-based embodied agents. Our code and dataset are available at https://embodiedbench.github.io.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Problem Formulation
  • 4. EmbodiedBench
  • 4.1. High-level and Low-level Tasks
  • 4.2. Capability-oriented Data Collection
  • 4.3. Vision-driven Agent Design
  • 5. Experiments
  • 5.1. Experimental Setups
  • 5.2. Benchmark Results
  • 5.3. Language-centric Ablation
  • 5.4. Visual-centric Ablation
  • 5.5. Error Analysis
  • 6. Conclusion
  • Limitations
  • Impact Statement
  • Acknowledgement
  • References
  • A. Additional Related Works
  • B. Future Research Directions
  • C. Details about EMBODIEDBENCH Environments and Datasets
  • C.1. EB-ALFRED
  • C.2. EB-Habitat
  • C.3. EB-Navigation
  • C.4. EB-Manipulation
  • D. Model Versions
  • E. Definitions and Examples of Capability-oriented Subsets
  • F. Additional Experiment Results
  • F.1. Subgoal Success Rate
  • F.2. Average Planner and Environment Steps
  • F.3. Camera Resolution
  • F.4. Detection Boxes
  • F.5. Multi-step Images
  • F.6. Multi-view Images
  • F.7. Visual In-context Learning (ICL)
  • F.8. Additional Ablation Study Conclusion
  • G. Further Discussion on Chat History as Input for EB-Navigation
  • H. Error Definitions and Additional Analysis
  • H.1. Error Type Definition
  • H.2. Error Analysis for EB-Navigation
  • H.3. Format Errors
  • I. Input of the Vision-driven Embodied Agent
  • I.1. Prompts
  • I.2. Skill Sets
  • I.3. In-context examples
  • I.4. Output JSON Schema
  • J. Supplementary Case Studies of Successful Planning
  • K. Supplementary Case Studies of Unsuccessful Planning

Knowls

  1. Knowl 1 — EMBODIEDBENCH spans four environments and two action levels

    model/method

    EMBODIEDBENCH evaluates vision-driven embodied agents on 1,128 test tasks in four simulated environments, covering high-level household planning and low-level navigation and manipulation. EB-ALFRED uses AI2-THOR for household tasks expressed through eight high-level skills; its revised simulator supports multiple instances of an object type and has a scene-dependent action space of 171–298 actions. EB-Habitat uses Habitat 2.0 for rearrangement tasks with 70 parameterized high-level skills; navigation is restricted to receptacles, so locating objects can require visiting multiple locations. EB-Navigation uses AI2-THOR to test low-level movement toward a target using visual observations and action-validity feedback. EB-Manipulation uses CoppeliaSim with a Franka Panda arm and tests low-level object interaction using a seven-component gripper action; its continuous position and orientation controls are discretized, and object detection boxes and indexed 3D positions are provided. The benchmark’s action-level distinction is that low-level actions are atomic executable commands, whereas high-level actions represent behavior sequences or skills.

  2. Knowl 2 — Six subsets probe distinct embodied-agent capabilities

    definition

    EMBODIEDBENCH partitions tasks by capability: Base tests basic task solving; Common Sense uses indirect object references grounded in everyday knowledge; Complex Instruction adds relevant or irrelevant context that can obscure the request; Spatial Awareness identifies objects by relations to other objects or locations; Visual Appearance describes targets by visual attributes such as color or shape; and Long Horizon requires extended action sequences. EB-ALFRED and EB-Habitat each contain 300 test instances, with 50 in each of the six subsets. EB-Navigation contains 300 instances across five subsets, 60 each, and omits Spatial Awareness. EB-Manipulation contains 228 instances: 48 per subset except Visual Appearance, which has 36; it omits Long Horizon.

  3. Knowl 3 — The agent uses iterative multimodal planning with executable action sequences

    model/method

    The unified EMBODIEDBENCH agent conditions its planner on the language instruction, visual observations, few-shot demonstrations, interaction history, environment feedback, and task-specific action information. At each planning call, the MLLM describes the current visual state, reflects on prior actions and feedback, reasons about how to reach the goal, writes a language plan, and returns an executable plan in structured JSON. Rather than requiring one action per call, the planner may emit a variable-length sequence of actions, which reduces redundant replanning and model calls when an individual low-level action produces little visual change. If a plan fails or contains an invalid action, planning resumes from the latest environment state. The standard input is mainly the current image; the authors report that tested MLLMs often struggle to interpret multiple historical images.

  4. Knowl 4 — Evaluation covers 24 MLLMs under standardized interaction settings

    experimental setup

    The experiments compare 24 MLLMs: eight proprietary models and 16 open-source models, including open-source model sizes from 7B to 90B parameters. All models use temperature 0 and a maximum completion length of 2,048 tokens; images are standardized to 500 × 500 pixels. The maximum environment horizon is 30 steps for high-level tasks, 20 for EB-Navigation, and 15 for EB-Manipulation. Task success rate is the primary evaluation metric.

  5. Knowl 5 — High-level task success exceeds low-level manipulation success

    empirical result

    Across the four environments, MLLMs generally perform better on high-level household tasks than on low-level manipulation. Claude-3.5-Sonnet has the highest average success across EB-ALFRED and EB-Habitat, with 64.0% and 68.0%, respectively; Claude-3.7-Sonnet scores 67.7% on EB-ALFRED. GPT-4o leads the proprietary models on EB-Navigation at 57.7% and EB-Manipulation at 28.9%. Among open-source models, InternVL3-78B scores 39.0% on EB-ALFRED, 55.0% on EB-Habitat, and 53.7% on EB-Navigation; Ovis2-34B reaches 26.8% on EB-Manipulation. The results also show an overall scaling trend among open-source models, while a gap remains between the strongest open and proprietary systems, particularly on high-level tasks.

  6. Knowl 6 — Visual input matters much more for low-level tasks

    empirical result

    A comparison of vision-enabled GPT-4o with its language-only variant shows a pronounced modality difference by action level. On EB-Navigation, success falls from 57.7% to 17.4% without vision, and on EB-Manipulation it falls from 28.9% to 16.2%. For EB-Navigation’s Long Horizon subset, the corresponding scores are 55.0% with vision and 0.0% without it. By contrast, the language-only GPT-4o scores 58.0% versus 56.3% on EB-ALFRED and 56.0% versus 59.0% on EB-Habitat, so removing images does not consistently reduce high-level task performance. These comparisons indicate that visual observations are particularly important for low-level control in the tested setup.

  7. Knowl 7 — Long-horizon planning is a major performance bottleneck

    empirical result

    The authors identify Long Horizon as the most challenging capability subset overall. In EB-Habitat, Claude-3.5-Sonnet’s task success decreases from 96% on Base to 58% on Long Horizon, while GPT-4o decreases from 86% to 64%. Long-horizon tasks include extended action sequences (typically more than 15 steps in EB-ALFRED); in EB-Navigation, the target is not visible in the initial view, and in EB-Habitat the subset includes multiple rearrangements or objects. The results show that success on basic tasks does not reliably carry over to tasks requiring longer planning and execution.

  8. Knowl 8 — Input ablations show that useful context depends on the task

    empirical result

    Ablations on EB-ALFRED and EB-Manipulation show that different input additions have different effects. On EB-ALFRED’s Base subset, removing environment feedback lowers success by 10 percentage points for GPT-4o and 8 points for Claude-3.5-Sonnet; using no in-context examples reduces success to around 40%. On EB-Manipulation’s Base subset, 500 × 500 images outperform both 300 × 300 and 700 × 700 images: GPT-4o scores 39.6%, 22.9%, and 35.4%, respectively, while Claude-3.5-Sonnet scores 37.5%, 27.1%, and 29.2%. Removing detection boxes lowers scores from 39.6% to 27.1% for GPT-4o and from 37.5% to 29.2% for Claude-3.5-Sonnet. Adding the two preceding images also reduces manipulation success: GPT-4o scores 31.3% with multi-step images versus 39.6% without them, and Claude-3.5-Sonnet scores 29.2% versus 37.5%. In contrast, two visual in-context examples outperform language-only examples on manipulation: 35.4% versus 27.1% for GPT-4o and 41.7% versus 25.0% for Claude-3.5-Sonnet. The visual-example comparison uses only two demonstrations, whereas the main experiments use more text-based examples.

  9. Knowl 9 — Planning failures dominate the analyzed episodes

    empirical result

    The authors analyze 110 failed GPT-4o episodes from EB-ALFRED and EB-Manipulation, sampling 10 failures from each evaluated subset. In EB-ALFRED, errors are classified as planning (55%), reasoning (41%), or perception (4%); missing steps (23%) and invalid actions (22%) are the most common reported planning failures, while reflection errors account for 17% and premature termination decisions for 13%. In EB-Manipulation, planning errors account for 44% of failures, perception errors 33%, and reasoning errors 23%. Inaccurate actions are prominent in manipulation (42% in the error breakdown), reflecting difficulty estimating gripper poses; wrong object recognition accounts for 22% of failures. Thus, planning is the largest error category in both environments, with perception failures more prominent in manipulation than in the high-level household setting.

  10. Knowl 10 — The evaluation is limited to simulated environments

    limitation

    EMBODIEDBENCH evaluates agents only in simulation and reports no real-world experiments. The authors note that simulation provides a reproducible, lower-cost, and safer evaluation setting, but it does not establish performance in practical deployment. They identify more realistic simulations and standardized, cost-effective real-world test suites as possible ways to address this gap.

Coverage note — Detailed action inventories, prompt templates, supplementary per-model tables, and illustrative planning episodes are omitted because they add implementation detail or examples rather than distinct core findings.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575, 2025.
  3. 3.Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691, 2022.
  4. 4.Ajay, A., Han, S., Du, Y., Li, S., Gupta, A., Jaakkola, T., Tenenbaum, J., Kaelbling, L., Srivastava, A., and Agrawal, P. Compositional foundation models for hierarchical planning. Advances in Neural Information Processing Systems, 36:22304–22325, 2023.
  5. 5.Anthropic. Claude 3.5 sonnet, 2024. URL https://www.anthropic.com/news/claude-3-5-sonnet.
  6. 6.Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023.
  7. 7.Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., and Lin, J. Qwen2.5-vl technical report, 2025.
  8. 8.Belkhale, S., Ding, T., Xiao, T., Sermanet, P., Vuong, Q., Tompson, J., Chebotar, Y., Dwibedi, D., and Sadigh, D. Rt-h: Action hierarchies using language. arXiv preprint arXiv:2403.01823, 2024.
  9. 9.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosse-lut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  10. 10.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022.
  11. 11.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023.
  12. 12.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546.
  13. 13.Calli, B., Singh, A., Walsman, A., Srinivasa, S., Abbeel, P., and Dollar, A. M. The ycb object and model set: Towards common benchmarks for manipulation research. In 2015 international conference on advanced robotics (ICAR), pp. 510–517. IEEE, 2015.
  14. 14.Chang, M., Chhablani, G., Clegg, A., Cote, M. D., Desai, R., Hlavac, M., Karashchuk, V., Krantz, J., Mottaghi, R., Parashar, P., et al. Partnr: A benchmark for planning and reasoning in embodied multi-agent tasks. arXiv preprint arXiv:2411.00081, 2024.
  15. 15.Chattopadhyay, P., Hoffman, J., Mottaghi, R., and Kembhavi, A. Robustnav: Towards benchmarking robustness in embodied navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15691–15700, 2021.
  16. 16.Chen, B., Xu, Z., Kirmani, S., Ichter, B., Sadigh, D., Guibas, L., and Xia, F. Spatialvlm: Endowing vision-language models with spatial reasoning capabilities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14455–14465, 2024a.
  17. 17.Chen, Y., Cui, W., Chen, Y., Tan, M., Zhang, X., Zhao, D., and Wang, H. Robogpt: an intelligent agent of making embodied long-term decisions for daily instruction tasks. arXiv preprint arXiv:2311.15649, 2023a.
  18. 18.Chen, Y., Wang, X., Li, M., Hoiem, D., and Ji, H. Vistruct: Visual structural knowledge extraction via curriculum guided code-vision representation. In Proc. The 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP2023), 2023b.
  19. 19.Chen, Y., Wang, X., Peng, H., and Ji, H. Solo: A single transformer for scalable vision-language modeling. In Transactions on Machine Learning Research, 2024b.
  20. 20.Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., Li, B., Luo, P., Lu, T., Qiao, Y., and Dai, J. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv preprint arXiv:2312.14238, 2023c.
  21. 21.Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Cui, E., Zhu, J., Ye, S., Tian, H., Liu, Z., Gu, L., Wang, X., Li, Q., Ren, Y., Chen, Z., Luo, J., Wang, J., Jiang, T., Wang, B., He, C., Shi, B., Zhang, X., Lv, H., Wang, Y., Shao, W., Chu, P., Tu, Z., He, T., Wu, Z., Deng, H., Ge, J., Chen, K., Zhang, K., Wang, L., Dou, M., Lu, L., Zhu, X., Lu, T., Lin, D., Qiao, Y., Dai, J., and Wang, W. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025. URL https://arxiv.org/abs/2412.05271.
  22. 22.Cheng, A.-C., Yin, H., Fu, Y., Guo, Q., Yang, R., Kautz, J., Wang, X., and Liu, S. Spatialrgpt: Grounded spatial reasoning in vision language model. arXiv preprint arXiv:2406.01584, 2024.
  23. 23.Cheng, Z., Tu, Y., Li, R., Dai, S., Hu, J., Hu, S., Li, J., Shi, Y., Yu, T., Chen, W., et al. Embodiedeval: Evaluate multimodal llms as embodied agents. arXiv preprint arXiv:2501.11858, 2025.
  24. 24.Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S. Diffusion policy: Visuomotor policy learning via action diffusion. The International Journal of Robotics Research, pp. 02783649241273668, 2023.
  25. 25.Choi, J.-W., Yoon, Y., Ong, H., Kim, J., and Jang, M. Lotabench: Benchmarking language-oriented task planners for embodied agents. arXiv preprint arXiv:2402.08178, 2024.
  26. 26.Contributors, L. Lmdeploy: A toolkit for compressing, deploying, and serving llm. https://github.com/InternLM/lmdeploy, 2023.
  27. 27.DeepMind, G. Introducing gemini 2.0: our new ai model for the agentic era, 2024. URL https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/.
  28. 28.Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al. Palm-e: an embodied multimodal language model. In Proceedings of the 40th International Conference on Machine Learning, pp. 8469–8488, 2023.
  29. 29.Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., Tang, C., Wang, C., Zhang, D., Yuan, E., Lu, E., Tang, F., Sung, F., Wei, G., Lai, G., Guo, H., Zhu, H., et al. Kimi k1.5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599, 2025.
  30. 30.Du, Y., Yang, M., Florence, P., Xia, F., Wahid, A., Ichter, B., Sermanet, P., Yu, T., Abbeel, P., Tenenbaum, J. B., et al. Video language planning. arXiv preprint arXiv:2310.10625, 2023.
  31. 31.Durante, Z., Huang, Q., Wake, N., Gong, R., Park, J. S., Sarkar, B., Taori, R., Noda, Y., Terzopoulos, D., Choi, Y., et al. Agent ai: Surveying the horizons of multimodal interaction. arXiv preprint arXiv:2401.03568, 2024.
  32. 32.Fu, Z., Zhao, T. Z., and Finn, C. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation. arXiv preprint arXiv:2401.02117, 2024.
  33. 33.Gao, C., Zhao, B., Zhang, W., Mao, J., Zhang, J., Zheng, Z., Man, F., Fang, J., Zhou, Z., Cui, J., et al. Embodiedcity: A benchmark platform for embodied agent in real-world city environment. arXiv preprint arXiv:2410.09604, 2024a.
  34. 34.Gao, J., Sarkar, B., Xia, F., Xiao, T., Wu, J., Ichter, B., Majumdar, A., and Sadigh, D. Physically grounded vision-language models for robotic manipulation. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pp. 12462–12469. IEEE, 2024b.
  35. 35.Gu, Q., Kuwajerwala, A., Morin, S., Jatavallabhula, K. M., Sen, B., Agarwal, A., Rivera, C., Paul, W., Ellis, K., Chellappa, R., et al. Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pp. 5021–5028. IEEE, 2024.
  36. 36.Gulino, C., Fu, J., Luo, W., Tucker, G., Bronstein, E., Lu, Y., Harb, J., Pan, X., Wang, Y., Chen, X., et al. Waymax: An accelerated, data-driven simulator for large-scale autonomous driving research. Advances in Neural Information Processing Systems, 36, 2024.
  37. 37.Huang, H., Lin, F., Hu, Y., Wang, S., and Gao, Y. Copa: General robotic manipulation through spatial constraints of parts with foundation models. arXiv preprint arXiv:2403.08248, 2024a.
  38. 38.Huang, J., Yong, S., Ma, X., Linghu, X., Li, P., Wang, Y., Li, Q., Zhu, S.-C., Jia, B., and Huang, S. An embodied generalist agent in 3d world. arXiv preprint arXiv:2311.12871, 2023a.
  39. 39.Huang, S., Jiang, Z., Dong, H., Qiao, Y., Gao, P., and Li, H. Instruct2act: Mapping multi-modality instructions to robotic actions with large language model. arXiv preprint arXiv:2305.11176, 2023b.
  40. 40.Huang, W., Abbeel, P., Pathak, D., and Mordatch, I. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, pp. 9118–9147. PMLR, 2022a.
  41. 41.Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., et al. Inner monologue: Embodied reasoning through planning with language models. arXiv preprint arXiv:2207.05608, 2022b.
  42. 42.Huang, W., Wang, C., Zhang, R., Li, Y., Wu, J., and Fei-Fei, L. Voxposer: Composable 3d value maps for robotic manipulation with language models. arXiv preprint arXiv:2307.05973, 2023c.
  43. 43.Huang, W., Xia, F., Shah, D., Driess, D., Zeng, A., Lu, Y., Florence, P., Mordatch, I., Levine, S., Hausman, K., et al. Grounded decoding: Guiding text generation with grounded models for robot control. arXiv preprint arXiv:2303.00855, 2023d.
  44. 44.Huang, W., Wang, C., Li, Y., Zhang, R., and Fei-Fei, L. Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation. arXiv preprint arXiv:2409.01652, 2024b.
  45. 45.James, S., Ma, Z., Arrojo, D. R., and Davison, A. J. Rlbench: The robot learning benchmark & learning environment. IEEE Robotics and Automation Letters, 5(2):3019–3026, 2020.
  46. 46.Jiang, H., Huang, B., Wu, R., Li, Z., Garg, S., Nayyeri, H., Wang, S., and Li, Y. Roboexp: Action-conditioned scene graph via interactive exploration for robotic manipulation. arXiv preprint arXiv:2402.15487, 2024.
  47. 47.Khanna, M., Ramrakhya, R., Chhablani, G., Yenamandra, S., Gervet, T., Chang, M., Kira, Z., Chaplot, D. S., Batra, D., and Mottaghi, R. Goat-bench: A benchmark for multi-modal lifelong navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16373–16383, 2024.
  48. 48.Kim, M. J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024.
  49. 49.Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., and Fried, D. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. arXiv preprint arXiv:2401.13649, 2024.
  50. 50.Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Deitke, M., Ehsani, K., Gordon, D., Zhu, Y., et al. Ai2-thor: An interactive 3d environment for visual ai. arXiv preprint arXiv:1712.05474, 2017.
  51. 51.Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023.
  52. 52.Li, C., Xia, F., Martín-Martín, R., Lingelbach, M., Srivastava, S., Shen, B., Vainio, K., Gokmen, C., Dharan, G., Jain, T., et al. igibson 2.0: Object-centric simulation for robot learning of everyday household tasks. arXiv preprint arXiv:2108.03272, 2021.
  53. 53.Li, C., Zhang, R., Wong, J., Gokmen, C., Srivastava, S., Martín-Martín, R., Wang, C., Levine, G., Lingelbach, M., Sun, J., et al. Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation. In Conference on Robot Learning, pp. 80–93. PMLR, 2023.
  54. 54.Li, K., Yu, B., Zheng, Q., Zhan, Y., Zhang, Y., Zhang, T., Yang, Y., Chen, Y., Sun, L., Cao, Q., Shen, L., Li, L., Tao, D., and He, X. Muep: A multimodal benchmark for embodied planning with foundation models. In Larson, K. (ed.), Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24, pp. 129–138. International Joint Conferences on Artificial Intelligence Organization, 8 2024a. doi: 10.24963/ijcai.2024/15. URL https://doi.org/10.24963/ijcai.2024/15. Main Track.
  55. 55.Li, M., Zhao, S., Wang, Q., Wang, K., Zhou, Y., Srivastava, S., Gokmen, C., Lee, T., Li, L. E., Zhang, R., et al. Embodied agent interface: Benchmarking llms for embodied decision making. arXiv preprint arXiv:2410.07166, 2024b.
  56. 56.Li, X., Hsu, K., Gu, J., Pertsch, K., Mees, O., Walke, H. R., Fu, C., Lunawat, I., Sieh, I., Kirmani, S., et al. Evaluating real-world robot manipulation policies in simulation. arXiv preprint arXiv:2405.05941, 2024c.
  57. 57.Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A. Code as policies: Language model programs for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 9493–9500. IEEE, 2023.
  58. 58.Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J. Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a. URL https://llava-vl.github.io/blog/2024-01-30-llava-next/.
  59. 59.Liu, J., Li, S., Wang, Z., Li, M., and Ji, H. A language first approach for procedural planning. In Proc. The 61st Annual Meeting of the Association for Computational Linguistics (ACL2023) Findings, 2023a.
  60. 60.Liu, S., Chen, J., Ruan, S., Su, H., and Yin, Z. Exploring the robustness of decision-level through adversarial attacks on llm-based embodied models. In Proceedings of the 32nd ACM International Conference on Multimedia, pp. 8120–8128, 2024b.
  61. 61.Liu, S., Wu, L., Li, B., Tan, H., Chen, H., Wang, Z., Xu, K., Su, H., and Zhu, J. Rdt-1b: a diffusion foundation model for bimanual manipulation. arXiv preprint arXiv:2410.07864, 2024c.
  62. 62.Liu, S., Ren, Z., Gupta, S., and Wang, S. Physgen: Rigid-body physics-grounded image-to-video generation. In European Conference on Computer Vision, pp. 360–378. Springer, 2025.
  63. 63.Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., et al. Agentbench: Evaluating llms as agents. arXiv preprint arXiv:2308.03688, 2023b.
  64. 64.Liu, X., Guo, D., Zhang, X., and Liu, H. Heterogeneous embodied multi-agent collaboration. IEEE Robotics and Automation Letters, 2024d.
  65. 65.Liu, X., Zhang, T., Gu, Y., Iong, I. L., Xu, Y., Song, X., Zhang, S., Lai, H., Liu, X., Zhao, H., et al. Visualagentbench: Towards large multimodal models as visual foundation agents. arXiv preprint arXiv:2408.06327, 2024e.
  66. 66.Lu, S., Li, Y., Chen, Q.-G., Xu, Z., Luo, W., Zhang, K., and Ye, H.-J. Ovis: Structural embedding alignment for multimodal large language model. arXiv preprint arXiv:2405.20797, 2024.
  67. 67.Luo, J., Xu, C., Liu, F., Tan, L., Lin, Z., Wu, J., Abbeel, P., and Levine, S. Fmb: a functional manipulation benchmark for generalizable robotic learning. The International Journal of Robotics Research, pp. 02783649241276017, 2023.
  68. 68.Ma, Y., Cui, C., Cao, X., Ye, W., Liu, P., Lu, J., Abdelraouf, A., Gupta, R., Han, K., Bera, A., et al. Lampilot: An open benchmark dataset for autonomous driving with language model programs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15141–15151, 2024a.
  69. 69.Ma, Y., Song, Z., Zhuang, Y., Hao, J., and King, I. A survey on vision-language-action models for embodied ai. arXiv preprint arXiv:2405.14093, 2024b.
  70. 70.Madan, N., Møgelmose, A., Modi, R., Rawat, Y. S., and Moeslund, T. B. Foundation models for video understanding: A survey. arXiv preprint arXiv:2405.03770, 2024.
  71. 71.Mao, J., Qian, Y., Zhao, H., and Wang, Y. Gpt-driver: Learning to drive with gpt. arXiv preprint arXiv:2310.01415, 2023.
  72. 72.Mazzaglia, P., Verbelen, T., Dhoedt, B., Courville, A., and Rajeswar, S. Genrl: Multimodal-foundation world models for generalization in embodied agents. arXiv preprint arXiv:2406.18043, 2024.
  73. 73.Meta. Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024. URL https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/.
  74. 74.Mu, Y., Zhang, Q., Hu, M., Wang, W., Ding, M., Jin, J., Wang, B., Dai, J., Qiao, Y., and Luo, P. Embodiedgpt: Vision-language pre-training via embodied chain of thought. Advances in Neural Information Processing Systems, 36, 2024.
  75. 75.Nasiriany, S., Maddukuri, A., Zhang, L., Parikh, A., Lo, A., Joshi, A., Mandlekar, A., and Zhu, Y. Robocasa: Large-scale simulation of everyday tasks for generalist robots. arXiv preprint arXiv:2406.02523, 2024a.
  76. 76.Nasiriany, S., Xia, F., Yu, W., Xiao, T., Liang, J., Dasgupta, I., Xie, A., Driess, D., Wahid, A., Xu, Z., et al. Pivot: Iterative visual prompting elicits actionable knowledge for vlms. arXiv preprint arXiv:2402.07872, 2024b.
  77. 77.OpenAI. Hello gpt-4o, 2024a. URL https://openai.com/index/hello-gpt-4o/.
  78. 78.OpenAI. Gpt-4o mini: advancing cost-efficient intelligence, 2024b. URL https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/.
  79. 79.Puig, X., Ra, K., Boben, M., Li, J., Wang, T., Fidler, S., and Torralba, A. Virtualhome: Simulating household activities via programs. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 8494–8502, 2018.
  80. 80.Qian, C., Han, P., Luo, Q., He, B., Chen, X., Zhang, Y., Du, H., Yao, J., Yang, X., Zhang, D., Li, Y., and Ji, H. Escapebench: Pushing language models to think outside the box. In arxiv, 2024.
  81. 81.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  82. 82.Rana, K., Haviland, J., Garg, S., Abou-Chakra, J., Reid, I. D., and Suenderhauf, N. Sayplan: Grounding large language models using 3d scene graphs for scalable task planning. CoRR, 2023.
  83. 83.Redmon, J. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2016.
  84. 84.Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  85. 85.Rohmer, E., Singh, S. P., and Freese, M. V-rep: A versatile and scalable robot simulation framework. In 2013 IEEE/RSJ international conference on intelligent robots and systems, pp. 1321–1326. IEEE, 2013.
  86. 86.Sarch, G., Somani, S., Kapoor, R., Tarr, M. J., and Fragkiadaki, K. Helper-x: A unified instructable embodied agent to tackle four interactive vision-language domains with memory-augmented language models. arXiv preprint arXiv:2404.19065, 2024a.
  87. 87.Sarch, G. H., Jang, L., Tarr, M. J., Cohen, W. W., Marino, K., and Fragkiadaki, K. Vlm agents generate their own memories: Distilling experience into embodied programs of thought. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024b.
  88. 88.Sharma, S., Huang, H., Shivakumar, K., Chen, L. Y., Hoque, R., Ichter, B., and Goldberg, K. Semantic mechanical search with large vision and language models. arXiv preprint arXiv:2302.12915, 2023.
  89. 89.Shen, B., Xia, F., Li, C., Martín-Martín, R., Fan, L., Wang, G., Perez-D’Arpino, C., Buch, S., Srivastava, S., Tchapmi, L., et al. igibson 1.0: A simulation environment for interactive tasks in large realistic scenes. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 7520–7527. IEEE, 2021.
  90. 90.Shridhar, M., Thomason, J., Gordon, D., Bisk, Y., Han, W., Mottaghi, R., Zettlemoyer, L., and Fox, D. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10740–10749, 2020a.
  91. 91.Shridhar, M., Yuan, X., Cote, M.-A., Bisk, Y., Trischler, A., and Hausknecht, M. Alfworld: Aligning text and embodied environments for interactive learning. In International Conference on Learning Representations, 2020b.
  92. 92.Shridhar, M., Manuelli, L., and Fox, D. Cliport: What and where pathways for robotic manipulation. In Conference on robot learning, pp. 894–906. PMLR, 2022.
  93. 93.Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A. Progp prompt: Generating situated robot task plans using large language models. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 11523–11530. IEEE, 2023.
  94. 94.Song, C. H., Wu, J., Washington, C., Sadler, B. M., Chao, W.-L., and Su, Y. Llm-planner: Few-shot grounded planning for embodied agents with large language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2998–3009, 2023.
  95. 95.Song, X., Chen, W., Liu, Y., Chen, W., Li, G., and Lin, L. Towards long-horizon vision-language navigation: Platform, benchmark and method. arXiv preprint arXiv:2412.09082, 2024.
  96. 96.Stone, A., Xiao, T., Lu, Y., Gopalakrishnan, K., Lee, K.-H., Vuong, Q., Wohlhart, P., Kirmani, S., Zitkovich, B., Xia, F., et al. Open-world object manipulation using pre-trained vision-language models. arXiv preprint arXiv:2303.00905, 2023.
  97. 97.Sun, H. Reinforcement learning in the era of llms: What is essential? what is needed? an rl perspective on rlhf, prompting, and beyond. arXiv preprint arXiv:2310.06147, 2023.
  98. 98.Szot, A., Clegg, A., Undersander, E., Wijmans, E., Zhao, Y., Turner, J., Maestre, N., Mukadam, M., Chaplot, D. S., Maksymets, O., et al. Habitat 2.0: Training home assistants to rearrange their habitat. Advances in neural information processing systems, 34:251–266, 2021.
  99. 99.Szot, A., Schwarzer, M., Agrawal, H., Mazoure, B., Metcalf, R., Talbott, W., Mackraz, N., Hjelm, R. D., and Toshev, A. T. Large language models as generalizable policies for embodied tasks. In The Twelfth International Conference on Learning Representations, 2023.
  100. 100.Szot, A., Mazoure, B., Attia, O., Timofeev, A., Agrawal, H., Hjelm, D., Gan, Z., Kira, Z., and Toshev, A. From multimodal llms to generalist embodied agents: Methods and lessons. arXiv preprint arXiv:2412.08442, 2024.
  101. 101.Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024a.
  102. 102.Team, G., Kamath, A., Ferret, J., Pathak, S., Vieillard, N., Merhej, R., Perrin, S., Matejovicova, T., Rame, A., Riviere, M., et al. Gemma 3 technical report. arXiv preprint arXiv:2503.19786, 2025.
  103. 103.Team, O. M., Ghosh, D., Walke, H., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Kreiman, T., Xu, C., et al. Octo: An open-source generalist robot policy. arXiv preprint arXiv:2405.12213, 2024b.
  104. 104.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  105. 105.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023a.
  106. 106.Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024.
  107. 107.Wang, Q., Li, M., Chan, H. P., Huang, L., Hockenmaier, J., Girish, C., and Ji, H. Multimedia generative script learning for task planning. In Proc. The 61st Annual Meeting of the Association for Computational Linguistics (ACL2023) Findings, 2023b.
  108. 108.Wang, Y., Xian, Z., Chen, F., Wang, T.-H., Wang, Y., Fragkiadaki, K., Erickson, Z., Held, D., and Gan, C. Robogen: Towards unleashing infinite data for automated robot learning via generative simulation. arXiv preprint arXiv:2311.01455, 2023c.
  109. 109.Wang, Z., Blume, A., Li, S., Liu, G., Cho, J., Tang, Z., Bansal, M., and Ji, H. Paxion: Patching video-language foundation models with action knowledge. In Proc. 2023 Conference on Neural Information Processing Systems (NeurIPS2023) [Spotlight Paper], 2023d.
  110. 110.Wang, Z., Cai, S., Chen, G., Liu, A., Ma, X., and Liang, Y. Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents. arXiv preprint arXiv:2302.01560, 2023e.
  111. 111.Wu, C. H., Shah, R. R., Koh, J. Y., Salakhutdinov, R., Fried, D., and Raghunathan, A. Dissecting adversarial robustness of multimodal lm agents. In NeurIPS 2024 Workshop on Open-World Agents, 2024a.
  112. 112.Wu, Z., Chen, X., Pan, Z., Liu, X., Liu, W., Dai, D., Gao, H., Ma, Y., Wu, C., Wang, B., et al. Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302, 2024b.
  113. 113.Xiang, F., Qin, Y., Mo, K., Xia, Y., Zhu, H., Liu, F., Liu, M., Jiang, H., Yuan, Y., Wang, H., et al. Sapien: A simulated part-based interactive environment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11097–11107, 2020.
  114. 114.Xiang, J., Liu, G., Gu, Y., Gao, Q., Ning, Y., Zha, Y., Feng, Z., Tao, T., Hao, S., Shi, Y., et al. Pandora: Towards general world model with natural language actions and video states. arXiv preprint arXiv:2406.09455, 2024.
  115. 115.Xiao, T., Chan, H., Sermanet, P., Wahid, A., Brohan, A., Hausman, K., Levine, S., and Tompson, J. Robotic skill acquisition via instruction augmentation with vision-language models. arXiv preprint arXiv:2211.11736, 2022.
  116. 116.Xie, J., Chen, Z., Zhang, R., Wan, X., and Li, G. Large multimodal agents: A survey. arXiv preprint arXiv:2402.15116, 2024.
  117. 117.Xu, J., Yang, R., Luo, F., Fang, M., Wang, B., and Han, L. Robust decision transformer: Tackling data corruption in offline rl via sequence modeling. arXiv preprint arXiv:2407.04285, 2024.
  118. 118.Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al. Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115, 2024a.
  119. 119.Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. arXiv preprint arXiv:2310.11441, 2023a.
  120. 120.Yang, R., Zhong, H., Xu, J., Zhang, A., Zhang, C., Han, L., and Zhang, T. Towards robust offline reinforcement learning under diverse data corruption. arXiv preprint arXiv:2310.12955, 2023b.
  121. 121.Yang, R., Ding, R., Lin, Y., Zhang, H., and Zhang, T. Regularizing hidden states enables learning generalizable reward model for llms. arXiv preprint arXiv:2406.10216, 2024b.
  122. 122.Yang, R., Pan, X., Luo, F., Qiu, S., Zhong, H., Yu, D., and Chen, J. Rewards-in-context: Multi-objective alignment of foundation models with dynamic preference adjustment. arXiv preprint arXiv:2402.10207, 2024c.
  123. 123.Yang, Y., Zhou, T., Li, K., Tao, D., Li, L., Shen, L., He, X., Jiang, J., and Shi, Y. Embodied multi-modal agent trained by an llm from a parallel textworld. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 26275–26285, 2024d.
  124. 124.Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, 2023.
  125. 125.Yin, Y., Wang, Z., Sharma, Y., Niu, D., Darrell, T., and Herzig, R. In-context learning enables robot action prediction in llms. arXiv preprint arXiv:2410.12782, 2024.
  126. 126.Zawalski, M., Chen, W., Pertsch, K., Mees, O., Finn, C., and Levine, S. Robotic control via embodied chain-of-thought reasoning. arXiv preprint arXiv:2407.08693, 2024.
  127. 127.Zhai, S., Bai, H., Lin, Z., Pan, J., Tong, P., Zhou, Y., Suhr, A., Xie, S., LeCun, Y., Ma, Y., et al. Fine-tuning large vision-language models as decision-making agents via reinforcement learning. Advances in Neural Information Processing Systems, 37:110935–110971, 2025.
  128. 128.Zhang, S., Xu, Z., Liu, P., Yu, X., Li, Y., Gao, Q., Fei, Z., Yin, Z., Wu, Z., Jiang, Y.-G., et al. Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks. arXiv preprint arXiv:2412.18194, 2024a.
  129. 129.Zhang, X., Li, J., Chu, W., Hai, J., Xu, R., Yang, Y., Guan, S., Xu, J., and Cui, P. On the out-of-distribution generalization of multimodal large language models. arXiv preprint arXiv:2402.06599, 2024b.
  130. 130.Zhao, T. Z., Kumar, V., Levine, S., and Finn, C. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023.
  131. 131.Zheng, K., Chen, X., Jenkins, O. C., and Wang, X. Vlmbench: A compositional benchmark for vision-and-language manipulation. Advances in Neural Information Processing Systems, 35:665–678, 2022.
  132. 132.Zhou, G., Hong, Y., and Wu, Q. Navgpt: Explicit reasoning in vision-and-language navigation with large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 7641–7649, 2024a.
  133. 133.Zhou, Y., Li, X., Wang, Q., and Shen, J. Visual in-context learning for large vision-language models. arXiv preprint arXiv:2402.11574, 2024b.
  134. 134.Zhu, J., Wang, W., Chen, Z., Liu, Z., Ye, S., Gu, L., Duan, Y., Tian, H., Su, W., Shao, J., et al. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479, 2025.
  135. 135.Zou, C., Guo, X., Yang, R., Zhang, J., Hu, B., and Zhang, H. Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models. arXiv preprint arXiv:2411.00836, 2024.

Citation

MLA
Yang, R., et al. “EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents”. arXiv, 2025, http://arxiv.org/abs/2502.09560v3.
APA
Yang, R., Chen, H., Zhang, J., Zhao, M., Qian, C., Wang, K., Wang, Q., Koripella, T. V., Movahedi, M., Li, M., Ji, H., Zhang, H., & Zhang, T. (2025). EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents. arXiv. http://arxiv.org/abs/2502.09560v3
Chicago
Yang, R., H. Chen, J. Zhang, et al. 2025. “EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents”. arXiv. http://arxiv.org/abs/2502.09560v3.
Harvard
Yang, R. et al. (2025) “EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.09560v3.
Vancouver
1. Yang R, Chen H, Zhang J, et al (2025) EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents. arXiv

BibTeX

@article{yang2025embodiedbench,
  title = {EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents},
  author = {Yang, Rui and Chen, Hanyang and Zhang, Junyu and Zhao, Mark and Qian, Cheng and Wang, Kangrui and Wang, Qineng and Koripella, Teja Venkat and Movahedi, Marziyeh and Li, Manling and Ji, Heng and Zhang, Huan and Zhang, Tong},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.09560v3},
  eprint = {2502.09560}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/