An Embodied Generalist Agent in 3D World
Jiangyong HuangSilong YongXiaojian MaXiongkun LinghuPuhao LiYan WangQing LiSong-Chun ZhuBaoxiong JiaSiyuan Huang
Presents LEO, a multimodal generalist agent trained on unified vision-language-action sequences to perform complex 3D perception, spatial reasoning, robotic manipulation, and embodied navigation.
Modern artificial intelligence models have achieved remarkable success in general-purpose reasoning across two-dimensional images and text. However, they struggle to perceive, reason about, and physically act within the three-dimensional physical environments essential to real-world tasks and embodied intelligence. Developing agents capable of operating in the physical world has been severely hindered by the scarcity of large-scale 3D vision-language datasets, fragmented architectures, and the absence of unified learning strategies bridging perception and physical action.
The article aims to introduce and evaluate LEO, an embodied multi-modal generalist agent designed to perceive, ground, reason, plan, and act in 3D environments using a single unified model architecture. Specifically, it assesses whether an autoregressive language modeling framework can directly align 3D point clouds, egocentric 2D images, text instructions, and discrete embodied action tokens without relying on separate, task-specific output heads.
To accomplish this, the authors implemented a two-stage training scheme comprising 3D vision-language alignment followed by vision-language-action instruction tuning. They curated two large-scale datasets, LEO-align and LEO-instruct, leveraging automated data generation pipelines that prompt large models with structured 3D scene graphs and object-centric reasoning chains, followed by rigorous filtering. The model processes object-centric 3D point clouds and 2D images via dedicated encoders and adapters, feeding interleaved multi-modal tokens into a seven-billion-parameter language model tuned efficiently using low-rank adaptation. The system was systematically evaluated across diverse benchmarks covering 3D captioning, situated question answering, multi-round dialogue, task planning, simulated robotic manipulation, and object navigation.
The evaluation revealed several critical findings. First, LEO established new state-of-the-art results on 3D dense captioning and question answering benchmarks without requiring task-specific fine-tuning; on the Scan2Cap captioning benchmark, it achieved a CIDEr score of 72.4 compared to 66.9 from the best previous fine-tuned generalist model, while on ScanQA it scored 101.4 versus the prior best of 69.6. Second, the agent demonstrated robust competence in embodied acting, matching specialist performance in robotic manipulation and successfully zero-shot transferring navigation capabilities to novel environments. Third, scaling experiments proved that instruction-tuning loss decreases log-linearly with increased data scale and model parameters, conforming to established scaling laws. Finally, the analysis showed that while vision-language pretraining aids downstream embodied control, adding action tokens slightly degraded language performance, identifying an asymmetric transfer tradeoff.
These findings demonstrate that integrating 3D object-centric spatial representations directly into large foundation models is a viable path toward generalist physical agents. In practice, this unified architecture can lower development and deployment costs by replacing multiple brittle, task-specific modules with a single general-purpose system across robotics, assistive technologies, and ambient computing. However, decision-makers should note that joint training of action policies with language models requires careful data balancing to prevent degradation of conversational and reasoning capabilities.
Moving forward, practitioners should explore scaling up multi-modal 3D training datasets and adopt balanced data schemes with negative sampling to mitigate agent hallucinations. For embodied action, future engineering efforts should integrate policy recurrence into the architecture, as LEO's current feed-forward action policy struggles to match human demonstrations on complex exploratory paths. Further research must also prioritize safety validations and alignment protocols before deploying these embodied agents in unconstrained physical environments.
Confidence in these findings is well-supported by thorough quantitative ablations and comparative benchmarks across established indoor datasets. Nevertheless, limitations remain regarding generalization to complex, out-of-distribution real-world scenes and the simplified feed-forward control policy. Stakeholders should maintain cautious optimism and conduct targeted pilot studies in real physical setups before broad commercial implementation.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). PaLM-E introduced the foundational concept of end-to-end embodied multimodal language models integrating continuous sensor and state inputs for robotic planning and decision-making.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). RT-2 established the vision-language-action (VLA) paradigm of fine-tuning web-scale multimodal models with robotic control tokens, which LEO directly generalizes to 3D domains.
- Paper: ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding, Le Xue et al. (2023). ULIP provides foundational pretraining techniques for learning aligned representations across 3D point clouds, 2D images, and text, underpinning 3D vision-language alignment.
- Paper: Language Conditioned Spatial Relation Reasoning for 3D Object Grounding, Shizhe Chen et al. (2022). ViL3DRel established key mechanisms for 3D language-conditioned spatial relation reasoning and object grounding essential for embodied agents in 3D environments.
- Paper: Visual Instruction Tuning, Haotian Liu et al. (2023). LLaVA pioneered visual instruction tuning to align visual encoders with large language backbones, directly inspiring LEO's multi-stage vision-language and vision-language-action instruction tuning.
- Paper: EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought, Yao Mu et al. (2023). EmbodiedGPT outlines how to connect multimodal foundation models to downstream robotic policies via embodied chain-of-thought planning.
- Paper: Context-aware Alignment and Mutual Masking for 3D-Language Pre-training, Zhao Jin et al. (2023). This work explores context-aware spatial-semantic alignment and mutual masking for 3D point cloud and language pretraining that precedes full generalist agent architectures.
- Paper: Multi-View Transformer for 3D Visual Grounding, Shijia Huang et al. (2022). Multi-View Transformer demonstrates view-robust 3D visual grounding by fusing multi-angle projections with text queries, a core capability required for 3D embodied agents.
- Paper: Habitat: A Platform for Embodied AI Research, Manolis Savva et al. (2019). Habitat establishes the foundational 3D simulation platform and benchmarks for training and evaluating embodied agents on navigation and manipulation.
- Paper: RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics, Chan Hee Song et al. (2025). RoboSpatial extends 3D embodied generalist concepts by systematically training and benchmarking foundational spatial reasoning across observer-, world-, and object-centric reference frames.
- Paper: UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, Jianke Zhang et al. (2025). UP-VLA builds on 3D generalist architectures by unifying high-level multimodal understanding with low-level future visual prediction in an autoregressive policy.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). CoA-VLA enhances VLA models like LEO by injecting explicit visual-textual chains of affordance into policy networks to resolve spatial ambiguity.
- Paper: π0.5: a Vision-Language-Action Model with Open-World Generalization, Physical Intelligence et al. (2025). pi0.5 extends vision-language-action generalist agents to open-world domestic settings using continuous action experts and flow matching.
- Paper: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces, Jihan Yang et al. (2025). Thinking in Space investigates and benchmarks the extent to which multimodal language models understand, remember, and recall 3D physical spaces from video.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). This survey synthesizes recent advances and architectural evolutions in Vision-Language-Action generalist models, placing systems like LEO into a broader taxonomic perspective.
- Paper: Visual Programming for Zero-Shot Open-Vocabulary 3D Visual Grounding, Zhihao Yuan et al. (2024). This work explores training-free visual programming to achieve zero-shot open-vocabulary 3D grounding, offering an alternative modular perspective to LEO's unified fine-tuning approach.
- Paper: GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation, Mukul Khanna et al. (2024). GOAT-Bench provides a lifelong multimodal navigation benchmark to evaluate continuous 3D embodied agent capabilities over extended horizons.
