Unified Human-Scene Interaction via Prompted Chain-of-Contacts
Zeqi XiaoTai WangJingbo WangJinkun CaoWenwei ZhangBo DaiDahua LinJiangmiao Pang
Develops UniHSI, a framework that uses large language models to translate text commands into sequential joint-object contact plans, enabling versatile and controllable human-scene interactions across diverse 3D environments.
Generating realistic human behaviors that interact naturally with indoor environments is essential for virtual reality and embodied artificial intelligence. However, existing approaches often struggle with practical deployment because they rely on separate, manually designed control policies for different tasks and require labor-intensive motion capture datasets paired with specific 3D objects. Furthermore, enabling characters to understand flexible natural language instructions across multi-step sequences remains a major hurdle.
The article introduces and evaluates UniHSI, a unified physics-based framework designed to execute diverse, multi-step human-scene interactions driven directly by natural language commands. The core objective is to demonstrate that complex interactions can be standardized into structured contact sequences, enabling a single reinforcement learning controller to execute versatile tasks without requiring paired interaction annotations.
To achieve this, the approach breaks down interactions into a "Chain of Contacts," which represents an action as ordered pairs of humanoid body joints and specific object parts. The framework uses a high-level large language model planner to translate user text instructions into structured contact chains, which are then passed to a unified low-level controller. This controller leverages reinforcement learning, an adversarial motion prior to ensure natural movement style, and an ego-centric heightmap to prevent environmental collisions in physics simulations. The authors evaluated the framework across thousands of generated plans using object models from PartNet and realistic indoor environments from ScanNet, comparing performance against established single-task and multi-task baselines.
The findings show that UniHSI achieves high execution accuracy across varied tasks, reaching an 85.5% success rate on simple single-object tasks and maintaining strong performance across multi-step sequences. In direct baseline comparisons, the unified model achieved an 81.5% success rate on complex "Lie Down" tasks—far surpassing standard single-task baselines (21.3%) and vanilla multi-task baseline combinations (20.1%)—while delivering a 94.3% success rate on "Sit" tasks. Ablation experiments demonstrated that dynamic reward balancing through adaptive contact weights is critical; removing this component caused success rates on simple tasks to plummet from 85.5% to 21.2%. Additionally, evaluations of language planners revealed that GPT-4 substantially outperformed GPT-3.5 in planning correctness (71.9% versus 49.1%) and execution success rate (57.3% versus 35.6%), although both trailed human planning (73.2% execution success).
These results indicate that decomposing physical interactions into contact-point representations allows diverse interactive skills to be learned within a single unified model without task-specific engineering. This significantly lowers data collection costs by removing the need for manual interaction annotations and reduces the risk of unnatural motion artifacts. Organizations developing interactive virtual environments and robotics simulations can adopt this structure to scale human-agent behaviors more efficiently.
Moving forward, developers should prioritize stronger spatial reasoning models or enhanced prompt verification systems to mitigate language planner failures, particularly when handling intricate multi-object sequences where task success rates drop to 40.5% in synthetic environments and 22.3% in complex real scans. Future research and pilot implementations should focus on integrating language models directly into the training loop and expanding capabilities to dynamic scenarios involving movable or carried objects.
Confidence in these findings is solid for static indoor environments, supported by extensive automated metrics and human user studies that rated the generated motions significantly higher in naturalness and semantic alignment than prior baselines. However, decision-makers should note key boundary conditions: the current system is restricted to interactions with static, fixed objects, and the overall system remains subject to occasional reasoning and spatial errors from upstream language models.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). Establishes foundational techniques for using pre-trained large language models to decompose high-level text instructions into actionable step-by-step plans, directly underlying UniHSI's LLM planning module.
- Paper: EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought, Yao Mu et al. (2023). Demonstrates embodied chain-of-thought planning to bridge high-level multimodal reasoning with low-level physical control, providing conceptual background for UniHSI's Chain-of-Contacts execution paradigm.
- Paper: SMPL, Matthew Loper et al. (2015). Introduces the SMPL 3D parametric human body model that provides the joint and mesh foundations used to represent human interaction dynamics in 3D scenes.
- Paper: Expressive Body Capture: 3D Hands, Face, and Body From a Single Image, Georgios Pavlakos et al. (2019). Extends parametric human body representation to expressive articulated body and hand joints (SMPL-X), essential for defining fine-grained joint-object contact regions.
- Paper: AMASS: Archive of Motion Capture As Surface Shapes, Naureen Mahmood et al. (2019). Supplies the standardized large-scale 3D human motion dataset (AMASS) necessary for learning and evaluating full-body kinematic interactions.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). Examines how vision-language web knowledge can be mapped directly into low-level control actions, motivating language-grounded embodied execution frameworks.
- Paper: AI2-THOR: An Interactive 3D Environment for Visual AI, Eric Kolve et al. (2017). Provides the interactive 3D simulation environment paradigm for benchmarked visual AI and embodied navigation-manipulation interactions in indoor scenes.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). Extends step-by-step physical interaction planning by integrating explicit visual-textual chains of affordances directly into vision-language-action policies.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). Applies sequential chain-of-thought reasoning to generate visual subgoals that guide step-by-step physical robot manipulation.
- Paper: Planning with Reasoning using Vision Language World Model, Delong Chen et al. (2025). Generalizes language-prompted embodied planning into predictive vision-language world models supporting hierarchical search and state prediction.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). Surveys the broader evolution and architectural paradigms of vision-language-action foundation models that connect high-level planning with low-level execution.
- Paper: π0.5: a Vision-Language-Action Model with Open-World Generalization, Physical Intelligence et al. (2025). Demonstrates open-world generalization in vision-language-action control by scaling hierarchical training across diverse real-world household environments.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). Builds on language-conditioned control to address low-latency continuous action execution in fast, dynamic object manipulation scenarios.
- Paper: RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation, Songming Liu et al. (2025). Applies unified action spaces and diffusion foundation models to complex, coordinated dual-arm physical manipulations.
