Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene Affordance
Zan WangYixin ChenBaoxiong JiaPuhao LiJinlu ZhangJingze ZhangTengyu LiuYixin ZhuWei LiangSiyuan Huang
Proposes a two-stage diffusion framework that uses explicit 3D scene affordance maps as an intermediate representation to generate physically plausible, language-guided human motions in complex environments even with limited paired training data.
Generating realistic 3D human animations from natural language descriptions is critical for animation synthesis, film production, and synthetic data generation. However, generating movement that accurately matches text while interacting plausibly with physical 3D environments presents significant technical hurdles. Existing generative approaches struggle because simultaneously aligning three distinct modalities—text, 3D scenes, and human motion—creates complex interdependencies, and the field suffers from a severe scarcity of high-quality paired training data combining all three elements.
The article develops and evaluates a two-stage generative framework that uses scene affordance maps as an intermediate bridge between text grounding in 3D space and conditional motion generation.
To address this challenge, the authors reformulated scene affordances—defined as distance fields measuring the spatial relationship between human skeletal joints and scene surfaces—to serve as an intermediate representation. The framework divides the task into two stages: first, an affordance diffusion model predicts interaction regions from text descriptions and 3D point clouds; second, an affordance-to-motion diffusion model synthesizes the temporal human movements guided by the predicted affordance map and text. The authors evaluated the approach across benchmark datasets and tested generalization on a curated evaluation suite of 16 novel indoor scenes with 80 unique text instructions.
The experimental findings show substantial improvements across key performance metrics. On the standard text-to-motion benchmark, the method achieved a distribution distance score (Fréchet Inception Distance, or FID) of 0.352 compared to 0.489–0.544 for leading diffusion baselines, indicating significantly higher motion quality and realism. On the human-scene interaction benchmark, the model reduced target grounding error to approximately 0.156 meters—a more than 50% improvement over prior variational autoencoder baselines (0.422 meters)—while increasing physical contact accuracy to approximately 96% compared to 84% for prior methods. Furthermore, tests on unseen scenes and open-ended descriptions confirmed that the two-stage model generalizes effectively to novel environments, whereas single-stage end-to-end models frequently produced severe collisions or missed contact targets.
These results demonstrate that decomposing complex multimodal generation into explicit spatial affordances and subsequent motion synthesis significantly mitigates the risk of physical implausibility in automated animation pipelines. By reducing the dependency on massive paired datasets, this approach lowers the cost and data-collection overhead required to build robust digital humans and interactive virtual agents.
Organizations developing virtual environments and animation tools should consider adopting affordance-based representations rather than direct end-to-end models for complex spatial tasks. Future engineering initiatives should prioritize reducing inference latency caused by the iterative diffusion denoising steps and developing strategies to expand high-coverage training data for highly intricate or multi-step human interactions.
Readers should note that while the results demonstrate strong statistical confidence across repeated benchmark trials, the model still exhibits limitations. Specifically, performance degrades when confronted with highly complex, multi-action textual instructions or entirely novel interaction types that require specific directional alignments (such as orienting the body correctly toward a sink). Latency constraints also remain an operational consideration for real-time applications.
- Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). Introduces the foundational Motion Diffusion Model (MDM) framework for synthesizing human motion from text descriptions that the source adapts for scene-grounded motion synthesis.
- Paper: Executing your Commands via Motion Diffusion in Latent Space, Xin Chen et al. (2023). Establishes latent-based diffusion modeling on standard text-to-motion benchmarks like HumanML3D, providing essential background on conditional motion diffusion architectures.
- Paper: MIME: Human-Aware 3D Scene Generation, Hongwei Yi et al. (2023). Examines the direct relationship between human body interactions and 3D indoor scene geometry, motivating affordance-based representations for scene-aware human movement.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Presents trajectory planning and behavior generation via diffusion processes, providing conceptual foundations for conditional spatial diffusion models.
- Paper: Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation, Jiaming Song et al. (2023). Provides techniques for enforcing test-time spatial and physical constraints during diffusion sampling, relevant to generating collision-free, affordance-guided motions.
- Paper: Unified Human-Scene Interaction via Prompted Chain-of-Contacts, Zeqi Xiao et al. (2024). Extends language-guided scene interaction by using large language models to decompose complex prompts into structured contact chains executed via physics simulation.
- Paper: CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control, Guy Tevet et al. (2025). Builds on kinematic motion diffusion by integrating real-time diffusion planners with closed-loop physics simulation controllers to eliminate penetration and contact artifacts in 3D scenes.
- Paper: AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and Beyond, Zixiang Zhou et al. (2024). Advances language-driven motion generation by unifying hierarchical planning, multi-step instruction understanding, and long-sequence synthesis into a single foundation model.
- Paper: Seamless Human Motion Composition with Blended Positional Encodings, Germán Barquero et al. (2024). Generalizes single-instruction motion diffusion frameworks to long, seamlessly composed sequences driven by continuous multi-prompt language instructions.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). Applies explicit visual-textual affordance reasoning as intermediate representations to guide sequential action generation in physical environments.
