Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene Affordance

Zan WangYixin ChenBaoxiong JiaPuhao LiJinlu ZhangJingze ZhangTengyu LiuYixin ZhuWei LiangSiyuan Huang

article2024CVPR111 citations

Proposes a two-stage diffusion framework that uses explicit 3D scene affordance maps as an intermediate representation to generate physically plausible, language-guided human motions in complex environments even with limited paired training data.

Listen

Generating realistic 3D human animations from natural language descriptions is critical for animation synthesis, film production, and synthetic data generation. However, generating movement that accurately matches text while interacting plausibly with physical 3D environments presents significant technical hurdles. Existing generative approaches struggle because simultaneously aligning three distinct modalities—text, 3D scenes, and human motion—creates complex interdependencies, and the field suffers from a severe scarcity of high-quality paired training data combining all three elements.

The article develops and evaluates a two-stage generative framework that uses scene affordance maps as an intermediate bridge between text grounding in 3D space and conditional motion generation.

To address this challenge, the authors reformulated scene affordances—defined as distance fields measuring the spatial relationship between human skeletal joints and scene surfaces—to serve as an intermediate representation. The framework divides the task into two stages: first, an affordance diffusion model predicts interaction regions from text descriptions and 3D point clouds; second, an affordance-to-motion diffusion model synthesizes the temporal human movements guided by the predicted affordance map and text. The authors evaluated the approach across benchmark datasets and tested generalization on a curated evaluation suite of 16 novel indoor scenes with 80 unique text instructions.

The experimental findings show substantial improvements across key performance metrics. On the standard text-to-motion benchmark, the method achieved a distribution distance score (Fréchet Inception Distance, or FID) of 0.352 compared to 0.489–0.544 for leading diffusion baselines, indicating significantly higher motion quality and realism. On the human-scene interaction benchmark, the model reduced target grounding error to approximately 0.156 meters—a more than 50% improvement over prior variational autoencoder baselines (0.422 meters)—while increasing physical contact accuracy to approximately 96% compared to 84% for prior methods. Furthermore, tests on unseen scenes and open-ended descriptions confirmed that the two-stage model generalizes effectively to novel environments, whereas single-stage end-to-end models frequently produced severe collisions or missed contact targets.

These results demonstrate that decomposing complex multimodal generation into explicit spatial affordances and subsequent motion synthesis significantly mitigates the risk of physical implausibility in automated animation pipelines. By reducing the dependency on massive paired datasets, this approach lowers the cost and data-collection overhead required to build robust digital humans and interactive virtual agents.

Organizations developing virtual environments and animation tools should consider adopting affordance-based representations rather than direct end-to-end models for complex spatial tasks. Future engineering initiatives should prioritize reducing inference latency caused by the iterative diffusion denoising steps and developing strategies to expand high-coverage training data for highly intricate or multi-step human interactions.

Readers should note that while the results demonstrate strong statistical confidence across repeated benchmark trials, the model still exhibits limitations. Specifically, performance degrades when confronted with highly complex, multi-action textual instructions or entirely novel interaction types that require specific directional alignments (such as orienting the body correctly toward a sink). Latency constraints also remain an operational consideration for real-time applications.

  • Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). Introduces the foundational Motion Diffusion Model (MDM) framework for synthesizing human motion from text descriptions that the source adapts for scene-grounded motion synthesis.
  • Paper: Executing your Commands via Motion Diffusion in Latent Space, Xin Chen et al. (2023). Establishes latent-based diffusion modeling on standard text-to-motion benchmarks like HumanML3D, providing essential background on conditional motion diffusion architectures.
  • Paper: MIME: Human-Aware 3D Scene Generation, Hongwei Yi et al. (2023). Examines the direct relationship between human body interactions and 3D indoor scene geometry, motivating affordance-based representations for scene-aware human movement.
  • Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Presents trajectory planning and behavior generation via diffusion processes, providing conceptual foundations for conditional spatial diffusion models.
  • Paper: Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation, Jiaming Song et al. (2023). Provides techniques for enforcing test-time spatial and physical constraints during diffusion sampling, relevant to generating collision-free, affordance-guided motions.
Cover for Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene Affordance

Abstract

Despite significant advancements in text-to-motion synthesis, generating language-guided human motion within 3D environments poses substantial challenges. These challenges stem primarily from (i) the absence of powerful generative models capable of jointly modeling natural language, 3D scenes, and human motion, and (ii) the generative models' intensive data requirements contrasted with the scarcity of comprehensive, high-quality, language-scene-motion datasets. To tackle these issues, we introduce a novel two-stage framework that employs scene affordance as an intermediate representation, effectively linking 3D scene grounding and conditional motion generation. Our framework comprises an Affordance Diffusion Model (ADM) for predicting explicit affordance map and an Affordance-to-Motion Diffusion Model (AMDM) for generating plausible human motions. By leveraging scene affordance maps, our method overcomes the difficulty in generating human motion under multimodal condition signals, especially when training with limited data lacking extensive language-scene-motion pairs. Our extensive experiments demonstrate that our approach consistently outperforms all baselines on established benchmarks, including HumanML3D and HUMANISE. Additionally, we validate our model's exceptional generalization capabilities on a specially curated evaluation set featuring previously unseen descriptions and scenes.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Language, Human Motion, and 3D Scene
  • 2.2. Conditional Human Motion Generation
  • 2.3. Scene Affordance
  • 3. Preliminaries
  • 4. Method
  • 4.1. Affordance Map
  • 4.2. Affordance Diffusion Model
  • 4.3. Affordance-to-Motion Diffusion Model
  • 4.4. Implementation Details
  • 5. Experiments
  • 5.1. Datasets
  • 5.2. Metrics and Baselines
  • 5.3. Results on HumanML3D
  • 5.4. Results on HUMANISE
  • 5.5. Results on Novel Evaluation Set
  • 5.6. Ablation Study
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Two-stage affordance-guided motion generation

    model/method

    The paper introduces a two-stage diffusion framework for generating human motion conditioned jointly on a 3D scene and a natural-language description. The Affordance Diffusion Model (ADM) first predicts a scene affordance map from the scene point cloud and language. The Affordance-to-Motion Diffusion Model (AMDM) then generates a human motion sequence from the predicted affordance map, the scene, and the language description. The affordance map acts as an intermediate representation: it localizes where the described interaction should occur while retaining geometric information about how the human body relates to nearby scene surfaces. The architecture diagram on page 4 depicts this sequence as input scene and text → affordance map → interactive motion.

  2. Knowl 2 — Distance-based scene affordance map

    definition

    Let the 3D scene be an RGB point cloud S=[s_1,dots,s_N]in\mathbb{R}^{N\times 6}, where each point sns_n contains 3D position and RGB information. Let a motion sequence contain FF frames, with pose xi∈RJ×3x_i\in\mathbb{R}^{J\times 3} at frame ii, where JJ is the number of human skeleton joints and xi,j∈R3x_{i,j}\in\mathbb{R}^3 is joint jj at that frame. For scene point nn and joint jj, the per-frame distance field is di(n,j)=∥snpos−xi,j∥2d_i(n,j)=\lVert s_n^{\mathrm{pos}}-x_{i,j}\rVert_2, where snpos∈R3s_n^{\mathrm{pos}}\in\mathbb{R}^3 is the positional component of sns_n. The distance is converted to a proximity value by

    ci(n,j)=exp⁡(−12di(n,j)σ2),c_i(n,j)=\exp\left(-\frac{1}{2}\frac{d_i(n,j)}{\sigma^2}\right),

    where sigma>0 is a fixed normalization factor. The sequence-level affordance map is obtained by temporal max pooling:

    C(n,j)=max⁡i∈{1,…,F}ci(n,j),C(n,j)=\max_{i\in\{1,\dots,F\}}c_i(n,j),

    so C∈RN×JC\in\mathbb{R}^{N\times J} records the strongest proximity of every scene point to every body joint over the motion. Larger values identify scene regions close to joints involved in the interaction. The implementation uses sigma=0.8.

  3. Knowl 3 — Affordance Diffusion Model

    model/method

    The Affordance Diffusion Model (ADM) learns the conditional distribution of affordance maps given an RGB point cloud SS and tokenized language description LL. Its reverse diffusion distribution is

    pθ(C0:T∣S,L)=p(CT)∏t=1Tpθ(Ct−1∣Ct,S,L),p_\theta(C_{0:T}\mid S,L)=p(C_T)\prod_{t=1}^{T}p_\theta(C_{t-1}\mid C_t,S,L),

    where C0C_0 is the clean affordance map, CtC_t is its noisy version at diffusion step tt, TT is the number of diffusion steps, and theta denotes learnable parameters. ADM uses a Perceiver backbone with Encode, Process, and Decode blocks. Point-wise scene features and the noisy affordance map are supplied as attention keys and values, while the concatenation of frozen text features and a diffusion-step embedding supplies the query. Multiple self-attention layers refine the latent representation in the Process block, and the Decode block uses the refined latent features to update per-point features before a linear projection produces the affordance estimate. Rather than predicting the noise added during diffusion, ADM directly predicts the clean input map through a function GθG_\theta. It is trained with

    LADM=EC0,t[∥C0−Gθ(Ct,t,S,L)∥22].\mathcal{L}_{\mathrm{ADM}}=\mathbb{E}_{C_0,t}\left[\left\lVert C_0-G_\theta(C_t,t,S,L)\right\rVert_2^2\right].

    The expectation is over training affordance maps and sampled diffusion steps.

  4. Knowl 4 — Affordance-to-Motion Diffusion Model

    model/method

    The Affordance-to-Motion Diffusion Model (AMDM) generates a motion sequence using the affordance map, scene, and language conditions. Its conditional reverse diffusion distribution is

    pϕ(X0:T∣C,S,L)=p(XT)∏t=1Tpϕ(Xt−1∣Xt,C,S,L),p_\phi(X_{0:T}\mid C,S,L)=p(X_T)\prod_{t=1}^{T}p_\phi(X_{t-1}\mid X_t,C,S,L),

    where X0X_0 is the clean sequence of joint positions, XtX_t is its noisy version at step tt, CC is the affordance map, SS is the scene point cloud, LL is the language description, and ϕ\phi denotes learnable parameters. An affordance encoder based on Point Transformer extracts multiscale affordance features with different cardinalities, and U-Net decoder layers process these features. A Transformer motion backbone receives the concatenation of noisy motion features, language features, and diffusion-step embeddings. Its cross-attention layers use this motion-language representation to attend to the affordance features, while self-attention models dependencies within the motion representation. A final linear layer maps the fused representation to joint-position space. As with ADM, AMDM directly estimates the clean sequence using GϕG_\phi and minimizes

    LAMDM=EX0,t[∥X0−Gϕ(Xt,t,C,S,L)∥22].\mathcal{L}_{\mathrm{AMDM}}=\mathbb{E}_{X_0,t}\left[\left\lVert X_0-G_\phi(X_t,t,C,S,L)\right\rVert_2^2\right].

    At inference time, ADM first produces CC, and AMDM uses that predicted map rather than a ground-truth affordance map.

  5. Knowl 5 — Datasets and training protocol

    experimental setup

    The experiments combine three sources of language-scene-motion data. HumanML3D supplies text-motion pairs but no scenes, so the authors add a floor and retain the dataset's original motion representation and train-test splits. HUMANISE aligns AMASS motions with ScanNet scenes; spatially referring descriptions are removed, scenes are segmented into chunks, and the original motions and splits are retained. The authors additionally create a generalization set containing 16 scenes from ScanNet, PROX, Replica, and Matterport3D and 80 interaction descriptions written by Turkers; this set contains unseen language-scene combinations and no ground-truth motions. A consolidated training set combines HumanML3D, HUMANISE, and PROX, converts annotations to joint-position representations, and randomly places furniture around HumanML3D motions. It contains 63,770 human-scene interactions, of which 48,470 have language annotations.

    Both diffusion models use frozen CLIP-ViT-B/32 text features, the AdamW optimizer, and a fixed learning rate of 10−410^{-4}. ADM is trained on two NVIDIA A100 GPUs with batch size 64 per GPU, while AMDM is trained on four NVIDIA A100 GPUs with batch size 32 per GPU; both are trained to convergence. HumanML3D is evaluated with R-Precision, FID, multimodal distance, diversity, and multimodality. HUMANISE and the novel set additionally use goal distance, contact, non-collision, perceptual quality, and action scores. Reported evaluations are repeated five times with 95% confidence intervals for quantitative metrics.

  6. Knowl 6 — HumanML3D generation results

    data/table

    On HumanML3D, the affordance-based method uses the Perceiver ADM and encoder-based AMDM. Lower FID and multimodal distance are preferred; diversity and multimodality should be close to or exceed the real-motion reference, while higher R-Precision is preferred. The results show that the proposed method improves over the comparable MDM diffusion baseline in R-Precision, FID, and multimodal distance while retaining comparable diversity. In particular, its FID is 0.352±0.1090.352\pm0.109, compared with 0.544±0.0440.544\pm0.044 for the uncorrected MDM row and 0.489±0.0250.489\pm0.025 for the adjusted MDM row.

    Could not parse LaTeX table

    The superscript-star rows use the authors' adjusted evaluation implementation. The results support the claim that scene-affordance conditioning enriches text-to-motion generation even when HumanML3D is augmented only with a floor rather than full scene geometry.

  7. Knowl 7 — HUMANISE scene-grounded motion results

    data/table

    On HUMANISE, the two-stage affordance model is compared with the original conditional VAE and one-stage diffusion models that bypass explicit affordance-map generation. Lower goal distance indicates better scene grounding; higher APD, contact, non-collision, quality, and action scores are preferred. Both AMDM variants substantially improve grounding, contact, perceptual quality, and action consistency over the baselines. The decoder-based affordance model obtains the best goal distance, APD, contact, and quality score among the learned systems, while the encoder-based variant obtains the highest action score.

    Could not parse LaTeX table

    Qualitative comparisons show that the cVAE often places the person at the wrong target object, whereas the one-stage diffusion model can produce scene collisions or fail to make contact. The affordance-conditioned models more consistently place the person at the semantically specified object and generate the required interaction.

  8. Knowl 8 — Generalization to unseen scenes and descriptions

    data/table

    The novel evaluation set contains unseen scenes and Turker-written interaction descriptions but no ground-truth motions. The row labeled Real is only a reference computed from language-motion pairs in the HumanML3D test set, not ground truth for the novel scenes. Compared with one-stage diffusion, the affordance-based encoder variant reduces FID from 11.848±1.63411.848\pm1.634 to 7.887±1.1897.887\pm1.189, improves multimodality from 4.966±0.3214.966\pm0.321 to 5.159±0.3565.159\pm0.356, and improves quality and action scores from 1.94±1.151.94\pm1.15 and 2.61±1.452.61\pm1.45 to 2.06±1.232.06\pm1.23 and 2.63±1.472.63\pm1.47. The affordance-based decoder variant achieves the highest learned contact score, 88.63±2.97588.63\pm2.975, but has weaker FID and multimodality than the encoder variant.

    Could not parse LaTeX table

    These results indicate that explicit affordance prediction improves robustness to novel language-scene combinations, particularly in distributional similarity, interaction quality, and action consistency, although the gain is not uniform across every metric or AMDM insertion variant.

  9. Knowl 9 — Perceiver affordance encoding is better than pointwise MLP encoding

    empirical result

    The authors compare three ADM architectures for predicting affordance maps: an MLP that processes scene points independently, a Point Transformer, and the proposed Perceiver. Grounding is measured by minimum joint-to-target distance, pelvis-to-target distance, and all-joint-to-target distance; lower values are better. The Perceiver is substantially closer to the ground-truth affordance geometry than the MLP and is slightly better than the Point Transformer on pelvis and all-joint distances.

    Could not parse LaTeX table

    When these affordance variants condition an encoder-based AMDM, the same ordering largely persists for grounding and contact. The Perceiver obtains goal distance 0.156±0.0060.156\pm0.006 and contact 95.86±0.32395.86\pm0.323, compared with 0.164±0.0100.164\pm0.010 and 94.39±0.40894.39\pm0.408 for the Point Transformer and 0.394±0.0100.394\pm0.010 and 73.96±0.43473.96\pm0.434 for the MLP. The ground-truth affordance condition gives 0.0170.017, 90.7990.79, and 99.8499.84 for goal distance, contact, and non-collision, respectively. The authors attribute the MLP's poorer grounding to its lack of information exchange between scene points and the Perceiver's advantage to cross-attention between point and language features.

  10. Knowl 10 — Failure modes and stated limitations

    limitation

    The method has two limitations identified by the authors. First, because both stages use iterative diffusion sampling, inference is relatively slow. Second, although the affordance representation reduces the difficulty caused by scarce paired language-scene-motion data, the amount and diversity of such data remain a critical bottleneck.

    The qualitative failure cases expose the remaining generalization weaknesses. For an unfamiliar interaction such as washing hands near a tap, the generated person can be placed in the correct region but fail to orient the body appropriately toward the sink. The system also fails on descriptions whose interaction structure is unusually complex, and on human-scene interactions that are entirely absent from training.

Coverage note — No substantial methodological or empirical contribution was omitted; image-only qualitative examples were folded into the corresponding generalization and failure-mode knowls rather than reproduced as separate knowls.

References

  1. 1.Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas Guibas. Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes. In European Conference on Computer Vision (ECCV), 2020. 2, 5
  2. 2.Hyemin Ahn, Timothy Ha, Yunho Choi, Hwiyeon Yoo, and Songhwai Oh. Text2action: Generative adversarial synthesis from language to action. In International Conference on Robotics and Automation (ICRA), 2018. 3
  3. 3.Chaitanya Ahuja and Louis-Philippe Morency. Language2pose: Natural language grounded pose forecasting. In International Conference on 3D Vision (3DV), 2019. 2, 3, 6
  4. 4.Joao Pedro Araújo, Jiaman Li, Karthik Vetrivel, Rishi Agarwal, Jiajun Wu, Deepak Gopinath, Alexander William Clegg, and Karen Liu. Circle: Capture in rich contextual environments. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3
  5. 5.Nikos Athanasiou, Mathis Petrovich, Michael J Black, and Gul Varol. Teach: Temporal action composition for 3d humans. In International Conference on 3D Vision (3DV), 2022. 2, 3
  6. 6.Nikos Athanasiou, Mathis Petrovich, Michael J. Black, and Gul Varol. SINC: Spatial composition of 3D human motions for simultaneous action generation. In International Conference on Computer Vision (ICCV), 2023. 2, 3
  7. 7.Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, and Motoaki Kawanabe. Scanqa: 3d question answering for spatial scene understanding. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  8. 8.Emad Barsoum, John Kender, and Zicheng Liu. Hp-gan: Probabilistic 3d human motion prediction via gan. In Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 3
  9. 9.Zhe Cao, Hang Gao, Karttikeya Mangalam, Qi-Zhi Cai, Minh Vo, and Jitendra Malik. Long-term human motion prediction with scene context. In European Conference on Computer Vision (ECCV), 2020. 2, 3
  10. 10.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. In International Conference on 3D Vision (3DV), 2017. 5
  11. 11.Yu-Wei Chao, Jimei Yang, Weifeng Chen, and Jia Deng. Learning to sit: Synthesizing human-chair interactions via hierarchical control. In AAAI Conference on Artificial Intelligence (AAAI), 2021. 3
  12. 12.Dave Zhenyu Chen, Qirui Wu, Matthias Nießner, and Angel X Chang. D3net: a speaker-listener architecture for semi-supervised dense captioning and visual grounding in rgb-d scans. In European Conference on Computer Vision (ECCV), 2022. 2
  13. 13.Sijin Chen, Hongyuan Zhu, Xin Chen, Yinjie Lei, Gang Yu, and Tao Chen. End-to-end 3d dense captioning with vote2cap-detr. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
  14. 14.Xin Chen, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, and Gang Yu. Executing your commands via motion diffusion in latent space. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3
  15. 15.Yixin Chen, Siyuan Huang, Tao Yuan, Siyuan Qi, Yixin Zhu, and Song-Chun Zhu. Holistic++ scene understanding: Single-view 3d holistic scene parsing and human pose estimation with human-object interaction and physical commonsense. In International Conference on Computer Vision (ICCV), 2019. 3
  16. 16.Yixin Chen, Qing Li, Deqian Kong, Yik Lun Kei, Tao Gao, Yixin Zhu, and Siyuan Huang. Yourefit: Embodied reference understanding with language and gesture. In International Conference on Computer Vision (ICCV), 2021. 2
  17. 17.Yixin Chen, Sai Kumar Dwivedi, Michael J Black, and Dimitrios Tzionas. Detecting human-object contact in images. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  18. 18.Zhenyu Chen, Ali Gholami, Matthias Nießner, and Angel X Chang. Scan2cap: Context-aware dense captioning in rgb-d scans. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  19. 19.Jieming Cui, Tengyu Liu, Nian Liu, Yaodong Yang, Yixin Zhu, and Siyuan Huang. Anyskill: Learning open-vocabulary physical skill for interactive agents. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 3
  20. 20.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 5
  21. 21.Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra. Embodied question answering. In Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 2
  22. 22.Shengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen, and Kui Jia. 3d affordancenet: A benchmark for visual object affordance understanding. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
  23. 23.Kuan Fang, Te-Lin Wu, Daniel Yang, Silvio Savarese, and Joseph J Lim. Demo2vec: Reasoning object affordances from online videos. In Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 3
  24. 24.Anindita Ghosh, Noshaba Cheema, Cennet Oguz, Christian Theobalt, and Philipp Slusallek. Synthesis of compositional animations from textual descriptions. In International Conference on Computer Vision (ICCV), 2021. 2, 3
  25. 25.Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Christian Theobalt, and Philipp Slusallek. Imos: Intent-driven full-body motion synthesis for human-object interactions. In Computer Graphics Forum, 2023. 3
  26. 26.James J Gibson. The theory of affordances. Hilldale, USA, 1(2):67–82, 1977. 3
  27. 27.Helmut Grabner, Juergen Gall, and Luc Van Gool. What makes a chair a chair? In Conference on Computer Vision and Pattern Recognition (CVPR), 2011. 3
  28. 28.Chuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, and Li Cheng. Action2motion: Conditioned generation of 3d human motions. In International Conference on Multimedia, 2020. 2, 3
  29. 29.Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2, 3, 4, 5, 6, A3
  30. 30.Abhinav Gupta, Scott Satkin, Alexei A. Efros, and Martial Hebert. From 3d scene geometry to human workspace. In Conference on Computer Vision and Pattern Recognition (CVPR), 2011. 3
  31. 31.Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J Black. Resolving 3d human pose ambiguities with 3d scene constraints. In International Conference on Computer Vision (ICCV), 2019. 2, 3, 5
  32. 32.Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael J Black. Stochastic scene-aware motion prediction. In International Conference on Computer Vision (ICCV), 2021. 2, 3
  33. 33.Mohamed Hassan, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas, and Michael J Black. Populating 3d scenes by learning human-scene interaction. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
  34. 34.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (NeurIPS), 2020. 3
  35. 35.Daniel Holden, Jun Saito, and Taku Komura. A deep learning framework for character motion synthesis and editing. ACM Transactions on Graphics (TOG), 35(4):1–11, 2016. 2
  36. 36.Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An embodied generalist agent in 3d world. arXiv preprint arXiv:2311.12871, 2023. 2
  37. 37.Siyuan Huang, Zan Wang, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu, Wei Liang, and Song-Chun Zhu. Diffusion-based generation, optimization, and planning in 3d scenes. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3
  38. 38.Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. Perceiver: General perception with iterative attention. In International Conference on Machine Learning (ICML), 2021. 2, 5
  39. 39.Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al. Perceiver io: A general architecture for structured inputs & outputs. In International Conference on Learning Representations (ICLR), 2022. 2, 5
  40. 40.Krishna Murthy Jatavallabhula, Alihusein Kuwajerwala, Qiao Gu, Mohd Omama, Tao Chen, Shuang Li, Ganesh Iyer, Soroush Saryazdi, Nikhil Keetha, Ayush Tewari, et al. Conceptfusion: Open-set multimodal 3d mapping. arXiv preprint arXiv:2302.07241, 2023. 2
  41. 41.Baoxiong Jia, Yixin Chen, Huangyue Yu, Yan Wang, Xuesong Niu, Tengyu Liu, Qing Li, and Siyuan Huang. Sceneverse: Scaling 3d vision-language learning for grounded scene understanding. arXiv preprint arXiv:2401.09340, 2024. 2
  42. 42.Hanwen Jiang, Shaowei Liu, Jiashun Wang, and Xiaolong Wang. Hand-object contact consistency reasoning for human grasps generation. In International Conference on Computer Vision (ICCV), 2021. 3
  43. 43.Nan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui, Zhiyuan Zhang, Yixin Chen, He Wang, Yixin Zhu, and Siyuan Huang. Full-body articulated human-object interaction. In International Conference on Computer Vision (ICCV), 2023. 3
  44. 44.Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction modeling. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 3
  45. 45.Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. Lerf: Language embedded radiance fields. In International Conference on Computer Vision (ICCV), 2023. 2
  46. 46.Mia Kokic, Danica Kragic, and Jeannette Bohg. Learning task-oriented grasping from human activity datasets. IEEE Robotics and Automation Letters (RA-L), 5(2):3352–3359, 2020. 3
  47. 47.Hema S Koppula and Ashutosh Saxena. Physically grounded spatio-temporal object affordances. In European Conference on Computer Vision (ECCV), 2014. 3
  48. 48.Sumith Kulal, Tim Brooks, Alex Aiken, Jiajun Wu, Jimei Yang, Jingwan Lu, Alexei A Efros, and Krishna Kumar Singh. Putting people in their place: Affordance-aware human insertion into scenes. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  49. 49.Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, and Jan Kautz. Dancing to music. In Advances in Neural Information Processing Systems (NeurIPS), 2019. 2, 3
  50. 50.Chen Li, Zhen Zhang, Wee Sun Lee, and Gim Hee Lee. Convolutional sequence to sequence model for human dynamics. In Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 3
  51. 51.Jing Li, Di Kang, Wenjie Pei, Xuefei Zhe, Ying Zhang, Zhenyu He, and Linchao Bao. Audio2gestures: Generating diverse gestures from speech audio with conditional variational autoencoders. In International Conference on Computer Vision (ICCV), 2021. 2, 3
  52. 52.Jiaman Li, Jiajun Wu, and C Karen Liu. Object motion guided human motion synthesis. ACM Transactions on Graphics (TOG), 42(6):1–11, 2023. 3
  53. 53.Puhao Li, Tengyu Liu, Yuyang Li, Yiran Geng, Yixin Zhu, Yaodong Yang, and Siyuan Huang. Gendexgrasp: Generalizable dexterous grasping. In International Conference on Robotics and Automation (ICRA), 2023. 3
  54. 54.Xueting Li, Sifei Liu, Kihwan Kim, Xiaolong Wang, Ming-Hsuan Yang, and Jan Kautz. Putting humans in a scene: Learning affordance in 3d indoor environments. In Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 3
  55. 55.Yuyang Li, Bo Liu, Puhao Li, Yaodong Yang, Yixin Zhu, Tengyu Liu, and Siyuan Huang. Grasp multiple objects with one hand. IEEE Robotics and Automation Letters (RA-L), 9(5):4027–4034, 2024. 3
  56. 56.Hung Yu Ling, Fabio Zinno, George Cheng, and Michiel Van De Panne. Character controllers using motion vaes. ACM Transactions on Graphics (TOG), 39(4):40–1, 2020. 3
  57. 57.Xiaojian Ma, Silong Yong, Zilong Zheng, Qing Li, Yitao Liang, Song-Chun Zhu, and Siyuan Huang. Sqa3d: Situated question answering in 3d scenes. In International Conference on Learning Representations (ICLR), 2023. 2
  58. 58.Naureen Mahmood, Nima Ghorbani, Nikolaus F Troje, Gerard Pons-Moll, and Michael J Black. Amass: Archive of motion capture as surface shapes. In International Conference on Computer Vision (ICCV), 2019. 5
  59. 59.Wei Mao, Richard I Hartley, Mathieu Salzmann, et al. Contact-aware human motion forecasting. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 3, 4
  60. 60.Tushar Nagarajan and Kristen Grauman. Learning affordance landscapes for interaction exploration in 3d environments. In Advances in Neural Information Processing Systems (NeurIPS), 2020. 3
  61. 61.Tushar Nagarajan, Christoph Feichtenhofer, and Kristen Grauman. Grounded human-object interaction hotspots from video. In Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 3
  62. 62.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. Expressive body capture: 3d hands, face, and body from a single image. In Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 4, A1
  63. 63.Songyou Peng, Kyle Genova, Chiyu Jiang, Andrea Tagliasacchi, Marc Pollefeys, Thomas Funkhouser, et al. Openscene: 3d scene understanding with open vocabularies. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
  64. 64.Xue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine, and Angjoo Kanazawa. Amp: Adversarial motion priors for stylized physics-based character control. ACM Transactions on Graphics (TOG), 40(4):1–20, 2021. 3
  65. 65.Xue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine, and Sanja Fidler. Ase: Large-scale reusable adversarial skill embeddings for physically simulated characters. ACM Transactions on Graphics (TOG), 41(4):1–17, 2022. 3
  66. 66.Mathis Petrovich, Michael J Black, and Gul Varol. Action-conditioned 3d human motion synthesis with transformer vae. In International Conference on Computer Vision (ICCV), 2021. 2, 3
  67. 67.Mathis Petrovich, Michael J Black, and Gul Varol. Temos: Generating diverse human motions from textual descriptions. In European Conference on Computer Vision (ECCV), 2022. 2, 3
  68. 68.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021. 2
  69. 69.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 5
  70. 70.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning (ICML), 2015. 3
  71. 71.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. In Advances in Neural Information Processing Systems (NeurIPS), 2019. 3
  72. 72.Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al. The replica dataset: A digital replica of indoor spaces. arXiv preprint arXiv:1906.05797, 2019. 5
  73. 73.Omid Taheri, Vasileios Choutas, Michael J Black, and Dimitrios Tzionas. Goal: Generating 4d whole-body motion for hand-object grasping. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
  74. 74.Ayça Takmaz, Elisabetta Fedele, Robert W Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann. OpenMask3D: Open-Vocabulary 3D Instance Segmentation. In Advances in Neural Information Processing Systems (NeurIPS), 2023. 2
  75. 75.Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. Motionclip: Exposing human motion generation to clip space. In European Conference on Computer Vision (ECCV), 2022. 2, 3
  76. 76.Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. Human motion diffusion model. In International Conference on Learning Representations (ICLR), 2023. 2, 3, 5, 6, A1
  77. 77.Jesse Thomason, Mohit Shridhar, Yonatan Bisk, Chris Paxton, and Luke Zettlemoyer. Language grounding with 3d objects. In Conference on Robot Learning (CoRL), 2022. 2
  78. 78.Matthew Thorne, David Burke, and Michiel Van De Panne. Motion doodles: an interface for sketching character motion. ACM Transactions on Graphics (TOG), 23(3):424–431, 2004. 2
  79. 79.Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Thanh Nguyen, and Sai-Kit Yeung. Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data. In International Conference on Computer Vision (ICCV), 2019. 5, A2
  80. 80.Jiashun Wang, Huazhe Xu, Jingwei Xu, Sifei Liu, and Xiaolong Wang. Synthesizing long-term 3d human motion and interaction in 3d scenes. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2, 3
  81. 81.Jingbo Wang, Sijie Yan, Bo Dai, and Dahua Lin. Scene-aware generative network for human motion synthesis. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
  82. 82.Jingbo Wang, Yu Rong, Jingyuan Liu, Sijie Yan, Dahua Lin, and Bo Dai. Towards diverse and natural scene-aware 3d human motion synthesis. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2, 3
  83. 83.Xiaolong Wang, Rohit Girdhar, and Abhinav Gupta. Binge watching: Scaling affordance learning from sitcoms. In Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 3
  84. 84.Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, Wei Liang, and Siyuan Huang. Humanise: Language-conditioned human motion generation in 3d scenes. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 2, 3, 5, 6, A1
  85. 85.Yueh-Hua Wu, Jiashun Wang, and Xiaolong Wang. Learning generalizable dexterous manipulation from human grasp affordance. In Conference on Robot Learning (CoRL), 2023. 3
  86. 86.Zeqi Xiao, Tai Wang, Jingbo Wang, Jinkun Cao, Wenwei Zhang, Bo Dai, Dahua Lin, and Jiangmiao Pang. Unified human-scene interaction via prompted chain-of-contacts. In International Conference on Learning Representations (ICLR), 2024. 3
  87. 87.Kevin Xie, Tingwu Wang, Umar Iqbal, Yunrong Guo, Sanja Fidler, and Florian Shkurti. Physics-based human motion estimation and synthesis from videos. In International Conference on Computer Vision (ICCV), 2021. 3
  88. 88.Shuquan Ye, Dongdong Chen, Songfang Han, and Jing Liao. 3d question answering. IEEE Transactions on Visualization and Computer Graph (TVCG), pages 1–16, 2022. 2
  89. 89.Hongwei Yi, Hualin Liang, Yifei Liu, Qiong Cao, Yandong Wen, Timo Bolkart, Dacheng Tao, and Michael J Black. Generating holistic 3d human motion from speech. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3
  90. 90.Ye Yuan and Kris Kitani. Dlow: Diversifying latent flows for diverse human motion prediction. In European Conference on Computer Vision (ECCV), 2020. 3
  91. 91.Zhihao Yuan, Xu Yan, Yinghong Liao, Yao Guo, Guanbin Li, Shuguang Cui, and Zhen Li. X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  92. 92.Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen. T2m-gpt: Generating human motion from textual descriptions with discrete representations. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3
  93. 93.Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. Motiondiffuse: Text-driven human motion generation with diffusion model. arXiv preprint arXiv:2208.15001, 2022. 2, 3
  94. 94.Siwei Zhang, Yan Zhang, Qianli Ma, Michael J Black, and Siyu Tang. Place: Proximity learning of articulation and contact in 3d environments. In International Conference on 3D Vision (3DV), 2020. 3
  95. 95.Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Vladimir Guzov, and Gerard Pons-Moll. Couch: Towards controllable human-chair interactions. In European Conference on Computer Vision (ECCV), 2022. 3
  96. 96.Yan Zhang and Siyu Tang. The wanderings of odysseus in 3d scenes. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
  97. 97.Yan Zhang, Mohamed Hassan, Heiko Neumann, Michael J Black, and Siyu Tang. Generating 3d people in scenes without people. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2, 3, 6
  98. 98.Yiming Zhang, ZeMing Gong, and Angel X Chang. Multi3drefer: Grounding text description to multiple 3d objects. In International Conference on Computer Vision (ICCV), 2023. 2
  99. 99.Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. Point transformer. In International Conference on Computer Vision (ICCV), 2021. 5, A1
  100. 100.Kaifeng Zhao, Shaofei Wang, Yan Zhang, Thabo Beeler, and Siyu Tang. Compositional human-scene interaction synthesis with semantic control. In European Conference on Computer Vision (ECCV), 2022. 2, 3
  101. 101.Kaifeng Zhao, Yan Zhang, Shaofei Wang, Thabo Beeler, and Siyu Tang. Synthesizing diverse human motions in 3d indoor scenes. In International Conference on Computer Vision (ICCV), 2023. 3
  102. 102.Wentao Zhu, Xiaoxuan Ma, Dongwoo Ro, Hai Ci, Jinlu Zhang, Jiaxin Shi, Feng Gao, Qi Tian, and Yizhou Wang. Human motion generation: A survey. Transactions on Pattern Analysis and Machine Intelligence (TPAMI), pages 1–20, 2023. 3
  103. 103.Yuke Zhu, Alireza Fathi, and Li Fei-Fei. Reasoning about object affordances in a knowledge base representation. In European Conference on Computer Vision (ECCV), 2014. 3
  104. 104.Yixin Zhu, Chenfanfu Jiang, Yibiao Zhao, Demetri Terzopoulos, and Song-Chun Zhu. Inferring forces and learning human utilities from videos. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 3
  105. 105.Ziyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng, Siyuan Huang, and Qing Li. 3d-vista: Pre-trained transformer for 3d vision and text alignment. In International Conference on Computer Vision (ICCV), 2023. 2

Citation

MLA
Wang, Z., et al. “Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance”. arXiv, 2024, http://arxiv.org/abs/2403.18036v1.
APA
Wang, Z., Chen, Y., Jia, B., Li, P., Zhang, J., Zhang, J., Liu, T., Zhu, Y., Liang, W., & Huang, S. (2024). Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance. arXiv. http://arxiv.org/abs/2403.18036v1
Chicago
Wang, Z., Y. Chen, B. Jia, et al. 2024. “Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance”. arXiv. http://arxiv.org/abs/2403.18036v1.
Harvard
Wang, Z. et al. (2024) “Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.18036v1.
Vancouver
1. Wang Z, Chen Y, Jia B, Li P, Zhang J, Zhang J, Liu T, Zhu Y, Liang W, Huang S (2024) Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance. arXiv

BibTeX

@article{wang2024move,
  title = {Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance},
  author = {Wang, Zan and Chen, Yixin and Jia, Baoxiong and Li, Puhao and Zhang, Jinlu and Zhang, Jingze and Liu, Tengyu and Zhu, Yixin and Liang, Wei and Huang, Siyuan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.18036v1},
  eprint = {2403.18036}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE