Unified Human-Scene Interaction via Prompted Chain-of-Contacts

Zeqi XiaoTai WangJingbo WangJinkun CaoWenwei ZhangBo DaiDahua LinJiangmiao Pang

article2024ICLR130 citations

Develops UniHSI, a framework that uses large language models to translate text commands into sequential joint-object contact plans, enabling versatile and controllable human-scene interactions across diverse 3D environments.

Listen

Generating realistic human behaviors that interact naturally with indoor environments is essential for virtual reality and embodied artificial intelligence. However, existing approaches often struggle with practical deployment because they rely on separate, manually designed control policies for different tasks and require labor-intensive motion capture datasets paired with specific 3D objects. Furthermore, enabling characters to understand flexible natural language instructions across multi-step sequences remains a major hurdle.

The article introduces and evaluates UniHSI, a unified physics-based framework designed to execute diverse, multi-step human-scene interactions driven directly by natural language commands. The core objective is to demonstrate that complex interactions can be standardized into structured contact sequences, enabling a single reinforcement learning controller to execute versatile tasks without requiring paired interaction annotations.

To achieve this, the approach breaks down interactions into a "Chain of Contacts," which represents an action as ordered pairs of humanoid body joints and specific object parts. The framework uses a high-level large language model planner to translate user text instructions into structured contact chains, which are then passed to a unified low-level controller. This controller leverages reinforcement learning, an adversarial motion prior to ensure natural movement style, and an ego-centric heightmap to prevent environmental collisions in physics simulations. The authors evaluated the framework across thousands of generated plans using object models from PartNet and realistic indoor environments from ScanNet, comparing performance against established single-task and multi-task baselines.

The findings show that UniHSI achieves high execution accuracy across varied tasks, reaching an 85.5% success rate on simple single-object tasks and maintaining strong performance across multi-step sequences. In direct baseline comparisons, the unified model achieved an 81.5% success rate on complex "Lie Down" tasks—far surpassing standard single-task baselines (21.3%) and vanilla multi-task baseline combinations (20.1%)—while delivering a 94.3% success rate on "Sit" tasks. Ablation experiments demonstrated that dynamic reward balancing through adaptive contact weights is critical; removing this component caused success rates on simple tasks to plummet from 85.5% to 21.2%. Additionally, evaluations of language planners revealed that GPT-4 substantially outperformed GPT-3.5 in planning correctness (71.9% versus 49.1%) and execution success rate (57.3% versus 35.6%), although both trailed human planning (73.2% execution success).

These results indicate that decomposing physical interactions into contact-point representations allows diverse interactive skills to be learned within a single unified model without task-specific engineering. This significantly lowers data collection costs by removing the need for manual interaction annotations and reduces the risk of unnatural motion artifacts. Organizations developing interactive virtual environments and robotics simulations can adopt this structure to scale human-agent behaviors more efficiently.

Moving forward, developers should prioritize stronger spatial reasoning models or enhanced prompt verification systems to mitigate language planner failures, particularly when handling intricate multi-object sequences where task success rates drop to 40.5% in synthetic environments and 22.3% in complex real scans. Future research and pilot implementations should focus on integrating language models directly into the training loop and expanding capabilities to dynamic scenarios involving movable or carried objects.

Confidence in these findings is solid for static indoor environments, supported by extensive automated metrics and human user studies that rated the generated motions significantly higher in naturalness and semantic alignment than prior baselines. However, decision-makers should note key boundary conditions: the current system is restricted to interactions with static, fixed objects, and the overall system remains subject to occasional reasoning and spatial errors from upstream language models.

Cover for Unified Human-Scene Interaction via Prompted Chain-of-Contacts

Abstract

Human-Scene Interaction (HSI) is a vital component of fields like embodied AI and virtual reality. Despite advancements in motion quality and physical plausibility, two pivotal factors, versatile interaction control and the development of a user-friendly interface, require further exploration before the practical application of HSI. This paper presents a unified HSI framework, UniHSI, which supports unified control of diverse interactions through language commands. This framework is built upon the definition of interaction as Chain of Contacts (CoC): steps of human joint-object part pairs, which is inspired by the strong correlation between interaction types and human-object contact regions. Based on the definition, UniHSI constitutes a Large Language Model (LLM) Planner to translate language prompts into task plans in the form of CoC, and a Unified Controller that turns CoC into uniform task execution. To facilitate training and evaluation, we collect a new dataset named ScenePlan that encompasses thousands of task plans generated by LLMs based on diverse scenarios. Comprehensive experiments demonstrate the effectiveness of our framework in versatile task execution and generalizability to real scanned scenes. The project page is at this https URL .

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Methodology
  • 3.1 Chain of Contacts
  • 3.2 Large Language Model Planner
  • 3.3 Unified Controller
  • 4 Experiments
  • 4.1 Datasets and Metrics
  • 4.2 Performance on ScenePlan
  • 4.3 Ablation Studies
  • 4.3.1 Key Components Ablation
  • 4.3.2 Design Comparison with Previous Methods
  • 5 Conclusion
  • A Limitations and Future Work.
  • B Implementation Details
  • C Detailed prompting example of the LLM Planner
  • D Details of the ScenePlan
  • E More Visualizations
  • F Demonstration of failure planning
  • G User Study on Motion Reality.
  • References

Knowls

  1. Knowl 1 — Chain of Contacts (CoC) Representation for Human-Scene Interaction

    definition

    In human-scene interaction (HSI), an interaction sequence is universally formulated as a Chain of Contacts (CoC), denoted by:

    C={S1,S2,…,SN}C = \{S_1, S_2, \dots, S_N\}

    where SiS_i represents the ii-th sequential contact step.

    Each contact step SS is composed of a set of contact tuples:

    S={{o1,p1,j1,c1,d1},{o2,p2,j2,c2,d2},…,{on,pn,jn,cn,dn}}S = \{\{o_1, p_1, j_1, c_1, d_1\}, \{o_2, p_2, j_2, c_2, d_2\}, \dots, \{o_n, p_n, j_n, c_n, d_n\}\}

    where:

    • oko_k identifies the target scene object.
    • pkp_k designates a specific semantic part of object oko_k (e.g., seat surface, mattress, pillow, keyboard) or another humanoid joint when modeling joint-to-joint contacts (such as placing a foot on a knee).
    • jkj_k designates the humanoid body joint involved in the interaction (e.g., pelvis, head, left hand, right foot).
    • ck∈{contact,not contact,not care}c_k \in \{\text{contact}, \text{not contact}, \text{not care}\} indicates the target contact state between joint jkj_k and part pkp_k.
    • dk∈{up,down,front,back,left,right,none}d_k \in \{\text{up}, \text{down}, \text{front}, \text{back}, \text{left}, \text{right}, \text{none}\} indicates the relative spatial approach or contact direction from joint jkj_k to part pkp_k.

    For non-contact navigation actions (e.g., approaching an object), pseudo-contact tuples of the form {ok,none,none,none,dk}\{o_k, \text{none}, \text{none}, \text{none}, d_k\} are defined, specifying that the humanoid's pelvis (root) should move toward object oko_k along direction dkd_k.

  2. Knowl 2 — Prompted LLM Planner for Synthesizing Chains of Contacts

    model/method

    UniHSI employs a Large Language Model (LLM) Planner to translate free-form natural language instructions into executable task plans represented as Chains of Contacts (CoC). The LLM is conditioned on:

    1. A natural language command L\mathcal{L} (e.g., "I want to play video games for a while, then go to sleep.").
    2. Background scene information B\mathcal{B}, specifying the list of objects O\mathcal{O}, their constituent parts P\mathcal{P}, and optional spatial relationships.
    3. The humanoid joint set J\mathcal{J} comprising 15 whole-body joints (pelvis, head, torso, left/right hips, knees, feet, shoulders, elbows, and hands).
    4. In-context rules and demonstration examples.

    The LLM decomposes the user command into sequential sub-goals. For navigation, it generates pseudo-contact pairs directing humanoid movement toward target objects. For physical interactions, it generates contact pairs {o,p,j,c,d}\{o, p, j, c, d\} specifying the exact joints, object parts, contact types, and approach directions. This architecture allows the generation of diverse interaction plans without requiring manually annotated motion-interaction datasets.

  3. Knowl 3 — TaskParser and Contact-Driven Goal Reward Formulation

    model/method

    The TaskParser converts the symbolic Chain of Contacts (CoC) step S={{ok,pk,jk,ck,dk}}k=1nS = \{\{o_k, p_k, j_k, c_k, d_k\}\}_{k=1}^n into uniform continuous state observations and reinforcement learning reward signals.

    For each contact pair kk:

    1. The 3D position of humanoid joint jkj_k, denoted vkj∈R3v^j_k \in \mathbb{R}^3, and the point cloud vkp∈Rm×3v^p_k \in \mathbb{R}^{m \times 3} of object part pkp_k (containing mm points) are extracted from the simulation.
    2. The nearest point vknp∈vkpv^{np}_k \in v^p_k to vkjv^j_k is identified.
    3. The distance vector is defined as dk=vknp−vkj∈R3d_k = v^{np}_k - v^j_k \in \mathbb{R}^3.
    4. The individual contact reward Rk∈[0,1]R_k \in [0, 1] is defined by:

    Rk={wdisexp⁡(−wdk∥dk∥)+wdirmax⁡(d^k⋅uk,0),if ck=contact1−exp⁡(−wdk∥dk∥),if ck=not contact1,if ck=not careR_k = \begin{cases} w_{\text{dis}} \exp(-w_{dk} \|d_k\|) + w_{\text{dir}} \max(\hat{d}_k \cdot \mathbf{u}_k, 0), & \text{if } c_k = \text{contact} \\ 1 - \exp(-w_{dk} \|d_k\|), & \text{if } c_k = \text{not contact} \\ 1, & \text{if } c_k = \text{not care} \end{cases}

    where d^k=dk∥dk∥\hat{d}_k = \frac{d_k}{\|d_k\|} is the unit distance vector, uk\mathbf{u}_k is the unit direction vector specified by dkd_k, and wdis,wdir,wdkw_{\text{dis}}, w_{\text{dir}}, w_{dk} are weighting coefficients.

    The composite task reward RGR^G for step SS is computed as:

    RG=∑k=1nwkRkR^G = \sum_{k=1}^n w_k R_k

    where wkw_k is the adaptive weight assigned to contact pair kk.

    The uniform task observation vector fed to the policy network is:

    sU={{v1np−v1j,c1,d1},{v2np−v2j,c2,d2},…,{vnnp−vnj,cn,dn}}s^U = \{\{v^{np}_1 - v^j_1, c_1, d_1\}, \{v^{np}_2 - v^j_2, c_2, d_2\}, \dots, \{v^{np}_n - v^j_n, c_n, d_n\}\}

  4. Knowl 4 — Adaptive Contact Weights for Multi-Contact Optimization

    equation

    To balance the optimization of multiple simultaneous contact pairs without manual tuning, UniHSI calculates the reward weight wkw_k of the kk-th contact pair adaptively based on the individual rewards Rk∈[0,1]R_k \in [0, 1]:

    wk=1−Rkn−∑i=1nRi+ϵw_k = \frac{1 - R_k}{n - \sum_{i=1}^n R_i + \epsilon}

    where:

    • Rk∈[0,1]R_k \in [0, 1] is the current reward achieved by the kk-th contact pair.
    • nn is the total number of contact pairs in the active contact step.
    • ϵ>0\epsilon > 0 is a small regularization constant to prevent division by zero.

    This formulation dynamically increases the relative weight for unsatisfied or difficult contact pairs (Rk≈0R_k \approx 0) while diminishing the weight of contact pairs that are already satisfied (Rk≈1R_k \approx 1).

  5. Knowl 5 — Step Completion Criterion for Sequential Contact Execution

    definition

    In the UniHSI TaskParser, a contact step SS comprising nn contact pairs {{ok,pk,jk,ck,dk}}k=1n\{\{o_k, p_k, j_k, c_k, d_k\}\}_{k=1}^n is evaluated as completed at simulation time tt if and only if all nn pairs satisfy their respective criteria simultaneously:

    • If ck=contactc_k = \text{contact}: ∥dk∥<0.1 m\|d_k\| < 0.1\,\text{m} and d^k⋅uk>0.8\hat{d}_k \cdot \mathbf{u}_k > 0.8, where dk=vknp−vkjd_k = v^{np}_k - v^j_k, d^k=dk∥dk∥\hat{d}_k = \frac{d_k}{\|d_k\|}, and uk\mathbf{u}_k is the specified unit contact direction.
    • If ck=not contactc_k = \text{not contact}: ∥dk∥>0.1 m\|d_k\| > 0.1\,\text{m}.
    • If ck=not carec_k = \text{not care}: evaluated as True unconditionally.

    When all conditions are met, the TaskParser sequentially advances the policy objective to step Si+1S_{i+1} in the Chain of Contacts.

  6. Knowl 6 — Unified Physical Controller and Motion Prior Architecture

    model/method

    UniHSI executes interaction tasks in the IsaacGym physics simulator using a reinforcement learning controller built on the Adversarial Motion Priors (AMP) framework with Proximal Policy Optimization (PPO). The total reward is:

    R(st,at,st+1,G)=wGRG(st,at,st+1,G)+wSRS(st,st+1)R(s_t, a_t, s_{t+1}, G) = w^G R^G(s_t, a_t, s_{t+1}, G) + w^S R^S(s_t, s_{t+1})

    where RGR^G is the contact-driven goal reward, wGw^G and wSw^S are empirical weights, and RSR^S is the motion style reward:

    RS(st,st+1)=−log⁡(1−D(stA,st+1A))R^S(s_t, s_{t+1}) = -\log(1 - D(s_t^A, s_{t+1}^A))

    The discriminator DD receives 10 consecutive frames sA=Φ(s)s^A = \Phi(s) and is trained with a least-squares GAN loss and gradient penalty:

    min⁡D−EdM[log⁡D(stA,st+1A)]−Edπ[log⁡(1−D(stA,st+1A))]+wgpEdM[∥∇ϕD(ϕ)∣ϕ=(stA,st+1A)∥2]\min_D -\mathbb{E}_{d^{\mathcal{M}}}[\log D(s_t^A, s_{t+1}^A)] - \mathbb{E}_{d^\pi}[\log(1 - D(s_t^A, s_{t+1}^A))] + w^{\text{gp}} \mathbb{E}_{d^{\mathcal{M}}}\left[\|\nabla_\phi D(\phi)|_{\phi=(s_t^A, s_{t+1}^A)}\|^2\right]

    where dMd^{\mathcal{M}} and dπd^\pi denote the state transition distributions from motion capture datasets M\mathcal{M} (SAMP and CIRCLE) and policy π\pi, respectively, and wgpw^{\text{gp}} regularizes the gradient penalty.

    To avoid obstacles in cluttered environments, the humanoid policy observes an ego-centric square heightmap of surrounding geometry. The policy network combines a CNN layer for the heightmap with two MLP layers ([1024,1024,512][1024, 1024, 512] units) for actor and critic, trained with the Adam optimizer (learning rate 2×10−52 \times 10^{-5}) across 8,192 parallel environments on a single NVIDIA A100 GPU.

  7. Knowl 7 — The ScenePlan Dataset

    experimental setup

    The ScenePlan dataset provides indoor 3D environments and corresponding LLM-generated Chain of Contacts (CoC) interaction plans:

    • Source data: 3D objects with part annotations from PartNet and 3D reconstructed rooms from ScanNet.
    • Training set: 40 PartNet objects, with 5 to 20 plausible interaction steps generated per object. In each training episode, 1 to 4 objects are randomly sampled and assembled into a scenario with combined multi-step plans.
    • Evaluation set: 40 PartNet objects and 10 real ScanNet room scenes. Contains 1,040 test plans for PartNet scenarios and 100 test plans for ScanNet scenarios.
    • Difficulty classification:
      • Simple: ≤3\le 3 contact steps involving 1 object.
      • Medium: >3> 3 contact steps involving 1 object.
      • Hard: >3> 3 contact steps involving multiple objects sequentially.
    • Motion datasets: Motion priors are trained using 100 minutes of MoCap clips from SAMP (walking, sitting, lying down) and 20 selected clips from CIRCLE (reaching).
  8. Knowl 8 — Performance Evaluation and Ablations on the ScenePlan Dataset

    data/table

    Quantitative evaluation of UniHSI on the ScenePlan dataset across difficulty levels (Simple, Mid, Hard). Metrics include Success Rate (percentage of episodes completing all nn steps within n×10n \times 10 seconds), Contact Error (average distance/non-contact error across active contact pairs in meters), and Success Steps (average count of successfully executed steps):

    Source / Variant Success Rate (%) ↑\uparrow Contact Error (m) ↓\downarrow Success Steps
    Simple Mid Hard Simple Mid Hard Simple Mid Hard
    PartNet 85.5 67.9 40.5 0.035 0.037 0.040 2.1 4.1 4.8
    wo Adaptive Weights 21.2 5.3 0.1 0.181 0.312 0.487 0.7 1.2 0.0
    wo Heightmap 61.6 45.7 0.0 0.068 0.076 - 1.8 3.4 0.0
    ScanNet 73.2 43.1 22.3 0.061 0.072 0.062 2.2 3.5 4.8

    Ablating Adaptive Weights severely impairs performance across all levels (e.g., Simple success rate drops from 85.5%85.5\% to 21.2%21.2\%), demonstrating its necessity for multi-contact coordination. Removing the ego-centric heightmap reduces Hard multi-object success to 0.0%0.0\% due to collisions. On unseen scanned ScanNet scenes with cluttered geometry, UniHSI demonstrates zero-shot generalizability.

  9. Knowl 9 — Benchmark Comparison on Standard Atomic Interaction Tasks

    data/table

    Performance comparison of UniHSI against kinematic and physics-based baselines across 70 objects (30 for sitting, 30 for lying down, 10 for reaching) with 4,096 evaluation trials per task under random orientations and object placements:

    Methods Success Rate (%) ↑\uparrow Contact Error (m) ↓\downarrow
    Sit Lie Down Reach Sit Lie Down Reach
    NSM (Starke et al., 2019) 75.0 - - 0.19 - -
    SAMP (Hassan et al., 2021a) 75.0 50.0 - 0.06 0.05 -
    InterPhys (Hassan et al., 2023) 93.7 80.0 - 0.09 0.30 -
    AMP (Peng et al., 2021) - Sit 77.3 - - 0.090 - -
    AMP - Lie Down - 21.3 - - 0.112 -
    AMP - Reach - - 98.1 - - 0.016
    AMP - Vanilla Combination (VC) 62.5 20.1 90.3 0.093 0.108 0.032
    UniHSI (Ours) 94.3 81.5 97.5 0.032 0.061 0.016

    While previous physics-based baselines require separate policy networks tailored to individual tasks, UniHSI unifies diverse tasks into a single policy that matches or outperforms task-specific policies. In contrast, naively training a single AMP model across multiple tasks (AMP Vanilla Combination) causes multi-task interference and significant performance drops (62.5%62.5\% on Sit, 20.1%20.1\% on Lie Down).

  10. Knowl 10 — LLM Planner Comparison for Plan Correctness and Physical Execution

    empirical result

    The effectiveness of different Large Language Models as high-level planners was evaluated across 100 language instructions with 10 physical simulation trials per generated plan:

    • Planning Correctness (PC, %): Percentage of plans evaluated by human annotators as semantically and logically valid.
    • Execution Success Rate (ESR, %): Percentage of trials where the low-level controller successfully executed the planned steps to completion.

    Experimental results:

    • Human Annotator Plans: ESR of 73.2%73.2\%.
    • GPT-4: Planning Correctness of 71.9%71.9\%, ESR of 57.3%57.3\%.
    • GPT-3.5: Planning Correctness of 49.1%49.1\%, ESR of 35.6%35.6\%.

    GPT-4 significantly outperformed GPT-3.5 in generating coherent contact chains and respecting spatial constraints. Common failure modes of LLMs included missing transition steps, suboptimal spatial direction assignment, and generating out-of-distribution manipulation actions (such as opening laptop lids).

  11. Knowl 11 — Limitations of UniHSI

    limitation

    The UniHSI framework has two primary limitations:

    1. Static Object Constraint: The controller only supports humanoid interactions with rigid, fixed scene objects. Manipulating, moving, or carrying dynamic objects is currently unsupported.
    2. Decoupled Offline Planning: The LLM Planner operates as an offline pre-generation module rather than an interactive, closed-loop planner embedded directly within the online simulation training and execution pipeline.

Coverage note — None was omitted; all key contributed definitions, mathematical formulations, architectural components, dataset details, experimental benchmarks, ablation studies, LLM evaluations, and limitations are fully represented.

References

  1. 1.Joao Pedro Araújo, Jiaman Li, Karthik Vetrivel, Rishi Agarwal, Jiajun Wu, Deepak Gopinath, Alexander William Clegg, and Karen Liu. Circle: Capture in rich contextual environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 21211–21221, 2023.
  2. 2.Nikos Athanasiou, Mathis Petrovich, Michael J Black, and Gül Varol. Sinc: Spatial composition of 3d human motions for simultaneous action generation. arXiv preprint arXiv:2304.10417, 2023.
  3. 3.Emad Barsoum, John Kender, and Zicheng Liu. Hp-gan: Probabilistic 3d human motion prediction via gan. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 1418–1427, 2018.
  4. 4.Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023.
  5. 5.Xin Chen, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, and Gang Yu. Executing your commands via motion diffusion in latent space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18000–18010, 2023.
  6. 6.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5828–5839, 2017.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  8. 8.Félix G Harvey, Mike Yurick, Derek Nowrouzezahrai, and Christopher Pal. Robust motion inbetweening. ACM Transactions on Graphics (TOG), 39(4):60–1, 2020.
  9. 9.Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael J Black. Stochastic scene-aware motion prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11374–11384, 2021a.
  10. 10.Mohamed Hassan, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas, and Michael J Black. Populating 3d scenes by learning human-scene interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14708–14718, 2021b.
  11. 11.Mohamed Hassan, Yunrong Guo, Tingwu Wang, Michael Black, Sanja Fidler, and Xue Bin Peng. Synthesizing physical character-scene interactions. arXiv preprint arXiv:2302.00883, 2023.
  12. 12.Daniel Holden, Taku Komura, and Jun Saito. Phase-functioned neural networks for character control. ACM Transactions on Graphics (TOG), 36(4):1–13, 2017.
  13. 13.Siyuan Huang, Zan Wang, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu, Wei Liang, and Song-Chun Zhu. Diffusion-based generation, optimization, and planning in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16750–16761, 2023.
  14. 14.Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen. Motiongpt: Human motion as a foreign language. arXiv preprint arXiv:2306.14795, 2023.
  15. 15.Jordan Juravsky, Yunrong Guo, Sanja Fidler, and Xue Bin Peng. Padl: Language-directed physics-based character control. In SIGGRAPH Asia 2022 Conference Papers, pp. 1–9, 2022.
  16. 16.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  17. 17.Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, et al. Isaac gym: High performance gpu-based physics simulation for robot learning. arXiv preprint arXiv:2108.10470, 2021.
  18. 18.Kaichun Mo, Shilin Zhu, Angel X Chang, Li Yi, Subarna Tripathi, Leonidas J Guibas, and Hao Su. Partnet: A large-scale benchmark for fine-grained and hierarchical part-level 3d object understanding. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 909–918, 2019.
  19. 19.OpenAI. Gpt-3: Generative pre-trained transformer 3. https://openai.com/research/gpt-3, 2020.
  20. 20.OpenAI. Gpt-4 technical report, 2023.
  21. 21.Liang Pan, Jingbo Wang, Buzhen Huang, Junyu Zhang, Haofan Wang, Xu Tang, and Yangang Wang. Synthesizing physically plausible human motions in 3d scenes. arXiv preprint arXiv:2308.09036, 2023.
  22. 22.Dario Pavllo, David Grangier, and Michael Auli. Quaternet: A quaternion-based recurrent model for human motion. arXiv preprint arXiv:1805.06485, 2018.
  23. 23.Xue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine, and Angjoo Kanazawa. Amp: Adversarial motion priors for stylized physics-based character control. ACM Transactions on Graphics (ToG), 40(4):1–20, 2021.
  24. 24.Xue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine, and Sanja Fidler. Ase: Large-scale reusable adversarial skill embeddings for physically simulated characters. ACM Transactions On Graphics (TOG), 41(4):1–17, 2022.
  25. 25.Juan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez, and David Lindner. Vision-language models are zero-shot reward models for reinforcement learning. arXiv preprint arXiv:2310.12921, 2023.
  26. 26.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  27. 27.Sebastian Starke, He Zhang, Taku Komura, and Jun Saito. Neural state machine for character-scene interactions. ACM Trans. Graph., 38(6):209–1, 2019.
  28. 28.Sebastian Starke, Yiwei Zhao, Taku Komura, and Kazi Zaman. Local motion phases for learning multi-contact character movements. ACM Transactions on Graphics (TOG), 39(4):54–1, 2020.
  29. 29.Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. Motionclip: Exposing human motion generation to clip space. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXII, pp. 358–374. Springer, 2022a.
  30. 30.Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. Human motion diffusion model. arXiv preprint arXiv:2209.14916, 2022b.
  31. 31.Jingbo Wang, Yu Rong, Jingyuan Liu, Sijie Yan, Dahua Lin, and Bo Dai. Towards diverse and natural scene-aware 3d human motion synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20460–20469, 2022a.
  32. 32.Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, Wei Liang, and Siyuan Huang. Humanise: Language-conditioned human motion generation in 3d scenes. Advances in Neural Information Processing Systems, 35:14959–14971, 2022b.
  33. 33.Jungdam Won, Deepak Gopinath, and Jessica Hodgins. Physics-based character controllers using conditional vaes. ACM Transactions on Graphics (TOG), 41(4):1–12, 2022.
  34. 34.Sijie Yan, Zhizhong Li, Yuanjun Xiong, Huahan Yan, and Dahua Lin. Convolutional sequence generation for skeleton-based action synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4394–4402, 2019.
  35. 35.Heyuan Yao, Zhenhua Song, Baoquan Chen, and Libin Liu. Controlvae: Model-based learning of generative controllers for physics-based characters. ACM Transactions on Graphics (TOG), 41(6):1–16, 2022.
  36. 36.Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen. T2m-gpt: Generating human motion from textual descriptions with discrete representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023a.
  37. 37.Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. Motiondiffuse: Text-driven human motion generation with diffusion model. arXiv preprint arXiv:2208.15001, 2022a.
  38. 38.Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Vladimir Guzov, and Gerard Pons-Moll. Couch: Towards controllable human-chair interactions. In European Conference on Computer Vision, pp. 518–535. Springer, 2022b.
  39. 39.Yaqi Zhang, Di Huang, Bin Liu, Shixiang Tang, Yan Lu, Lu Chen, Lei Bai, Qi Chu, Nenghai Yu, and Wanli Ouyang. Motiongpt: Finetuned llms are general-purpose motion generators. arXiv preprint arXiv:2306.10900, 2023b.
  40. 40.Kaifeng Zhao, Shaofei Wang, Yan Zhang, Thabo Beeler, and Siyu Tang. Compositional human-scene interaction synthesis with semantic control. In European Conference on Computer Vision, pp. 311–327. Springer, 2022.
  41. 41.Kaifeng Zhao, Yan Zhang, Shaofei Wang, Thabo Beeler, and Siyu Tang. Synthesizing diverse human motions in 3d indoor scenes. arXiv preprint arXiv:2305.12411, 2023.

Citation

MLA
Xiao, Z., et al. “Unified Human-Scene Interaction via Prompted Chain-of-Contacts”. arXiv, 2023, http://arxiv.org/abs/2309.07918v5.
APA
Xiao, Z., Wang, T., Wang, J., Cao, J., Zhang, W., Dai, B., Lin, D., & Pang, J. (2023). Unified Human-Scene Interaction via Prompted Chain-of-Contacts. arXiv. http://arxiv.org/abs/2309.07918v5
Chicago
Xiao, Z., T. Wang, J. Wang, et al. 2023. “Unified Human-Scene Interaction via Prompted Chain-of-Contacts”. arXiv. http://arxiv.org/abs/2309.07918v5.
Harvard
Xiao, Z. et al. (2023) “Unified Human-Scene Interaction via Prompted Chain-of-Contacts”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2309.07918v5.
Vancouver
1. Xiao Z, Wang T, Wang J, Cao J, Zhang W, Dai B, Lin D, Pang J (2023) Unified Human-Scene Interaction via Prompted Chain-of-Contacts. arXiv

BibTeX

@article{xiao2023unified,
  title = {Unified Human-Scene Interaction via Prompted Chain-of-Contacts},
  author = {Xiao, Zeqi and Wang, Tai and Wang, Jingbo and Cao, Jinkun and Zhang, Wenwei and Dai, Bo and Lin, Dahua and Pang, Jiangmiao},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2309.07918v5},
  eprint = {2309.07918}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors