MIME: Human-Aware 3D Scene Generation

Hongwei YiChun-Hao P. HuangShashank TripathiLea HeringJustus ThiesMichael J. Black

article2023CVPR73 citations

Proposes an auto-regressive transformer model that synthesizes plausible indoor 3D scenes by utilizing human motion sequences to infer room layouts, object placements, and free-space constraints.

Listen

Generating realistic 3D indoor environments populated by moving humans is essential for video games, architectural planning, and artificial intelligence training data generation, but traditional scene modeling remains expensive and labor-intensive. While previous research has focused on synthesizing human movement within pre-existing 3D spaces, relatively little work explores the reverse challenge of inferring a complete environment from human body movements alone. The article addresses this gap by developing MIME (Mining Interaction and Movement to infer 3D Environments), a computational framework that treats human body motion as an active scanner to predict full, plausible 3D room layouts.

The article's main objective is to design and evaluate an auto-regressive generative framework that uses 3D human motion and floor plans to synthesize complete indoor furniture layouts that accurately support human interactions and respect open space constraints.

To achieve this, the authors created a large synthetic training dataset called 3D-FRONT HUMAN by populating existing 3D room floor plans with moving and interacting virtual humans across four room categories: bedrooms, living rooms, dining rooms, and libraries. The methodology divides human motion into two components: free-space movement, which maps out walkable floor areas where objects cannot exist, and contact interactions (such as sitting, lying, or touching), which signal the presence and type of furniture. A transformer-based neural network processes these motion inputs alongside an empty floor plan to generate furniture bounding boxes sequentially. A post-processing refinement step then retrieves matching 3D furniture models from a catalog and optimizes their placement using geometric contact and collision rules.

The experimental findings show that MIME outperforms existing baseline models in physical plausibility and interaction accuracy. First, MIME drastically cut physical collisions between generated furniture and walkable space compared to human-unaware baseline models, achieving less than half the room interpenetration rates across living rooms (0.050 vs. 0.129), dining rooms (0.047 vs. 0.121), and bedrooms (0.129 vs. 0.348). Second, MIME greatly improved 3D spatial alignment between humans and interacting furniture, achieving contact intersection-over-union scores of 0.756 to 0.920 in common rooms compared to baseline scores ranging from 0.122 to 0.376. Third, when tested on a real-world motion-capture dataset without fine-tuning, MIME achieved an 8.47 object detection accuracy score, substantially higher than the 5.36 achieved by prior motion-conditioned reconstruction methods, while uniquely generating full room layouts rather than isolated objects in contact with the human.

These results demonstrate that human motion provides strong geometric constraints that allow automated systems to reconstruct complete, functional indoor scenes at scale. By enabling the conversion of archival motion capture data into realistic 3D environments, this approach can lower the cost and development timelines required to create synthetic training data for computer vision, architectural design, and interactive virtual reality.

The authors recommend integrating motion-aware generative methods into pipelines for large-scale synthetic data generation and digital room layout design. For future technical deployments, engineering teams should implement higher-resolution floor plan encodings and develop end-to-end models that jointly predict room boundaries and 3D shapes alongside furniture layouts.

A primary limitation of the study is its reliance on static scenes and coarse floor plan grid resolutions (where one grid unit represents approximately 10 centimeters), which can occasionally induce minor spatial collisions. Additionally, the system currently assumes all objects remain stationary, leaving dynamic interactions—such as opening doors or moving handheld objects—for future work. Despite these limitations, the quantitative improvements on both synthetic and real-world datasets support high confidence in MIME's ability to generate plausible 3D layouts from human motion.

Cover for MIME: Human-Aware 3D Scene Generation

Abstract

Generating realistic 3D worlds occupied by moving humans has many applications in games, architecture, and synthetic data creation. But generating such scenes is expensive and labor intensive. Recent work generates human poses and motions given a 3D scene. Here, we take the opposite approach and generate 3D indoor scenes given 3D human motion. Such motions can come from archival motion capture or from IMU sensors worn on the body, effectively turning human movement into a “scanner” of the 3D world. Intuitively, human movement indicates the free-space in a room and human contact indicates surfaces or objects that support activities such as sitting, lying or touching. We propose MIME (Mining Interaction and Movement to infer 3D Environments), which is a generative model of indoor scenes that produces furniture layouts that are consistent with the human movement. MIME uses an auto-regressive transformer architecture that takes the already generated objects in the scene as well as the human motion as input, and outputs the next plausible object. To train MIME, we build a dataset by populating the 3D FRONT scene dataset with 3D humans. Our experiments show that MIME produces more diverse and plausible 3D scenes than a recent generative scene method that does not know about human movement. Code and data are available for research at https://mime.is.tue.mpg.de.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Generative Human-aware Scene Synthesis
  • 3.2. Training and Inference
  • 3.3. 3D Scene Refinement
  • 4. Dataset Generation of 3D-FRONT HUMAN
  • 5. Experiments
  • 5.1. Human-aware Scene Synthesis
  • 5.2. Ablation Study
  • 6. Limitation and Discussion
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — MIME Human-Scene Representation and Conditioning Formulation

    model/method

    MIME models indoor 3D scene synthesis conditioned on 3D human motion and a 2D floor plan layout. Given human motion sequences and a floor plan F\mathcal{F}, the human input is split into two representations:

    1. Free-space humans: Poses that only contact the ground floor plane F\mathcal{F} (such as walking or standing). A 2D binary free-space mask E∈{0,1}H×W\mathcal{E} \in \{0, 1\}^{H \times W} is constructed by taking the union of all projected foot contact points onto F\mathcal{F}. A value of E(p)=1\mathcal{E}(p) = 1 indicates that pixel location pp must remain clear of furniture.
    2. Contact humans: Posed 3D human body meshes (represented using SMPL-X) are processed using POSA (Pose with Surfaces and Affordances) to predict surface vertices in contact with scene objects. For each contact human, a 3D bounding box is computed around their contact vertices. Non-maximum suppression (NMS) over 3D Intersection-over-Union (IoU) aggregates multiple interacting humans into a set of NN non-overlapping contact bounding boxes C={ci}i=1N\mathcal{C} = \{c_i\}_{i=1}^N.

    A generated 3D scene S\mathcal{S} is represented as an unordered set of 3D bounding boxes partitioned into two disjoint subsets: contact objects O={oi}i=1N\mathcal{O} = \{o_i\}_{i=1}^N (objects interacting with the contact humans C\mathcal{C}) and non-contact objects Q={qi}i=1M\mathcal{Q} = \{q_i\}_{i=1}^M (objects placed in free space without direct human interaction), such that S=O∪Q\mathcal{S} = \mathcal{O} \cup \mathcal{Q}.

  2. Knowl 2 — Autoregressive Scene Likelihood and Optimization Objective in MIME

    equation

    In MIME, the log-likelihood of generating a complete 3D scene S=O∪Q\mathcal{S} = \mathcal{O} \cup \mathcal{Q} conditioned on the 2D floor plan F\mathcal{F}, binary free-space mask E\mathcal{E}, and contact human representations C={ci}i=1N\mathcal{C} = \{c_i\}_{i=1}^N is factorized into contact object generation and non-contact object generation:

    log⁡p(S)=log⁡p(O∣F,E,C)+log⁡p(Q∣F,E,C)\log p(\mathcal{S}) = \log p(\mathcal{O} \mid \mathcal{F}, \mathcal{E}, \mathcal{C}) + \log p(\mathcal{Q} \mid \mathcal{F}, \mathcal{E}, \mathcal{C})

    The likelihood of the contact objects O\mathcal{O} is computed by accumulating probabilities across random permutations π(O)\pi(\mathcal{O}):

    p(O∣F,E,C)=∑O^∈π(O)∏j∈O^p(oj∣o<j,F,E,c≥j)p(\mathcal{O} \mid \mathcal{F}, \mathcal{E}, \mathcal{C}) = \sum_{\hat{\mathcal{O}} \in \pi(\mathcal{O})} \prod_{j \in \hat{\mathcal{O}}} p\left(o_j \mid o_{<j}, \mathcal{F}, \mathcal{E}, c_{\geq j}\right)

    where ojo_j is the jj-th generated contact object, o<jo_{<j} denotes previously generated objects, and c≥jc_{\geq j} denotes the remaining unassigned contact humans.

    Once contact objects O\mathcal{O} are placed, they are treated as existing scene entities Q′\mathcal{Q}', and the likelihood of non-contact objects Q\mathcal{Q} is formulated as:

    p(Q∣F,E,C)=p(Q∣F,E,O)=∑Q^∈π(Q∪Q′)∏j∈Q^p(qj∣q<j,F,E)p(\mathcal{Q} \mid \mathcal{F}, \mathcal{E}, \mathcal{C}) = p(\mathcal{Q} \mid \mathcal{F}, \mathcal{E}, \mathcal{O}) = \sum_{\hat{\mathcal{Q}} \in \pi(\mathcal{Q} \cup \mathcal{Q}')} \prod_{j \in \hat{\mathcal{Q}}} p\left(q_j \mid q_{<j}, \mathcal{F}, \mathcal{E}\right)

    Monte Carlo sampling is used during training to approximate permutations, ensuring permutation invariance over generated objects. Training minimizes the negative log-likelihood of object attributes across autoregressive steps.

  3. Knowl 3 — MIME Transformer Architecture and Attribute Decoding

    model/method

    MIME uses an autoregressive transformer architecture composed of input encoders, a transformer encoder, and multi-layer perceptron (MLP) attribute decoders:

    • Free-Space and Floor Plan Encoder: A ResNet-18 encodes the concatenation of the 2D floor plan F\mathcal{F} and the 2D free-space mask E\mathcal{E} into a spatial conditioning feature FF.
    • Contact and Furniture Encoder: A shared-weight encoder EθE_\theta embeds both contact human bounding boxes and already-generated scene objects. For the jj-th entity, its embedding is:

    Eθ(Ij,kj,tj,rj,sj)→(Ij,λ(kj),p(tj),p(rj),p(sj))E_\theta(I_j, k_j, t_j, r_j, s_j) \rightarrow \left(I_j, \lambda(k_j), p(t_j), p(r_j), p(s_j)\right)

    where Ij∈{0,1}I_j \in \{0, 1\} is a binary contact indicator (Ij=1I_j = 1 for the specific contact human targeted in the current generation step, and 00 for other humans and all existing furniture), λ(kj)\lambda(k_j) is a learnable embedding of semantic/interaction class kjk_j, and p(⋅)p(\cdot) is sinusoidal positional encoding applied to 3D translation tj∈R3t_j \in \mathbb{R}^3, rotation rj∈Rr_j \in \mathbb{R}, and bounding box size sj∈R3s_j \in \mathbb{R}^3.

    • Transformer Encoder: The sequence of M+NM+N object and human embeddings Ti=1M+NT_{i=1}^{M+N}, the spatial feature FF, and a learnable query vector q∈R64q \in \mathbb{R}^{64} are processed by a transformer encoder τθ(F,Ti=1M+N,q)→q^\tau_\theta(F, T_{i=1}^{M+N}, q) \rightarrow \hat{q} without sequence positional encodings.
    • Sequential Attribute Prediction: Given the predicted query feature q^\hat{q}, four sequential MLPs predict the parameters of the new object (k^,t^,r^,s^)(\hat{k}, \hat{t}, \hat{r}, \hat{s}):
      1. Semantic class category k^\hat{k} (including an end-of-scene token);
      2. Translation t^\hat{t}, conditioned on [q^,k^][\hat{q}, \hat{k}];
      3. Rotation r^\hat{r}, conditioned on [q^,k^,t^][\hat{q}, \hat{k}, \hat{t}];
      4. Size s^\hat{s}, conditioned on [q^,k^,t^,r^][\hat{q}, \hat{k}, \hat{t}, \hat{r}].
  4. Knowl 4 — Autoregressive Inference and Contact Human Pruning in MIME

    algorithm

    During inference, MIME generates 3D object bounding boxes conditioned on a floor plan F\mathcal{F}, free-space mask E\mathcal{E}, and input contact humans C\mathcal{C}. At each step, humans whose contact constraints are fulfilled by the newly generated object are pruned using ground-plane 2D Intersection-over-Union (IoU).

    Input: Floor plan F\mathcal{F}, free-space mask E\mathcal{E}, set of contact humans C={ci}i=1N\mathcal{C} = \{c_i\}_{i=1}^N
    Output: Set of 3D object bounding boxes S={oj}j=1K\mathcal{S} = \{o_j\}_{j=1}^K
    Initialize generated scene S←∅\mathcal{S} \leftarrow \emptyset
    Initialize active contact humans Cactive←C\mathcal{C}_{\text{active}} \leftarrow \mathcal{C}
    while True do
        if Cactive≠∅\mathcal{C}_{\text{active}} \neq \emptyset then
            Set contact indicator I1←1I_1 \leftarrow 1 for the first human in Cactive\mathcal{C}_{\text{active}}
            Set Ij←0I_j \leftarrow 0 for all other humans j>1j > 1 in Cactive\mathcal{C}_{\text{active}}
        end if
        
        Encode F\mathcal{F} and E\mathcal{E} into feature map FF using ResNet-18
        Encode each object in S\mathcal{S} and each human in Cactive\mathcal{C}_{\text{active}} using EθE_\theta into embeddings TT
        Compute query embedding q^←τθ(F,T,q)\hat{q} \leftarrow \tau_\theta(F, T, q)
        Predict class label k^\hat{k} from q^\hat{q}
        
        if k^\hat{k} is the end-of-scene token then
            break
        end if
        
        Predict translation t^\hat{t}, rotation r^\hat{r}, and size s^\hat{s} consecutively from q^\hat{q}
        Construct object bounding box onew←(k^,t^,r^,s^)o_{\text{new}} \leftarrow (\hat{k}, \hat{t}, \hat{r}, \hat{s})
        S←S∪{onew}\mathcal{S} \leftarrow \mathcal{S} \cup \{o_{\text{new}}\}
        
        for each contact human c∈Cactivec \in \mathcal{C}_{\text{active}} do
            Project 3D bounding box of cc and onewo_{\text{new}} onto the ground floor plane
            Compute 2D Intersection-over-Union IoU2D(c,onew)\text{IoU}_{\text{2D}}(c, o_{\text{new}})
            if IoU2D(c,onew)>0.5\text{IoU}_{\text{2D}}(c, o_{\text{new}}) > 0.5 then
                Cactive←Cactive∖{c}\mathcal{C}_{\text{active}} \leftarrow \mathcal{C}_{\text{active}} \setminus \{c\}
            end if
        end for
    end while
    return S\mathcal{S}
  5. Knowl 5 — 3D Scene Mesh Retrieval and Geometric Optimization Refinement

    model/method

    After MIME predicts 3D bounding boxes (k^,t^,r^,s^)(\hat{k}, \hat{t}, \hat{r}, \hat{s}), 3D CAD mesh models are retrieved from the 3D-FUTURE dataset by matching semantic class k^\hat{k} and bounding box size s^\hat{s}. To eliminate interpenetrations and ensure realistic human-scene contact, object poses are post-processed via continuous geometric refinement:

    1. A unified signed distance field (SDF) volume of the scene is calculated.
    2. Contact vertices from all SMPL-X human body meshes are accumulated in 3D space.
    3. Object translations and rotations are jointly optimized using the contact loss Lcontact\mathcal{L}_{\text{contact}} and collision loss Lcollision\mathcal{L}_{\text{collision}} from MOVER:

    Lrefine=wcontactLcontact+wcollisionLcollision\mathcal{L}_{\text{refine}} = w_{\text{contact}} \mathcal{L}_{\text{contact}} + w_{\text{collision}} \mathcal{L}_{\text{collision}}

    where wcontact=105w_{\text{contact}} = 10^5 pulls object surfaces into alignment with contacting human body vertices and wcollision=103w_{\text{collision}} = 10^3 penalizes human-furniture interpenetrations.

  6. Knowl 6 — 3D-FRONT HUMAN Dataset Construction

    experimental setup

    To train human-aware 3D scene synthesis models, the 3D-FRONT HUMAN dataset extends synthetic room layouts from 3D-FRONT by automatically populating them with SMPL-X human bodies:

    • Room Categories and Scene Counts:
      • Bedrooms: 5,689 scenes across 21 object categories
      • Living rooms: 2,987 scenes across 24 object categories
      • Dining rooms: 2,549 scenes across 24 object categories
      • Libraries: 679 scenes across 25 object categories
    • Population Procedure:
      1. Contact humans: Static 3D human meshes from RenderPeople are assigned plausible interactions (touching, sitting, lying) with compatible furniture (e.g., lying on beds, sitting on chairs/sofas, touching nightstands/wardrobes).
      2. Free-space humans: A randomized number of static standing bodies and continuous walking motion clips from the AMASS dataset are placed with random initial positions and trajectories in open floor space. Humans intersecting with room objects are removed.
    • Data Splitting: For each room category, scenes are partitioned into 80% training, 10% validation, and 10% testing splits.
  7. Knowl 7 — Free-Space Interpenetration Evaluation Metric

    definition

    The interpenetration metric LinterL_{\text{inter}} measures spatial collision between generated 3D room objects and human movement trajectories (free space). It computes the proportion of the 2D ground-plane free-space mask E\mathcal{E} covered by the orthogonal 2D bounding projections of all MM generated objects {Oj}j=1M\{O_j\}_{j=1}^M:

    Linter=∑j=1M∑p∈OjE(p)∑p∈EE(p)L_{\text{inter}} = \frac{\sum_{j=1}^{M} \sum_{p \in O_j} \mathcal{E}(p)}{\sum_{p \in \mathcal{E}} \mathcal{E}(p)}

    where pp denotes pixel coordinates on the floor plan image, OjO_j is the set of pixels covered by object jj's ground projection, and E(p)=1\mathcal{E}(p) = 1 if pixel pp is traversed by walking or standing humans (00 otherwise). Lower values indicate superior preservation of human movement pathways.

  8. Knowl 8 — Quantitative Evaluation on 3D-FRONT HUMAN Benchmark

    data/table

    MIME was quantitatively evaluated against the ATISS baseline on the test split of 3D-FRONT HUMAN without refinement. Interaction consistency was measured using Free-Space Interpenetration (Linter↓L_{\text{inter}} \downarrow), 2D IoU (↑\uparrow), and 3D IoU (↑\uparrow) between generated objects and contact human bounding boxes. Scene realism and diversity were measured using Fréchet Inception Distance (FID ↓\downarrow computed at 2562256^2 resolution over 10 runs) and Category KL Divergence (↓\downarrow).

    Room Type Interpenetration (↓\downarrow) 2D IoU (↑\uparrow) 3D IoU (↑\uparrow) FID Score (↓\downarrow) Category KL Div. (↓\downarrow)
    ATISS Ours ATISS Ours ATISS Ours ATISS Ours ATISS Ours
    Bedroom 0.348 0.129 0.472 0.939 0.376 0.756 70.21 ±\pm 1.80 74.18 ±\pm 2.19 0.028 0.044
    Living 0.129 0.050 0.480 0.971 0.360 0.920 130.61 ±\pm 1.27 150.03 ±\pm 1.00 0.004 0.053
    Dining 0.121 0.047 0.163 0.959 0.122 0.769 45.99 ±\pm 0.90 76.75 ±\pm 1.45 0.004 0.037
    Library 0.139 0.106 0.351 0.725 0.390 0.570 93.16 ±\pm 2.59 118.34 ±\pm 2.94 0.066 0.093

    MIME achieves large improvements in 2D IoU, 3D IoU, and interpenetration reduction across all room types. Because human motion and contact explicitly constrain the allowable object arrangements, unconditioned ATISS achieves lower (better) FID and KL divergence scores due to higher unconstrained variation.

  9. Knowl 9 — Real-World Generalization on PROX-D Dataset

    empirical result

    To evaluate transfer to real captured human motion, MIME (trained on 3D-FRONT HUMAN without finetuning or mesh refinement) was evaluated against Pose2Room (P2R-Net) on the real-world PROX-D RGB-D capture dataset using 3D bounding box annotations. 3D object detection accuracy was evaluated using mean average precision at 3D IoU threshold 0.5 ([email protected]) over 10 sampled scenes for each of 5 motion sequences:

    Method 3D IoU ([email protected])
    P2R-Net (Pose2Room) w/o pretrain 5.36
    Ours (MIME) w/o pretrain 8.47

    MIME achieves higher 3D object detection accuracy on contacted objects (8.47 vs. 5.36 [email protected]) while additionally synthesizing the non-contact background room layout, which Pose2Room cannot generate.

  10. Knowl 10 — Limitations of MIME

    limitation

    The MIME framework has four primary stated limitations:

    1. Static Scene Assumption: MIME assumes scenes and objects are rigid and static, preventing the generation of dynamic interactions such as moving furniture, grasping small objects, or opening doors.
    2. Floor Plan Resolution: Floor plans are encoded at a coarse 64×6464 \times 64 grid representing a 6.2×6.2 m26.2 \times 6.2\text{ m}^2 room (approximately 10 cm10\text{ cm} per pixel), which limits fine boundary alignment.
    3. Heuristic Pruning Threshold: The autoregressive pruning of resolved contact humans relies on a fixed 2D IoU threshold (>0.5> 0.5) rather than a learned contact satisfaction model.
    4. Two-Stage Mesh Generation: The pipeline separates bounding box generation from CAD model retrieval and continuous optimization refinement, rather than performing end-to-end mesh synthesis.

Coverage note — No substantial contributed material was omitted from the extracted knowls.

References

  1. 1.Eduard Gabriel Bazavan, Andrei Zanfir, Mihai Zanfir, William T Freeman, Rahul Sukthankar, and Cristian Sminchisescu. HSPACE: Synthetic parametric humans animated in complex environments. arXiv, 2021. 3
  2. 2.Bharat Lal Bhatnagar, Xianghui Xie, Ilya Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. BEHAVE: Dataset and method for tracking human object interactions. In Computer Vision and Pattern Recognition (CVPR). IEEE, 2022. 3
  3. 3.Zhongang Cai, Mingyuan Zhang, Jiawei Ren, Chen Wei, Daxuan Ren, Zhengyu Lin, Haiyu Zhao, Lei Yang, and Ziwei Liu. Playing for 3d human recovery. arXiv, 2021. 3
  4. 4.Zhe Cao, Hang Gao, Karttikeya Mangalam, Qizhi Cai, Minh Vo, and Jitendra Malik. Long-term human motion prediction with scene context. In European Conference on Computer Vision (ECCV), 2020. 3
  5. 5.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3D: Learning from rgb-d data in indoor environments. In International Conference on 3D Vision (3DV), 2017. 3
  6. 6.Angel Chang, Will Monroe, Manolis Savva, Christopher Potts, and Christopher D Manning. Text to 3d scene generation with rich lexical grounding. arXiv, 2015. 2
  7. 7.Angel X Chang, Mihail Eric, Manolis Savva, and Christopher D Manning. Sceneseer: 3d scene design with natural language. arXiv, 2017. 2
  8. 8.CMU Graphics Lab. CMU Graphics Lab Motion Capture Database. http://mocap.cs.cmu.edu/, 2000. 3
  9. 9.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. ScanNet: Richly-annotated 3d reconstructions of indoor scenes. In Computer Vision and Pattern Recognition (CVPR), 2017. 3
  10. 10.Yudi Dai, YiTai Lin, XiPing Lin, Chenglu Wen, Lan Xu, Hongwei Yi, Siqi Shen, Yuexin Ma, and Cheng Wang. Sloper4d: A scene-aware dataset for global 4d human pose estimation in urban environments. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  11. 11.Jeevan Devaranjan, Amlan Kar, and Sanja Fidler. Meta-sim2: Learning to generate synthetic datasets. In European Conference on Computer Vision (ECCV), 2020. 2
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv, 2018. 5
  13. 13.Xinhan Di, Pengqian Yu, Hong Zhu, Lei Cai, Qiuyan Sheng, Changyu Sun, and Lingqiang Ran. Structural plan of indoor scenes with personalized preferences. In European Conference on Computer Vision (ECCV), pages 455–468. Springer, 2020. 2
  14. 14.Mingyu Ding, Yuqi Huo, Hongwei Yi, Zhe Wang, Jianping Shi, Zhiwu Lu, and Ping Luo. Learning depth-guided convolutions for monocular 3d object detection. In Computer Vision and Pattern Recognition (CVPR), pages 1000–1001, 2020. 3
  15. 15.Qi Fang, Kang Chen, Yinghui Fan, Qing Shuai, Jiefeng Li, and Weidong Zhang. Learning analytical posterior probability for human mesh recovery. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  16. 16.Matthew Fisher, Daniel Ritchie, Manolis Savva, Thomas Funkhouser, and Pat Hanrahan. Example-based synthesis of 3d object arrangements. Transactions on Graphics (TOG), 31(6):1–11, 2012. 2
  17. 17.Matthew Fisher, Manolis Savva, Yangyan Li, Pat Hanrahan, and Matthias Nießner. Activity-centric scene synthesis for functional 3d scene modeling. Transactions on Graphics (TOG), 34(6):1–13, 2015. 2
  18. 18.Huan Fu, Bowen Cai, Lin Gao, Ling-Xiao Zhang, Jiaming Wang, Cao Li, Qixun Zeng, Chengyue Sun, Rongfei Jia, Binqiang Zhao, et al. 3d-front: 3d furnished rooms with layouts and semantics. In International Conference on Computer Vision (ICCV), pages 10933–10942, 2021. 2, 3, 5, 8
  19. 19.Huan Fu, Rongfei Jia, Lin Gao, Mingming Gong, Binqiang Zhao, Steve Maybank, and Dacheng Tao. 3d-future: 3d furniture shape with texture. International Journal of Computer Vision (IJCV), pages 1–25, 2021. 2, 5
  20. 20.Vladimir Guzov, Aymen Mir, Torsten Sattler, and Gerard Pons-Moll. Human poseitioning system (HPS): 3d human pose estimation and self-localization in large scenes from body-mounted sensors. In Computer Vision and Pattern Recognition (CVPR), pages 4318–4329, 2021. 3
  21. 21.Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael Black. Stochastic scene-aware motion prediction. In International Conference on Computer Vision (ICCV), Oct. 2021. 3
  22. 22.Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J. Black. Resolving 3D human pose ambiguities with 3D scene constraints. In International Conference on Computer Vision (ICCV), pages 2282–2292, 2019. 2, 3, 6, 7, 8
  23. 23.Mohamed Hassan, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas, and Michael J. Black. Populating 3D scenes by learning human-scene interaction. In Computer Vision and Pattern Recognition (CVPR), June 2021. 2, 4
  24. 24.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. 4
  25. 25.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Conference on Neural Information Processing Systems (NeurIPS), volume 30, 2017. 6
  26. 26.Chun-Hao P. Huang, Hongwei Yi, Markus Höschle, Matvey Safroshkin, Tsvetelina Alexiadis, Senya Polikovsky, Daniel Scharstein, and Michael J. Black. Capturing and inferring dense full-body human-scene contact. In Computer Vision and Pattern Recognition (CVPR), 2022. 3
  27. 27.Siyuan Huang, Zan Wang, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu, Wei Liang, and Song-Chun Zhu. Diffusion-based generation, optimization, and planning in 3d scenes. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  28. 28.Yinghao Huang, Manuel Kaufmann, Emre Aksan, Michael J Black, Otmar Hilliges, and Gerard Pons-Moll. Deep inertial poser: Learning to reconstruct human pose from sparse inertial measurements in real time. Transactions on Graphics (TOG), 37(6):1–15, 2018. 3
  29. 29.Yangyi Huang, Hongwei Yi, Weiyang Liu, Haofan Wang, Boxi Wu, Wenxiao Wang, Binbin Lin, Debing Zhang, and Deng Cai. One-shot implicit animatable avatars with model-based priors. arXiv, 2022. 3
  30. 30.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 36(7):1325–1339, jul 2014. 3
  31. 31.Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, Timothy Scott Godisart, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh. Panoptic studio: A massively multiview system for social interaction capture. Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2017. 3
  32. 32.Amlan Kar, Aayush Prakash, Ming-Yu Liu, Eric Cameracci, Justin Yuan, Matt Rusiniak, David Acuna, Antonio Torralba, and Sanja Fidler. Meta-sim: Learning to generate synthetic datasets. In International Conference on Computer Vision (ICCV), 2019. 2
  33. 33.Mohammad Keshavarzi, Aakash Parikh, Xiyu Zhai, Melody Mao, Luisa Caldas, and Allen Y Yang. SceneGen: Generative contextual scene augmentation using scene graph priors. arXiv, 2020. 2
  34. 34.Jiefeng Li, Siyuan Bian, Qi Liu, Jiasheng Tang, Fan Wang, and Cewu Lu. NIKI: Neural inverse kinematics with invertible neural networks for 3d human pose and shape estimation. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  35. 35.Manyi Li, Akshay Gadi Patil, Kai Xu, Siddhartha Chaudhuri, Owais Khan, Ariel Shamir, Changhe Tu, Baoquan Chen, Daniel Cohen-Or, and Hao Zhang. Grains: Generative recursive autoencoders for indoor scenes. Transactions on Graphics (TOG), 38(2):1–16, 2019. 2
  36. 36.Tingting Liao, Xiaomei Zhang, Yuliang Xiu, Hongwei Yi, Xudong Liu, Guo-Jun Qi, Yong Zhang, Xuan Wang, Xiangyu Zhu, and Zhen Lei. High-Fidelity Clothed Avatar Reconstruction from a Single Image. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  37. 37.Andrew Luo, Zhoutong Zhang, Jiajun Wu, and Joshua B Tenenbaum. End-to-end optimization of scene layout. In Computer Vision and Pattern Recognition (CVPR), pages 3754–3763, 2020. 2
  38. 38.Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: Archive of motion capture as surface shapes. In International Conference on Computer Vision (ICCV), pages 5442–5451, Oct. 2019. 1, 2, 6
  39. 39.Dushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu, Srinath Sridhar, Gerard Pons-Moll, and Christian Theobalt. Single-shot multi-person 3D pose estimation from monocular rgb. In International Conference on 3D Vision (3DV). IEEE, sep 2018. 3
  40. 40.Aron Monszpart, Paul Guerrero, Duygu Ceylan, Ersin Yumer, and Niloy J. Mitra. iMapper: interaction-guided scene mapping from monocular videos. Transactions on Graphics (TOG), 38(4):92:1–92:15, 2019. 3
  41. 41.Pascal Müller, Peter Wonka, Simon Haegler, Andreas Ulmer, and Luc Van Gool. Procedural modeling of buildings. In Proceedings of the international conference on Computer graphics and interactive techniques (SIGGRAPH), pages 614–623, 2006. 2
  42. 42.Claudio Mura, Renato Pajarola, Konrad Schindler, and Niloy Mitra. Walk2map: Extracting floor plans from indoor walk trajectories. In Computer Graphics Forum (CGF), volume 40, pages 375–388, 2021. 3
  43. 43.Yinyu Nie, Angela Dai, Xiaoguang Han, and Matthias Nießner. Pose2room: understanding 3d scenes from human activities. In European Conference on Computer Vision, pages 425–443. Springer, 2022. 2, 3, 6, 7, 8
  44. 44.Wamiq Reyaz Para, Paul Guerrero, Niloy Mitra, and Peter Wonka. COFS: Controllable furniture layout synthesis. In SIGGRAPH Conference Papers, 2023. 2, 4
  45. 45.Yoav IH Parish and Pascal Müller. Procedural modeling of cities. In Proceedings of the international conference on Computer graphics and interactive techniques (SIGGRAPH), pages 301–308, 2001. 2
  46. 46.Despoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis, Andreas Geiger, and Sanja Fidler. ATISS: Autoregressive transformers for indoor scene synthesis. In Conference on Neural Information Processing Systems (NeurIPS), 2021. 2, 4, 5, 6, 7
  47. 47.Priyanka Patel, Chun-Hao P. Huang, Joachim Tesch, David T. Hoffmann, Shashank Tripathi, and Michael J. Black. AGORA: Avatars in geography optimized for regression analysis. In Computer Vision and Pattern Recognition (CVPR), June 2021. 2, 3, 5
  48. 48.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3d hands, face, and body from a single image. In Computer Vision and Pattern Recognition (CVPR), 2019. 5
  49. 49.Aayush Prakash, Shaad Boochoon, Mark Brophy, David Acuna, Eric Cameracci, Gavriel State, Omer Shapira, and Stan Birchfield. Structured domain randomization: Bridging the reality gap by context-aware synthetic data. In International Conference on Robotics and Automation. (ICRA), pages 7249–7255. IEEE, 2019. 2
  50. 50.Pulak Purkait, Christopher Zach, and Ian Reid. Sg-vae: Scene grammar variational autoencoder to generate new indoor scenes. In European Conference on Computer Vision (ECCV), pages 155–171. Springer, 2020. 2
  51. 51.Siyuan Qi, Yixin Zhu, Siyuan Huang, Chenfanfu Jiang, and Song-Chun Zhu. Human-centric indoor scene synthesis using stochastic grammar. In Computer Vision and Pattern Recognition (CVPR), pages 5899–5908, 2018. 3
  52. 52.René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In International Conference on Computer Vision (ICCV), pages 12179–12188, 2021. 8
  53. 53.Daniel Ritchie, Kai Wang, and Yu-an Lin. Fast and flexible indoor scene synthesis via deep convolutional generative models. In Computer Vision and Pattern Recognition (CVPR), pages 6182–6190, 2019. 2
  54. 54.Manolis Savva, Angel X. Chang, Pat Hanrahan, Matthew Fisher, and Matthias Nießner. PiGraphs: Learning Interaction Snapshots from Observations. ACM Transactions on Graphics (TOG), 35(4), 2016. 3
  55. 55.Ruizhi Shao, Zerong Zheng, Hanzhang Tu, Boning Liu, Hongwen Zhang, and Yebin Liu. Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  56. 56.Leonid Sigal, Alexandru O Balan, and Michael J Black. Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. International Journal of Computer Vision (IJCV), 87(1):4–27, 2010. 3
  57. 57.Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, Anton Clarkson, Mingfei Yan, Brian Budge, Yajie Yan, Xiaqing Pan, June Yon, Yuyang Zou, Kimberly Leon, Nigel Carter, Jesus Briales, Tyler Gillingham, Elias Mueggler, Luis Pesqueira, Manolis Savva, Dhruv Batra, Hauke M. Strasdat, Renzo De Nardi, Michael Goesele, Steven Lovegrove, and Richard Newcombe. The Replica dataset: A digital replica of indoor spaces. arXiv, 2019. 3
  58. 58.Jingxiang Sun, Xuan Wang, Lizhen Wang, Xiaoyu Li, Yong Zhang, Hongwen Zhang, and Yebin Liu. Next3d: Generative neural texture rasterization for 3d-aware head avatars. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  59. 59.Yu Sun, Qian Bao, Wu Liu, Tao Mei, and Michael J Black. TRACE: 5D Temporal Regression of Avatars with Dynamic Cameras in 3D Environments. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  60. 60.Jerry O Talton, Yu Lou, Steve Lesser, Jared Duke, Radomír Měch, and Vladlen Koltun. Metropolis procedural modeling. Transactions on Graphics (TOG), 30(2):1–14, 2011. 2
  61. 61.Yating Tian, Hongwen Zhang, Yebin Liu, and Limin Wang. Recovering 3d human mesh from monocular images: A survey. arXiv, 2022. 3
  62. 62.Shashank Tripathi, Lea Müller, Chun-Hao P. Huang, Taheri Omid, Michael J. Black, and Dimitrios Tzionas. 3D human pose estimation via intuitive physics. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  63. 63.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Conference on Neural Information Processing Systems (NeurIPS), 2017. 4, 5
  64. 64.Timo von Marcard, Roberto Henschel, Michael J. Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3D human pose in the wild using IMUs and a moving camera. In European Conference on Computer Vision (ECCV), pages 614–631, 2018. 3
  65. 65.Kai Wang, Yu-An Lin, Ben Weissmann, Manolis Savva, Angel X Chang, and Daniel Ritchie. Planit: Planning and instantiating indoor scenes with relation graph and spatial prior networks. Transactions on Graphics (TOG), 38(4):1–15, 2019. 2
  66. 66.Kai Wang, Manolis Savva, Angel X Chang, and Daniel Ritchie. Deep convolutional priors for indoor scene synthesis. Transactions on Graphics (TOG), 37(4):1–14, 2018. 2
  67. 67.Xinpeng Wang, Chandan Yeshwanth, and Matthias Nießner. Sceneformer: Indoor scene generation with transformers. In International Conference on 3D Vision (3DV), 2021. 2
  68. 68.Zhe Wang, Liyan Chen, Shaurya Rathore, Daeyun Shin, and Charless Fowlkes. Geometric pose affordance: 3D human pose with scene constraints. arXiv, 2019. 3
  69. 69.Qiuhong Anna Wei, Sijie Ding, Jeong Joon Park, Rahul Sajnani, Adrien Poulenard, Srinath Sridhar, and Leonidas Guibas. Lego-net: Learning regular rearrangements of objects in rooms. arXiv, 2023. 8
  70. 70.Xianghui Xie, Bharat Lal Bhatnagar, and Gerard Pons-Moll. Visibility aware human-object interaction tracking from single rgb camera. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  71. 71.Yuliang Xiu, Jinlong Yang, Xu Cao, Dimitrios Tzionas, and Michael J. Black. ECON: Explicit Clothed humans Optimized via Normal Integration. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  72. 72.Jianfeng Yan, Zizhuang Wei, Hongwei Yi, Mingyu Ding, Runze Zhang, Yisong Chen, Guoping Wang, and Yu-Wing Tai. Dense hybrid recurrent multi-view stereo net with dynamic consistency checking. In European Conference on Computer Vision (ECCV), pages 674–689. Springer, 2020. 3
  73. 73.Sifan Ye, Yixing Wang, Jiaman Li, Dennis Park, C Karen Liu, Huazhe Xu, and Jiajun Wu. Scene synthesis from human motion. In SIGGRAPH Asia Conference Papers, pages 1–9, 2022. 3
  74. 74.Hongwei Yi, Chun-Hao P. Huang, Dimitrios Tzionas, Muhammed Kocabas, Mohamed Hassan, Siyu Tang, Justus Thies, and Michael J. Black. Human-aware object placement for visual environment reconstruction. In Computer Vision and Pattern Recognition (CVPR), 2022. 3, 5, 7, 8
  75. 75.Hongwei Yi, Hualin Liang, Yifei Liu, Qiong Cao, Yandong Wen, Timo Bolkart, Dacheng Tao, and Michael J. Black. Generating holistic 3D human motion from speech. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  76. 76.Hongwei Yi, Shaoshuai Shi, Mingyu Ding, Jiankai Sun, Kui Xu, Hui Zhou, Zhe Wang, Sheng Li, and Guoping Wang. Segvoxelnet: Exploring semantic context and depth-aware features for 3d vehicle detection from point cloud. In International Conference on Robotics and Automation. (ICRA), pages 2274–2280. IEEE, 2020. 3
  77. 77.Hongwei Yi, Zizhuang Wei, Mingyu Ding, Runze Zhang, Yisong Chen, Guoping Wang, and Yu-Wing Tai. Pyramid multi-view stereo net with self-adaptive view aggregation. In European Conference on Computer Vision (ECCV), pages 766–782. Springer, 2020. 3
  78. 78.Zhixuan Yu, Jae Shin Yoon, In Kyu Lee, Prashanth Venkatesh, Jaesik Park, Jihun Yu, and Hyun Soo Park. HUMBI: A large multiview dataset of human body expressions. In Computer Vision and Pattern Recognition (CVPR), June 2020. 3
  79. 79.Hongwen Zhang, Siyou Lin, Ruizhi Shao, Yuxiang Zhang, Zerong Zheng, Han Huang, Yandong Guo, and Yebin Liu. Closet: Modeling clothed humans on continuous surface with explicit template decomposition. In Computer Vision and Pattern Recognition (CVPR), June 2023. 3
  80. 80.Song-Hai Zhang, Shao-Kui Zhang, Wei-Yu Xie, Cheng-Yang Luo, and Hong-Bo Fu. Fast 3d indoor scene synthesis with discrete and exact layout pattern extraction. arXiv, 2020. 2, 6
  81. 81.Zaiwei Zhang, Zhenpei Yang, Chongyang Ma, Linjie Luo, Alexander Huth, Etienne Vouga, and Qixing Huang. Deep generative modeling for scene synthesis via hybrid representations. Transactions on Graphics (TOG), 39(2):1–21, 2020.
  82. 82.Yang Zhou, Zachary While, and Evangelos Kalogerakis. Scenegraphnet: Neural message passing for 3d indoor scene augmentation. In International Conference on Computer Vision (ICCV), pages 7384–7392, 2019. 2

Citation

MLA
Yi, H., et al. “MIME: Human-Aware 3D Scene Generation”. arXiv, 2022, http://arxiv.org/abs/2212.04360v1.
APA
Yi, H., Huang, C.-H. P., Tripathi, S., Hering, L., Thies, J., & Black, M. J. (2022). MIME: Human-Aware 3D Scene Generation. arXiv. http://arxiv.org/abs/2212.04360v1
Chicago
Yi, H., C.-H. P. Huang, S. Tripathi, L. Hering, J. Thies, and M. J. Black. 2022. “MIME: Human-Aware 3D Scene Generation”. arXiv. http://arxiv.org/abs/2212.04360v1.
Harvard
Yi, H. et al. (2022) “MIME: Human-Aware 3D Scene Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.04360v1.
Vancouver
1. Yi H, Huang C-HP, Tripathi S, Hering L, Thies J, Black MJ (2022) MIME: Human-Aware 3D Scene Generation. arXiv

BibTeX

@article{yi2022mime,
  title = {MIME: Human-Aware 3D Scene Generation},
  author = {Yi, Hongwei and Huang, Chun-Hao P. and Tripathi, Shashank and Hering, Lea and Thies, Justus and Black, Michael J.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.04360v1},
  eprint = {2212.04360}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE