MIME: Human-Aware 3D Scene Generation
Hongwei YiChun-Hao P. HuangShashank TripathiLea HeringJustus ThiesMichael J. Black
Proposes an auto-regressive transformer model that synthesizes plausible indoor 3D scenes by utilizing human motion sequences to infer room layouts, object placements, and free-space constraints.
Generating realistic 3D indoor environments populated by moving humans is essential for video games, architectural planning, and artificial intelligence training data generation, but traditional scene modeling remains expensive and labor-intensive. While previous research has focused on synthesizing human movement within pre-existing 3D spaces, relatively little work explores the reverse challenge of inferring a complete environment from human body movements alone. The article addresses this gap by developing MIME (Mining Interaction and Movement to infer 3D Environments), a computational framework that treats human body motion as an active scanner to predict full, plausible 3D room layouts.
The article's main objective is to design and evaluate an auto-regressive generative framework that uses 3D human motion and floor plans to synthesize complete indoor furniture layouts that accurately support human interactions and respect open space constraints.
To achieve this, the authors created a large synthetic training dataset called 3D-FRONT HUMAN by populating existing 3D room floor plans with moving and interacting virtual humans across four room categories: bedrooms, living rooms, dining rooms, and libraries. The methodology divides human motion into two components: free-space movement, which maps out walkable floor areas where objects cannot exist, and contact interactions (such as sitting, lying, or touching), which signal the presence and type of furniture. A transformer-based neural network processes these motion inputs alongside an empty floor plan to generate furniture bounding boxes sequentially. A post-processing refinement step then retrieves matching 3D furniture models from a catalog and optimizes their placement using geometric contact and collision rules.
The experimental findings show that MIME outperforms existing baseline models in physical plausibility and interaction accuracy. First, MIME drastically cut physical collisions between generated furniture and walkable space compared to human-unaware baseline models, achieving less than half the room interpenetration rates across living rooms (0.050 vs. 0.129), dining rooms (0.047 vs. 0.121), and bedrooms (0.129 vs. 0.348). Second, MIME greatly improved 3D spatial alignment between humans and interacting furniture, achieving contact intersection-over-union scores of 0.756 to 0.920 in common rooms compared to baseline scores ranging from 0.122 to 0.376. Third, when tested on a real-world motion-capture dataset without fine-tuning, MIME achieved an 8.47 object detection accuracy score, substantially higher than the 5.36 achieved by prior motion-conditioned reconstruction methods, while uniquely generating full room layouts rather than isolated objects in contact with the human.
These results demonstrate that human motion provides strong geometric constraints that allow automated systems to reconstruct complete, functional indoor scenes at scale. By enabling the conversion of archival motion capture data into realistic 3D environments, this approach can lower the cost and development timelines required to create synthetic training data for computer vision, architectural design, and interactive virtual reality.
The authors recommend integrating motion-aware generative methods into pipelines for large-scale synthetic data generation and digital room layout design. For future technical deployments, engineering teams should implement higher-resolution floor plan encodings and develop end-to-end models that jointly predict room boundaries and 3D shapes alongside furniture layouts.
A primary limitation of the study is its reliance on static scenes and coarse floor plan grid resolutions (where one grid unit represents approximately 10 centimeters), which can occasionally induce minor spatial collisions. Additionally, the system currently assumes all objects remain stationary, leaving dynamic interactions—such as opening doors or moving handheld objects—for future work. Despite these limitations, the quantitative improvements on both synthetic and real-world datasets support high confidence in MIME's ability to generate plausible 3D layouts from human motion.
- Paper: AMASS: Archive of Motion Capture As Surface Shapes, Naureen Mahmood et al. (2019). Provides the foundational large-scale parametric 3D human motion representation (AMASS/SMPL) used to represent human movement and body geometry in scene-interaction pipelines.
- Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). Introduces key transformer-based motion modeling and physical contact regularizers that underpin modern generative representations of 3D human movement.
- Paper: End-to-End Recovery of Human Shape and Pose, Angjoo Kanazawa et al. (2017). Establishes standard methods for 3D human mesh recovery and pose parameter estimation essential for extracting 3D body motion as spatial signals.
- Paper: Matterport3D: Learning from RGB-D Data in Indoor Environments, Angel Chang et al. (2017). Supplies fundamental datasets and geometric priors for indoor room structures and 3D bounding-box spatial layouts.
- Paper: 3D Semantic Parsing of Large-Scale Indoor Spaces, Iro Armeni et al. (2016). Pioneers hierarchical spatial parsing and floor-plan geometry analysis for indoor environments.
- Paper: Unified Human-Scene Interaction via Prompted Chain-of-Contacts, Zeqi Xiao et al. (2024). Extends human-scene contact modeling by using structured chain-of-contact sequences to drive dynamic, physics-based multi-step humanoid interactions in 3D scenes.
- Paper: Seamless Human Motion Composition with Blended Positional Encodings, Germán Barquero et al. (2024). Advances the synthesis of long-duration, multi-action sequential human motion that can serve as richer trajectory inputs for motion-conditioned scene reasoning.
- Paper: VOID: Video Object and Interaction Deletion, Saman Motamed et al. (2026). Applies human-object interaction reasoning to counterfactual video generation and dynamic causal simulation.
