PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
Yandan YangBaoxiong JiaPeiyuan ZhiSiyuan Huang
Proposes a physics-guided diffusion model that synthesizes functional 3D indoor environments with articulated objects by directly enforcing collision avoidance, room boundaries, and agent reachability constraints during generation.
Embodied artificial intelligence requires simulated 3D environments where virtual agents and robots can learn navigation and manipulation skills through physical interaction. However, existing automated scene synthesis techniques focus primarily on visual realism and perceptual quality. They rely heavily on datasets filled with static, non-interactive objects and frequently produce physically implausible arrangements, such as colliding furniture, blocked pathways, and objects protruding through walls. These flaws severely undermine the utility of synthetic environments for physics-based training and simulation.
The article aims to resolve this bottleneck by developing and evaluating PHYSCENE, a generative framework that creates realistic, physically plausible, and interactive 3D indoor scenes populated with articulated, interactable objects tailored for embodied AI agents.
The approach integrates a conditional diffusion model—a machine learning method that iteratively denoises random inputs into structured room layouts—with physics-based and interactivity guidance mechanisms during inference. The system models objects using labels, dimensions, orientations, and latent geometric shape features. These shape features allow the framework to retrieve matching articulated objects, such as openable cabinets and wardrobes, from external interactive asset repositories. To maintain physical plausibility, the method applies three core guidance functions: collision avoidance between object bounding boxes, room layout alignment to keep objects within designated floor plans, and reachability path planning to guarantee that simulated agents can navigate the room and access objects. The authors evaluated the system on thousands of indoor rooms from standard benchmarks, comparing it against leading autoregressive and diffusion-based baseline models.
Across extensive experiments, PHYSCENE established state-of-the-art results on standard visual quality metrics while significantly outperforming previous methods on physical plausibility. In unconditional generation across living room environments, the proposed model reduced object collision rates to 13.0%, compared to 18.3% for the leading diffusion baseline and 37.2% for the autoregressive baseline. Scene-level collision incidence dropped to 47.7%, substantially lower than the baselines' 57.0% and 87.0%. In floor-plan-conditioned tasks, the framework consistently achieved superior walkable area ratios and lower collision rates across bedrooms, living rooms, and dining rooms. Furthermore, when incorporating articulated objects manipulated to their fullest extension, the method maintained the lowest overall collision rates (76% of scenes containing collisions, versus 78% and 86% in baselines) while preserving high object reachability.
These findings demonstrate that automated scene synthesis can enforce physical commonsense rules alongside visual aesthetics. For organizations developing robotic systems and embodied AI, this approach reduces the manual labor and high costs of building interactive 3D simulation assets. It also mitigates the risk of agents learning flawed behaviors in unfeasible environments. An ablation analysis showed that individual guidance constraints naturally compete—such as collision avoidance pushing objects apart while boundary constraints push them inward—but balancing these functions successfully optimizes the entire layout.
For practical implementation, teams should adopt guided diffusion frameworks when generating large-scale synthetic training data for physical simulations, balancing guidance weights to fit specific room constraints. However, decision-makers should note certain limitations: the current implementation is restricted to major furniture categories within a limited set of room types and does not yet include small manipulable objects, such as items used in pick-and-place tasks. While confidence in the framework's physical layout optimization is high, expanding asset libraries to support fine-grained object manipulation remains a necessary next step.
- Paper: MIME: Human-Aware 3D Scene Generation, Hongwei Yi et al. (2023). Provides foundational principles on human-centric indoor scene synthesis and spatial collision avoidance that motivate PhyScene's interactivity-guided layout generation.
- Paper: AI2-THOR: An Interactive 3D Environment for Visual AI, Eric Kolve et al. (2017). Introduces the standard benchmark environment and interaction paradigms for embodied AI simulation that PhyScene aims to automatically populate with physically interactable assets.
- Paper: Habitat: A Platform for Embodied AI Research, Manolis Savva et al. (2019). Establishes the embodied AI simulation platform requirements for navigable indoor scenes that PhyScene directly targets for synthetic environment generation.
- Paper: DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis, Yinghao Xu et al. (2023). Demonstrates object-centric 3D layout disentanglement and bounding-box spatial priors, serving as a core conceptual precursor to guided generative scene synthesis.
- Paper: Diffusion-based Molecule Generation with Informative Prior Bridges, Lemeng Wu et al. (2022). Explores the integration of physical potential priors directly into diffusion trajectories, underpinning the physics-guided diffusion mechanisms used in PhyScene.
- Paper: The Replica Dataset: A Digital Replica of Indoor Spaces, Julian Straub et al. (2019). Establishes high-fidelity indoor 3D reconstructions and physical navigation standards that highlight the need for scalable synthetic alternatives like PhyScene.
- Paper: Unified Human-Scene Interaction via Prompted Chain-of-Contacts, Zeqi Xiao et al. (2024). Applies physically plausible human-scene interaction policies in synthesized 3D indoor environments populated by articulated objects.
- Paper: Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene Affordance, Zan Wang et al. (2024). Builds upon physical layout and reachability affordances to generate language-guided full-body human motion across 3D indoor spaces.
- Paper: SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World, Kiana Ehsani et al. (2024). Scales navigation and manipulation training across procedurally generated, interactable simulated homes using imitation learning.
- Paper: RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics, Chan Hee Song et al. (2025). Evaluates spatial reasoning, free-space identification, and placement compatibility for robotics downstream of structured 3D scene generation.
- Paper: CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control, Guy Tevet et al. (2025). Extends diffusion-driven planning into real-time closed-loop physics simulation and multi-task character control.
- Paper: WorldSimBench: Towards Video Generation Models as World Simulators, Yiran Qin et al. (2025). Provides a comprehensive benchmark evaluating physical plausibility and manipulative action generation in simulated 3D world models.
