PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI

Yandan YangBaoxiong JiaPeiyuan ZhiSiyuan Huang

article2024CVPR158 citations

Proposes a physics-guided diffusion model that synthesizes functional 3D indoor environments with articulated objects by directly enforcing collision avoidance, room boundaries, and agent reachability constraints during generation.

Listen

Embodied artificial intelligence requires simulated 3D environments where virtual agents and robots can learn navigation and manipulation skills through physical interaction. However, existing automated scene synthesis techniques focus primarily on visual realism and perceptual quality. They rely heavily on datasets filled with static, non-interactive objects and frequently produce physically implausible arrangements, such as colliding furniture, blocked pathways, and objects protruding through walls. These flaws severely undermine the utility of synthetic environments for physics-based training and simulation.

The article aims to resolve this bottleneck by developing and evaluating PHYSCENE, a generative framework that creates realistic, physically plausible, and interactive 3D indoor scenes populated with articulated, interactable objects tailored for embodied AI agents.

The approach integrates a conditional diffusion model—a machine learning method that iteratively denoises random inputs into structured room layouts—with physics-based and interactivity guidance mechanisms during inference. The system models objects using labels, dimensions, orientations, and latent geometric shape features. These shape features allow the framework to retrieve matching articulated objects, such as openable cabinets and wardrobes, from external interactive asset repositories. To maintain physical plausibility, the method applies three core guidance functions: collision avoidance between object bounding boxes, room layout alignment to keep objects within designated floor plans, and reachability path planning to guarantee that simulated agents can navigate the room and access objects. The authors evaluated the system on thousands of indoor rooms from standard benchmarks, comparing it against leading autoregressive and diffusion-based baseline models.

Across extensive experiments, PHYSCENE established state-of-the-art results on standard visual quality metrics while significantly outperforming previous methods on physical plausibility. In unconditional generation across living room environments, the proposed model reduced object collision rates to 13.0%, compared to 18.3% for the leading diffusion baseline and 37.2% for the autoregressive baseline. Scene-level collision incidence dropped to 47.7%, substantially lower than the baselines' 57.0% and 87.0%. In floor-plan-conditioned tasks, the framework consistently achieved superior walkable area ratios and lower collision rates across bedrooms, living rooms, and dining rooms. Furthermore, when incorporating articulated objects manipulated to their fullest extension, the method maintained the lowest overall collision rates (76% of scenes containing collisions, versus 78% and 86% in baselines) while preserving high object reachability.

These findings demonstrate that automated scene synthesis can enforce physical commonsense rules alongside visual aesthetics. For organizations developing robotic systems and embodied AI, this approach reduces the manual labor and high costs of building interactive 3D simulation assets. It also mitigates the risk of agents learning flawed behaviors in unfeasible environments. An ablation analysis showed that individual guidance constraints naturally compete—such as collision avoidance pushing objects apart while boundary constraints push them inward—but balancing these functions successfully optimizes the entire layout.

For practical implementation, teams should adopt guided diffusion frameworks when generating large-scale synthetic training data for physical simulations, balancing guidance weights to fit specific room constraints. However, decision-makers should note certain limitations: the current implementation is restricted to major furniture categories within a limited set of room types and does not yet include small manipulable objects, such as items used in pick-and-place tasks. While confidence in the framework's physical layout optimization is high, expanding asset libraries to support fine-grained object manipulation remains a necessary next step.

arXiv: 2404.09465
Cover for PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI

Abstract

With recent developments in Embodied Artificial Intelligence (EAI) research, there has been a growing demand for high-quality, large-scale interactive scene generation. While prior methods in scene synthesis have prioritized the naturalness and realism of the generated scenes, the physical plausibility and interactivity of scenes have been largely left unexplored. To address this disparity, we introduce PHYSCENE, a novel method dedicated to generating interactive 3D scenes characterized by realistic layouts, articulated objects, and rich physical interactivity tailored for embodied agents. Based on a conditional diffusion model for capturing scene layouts, we devise novel physics- and interactivity-based guidance mechanisms that integrate constraints from object collision, room layout, and object reachability. Through extensive experiments, we demonstrate that PHYSCENE effectively leverages these guidance functions for physically interactable scene synthesis, outperforming existing state-of-the-art scene synthesis methods by a large margin. Our findings suggest that the scenes generated by PHYSCENE hold considerable potential for facilitating diverse skill acquisition among agents within interactive environments, thereby catalyzing further advancements in embodied AI research.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. PHYSCENE
  • 3.1. Object representation
  • 3.2. Conditional Diffusion for Layout Modeling
  • 3.3. Guidance for Physical Interactivity
  • 4. Experiment
  • 4.1. Unconditioned Scene Synthesis
  • 4.2. Floor-conditioned Scene Synthesis
  • 4.3. Scene Synthesis with Articulated Objects
  • 4.4. Ablation Study on Guidance
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Posterior Guidance Formulation for Scene Diffusion Models

    model/method

    To incorporate physical plausibility and room layout constraints into 3D scene generation, PHYSCENE formulates generative layout modeling as a conditional diffusion process guided by posterior constraint optimization.

    Let x0={o1,…,oN}\mathbf{x}_0 = \{o_1, \dots, o_N\} represent a scene layout consisting of NN objects, conditioned on a 2D floor plan F\mathcal{F}. The forward diffusion process adds Gaussian noise according to schedule α^t\hat{\alpha}_t:

    q(xt∣x0)=N(xt;α^tx0,(1−α^t)I)q(\mathbf{x}_t | \mathbf{x}_0) = \mathcal{N}(\mathbf{x}_t; \sqrt{\hat{\alpha}_t}\mathbf{x}_0, (1 - \hat{\alpha}_t)\mathbf{I})

    The conditional reverse transition parameterized by θ\theta is:

    pθ(xt−1∣xt,F)=N(xt−1;μθ(xt,t,F),Σθ(xt,t,F))p_\theta(\mathbf{x}_{t-1} | \mathbf{x}_t, \mathcal{F}) = \mathcal{N}(\mathbf{x}_{t-1}; \boldsymbol{\mu}_\theta(\mathbf{x}_t, t, \mathcal{F}), \boldsymbol{\Sigma}_\theta(\mathbf{x}_t, t, \mathcal{F}))

    Let O∈{0,1}O \in \{0, 1\} be an optimality indicator denoting whether a generated layout satisfies a scalar constraint function φ(x,F)\varphi(\mathbf{x}, \mathcal{F}). The conditioned posterior distribution is expressed as:

    p(x0∣F,O=1)∝pθ(x0∣F)⋅exp⁡(φ(x0,F))p(\mathbf{x}_0 | \mathcal{F}, O=1) \propto p_\theta(\mathbf{x}_0 | \mathcal{F}) \cdot \exp(\varphi(\mathbf{x}_0, \mathcal{F}))

    Using a first-order Taylor expansion around xt=μ=μθ(xt,t,F)\mathbf{x}_t = \boldsymbol{\mu} = \boldsymbol{\mu}_\theta(\mathbf{x}_t, t, \mathcal{F}), the gradient of the constraint log-likelihood is approximated by:

    g=∇xtφ(xt,F)∣xt=μ\mathbf{g} = \left. \nabla_{\mathbf{x}_t} \varphi(\mathbf{x}_t, \mathcal{F}) \right|_{\mathbf{x}_t = \boldsymbol{\mu}}

    The guided reverse transition is tilted as a perturbed Gaussian:

    pθ(xt−1∣xt,F,O=1)=N(xt−1;μ+λΣg,Σ)p_\theta(\mathbf{x}_{t-1} | \mathbf{x}_t, \mathcal{F}, O=1) = \mathcal{N}(\mathbf{x}_{t-1}; \boldsymbol{\mu} + \lambda \boldsymbol{\Sigma} \mathbf{g}, \boldsymbol{\Sigma})

    where Σ=Σθ(xt,t,F)\boldsymbol{\Sigma} = \boldsymbol{\Sigma}_\theta(\mathbf{x}_t, t, \mathcal{F}) and λ>0\lambda > 0 is a guidance scale parameter. During training, the guidance objective is optimized via:

    Lθ(x0∣F,O=1)=Et,ϵ,x0[∥ϵ−ϵθ(xt,t,F)−λΣg∥22]\mathcal{L}_\theta(\mathbf{x}_0 | \mathcal{F}, O=1) = \mathbb{E}_{t, \boldsymbol{\epsilon}, \mathbf{x}_0}\left[\|\boldsymbol{\epsilon} - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t, \mathcal{F}) - \lambda \boldsymbol{\Sigma} \mathbf{g}\|_2^2\right]

    Because intermediate noisy latents xt\mathbf{x}_t do not represent valid physical geometry, the constraint gradient g\mathbf{g} is evaluated on the one-step predicted clean scene layout x~0t=1α^t(xt−1−α^tϵθ(xt,t,F))\tilde{\mathbf{x}}_0^t = \frac{1}{\sqrt{\hat{\alpha}_t}}(\mathbf{x}_t - \sqrt{1 - \hat{\alpha}_t}\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t, \mathcal{F})), computing φ(x~0t,F)\varphi(\tilde{\mathbf{x}}_0^t, \mathcal{F}).

  2. Knowl 2 — Physical and Navigational Guidance Functions for Scene Synthesis

    model/method

    PHYSCENE defines three explicit differentiable guidance functions to ensure physical plausibility, layout compliance, and embodied agent navigability in synthesized 3D scenes.

    1. Object Collision Avoidance Guidance (φcoll\varphi_{\text{coll}}): To prevent inter-penetration between objects without expensive mesh-level collision checks, 3D oriented bounding boxes bi=[ti,ri,si]\mathbf{b}_i = [\mathbf{t}_i, \mathbf{r}_i, \mathbf{s}_i] (center ti∈R3\mathbf{t}_i \in \mathbb{R}^3, orientation ri∈R2\mathbf{r}_i \in \mathbb{R}^2, size si∈R3\mathbf{s}_i \in \mathbb{R}^3) are used to penalize pairwise volumetric overlap:

    φcoll(x)=−∑i=1N∑j≠iIoU3D(bi,bj)\varphi_{\text{coll}}(\mathbf{x}) = -\sum_{i=1}^N \sum_{j \neq i} \mathbf{IoU}_{3D}(\mathbf{b}_i, \mathbf{b}_j)

    where IoU3D\mathbf{IoU}_{3D} is the 3D bounding box Intersection over Union.

    1. Room-Layout Boundary Guidance (φlayout\varphi_{\text{layout}}): Given a 2D floor plan polygon F\mathcal{F}, a set of WW outer wall boundaries is represented as infinite-thickness 3D wall bounding boxes {bwwall}w=1W\{\mathbf{b}_w^{\text{wall}}\}_{w=1}^W. Objects protruding outside room boundaries are penalized via:

    φlayout(x∣F)=−∑i=1N∑j=1WIoU3D(bi,bjwall)\varphi_{\text{layout}}(\mathbf{x} | \mathcal{F}) = -\sum_{i=1}^{N} \sum_{j=1}^{W} \mathbf{IoU}_{3D}(\mathbf{b}_i, \mathbf{b}_j^{\text{wall}})

    1. Agent Reachability Guidance (φreach\varphi_{\text{reach}}): To prevent furniture arrangements from partitioning a room into disjoint inaccessible regions, the room is mapped to a 2D grid where Gaussian cost distributions are placed around object centers. An A∗A^* path planner finds the minimum-cost traversal path between the centers of the two largest disconnected walkable components. Along this path, LL agent footprint bounding boxes {blagent}l=1L\{\mathbf{b}_l^{\text{agent}}\}_{l=1}^L of agent dimensions are sampled, penalizing object overlap along the traversal route:

    φreach(x∣F)=−∑i=1N∑j=1LIoU3D(bi,bjagent)\varphi_{\text{reach}}(\mathbf{x} | \mathcal{F}) = -\sum_{i=1}^{N} \sum_{j=1}^{L} \mathbf{IoU}_{3D}(\mathbf{b}_i, \mathbf{b}_j^{\text{agent}})

    The total composite guidance function is linearly combined with balancing weights γ1,γ2,γ3\gamma_1, \gamma_2, \gamma_3:

    φ(x,F)=γ1φcoll(x)+γ2φlayout(x,F)+γ3φreach(x,F)\varphi(\mathbf{x}, \mathcal{F}) = \gamma_1 \varphi_{\text{coll}}(\mathbf{x}) + \gamma_2 \varphi_{\text{layout}}(\mathbf{x}, \mathcal{F}) + \gamma_3 \varphi_{\text{reach}}(\mathbf{x}, \mathcal{F})

  3. Knowl 3 — Object Representation and Latent Shape Matching for Cross-Asset Articulated Scene Generation

    model/method

    In PHYSCENE, an indoor scene x\mathbf{x} of NN objects is parameterized as x={o1,…,oN}\mathbf{x} = \{o_1, \dots, o_N\}. Each object oio_i is represented by a tuple:

    oi=[ci,si,ri,ti,fi]o_i = [\mathbf{c}_i, \mathbf{s}_i, \mathbf{r}_i, \mathbf{t}_i, \mathbf{f}_i]

    where ci∈RC\mathbf{c}_i \in \mathbb{R}^C is a one-hot semantic category vector over CC classes, si∈R3\mathbf{s}_i \in \mathbb{R}^3 represents bounding box dimensions (width, height, depth), ri=(cos⁡θi,sin⁡θi)∈R2\mathbf{r}_i = (\cos \theta_i, \sin \theta_i) \in \mathbb{R}^2 represents planar rotation (yaw), ti∈R3\mathbf{t}_i \in \mathbb{R}^3 denotes 3D translation/position, and fi∈R32\mathbf{f}_i \in \mathbb{R}^{32} is a latent geometric shape feature vector extracted from a 3D Variational Autoencoder (VAE) shape encoder.

    Because scene synthesis datasets (such as 3D-FRONT) contain only static CAD models whereas Embodied AI simulation requires interactable objects (such as those in GAPartNet), direct retrieval using bounding box dimensions alone fails across disparate asset libraries. PHYSCENE matches the predicted latent geometric shape embedding fi\mathbf{f}_i against a pre-encoded latent codebook of articulated objects from GAPartNet (specifically for storage and table categories like wardrobes and drawers) using nearest-neighbor retrieval. This enables generation of functional, articulated 3D environments from models trained on static room datasets.

  4. Knowl 4 — PHYSCENE Training and Guided Sampling Algorithm

    algorithm

    The training and inference procedures of PHYSCENE integrate physical guidance functions into the conditional diffusion network optimization and sampling steps.

    Input: Training scene layouts x0={o1,…,oN}x_0 = \{o_1, \dots, o_N\} with floor plans F\mathcal{F}, total diffusion steps TT, noise schedule α^t\hat{\alpha}_t, learning rate η\eta, guidance weights γ1,γ2,γ3\gamma_1, \gamma_2, \gamma_3, scaling factor λ\lambda
    Output: Trained conditional diffusion denoiser ϵθ\epsilon_\theta and generated scene layout x0x_0
    // Phase 1: Constraint-Guided Training
    repeat
        Sample x0∼p(x0∣F)x_0 \sim p(x_0 | \mathcal{F}), ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, \mathbf{I}), t∼Uniform({1,…,T})t \sim \text{Uniform}(\{1, \dots, T\})
        xt=α^tx0+1−α^tϵx_t = \sqrt{\hat{\alpha}_t} x_0 + \sqrt{1 - \hat{\alpha}_t} \epsilon
        Estimate clean layout x~0t\tilde{x}_0^t using model prediction pθ(⋅∣F)p_\theta(\cdot | \mathcal{F})
        Calculate constraint φ(x~0t,F)=γ1φcoll(x~0t)+γ2φlayout(x~0t,F)+γ3φreach(x~0t,F)\varphi(\tilde{x}_0^t, \mathcal{F}) = \gamma_1 \varphi_{\text{coll}}(\tilde{x}_0^t) + \gamma_2 \varphi_{\text{layout}}(\tilde{x}_0^t, \mathcal{F}) + \gamma_3 \varphi_{\text{reach}}(\tilde{x}_0^t, \mathcal{F})
        Compute gradient g=∇xtφ(x~0t,F)g = \nabla_{x_t} \varphi(\tilde{x}_0^t, \mathcal{F})
        Update network parameters: θ←θ−η∇θ∥ϵ−ϵθ(xt,t,F)−λΣθg∥22\theta \leftarrow \theta - \eta \nabla_\theta \|\epsilon - \epsilon_\theta(x_t, t, \mathcal{F}) - \lambda \boldsymbol{\Sigma}_\theta g\|_2^2
    until converged
    // Phase 2: Constraint-Guided Sampling / Generation
    Input: Gaussian noise layout xT∼N(0,I)x_T \sim \mathcal{N}(0, \mathbf{I}), floor plan F\mathcal{F}
    for t=Tt = T down to 1 do
        μ=μθ(xt,t,F)\mu = \mu_\theta(x_t, t, \mathcal{F}), Σ=Σθ(xt,t,F)\boldsymbol{\Sigma} = \boldsymbol{\Sigma}_\theta(x_t, t, \mathcal{F})
        Estimate clean layout x~0t\tilde{x}_0^t from xtx_t
        Evaluate composite guidance φ(x~0t,F)=γ1φcoll(x~0t)+γ2φlayout(x~0t,F)+γ3φreach(x~0t,F)\varphi(\tilde{x}_0^t, \mathcal{F}) = \gamma_1 \varphi_{\text{coll}}(\tilde{x}_0^t) + \gamma_2 \varphi_{\text{layout}}(\tilde{x}_0^t, \mathcal{F}) + \gamma_3 \varphi_{\text{reach}}(\tilde{x}_0^t, \mathcal{F})
        Compute gradient g=∇xtφ(x~0t,F)∣xt=μg = \left. \nabla_{x_t} \varphi(\tilde{x}_0^t, \mathcal{F}) \right|_{x_t = \mu}
        Sample xt−1∼N(xt−1;μ+λΣg,Σ)x_{t-1} \sim \mathcal{N}(x_{t-1}; \mu + \lambda \boldsymbol{\Sigma} g, \boldsymbol{\Sigma})
    end
    return x0x_0
  5. Knowl 5 — Kinematic Envelope Bounding Box Expansion for Articulated Objects

    model/method

    When placing articulated objects (e.g., cabinets with swing doors, drawers, or wardrobes) in interactive 3D environments, evaluating collision avoidance using resting/closed bounding boxes results in operational collisions during agent manipulation.

    To preserve full physical interactivity, PHYSCENE expands each articulated object's 3D bounding box bi\mathbf{b}_i to encompass its entire kinematic swept volume—the spatial domain occupied when all revolute and prismatic joints are articulated to their maximum mechanical limits. These expanded bounding boxes are used during collision guidance φcoll\varphi_{\text{coll}} and reachability guidance φreach\varphi_{\text{reach}}, ensuring generated layouts provide clearance for opening, closing, and agent interaction.

  6. Knowl 6 — Physical Plausibility and Interactivity Metrics for Scene Synthesis

    definition

    To evaluate scene synthesis for Embodied AI beyond visual perceptual metrics (FID, KID, SCA, CKL), PHYSCENE defines five physical plausibility and interactivity metrics computed across generated 3D environments:

    • Object Collision Rate (Colobj↓Col_{\text{obj}} \downarrow): The fraction of 3D object instances in a scene that collide with at least one other object. Evaluated on watertight re-meshed 3D CAD models.
    • Scene Collision Rate (Colscene↓Col_{\text{scene}} \downarrow): The proportion of generated scenes that exhibit at least one inter-object collision among all generated test scenes.
    • Floor Plan Violation Rate (Rout↓R_{\text{out}} \downarrow): The percentage of objects in a scene whose 3D bounding boxes intersect or extend beyond the pre-defined 2D room floor plan boundary.
    • Object Reachability Rate (Rreach↑R_{\text{reach}} \uparrow): The average fraction of objects in a scene reachable by an embodied agent navigating from a random starting point on the walkable floor plane.
    • Walkable Area Connectivity (Rwalkable↑R_{\text{walkable}} \uparrow): The ratio of the surface area of the largest connected 2D walkable region to the total walkable floor area within the room, indicating the absence of fragmented dead zones.
  7. Knowl 7 — Performance Comparison on Unconditional Scene Synthesis

    data/table

    PHYSCENE was evaluated against state-of-the-art baselines ATISS (autoregressive transformer) and DiffuScene (diffusion without layout constraints) on unconditional scene generation using 3D-FRONT rooms (evaluated on 1000 generated scenes per room category). Perceptual quality was assessed via Fréchet Inception Distance (FID), Scene Classification Accuracy (SCA), and Category KL divergence (CKL ×0.01\times 0.01). Physical plausibility was measured via object collision rate (ColobjCol_{\text{obj}}) and scene collision rate (ColsceneCol_{\text{scene}}).

    Method FID ↓\downarrow SCA ↓\downarrow CKL ↓\downarrow Colobj↓Col_{\text{obj}} \downarrow Colscene↓Col_{\text{scene}} \downarrow
    Bedroom
    ATISS 36.92 49.24 0.0036 0.255 0.50
    DiffuScene 28.63 51.33 0.0031 0.238 0.42
    PHYSCENE (Ours) 28.56 55.71 0.0030 0.187 0.35
    Living Room
    ATISS 55.76 53.33 0.0016 0.372 0.870
    DiffuScene 54.36 50.24 0.0010 0.183 0.570
    PHYSCENE (Ours) 40.67 56.20 0.0015 0.130 0.477
    Dining Room
    ATISS 41.89 58.20 0.0028 0.483 0.91
    DiffuScene 37.68 57.60 0.0031 0.253 0.63
    PHYSCENE (Ours) 37.88 58.74 0.0022 0.134 0.40

    PHYSCENE achieves lowest collision rates across all room types (ColobjCol_{\text{obj}} reduced by 21.4% to 47.0% compared to DiffuScene) while maintaining or improving visual perceptual quality (e.g., FID drops from 54.36 to 40.67 in Living Rooms).

  8. Knowl 8 — Performance Comparison on Floor-Conditioned Scene Synthesis

    data/table

    PHYSCENE was evaluated against ATISS and DiffuScene on floor-plan conditioned 3D scene generation across Bedroom, Living Room, and Dining Room splits of the 3D-FRONT dataset. Metrics evaluate perceptual realism (FID, KID ×0.001\times 0.001, SCA, CKL ×0.01\times 0.01) alongside physical plausibility and interactivity (ColobjCol_{\text{obj}}, ColsceneCol_{\text{scene}}, RoutR_{\text{out}}, RwalkableR_{\text{walkable}}, RreachR_{\text{reach}}).

    Room Type Method FID ↓\downarrow KID ↓\downarrow SCA ↓\downarrow CKL ↓\downarrow Colobj↓Col_{\text{obj}} \downarrow Colscene↓Col_{\text{scene}} \downarrow Rout↓R_{\text{out}} \downarrow Rwalkable↑R_{\text{walkable}} \uparrow Rreach↑R_{\text{reach}} \uparrow
    Bedroom ATISS 30.19 0.0010 49.14 0.0028 0.248 0.46 0.286 0.839 0.736
    DiffuScene 25.00 0.0004 51.78 0.0031 0.228 0.43 0.272 0.827 0.755
    PHYSCENE 25.52 0.0006 50.10 0.0025 0.187 0.36 0.245 0.865 0.762
    Living Room ATISS 45.66 0.0035 51.64 0.0016 0.316 0.85 0.136 0.814 0.791
    DiffuScene 38.69 0.0012 54.06 0.0017 0.198 0.69 0.238 0.790 0.756
    PHYSCENE 43.33 0.0031 53.50 0.0015 0.191 0.63 0.219 0.815 0.771
    Dining Room ATISS 41.66 0.0039 64.57 0.0040 0.591 0.96 0.132 0.874 0.848
    DiffuScene 38.31 0.0020 60.19 0.0013 0.160 0.55 0.244 0.787 0.847
    PHYSCENE 39.90 0.0026 60.00 0.0013 0.151 0.53 0.217 0.852 0.789

    Compared to DiffuScene, PHYSCENE consistently improves physical feasibility across all room categories by reducing collision rates (ColobjCol_{\text{obj}}, ColsceneCol_{\text{scene}}), boundary clipping (RoutR_{\text{out}}), and fragmented floor space (increasing RwalkableR_{\text{walkable}}). While ATISS achieves low RoutR_{\text{out}} by strictly clamping objects to interior boundaries, it incurs significantly worse collision rates (up to 0.591 ColobjCol_{\text{obj}} and 0.96 ColsceneCol_{\text{scene}} in dining rooms).

  9. Knowl 9 — Ablation of Guidance Functions and Articulated Asset Embedding

    data/table

    Ablation experiments conducted on Living Room layouts evaluate the individual impact of guidance terms (collision φcoll\varphi_{\text{coll}}, layout φlayout\varphi_{\text{layout}}, and reachability/interaction φreach\varphi_{\text{reach}}), alongside comparisons of articulated object integration against ATISS and DiffuScene.

    Articulated Embedding (Living Room) Colobj↓Col_{\text{obj}} \downarrow Colscene↓Col_{\text{scene}} \downarrow Rout↓R_{\text{out}} \downarrow Rreach↑R_{\text{reach}} \uparrow
    ATISS 0.360 0.86 0.154 0.758
    DiffuScene 0.262 0.78 0.237 0.702
    PHYSCENE (Ours) 0.251 0.76 0.229 0.755
    φcoll\varphi_{\text{coll}} φlayout\varphi_{\text{layout}} φreach\varphi_{\text{reach}} Colobj↓Col_{\text{obj}} \downarrow Rout↓R_{\text{out}} \downarrow Rwalkable↑R_{\text{walkable}} \uparrow Rreach↑R_{\text{reach}} \uparrow
    - - - 0.200 0.240 0.808 0.763
    ✓ - - 0.111 0.354 0.832 0.793
    - ✓ - 0.279 0.110 0.774 0.742
    - - ✓ 0.239 0.260 0.927 0.813
    ✓ ✓ ✓ 0.191 0.219 0.815 0.771

    The ablation reveals a structural trade-off: collision guidance alone pushes objects apart, increasing room boundary violations (RoutR_{\text{out}} rises from 0.240 to 0.354), whereas layout guidance compresses objects inward, raising inter-object collisions (ColobjCol_{\text{obj}} rises from 0.200 to 0.279). Combining all three guidance functions achieves balanced joint optimization.

  10. Knowl 10 — Limitations in Scene Diversity and Object Granularity

    limitation

    PHYSCENE is subject to two main limitations:

    1. Room Category Constraints: Training and evaluation are restricted to three residential indoor room types (bedrooms, living rooms, and dining rooms) available in structured 3D synthesis datasets, limiting direct zero-shot generalization to non-residential or complex multi-room environments.
    2. Absence of Small Manipulable Objects: The generation pipeline models major rigid and articulated furniture assets but does not synthesize small tabletop, countertop, or container-held objects (such as cups, books, or utensils). Consequently, the generated scenes cannot directly support fine-grained embodied manipulation tasks (e.g., tabletop pick-and-place) without an auxiliary secondary object placement stage.

Coverage note — Table 1 from the paper (ground truth 3D-FRONT constraint violation statistics) is omitted as a separate knowl because it serves as dataset motivation rather than a contributed model or experimental outcome. Supplementary extensions to specific human/agent poses (e.g., sitting and grasping) are omitted as they are not detailed in the main paper.

References

  1. 1.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. Do as i can, not as i say: Grounding language in robotic affordances. In Conference on Robot Learning (CoRL), 2022. 1
  2. 2.Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sunderhauf, Ian Reid, Stephen Gould, and ¨ Anton Van Den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 1
  3. 3.Arpit Bansal, Hong-Min Chu, Avi Schwarzschild, Soumyadip Sengupta, Micah Goldblum, Jonas Geiping, and Tom Goldstein. Universal guidance for diffusion models. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3
  4. 4.Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, et al. Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2021. 2
  5. 5.Mikołaj Binkowski, Danica J Sutherland, Michael Arbel, and ´ Arthur Gretton. Demystifying mmd gans. arXiv preprint arXiv:1801.01401, 2018. 2, 6
  6. 6.Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  7. 7.Angel Chang, Manolis Savva, and Christopher D Manning. Learning spatial knowledge for text to 3d scene generation. In Proceedings of the conference on Empirical Methods in Natural Language Processing (EMNLP), 2014. 1, 2
  8. 8.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 2
  9. 9.Matt Deitke, Winson Han, Alvaro Herrasti, Aniruddha Kembhavi, Eric Kolve, Roozbeh Mottaghi, Jordi Salvador, Dustin Schwenk, Eli VanderBilt, Matthew Wallingford, et al. Robothor: An open simulation-to-real embodied ai platform. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 1
  10. 10.Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Procthor: Large-scale embodied ai using procedural generation. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2022. 1, 2
  11. 11.Helisa Dhamo, Fabian Manhardt, Nassir Navab, and Federico Tombari. Graph-to-3d: End-to-end generation and manipulation of 3d scenes using scene graphs. In Proceedings of International Conference on Computer Vision (ICCV), 2021. 1, 2
  12. 12.Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. In Proceedings of International Conference on Machine Learning (ICML), 2023. 1
  13. 13.Matthew Fisher, Manolis Savva, Yangyan Li, Pat Hanrahan, and Matthias Nießner. Activity-centric scene synthesis for functional 3d scene modeling. ACM Transactions on Graphics (TOG), 34(6):1–13, 2015. 2
  14. 14.Huan Fu, Bowen Cai, Lin Gao, Ling-Xiao Zhang, Jiaming Wang, Cao Li, Qixun Zeng, Chengyue Sun, Rongfei Jia, Binqiang Zhao, et al. 3d-front: 3d furnished rooms with layouts and semantics. In Proceedings of International Conference on Computer Vision (ICCV), 2021. 1, 2, 3, 5
  15. 15.Huan Fu, Rongfei Jia, Lin Gao, Mingming Gong, Binqiang Zhao, Steve Maybank, and Dacheng Tao. 3d-future: 3d furniture shape with texture. International Journal of Computer Vision (IJCV), 129:3313–3337, 2021. 2, 3, 5
  16. 16.Qiang Fu, Xiaowu Chen, Xiaotian Wang, Sijia Wen, Bin Zhou, and Hongbo Fu. Adaptive synthesis of indoor scenes via activity-associated object relation graphs. ACM Transactions on Graphics (TOG), 36(6):1–13, 2017. 1, 2
  17. 17.Haoran Geng, Helin Xu, Chengyang Zhao, Chao Xu, Li Yi, Siyuan Huang, and He Wang. Gapartnet: Cross-category domain-generalizable object perception and manipulation via generalizable and actionable parts. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3, 5
  18. 18.Ran Gong, Jiangyong Huang, Yizhou Zhao, Haoran Geng, Xiaofeng Gao, Qingyang Wu, Wensi Ai, Ziheng Zhou, Demetri Terzopoulos, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. Arnold: A benchmark for language-grounded task learning with continuous states in realistic 3d scenes. In Proceedings of International Conference on Computer Vision (ICCV), 2023. 1
  19. 19.Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, et al. Maniskill2: A unified benchmark for generalizable manipulation skills. arXiv preprint arXiv:2302.04659, 2023. 2
  20. 20.Peter Hart, Nils Nilsson, and Bertram Raphael. A formal basis for the heuristic determination of minimum cost paths. IEEE Transactions on Systems Science and Cybernetics, 4 (2):100–107, 1968. 5
  21. 21.Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J Black. Resolving 3d human pose ambiguities with 3d scene constraints. In Proceedings of International Conference on Computer Vision (ICCV), 2019. 3
  22. 22.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2017. 2, 6
  23. 23.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 2, 3
  24. 24.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2020. 3, 4
  25. 25.Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An embodied generalist agent in 3d world. arXiv preprint arXiv:2311.12871, 2023. 1
  26. 26.Siyuan Huang, Zan Wang, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu, Wei Liang, and Song-Chun Zhu. Diffusion-based generation, optimization, and planning in 3d scenes. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3, 4
  27. 27.Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al. Inner monologue: Embodied reasoning through planning with language models. In Conference on Robot Learning (CoRL), 2022. 1
  28. 28.Naoto Inoue, Kotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, and Kota Yamaguchi. Layoutdm: Discrete diffusion model for controllable layout generation. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  29. 29.Baoxiong Jia, Yixin Chen, Huangyue Yu, Yan Wang, Xuesong Niu, Tengyu Liu, Qing Li, and Siyuan Huang. Sceneverse: Scaling 3d vision-language learning for grounded scene understanding. arXiv preprint arXiv:2401.09340, 2024. 2
  30. 30.Chenfanfu Jiang, Siyuan Qi, Yixin Zhu, Siyuan Huang, Jenny Lin, Lap-Fai Yu, Demetri Terzopoulos, and Song-Chun Zhu. Configurable 3d scene synthesis and 2d image rendering with per-pixel ground truth using stochastic grammars. International Journal of Computer Vision (IJCV), pages 920–941, 2018. 1
  31. 31.Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan. Vima: General robot manipulation with multimodal prompts. arXiv, 2022. 1
  32. 32.Mukul Khanna, Yongsen Mao, Hanxiao Jiang, Sanjay Haresh, Brennan Schacklett, Dhruv Batra, Alexander Clegg, Eric Undersander, Angel X Chang, and Manolis Savva. Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation. arXiv preprint arXiv:2306.11290, 2023. 2
  33. 33.Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, et al. Ai2-thor: An interactive 3d environment for visual ai. arXiv preprint arXiv:1712.05474, 2017. 1
  34. 34.Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee. Beyond the nav-graph: Vision-and-language navigation in continuous environments. In Proceedings of European Conference on Computer Vision (ECCV), 2020. 1
  35. 35.Chengshu Li, Fei Xia, Roberto Mart´ın-Mart´ın, Michael Lingelbach, Sanjana Srivastava, Bokui Shen, Kent Vainio, Cem Gokmen, Gokul Dharan, Tanish Jain, et al. igibson 2.0: Object-centric simulation for robot learning of everyday household tasks. arXiv preprint arXiv:2108.03272, 2021. 1
  36. 36.Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Mart´ın-Mart´ın, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, et al. Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation. In Conference on Robot Learning, 2023. 2
  37. 37.Xingyu Lin, Yufei Wang, Jake Olkin, and David Held. Softgym: Benchmarking deep reinforcement learning for deformable object manipulation. In Conference on Robot Learning, 2021. 2
  38. 38.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models. arXiv preprint arXiv:2211.01095, 2022. 2, 3
  39. 39.Mayank Mittal, Calvin Yu, Qinxi Yu, Jingzhou Liu, Nikita Rudin, David Hoeller, Jia Lin Yuan, Ritvik Singh, Yunrong Guo, Hammad Mazhar, et al. Orbit: A unified simulation framework for interactive robot learning environments. IEEE Robotics and Automation Letters, 2023. 2
  40. 40.Tongzhou Mu, Zhan Ling, Fanbo Xiang, Derek Yang, Xuanlin Li, Stone Tao, Zhiao Huang, Zhiwei Jia, and Hao Su. Maniskill: Generalizable manipulation skill benchmark with large-scale demonstrations. arXiv preprint arXiv:2107.14483, 2021. 2
  41. 41.Yinyu Nie, Angela Dai, Xiaoguang Han, and Matthias Nießner. Learning 3d scene priors with 2d supervision. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
  42. 42.Despoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis, Andreas Geiger, and Sanja Fidler. Atiss: Autoregressive transformers for indoor scene synthesis. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2021. 2, 5, 6
  43. 43.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 3
  44. 44.Pulak Purkait, Christopher Zach, and Ian Reid. Sg-vae: Scene grammar variational autoencoder to generate new indoor scenes. In Proceedings of European Conference on Computer Vision (ECCV), 2020. 2
  45. 45.Siyuan Qi, Yixin Zhu, Siyuan Huang, Chenfanfu Jiang, and Song-Chun Zhu. Human-centric indoor scene synthesis using stochastic grammar. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 1, 2, 3
  46. 46.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1 (2):3, 2022. 3
  47. 47.Ludan Ruan, Yiyang Ma, Huan Yang, Huiguo He, Bei Liu, Jianlong Fu, Nicholas Jing Yuan, Qin Jin, and Baining Guo. Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  48. 48.Mohit Shridhar, Lucas Manuelli, and Dieter Fox. Cliport: What and where pathways for robotic manipulation. In Conference on Robot Learning, 2022. 1
  49. 49.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 3
  50. 50.Andrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Singh Chaplot, Oleksandr Maksymets, et al. Habitat 2.0: Training home assistants to rearrange their habitat. Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 34:251–266, 2021. 1
  51. 51.Jiapeng Tang, Yinyu Nie, Lev Markhasin, Angela Dai, Justus Thies, and Matthias Nießner. Diffuscene: Denoising diffusion models for gerative indoor scene synthesis. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2, 3, 5, 6
  52. 52.Jiashun Wang, Huazhe Xu, Jingwei Xu, Sifei Liu, and Xiaolong Wang. Synthesizing long-term 3d human motion and interaction in 3d scenes. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
  53. 53.Weiqi Wang, Zihang Zhao, Ziyuan Jiao, Yixin Zhu, Song-Chun Zhu, and Hangxin Liu. Rearrange indoor scenes for human-robot co-activity. In Proceedings of International Conference on Robotics and Automation (ICRA), 2023. 3
  54. 54.Xinpeng Wang, Chandan Yeshwanth, and Matthias Nießner. Sceneformer: Indoor scene generation with transformers. In Proceedings of International Conference on 3D Vision (3DV), 2021. 1, 2
  55. 55.Zan Wang, Yixin Chen, Baoxiong Jia, Puhao Li, Jinlu Zhang, Jingze Zhang, Tengyu Liu, Yixin Zhu, Wei Liang, and Siyuan Huang. Move as you say, interact as you can: Language-guided human motion generation with scene affordance. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 3
  56. 56.Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of International Conference on Computer Vision (ICCV), 2023. 3
  57. 57.Fanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia, Hao Zhu, Fangchen Liu, Minghua Liu, Hanxiao Jiang, Yifu Yuan, He Wang, et al. Sapien: A simulated part-based interactive environment. In Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 1
  58. 58.Kun Xu, Kang Chen, Hongbo Fu, Wei-Lun Sun, and Shi-Min Hu. Sketch2scene: Sketch-based co-retrieval and co-placement of 3d models. ACM Transactions on Graphics (TOG), 32(4):1–15, 2013. 1, 2
  59. 59.Ming-Jia Yang, Yu-Xiao Guo, Bin Zhou, and Xin Tong. Indoor scene generation from a collection of semantic-segmented depth images. In Proceedings of International Conference on Computer Vision (ICCV), 2021. 2
  60. 60.Peiyu Yu, Sirui Xie, Xiaojian Ma, Baoxiong Jia, Bo Pang, Ruiqi Gao, Yixin Zhu, Song-Chun Zhu, and Ying Nian Wu. Latent diffusion energy-based model for interpretable text modeling. In Proceedings of International Conference on Machine Learning (ICML), 2022. 3
  61. 61.Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. Physdiff: Physics-guided human motion diffusion model. In Proceedings of International Conference on Computer Vision (ICCV), 2023. 3
  62. 62.Guangyao Zhai, Evin Pınar Ornek, Shun-Cheng Wu, Yan ¨ Di, Federico Tombari, Nassir Navab, and Benjamin Busam. Commonscenes: Generating commonsense 3d indoor scenes with scene graphs. Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2024. 1, 2
  63. 63.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of International Conference on Computer Vision (ICCV), 2023. 3
  64. 64.Zaiwei Zhang, Zhenpei Yang, Chongyang Ma, Linjie Luo, Alexander Huth, Etienne Vouga, and Qixing Huang. Deep generative modeling for scene synthesis via hybrid representations. ACM Transactions on Graphics (TOG), 39(2):1–21, 2020. 2
  65. 65.Kaizhi Zheng, Xiaotong Chen, Odest Chadwicke Jenkins, and Xin Wang. Vlmbench: A compositional benchmark for vision-and-language manipulation. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2022. 2
  66. 66.Dingfu Zhou, Jin Fang, Xibin Song, Chenye Guan, Junbo Yin, Yuchao Dai, and Ruigang Yang. Iou loss for 2d/3d object detection. In Proceedings of International Conference on 3D Vision (3DV), 2019. 5
  67. 67.Yang Zhou, Zachary While, and Evangelos Kalogerakis. Scenegraphnet: Neural message passing for 3d indoor scene augmentation. In Proceedings of International Conference on Computer Vision (ICCV), 2019. 1, 2

Citation

MLA
Yang, Y., et al. “PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI”. arXiv, 2024, http://arxiv.org/abs/2404.09465v2.
APA
Yang, Y., Jia, B., Zhi, P., & Huang, S. (2024). PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI. arXiv. http://arxiv.org/abs/2404.09465v2
Chicago
Yang, Y., B. Jia, P. Zhi, and S. Huang. 2024. “PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI”. arXiv. http://arxiv.org/abs/2404.09465v2.
Harvard
Yang, Y. et al. (2024) “PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.09465v2.
Vancouver
1. Yang Y, Jia B, Zhi P, Huang S (2024) PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI. arXiv

BibTeX

@article{yang2024physcene,
  title = {PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI},
  author = {Yang, Yandan and Jia, Baoxiong and Zhi, Peiyuan and Huang, Siyuan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.09465v2},
  eprint = {2404.09465}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE