NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions
Juze ZhangHaimin LuoHongdi YangXinru XuQianyang WuYe ShiJingyi YuLan XuJingya Wang
Introduces a large-scale multi-view dataset alongside a layer-wise neural radiance pipeline to accurately decouple, track, and photorealistically render dynamic human-object interactions under severe occlusions.
Accurately capturing, modeling, and rendering complex human-object interactions in three dimensions is vital for advancing fields such as digital entertainment, vocational training, tele-medicine, and sports analytics. However, existing visual capture methods struggle significantly with mutual occlusions, complex motions, and texture ambiguities when humans interact with physical items. Progress in this domain has been largely bottlenecked by the absence of dense-view appearance datasets and processing methods capable of cleanly separating interacting subjects from objects.
The main objective of the article is to establish a comprehensive data and modeling framework that captures dense multi-view human-object interactions and evaluates a specialized neural pipeline for motion tracking, geometry reconstruction, and photorealistic novel-view rendering.
To accomplish this, the authors constructed a dense camera capture dome comprising 76 high-resolution cinema cameras synchronized with 16 optical motion tracking cameras. Using this setup, they collected the HODome dataset, encompassing 274 interaction sequences across 10 diverse subjects and 23 physical objects, totaling approximately 71 million video frames. To process this massive data stream, the authors developed NeuralDome, a layer-wise neural modeling pipeline. The pipeline jointly tracks human skeletal motion alongside rigid object poses, decomposes scenes into dynamic human and static rigid neural radiance fields, applies specialized geometry and contact regularizers, and utilizes image blending to produce decoupled, high-fidelity digital assets.
The findings demonstrate substantial performance advantages over existing approaches. In novel-view appearance synthesis, NeuralDome achieved an average peak signal-to-noise ratio of 31.93 dB, notably outperforming alternative neural rendering baselines which achieved 22.67 dB and 24.99 dB. When benchmarking 3D geometry reconstruction, training an existing baseline model on this data reduced surface-to-point reconstruction error by roughly 62% for single-view inputs and by over 90% for multi-view inputs compared to the pre-trained baseline. Furthermore, the decoupling strategy successfully separated human bodies from contacted objects without the severe visual artifacts and blending errors typical of prior methods.
These results show that layer-wise neural decomposition, combined with high-density visual supervision, resolves longstanding occlusion and separation challenges in visual computing. For decision-makers, this framework lowers the technical risk and labor associated with digital asset creation, opening practical pathways for training high-performing computer vision models using fewer cameras during downstream deployment.
Looking ahead, development teams and researchers should leverage the publicly released dataset and processing tools to train generalizable interaction models and develop downstream sparse-view rendering applications. Further engineering work is needed to expand the framework beyond controlled studio setups, particularly to handle multi-person interactions, holistic background environments, and varied lighting conditions.
The primary limitations of this work are its restriction to single-person interactions within a static indoor lighting environment, as well as the absence of full surrounding room scene reconstruction. Despite these boundary conditions, the empirical evidence provides strong confidence in the pipeline's effectiveness for capturing and decoupling complex human-object interactions.
- Paper: D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video, Tianhao Wu et al. (2022). Presents foundational methods for decomposing scenes into separate static and dynamic neural radiance fields with explicit regularization, directly informing NeuralDome's layer-wise neural field decomposition.
- Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). Introduces continuous volumetric deformation fields and regularization for dynamic NeRFs, providing core conceptual building blocks for dynamic human neural modeling.
- Paper: PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization, Shunsuke Saito et al. (2019). Establishes pixel-aligned implicit functions for high-resolution 3D human shape reconstruction, a foundational baseline and technique referenced in neural geometry modeling.
- Paper: AMASS: Archive of Motion Capture As Surface Shapes, Naureen Mahmood et al. (2019). Provides the foundational parametric body models and motion capture normalization techniques essential for tracking articulated human skeletal motion and geometry.
- Paper: GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields, Michael Niemeyer et al. (2021). Introduces compositional neural feature field representations for independently controlling and rendering distinct dynamic and static objects in a scene.
- Paper: End-to-End Recovery of Human Shape and Pose, Angjoo Kanazawa et al. (2017). Establishes end-to-end parametric human body mesh recovery from images, serving as a primary prerequisite for vision-based human motion and pose tracking.
- Paper: Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields, Jonathan T. Barron et al. (2022). Develops anti-aliased, unbounded 360-degree neural radiance field rendering strategies and regularizers widely utilized in dense multi-camera capture pipelines.
- Paper: 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis, Zhicheng Lu et al. (2024). Extends dynamic novel view synthesis beyond volumetric neural radiance fields to real-time 3D deformable Gaussian splatting.
- Paper: TAPVid-3D: A Benchmark for Tracking Any Point in 3D, Skanda Koppula et al. (2024). Generalizes multi-view dynamic capture and tracking by establishing a broad benchmark for tracking any physical point in 3D across dynamic manipulation sequences.
- Paper: DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision, Lu Ling et al. (2024). Scales multi-view scene capture and novel view synthesis benchmarking to massive, diverse real-world environments.
- Paper: Free3D: Consistent Novel View Synthesis Without 3D Representation, Chuanxia Zheng et al. (2024). Explores synthesizing consistent 360-degree novel views of interacting objects without requiring explicit multi-camera 3D reconstruction pipelines.
- Paper: Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors, Guocheng Qian et al. (2024). Applies combined 2D and 3D diffusion priors to reconstruct high-fidelity textured 3D assets from single unposed images, advancing single-view interaction asset generation.
- Paper: VOID: Video Object and Interaction Deletion, Saman Motamed et al. (2026). Applies human-object interaction decomposition to physically plausible video editing and counterfactual object deletion.
