NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions

Juze ZhangHaimin LuoHongdi YangXinru XuQianyang WuYe ShiJingyi YuLan XuJingya Wang

article2023CVPR80 citations

Introduces a large-scale multi-view dataset alongside a layer-wise neural radiance pipeline to accurately decouple, track, and photorealistically render dynamic human-object interactions under severe occlusions.

Listen

Accurately capturing, modeling, and rendering complex human-object interactions in three dimensions is vital for advancing fields such as digital entertainment, vocational training, tele-medicine, and sports analytics. However, existing visual capture methods struggle significantly with mutual occlusions, complex motions, and texture ambiguities when humans interact with physical items. Progress in this domain has been largely bottlenecked by the absence of dense-view appearance datasets and processing methods capable of cleanly separating interacting subjects from objects.

The main objective of the article is to establish a comprehensive data and modeling framework that captures dense multi-view human-object interactions and evaluates a specialized neural pipeline for motion tracking, geometry reconstruction, and photorealistic novel-view rendering.

To accomplish this, the authors constructed a dense camera capture dome comprising 76 high-resolution cinema cameras synchronized with 16 optical motion tracking cameras. Using this setup, they collected the HODome dataset, encompassing 274 interaction sequences across 10 diverse subjects and 23 physical objects, totaling approximately 71 million video frames. To process this massive data stream, the authors developed NeuralDome, a layer-wise neural modeling pipeline. The pipeline jointly tracks human skeletal motion alongside rigid object poses, decomposes scenes into dynamic human and static rigid neural radiance fields, applies specialized geometry and contact regularizers, and utilizes image blending to produce decoupled, high-fidelity digital assets.

The findings demonstrate substantial performance advantages over existing approaches. In novel-view appearance synthesis, NeuralDome achieved an average peak signal-to-noise ratio of 31.93 dB, notably outperforming alternative neural rendering baselines which achieved 22.67 dB and 24.99 dB. When benchmarking 3D geometry reconstruction, training an existing baseline model on this data reduced surface-to-point reconstruction error by roughly 62% for single-view inputs and by over 90% for multi-view inputs compared to the pre-trained baseline. Furthermore, the decoupling strategy successfully separated human bodies from contacted objects without the severe visual artifacts and blending errors typical of prior methods.

These results show that layer-wise neural decomposition, combined with high-density visual supervision, resolves longstanding occlusion and separation challenges in visual computing. For decision-makers, this framework lowers the technical risk and labor associated with digital asset creation, opening practical pathways for training high-performing computer vision models using fewer cameras during downstream deployment.

Looking ahead, development teams and researchers should leverage the publicly released dataset and processing tools to train generalizable interaction models and develop downstream sparse-view rendering applications. Further engineering work is needed to expand the framework beyond controlled studio setups, particularly to handle multi-person interactions, holistic background environments, and varied lighting conditions.

The primary limitations of this work are its restriction to single-person interactions within a static indoor lighting environment, as well as the absence of full surrounding room scene reconstruction. Despite these boundary conditions, the empirical evidence provides strong confidence in the pipeline's effectiveness for capturing and decoupling complex human-object interactions.

Cover for NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions

Abstract

Humans constantly interact with objects in daily life tasks. Capturing such processes and subsequently conducting visual inferences from a fixed viewpoint suffers from occlusions, shape and texture ambiguities, motions, etc. To mitigate the problem, it is essential to build a training dataset that captures free-viewpoint interactions. We construct a dense multi-view dome to acquire a complex human object interaction dataset, named HODome, that consists of ~71M frames on 10 subjects interacting with 23 objects. To process the HODome dataset, we develop NeuralDome, a layer-wise neural processing pipeline tailored for multi-view video inputs to conduct accurate tracking, geometry reconstruction and free-view rendering, for both human subjects and objects. Extensive experiments on the HODome dataset demonstrate the effectiveness of NeuralDome on a variety of inference, modeling, and rendering tasks. Both the dataset and the NeuralDome tools will be disseminated to the community for further development, which can be found at https://juzezhang.github.io/NeuralDome

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Neural Human Rendering
  • 2.2. Human-object Modeling
  • 2.3. Human-centric Dataset
  • 3. HODome Dataset
  • 3.1. Data Capturing System
  • 3.2. Dataset Modality
  • 4. Neural Modeling on HODome
  • 4.1. Human-object Tracking
  • 4.2. Layer-wise Neural Human-Object Rendering
  • 5. Experiments
  • 5.1. Analysis of Neural Rendering
  • 5.2. Task and benchmark
  • 5.3. Limitations
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — HODome Dataset Specification

    experimental setup

    The HODome dataset is a multi-view video and motion capture dataset specifically designed for human-object interaction (HOI) modeling, tracking, and neural rendering.

    Key specifications and setup parameters include:

    • Capture Hardware: 76 synchronized Z-CAM cinema RGB cameras capturing 3840×21603840 \times 2160 resolution video at 60 frames per second (fps)60\text{ frames per second (fps)}, synchronized with 16 OptiTrack optical motion capture cameras.
    • Dataset Scale: 274 interaction sequences totaling approximately 71 million video frames across all camera views, where each sequence lasts approximately 60 seconds60\text{ seconds}.
    • Subjects and Objects: 10 human subjects (5 males, 5 females) wearing varied apparel interacting with 23 distinct rigid 3D objects spanning diverse geometries and interaction types.
    • Annotations and Digital Assets: 3D textured mesh templates pre-scanned for all 23 objects, optical-marker-based 6-DoF rigid object poses, SMPL-X parametric human body/face/hand motion parameters, 2D human keypoints (25 body joints, 42 hand joints), pseudo-contact masks, human-annotated segmentation masks on a dedicated evaluation benchmark subset, and decoupled dynamic neural radiance field representations.
  2. Knowl 2 — Layer-Wise Neural Human-Object Representation

    model/method

    NeuralDome models dynamic human-object interactions as decoupled layered neural radiance fields consisting of a dynamic human layer and a static rigid object layer:

    1. Dynamic Human Layer: Formulated as a pose-embedded dynamic Neural Radiance Field (NeRF) defined in canonical space. Skeletal pose parameters derived from the parametric SMPL-X model transform live frame coordinates to canonical space. A non-rigid deformation multi-layer perceptron (MLP) conditioned on skeletal pose embeddings EpE_p predicts residual deformations. Time-varying human appearance and clothing dynamics are conditioned on a human appearance latent code lhcl_h^c.

    2. Rigid Object Layer: Formulated as a static neural radiance field in canonical object space. Live points are mapped to canonical space using tracked 6-DoF rigid object transformations (Rt,Tt)∈SO(3)×R3(R_t, T_t) \in \mathrm{SO}(3) \times \mathbb{R}^3. Time-varying shadows cast during interactions are modeled by conditioning the canonical object radiance field on an object appearance latent code locl_o^c.

  3. Knowl 3 — Dynamic Human-Object Volume Rendering with Object-Aware Ray Sampling

    model/method

    To render composite human-object scenes, camera rays are integrated across segmented depth intervals corresponding to the human and object volumes.

    For a ray intersecting the bounding volume of the ii-th entity between near depth dnid_n^i and far depth dfid_f^i, the ray segment is divided into NN equal bins, and depth points pjip_j^i are uniformly sampled within each bin:

    pji∼U(dni+j−1N(dfi−dni),  dni+jN(dfi−dni)),j∈{1,2,…,N}p_j^i \sim \mathcal{U}\left(d_n^i + \frac{j-1}{N}(d_f^i - d_n^i), \; d_n^i + \frac{j}{N}(d_f^i - d_n^i)\right), \quad j \in \{1, 2, \dots, N\}

    To increase sampling efficiency for objects, the tracked 3D object template mesh is used to compute exact ray-surface intersections, restricting point sampling to a narrow interval around the first intersection. Ray samples across human and object intervals are merged and sorted by ascending depth into a single sequence of MM sample points (p1,p2,…,pM)(p_1, p_2, \dots, p_M). The composite pixel color CC is computed via numerical quadrature:

    C=∑i=1MT(pi)(1−exp⁡(−σpiδpi))cpiC = \sum_{i=1}^M T(p_i) \left(1 - \exp(-\sigma_{p_i} \delta_{p_i})\right) c_{p_i}

    where δpi\delta_{p_i} is the distance between adjacent sample points pip_i and pi+1p_{i+1}, σpi\sigma_{p_i} is volume density, cpic_{p_i} is RGB color, and the accumulated transmittance T(pi)T(p_i) is:

    T(pi)=∏j=1i−1exp⁡(−σpjδpj)T(p_i) = \prod_{j=1}^{i-1} \exp(-\sigma_{p_j} \delta_{p_j})

  4. Knowl 4 — Joint Optimization Objective for Human-Object Tracking

    equation

    To accurately register interacting humans and objects while preventing physical interpenetration and tracking drift, human SMPL-X parameters and object rigid poses are jointly optimized per frame by minimizing the overall energy:

    E(βt,θt,ψt,γt,Rt,Tt)=Esmpl+λcontactEcontact+λhomaskEhomask+λmakerEmakerE(\beta_t, \theta_t, \psi_t, \gamma_t, R_t, T_t) = E_{\mathrm{smpl}} + \lambda_{\mathrm{contact}} E_{\mathrm{contact}} + \lambda_{\mathrm{homask}} E_{\mathrm{homask}} + \lambda_{\mathrm{maker}} E_{\mathrm{maker}}

    where:

    • βt∈R10\beta_t \in \mathbb{R}^{10} represents SMPL-X human body shape parameters at frame tt.
    • θt\theta_t represents SMPL-X human body, jaw, and finger joint pose parameters at frame tt.
    • ψt\psi_t represents SMPL-X facial expression parameters at frame tt.
    • γt∈R3\gamma_t \in \mathbb{R}^3 is the global human root translation at frame tt.
    • Rt∈SO(3)R_t \in \mathrm{SO}(3) and Tt∈R3T_t \in \mathbb{R}^3 are the 3D rotation matrix and translation vector of the rigid object relative to its canonical pre-scanned mesh template.
    • EsmplE_{\mathrm{smpl}} is a multi-view 2D reprojection data-fitting loss measuring the ℓ2\ell_2 distance between detected 2D body/hand keypoints across all camera views and projected SMPL-X 3D joints.
    • EcontactE_{\mathrm{contact}} enforces valid contact surfaces and penalizes geometric penetration between the human mesh and object template.
    • EhomaskE_{\mathrm{homask}} measures silhouette alignment with foreground segmentation masks across camera views.
    • EmakerE_{\mathrm{maker}} enforces geometric alignment between tracked optical markers and corresponding marker positions on the object template mesh.
    • λcontact,λhomask,λmaker\lambda_{\mathrm{contact}}, \lambda_{\mathrm{homask}}, \lambda_{\mathrm{maker}} are positive balancing scalar weights.
  5. Knowl 5 — Template-Aware Geometry Regularizers for Layered HOI Radiance Fields

    equation

    To prevent density ambiguities and suppress geometric artifacts where human and object bounding volumes overlap, two geometry regularizers are applied based on the tracked object template mesh M\mathcal{M}:

    1. Object Occupancy and Sparsity Loss: Forces the object radiance field to exist as a solid surface inside the object template and remain zero outside:

    Lo=∑p∈O(Ω−(p,M)∥exp⁡(−σp)∥22+Ω+(p,M)∥σp∥22)\mathcal{L}_o = \sum_{p \in \mathcal{O}} \left( \Omega^-(p, \mathcal{M}) \|\exp(-\sigma_p)\|_2^2 + \Omega^+(p, \mathcal{M}) \|\sigma_p\|_2^2 \right)

    where O\mathcal{O} is a set of points randomly sampled inside the object bounding box, σp\sigma_p is the predicted object density at point pp, Ω−(p,M)\Omega^-(p, \mathcal{M}) is an indicator function evaluating to 11 if point pp is inside mesh M\mathcal{M} and 00 otherwise, and Ω+(p,M)\Omega^+(p, \mathcal{M}) evaluates to 11 if pp is outside mesh M\mathcal{M} and 00 otherwise.

    1. Human Non-Penetration Regularizer: Prevents the human non-rigid deformation network from placing human density inside the object volume:

    Lh=∑p∈HΩ−(p,M)∥σp∥22\mathcal{L}_h = \sum_{p \in \mathcal{H}} \Omega^-(p, \mathcal{M}) \|\sigma_p\|_2^2

    where H\mathcal{H} is a set of points randomly sampled in the human bounding box, and σp\sigma_p is the human volume density at point pp.

  6. Knowl 6 — Weakly-Supervised Semantic Layer Decoupling

    model/method

    To cleanly separate the human and object layers without per-frame manual segmentation, NeuralDome uses a weakly-supervised semantic regularization scheme:

    1. Pseudo-Mask Extraction: A preliminary global radiance field is optimized over the temporal sequence. Because rigid object motion is constrained across frames, rendering the object field independently yields reliable appearance reconstructions in regions visible to cameras. Pixels whose rendered colors match the observed RGB images within a color threshold are identified as high-confidence object pixels, forming a coarse pseudo-segmentation mask So\mathcal{S}_o.

    2. Volumetric Semantic Integration: For each ray rr, a continuous layer semantic label ss is rendered via volumetric alpha-compositing:

    s=∑i=1MT(pi)(1−exp⁡(−σpiδpi))lpis = \sum_{i=1}^M T(p_i) \left(1 - \exp(-\sigma_{p_i} \delta_{p_i})\right) l_{p_i}

    where lpil_{p_i} is a one-hot vector indicating whether sample point pip_i belongs to the human layer or the object layer.

    1. Semantic Loss: Supervised training is applied across rays intersecting the pseudo-labeled object region So\mathcal{S}_o:

    Ls=∑r∈So∥sr−s^r∥22\mathcal{L}_s = \sum_{r \in \mathcal{S}_o} \|s_r - \hat{s}_r\|_2^2

    where s^r\hat{s}_r is the one-hot pseudo-ground-truth semantic label for ray rr.

  7. Knowl 7 — Texture Blending and Per-Frame Neural Asset Enhancement

    model/method

    Because temporally aggregated radiance fields can suffer from high-frequency texture blurriness caused by subtle non-rigid cloth dynamics and tracking inaccuracies, NeuralDome employs an image-blending enhancement step to construct per-frame neural representations:

    1. Segmentation Mask Rendering: The trained layered model renders crisp, occlusion-free semantic masks for the individual human and object instances from each camera view.
    2. Image Projection and Blending: The unoccluded, high-resolution observed pixels from the 76 synchronized camera views are masked and blended onto the rendered individual human and object image planes.
    3. Per-Frame Asset Reconstruction: Instant-NGP (for fast novel view synthesis) and Instant-NSR (for surface geometry extraction via neural implicit surfaces) are trained on the decoupled, high-frequency blended multi-view images. This process runs in seconds per frame to produce decoupled, photo-realistic human and object neural assets.
  8. Knowl 8 — Novel-View Synthesis Evaluation on Human-Object Interactions

    data/table

    NeuralDome's layered neural rendering approach was evaluated against NeuralBody (NB) and ST-NeRF for novel view synthesis on HOI sequences captured across 76 viewpoints (20 viewpoints used for training, 56 held out for testing). Quality is measured using Peak Signal-to-Noise Ratio (PSNR in dB) and Structural Similarity Index Measure (SSIM).

    Methods NB [45] ST-NeRF [77] Ours
    Scenes PSNR ↑\uparrow SSIM ↑\uparrow PSNR ↑\uparrow SSIM ↑\uparrow PSNR ↑\uparrow SSIM ↑\uparrow
    Bigsofa 19.33 0.896 24.02 0.886 32.28 0.958
    Sofa 26.73 0.965 28.49 0.958 35.62 0.987
    Table 21.94 0.933 22.44 0.900 27.88 0.949
    Average 22.67 0.931 24.99 0.915 31.93 0.964

    NeuralBody lacks explicit object modeling, resulting in floating artifacts in the reconstructed density field. ST-NeRF struggles when human and object bounding boxes heavily overlap without semantic guidance. NeuralDome outperforms both baselines by +6.94 dB PSNR and +0.049 SSIM on average over ST-NeRF.

  9. Knowl 9 — Evaluation on Human-Object Capture, 3D Reconstruction, and Sparse Rendering

    data/table

    The HODome dataset was used to benchmark three downstream visual inference tasks under complex interactions:

    1. Human-Object Pose Capture: Evaluated on Mean Per Joint Position Error (MPJPE in mm), Procrustes-Aligned MPJPE (PA-MPJPE in mm), human Chamfer distance (Chamferh_h in mm), object Chamfer distance (Chamfero_o in mm), Vertex-to-Vertex error (V2V in mm), and Procrustes-aligned V2V (p.V2V in mm).
    Method MPJPE ↓\downarrow PA-MPJPE ↓\downarrow Chamferh_h ↓\downarrow Chamfero_o ↓\downarrow V2V ↓\downarrow p.V2V ↓\downarrow
    Fit to input 16.11 6.94 86.66 18.14 61.57 28.92
    PHOSA 14.88 6.94 77.62 16.64 54.59 26.09
    CHORE 10.21 6.14 86.58 7.69 44.93 14.93
    1. Geometry Reconstruction (PIFu): Evaluated using Point-to-Surface (P2S ×10−4\times 10^{-4}) and Chamfer Distance (CD ×10−4\times 10^{-4}).
    Method P2S ×10−4↓\times 10^{-4} \downarrow Chamfer ×10−4↓\times 10^{-4} \downarrow
    Origin PIFu 38.726 40.947
    Monocular PIFu-trained 14.653 14.483
    6-View PIFu-trained 3.376 4.901
    1. Sparse-View Novel View Rendering:
    Method PSNR ↑\uparrow SSIM ↑\uparrow
    IBRNet 21.43 0.892
    NeuRay 23.34 0.909
    NeuralHumanFVV 21.69 0.914
    NeuralHOIFVV (Ours) 23.10 0.912

    Training PIFu and neural rendering models on HODome significantly improves geometric accuracy and novel-view quality under severe human-object occlusions compared to pre-trained single-human baselines.

  10. Knowl 10 — Limitations of the NeuralDome Pipeline and Dataset

    limitation

    The NeuralDome pipeline and HODome dataset have three primary limitations:

    1. Single-Subject Interaction Assumption: The framework explicitly tracks and decomposes one human interacting with rigid objects; it cannot directly handle multi-person interactions or crowd scenarios.
    2. Absence of Holistic 3D Environment Modeling: The pipeline models only the foreground human and object assets; reconstruction of surrounding background static scenes and environment geometry is not included.
    3. Studio Environment Invariance: Capture was performed within a controlled multi-camera dome under fixed studio lighting and uniform green background conditions, limiting zero-shot generalization to diverse in-the-wild illumination and unstructured environments.

Coverage note — No substantial contributed material was omitted; detailed mathematical expansions of the tracking loss sub-terms ($E_{smpl}, E_{contact}, E_{homask}, E_{maker}$) and standard off-the-shelf Instant-NGP/Instant-NSR architectures were summarized according to their standalone contributions.

References

  1. 1.Naturalpoint, inc. motion capture systems. https://optitrack.com/. 6. 3
  2. 2.Reality capture. https://www.capturingreality.com/realitycapture. 3, 4
  3. 3.Easymocap - make human motion capture easier. Github, 2021. 4
  4. 4.Kara-Ali Aliev, Artem Sevastopolsky, Maria Kolos, Dmitry Ulyanov, and Victor Lempitsky. Neural point-based graphics. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXII 16, pages 696–712. Springer, 2020. 2
  5. 5.Paul J Besl and Neil D McKay. Method for registration of 3-d shapes. In Sensor fusion IV: control paradigms and data structures, volume 1611, pages 586–606. Spie, 1992. 4
  6. 6.Bharat Lal Bhatnagar, Xianghui Xie, Ilya A Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Behave: Dataset and method for tracking human object interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15935–15946, 2022. 2, 3, 4, 7
  7. 7.Hongrui Cai, Wanquan Feng, Xuetao Feng, Yan Wang, and Juyong Zhang. Neural surface reconstruction of dynamic scenes with monocular rgb-d camera. In Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS), 2022. 2
  8. 8.Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. Realtime multi-person 2d pose estimation using part affinity fields. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7291–7299, 2017. 4
  9. 9.Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. arXiv preprint arXiv:2203.09517, 2022. 2
  10. 10.Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14124–14133, 2021. 2
  11. 11.Alvaro Collet, Ming Chuang, Pat Sweeney, Don Gillett, Dennis Evseev, David Calabrese, Hugues Hoppe, Adam Kirk, and Steve Sullivan. High-quality streamable free-viewpoint video. ACM Transactions on Graphics (TOG), 34(4):69, 2015. 2, 3
  12. 12.Mingsong Dou, Philip Davidson, Sean Ryan Fanello, Sameh Khamis, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, and Shahram Izadi. Motion2fusion: Real-time volumetric performance capture. ACM Trans. Graph., 36(6):246:1–246:16, Nov. 2017. 2, 3
  13. 13.Kaiwen Guo, Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts-Escolano, Rohit Pandey, Jason Dourgarian, et al. The relightables: Volumetric performance capture of humans with realistic relighting. ACM Transactions on Graphics (ToG), 38(6):1–19, 2019. 3
  14. 14.Marc Habermann, Lingjie Liu, Weipeng Xu, Michael Zollhoefer, Gerard Pons-Moll, and Christian Theobalt. Real-time deep dynamic characters. ACM Transactions on Graphics (TOG), 40(4):1–16, 2021. 2
  15. 15.Nils Hasler, Carsten Stoll, Martin Sunkel, Bodo Rosenhahn, and H-P Seidel. A statistical model of human pose and body shape. In Computer graphics forum, volume 28, pages 337–346. Wiley Online Library, 2009. 3
  16. 16.Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J Black. Resolving 3d human pose ambiguities with 3d scene constraints. In Proceedings of the IEEE/CVF international conference on computer vision, pages 2282–2292, 2019. 3, 5
  17. 17.Tong He, Yuanlu Xu, Shunsuke Saito, Stefano Soatto, and Tony Tung. Arch++: Animation-ready clothed human reconstruction revisited. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11046–11056, 2021. 2
  18. 18.Yinghao Huang, Omid Taheri, Michael J Black, and Dimitrios Tzionas. Intercap: Joint markerless 3d tracking of humans and objects in interaction. In DAGM German Conference on Pattern Recognition, pages 281–299. Springer, 2022. 3
  19. 19.Zeng Huang, Yuanlu Xu, Christoph Lassner, Hao Li, and Tony Tung. Arch: Animatable reconstruction of clothed humans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3093–3102, 2020. 2
  20. 20.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7):1325–1339, jul 2014. 3
  21. 21.Yuheng Jiang, Suyi Jiang, Guoxing Sun, Zhuo Su, Kaiwen Guo, Minye Wu, Jingyi Yu, and Lan Xu. Neuralhofusion: Neural volumetric rendering under human-object interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6155–6165, 2022. 2, 3
  22. 22.Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, Timothy Scott Godisart, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh. Panoptic studio: A massively multiview system for social interaction capture. 2017. 3
  23. 23.Youngjoong Kwon, Dahun Kim, Duygu Ceylan, and Henry Fuchs. Neural human performer: Learning generalizable radiance fields for human performance rendering. Advances in Neural Information Processing Systems, 34:24741–24752, 2021. 2
  24. 24.Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J. Black, and Peter V. Gehler. Unite the people: Closing the loop between 3d and 2d human representations. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), July 2017. 3
  25. 25.Ruilong Li, Julian Tanke, Minh Vo, Michael Zollhofer, Jurgen Gall, Angjoo Kanazawa, and Christoph Lassner. Tava: Template-free animatable volumetric actors. arXiv preprint arXiv:2206.08929, 2022. 2
  26. 26.Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding. IEEE transactions on pattern analysis and machine intelligence, 42(10):2684–2701, 2019. 2, 3
  27. 27.Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, and Christian Theobalt. Neural actor: Neural free-view synthesis of human actors with pose control. ACM Trans. Graph.(ACM SIGGRAPH Asia), 2021. 2
  28. 28.Lingjie Liu, Weipeng Xu, Marc Habermann, Michael Zollhöfer, Florian Bernard, Hyeongwoo Kim, Wenping Wang, and Christian Theobalt. Neural human video rendering by learning dynamic textures and rendering-to-video translation. arXiv preprint arXiv:2001.04947, 2020. 2
  29. 29.Lingjie Liu, Weipeng Xu, Michael Zollhoefer, Hyeongwoo Kim, Florian Bernard, Marc Habermann, Wenping Wang, and Christian Theobalt. Neural rendering and reenactment of human actor videos. ACM Transactions on Graphics (TOG), 38(5):1–14, 2019. 2
  30. 30.Yuan Liu, Sida Peng, Lingjie Liu, Qianqian Wang, Peng Wang, Christian Theobalt, Xiaowei Zhou, and Wenping Wang. Neural rays for occlusion-aware image-based rendering. In CVPR, 2022. 2, 8
  31. 31.Stephen Lombardi, Tomas Simon, Jason Saragih, Gabriel Schwartz, Andreas Lehrmann, and Yaser Sheikh. Neural volumes: Learning dynamic renderable volumes from images. arXiv preprint arXiv:1906.07751, 2019. 2
  32. 32.Stephen Lombardi, Tomas Simon, Gabriel Schwartz, Michael Zollhoefer, Yaser Sheikh, and Jason Saragih. Mixture of volumetric primitives for efficient neural rendering. arXiv preprint arXiv:2103.01954, 2021. 2
  33. 33.Xiaoxiao Long, Cheng Lin, Peng Wang, Taku Komura, and Wenping Wang. Sparseneus: Fast generalizable neural surface reconstruction from sparse views. ECCV, 2022. 2
  34. 34.H. Luo, A. Chen, Q. Zhang, B. Pang, M. Wu, L. Xu, and J. Yu. Convolutional neural opacity radiance fields. In 2021 IEEE International Conference on Computational Photography (ICCP), pages 1–12, Los Alamitos, CA, USA, may 2021. IEEE Computer Society. 2
  35. 35.Haimin Luo, Teng Xu, Yuheng Jiang, Chenglin Zhou, Qiwei Qiu, Yingliang Zhang, Wei Yang, Lan Xu, and Jingyi Yu. Artemis: Articulated neural pets with appearance and motion synthesis. ACM Trans. Graph., 41(4), jul 2022. 2
  36. 36.Quan Meng, Anpei Chen, Haimin Luo, Minye Wu, Hao Su, Lan Xu, Xuming He, and Jingyi Yu. Gnerf: Gan-based neural radiance field without posed camera. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6351–6361, 2021. 2
  37. 37.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European conference on computer vision, pages 405–421. Springer, 2020. 2
  38. 38.Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph., 41(4):102:1–102:15, July 2022. 2, 4, 6
  39. 39.Michael Niemeyer, Jonathan T. Barron, Ben Mildenhall, Mehdi S. M. Sajjadi, Andreas Geiger, and Noha Radwan. Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  40. 40.Atsuhiro Noguchi, Xiao Sun, Stephen Lin, and Tatsuya Harada. Neural articulated radiance field. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5762–5772, 2021. 2
  41. 41.Anqi Pang, Xin Chen, Haimin Luo, Minye Wu, Jingyi Yu, and Lan Xu. Few-shot neural human performance rendering from sparse rgbd videos, 2021. 2
  42. 42.Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M. Seitz. Hypernerf: A higher-dimensional representation for topologically varying neural radiance fields. ACM Trans. Graph., 40(6), dec 2021. 2
  43. 43.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10975–10985, 2019. 2, 5
  44. 44.Sida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang, Qing Shuai, Xiaowei Zhou, and Hujun Bao. Animatable neural radiance fields for modeling dynamic human bodies. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14314–14323, 2021. 2
  45. 45.Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In CVPR, 2021. 2, 6, 7
  46. 46.Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, jun 2021. 2
  47. 47.Yu Rong, Takaaki Shiratori, and Hanbyul Joo. Frankmocap: A monocular 3d whole-body pose estimation system via regression and integration. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1749–1759, 2021. 7
  48. 48.Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Angjoo Kanazawa, and Hao Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2304–2314, 2019. 2, 7, 8
  49. 49.Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul Joo. Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 84–93, 2020. 2
  50. 50.Manolis Savva, Angel X Chang, Pat Hanrahan, Matthew Fisher, and Matthias Nießner. Pigraphs: learning interaction snapshots from observations. ACM Transactions on Graphics (TOG), 35(4):1–12, 2016. 3
  51. 51.Johannes Lutz Schönberger and Jan-Michael Frahm. Structure-from-Motion Revisited. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 2, 3
  52. 52.Qing Shuai, Chen Geng, Qi Fang, Sida Peng, Wenhao Shen, Xiaowei Zhou, and Hujun Bao. Novel view synthesis of human interactions from sparse multi-view videos. In ACM SIGGRAPH 2022 Conference Proceedings, pages 1–10, 2022. 2
  53. 53.Qing Shuai, Chen Geng, Qi Fang, Sida Peng, Wenhao Shen, Xiaowei Zhou, and Hujun Bao. Novel view synthesis of human interactions from sparse multi-view videos. In ACM SIGGRAPH 2022 Conference Proceedings, SIGGRAPH '22, New York, NY, USA, 2022. Association for Computing Machinery. 3, 6
  54. 54.Aliaksandra Shysheya, Egor Zakharov, Kara-Ali Aliev, Renat Bashirov, Egor Burkov, Karim Iskakov, Aleksei Ivakhnenko, Yury Malkov, Igor Pasechnik, Dmitry Ulyanov, et al. Textured neural avatars. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2387–2397, 2019. 2
  55. 55.Leonid Sigal, Alexandru O Balan, and Michael J Black. Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. International journal of computer vision, 87(1):4–27, 2010. 3
  56. 56.Zhuo Su, Lan Xu, Zerong Zheng, Tao Yu, Yebin Liu, and Lu Fang. Robustfusion: Human volumetric capture with data-driven visual cues using a rgbd camera. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision – ECCV 2020, pages 246–264, Cham, 2020. Springer International Publishing. 3
  57. 57.Guoxing Sun, Xin Chen, Yizhang Chen, Anqi Pang, Pei Lin, Yuheng Jiang, Lan Xu, Jingya Wang, and Jingyi Yu. Neural free-viewpoint performance rendering under complex human-object interactions. In Proceedings of the 29th ACM International Conference on Multimedia, 2021. 2, 3
  58. 58.Guoxing Sun, Xin Chen, Yizhang Chen, Anqi Pang, Pei Lin, Yuheng Jiang, Lan Xu, Jingyi Yu, and Jingya Wang. Neural free-viewpoint performance rendering under complex human-object interactions. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4651–4660, 2021. 3
  59. 59.Xin Suo, Yuheng Jiang, Pei Lin, Yingliang Zhang, Minye Wu, Kaiwen Guo, and Lan Xu. Neuralhumanfvv: Real-time neural volumetric human performance rendering using rgb cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6226–6237, 2021. 2
  60. 60.Xin Suo, Yuheng Jiang, Pei Lin, Yingliang Zhang, Minye Wu, Kaiwen Guo, and Lan Xu. Neuralhumanfvv: Real-time neural volumetric human performance rendering using rgb cameras. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6222–6233, 2021. 8
  61. 61.Omid Taheri, Nima Ghorbani, Michael J Black, and Dimitrios Tzionas. Grab: A dataset of whole-body human grasping of objects. In European conference on computer vision, pages 581–600. Springer, 2020. 2, 3, 4
  62. 62.Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, Yifan Wang, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, Tomas Simon, Christian Theobalt, Matthias Niessner, Jonathan T. Barron, Gordon Wetzstein, Michael Zollhoefer, and Vladislav Golyanik. Advances in neural rendering, 2021. 2
  63. 63.Edgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhofer, Christoph Lassner, and Christian Theobalt. Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monocular video. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12959–12970, 2021. 2
  64. 64.Dor Verbin, Peter Hedman, Ben Mildenhall, Todd Zickler, Jonathan T Barron, and Pratul P Srinivasan. Ref-nerf: Structured view-dependent appearance for neural radiance fields. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5481–5490. IEEE, 2022. 2
  65. 65.Timo Von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. In Proceedings of the European Conference on Computer Vision (ECCV), pages 601–617, 2018. 3
  66. 66.Liao Wang, Ziyu Wang, Pei Lin, Yuheng Jiang, Xin Suo, Minye Wu, Lan Xu, and Jingyi Yu. ibutter: Neural interactive bullet time generator for human free-viewpoint rendering. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4641–4650, 2021. 2
  67. 67.Liao Wang, Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu, and Lan Xu. Fourier plenoctrees for dynamic radiance field rendering in real-time. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13524–13534, 2022. 2
  68. 68.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. NeurIPS, 2021. 2
  69. 69.Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul Srinivasan, Howard Zhou, Jonathan T. Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser. Ibrnet: Learning multi-view image-based rendering. In CVPR, 2021. 2, 8
  70. 70.Shaofei Wang, Katja Schwarz, Andreas Geiger, and Siyu Tang. Arah: animatable volume rendering of articulated human sdfs. arXiv preprint arXiv:2210.10036, 2022. 2
  71. 71.Minye Wu, Yuehao Wang, Qiang Hu, and Jingyi Yu. Multi-view neural human rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2
  72. 72.Minye Wu, Yuehao Wang, Qiang Hu, and Jingyi Yu. Multi-view neural human rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1682–1691, 2020. 2
  73. 73.Xianghui Xie, Bharat Lal Bhatnagar, and Gerard Pons-Moll. Chore: Contact, human and object reconstruction from a single rgb image. arXiv preprint arXiv:2204.02445, 2022. 3, 7
  74. 74.Hongwei Yi, Chun-Hao P Huang, Dimitrios Tzionas, Muhammed Kocabas, Mohamed Hassan, Siyu Tang, Justus Thies, and Michael J Black. Human-aware object placement for visual environment reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3959–3970, 2022. 3, 7
  75. 75.Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. Plenoctrees for real-time rendering of neural radiance fields. In arXiv, 2021. 2
  76. 76.Chao Zhang, Sergi Pujades, Michael J. Black, and Gerard Pons-Moll. Detailed, accurate, human shape estimation from clothed 3d scan sequences. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017. 3
  77. 77.Jiakai Zhang, Xinhang Liu, Xinyi Ye, Fuqiang Zhao, Yanshun Zhang, Minye Wu, Yingliang Zhang, Lan Xu, and Jingyi Yu. Editable free-viewpoint video using a layered neural representation. ACM Trans. Graph., 40(4), July 2021. 2, 3, 5, 6, 7
  78. 78.Jason Y Zhang, Sam Pepose, Hanbyul Joo, Deva Ramanan, Jitendra Malik, and Angjoo Kanazawa. Perceiving 3d human-object spatial arrangements from a single image in the wild. In European conference on computer vision, pages 34–51. Springer, 2020. 3, 7
  79. 79.Fuqiang Zhao, Yuheng Jiang, Kaixin Yao, Jiakai Zhang, Liao Wang, Haizhao Dai, Yuhui Zhong, Yingliang Zhang, Minye Wu, Lan Xu, et al. Human performance modeling and rendering via neural animated mesh. arXiv preprint arXiv:2209.08468, 2022. 2, 4, 6, 8
  80. 80.Fuqiang Zhao, Wei Yang, Jiakai Zhang, Pei Lin, Yingliang Zhang, Jingyi Yu, and Lan Xu. Humannerf: Efficiently generated human radiance field from sparse inputs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7743–7753, 2022. 2, 5
  81. 81.Qian-Yi Zhou, Jaesik Park, and Vladlen Koltun. Open3d: A modern library for 3d data processing. arXiv preprint arXiv:1801.09847, 2018. 4

Citation

MLA
Zhang, J., et al. “NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions”. arXiv, 2022, http://arxiv.org/abs/2212.07626v1.
APA
Zhang, J., Luo, H., Yang, H., Xu, X., Wu, Q., Shi, Y., Yu, J., Xu, L., & Wang, J. (2022). NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions. arXiv. http://arxiv.org/abs/2212.07626v1
Chicago
Zhang, J., H. Luo, H. Yang, et al. 2022. “NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions”. arXiv. http://arxiv.org/abs/2212.07626v1.
Harvard
Zhang, J. et al. (2022) “NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.07626v1.
Vancouver
1. Zhang J, Luo H, Yang H, Xu X, Wu Q, Shi Y, Yu J, Xu L, Wang J (2022) NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions. arXiv

BibTeX

@article{zhang2022neuraldome,
  title = {NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions},
  author = {Zhang, Juze and Luo, Haimin and Yang, Hongdi and Xu, Xinru and Wu, Qianyang and Shi, Ye and Yu, Jingyi and Xu, Lan and Wang, Jingya},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.07626v1},
  eprint = {2212.07626}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE