FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
Alexandros DelitzasChenyangguang ZhangAlexey GavryushinTommaso Di MarioBoyang SunRishabh DabralLeonidas GuibasChristian TheobaltMarc PollefeysFrancis Engelmann
Develops a framework for reconstructing simulation-ready articulated 3D digital twins directly from in-the-wild egocentric interaction videos, recovering part kinematics and dynamic geometry without requiring CAD priors or controlled multi-state captures.
Standard 3D computer vision techniques capture indoor environments as static, frozen scenes. However, robots and embodied AI systems must interact with physical spaces by opening doors, pulling drawers, and manipulating objects. Existing approaches to modeling dynamic environments rely on labor-intensive multi-state captures, synthetic CAD model substitutions, or controlled laboratory setups with fixed cameras. As a result, autonomous systems lack reliable methods to reconstruct interactive, physically consistent 3D digital twins directly from everyday human interactions in the real world.
The article introduces and evaluates FunREC, a training-free framework designed to reconstruct functional 3D digital twins directly from single egocentric (first-person) RGB-D interaction videos. The main objective is to automatically identify movable parts, estimate their kinematic joint parameters, track their 3D motion, and reconstruct both static and moving geometry—including occluded interiors—to produce simulation-ready 3D models.
The authors approach the problem through a modular optimization pipeline that combines geometric reasoning with foundation vision-language and segmentation models. The system splits video sequences into fragments, classifies interactions, clusters sparse 3D point trajectories into articulated components, extracts pixel-accurate masks, and jointly optimizes part poses and joint parameters. The article evaluates this method against leading dynamic reconstruction and tracking baselines across three benchmarks: HOI4D (30 single-object lab interactions) and two newly introduced benchmarks, RealFun4D (351 real-world interactions across 60 apartments) and OmniFun4D (127 photorealistic simulated sequences).
The findings show that FunREC significantly outperforms existing baselines across all core evaluation metrics. First, in articulated motion estimation, the framework achieves axis direction errors of roughly 5 to 12 degrees and position errors of 0.03 to 0.06 meters across datasets, reducing errors by roughly five- to ten-fold compared to prior methods while maintaining a 0% failure rate on real-world data. Second, FunREC achieves moving part segmentation accuracy of 74.8 to 77.9 mean Intersection-over-Union, improving performance by approximately 50 points over dynamic tracking pipelines. Third, 6D part pose tracking accuracy reaches 75.6% to 79.5%, more than doubling the precision of leading baselines. Finally, 3D surface reconstruction errors were substantially lower, achieving Chamfer Distances of 0.7 cm on HOI4D, 3.2 cm on OmniFun4D, and 6.1 cm on RealFun4D.
These results demonstrate that human-scene interaction provides rich natural supervision to recover physical functionality without manual annotation. In practical terms, this lowers the cost, time, and complexity of building interactive virtual environments for robotics. The reconstructed scenes can be directly exported into standard simulation formats (such as URDF and USD), allowing physical properties, hand affordance contact maps, and robotic manipulation policies to transfer directly from human demonstration to real hardware, such as mobile robotic arms.
Decision-makers and engineering teams should consider piloting this approach to automate the creation of digital twins for simulation-based robot training and spatial computing. Next operational steps should focus on integrating automated physical property estimation (such as mass and friction) and establishing end-to-end data pipelines from casual wearable camera captures to virtual simulation environments.
Confidence in these results is high across tested indoor categories, supported by consistent performance across both synthetic benchmarks and diverse real-world apartments. Nevertheless, readers should account for current operational limitations: the pipeline assumes known camera intrinsics, clear observation of the moving parts during interaction, and relies on pre-trained vision models for semantic segmentation and feature tracking, which could encounter edge-case failures in extreme lighting or severe visual occlusion.
- Paper: HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from Video, Zicong Fan et al. (2024). HOLD establishes foundational techniques for joint category-agnostic 3D reconstruction of interacting hands and objects from video by leveraging physical contact and spatial alignment constraints.
- Paper: gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction, Zerui Chen et al. (2023). gSDF introduces the core methodology of guiding neural signed distance functions with explicit kinematic chains to disentangle pose tracking and shape reconstruction during hand-object manipulation.
- Paper: GART: Gaussian Articulated Template Models, Jiahui Lei et al. (2024). GART provides the mathematical formulation for modeling moving articulated geometry in canonical space and animating it via explicit kinematic skinning under differentiable rendering.
- Paper: D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video, Tianhao Wu et al. (2022). D2NeRF demonstrates how to decouple dynamic moving objects from static background geometry in unconstrained video sequences using self-supervised neural fields.
- Paper: Neural RGB-D Surface Reconstruction, Dejan Azinovic et al. (2022). This work establishes the neural signed distance function formulation for high-quality, room-scale 3D surface reconstruction from consumer RGB-D sequences with joint pose refinement.
- Paper: PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI, Yandan Yang et al. (2024). PhyScene provides key concepts for modeling interactable, articulated 3D indoor objects and evaluating them for physics-based embodied AI simulation.
- Paper: NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions, Juze Zhang et al. (2023). NeuralDome details pipelines for decomposing complex human-object interactions into dynamic and static neural fields under contact constraints.
No sufficiently relevant recommendations were found.
