FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

Alexandros DelitzasChenyangguang ZhangAlexey GavryushinTommaso Di MarioBoyang SunRishabh DabralLeonidas GuibasChristian TheobaltMarc PollefeysFrancis Engelmann

article2026arXiv9 citations

Develops a framework for reconstructing simulation-ready articulated 3D digital twins directly from in-the-wild egocentric interaction videos, recovering part kinematics and dynamic geometry without requiring CAD priors or controlled multi-state captures.

Listen

Standard 3D computer vision techniques capture indoor environments as static, frozen scenes. However, robots and embodied AI systems must interact with physical spaces by opening doors, pulling drawers, and manipulating objects. Existing approaches to modeling dynamic environments rely on labor-intensive multi-state captures, synthetic CAD model substitutions, or controlled laboratory setups with fixed cameras. As a result, autonomous systems lack reliable methods to reconstruct interactive, physically consistent 3D digital twins directly from everyday human interactions in the real world.

The article introduces and evaluates FunREC, a training-free framework designed to reconstruct functional 3D digital twins directly from single egocentric (first-person) RGB-D interaction videos. The main objective is to automatically identify movable parts, estimate their kinematic joint parameters, track their 3D motion, and reconstruct both static and moving geometry—including occluded interiors—to produce simulation-ready 3D models.

The authors approach the problem through a modular optimization pipeline that combines geometric reasoning with foundation vision-language and segmentation models. The system splits video sequences into fragments, classifies interactions, clusters sparse 3D point trajectories into articulated components, extracts pixel-accurate masks, and jointly optimizes part poses and joint parameters. The article evaluates this method against leading dynamic reconstruction and tracking baselines across three benchmarks: HOI4D (30 single-object lab interactions) and two newly introduced benchmarks, RealFun4D (351 real-world interactions across 60 apartments) and OmniFun4D (127 photorealistic simulated sequences).

The findings show that FunREC significantly outperforms existing baselines across all core evaluation metrics. First, in articulated motion estimation, the framework achieves axis direction errors of roughly 5 to 12 degrees and position errors of 0.03 to 0.06 meters across datasets, reducing errors by roughly five- to ten-fold compared to prior methods while maintaining a 0% failure rate on real-world data. Second, FunREC achieves moving part segmentation accuracy of 74.8 to 77.9 mean Intersection-over-Union, improving performance by approximately 50 points over dynamic tracking pipelines. Third, 6D part pose tracking accuracy reaches 75.6% to 79.5%, more than doubling the precision of leading baselines. Finally, 3D surface reconstruction errors were substantially lower, achieving Chamfer Distances of 0.7 cm on HOI4D, 3.2 cm on OmniFun4D, and 6.1 cm on RealFun4D.

These results demonstrate that human-scene interaction provides rich natural supervision to recover physical functionality without manual annotation. In practical terms, this lowers the cost, time, and complexity of building interactive virtual environments for robotics. The reconstructed scenes can be directly exported into standard simulation formats (such as URDF and USD), allowing physical properties, hand affordance contact maps, and robotic manipulation policies to transfer directly from human demonstration to real hardware, such as mobile robotic arms.

Decision-makers and engineering teams should consider piloting this approach to automate the creation of digital twins for simulation-based robot training and spatial computing. Next operational steps should focus on integrating automated physical property estimation (such as mass and friction) and establishing end-to-end data pipelines from casual wearable camera captures to virtual simulation environments.

Confidence in these results is high across tested indoor categories, supported by consistent performance across both synthetic benchmarks and diverse real-world apartments. Nevertheless, readers should account for current operational limitations: the pipeline assumes known camera intrinsics, clear observation of the moving parts during interaction, and relies on pre-trained vision models for semantic segmentation and feature tracking, which could encounter edge-case failures in extreme lighting or severe visual occlusion.

arXiv: 2604.05621

No sufficiently relevant recommendations were found.

Cover for FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

Abstract

We present FunRec, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunRec operates directly on in-the-wild human interaction sequences to recover interactable 3D scenes. It automatically discovers articulated parts, estimates their kinematic parameters, tracks their 3D motion, and reconstructs static and moving geometry in canonical space, yielding simulation-compatible meshes. Across new real and simulated benchmarks, FunRec surpasses prior work by a large margin, achieving up to +50 mIoU improvement in part segmentation, 5-10 times lower articulation and pose errors, and significantly higher reconstruction accuracy. We further demonstrate applications on URDF/USD export for simulation, hand-guided affordance mapping and robot-scene interaction.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Proposed Method
  • 3.1 Fragment Construction
  • 3.2 Dynamic Fragment Reconstruction
  • 3.3 Global Fragment Alignment
  • 4 Data Collection
  • 5 Experiments
  • 6 Applications
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Problem Formulation and Articulated Kinematic Parameterization

    definition

    Let an egocentric RGB-D video with known camera intrinsics be defined as V={(Ii,Di)}i=1NV = \{(I_i, D_i)\}_{i=1}^N, where IiI_i is the RGB frame and DiD_i is the depth map at time step i∈{1,…,N}i \in \{1, \dots, N\}. The 3D scene consists of a static background geometry PsP^s and a set of KK movable articulated components {Pmk}k=1K\{P^{m_k}\}_{k=1}^K.

    Each articulated component mkm_k is parameterized by a joint model ϕk\phi_k and a sequence of rigid transformations {Timk}i=1N⊂SE(3)\{T_i^{m_k}\}_{i=1}^N \subset SE(3) mapping canonical component coordinates to the global world frame at frame ii. Two joint kinematics are modeled:

    1. Prismatic Joint (Translation): Specified by a unit translation axis a∈S2a \in \mathbb{S}^2 on the unit sphere and scalar displacements λi∈R\lambda_i \in \mathbb{R}. The transformation at frame ii is given by: Tpris(a,λi)=[I3×3λia0⊤1]T_{\text{pris}}(a, \lambda_i) = \begin{bmatrix} I_{3 \times 3} & \lambda_i a \\ 0^\top & 1 \end{bmatrix} where I3×3I_{3 \times 3} is the identity matrix.

    2. Revolute Joint (Rotation): Specified by a unit rotation axis a∈S2a \in \mathbb{S}^2, a pivot position p∈R3p \in \mathbb{R}^3 constrained to be orthogonal to the axis (a⊤p=0a^\top p = 0, the closest point on the line of rotation to the origin), and rotation angles θi∈S1\theta_i \in \mathbb{S}^1. The rigid transformation at frame ii is given by: Trev(a,p,θi)=[R(a,θi)(I3×3−R(a,θi))p0⊤1]T_{\text{rev}}(a, p, \theta_i) = \begin{bmatrix} R(a, \theta_i) & (I_{3 \times 3} - R(a, \theta_i))p \\ 0^\top & 1 \end{bmatrix} where R(a,θ)∈SO(3)R(a, \theta) \in SO(3) denotes the 3D rotation matrix about axis aa by angle θ\theta.

  2. Knowl 2 — FunREC Pipeline Architecture

    model/method

    FunREC is a training-free framework that reconstructs functional, interactable 3D digital twins directly from an unconstrained egocentric RGB-D interaction sequence VV. The pipeline operates in three stages:

    1. Fragment Construction and Interaction Identification: The video VV is segmented into temporal fragments {Vk}k=1K\{V_k\}_{k=1}^K. A video-language model (VLM) classifies each fragment as static or dynamic and predicts the joint type σk∈{prismatic,revolute}\sigma_k \in \{\text{prismatic}, \text{revolute}\}.

    2. Dynamic Fragment Processing: For each dynamic fragment:

    • Dense correspondences from RoMA are masked to remove hands and downweight moving objects, followed by SupeRANSAC and fragment-level pose graph optimization to yield camera poses {Tic}i=1N\{T_i^c\}_{i=1}^N.
    • Sparse 3D trajectories τ∈RT×N×3\tau \in \mathbb{R}^{T \times N \times 3} with visibility scores o∈RT×No \in \mathbb{R}^{T \times N} are tracked via TAPIP3D and periodically re-initialized.
    • Articulation-aware clustering assigns tracks to the manipulated part.
    • Track-guided semantic region voting on keyframe over-segmentations combined with SAM2 video propagation extracts dense per-frame articulated part masks {Mim}\{M_i^m\}.
    • Part poses {Tim}\{T_i^m\} and kinematic parameters ϕm\phi^m are jointly optimized across the sequence using non-linear manifold optimization.
    • Two separate TSDF volumes are integrated: a world-frame static scene volume (excluding hand and moving-part pixels) and a canonical-frame dynamic part volume (fusing moving-part pixels transformed by (Tim)−1Tic(T_i^m)^{-1} T_i^c).
    1. Global Fragment Alignment: Geometric feature correspondences between submaps SkS_k are extracted via PREDATOR, verified with loop closure RMSE filtering, and aligned via pose graph optimization to construct a globally coherent 3D scene with interactable articulated parts.
  3. Knowl 3 — Articulation-Aware 3D Motion Clustering and Part Selection

    algorithm

    Given a set of sparse 3D point tracks τl={τl,i}i=1N\tau_l = \{\tau_{l,i}\}_{i=1}^N in the world frame and fragment articulation type σk∈{prismatic,revolute}\sigma_k \in \{\text{prismatic}, \text{revolute}\}, FunREC isolates the interacting articulated part tracks from static background and outlier tracks.

    Input: 3D point tracks τl={τl,i}i=1N\tau_l = \{\tau_{l,i}\}_{i=1}^N, visibility scores ol,io_{l,i}, 2D interacted-object masks MiobjM_i^{\text{obj}} with confidences CiobjC_i^{\text{obj}}, joint type σk\sigma_k, static threshold ϵs\epsilon_s, fitting threshold ϵf\epsilon_f
    Output: Interacted moving track set τm\tau^m, static track set τs\tau^s
    for each track τl\tau_l do
        if max⁡i,j∥τl,i−τl,j∥2<ϵs\max_{i,j} \|\tau_{l,i} - \tau_{l,j}\|_2 < \epsilon_s then
            discard τl\tau_l
        else
            Fit per-track kinematic hypothesis ϕ^l\hat{\phi}_l to τl\tau_l based on σk\sigma_k
            if σk==prismatic\sigma_k == \text{prismatic} then
                τ^l,i←Tpris(a^l,λ^l,i)τl,1\hat{\tau}_{l,i} \leftarrow T_{\text{pris}}(\hat{a}_l, \hat{\lambda}_{l,i}) \tau_{l,1}
            else
                τ^l,i←Trev(a^l,p^l,θ^l,i)τl,1\hat{\tau}_{l,i} \leftarrow T_{\text{rev}}(\hat{a}_l, \hat{p}_l, \hat{\theta}_{l,i}) \tau_{l,1}
            end if
            Compute mean fitting error el←1∣τl∣∑i∥τl,i−τ^l,i∥2e_l \leftarrow \frac{1}{|\tau_l|} \sum_{i} \|\tau_{l,i} - \hat{\tau}_{l,i}\|_2
            if el≥ϵfe_l \ge \epsilon_f then
                discard τl\tau_l
            end if
        end if
    end for
    Cluster retained tracks into clusters {γ}\{\gamma\} via HDBSCAN based on joint parameters (a^l,p^l,ϕ^l)(\hat{a}_l, \hat{p}_l, \hat{\phi}_l)
    for each cluster γ\gamma do
        sγ←∑l∈γ∑i=1Nol,i⋅Ciobj⋅I[π(τl,i)∈Miobj]s_\gamma \leftarrow \sum_{l \in \gamma} \sum_{i=1}^N o_{l,i} \cdot C_i^{\text{obj}} \cdot \mathbb{I}[\pi(\tau_{l,i}) \in M_i^{\text{obj}}]
    end for
    γ∗←arg⁡max⁡γsγ\gamma^* \leftarrow \arg\max_\gamma s_\gamma
    τm←γ∗\tau^m \leftarrow \gamma^*
    τs←⋃γ≠γ∗γ∪{static tracks}\tau^s \leftarrow \bigcup_{\gamma \ne \gamma^*} \gamma \cup \{\text{static tracks}\}
    return τm,τs\tau^m, \tau^s
  4. Knowl 4 — Pixel-Aligned Part Segmentation via Geometric Prompting of SAM2

    model/method

    To lift sparse moving 3D tracks τm\tau^m and static tracks τs\tau^s into dense, temporally continuous binary segmentation masks {Mim}i=1N\{M_i^m\}_{i=1}^N without relying on fragile direct point prompting, FunREC combines keyframe over-segmentation voting with temporal video propagation:

    1. Keyframe Over-segmentation: A uniform subset of keyframes {Iq}q=1Q\{I_q\}_{q=1}^Q is selected. SAM's automatic mask generator produces a set of disjoint region hypotheses Gq(p)∈NG_q(p) \in \mathbb{N} for every pixel pp in frame IqI_q.

    2. Track Projection and Motion Ratio Voting: 3D tracks are projected into the keyframe image plane using camera pose TqcT_q^c: π(τqm)\pi(\tau_q^m) and π(τqs)\pi(\tau_q^s). For each semantic region rr in GqG_q, the counts of interior moving and static tracks are computed as: nrm=∑p∈rI[p∈π(τqm)],nrs=∑p∈rI[p∈π(τqs)]n_r^m = \sum_{p \in r} \mathbb{I}[p \in \pi(\tau_q^m)], \quad n_r^s = \sum_{p \in r} \mathbb{I}[p \in \pi(\tau_q^s)] The motion ratio γr\gamma_r is evaluated with a small stabilizer ϵ>0\epsilon > 0: γr=nrmnrm+nrs+ϵ\gamma_r = \frac{n_r^m}{n_r^m + n_r^s + \epsilon}

    3. Keyframe Mask Formulation: Regions with γr>ηm\gamma_r > \eta_m (where ηm\eta_m is a predetermined threshold) form the keyframe articulated part mask: Mqsm(p)={1,if Gq(p)∈{r∣γr>ηm}0,otherwiseM_q^{sm}(p) = \begin{cases} 1, & \text{if } G_q(p) \in \{r \mid \gamma_r > \eta_m\} \\ 0, & \text{otherwise} \end{cases}

    4. Temporal Propagation: The keyframe masks {Mqsm}q=1Q\{M_q^{sm}\}_{q=1}^Q serve as spatial memory prompts for SAM2's video propagation module, generating pixel-accurate dynamic part masks {Mim}i=1N\{M_i^m\}_{i=1}^N across the full video fragment.

  5. Knowl 5 — Joint Optimization of Articulated Part Poses and Kinematic Parameters

    equation

    Given pairwise relative part transformations Ti→jm∈SE(3)T_{i \to j}^m \in SE(3) estimated via SupeRANSAC over 3D-3D correspondences in the moving track set τm\tau^m (weighted by joint visibility ol,imol,jmo_{l,i}^m o_{l,j}^m), the absolute part poses {Tim}i=1N\{T_i^m\}_{i=1}^N, kinematic parameters ϕm\phi^m, and loop closure confidences lijm∈[0,1]l_{ij}^m \in [0, 1] are jointly recovered by minimizing the objective: L({Tim},{lijm},ϕm)=∑i=1N−1f(Tim,Ti+1m,Ti→i+1m)+∑i,jlijmf(Tim,Tjm,Ti→jm)+μ∑i,j(lijm−1)2\mathcal{L}(\{T_i^m\}, \{l_{ij}^m\}, \phi^m) = \sum_{i=1}^{N-1} f(T_i^m, T_{i+1}^m, T_{i \to i+1}^m) + \sum_{i,j} l_{ij}^m f(T_i^m, T_j^m, T_{i \to j}^m) + \mu \sum_{i,j} (\sqrt{l_{ij}^m} - 1)^2 where μ\mu balances the loop closure regularization term, and the residual cost function is: f(Tim,Tjm,Ti→jm)=eij⊤Ωijeijf(T_i^m, T_j^m, T_{i \to j}^m) = e_{ij}^\top \Omega_{ij} e_{ij} with the Lie algebra error vector: eij=log⁡SE(3)((Tim)−1TjmTi→jm)e_{ij} = \log_{SE(3)}\left((T_i^m)^{-1} T_j^m T_{i \to j}^m\right) where Ωij\Omega_{ij} is the information matrix and log⁡SE(3):SE(3)→se(3)\log_{SE(3)}: SE(3) \to \mathfrak{se}(3) maps group transformation errors to the tangent space.

    The initial parameters ϕm\phi^m are obtained by least-squares fitting to τm\tau^m. The objective is minimized using Ceres Solver with manifold optimization, enforcing unit norm constraints on the axis a∈S2a \in \mathbb{S}^2, orthogonality for the revolute pivot a⊤p=0a^\top p = 0, and valid rotation angles θi∈S1\theta_i \in \mathbb{S}^1.

  6. Knowl 6 — Dual-TSDF Reconstruction and Functional Interactive Scene Generation

    model/method

    To reconstruct complete scenes including occluded interior volumes revealed during motion, FunREC fuses depth frames into two decoupled Truncated Signed Distance Function (TSDF) volumes:

    1. Static Background Volume (PsP^s): Integrated in the global world frame using estimated camera poses {Tic}\{T_i^c\}, masking out dynamic pixels associated with the 2D hand mask MihM_i^h and the articulated part mask MimM_i^m.

    2. Canonical Part Volume (PmP^m): Integrated in the canonical frame of the moving component. For each frame ii, the camera depth map is mapped to canonical coordinates via (Tim)−1Tic(T_i^m)^{-1} T_i^c, and only pixels inside the moving mask MimM_i^m are fused. This aggregates observations across all articulated states into a clean, occlusion-free 3D mesh PmP^m.

    3. Dynamic Compositing and Range Estimation: The complete 3D scene at any time step ii is represented as: Pi=Ps∪Tim(Pm)P_i = P^s \cup T_i^m(P^m) The operational kinematic range is extracted from the bounding extrema of the tracked state sequence: λ∈[λmin⁡,λmax⁡]\lambda \in [\lambda_{\min}, \lambda_{\max}] for prismatic joints, or θ∈[θmin⁡,θmax⁡]\theta \in [\theta_{\min}, \theta_{\max}] for revolute joints. This parameterization enables direct export to URDF and USD scene descriptions for simulation environments (e.g., Isaac Sim).

  7. Knowl 7 — RealFun4D and OmniFun4D Datasets for Functional Scene Reconstruction

    experimental setup

    To evaluate scene-scale functional reconstruction from egocentric video, two new datasets were developed:

    1. RealFun4D: Comprises 351 in-the-wild RGB-D interaction sequences captured in 60 real apartments across 4 countries using a head-mounted Azure Kinect DK (1920×10801920 \times 1080 resolution at 15 FPS). Annotations include temporal interaction intervals, textual descriptions, ground-truth 2D hand masks, 2D part masks, 3D articulation joint types/axes, and 2D part tracks.

    2. OmniFun4D: Comprises 127 photorealistic interaction trajectories across 12 OmniGibson synthetic indoor scenes. RGB-D videos and object masks are generated using NVIDIA RTX Path Tracing. Natural human head motion is simulated by applying stochastic Gaussian perturbations to camera trajectories.

    3. Evaluation Benchmark: A standardized test split of 60 interaction sequences is drawn from RealFun4D and OmniFun4D, complemented by 30 interactions across 4 articulated object categories (laptop, cabinet, safe, trash can) extracted from HOI4D.

  8. Knowl 8 — Articulated Motion Estimation Evaluation

    data/table

    Articulated motion estimation accuracy is evaluated on OmniFun4D, HOI4D, and RealFun4D. Metrics include articulation axis direction error (∘^\circ), revolute axis position error (meters), joint state error (∘^\circ for revolute, meters for prismatic), and failure rate (percentage of sequences where joint type was misidentified or processing failed).

    Method OmniFun4D HOI4D RealFun4D
    Axis (∘^\circ) Pos (m) State (∘^\circ/m) Fail (%) Axis (∘^\circ) Pos (m) State (∘^\circ/m) Fail (%) Axis (∘^\circ) Pos (m) State (∘^\circ/m) Fail (%)
    MonST3R (ICP) 52.7 / 61.5 1.14 66.2 / 0.27 11.7 67.3 / 66.1 0.45 60.8 / 0.06 0.0 54.6 / 59.1 0.68 65.9 / 0.21 0.0
    MonST3R (CoTr3) 46.8 / 58.9 1.20 45.3 / 0.18 11.7 54.7 / 51.5 0.44 44.8 / 0.07 0.0 56.8 / 52.4 0.70 57.4 / 0.22 0.0
    MonST3R (GT depth+CoTr3) 57.6 / 63.5 1.10 46.5 / 0.18 11.7 48.6 / 30.1 0.41 37.6 / 0.06 0.0 51.4 / 56.2 0.55 54.3 / 0.21 0.0
    SpatialTrackerV2 59.8 / 69.1 1.22 48.6 / 0.19 30.0 57.9 / 44.6 0.46 41.9 / 0.07 0.0 59.4 / 49.8 0.49 57.5 / 0.22 20.0
    SpatialTrackerV2 (GT depth) 48.7 / 69.3 1.10 46.8 / 0.19 30.0 54.3 / 38.8 0.43 38.7 / 0.07 0.0 60.8 / 42.6 0.58 55.6 / 0.22 20.0
    BundleSDF (GT mask) 38.2 / 55.9 0.95 23.4 / 0.20 55.0 26.3 / 24.5 0.24 9.1 / 0.07 43.3 32.0 / 52.1 0.58 15.6 / 0.22 36.7
    ArtGS 64.3 / 23.4 1.69 – / – 95.0 60.5 / 53.2 0.35 – / – 66.7 70.1 / 18.2 0.34 – / – 66.7
    FunREC (Ours) 5.3 / 5.4 0.03 5.0 / 0.02 1.7 12.4 / 1.3 0.06 9.1 / 0.02 0.0 5.6 / 5.5 0.05 8.4 / 0.03 0.0

    Values are presented as revolute / prismatic joint results. FunREC achieves 5-10×\times lower axis and position errors than baselines and reduces failure rates from 20.0%−95.0%20.0\%-95.0\% down to 0.0%−1.7%0.0\%-1.7\% across datasets.

  9. Knowl 9 — 6D Part Pose Estimation and Surface Reconstruction Performance

    data/table

    6D part pose tracking and 3D surface mesh reconstruction accuracy are evaluated across OmniFun4D, HOI4D, and RealFun4D using Average Distance of Model Points (ADD, %), Symmetric Average Distance (ADD-S, %), and Chamfer Distance (CD, cm).

    Method OmniFun4D HOI4D RealFun4D
    ADD-S (%) ↑\uparrow ADD (%) ↑\uparrow CD (cm) ↓\downarrow ADD-S (%) ↑\uparrow ADD (%) ↑\uparrow CD (cm) ↓\downarrow ADD-S (%) ↑\uparrow ADD (%) ↑\uparrow CD (cm) ↓\downarrow
    MonST3R (ICP) 14.36 9.62 29.8 29.41 19.12 6.4 13.21 9.11 13.4
    MonST3R (CoTr3) 33.58 21.94 23.6 44.72 25.98 2.1 18.64 11.35 17.6
    MonST3R (GT depth+CoTr3) 37.12 30.48 13.9 54.83 37.11 1.3 20.11 14.62 13.2
    SpatialTrackerV2 26.94 13.11 34.7 61.25 41.02 0.9 19.77 13.92 18.5
    SpatialTrackerV2 (GT depth) 29.71 15.43 9.88 60.98 40.21 0.8 22.94 15.22 13.5
    BundleSDF (GT mask) 22.84 12.37 17.1 53.12 35.22 1.4 20.88 12.91 10.6
    ArtGS – – 20.3 – – 6.8 – – 20.6
    FunREC (Ours) 78.96 71.28 3.2 79.43 69.85 0.7 75.62 68.11 6.1

    FunREC achieves more than twofold improvement in ADD-S and ADD over prior state-of-the-art pose trackers while attaining the lowest Chamfer Distance error on all datasets (3.2 cm on OmniFun4D, 0.7 cm on HOI4D, and 6.1 cm on RealFun4D).

  10. Knowl 10 — Moving Part Segmentation Performance

    data/table

    The quality of 2D moving part segmentation is quantified using mean Intersection-over-Union (mIoU, %) across all video frames in OmniFun4D, HOI4D, and RealFun4D.

    Method OmniFun4D mIoU (%) ↑\uparrow HOI4D mIoU (%) ↑\uparrow RealFun4D mIoU (%) ↑\uparrow
    MonST3R 23.6 26.8 23.7
    SpatialTrackerV2 (SAM2) 6.2 5.8 13.4
    FunREC (Ours) 77.9 76.4 74.8

    FunREC improves moving part segmentation by over +50 mIoU compared to MonST3R and SpatialTrackerV2, demonstrating accurate part delineation across synthetic, lab, and real-world environments.

Coverage note — No substantial contributed material was omitted; high-level qualitative affordance and robot interaction demonstration details are summarized within the pipeline and experimental scope.

References

  1. 1.Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, Linguang Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, Jakob Julian Engel, and Tomas Hodan. HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  2. 2.Daniel Barath. Superansac: One ransac to rule them all. arXiv preprint arXiv:2506.04803, 2025.
  3. 3.Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, et al. ARKitScenes: A Diverse Real-world Dataset for 3D Indoor Scene Understanding Using Mobile RGB-D Data. In International Conference on Neural Information Processing Systems (NeurIPS), 2021.
  4. 4.Tjark Behrens, Rene Zurbrugg, Marc Pollefeys, Zuria Bauer, and Hermann Blum. Lost & Found: Tracking Changes from Egocentric Observations in 3D Dynamic Scene Graphs. IEEE Robotics and Automation Letters (RA-L), 2025.
  5. 5.Ricardo J. G. B. Campello, Davoud Moulavi, and Joerg Sander. Density-Based Clustering Based on Hierarchical Density Estimates. In Advances in Knowledge Discovery and Data Mining, 2013.
  6. 6.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niener, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3D: Learning from RGB-D Data in Indoor Environments. In International Conference on 3d Vision (3dV), 2017.
  7. 7.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Niener. ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  8. 8.Tianyuan Dai, Josiah Wong, Yunfan Jiang, Chen Wang, Cem Gokmen, Ruohan Zhang, Jiajun Wu, and Li Fei-Fei. Automated Creation of Digital Cousins for Robust Policy Learning. In Conference on Robot Learning (CoRL), 2024.
  9. 9.Ahmad Darkhalil, Dandan Shan, Bin Zhu, Jian Ma, Amlan Kar, Richard Higgins, Sanja Fidler, David Fouhey, and Dima Damen. EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations. In International Conference on Neural Information Processing Systems (NeurIPS), 2022.
  10. 10.Google Deepmind. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities. arXiv preprint arXiv:2507.06261, 2025.
  11. 11.Alexandros Delitzas, Ayca Takmaz, Federico Tombari, Robert Sumner, Marc Pollefeys, and Francis Engelmann. SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D Scenes. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  12. 12.Johan Edstedt, Qiyu Sun, Georg Bokman, Marten Wadenback, and Michael Felsberg. RoMa: Robust Dense Feature Matching. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  13. 13.Chengshu Li et al. BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation. arXiv preprint arXiv:2403.09227, 2024.
  14. 14.Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Otmar Hilliges. ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  15. 15.Haiwen Feng, Junyi Zhang, Qianqian Wang, Yufei Ye, Pengcheng Yu, Michael J. Black, Trevor Darrell, and Angjoo Kanazawa. St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  16. 16.Pradyumn Goyal, Dmitry Petrov, Sheldon Andrews, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu, and Evangelos Kalogerakis. GEOPARD: Geometric Pretraining for Articulation Prediction in 3D Shapes. In International Conference on Computer Vision (ICCV), 2025.
  17. 17.Vladimir Guzov, Julian Chibane, Riccardo Marin, Yannan He, Yunus Saracoglu, Torsten Sattler, and Gerard Pons-Moll. Interaction Replica: Tracking human–object interaction and scene changes from human motion. In International Conference on 3d Vision (3dV), 2024.
  18. 18.Anna-Maria Halacheva, Yang Miao, Jan-Nico Zaech, Xi Wang, Luc Van Gool, and Danda Pani Paudel. Holistic Understanding of 3D Scenes as Universal Scene Description. In International Conference on Computer Vision (ICCV), 2025.
  19. 19.Adam W. Harley, Yang You, Xinglong Sun, Yang Zheng, Nikhil Raghuraman, Yunqi Gu, Sheldon Liang, Wen-Hsuan Chu, Achal Dave, Pavel Tokmakov, Suya You, Rares Ambrus, Katerina Fragkiadaki, and Leonidas J. Guibas. All-Tracker: Efficient Dense Point Tracking at High Resolution. In International Conference on Computer Vision (ICCV), 2025.
  20. 20.Nick Heppert, Muhammad Zubair Irshad, Sergey Zakharov, Katherine Liu, Rares Andrei Ambrus, Jeannette Bohg, Abhinav Valada, and Thomas Kollar. CARTO: Category and Joint Agnostic Reconstruction of ARTiculated Objects. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  21. 21.Shengyu Huang, Zan Gojcic, Mikhail Usvyatsov, and Konrad Schindler Andreas Wieser. PREDATOR: Registration of 3D Point Clouds with Low Overlap. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  22. 22.Zhening Huang, Xiaoyang Wu, Fangcheng Zhong, Hengshuang Zhao, Matthias Niener, and Joan Lasenby. LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D Scans. arXiv preprint arXiv:2507.02861, 2025.
  23. 23.Zhao Huang, Boyang Sun, Alexandros Delitzas, Jiaqi Chen, and Marc Pollefeys. React3d: Recovering articulations for interactive physical 3d scenes. IEEE Robotics and Automation Letters (RA-L), 2026.
  24. 24.Allison Janoch, Sergey Karayev, Yangqing Jia, Jonathan T. Barron, Mario Fritz, Kate Saenko, and Trevor Darrell. A Category-level 3D Object Dataset: Putting the Kinect to Work. In International Conference on Computer Vision (ICCV) Workshops, 2011.
  25. 25.Hanxiao Jiang, Yongsen Mao, Manolis Savva, and Angel X Chang. OPD: Single-view 3D Openable Part Detection. In European Conference on Computer Vision (ECCV), 2022.
  26. 26.Zhenyu Jiang, Cheng-Chun Hsu, and Yuke Zhu. Ditto: Building Digital Twins of Articulated Objects from Interaction. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  27. 27.Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co-Tracker: It is Better to Track Together. In European Conference on Computer Vision (ECCV), 2024.
  28. 28.Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co-Tracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos. In International Conference on Computer Vision (ICCV), 2025.
  29. 29.Justin Kerr, Chung Min Kim, Mingxuan Wu, Brent Yi, Qianqian Wang, Ken Goldberg, and Angjoo Kanazawa. Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction. In Conference on Robot Learning (CoRL), 2024.
  30. 30.Jeonghwan Kim, Jisoo Kim, Jeonghyeon Na, and Hanbyul Joo. Parahome: Parameterizing everyday home activities towards 3d generative modeling of human-object interactions. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  31. 31.Maxim Kolodiazhnyi, Anna Vorontsova, Anton Konushin, and Danila Rukhovich. OneFormer3D: One transformer for Unified Point Cloud Segmentation. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  32. 32.Taein Kwon, Bugra Tekin, Jan Stuhmer, Federica Bogo, and Marc Pollefeys. H2O: Two Hands Manipulating Objects for First Person Interaction Recognition. In International Conference on Computer Vision (ICCV), 2021.
  33. 33.Justin Lazarow, David Griffiths, Gefen Kohavi, Francisco Crespo, and Afshin Dehghan. Cubify Anything: Scaling Indoor 3D Object Detection. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  34. 34.Long Le, Jason Xie, William Liang, Hung-Ju Wang, Yue Yang, Yecheng Jason Ma, Kyle Vedder, Arjun Krishna, Dinesh Jayaraman, and Eric Eaton. Articulate-anything: Automatic modeling of articulated objects via a vision-language foundation model. In International Conference on Learning Representations (ICLR), 2025.
  35. 35.Jiahui Lei, Yijia Weng, Adam Harley, Leonidas Guibas, and Kostas Daniilidis. MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  36. 36.Zhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang, Linyi Jin, Vickie Ye, Angjoo Kanazawa, Aleksander Holynski, and Noah Snavely. MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  37. 37.Jiayi Liu, Ali Mahdavi-Amiri, and Manolis Savva. PARIS: Part-level Reconstruction and Motion Analysis for Articulated Objects. In International Conference on Computer Vision (ICCV), 2023.
  38. 38.Jiayi Liu, Manolis Savva, and Ali Mahdavi-Amiri. Survey on Modeling of Human-made Articulated Objects. arXiv preprint arXiv:2403.14937, 2025.
  39. 39.Xueyi Liu, Ji Zhang, Ruizhen Hu, Haibin Huang, He Wang, and Li Yi. Self-Supervised Category-Level Articulated Object Pose Estimation with Part-Level SE(3) Equivariance. In International Conference on Learning Representations (ICLR), 2023.
  40. 40.Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  41. 41.Yu Liu, Baoxiong Jia, Ruijie Lu, Chuyue Gan, Huayu Chen, Junfeng Ni, Song-Chun Zhu, and Siyuan Huang. Videoartgs: Building digital twins of articulated objects from monocular video. arXiv preprint arXiv:2509.17647, 2025.
  42. 42.Yu Liu, Baoxiong Jia, Ruijie Lu, Junfeng Ni, Song-Chun Zhu, and Siyuan Huang. Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting. In International Conference on Learning Representations (ICLR), 2025.
  43. 43.Xianzheng Ma, Yash Bhalgat, Brandon Smart, Shuai Chen, Xinghui Li, Jian Ding, Jindong Gu, Dave Zhenyu Chen, Songyou Peng, Jia-Wang Bian, Philip H Torr, Marc Pollefeys, Matthias Niener, Ian D Reid, Angel X. Chang, Iro Laina, and Victor Adrian Prisacariu. When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models, 2024.
  44. 44.Yongsen Mao, Yiming Zhang, Hanxiao Jiang, Angel X Chang, and Manolis Savva. MultiScan: Scalable RGBD Scanning for 3D Environments With Articulated Objects. In International Conference on Neural Information Processing Systems (NeurIPS), 2022.
  45. 45.Muzammal Naseer, Salman Khan, and Fatih Porikli. Indoor Scene Understanding in 2.5/3D for Autonomous Agents: A Survey. IEEE Access, 2018.
  46. 46.Tuan Ngo, Peiye Zhuang, Evangelos Kalogerakis, Chuang Gan, Sergey Tulyakov, Hsin-Ying Lee, and Chaoyang Wang. DELTA: Dense Efficient Long-Range 3D Tracking for Any Video. In International Conference on Learning Representations (ICLR), 2025.
  47. 47.NVIDIA. Isaac Sim.
  48. 48.NVIDIA. NVIDIA RTX Path Tracing, 2023.
  49. 49.Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing Hands in 3D With Transformers. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  50. 50.Weikun Peng, Jun Lv, Cewu Lu, and Manolis Savva. iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos. In International Conference on 3d Vision (3dV), 2026.
  51. 51.Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End-to-end 3d hand localization and reconstruction in-the-wild. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  52. 52.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  53. 53.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep Hugh Voting for 3D Object Detection in Point Clouds. In International Conference on Computer Vision (ICCV), 2019.
  54. 54.Frano Rajic, Haofei Xu, Marko Mihajlovic, Siyuan Li, Irem Demir, Emircan Gundogdu, Lei Ke, Sergey Prokudin, Marc Pollefeys, and Siyu Tang. Multi-view 3d point tracking. In International Conference on Computer Vision (ICCV), 2025.
  55. 55.Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Radle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollar, and Christoph Feichtenhofer. SAM 2: Segment Anything in Images and Videos. arXiv preprint arXiv:2510.11340, 2024.
  56. 56.Bokui Shen, Fei Xia, Chengshu Li, Roberto Martın-Martın, Linxi Fan, Guanzhi Wang, Claudia Perez-D’Arpino, Shyamal Buch, Sanjana Srivastava, Lyne P. Tchapmi, Micael E. Tchapmi, Kent Vainio, Josiah Wong, Li Fei-Fei, and Silvio Savarese. iGibson 1.0: A Simulation Environment for Interactive Tasks in Large Realistic Scenes. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021.
  57. 57.Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor Segmentation and Support Inference from RGBD Images. In European Conference on Computer Vision (ECCV), 2012.
  58. 58.Shuran Song, Samuel P. Lichtenberg, and Jianxiong Xiao. SUN RGB-D: A RGB-D Scene Understanding Benchmark Suite. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  59. 59.Tao Sun, Yan Hao, Shengyu Huang, Silvio Savarese, Konrad Schindler, Marc Pollefeys, and Iro Armeni. Nothing Stands Still: A Spatiotemporal Benchmark on 3D Point Cloud Registration Under Large Geometric and Temporal Change. ISPRS Journal of Photogrammetry and Remote Sensing, 2025.
  60. 60.Xiaohao Sun, Hanxiao Jiang, Manolis Savva, and Angel Chang. OPDMulti: Openable Part Detection for Multiple Objects. In International Conference on 3d Vision (3dV), 2024.
  61. 61.Johanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari, and Matthias Niessner. RIO: 3D Object Instance Re-Localization in Changing Indoor Environments. In International Conference on Computer Vision (ICCV), 2019.
  62. 62.Jikai Wang, Qifan Zhang, Yu-Wei Chao, Bowen Wen, Xiaohu Guo, and Yu Xiang. HO-Cap: A Capture System and Dataset for 3D Reconstruction and Pose Tracking of Hand-Object Interaction. In International Conference on Neural Information Processing Systems (NeurIPS), 2025.
  63. 63.Qianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li, Bharath Hariharan, Aleksander Holynski, and Noah Snavely. Tracking Everything Everywhere All at Once. In International Conference on Computer Vision (ICCV), 2023.
  64. 64.Qianqian Wang, Vickie Ye, Hang Gao, Weijia Zeng, Jake Austin, Zhengqi Li, and Angjoo Kanazawa. Shape of Motion: 4D Reconstruction from a Single Video. In International Conference on Computer Vision (ICCV), 2025.
  65. 65.Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A. Efros, and Angjoo Kanazawa. Continuous 3D Perception Model with Persistent State. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  66. 66.B Wen and Kostas E Bekris. BundleTrack: 6D Pose Tracking for Novel Objects without Instance or Category-Level 3D Models. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021.
  67. 67.Bowen Wen, Jonathan Tremblay, Valts Blukis, Stephen Tyree, Thomas Muller, Alex Evans, Dieter Fox, Jan Kautz, and Stan Birchfield. BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  68. 68.Bowen Wen, Wei Yang, Jan Kautz, and Stan Birchfield. FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  69. 69.Yijia Weng, Bowen Wen, Jonathan Tremblay, Valts Blukis, Dieter Fox, Leonidas Guibas, and Stan Birchfield. Neural Implicit Representation for Building Digital Twins of Unknown Articulated Objects. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  70. 70.Abdelrhman Werby, Martin Buechner, Adrian Roefer, Chenguang Huang, Wolfram Burgard, and Abhinav Valada. Articulated object estimation in the wild. In Conference on Robot Learning (CoRL), 2025.
  71. 71.Di Wu, Liu Liu, Zhou Linli, Anran Huang, Liangtu Song, Qiaojun Yu, Qi Wu, and Cewu Lu. Reartgs: Reconstructing and generating articulated objects via 3d gaussian splatting with geometric and motion constraints. In International Conference on Neural Information Processing Systems (NeurIPS), 2025.
  72. 72.Mingxuan Wu, Huang Huang, Justin Kerr, Chung Min Kim, Anthony Zhang, Brent Yi, and Angjoo Kanazawa. Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding. In International Conference on Computer Vision (ICCV), 2025.
  73. 73.Hongchi Xia, Entong Su, Marius Memmel, Arhan Jain, Raymond Yu, Numfor Mbiziwo-Tiapo, Ali Farhadi, Abhishek Gupta, Shenlong Wang, and Wei-Chiu Ma. Drawer: Digital Reconstruction and Articulation with Environment Realism. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  74. 74.Jianxiong Xiao, Andrew Owens, and Antonio Torralba. SUN3D: A Database of Big Spaces Reconstructed Using SfM and Object Labels. In International Conference on Computer Vision (ICCV), 2013.
  75. 75.Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou. SpatialTracker: Tracking Any 2D Pixels in 3D Space. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  76. 76.Yuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev, Yuri Makarov, Bingyi Kang, Xing Zhu, Hujun Bao, Yujun Shen, and Xiaowei Zhou. Spatialtrackerv2: Advancing 3d point tracking with explicit camera motion. In International Conference on Computer Vision (ICCV), 2025.
  77. 77.Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Niener, and Angela Dai. ScanNet++: A High-fidelity Dataset of 3D Indoor Scenes. In International Conference on Computer Vision (ICCV), 2023.
  78. 78.Huangyue Yu, Baoxiong Jia, Yixin Chen, Yandan Yang, Puhao Li, Rongpeng Su, Jiaxin Li, Qing Li, Wei Liang, Zhu Song-Chun, Tengyu Liu, and Siyuan Huang. METASCENES: Towards Automated Replica Creation for Real-world 3D Scans. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  79. 79.Chengbo Yuan, Geng Chen, Li Yi, and Yang Gao. Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos. In International Conference on Computer Vision (ICCV), 2025.
  80. 80.Bowei Zhang, Lei Ke, Adam W Harley, and Katerina Fragkiadaki. TAPIP3D: Tracking Any Point in Persistent 3D Geometry. In International Conference on Neural Information Processing Systems (NeurIPS), 2025.
  81. 81.Daiwei Zhang, Gengyan Li, Jiajie Li, Mickael Bressieux, Otmar Hilliges, Marc Pollefeys, Luc Van Gool, and Xi Wang. EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting. In International Conference on 3d Vision (3dV), 2025.
  82. 82.Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming-Hsuan Yang. MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion. In International Conference on Learning Representations (ICLR), 2025.
  83. 83.Mandi Zhao, Yijia Weng, Dominik Bauer, and Shuran Song. Real2Code: Reconstruct Articulated Objects via Code Generation. In International Conference on Learning Representations (ICLR), 2025.

Citation

MLA
Delitzas, A., et al. “FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos”. arXiv, 2026, http://arxiv.org/abs/2604.05621v2.
APA
Delitzas, A., Zhang, C., Gavryushin, A., Mario, T. D., Sun, B., Dabral, R., Guibas, L., Theobalt, C., Pollefeys, M., Engelmann, F., & Barath, D. (2026). FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos. arXiv. http://arxiv.org/abs/2604.05621v2
Chicago
Delitzas, A., C. Zhang, A. Gavryushin, et al. 2026. “FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos”. arXiv. http://arxiv.org/abs/2604.05621v2.
Harvard
Delitzas, A. et al. (2026) “FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.05621v2.
Vancouver
1. Delitzas A, Zhang C, Gavryushin A, et al (2026) FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos. arXiv

BibTeX

@article{delitzas2026funrec,
  title = {FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos},
  author = {Delitzas, Alexandros and Zhang, Chenyangguang and Gavryushin, Alexey and Mario, Tommaso Di and Sun, Boyang and Dabral, Rishabh and Guibas, Leonidas and Theobalt, Christian and Pollefeys, Marc and Engelmann, Francis and Barath, Daniel},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.05621v2},
  eprint = {2604.05621}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/