HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from Video
Zicong FanMaria ParelliMaria Eleni KadoglouXu ChenMuhammed KocabasMichael J. BlackOtmar Hilliges
Presents HOLD, the first category-agnostic approach to jointly reconstruct 3D articulated hands and arbitrary manipulated objects from monocular video using compositional neural implicit models and physical interaction constraints without requiring 3D annotations or pre-scanned object templates.
Capturing realistic three-dimensional (3D) representations of human hands interacting with physical objects is essential for virtual reality, robotics, and behavioral modeling. However, existing computer vision approaches face significant hurdles: they typically require pre-scanned 3D geometric templates of the objects, rely on constrained multi-view camera rigs, assume rigid non-articulated hands, or depend heavily on supervised models limited to a small set of pre-trained object classes. These dependencies severely restrict their practical use in unconstrained real-world environments.
The article demonstrates that high-quality, category-agnostic 3D surfaces of both articulated hands and arbitrary manipulated objects can be jointly reconstructed directly from standard monocular (single-camera) video recordings without requiring any pre-scanned object templates, prior category knowledge, or 3D training annotations.
The researchers developed an approach named HOLD (Hand and Object reconstruction by Leveraging interaction constraints in three Dimensions). The framework initializes hand and object poses using standard 2D hand regression and classical structure-from-motion techniques. It then constructs a compositional implicit neural representation that simultaneously models the articulated hand, the object, and dynamic background elements. Hand and object poses are iteratively refined by enforcing physical contact and spatial alignment constraints, after which the neural model is fully trained to produce fine geometric detail. The approach was validated quantitatively on benchmark laboratory datasets (HO3D-v3) and qualitatively across newly collected real-world video sequences featuring diverse lighting, indoor and outdoor settings, and both static and moving first-person (egocentric) views.
The evaluation yielded several key findings. First, HOLD significantly outperformed existing state-of-the-art baselines in object surface accuracy and spatial positioning: on benchmark data, it achieved a Chamfer distance of 0.4 cm² compared to 3.8–4.3 cm² for baselines, an F-score of 96.5% compared to 68.8–75.8%, and cut hand pose errors down to 24.2 mm. Second, while prior category-dependent models suffered major performance degradation on novel objects (dropping from 83.5% to 57.8% F-score), HOLD maintained consistently high accuracy (above 95% F-score) across both familiar and completely unseen items. Third, ablation analyses confirmed that jointly modeling the hand and object provides critical complementary cues; omitting hand modeling led to severe object artifacts such as holes at grasp points, while omitting contact-based pose refinement resulted in large spatial misalignments.
These findings establish that high-fidelity 3D interaction capture can be achieved from consumer-grade monocular video (such as a smartphone camera) without expensive 3D scanning equipment or restrictive training datasets. This substantially lowers the cost, hardware requirements, and complexity of digitizing human interactions, making scalable deployment feasible for consumer augmented reality, spatial computing, and scalable robot imitation learning.
Organizations seeking to implement scalable 3D interaction capture should adopt joint hand-object modeling architectures and leverage contact-based constraints rather than treating hand tracking and object reconstruction independently. Future development efforts should focus on integrating detector-free structure-from-motion to better handle thin or featureless items, incorporating generative 2D priors to hallucinate unobserved object surfaces, and adopting faster rendering primitives (such as Gaussian splatting) to reduce computational overhead.
The findings are supported by strong benchmark metrics and convincing qualitative demonstrations across diverse in-the-wild video feeds. Nonetheless, decision-makers should recognize current limitations: the initialization pipeline struggles with textureless or extremely thin objects where structure-from-motion fails, fully unobserved object regions cannot be reconstructed from visual data alone, and full-sequence neural optimization currently requires substantial graphics processing time (approximately 10 hours on high-end hardware).
- Paper: gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction, Zerui Chen et al. (2023). This paper establishes the core paradigm of combining articulated kinematic chains with implicit signed distance functions for joint 3D hand and unknown-object reconstruction.
- Paper: ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation, Zicong Fan et al. (2023). This work introduces benchmark data and contact modeling formulations for dynamic hand-object interactions that foundational 3D reconstruction methods build upon.
- Paper: Volume Rendering of Neural Implicit Surfaces, Lior Yariv et al. (2021). This paper establishes the neural implicit surface volume rendering framework (VolSDF) that enables category-agnostic 3D geometry reconstruction from 2D views.
- Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). This foundational work introduces deformable neural radiance fields for reconstructing non-rigidly moving subjects from monocular video sequences.
- Paper: ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis, Lixin Yang et al. (2022). This paper provides essential background on handling data scarcity and contact constraints in articulated 3D hand-object pose estimation.
- Paper: Expressive Body Capture: 3D Hands, Face, and Body From a Single Image, Georgios Pavlakos et al. (2019). This paper establishes the expressive SMPL-X parametric representation for articulated hands that underpins modern hand mesh recovery pipelines.
- Paper: HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos, Prithviraj Banerjee et al. (2025). This work extends hand-object tracking research by providing a large-scale 3D egocentric benchmark with multi-view streams and ground-truth poses to evaluate interaction modeling in real-world environments.
