gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction
Zerui ChenShizhe ChenCordelia SchmidIvan Laptev
Proposes a geometry-driven signed distance function framework that reconstructs high-fidelity 3D hands and manipulated objects from monocular RGB inputs by aligning neural implicit shapes with articulated kinematic chains and spatio-temporal features.
Accurate 3D reconstruction of human hands and manipulated objects is critical for emerging applications in virtual reality, robotics, and human-computer interaction. While multi-view or depth cameras are commonly used to capture these interactions, real-world deployment frequently demands reconstructing detailed shapes from standard, single-view 2D color images (monocular RGB). Existing techniques either rely on rigid parametric templates that lack fine surface detail or use flexible neural implicit surface models—specifically signed distance functions—that fail to capture underlying skeletal geometry and articulation.
The article develops and evaluates a framework called geometry-driven signed distance functions (gSDF). The primary objective is to demonstrate that integrating explicit structural pose guidance into implicit surface representations enables highly accurate, joint 3D reconstruction of hands and unknown manipulated objects from monocular images.
To achieve this, the approach first predicts sparse 3D hand joint locations and object centers from visual data. It then applies inverse kinematics to derive the full articulated kinematic chains for all finger joints. These poses form canonical reference frames that generate structured kinematic features for 3D query points, disentangling pose estimation from shape reconstruction. In addition, the framework projects query points onto 2D image planes to sample local visual features and uses a spatial-temporal transformer to aggregate context across neighboring video frames, mitigating issues like motion blur and hand-object occlusion. The method is validated on the synthetic ObMan benchmark (over 87,000 training meshes and 6,285 test samples) and the real-world DexYCB video dataset (nearly 30,000 training samples and 5,928 test samples).
The evaluation yields several key findings. First, aligning signed distance functions to full finger kinematic chains significantly improves accuracy, reducing hand surface reconstruction error (Chamfer distance) by 7.8% compared to aligning only the wrist, and by over 12% compared to using no pose priors. Second, incorporating estimated hand poses into object reconstruction reduces object surface error by more than 11% compared to using object translation alone. Third, the full framework establishes a new state of the art on both benchmarks, outperforming existing techniques on the synthetic dataset by 17.6% on hand error and 7.1% on object error, and on real-world video data by 12.2% on hand error and 14.4% on object error. Finally, architectural ablations show that using an asymmetric image backbone—separating hand pose estimation from shared object pose and shape features—achieves superior performance over fully unified or fully split alternatives.
These findings indicate that implicit shape modeling performs substantially better when guided by explicit kinematic structures, proving that high-fidelity hand-object capture is achievable from low-cost, monocular camera inputs without pre-scanned 3D object models. This enhances performance and reduces hardware costs for downstream systems in robotic manipulation and immersive digital environments.
Organizations developing spatial computing, interactive systems, or robotic grasping pipelines should consider adopting kinematically guided implicit frameworks to improve visual interaction fidelity. When implementing such pipelines, engineering teams should prioritize asymmetric neural network architectures and multi-frame temporal attention modules to ensure robustness against motion artifacts. Because the system assumes default limb proportions and focuses primarily on single-hand interactions, future initiatives should test deployments across broader demographic variations and two-hand manipulation scenarios before wide-scale operational deployment.
- Paper: DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Jeong Joon Park et al. (2019). It introduces continuous signed distance functions (SDFs) for learning 3D implicit shapes, which provides the foundational surface representation that gSDF augments with kinematic hand geometry.
- Paper: PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization, Shunsuke Saito et al. (2019). It establishes pixel-aligned implicit function modeling from monocular RGB inputs, laying essential groundwork for aligning visual features with 3D coordinate queries.
- Paper: ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis, Lixin Yang et al. (2022). It explores articulated 3D hand-object pose estimation from single RGB images, addressing key pose priors and interaction dynamics leveraged in gSDF's geometry guidance.
- Paper: Expressive Body Capture: 3D Hands, Face, and Body From a Single Image, Georgios Pavlakos et al. (2019). It develops articulated parametric models for capturing hands and bodies from images, which underpin the kinematic pose representations used to structure SDFs.
- Paper: PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes, Yu Xiang et al. (2017). It details 6D object pose estimation from RGB images in cluttered interaction contexts, supplying techniques crucial for joint hand-object alignment.
- Paper: HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos, Prithviraj Banerjee et al. (2025). It provides a large-scale egocentric multi-view video benchmark for 3D hand-object tracking that can be used to evaluate and extend monocular reconstruction methods like gSDF.
- Paper: ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation, Zicong Fan et al. (2023). It introduces a dynamic dataset and interaction benchmark of bimanual dexterous hand-object manipulation, offering a natural domain to extend kinematic-guided implicit reconstruction.
