gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction

Zerui ChenShizhe ChenCordelia SchmidIvan Laptev

article2023CVPR64 citations

Proposes a geometry-driven signed distance function framework that reconstructs high-fidelity 3D hands and manipulated objects from monocular RGB inputs by aligning neural implicit shapes with articulated kinematic chains and spatio-temporal features.

Listen

Accurate 3D reconstruction of human hands and manipulated objects is critical for emerging applications in virtual reality, robotics, and human-computer interaction. While multi-view or depth cameras are commonly used to capture these interactions, real-world deployment frequently demands reconstructing detailed shapes from standard, single-view 2D color images (monocular RGB). Existing techniques either rely on rigid parametric templates that lack fine surface detail or use flexible neural implicit surface models—specifically signed distance functions—that fail to capture underlying skeletal geometry and articulation.

The article develops and evaluates a framework called geometry-driven signed distance functions (gSDF). The primary objective is to demonstrate that integrating explicit structural pose guidance into implicit surface representations enables highly accurate, joint 3D reconstruction of hands and unknown manipulated objects from monocular images.

To achieve this, the approach first predicts sparse 3D hand joint locations and object centers from visual data. It then applies inverse kinematics to derive the full articulated kinematic chains for all finger joints. These poses form canonical reference frames that generate structured kinematic features for 3D query points, disentangling pose estimation from shape reconstruction. In addition, the framework projects query points onto 2D image planes to sample local visual features and uses a spatial-temporal transformer to aggregate context across neighboring video frames, mitigating issues like motion blur and hand-object occlusion. The method is validated on the synthetic ObMan benchmark (over 87,000 training meshes and 6,285 test samples) and the real-world DexYCB video dataset (nearly 30,000 training samples and 5,928 test samples).

The evaluation yields several key findings. First, aligning signed distance functions to full finger kinematic chains significantly improves accuracy, reducing hand surface reconstruction error (Chamfer distance) by 7.8% compared to aligning only the wrist, and by over 12% compared to using no pose priors. Second, incorporating estimated hand poses into object reconstruction reduces object surface error by more than 11% compared to using object translation alone. Third, the full framework establishes a new state of the art on both benchmarks, outperforming existing techniques on the synthetic dataset by 17.6% on hand error and 7.1% on object error, and on real-world video data by 12.2% on hand error and 14.4% on object error. Finally, architectural ablations show that using an asymmetric image backbone—separating hand pose estimation from shared object pose and shape features—achieves superior performance over fully unified or fully split alternatives.

These findings indicate that implicit shape modeling performs substantially better when guided by explicit kinematic structures, proving that high-fidelity hand-object capture is achievable from low-cost, monocular camera inputs without pre-scanned 3D object models. This enhances performance and reduces hardware costs for downstream systems in robotic manipulation and immersive digital environments.

Organizations developing spatial computing, interactive systems, or robotic grasping pipelines should consider adopting kinematically guided implicit frameworks to improve visual interaction fidelity. When implementing such pipelines, engineering teams should prioritize asymmetric neural network architectures and multi-frame temporal attention modules to ensure robustness against motion artifacts. Because the system assumes default limb proportions and focuses primarily on single-hand interactions, future initiatives should test deployments across broader demographic variations and two-hand manipulation scenarios before wide-scale operational deployment.

Cover for gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction

Abstract

Signed distance functions (SDFs) is an attractive framework that has recently shown promising results for 3D shape reconstruction from images. SDFs seamlessly generalize to different shape resolutions and topologies but lack explicit modelling of the underlying 3D geometry. In this work, we exploit the hand structure and use it as guidance for SDF-based shape reconstruction. In particular, we address reconstruction of hands and manipulated objects from monocular RGB images. To this end, we estimate poses of hands and objects and use them to guide 3D reconstruction. More specifically, we predict kinematic chains of pose transformations and align SDFs with highly-articulated hand poses. We improve the visual features of 3D points with geometry alignment and further leverage temporal information to enhance the robustness to occlusion and motion blurs. We conduct extensive experiments on the challenging ObMan and DexYCB benchmarks and demonstrate significant improvements of the proposed method over the state of the art.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Overview of gSDF
  • 3.2. Kinematic Feature Encoding
  • 3.3. Visual Feature Encoding
  • 3.4. Image Backbone Sharing Strategy
  • 3.5. Training
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Evaluation metrics
  • 4.3. Implementation details
  • 4.4. Ablation studies
  • 4.5. Comparison with state of the art
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Geometry-Driven Signed Distance Function Framework for 3D Hand-Object Reconstruction

    model/method

    The geometry-driven Signed Distance Function (gSDF) framework jointly reconstructs continuous 3D hand and unknown hand-held object surfaces from a single monocular RGB image It∈RH×W×3I_t \in \mathbb{R}^{H \times W \times 3} or video sequence. The 3D shapes of the hand and object are represented implicitly by two signed distance decoders fhf_h and fof_o mapping a 3D query point x∈R3x \in \mathbb{R}^3 to signed distance values:

    SDFhand(x)=fh([ev;eh])\text{SDF}_{hand}(x) = f_h([e_v; e_h]) SDFobject(x)=fo([ev;eo])\text{SDF}_{object}(x) = f_o([e_v; e_o])

    where:

    • eh∈R51e_h \in \mathbb{R}^{51} is a hand kinematic feature capturing canonical coordinates of xx transformed relative to all articulated hand joint frames along the kinematic chain.
    • eo∈R72e_o \in \mathbb{R}^{72} is an object kinematic feature capturing coordinates of xx relative to the predicted object center, 21 hand joints, and the hand wrist canonical frame.
    • ev∈Rde_v \in \mathbb{R}^d is a local visual feature sampled by projecting xx onto the 2D image plane, refined across video frames via a spatio-temporal transformer.
    • [⋅;⋅][\cdot ; \cdot] denotes feature concatenation.

    At test time, 3D meshes for the hand and object are extracted from the zero-level sets of SDFhand\text{SDF}_{hand} and SDFobject\text{SDF}_{object} using the Marching Cubes algorithm.

  2. Knowl 2 — Hand Kinematic Feature Formulation via Kinematic Chain Alignment

    equation

    To encode strong pose priors and disentangle hand pose from hand shape reconstruction, a 3D query point x∈R3x \in \mathbb{R}^3 is transformed into canonical coordinate frames defined by the hand kinematic chain.

    From predicted 3D joint coordinates ψh∈R21×3\psi_h \in \mathbb{R}^{21 \times 3} (obtained via soft-argmax on 64×64×6464 \times 64 \times 64 volumetric heatmaps), inverse kinematics calculates joint rotations θh,i∈R3\theta_{h,i} \in \mathbb{R}^3 and translations ϕh,i∈R3\phi_{h,i} \in \mathbb{R}^3 for 16 MANO joints relative to their parent joints. The forward transformation matrix Gh,i∈R4×4G_{h,i} \in \mathbb{R}^{4 \times 4} from the root frame to joint ii is:

    Gh,i=∏j∈A(i)[exp⁡(θh,j)ϕh,j0T1]G_{h, i} = \prod_{j \in A(i)} \begin{bmatrix} \exp(\theta_{h, j}) & \phi_{h, j} \\ \mathbf{0}^T & 1 \end{bmatrix}

    where A(i)A(i) denotes the ordered sequence of ancestor joints of joint ii, and exp⁡(⋅)\exp(\cdot) converts axis-angle rotation vectors into 3×33 \times 3 rotation matrices via Rodrigues' formula.

    The local kinematic feature eh,i∈R3e_{h,i} \in \mathbb{R}^3 for joint ii transforms xx into joint canonical coordinates:

    eh,i=Tp(x,θh,i,ϕh,i)=H~(Gh,i−1⋅H(x))e_{h,i} = T_p(x, \theta_{h,i}, \phi_{h,i}) = \widetilde{H}(G_{h,i}^{-1} \cdot H(x))

    where H(x)=[xT,1]T∈R4H(x) = [x^T, 1]^T \in \mathbb{R}^4 converts Euclidean coordinates to homogeneous coordinates, and H~([uT,w]T)=u/w∈R3\widetilde{H}([u^T, w]^T) = u/w \in \mathbb{R}^3 projects homogeneous coordinates back to Euclidean space.

    The complete hand kinematic embedding eh∈R51e_h \in \mathbb{R}^{51} concatenates the original query point and all 16 local joint canonical coordinates:

    eh=[x,eh,1,eh,2,…,eh,16]e_h = [x, e_{h,1}, e_{h,2}, \dots, e_{h,16}]

  3. Knowl 3 — Object Kinematic Feature Formulation

    equation

    To guide implicit object surface reconstruction with spatial priors and hand-grasp context, a 3D query point x∈R3x \in \mathbb{R}^3 is mapped into canonical frames defined by the predicted object center, individual hand joints, and the hand wrist.

    Let ψo∈R3\psi_o \in \mathbb{R}^3 denote the predicted object center position relative to the hand wrist, ψh,i∈R3\psi_{h,i} \in \mathbb{R}^3 (i∈{1,…,21}i \in \{1, \dots, 21\}) denote predicted 3D locations of the 21 hand joints, and Gh,1∈R4×4G_{h,1} \in \mathbb{R}^{4 \times 4} denote the transformation matrix of the hand wrist root frame.

    The object kinematic feature combines three transformations:

    1. Object-centered translation: xoc=Tt(x,ψo)=x−ψo∈R3x_{oc} = T_t(x, \psi_o) = x - \psi_o \in \mathbb{R}^3.
    2. Joint-relative translations for all 21 hand joints: eo,i=Tt(x,ψh,i)=x−ψh,i∈R3e_{o,i} = T_t(x, \psi_{h,i}) = x - \psi_{h,i} \in \mathbb{R}^3 for i∈{1,…,21}i \in \{1, \dots, 21\}.
    3. Wrist-canonical transformation: xow=Tp(x,θh,1,ϕh,1)=H~(Gh,1−1⋅H(x))∈R3x_{ow} = T_p(x, \theta_{h,1}, \phi_{h,1}) = \widetilde{H}(G_{h,1}^{-1} \cdot H(x)) \in \mathbb{R}^3, where H(⋅)H(\cdot) converts to homogeneous coordinates and H~(⋅)\widetilde{H}(\cdot) maps back to Euclidean coordinates.

    The aggregated object kinematic feature eo∈R72e_o \in \mathbb{R}^{72} is formed by concatenation:

    eo=[x,xoc,eo,1,eo,2,…,eo,21,xow]e_o = [x, x_{oc}, e_{o,1}, e_{o,2}, \dots, e_{o,21}, x_{ow}]

  4. Knowl 4 — Geometry-Aligned and Spatio-Temporally Enhanced Visual Feature Extraction

    model/method

    Rather than relying on globally pooled image features, gSDF derives point-specific visual features aligned with 3D geometry and aggregates temporal context across video frames:

    1. Geometry-Aligned Feature Sampling: A 3D query point x∈R3x \in \mathbb{R}^3 is projected onto the 2D image plane as x^\hat{x} using the camera projection matrix. Bilinear interpolation samples a local visual feature ev∈Rde_v \in \mathbb{R}^d from a 2D feature map vt∈R16×16×dv_t \in \mathbb{R}^{16 \times 16 \times d} at coordinate x^\hat{x}.

    2. Spatio-Temporal Feature Enhancement: When processing video sequences, feature maps vt−1r,vtr,vt+1r∈R16×16×dv_{t-1}^r, v_t^r, v_{t+1}^r \in \mathbb{R}^{16 \times 16 \times d} are extracted from consecutive frames using a ResNet backbone. The spatial-temporal tokens across the 3 frames (3×16×163 \times 16 \times 16 tokens) are fed into a 16-layer transformer with 4 attention heads. The refined output for time tt is reshaped into vt∈R16×16×dv_t \in \mathbb{R}^{16 \times 16 \times d} before bilinear sampling, enhancing robustness to occlusion and motion blur.

  5. Knowl 5 — Asymmetric Backbone Architecture and Two-Stage Training Scheme

    model/method

    gSDF separates the network backbones and trains them in two stages to resolve multi-task competition between sparse pose estimation and dense shape prediction:

    1. Backbone Partitioning: A standalone ResNet-18 backbone extracts features for the 3D hand pose predictor. A second, shared ResNet-18 backbone is used jointly for the object pose predictor and the SDF visual feature encoder.

    2. Stage 1 (Hand Joint Pretraining): The hand pose predictor is trained to predict 3D joint positions ψh∈R21×3\psi_h \in \mathbb{R}^{21 \times 3} using an ℓ2\ell_2 coordinate loss Lhp\mathcal{L}_{hp} and an ordinal depth loss Lord\mathcal{L}_{ord} over 20 sampled virtual viewpoints n⃗∈R3\vec{n} \in \mathbb{R}^3:

    Lhp=121∑i=121∥ψh,i−ψ^h,i∥22\mathcal{L}_{hp} = \frac{1}{21} \sum_{i=1}^{21} \|\psi_{h,i} - \hat{\psi}_{h,i}\|_2^2 Lord=∑j=221∑i=1j−11i,jord⋅∣(ψh,i−ψh,j)⋅n⃗∣\mathcal{L}_{ord} = \sum_{j=2}^{21} \sum_{i=1}^{j-1} \mathbf{1}^{ord}_{i,j} \cdot |(\psi_{h,i} - \psi_{h,j}) \cdot \vec{n}|

    where ψ^h,i\hat{\psi}_{h,i} is the ground-truth joint location, and 1i,jord=1\mathbf{1}^{ord}_{i,j} = 1 when the predicted relative depth order of joints ii and jj along n⃗\vec{n} contradicts the ground truth. Once trained, the hand pose predictor parameters are frozen.

    1. Stage 2 (Joint Pose and SDF Optimization): All remaining components are trained end-to-end using the composite shape loss Lshape\mathcal{L}_{shape}:

    Lshape=Lop+0.5Lhsdf+0.5Losdf\mathcal{L}_{shape} = \mathcal{L}_{op} + 0.5 \mathcal{L}_{hsdf} + 0.5 \mathcal{L}_{osdf}

    where Lop=∥ψo−ψ^o∥22\mathcal{L}_{op} = \|\psi_o - \hat{\psi}_o\|_2^2 supervises the object center ψo\psi_o, and the SDF decoders are supervised with ℓ1\ell_1 distance losses on 1000 sampled points per mesh (500 inside, 500 outside):

    Lhsdf=∥SDFhand−SDF^hand∥1,Losdf=∥SDFobj−SDF^obj∥1\mathcal{L}_{hsdf} = \|\text{SDF}_{hand} - \widehat{\text{SDF}}_{hand}\|_1, \quad \mathcal{L}_{osdf} = \|\text{SDF}_{obj} - \widehat{\text{SDF}}_{obj}\|_1

  6. Knowl 6 — 3D Hand-Object Reconstruction Performance on ObMan and DexYCB Benchmarks

    data/table

    gSDF was evaluated on the synthetic ObMan benchmark (6,285 test samples) and the real-world video DexYCB benchmark (s0 split, 5,928 test samples). Metrics include median Hand Chamfer Distance (CDhCD_h, cm2\text{cm}^2), Hand F-score at 1mm (FSh@1FS_h@1) and 5mm (FSh@5FS_h@5), Object Chamfer Distance (CDoCD_o, cm2\text{cm}^2), Object F-score at 5mm (FSo@5FS_o@5) and 10mm (FSo@10FS_o@10), Hand Joint Error (EhE_h, cm), and Object Center Error (EoE_o, cm).

    Dataset / Method CDh↓CD_h \downarrow FSh@1↑FS_h@1 \uparrow FSh@5↑FS_h@5 \uparrow CDo↓CD_o \downarrow FSo@5↑FS_o@5 \uparrow FSo@10↑FS_o@10 \uparrow Eh↓E_h \downarrow Eo↓E_o \downarrow
    ObMan
    Hasson et al. 0.415 0.138 0.751 3.60 0.359 0.590 1.13 -
    Karunratanakul et al. 0.261 - - 6.80 - - - -
    Ye et al. - - - - 0.420 0.630 - -
    Chen et al. (AlignSDF) 0.136 0.302 0.913 3.38 0.404 0.636 1.27 3.29
    gSDF (Ours) 0.112 0.332 0.935 3.14 0.438 0.660 0.93 3.43
    DexYCB
    Hasson et al. 0.537 0.115 0.647 1.94 0.383 0.642 1.67 -
    Karunratanakul et al. 0.364 0.154 0.764 2.06 0.392 0.660 - -
    Chen et al. (AlignSDF) 0.358 0.162 0.767 1.83 0.410 0.679 1.58 1.78
    Chen et al. (AlignSDF-2BB) 0.344 0.167 0.776 1.81 0.413 0.687 1.57 1.93
    gSDF (Ours) 0.302 0.177 0.801 1.55 0.437 0.709 1.44 1.96

    gSDF outperforms existing methods across both benchmarks, improving CDhCD_h by 17.6% on ObMan and by 12.2% on DexYCB over prior state-of-the-art results.

  7. Knowl 7 — Ablation of Hand and Object Kinematic Feature Configurations

    data/table

    Ablation experiments on DexYCB evaluate the effect of pose alignment levels within hand kinematic features (KhK_h) and object kinematic features (KoK_o) using global visual feature V1V_1:

    Hand Feature Wrist only All joints CDh↓(cm2)CD_h \downarrow (\text{cm}^2) FSh@1↑FS_h@1 \uparrow FSh@5↑FS_h@5 \uparrow
    Kh1K_h^1 (No pose prior) ×\times ×\times 0.364 0.154 0.764
    Kh2K_h^2 (Wrist only) ✓\checkmark ×\times 0.344 0.167 0.776
    Kh3K_h^3 (All joints) ×\times ✓\checkmark 0.317 0.171 0.788
    Object Feature Obj pose Hand pose CDo↓(cm2)CD_o \downarrow (\text{cm}^2) FSo@5↑FS_o@5 \uparrow FSo@10↑FS_o@10 \uparrow
    Ko1K_o^1 (No pose prior) ×\times ×\times 2.06 0.392 0.660
    Ko2K_o^2 (Obj center only) ✓\checkmark ×\times 1.93 0.396 0.668
    Ko3K_o^3 (Obj center + Hand joints) ✓\checkmark ✓\checkmark 1.71 0.418 0.689

    Aligning with all hand joints (Kh3K_h^3) reduces hand Chamfer distance by 7.8% relative to wrist-only alignment (Kh2K_h^2). Conditioning object SDF on hand joint positions (Ko3K_o^3) improves object Chamfer distance by 11.4% relative to object center alignment alone (Ko2K_o^2).

  8. Knowl 8 — Ablation of Visual Feature Encodings and Spatio-Temporal Refinement

    data/table

    Ablation experiments on DexYCB evaluate the impact of local point sampling and spatio-temporal transformer refinement when combined with full kinematic features (Kh3,Ko3K_h^3, K_o^3):

    Visual Feature Configuration CDh↓CD_h \downarrow FSh@1↑FS_h@1 \uparrow FSh@5↑FS_h@5 \uparrow CDo↓CD_o \downarrow FSo@5↑FS_o@5 \uparrow FSo@10↑FS_o@10 \uparrow Eh↓E_h \downarrow Eo↓E_o \downarrow
    V1V_1 (Global visual feature) 0.317 0.171 0.788 1.71 0.418 0.689 1.44 1.91
    V2V_2 (Geometry-aligned local) 0.310 0.172 0.795 1.71 0.426 0.694 1.44 1.98
    V3V_3 (Local + Spatial Transf.) 0.304 0.174 0.797 1.60 0.434 0.703 1.44 1.94
    V4V_4 (Local + Spatio-Temporal Transf.) 0.302 0.177 0.801 1.55 0.437 0.709 1.44 1.96

    Replacing global visual pooling (V1V_1) with local projected features (V2V_2) improves hand accuracy from CDh=0.317CD_h = 0.317 to 0.310 cm20.310\ \text{cm}^2. Adding the spatio-temporal transformer across video frames (V4V_4) produces large improvements on object shape accuracy, lowering CDoCD_o from 1.711.71 to 1.55 cm21.55\ \text{cm}^2.

  9. Knowl 9 — Ablation of Backbone Sharing Strategies

    data/table

    An ablation study on DexYCB compares three architectural backbone sharing strategies across the hand pose predictor, object pose predictor, and SDF feature encoder (using features V1,Kh3,Ko3V_1, K_h^3, K_o^3):

    Backbone Strategy CDh↓CD_h \downarrow FSh@1↑FS_h@1 \uparrow FSh@5↑FS_h@5 \uparrow CDo↓CD_o \downarrow FSo@5↑FS_o@5 \uparrow FSo@10↑FS_o@10 \uparrow Eh↓E_h \downarrow Eo↓E_o \downarrow
    Single backbone (shared by all) 0.411 0.148 0.741 1.88 0.402 0.674 1.72 1.83
    Symmetric (Pose shared, SDF separate) 0.324 0.168 0.779 1.84 0.405 0.672 1.46 1.93
    Asymmetric (Hand separate, Obj+SDF shared) 0.317 0.171 0.788 1.71 0.418 0.689 1.44 1.91

    The single backbone architecture underperforms due to task interference between pose estimation and SDF shape learning. The asymmetric strategy, which decouples hand pose regression into a dedicated backbone while sharing representations between object pose estimation and the SDF encoder, attains the best overall reconstruction quality.

Coverage note — None was omitted; all primary methodology, equations, training procedures, benchmark comparisons, and ablation studies from the main paper are fully represented.

References

  1. 1.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lućić, and Cordelia Schmid. ViViT: A video vision transformer. In ICCV, 2021. 5
  2. 2.Seungryul Baek, Kwang In Kim, and Tae-Kyun Kim. Pushing the envelope for RGB-based dense 3D hand pose estimation via neural rendering. In CVPR, 2019. 2
  3. 3.Luca Ballan, Aparna Taneja, Jurgen Gall, Luc Van Gool, and Marc Pollefeys. Motion capture of hands in action using discriminative salient points. In ECCV, 2012. 2
  4. 4.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, 2021. 5
  5. 5.Adnane Boukhayma, Rodrigo de Bem, and Philip HS Torr. 3D hand shape and pose from images in the wild. In CVPR, 2019. 2
  6. 6.Romain Bregier. Deep regression on manifolds: a 3D rotation case study. In 3DV, 2021. 3, 4
  7. 7.Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa, and Jitendra Malik. Reconstructing hand-object interactions in the wild. In ICCV, 2021. 1, 2
  8. 8.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. ShapeNet: An information-rich 3D model repository. arXiv preprint arXiv:1512.03012, 2015. 6
  9. 9.Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, et al. DexYCB: A benchmark for capturing hand grasping of objects. In CVPR, 2021. 1, 2, 6
  10. 10.Xingyu Chen, Yufeng Liu, Chongyang Ma, Jianlong Chang, Huayan Wang, Tian Chen, Xiaoyan Guo, Pengfei Wan, and Wen Zheng. Camera-space hand mesh recovery via semantic aggregation and adaptive 2D-1D registration. In CVPR, 2021. 2
  11. 11.Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. AlignSDF: Pose-Aligned signed distance fields for hand-object reconstruction. In ECCV, 2022. 1, 2, 3, 4, 5, 6, 7, 8
  12. 12.Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In CVPR, 2019. 2
  13. 13.Enric Corona, Tomas Hodan, Minh Vo, Francesc Moreno-Noguer, Chris Sweeney, Richard Newcombe, and Lingni Ma. LISA: Learning implicit shape and appearance of hands. In CVPR, 2022. 2
  14. 14.Thibault Groueix, Matthew Fisher, Vladimir G Kim, Bryan C Russell, and Mathieu Aubry. A papier-mache approach to learning 3D surface generation. In CVPR, 2018. 2
  15. 15.Henning Hamer, Juergen Gall, Thibaut Weise, and Luc Van Gool. An object-dependent hand pose prior from sparse training data. In CVPR, 2010. 2
  16. 16.Henning Hamer, Konrad Schindler, Esther Koller-Meier, and Luc Van Gool. Tracking a hand manipulating an object. In ICCV, 2009. 2
  17. 17.Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D annotation of hand and object poses. In CVPR, 2020. 1, 2
  18. 18.Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint Transformer: Solving joint identification in challenging hands and object interactions for accurate 3D pose estimation. In CVPR, 2022. 1, 2
  19. 19.Yana Hasson, Bugra Tekin, Federica Bogo, Ivan Laptev, Marc Pollefeys, and Cordelia Schmid. Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction. In CVPR, 2020. 1, 2, 4
  20. 20.Yana Hasson, Gul Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object reconstruction from RGB videos. In 3DV, 2021. 2
  21. 21.Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. In CVPR, 2019. 1, 2, 4, 6, 8
  22. 22.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 4, 6
  23. 23.Tony Heap and David Hogg. Towards 3D hand tracking using a deformable model. In FG, 1996. 2
  24. 24.Umar Iqbal, Pavlo Molchanov, Thomas Breuel Juergen Gall, and Jan Kautz. Hand pose estimation via latent 2.5D heatmap regression. In ECCV, 2018. 1, 2
  25. 25.Korrawe Karunratanakul, Adrian Spurr, Zicong Fan, Otmar Hilliges, and Siyu Tang. A skeleton-driven neural occupancy representation for articulated hands. In 3DV, 2021. 2
  26. 26.Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasping Field: Learning implicit representations for human grasps. In 3DV, 2020. 1, 2, 3, 4, 6, 7, 8
  27. 27.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 6
  28. 28.Nikos Kolotouros, Georgios Pavlakos, Michael J Black, and Kostas Daniilidis. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. In ICCV, 2019. 3, 4
  29. 29.Dominik Kulon, Riza Alp Guler, I. Kokkinos, M. Bronstein, and S. Zafeiriou. Weakly-supervised mesh-convolutional hand reconstruction in the wild. In CVPR, 2020. 2
  30. 30.Dominik Kulon, Haoyang Wang, Riza Alp Guler, Michael M. Bronstein, and Stefanos Zafeiriou. Single image 3D hand reconstruction with mesh convolutions. In BMVC, 2019. 2
  31. 31.Vincent Lepetit. Recent advances in 3D object and hand pose estimation. arXiv preprint arXiv:2006.05927, 2020. 2
  32. 32.Mengcheng Li, Liang An, Hongwen Zhang, Lianpeng Wu, Feng Chen, Tao Yu, and Yebin Liu. Interacting attention graph for single image two-hand reconstruction. In CVPR, 2022. 2
  33. 33.William E Lorensen and Harvey E Cline. Marching Cubes: A high resolution 3D surface construction algorithm. TOG, 1987. 3
  34. 34.Jun Lv, Wenqiang Xu, Lixin Yang, Sucheng Qian, Chongzhao Mao, and Cewu Lu. HandTailor: Towards high-precision monocular 3D hand recovery. In BMVC, 2021. 2
  35. 35.Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3D human pose estimation in the wild using improved CNN supervision. In 3DV, 2017. 6
  36. 36.Hao Meng, Sheng Jin, Wentao Liu, Chen Qian, Mengxiang Lin, Wanli Ouyang, and Ping Luo. 3D interacting hand pose estimation by hand de-occlusion and removal. In ECCV, 2022. 2
  37. 37.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy Networks: Learning 3D reconstruction in function space. In CVPR, 2019. 2
  38. 38.Gyeongsik Moon, Ju Yong Chang, and Kyoung Mu Lee. V2V-PoseNet: Voxel-to-voxel prediction network for accurate 3D hand and human pose estimation from a single depth map. In CVPR, 2018. 1, 2, 4
  39. 39.Franziska Mueller, Florian Bernard, Oleksandr Sotnychenko, Dushyant Mehta, Srinath Sridhar, Dan Casas, and Christian Theobalt. Ganerated hands for real-time 3D hand tracking from monocular RGB. In CVPR, 2018. 2
  40. 40.Franziska Mueller, Micah Davis, Florian Bernard, Oleksandr Sotnychenko, Mickeal Verschoor, Miguel A Otaduy, Dan Casas, and Christian Theobalt. Real-time pose and shape reconstruction of two interacting hands with a single depth camera. TOG, 2019. 2
  41. 41.Iason Oikonomidis, Nikolaos Kyriazis, and Antonis A Argyros. Full DOF tracking of a hand interacting with an object by modeling occlusions and physical constraints. In ICCV, 2011. 2
  42. 42.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. DeepSDF: Learning continuous signed distance functions for shape representation. In CVPR, 2019. 2, 6
  43. 43.Georgios Pavlakos, Xiaowei Zhou, and Kostas Daniilidis. Ordinal depth supervision for 3D human pose estimation. In CVPR, 2018. 5
  44. 44.Georgios Pavlakos, Xiaowei Zhou, Konstantinos G Derpanis, and Kostas Daniilidis. Coarse-to-fine volumetric prediction for single-image 3D human pose. In CVPR, 2017. 4
  45. 45.James M Rehg and Takeo Kanade. Visual tracking of high DOF articulated structures: an application to human hand tracking. In ECCV, 1994. 2
  46. 46.Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied Hands: Modeling and capturing hands and bodies together. TOG, 2017. 1, 2
  47. 47.Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Angjoo Kanazawa, and Hao Li. PiFu: Pixel-aligned implicit function for high-resolution clothed human digitization. In ICCV, 2019. 3, 4
  48. 48.Adrian Spurr, Aneesh Dahiya, Xi Wang, Xucong Zhang, and Otmar Hilliges. Self-supervised 3D hand pose estimation from monocular RGB via contrastive learning. In ICCV, 2021. 2
  49. 49.Srinath Sridhar, Franziska Mueller, Michael Zollhofer, Dan Casas, Antti Oulasvirta, and Christian Theobalt. Real-time joint tracking of a hand manipulating an object from RGB-D input. In ECCV, 2016. 2
  50. 50.Xiao Sun, Bin Xiao, Fangyin Wei, Shuang Liang, and Yichen Wei. Integral human pose regression. In ECCV, 2018. 2, 4
  51. 51.Danhang Tang, Hyung Jin Chang, Alykhan Tejani, and Tae-Kyun Kim. Latent regression forest: Structured estimation of 3D articulated hand posture. In CVPR, 2014. 1, 2
  52. 52.Maxim Tatarchenko, Stephan R Richter, Rene Ranftl, Zhuwen Li, Vladlen Koltun, and Thomas Brox. What do single-view 3D reconstruction networks learn? In CVPR, 2019. 6
  53. 53.Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+O: Unified egocentric recognition of 3D hand-object poses and interactions. In CVPR, 2019. 1, 2
  54. 54.Tze Ho Elden Tse, Kwang In Kim, Ales Leonardis, and Hyung Jin Chang. Collaborative learning for hand and object reconstruction with attention-guided graph convolution. In CVPR, 2022. 1, 2
  55. 55.Aggeliki Tsoli and Antonis A Argyros. Joint 3D tracking of a deformable object in interaction with a hand. In ECCV, 2018. 2
  56. 56.Dimitrios Tzionas and Juergen Gall. 3D object reconstruction from hand-object interactions. In ICCV, 2015. 2
  57. 57.Jiayi Wang, Franziska Mueller, Florian Bernard, Suzanne Sorli, Oleksandr Sotnychenko, Neng Qian, Miguel A Otaduy, Dan Casas, and Christian Theobalt. RGB2Hands: Real-time tracking of 3D hand interactions from monocular RGB video. TOG, 2020. 2
  58. 58.Yangang Wang, Jianyuan Min, Jianjie Zhang, Yebin Liu, Feng Xu, Qionghai Dai, and Jinxiang Chai. Video-based hand manipulation capture through composite motion control. TOG, 2013. 2
  59. 59.Fu Xiong, Boshen Zhang, Yang Xiao, Zhiguo Cao, Taidong Yu, Joey Tianyi Zhou, and Junsong Yuan. A2J: Anchor-to-joint regression network for 3D articulated pose estimation from a single depth image. In ICCV, 2019. 2
  60. 60.Lixin Yang, Kailin Li, Xinyu Zhan, Jun Lv, Wenqiang Xu, Jiefeng Li, and Cewu Lu. ArtiBoost: Boosting articulated 3D hand-object pose estimation via online exploration and synthesis. In CVPR, 2022. 2, 6
  61. 61.Lixin Yang, Kailin Li, Xinyu Zhan, Fei Wu, Anran Xu, Liu Liu, and Cewu Lu. OakInk: A large-scale knowledge repository for understanding hand-object interaction. In CVPR, 2022. 1, 2
  62. 62.Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu. CPF: Learning a contact potential field to model the hand-object interaction. In ICCV, 2021. 1, 2, 6
  63. 63.Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What’s in your hands? 3D reconstruction of generic objects in hands. In CVPR, 2022. 2, 3, 4, 6, 8
  64. 64.Frank Yu, Mathieu Salzmann, Pascal Fua, and Helge Rhodin. PCLs: Geometry-aware neural reconstruction of 3D pose with perspective crop layers. In CVPR, 2021. 6
  65. 65.Shanxin Yuan, Guillermo Garcia-Hernando, Bjorn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan, Xinghao Chen, Guijin Wang, Fan Yang, Kai Akiyama, Yang Wu, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iason Oikonomidis, Antonis Argyros, and Tae-Kyun Kim. Depth-based 3D hand pose estimation: From current achievements to future goals. In CVPR, June 2018. 2
  66. 66.Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the continuity of rotation representations in neural networks. In CVPR, 2019. 3, 4
  67. 67.Christian Zimmermann and Thomas Brox. Learning to estimate 3D hand pose from single RGB images. In ICCV, 2017. 1, 2

Citation

MLA
Chen, Z., et al. “gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction”. arXiv, 2023, http://arxiv.org/abs/2304.11970v1.
APA
Chen, Z., Chen, S., Schmid, C., & Laptev, I. (2023). gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction. arXiv. http://arxiv.org/abs/2304.11970v1
Chicago
Chen, Z., S. Chen, C. Schmid, and I. Laptev. 2023. “gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction”. arXiv. http://arxiv.org/abs/2304.11970v1.
Harvard
Chen, Z. et al. (2023) “gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.11970v1.
Vancouver
1. Chen Z, Chen S, Schmid C, Laptev I (2023) gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction. arXiv

BibTeX

@article{chen2023gsdf,
  title = {gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction},
  author = {Chen, Zerui and Chen, Shizhe and Schmid, Cordelia and Laptev, Ivan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.11970v1},
  eprint = {2304.11970}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE