HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from Video

Zicong FanMaria ParelliMaria Eleni KadoglouXu ChenMuhammed KocabasMichael J. BlackOtmar Hilliges

article2024CVPR69 citationsCVPR 2024 Highlight

Presents HOLD, the first category-agnostic approach to jointly reconstruct 3D articulated hands and arbitrary manipulated objects from monocular video using compositional neural implicit models and physical interaction constraints without requiring 3D annotations or pre-scanned object templates.

Listen

Capturing realistic three-dimensional (3D) representations of human hands interacting with physical objects is essential for virtual reality, robotics, and behavioral modeling. However, existing computer vision approaches face significant hurdles: they typically require pre-scanned 3D geometric templates of the objects, rely on constrained multi-view camera rigs, assume rigid non-articulated hands, or depend heavily on supervised models limited to a small set of pre-trained object classes. These dependencies severely restrict their practical use in unconstrained real-world environments.

The article demonstrates that high-quality, category-agnostic 3D surfaces of both articulated hands and arbitrary manipulated objects can be jointly reconstructed directly from standard monocular (single-camera) video recordings without requiring any pre-scanned object templates, prior category knowledge, or 3D training annotations.

The researchers developed an approach named HOLD (Hand and Object reconstruction by Leveraging interaction constraints in three Dimensions). The framework initializes hand and object poses using standard 2D hand regression and classical structure-from-motion techniques. It then constructs a compositional implicit neural representation that simultaneously models the articulated hand, the object, and dynamic background elements. Hand and object poses are iteratively refined by enforcing physical contact and spatial alignment constraints, after which the neural model is fully trained to produce fine geometric detail. The approach was validated quantitatively on benchmark laboratory datasets (HO3D-v3) and qualitatively across newly collected real-world video sequences featuring diverse lighting, indoor and outdoor settings, and both static and moving first-person (egocentric) views.

The evaluation yielded several key findings. First, HOLD significantly outperformed existing state-of-the-art baselines in object surface accuracy and spatial positioning: on benchmark data, it achieved a Chamfer distance of 0.4 cm² compared to 3.8–4.3 cm² for baselines, an F-score of 96.5% compared to 68.8–75.8%, and cut hand pose errors down to 24.2 mm. Second, while prior category-dependent models suffered major performance degradation on novel objects (dropping from 83.5% to 57.8% F-score), HOLD maintained consistently high accuracy (above 95% F-score) across both familiar and completely unseen items. Third, ablation analyses confirmed that jointly modeling the hand and object provides critical complementary cues; omitting hand modeling led to severe object artifacts such as holes at grasp points, while omitting contact-based pose refinement resulted in large spatial misalignments.

These findings establish that high-fidelity 3D interaction capture can be achieved from consumer-grade monocular video (such as a smartphone camera) without expensive 3D scanning equipment or restrictive training datasets. This substantially lowers the cost, hardware requirements, and complexity of digitizing human interactions, making scalable deployment feasible for consumer augmented reality, spatial computing, and scalable robot imitation learning.

Organizations seeking to implement scalable 3D interaction capture should adopt joint hand-object modeling architectures and leverage contact-based constraints rather than treating hand tracking and object reconstruction independently. Future development efforts should focus on integrating detector-free structure-from-motion to better handle thin or featureless items, incorporating generative 2D priors to hallucinate unobserved object surfaces, and adopting faster rendering primitives (such as Gaussian splatting) to reduce computational overhead.

The findings are supported by strong benchmark metrics and convincing qualitative demonstrations across diverse in-the-wild video feeds. Nonetheless, decision-makers should recognize current limitations: the initialization pipeline struggles with textureless or extremely thin objects where structure-from-motion fails, fully unobserved object regions cannot be reconstructed from visual data alone, and full-sequence neural optimization currently requires substantial graphics processing time (approximately 10 hours on high-end hardware).

Cover for HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from Video

Abstract

Since humans interact with diverse objects every day, the holistic 3D capture of these interactions is important to understand and model human behaviour. However, most existing methods for hand-object reconstruction from RGB either assume pre-scanned object templates or heavily rely on limited 3D hand-object data, restricting their ability to scale and generalize to more unconstrained interaction settings. To address this, we introduce HOLD – the first category-agnostic method that reconstructs an articulated hand and an object jointly from a monocular interaction video. We develop a compositional articulated implicit model that can reconstruct disentangled 3D hands and objects from 2D images. We also further incorporate hand-object constraints to improve hand-object poses and consequently the reconstruction quality. Our method does not rely on any 3D hand-object annotations while significantly outperforming fully-supervised baselines in both in-the-lab and challenging in-the-wild settings. Moreover, we qualitatively show its robustness in reconstructing from in-the-wild videos. See here for code, data, models, and updates.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method: HOLD
  • 3.1. Pose initialization
  • 3.2. HOLD-Net training
  • 3.2.1 HOLD-Net
  • 3.2.2 Training losses
  • 3.3. Pose refinement
  • 3.4. Final training
  • 4. Experiments
  • 4.1. State-of-the-art comparison
  • 4.2. Ablation
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Category-agnostic joint reconstruction pipeline

    model/method

    HOLD reconstructs an articulated hand and an unknown rigid object jointly from a monocular RGB interaction video. It requires neither a pre-scanned object template, a known object category, nor 3D hand–object annotations. The output consists of disentangled 3D hand and object surfaces, canonical geometries shared across video frames, and frame-specific hand and object poses.

    The pipeline has four stages: (1) initialize hand and object poses from off-the-shelf estimators; (2) briefly train a compositional implicit model, HOLD-Net, using the noisy poses; (3) refine the hand–object poses using contact and image-alignment constraints; and (4) retrain HOLD-Net with the refined poses to obtain the final geometries. The central modeling assumption is that hand shape, object shape, and their physical interaction provide complementary evidence for resolving monocular shape and depth ambiguities.

  2. Knowl 2 — Pose initialization without an object template

    model/method

    For every video frame, HOLD initializes the hand with an off-the-shelf hand pose estimator, obtaining MANO pose parameters theta in R^{48}, hand shape parameters beta, and hand translation t_h in R^3. The 48-dimensional pose includes the global hand orientation.

    Because no category-level object pose estimator is assumed, HOLD first removes the hand and other pixels using an off-the-shelf segmentation model. It applies HLoc-based structure-from-motion to the resulting object-only images, producing an object point cloud and per-frame object rotation R_o in SO(3) and translation t_o in R^3. Structure-from-motion determines the object only up to an unknown scale s in R.

    HOLD aligns the hand and object coordinate systems and estimates ss by optimizing the per-frame translations {t_h,t_o}, global hand shape beta, and object scale ss. The optimization encourages hand–object contact while requiring the projected hand joints and reconstructed object points to agree with their original 2D image projections. This initialization is category-agnostic but depends on structure-from-motion producing a usable object point cloud.

  3. Knowl 3 — Compositional canonical implicit representations

    model/method

    HOLD-Net represents the hand and object with separate canonical signed-distance and RGB texture fields, allowing their surfaces to be disentangled while sharing observations across frames. The hand field is an MLP with learnable parameters psi_h:

    f_h: R^3 to R times R^3, qquad x mapsto (d_h(x),c_h(x)),

    where x in R^3 is a canonical-space point, d_h(x) in R is its signed distance to the hand surface, and c_h(x) in R^3 is its RGB color. The hand field is conditioned on MANO pose theta and translation t_h. A point x' in R^3 in the posed observation space is mapped to canonical space by inverse linear blend skinning:

    x=(∑i=1nbwi(x′)Bi)−1x′,x=\left(\sum_{i=1}^{n_b}w_i(x')\mathbf B_i\right)^{-1}x',

    where nbn_b is the number of MANO bones, B_i is the bone transformation computed from theta, and wi(x′)w_i(x') is the skinning weight assigned to the deformed point. The weights are obtained from the distance-weighted skinning weights of the KK nearest MANO vertices.

    The object field has learnable parameters psi_o and a per-frame 32-dimensional latent code z_o in R^{32}:

    f_o: R^3 times R^{32} to R times R^3, qquad (x,z_o) mapsto(d_o(x,z_o),c_o(x,z_o)).

    For object scale ss, rotation R_o, translation t_o, and observation-space point x', the canonical object point is

    x=(sRo)−1(x′−to).x=(s\mathbf R_o)^{-1}(x'-t_o).

    The latent code zoz_o accounts for appearance changes caused by pose, occlusion, and shadows. A separate background field with learnable parameters psi_b models the dynamic scene and partially visible body regions; it takes a point xx, viewing direction v in R^3, and a distinct per-frame code z_b in R^{32}, and predicts signed distance and RGB color. The hand and object canonical geometries are time-independent across frames, whereas the object appearance and background are allowed to vary.

  4. Knowl 4 — Compositional volumetric rendering of hand, object, and background

    model/method

    For each camera ray rr with camera center oo and viewing direction vv, HOLD-Net independently samples points for the hand and object using error-bounded sampling. Hand samples are mapped to canonical space by inverse skinning, object samples by the rigid transformation, and background samples by the background parameterization. Signed distances are converted to volume densities sigma using the cumulative distribution function of a scaled Laplace distribution with positive learnable scale parameters.

    The hand and object samples are merged and sorted by depth. If the merged samples are indexed by i=1, ldots,2n, with density sigma_i, RGB color cic_i, and inter-sample distance delta_i, the foreground color is

    C_F(r)=\sum_{i=1}^{2n}\tau_i c_i, qquad \tau_i=\exp\left(-\sum_{j<i}\sigma_j\delta_j\right)\left(1-\exp(-\sigma_i\delta_i)\right).

    Here tau_i is the rendering weight of sample ii. The foreground occupancy probability is MF(r)=∑i=12nτiM_F(r)=\sum_{i=1}^{2n}\tau_i, and if CB(r)C_B(r) is the independently rendered background color, the final pixel color is

    C(r)=CF(r)+(1−MF(r))CB(r).C(r)=C_F(r)+(1-M_F(r))C_B(r).

    Accumulating only hand samples or only object samples produces amodal hand and object mask probabilities Mh(r)M_h(r) and Mo(r)M_o(r). Replacing each sample color with a one-hot class vector similarly produces per-pixel probabilities for hand, object, and background. These separate rendered quantities enable RGB supervision, semantic disentanglement, and sparsity constraints even when one surface occludes the other.

  5. Knowl 5 — Multi-term self-supervised training objective

    equation

    HOLD-Net is trained from RGB video, rendered masks, and geometric priors rather than 3D hand–object annotations. The optimized parameters include the hand, object, and background field parameters {\psi_h,\psi_o,\psi_b}; frame-specific pose, translation, and latent parameters {\theta,t_h,R_o,t_o,z_o,z_b}; and global hand-shape and object-scale parameters {\beta,s}.

    For rays rr sampled from the input images, the RGB term compares rendered color C(r)C(r) with observed color C^(r)\hat C(r), and the segmentation term compares the rendered three-class probability S(r)S(r) with the one-hot segmentation label S^(r)\hat S(r):

    \mathcal L_{\mathrm{rgb}}=\sum_r\lVert C(r)-\hat C(r)\rVert, qquad \mathcal L_{\mathrm{segm}}=\sum_r\lVert S(r)-\hat S(r)\rVert.

    For canonical hand points x∈Xx\in\mathcal X, where X is a random set of points sampled uniformly and near surfaces, the hand signed-distance prediction dh(x)d_h(x) is regularized toward the signed distance SDF⁡MANO(x)\operatorname{SDF}_{\mathrm{MANO}}(x) of a Loop-subdivided MANO mesh:

    Lsdf=∑x∈X∥dh(x)−SDF⁡MANO(x)∥.\mathcal L_{\mathrm{sdf}}=\sum_{x\in\mathcal X}\lVert d_h(x)-\operatorname{SDF}_{\mathrm{MANO}}(x)\rVert.

    An eikonal loss regularizes the gradients of the canonical hand and object signed-distance fields. A sparsity loss suppresses hand occupancy on rays rr in a set Fh\mathcal F_h whose closest point is farther than a threshold from the MANO mesh, and object occupancy on rays rr in a corresponding set Fo\mathcal F_o defined from a periodically extracted object mesh:

    Lsparse=∑r∈Fh∥Mh(r)∥+∑r∈Fo∥Mo(r)∥.\mathcal L_{\mathrm{sparse}}=\sum_{r\in\mathcal F_h}\lVert M_h(r)\rVert+\sum_{r\in\mathcal F_o}\lVert M_o(r)\rVert.

    The total objective is

    L=Lrgb+λsegmLsegm+λsdfLsdf+λsparseLsparse+λeikonalLeikonal,\mathcal L=\mathcal L_{\mathrm{rgb}}+\lambda_{\mathrm{segm}}\mathcal L_{\mathrm{segm}}+\lambda_{\mathrm{sdf}}\mathcal L_{\mathrm{sdf}}+\lambda_{\mathrm{sparse}}\mathcal L_{\mathrm{sparse}}+\lambda_{\mathrm{eikonal}}\mathcal L_{\mathrm{eikonal}},

    where each lambda is a loss weight. Because automatic segmentation is noisy, HOLD gradually decreases lambda_{\mathrm{segm}} during training while increasing lambda_{\mathrm{sdf}} and lambda_{\mathrm{sparse}}.

  6. Knowl 6 — Contact-based pose refinement and final reconstruction

    algorithm

    HOLD refines pose after a short pretraining stage has produced a coarse object surface. Jointly learning poses and shapes from the beginning is inefficient because each frame receives training signals only when it is sampled, so the coarse learned object mesh is instead used for a separate pose-refinement stage.

    The refinement optimizes hand translation tht_h, object rotation RoR_o, object translation tot_o, hand shape beta, and object scale ss using the extracted object mesh and the MANO hand model. Let VtipsiV_{\mathrm{tips}}^i be frequently contacting hand vertices and VojV_o^j be object vertices. The contact loss is

    Lcontact=∑imin⁡j∥Vtipsi−Voj∥.\mathcal L_{\mathrm{contact}}=\sum_i\min_j\lVert V_{\mathrm{tips}}^i-V_o^j\rVert.

    Soft Rasterizer renders amodal hand and object masks, and an occlusion-aware mask loss aligns those projections with the segmentation masks. The contact term reduces relative hand–object depth ambiguity, while the mask term improves 2D image alignment.

    After refinement, HOLD fully retrains the implicit fields with the refined poses. The shape and background networks and the object and background latent codes are reinitialized and trained from scratch to prevent artifacts caused by inaccurate poses during pretraining. The final training uses the same loss formulation as pretraining but runs for twice as many epochs; in the reported implementation, pretraining uses 100 epochs and final training uses 200 epochs.

  7. Knowl 7 — Evaluation protocol and implementation

    experimental setup

    HOLD is evaluated on HO3D-v3 and on a newly captured HOLD dataset. HO3D-v3 contains RGB videos of articulated hands manipulating rigid objects, with MANO parameters and 6D object poses. Because the HO3D test annotations are unavailable, evaluation uses two annotated training sequences per object: one selected to match prior work and one random sequence in which the hand and object remain visible. Banana and scissors are excluded from the main evaluation because structure-from-motion fails on their textureless or thin structures.

    The HOLD dataset contains household-object videos recorded with an iPhone 14 in indoor and outdoor scenes, using both moving first-person views and static third-person views under varied lighting. Videos are downsampled by retaining every tenth frame. Qualitative results show that HOLD reconstructs diverse objects and articulated hands in both static and moving egocentric views despite changing backgrounds and illumination.

    Hand accuracy is measured by root-relative mean-per-joint error (MPJPE) in millimeters. Object shape is measured by Chamfer distance (CD) in squared centimeters and F-score at 5 mm and 10 mm, reported as F5 and F10 percentages. For canonical object-shape evaluation, iterative closest point alignment to the ground-truth mesh allows scale, rotation, and translation. Hand-relative Chamfer distance CDhCD_h is computed after subtracting the predicted hand-root position from each object mesh, thereby measuring the object’s shape and position relative to the hand.

    Each sequence is optimized with Adam using 10 randomly sampled images per iteration and gradient clipping. Initial training takes approximately 10 hours for 100 epochs on an A100 GPU; final training uses 200 epochs. Hand and object masks are obtained with SAM-track, initialized by point prompts in the first frame.

  8. Knowl 8 — Performance on annotated hand–object reconstruction

    data/table

    On the HO3D evaluation sequences, HOLD achieves lower hand-pose error and object reconstruction error, higher local-shape F-score, and more accurate hand-relative object placement than the compared methods. The comparison is particularly notable because HOMan assumes a ground-truth object template, iHOI uses 3D annotations of test objects during training, whereas HOLD and DiffHOI do not use such test-object information.

    Could not parse LaTeX table

    HOLD also produces visibly finer geometry and more accurate poses than the baselines, including details such as mug handles and car frames. It remains qualitatively stable across in-the-lab and in-the-wild videos, varied backgrounds and lighting, and moving egocentric cameras.

  9. Knowl 9 — Category generalization and canonical object scanning

    data/table

    HOLD maintains its reconstruction quality on object categories outside the training distribution of DiffHOI, whereas DiffHOI degrades substantially on those unseen categories. HOLD also outperforms an in-hand object-scanning method on canonical object shape and local details.

    For objects belonging to DiffHOI’s training categories, the reported results are:

    Could not parse LaTeX table

    For canonical object scanning, comparison with Hampali et al. uses the released object point clouds because their implementation is unavailable:

    Could not parse LaTeX table

    These results support the claim that HOLD does not merely retrieve a category prior: it reconstructs instance-specific object details, including objects outside the categories represented by competing training data.

  10. Knowl 10 — Ablation evidence for joint modeling and contact refinement

    empirical result

    Ablations show that both hand modeling and contact-based pose refinement materially improve reconstruction. In the no-hand variant, the hand is masked from every frame and only an object network is trained. This produces a small degradation in object metrics and can leave a hole near the grasping region because the object field must explain pixels that actually belong to the occluding hand. Omitting pose refinement leaves the hand and object spatially misaligned under monocular depth ambiguity and greatly worsens their hand-relative placement.

    Could not parse LaTeX table

    The full model therefore improves not only the isolated object surface but also the relative 3D arrangement of the hand and object.

  11. Knowl 11 — Stated limitations

    limitation

    HOLD’s category-agnostic reconstruction remains limited in three situations. First, thin or textureless objects can defeat the detector-based structure-from-motion used for object-pose initialization; detector-free structure-from-motion is suggested as a possible remedy. Second, rarely observed object regions may be poorly reconstructed because the method relies on RGB supervision and receives little evidence for those surfaces; stronger shape priors could help. Third, sequence-specific neural implicit training is computationally expensive, and faster scene representations are identified as a possible way to reduce training time.

Coverage note — No substantial contributed material was deliberately omitted; implementation details and qualitative findings are included with the main method and evaluation knowls.

References

  1. 1.Adnane Boukhayma, Rodrigo de Bem, and Philip H. S. Torr. 3D hand shape and pose from images in the wild. In Computer Vision and Pattern Recognition (CVPR), 2019. 2
  2. 2.Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa, and Jitendra Malik. Reconstructing hand-object interactions in the wild. In International Conference on Computer Vision (ICCV), 2021. 1, 2
  3. 3.Xu Chen, Zijian Dong, Jie Song, Andreas Geiger, and Otmar Hilliges. Category level object pose estimation via neural analysis-by-synthesis. In European Conference on Computer Vision (ECCV). Springer, 2020. 3
  4. 4.Xingyu Chen, Baoyuan Wang, and Heung-Yeung Shum. Hand avatar: Free-pose hand animation and rendering from monocular video. In Computer Vision and Pattern Recognition (CVPR), 2023. 2
  5. 5.Zerui Chen, Shizhe Chen, Cordelia Schmid, and Ivan Laptev. gSDF: Geometry-driven signed distance functions for 3D hand-object reconstruction. In Computer Vision and Pattern Recognition (CVPR), 2023. 2, 6
  6. 6.Yangming Cheng, Liulei Li, Yuanyou Xu, Xiaodi Li, Zongxin Yang, Wenguan Wang, and Yi Yang. Segment and track anything. arXiv preprint arXiv:2305.06558, 2023. 6
  7. 7.Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. D-Grasp: Physically plausible dynamic grasp synthesis for hand-object interactions. In Computer Vision and Pattern Recognition (CVPR), 2022. 1
  8. 8.Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gregory Rogez. GanHand: Predicting human grasp affordances in multi-object scenes. In Computer Vision and Pattern Recognition (CVPR), 2020. 2
  9. 9.Markos Diomataris, Nikos Athanasiou, Omid Taheri, Xi Wang, Otmar Hilliges, and Michael J. Black. WANDR: Intention-guided human motion generation. In Computer Vision and Pattern Recognition (CVPR), 2024. 1
  10. 10.Enes Duran, Muhammed Kocabas, Vasileios Choutas, Zicong Fan, and Michael J. Black. HMP: Hand motion priors for pose and shape estimation from video. In Winter Conference on Applications of Computer Vision (WACV), 2024. 2
  11. 11.Zicong Fan, Adrian Spurr, Muhammed Kocabas, Siyu Tang, Michael J. Black, and Otmar Hilliges. Learning to disambiguate strongly interacting hands via probabilistic per-pixel part segmentation. In International Conference on 3D Vision (3DV), 2021. 2
  12. 12.Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Otmar Hilliges. ARCTIC: A dataset for dexterous bimanual hand-object manipulation. In Proceedings IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 1, 2
  13. 13.Zicong Fan, Takehiko Ohkawa, Linlin Yang, Nie Lin, Zhishan Zhou, Shihao Zhou, Jiajun Liang, Zhong Gao, Xunyang Zhang, Xue Zhang, Fei Li, Liu Zheng, Feng Lu, Karim Abou Zeid, Bastian Leibe, Jeongwan On, Seungryul Baek, Aditya Prakash, Saurabh Gupta, Kun He, Yoichi Sato, Otmar Hilliges, Hyung Jin Chang, and Angela Yao. Benchmarks and challenges in pose estimation for egocentric hand interactions with objects. arXiv preprint arXiv: 2403.16428, 2024. 2
  14. 14.Qichen Fu, Xingyu Liu, Ran Xu, Juan Carlos Niebles, and Kris M. Kitani. Deformer: Dynamic fusion transformer for robust hand pose estimation. In International Conference on Computer Vision (ICCV), 2023. 2
  15. 15.Patrick Grady, Chengcheng Tang, Christopher D. Twigg, Minh Vo, Samarth Brahmbhatt, and Charles C. Kemp. ContactOpt: Optimizing contact to improve grasps. In Computer Vision and Pattern Recognition (CVPR), 2021. 2
  16. 16.Amos Gropp, Lior Yariv, Niv Haim, Matan Atzmon, and Yaron Lipman. Implicit geometric regularization for learning shapes. In International Conference on Machine Learning (ICML), 2020. 5
  17. 17.Chen Guo, Tianjian Jiang, Xu Chen, Jie Song, and Otmar Hilliges. Vid2Avatar: 3D avatar reconstruction from videos in the wild via self-supervised scene decomposition. In Computer Vision and Pattern Recognition (CVPR), 2023. 3, 4
  18. 18.Zhiyang Guo, Wengang Zhou, Min Wang, Li Li, and Houqiang Li. HandNeRF: Neural radiance fields for animatable interacting hands. In Computer Vision and Pattern Recognition (CVPR), 2023. 2
  19. 19.Shreyas Hampali. 3D Pose and Shape Estimation of Objects and Hands in Challenging Scenarios. PhD thesis, TU Graz, 2023. 6
  20. 20.Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D annotation of hand and object poses. In Computer Vision and Pattern Recognition (CVPR), 2020. 6
  21. 21.Shreyas Hampali, Tomas Hodan, Luan Tran, Lingni Ma, Cem Keskin, and Vincent Lepetit. In-hand 3D object scanning from an RGB sequence. Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3, 6, 7
  22. 22.Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J. Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. In Computer Vision and Pattern Recognition (CVPR), 2019. 1, 2, 5
  23. 23.Yana Hasson, Bugra Tekin, Federica Bogo, Ivan Laptev, Marc Pollefeys, and Cordelia Schmid. Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction. In Computer Vision and Pattern Recognition (CVPR), 2020. 1
  24. 24.Yana Hasson, Gul Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object reconstruction from RGB videos. In International Conference on 3D Vision (3DV). IEEE, 2021. 1, 2, 6
  25. 25.Xingyi He, Jiaming Sun, Yifan Wang, Sida Peng, Qixing Huang, Hujun Bao, and Xiaowei Zhou. Detector-free structure from motion. In arxiv, 2023. 8
  26. 26.Di Huang, Xiaopeng Ji, Xingyi He, Jiaming Sun, Tong He, Qing Shuai, Wanli Ouyang, and Xiaowei Zhou. Reconstructing hand-held objects from monocular video. In SIGGRAPH Asia 2022 Conference Papers, 2022. 2, 3
  27. 27.Umar Iqbal, Pavlo Molchanov, Thomas Breuel Juergen Gall, and Jan Kautz. Hand pose estimation via latent 2.5D heatmap regression. In European Conference on Computer Vision (ECCV), 2018. 2
  28. 28.Manuel Kaufmann, Velko Vechev, and Dario Mylonopoulos. aitviewer, 2022. 6
  29. 29.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In International Conference on Computer Vision (ICCV), 2023. 3, 5
  30. 30.Muhammed Kocabas, Jen-Hao Rick Chang, James Gabriel, Oncel Tuzel, and Anurag Ranjan. HUGS: Human gaussian splats. In Computer Vision and Pattern Recognition (CVPR), 2024. 8
  31. 31.Jihyun Lee, Minhyuk Sung, Honggyu Choi, and Tae-Kyun Kim. Im2hands: Learning attentive implicit representation of interacting two-hand shapes. In Computer Vision and Pattern Recognition (CVPR), 2023. 2
  32. 32.Ke Li and Jitendra Malik. Amodal instance segmentation. In European Conference on Computer Vision (ECCV). Springer, 2016. 5
  33. 33.Lijun Li, Linrui Tian, Xindi Zhang, Qi Wang, Bang Zhang, Liefeng Bo, Mengyuan Liu, and Chen Chen. RenderIH: A large-scale synthetic dataset for 3D interacting hand pose estimation. In International Conference on Computer Vision (ICCV), 2023. 2
  34. 34.Mengcheng Li, Liang An, Hongwen Zhang, Lianpeng Wu, Feng Chen, Tao Yu, and Yebin Liu. Interacting attention graph for single image two-hand reconstruction. In Computer Vision and Pattern Recognition (CVPR), 2022. 2
  35. 35.Kevin Lin, Lijuan Wang, and Zicheng Liu. End-to-end human pose and mesh reconstruction with transformers. In Computer Vision and Pattern Recognition (CVPR), 2021. 3
  36. 36.Shichen Liu, Tianye Li, Weikai Chen, and Hao Li. Soft rasterizer: A differentiable renderer for image-based 3D reasoning. In International Conference on Computer Vision (ICCV), 2019. 5
  37. 37.Shaowei Liu, Hanwen Jiang, Jiarui Xu, Sifei Liu, and Xiaolong Wang. Semi-supervised 3D hand-object poses estimation with interactions in time. In Computer Vision and Pattern Recognition (CVPR), 2021. 2
  38. 38.Charles Loop. Smooth subdivision surfaces based on triangles. 1987. 5
  39. 39.Hao Meng, Sheng Jin, Wentao Liu, Chen Qian, Mengxiang Lin, Wanli Ouyang, and Ping Luo. 3D interacting hand pose estimation by hand de-occlusion and removal. In European Conference on Computer Vision (ECCV). Springer, 2022. 2
  40. 40.Gyeongsik Moon. Bringing inputs to shared domains for 3D interacting hands recovery in the wild. In Computer Vision and Pattern Recognition (CVPR), 2023.
  41. 41.Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Kyoung Mu Lee. InterHand2.6M: A dataset and baseline for 3D interacting hand pose estimation from a single RGB image. In European Conference on Computer Vision (ECCV), 2020. 2
  42. 42.Gyeongsik Moon, Shunsuke Saito, Weipeng Xu, Rohan Joshi, Julia Buffalini, Harley Bellan, Nicholas Rosen, Jesse Richardson, Mallorie Mize, Philippe De Bree, et al. A dataset of relighted 3D interacting hands. NeurIPS, 2023. 2
  43. 43.Franziska Mueller, Florian Bernard, Oleksandr Sotnychenko, Dushyant Mehta, Srinath Sridhar, Dan Casas, and Christian Theobalt. GANerated hands for real-time 3D hand tracking from monocular RGB. In Computer Vision and Pattern Recognition (CVPR), 2018. 2
  44. 44.Takehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan, Luan Tran, and Cem Keskin. AssemblyHands: towards egocentric activity understanding via 3d hand pose estimation. In Computer Vision and Pattern Recognition (CVPR), 2023. 2
  45. 45.Valeria Perasso. What have you touched today?, 2015. 1
  46. 46.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. DreamFusion: Text-to-3D using 2D diffusion. International Conference on Learning Representations (ICLR), 2022. 8
  47. 47.Aditya Prakash, Matthew Chang, Matthew Jin, and Saurabh Gupta. Learning hand-held object reconstruction from in-the-wild videos. arXiv, 2305.03036, 2023. 2
  48. 48.Wentian Qu, Zhaopeng Cui, Yinda Zhang, Chenyu Meng, Cuixia Ma, Xiaoming Deng, and Hongan Wang. Novel-view synthesis and pose estimation for hand-object interaction from sparse views. In International Conference on Computer Vision (ICCV), 2023. 2
  49. 49.James M. Rehg and Takeo Kanade. Visual tracking of high DOF articulated structures: An application to human hand tracking. In European Conference on Computer Vision (ECCV), 1994. 2
  50. 50.Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bodies together. Transactions on Graphics (TOG), 36(6), 2017. 3, 4
  51. 51.Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. In Computer Vision and Pattern Recognition (CVPR), 2019. 3
  52. 52.Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperGlue: Learning feature matching with graph neural networks. In Computer Vision and Pattern Recognition (CVPR), 2020. 3
  53. 53.Tomas Simon, Hanbyul Joo, Iain Matthews, and Yaser Sheikh. Hand keypoint detection in single images using multiview bootstrapping. In Computer Vision and Pattern Recognition (CVPR), 2017. 2
  54. 54.Adrian Spurr, Jie Song, Seonwook Park, and Otmar Hilliges. Cross-modal deep variational hand pose estimation. In Computer Vision and Pattern Recognition (CVPR), 2018.
  55. 55.Adrian Spurr, Umar Iqbal, Pavlo Molchanov, Otmar Hilliges, and Jan Kautz. Weakly supervised 3D hand pose estimation via biomechanical constraints. In European Conference on Computer Vision (ECCV), 2020.
  56. 56.Adrian Spurr, Aneesh Dahiya, Xi Wang, Xucong Zhang, and Otmar Hilliges. Self-supervised 3D hand pose estimation from monocular RGB via contrastive learning. In International Conference on Computer Vision (ICCV), 2021. 2
  57. 57.Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. LoFTR: Detector-free local feature matching with transformers. Computer Vision and Pattern Recognition (CVPR), 2021. 8
  58. 58.Anilkumar Swamy, Vincent Leroy, Philippe Weinzaepfel, Fabien Baradel, Salma Galaaoui, Romain Bregier, Matthieu Armando, Jean-Sebastien Franco, and Gregory Rogez. SHOWMe: Benchmarking object-agnostic hand-object 3D reconstruction. In International Conference on Computer Vision (ICCV), 2023. 2
  59. 59.Omid Taheri, Yi Zhou, Dimitrios Tzionas, Yang Zhou, Duygu Ceylan, Soren Pirk, and Michael J. Black. GRIP: Generating interaction poses using latent consistency and spatial cues. In International Conference on 3D Vision (3DV), 2024. 1
  60. 60.Maxim Tatarchenko, Stephan R Richter, Rene Ranftl, Zhuwen Li, Vladlen Koltun, and Thomas Brox. What do single-view 3D reconstruction networks learn? In Computer Vision and Pattern Recognition (CVPR), 2019. 6
  61. 61.Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+O: Unified egocentric recognition of 3D hand-object poses and interactions. In Computer Vision and Pattern Recognition (CVPR), 2019. 2
  62. 62.Tze Ho Elden Tse, Kwang In Kim, Ales Leonardis, and Hyung Jin Chang. Collaborative learning for hand and object reconstruction with attention-guided graph convolution. In Computer Vision and Pattern Recognition (CVPR), 2022. 2
  63. 63.Tze Ho Elden Tse, Franziska Mueller, Zhengyang Shen, Danhang Tang, Thabo Beeler, Mingsong Dou, Yinda Zhang, Sasa Petrovic, Hyung Jin Chang, Jonathan Taylor, et al. Spectral graphormer: Spectral graph-based transformer for egocentric two-hand reconstruction using multi-view color images. In International Conference on Computer Vision (ICCV), 2023. 2
  64. 64.Dimitrios Tzionas and Juergen Gall. A comparison of directional distances for hand pose estimation. In German Conference on Pattern Recognition (GCPR), 2013. 2
  65. 65.Dimitrios Tzionas and Juergen Gall. 3D object reconstruction from hand-object interactions. In International Conference on Computer Vision (ICCV), 2015. 3
  66. 66.He Wang, Srinath Sridhar, Jingwei Huang, Julien Valentin, Shuran Song, and Leonidas J Guibas. Normalized object coordinate space for category-level 6d object pose and size estimation. In Computer Vision and Pattern Recognition (CVPR), 2019. 3
  67. 67.Bowen Wen, Jonathan Tremblay, Valts Blukis, Stephen Tyree, Thomas Muller, Alex Evans, Dieter Fox, Jan Kautz, and Stan Birchfield. BundleSDF: Neural 6-dof tracking and 3d reconstruction of unknown objects. In Computer Vision and Pattern Recognition (CVPR), 2023. 3
  68. 68.Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu. CPF: Learning a contact potential field to model the hand-object interaction. In International Conference on Computer Vision (ICCV), 2021. 1, 2
  69. 69.Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. Volume rendering of neural implicit surfaces. In NeurIPS, 2021. 4
  70. 70.Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What’s in your hands? 3D reconstruction of generic objects in hands. In Computer Vision and Pattern Recognition (CVPR), 2022. 1, 2, 6
  71. 71.Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shubham Tulsiani. Diffusion-guided reconstruction of everyday hand-object interaction clips. In International Conference on Computer Vision (ICCV), 2023. 1, 2, 6
  72. 72.Yifei Yin, Chen Guo, Manuel Kaufmann, Juan Zarate, Jie Song, and Otmar Hilliges. Hi4D: 4D instance segmentation of close human interaction. In Computer Vision and Pattern Recognition (CVPR), 2023. 5
  73. 73.Baowen Zhang, Yangang Wang, Xiaoming Deng, Yinda Zhang, Ping Tan, Cuixia Ma, and Hongan Wang. Interacting two-hand 3D pose and shape reconstruction from single color image. In International Conference on Computer Vision (ICCV), 2021. 2
  74. 74.Hui Zhang, Sammy Christen, Zicong Fan, Otmar Hilliges, and Jie Song. GraspXL: Generating grasping motions for diverse objects at scale. arXiv preprint, 2024. 1
  75. 75.Hui Zhang, Sammy Christen, Zicong Fan, Luocheng Zheng, Jemin Hwangbo, Jie Song, and Otmar Hilliges. ArtiGrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation. In International Conference on 3D Vision (3DV), 2024. 1
  76. 76.Jason Y Zhang, Sam Pepose, Hanbyul Joo, Deva Ramanan, Jitendra Malik, and Angjoo Kanazawa. Perceiving 3D human-object spatial arrangements from a single image in the wild. In European Conference on Computer Vision (ECCV). Springer, 2020. 5
  77. 77.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020. 3, 4
  78. 78.Xiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang, and Wen Zheng. End-to-end hand mesh recovery from a monocular RGB image. In International Conference on Computer Vision (ICCV), 2019. 2
  79. 79.Licheng Zhong, Lixin Yang, Kailin Li, Haoyu Zhen, Mei Han, and Cewu Lu. Color-NeuS: Reconstructing neural implicit surfaces with color. In International Conference on 3D Vision (3DV), 2024. 2, 3
  80. 80.Yuxiao Zhou, Marc Habermann, Weipeng Xu, Ikhsanul Habibie, Christian Theobalt, and Feng Xu. Monocular real-time hand shape and motion capture using multi-modal data. In Computer Vision and Pattern Recognition (CVPR), 2020. 2
  81. 81.Andrea Ziani, Zicong Fan, Muhammed Kocabas, Sammy Christen, and Otmar Hilliges. TempCLR: Reconstructing hands via time-coherent contrastive learning. In International Conference on 3D Vision (3DV), 2022. 2
  82. 82.Christian Zimmermann and Thomas Brox. Learning to estimate 3D hand pose from single RGB images. In International Conference on Computer Vision (ICCV), 2017. 2

Citation

MLA
Fan, Z., et al. “HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video”. arXiv, 2023, http://arxiv.org/abs/2311.18448v1.
APA
Fan, Z., Parelli, M., Kadoglou, M. E., Kocabas, M., Chen, X., Black, M. J., & Hilliges, O. (2023). HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video. arXiv. http://arxiv.org/abs/2311.18448v1
Chicago
Fan, Z., M. Parelli, M. E. Kadoglou, et al. 2023. “HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video”. arXiv. http://arxiv.org/abs/2311.18448v1.
Harvard
Fan, Z. et al. (2023) “HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.18448v1.
Vancouver
1. Fan Z, Parelli M, Kadoglou ME, Kocabas M, Chen X, Black MJ, Hilliges O (2023) HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video. arXiv

BibTeX

@article{fan2023hold,
  title = {HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video},
  author = {Fan, Zicong and Parelli, Maria and Kadoglou, Maria Eleni and Kocabas, Muhammed and Chen, Xu and Black, Michael J. and Hilliges, Otmar},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.18448v1},
  eprint = {2311.18448}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE