ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis

Lixin YangKailin LiXinyu ZhanJun LvWenqiang XuJiefeng LiCewu Lu

article2022CVPR120 citations

Proposes an online synthetic data generation framework that adaptively samples and renders diverse, physically valid hand-object interactions based on training loss feedback to improve single-image 3D hand-object pose estimation.

Listen

Estimating the articulated three-dimensional poses of hands and objects from a single standard camera image is vital for applications in robotics and augmented reality. However, real-world data collection and manual annotation for this task are exceptionally difficult, slow, and expensive. Because human hands have high degrees of freedom and interact closely with objects, existing datasets suffer from limited diversity in hand poses, object configurations, and camera viewpoints.

The article demonstrates an online data enhancement framework called ArtiBoost, which systematically boosts hand-object pose estimation performance by continuously exploring and synthesizing diverse, physically plausible interaction data during model training.

The approach constructs a structured search space encompassing object types, anatomically valid hand grasp configurations, and camera viewpoints. Grasp poses are generated using contact constraints to ensure realism while preventing impossible hand-object intersections. Rather than generating a static synthetic dataset offline, the framework operates in real time alongside model training. It renders synthetic images, mixes them into batches of real images, and uses training error feedback to adaptively re-weight and sample difficult hand-object configurations that the machine learning model struggles to discern.

The findings show substantial improvements across standard benchmarks. Integrating ArtiBoost into standard classification and regression baseline models allowed them to outperform previous state-of-the-art architectures, improving hand pose error by approximately 10% to 28% on the HO3D benchmark. The contact-guided grasp synthesis significantly outperformed conventional offline grasp datasets. Crucially, a baseline model trained on only 10% of real annotated data combined with synthetic data achieved better accuracy than the same model trained on 100% of real annotated data.

These results demonstrate that online, targeted synthetic data generation can drastically lower the cost and operational bottlenecks associated with collecting massive real-world training datasets. By prioritizing harder examples and ensuring anatomical plausibility, machine learning systems can achieve higher precision and faster convergence without requiring overly complex neural network architectures.

Organizations developing vision-based manipulation or tracking systems should consider adopting dynamic, contact-aware data synthesis to augment scarce labeled data and reduce annotation costs. For deployment, engineering teams should evaluate integrating adaptive sampling into existing training pipelines.

The primary limitations include a persistent visual domain gap between synthetic renderings and real images, as well as dependence on a predefined lookup space rather than a fully differentiable rendering pipeline. Nonetheless, the experimental evidence strongly supports that increasing pose diversity via targeted online synthesis is a highly effective, reliable strategy for improving 3D hand-object pose estimation.

Cover for ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis

Abstract

Estimating the articulated 3D hand-object pose from a single RGB image is a highly ambiguous and challenging problem, requiring large-scale datasets that contain diverse hand poses, object types, and camera viewpoints. Most real-world datasets lack these diversities. In contrast, data synthesis can easily ensure those diversities separately. However, constructing both valid and diverse hand-object interactions and efficiently learning from the vast synthetic data is still challenging. To address the above issues, we propose ArtiBoost, a lightweight online data enhancement method. ArtiBoost can cover diverse hand-object poses and camera viewpoints through sampling in a Composited hand-object Configuration and Viewpoint space (CCV-space) and can adaptively enrich the current hard-discernable items by loss-feedback and sample re-weighting. ArtiBoost alternatively performs data exploration and synthesis within a learning pipeline, and those synthetic data are blended into real-world source data for training. We apply ArtiBoost on a simple learning baseline network and witness the performance boost on several hand-object benchmarks. Our models and code are available at https://github.com/lixiny/ArtiBoost.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Online Exploration in CCV-Space
  • 3.2. Online Synthesis for HOPE task
  • 3.3. Learning Framework
  • 4. Experiment and Result
  • 4.1. Dataset and Metrics
  • 4.2. HOPE Network Performance
  • 4.3. Ablation Study
  • 4.4. Applications
  • 5. Discussion
  • References

Knowls

  1. Knowl 1 — ArtiBoost alternates online exploration, synthesis, and loss-guided training

    model/method

    ArtiBoost is a model-agnostic online data-enhancement method for articulated 3D hand-object pose estimation. Given a real training set DrealD_{\mathrm{real}}, a discrete lookup space C\mathcal{C} of hand-object-viewpoint triplets, and a sampling-weight map MM, it repeatedly performs the following loop: (1) sample hand-object-viewpoint triplets from C\mathcal{C} according to MM; (2) render the sampled triplets into labeled synthetic RGB images; (3) mix the synthetic images with real images in training batches; (4) train any HOPE network by forward and backward propagation; and (5) after each training epoch, update MM using the synthetic samples' pose-estimation losses. Thus, the data generator and the pose-estimation model communicate during training, causing later synthetic batches to emphasize triplets that are difficult for the current model. The method is integrated by modifying the data loader rather than the network architecture.

  2. Knowl 2 — Composited hand-object configuration and viewpoint space

    definition

    The Composited Configuration and Viewpoint space (CCV-space) is a finite representation of the HOPE input domain. It contains object type, valid hand-object configuration, and camera viewpoint as three discrete dimensions:

    S={(no,np,nv)∈N+3∣no≤No,  np≤Np,  nv≤Nv}.\mathcal{S}=\{(n_o,n_p,n_v)\in\mathbb{N}_{+}^{3}\mid n_o\leq N_o,\;n_p\leq N_p,\;n_v\leq N_v\}.

    Here, NoN_o is the number of object types, NpN_p is the number of stored interacting hand poses per object, NvN_v is the number of camera viewpoints, and (no,np,nv)(n_o,n_p,n_v) denotes observing the npn_p-th hand-object configuration for object non_o from viewpoint nvn_v. Object type and hand pose are deliberately coupled: a valid grasp depends on the geometry of the selected object, so the space stores composited hand-object configurations rather than independently combining arbitrary objects and hand poses.

    Camera directions are sampled uniformly on the unit sphere using u∼U[−1,1]u\sim U[-1,1] and ϕ∼U[0,2π]\phi\sim U[0,2\pi]:

    nv=(1−u2cos⁡ϕ,  1−u2sin⁡ϕ,  u)T,\mathbf{n}_v=\left(\sqrt{1-u^2}\cos\phi,\;\sqrt{1-u^2}\sin\phi,\;u\right)^{\mathsf T},

    where nv\mathbf{n}_v is the viewpoint direction. The implementation uses Nu=12N_u=12 elevation samples and Nϕ=24N_\phi=24 azimuth samples, giving Nv=288N_v=288 viewpoints. It fits 300 candidate interactions per object, manually removes severely intersecting or unnatural grasps, and retains at most Np=100N_p=100 poses per object. For the 20 objects in DexYCB, this yields 20×100×288=576,00020\times100\times288=576{,}000 possible triplets.

  3. Knowl 3 — Anatomically constrained hand configuration space

    definition

    ArtiBoost represents hand geometry with the MANO model but does not interpolate its unconstrained 48-dimensional joint-rotation space. Instead, it constructs a valid hand configuration space with 21 degrees of freedom using an axis-adapted twist-splay-bend coordinate system. The five non-metacarpal joint chains are restricted to bending, while the five metacarpal joints can bend and splay; twisting along the finger-pointing direction and splaying at non-metacarpal joints are prohibited. Within each finger, proximal and distal bending angles are linked, whereas metacarpal bending is independent. Different fingers are independent provided that these anatomical restrictions are respected. The resulting degrees of freedom comprise one splay angle and two independent bending degrees of freedom for each of five fingers, plus six wrist degrees of freedom.

    The hand configuration is therefore described by the selected bending or splaying angles of 15 joints together with a wrist pose ξw∈se(3)\xi_w\in\mathrm{se}(3), while MANO shape parameters remain separate. These restrictions are used both to avoid anatomically abnormal interpolations and to preserve visually plausible prehensile configurations.

  4. Knowl 4 — Contact-guided synthesis of diverse hand-object grasps

    model/method

    ArtiBoost constructs the composited hand-object configuration space by fitting MANO hands to object contacts. An interaction is accepted only when the thumb and at least one other finger contact the object surface and the hand and object meshes do not intersect.

    For each object, the synthesis procedure is: (1) construct an offset surface outside the object, uniformly sample wrist-control points pwp_w, find the closest object vertex vov_o for each point, and use vo−pwv_o-p_w as the hand approach direction; initialize a prehensile hand resembling a parallel-jaw-gripper rest pose at pwp_w; (2) define the contact-feasible region between vov_o and the farthest object vertex reachable by the fingers, randomly select the thumb plus NN additional fingers with 1≤N≤41\leq N\leq4, choose a random minimum reaching radius rcr_c, and pair each selected fingertip pfp_f with a contact point vcv_c satisfying ∥vc−pw∥2≥rc\|v_c-p_w\|_2\geq r_c and minimizing its distance to pfp_f; and (3) optimize the hand using an anchor-based, contact-based objective in which unattached fingertip anchors are attracted to their paired object contacts and intersecting anchors are pushed outward. Optimization is restricted to the valid 21-DoF hand configuration space, preserving the hand-pose constraints. The procedure generates 300 candidates per object before manual rejection of severe interpenetration and unnatural grasps, after which up to 100 diverse interacting poses are retained.

  5. Knowl 5 — Loss-percentile sampling emphasizes hard CCV triplets

    algorithm

    ArtiBoost assigns every CCV-space triplet ii a nonnegative sampling weight wiw_i in a map MM. Its sampling probability is

    pi=wi∑jwj,p_i=\frac{w_i}{\sum_j w_j},

    where the sum is over all CCV-space triplets. Synthetic triplets are sampled according to the resulting multinomial distribution, with sampling performed without replacement within the current exploration batch.

    After each training epoch, the method computes the wrist-aligned mean per joint position error eie_i for every synthetic sample. Let emin⁡e_{\min} and emax⁡e_{\max} be the minimum and maximum errors in that epoch. The normalized error percentile and multiplicative weight update are

    qi=emax⁡−eiemax⁡−emin⁡,δwi=1qi+0.5.q_i=\frac{e_{\max}-e_i}{e_{\max}-e_{\min}},\qquad \delta w_i=\frac{1}{q_i+0.5}.

    The updated weight is the original weight multiplied by δwi\delta w_i, then clamped to the interval [0.1,2.0][0.1,2.0]. Consequently, the sample with maximum error receives factor 22, while the sample with minimum error receives factor 2/32/3. High-loss, hard-to-discern triplets therefore become more likely in later exploration rounds, while already easy triplets are down-weighted without allowing the distribution to become arbitrarily imbalanced.

  6. Knowl 6 — Online synthesis adds pose, viewpoint, shape, and appearance variation

    model/method

    Before rendering a sampled CCV triplet, ArtiBoost perturbs it while retaining the main validity constraints. The dependency between proximal and distal finger bending angles is relaxed for this disturbance, and Gaussian noise N(0,σ12)\mathcal{N}(0,\sigma_1^2) is added to each of the 15 bending angles with σ1=3\sigma_1=3 degrees. Gaussian noise N(0,σ22)\mathcal{N}(0,\sigma_2^2) is added to the five metacarpal splay angles with σ2=1.5\sigma_2=1.5 degrees. The perturbed hand is passed through the GrabNet RefineNet module to reduce newly introduced hand-object interpenetration. MANO shape parameters β∈R10\beta\in\mathbb{R}^{10} are sampled from the reported distribution N(0,0.5)\mathcal{N}(0,0.5).

    For the camera, the elevation variable receives uniform noise U(−δu,δu)U(-\delta_u,\delta_u) with δu=0.05\delta_u=0.05, the azimuth receives U(−δϕ,δϕ)U(-\delta_\phi,\delta_\phi) with δϕ=7.5\delta_\phi=7.5 degrees, and the in-plane camera rotation receives U(0,2π)U(0,2\pi). Hand skin tone and texture are sampled using the HTML parametric hand-texture model; the paper reports that omitting HTML's shadow-removal operation produces more plausible results. PyRender composites the textured hand and object over COCO backgrounds onto 224×224224\times224 images. Rendering runs at approximately 120 frames per second per Titan X graphics card, allowing it to proceed in parallel with network training.

  7. Knowl 7 — Classification and regression HOPE baselines share a four-term objective

    model/method

    ArtiBoost is evaluated with two simple HOPE networks, both using a ResNet-34 backbone. The classification model, Clas, predicts 22 volumetric heatmaps for 21 hand joints and the object centroid in a restricted uu-vv-depth coordinate system, applies soft-argmax, and converts the resulting coordinates to camera-space 3D points using camera intrinsics and the known wrist location. The regression model, Reg, uses multilayer perceptrons to predict MANO pose θ\theta, MANO shape β\beta, the wrist-relative object centroid, and object rotation ror_o; MANO converts θ\theta and β\beta to wrist-relative hand joints, which are translated to camera space using the known wrist location.

    Both models are trained with a weighted sum of location, object-corner, ordinal-depth, and symmetry-aware losses. For predicted points pip_i and ground-truth points p^i\hat p_i over the 21 hand joints plus object centroid, the location term is

    Lloc=122∑i=122∥pi−p^i∥22.\mathcal{L}_{\mathrm{loc}}=\frac{1}{22}\sum_{i=1}^{22}\|p_i-\hat p_i\|_2^2.

    For the eight object corners, cˉi\bar c_i is a canonical corner, c^i\hat c_i is its ground-truth camera-space position, and exp⁡(ro)\exp(r_o) is the predicted rotation matrix:

    Lcor=18∑i=18∥exp⁡(ro)cˉi−c^i∥22.\mathcal{L}_{\mathrm{cor}}=\frac{1}{8}\sum_{i=1}^{8}\|\exp(r_o)\bar c_i-\hat c_i\|_2^2.

    Let cjc_j be a predicted object corner, n⊥\mathbf n_\perp the viewing direction, and 11i,jord\mathbb{1}\mkern-6mu 1^{\mathrm{ord}}_{i,j} equal to one when the predicted depth ordering of hand joint pip_i and corner cjc_j disagrees with the ground truth and zero otherwise. The ordinal loss is

    Lord=∑j=18∑i=121\mathds1i,jord∣(pi−cj)⋅n⊥∣.\mathcal{L}_{\mathrm{ord}}=\sum_{j=1}^{8}\sum_{i=1}^{21}\mathds{1}^{\mathrm{ord}}_{i,j}\left|(p_i-c_j)\cdot\mathbf n_\perp\right|.

    For an object-specific symmetry set S\mathcal{S} of valid rotation matrices and ground-truth rotation r^o\hat r_o, the symmetry-aware term is

    Lsym=min⁡R∈S18∑i=18∥exp⁡(ro)cˉi−exp⁡(r^o)Rcˉi∥22.\mathcal{L}_{\mathrm{sym}}=\min_{R\in\mathcal{S}}\frac{1}{8}\sum_{i=1}^{8}\left\|\exp(r_o)\bar c_i-\exp(\hat r_o)R\bar c_i\right\|_2^2.

    The final objective is

    LHOPE=Lloc+λ1Lcor+λ2Lord+λ3Lsym,\mathcal{L}_{\mathrm{HOPE}}=\mathcal{L}_{\mathrm{loc}}+\lambda_1\mathcal{L}_{\mathrm{cor}}+\lambda_2\mathcal{L}_{\mathrm{ord}}+\lambda_3\mathcal{L}_{\mathrm{sym}},

    where λ1,λ2,λ3\lambda_1,\lambda_2,\lambda_3 are loss weights. For ordinary MPCPE evaluation the experiments use λ1=λ2=1\lambda_1=\lambda_2=1 and λ3=0\lambda_3=0; for symmetry-aware training they use λ1=λ2=0\lambda_1=\lambda_2=0 and λ3=1\lambda_3=1.

  8. Knowl 8 — Evaluation protocol covers three hand-object benchmarks and three pose metrics

    experimental setup

    The method is evaluated on FHAB, HO3D, HO3D v3, and DexYCB. FHAB contains approximately 20K manipulation samples; the action split used here has 10,503 training and 10,998 testing samples and is used mainly to verify the learning framework. HO3D testing is evaluated through its online server, while HO3D v3 uses its revised training/testing split. DexYCB contains 582K frames involving 20 YCB objects; the experiments use the official right-hand S0 split and remove frames whose minimum hand-object distance exceeds 5 cm.

    Hand accuracy is measured by wrist-aligned mean per joint position error (MPJPE). Object accuracy is measured by mean per corner position error (MPCPE) and maximum symmetry-aware surface distance (MSSD). MPCPE evaluates a unique object pose, whereas MSSD compares the predicted pose with the closest pose under the object's rotational symmetries, which is more appropriate for symmetric or revolution-invariant objects and for objects heavily occluded by the hand.

  9. Knowl 9 — ArtiBoost improves classification and regression baselines across benchmarks

    empirical result

    The reported benchmark comparisons show that adding ArtiBoost synthetic data improves both the classification and regression baselines. Errors are in millimeters for FHAB and DexYCB and centimeters for HO3D and HO3D v3. MPJPE and MPCPE are wrist-aligned unless otherwise stated.

    Could not parse LaTeX table
    Could not parse LaTeX table

    With symmetry-aware training on HO3D, the reported values are MPJPE followed by MSSD for mustard bottle, bleach cleanser, and potted meat can: Hampali et al. achieves 2.57,4.41,6.03,9.082.57,4.41,6.03,9.08; Clas sym achieves 3.10,4.07,6.56,8.703.10,4.07,6.56,8.70; and Clas sym + ArtiBoost achieves 2.53,3.14,5.72,6.362.53,3.14,5.72,6.36 cm. On DexYCB, Clas sym obtains MPJPE 13.0013.00 mm and MSSD values 74.9574.95, 63.6863.68, 88.1088.10, and 91.6691.66 mm for power drill, cracker box, scissors, and bleach cleanser; Clas sym + ArtiBoost obtains 12.8012.80 mm and 52.7052.70, 46.1346.13, 66.5266.52, and 72.3172.31 mm. On HO3D v3, Clas + ArtiBoost improves over Clas from MPJPE/MPCPE 2.94/7.532.94/7.53 to 2.50/5.882.50/5.88 cm, while Clas sym + ArtiBoost achieves the best reported MPJPE of 2.342.34 cm and MSSD values 2.662.66, 5.235.23, and 5.825.82 cm for the three listed object categories.

  10. Knowl 10 — Ablations show that CCV diversity and online feedback are both important

    empirical result

    The ablations isolate the contributions of contact-guided CCV grasps, online re-weighting, and synthetic data volume. On HO3D, all values are in centimeters:

    Could not parse LaTeX table

    Using the same rendering pipeline and the same amount of synthetic data, CCV-space grasps outperform manually annotated YCBAfford grasps, supporting the paper's claim that valid and diverse interaction configurations matter more than simply adding grasp examples. The full method improves further through online re-weighting.

    When only 10% of HO3D's real training data is used, Reg obtains 3.81/8.773.81/8.77 MPJPE/MPCPE, whereas Reg with ArtiBoost obtains 3.29/6.873.29/6.87; Clas obtains 3.63/7.663.63/7.66, whereas Clas with ArtiBoost obtains 3.05/6.023.05/6.02. Thus, synthetic data partly compensates for reduced real-world supervision, although the 100%-real-data Clas baseline remains at 3.06/7.243.06/7.24.

    ArtiBoost also transfers to the camera-space hand-object model of Hasson et al. On HO3D v1, the reproduced baseline changes from MPJPE/MPCPE/ camera-space MPJPE 5.75/9.61/6.245.75/9.61/6.24 to 3.67/3.24/3.573.67/3.24/3.57 with ArtiBoost. On HO3D, it changes from 3.69/12.38/5.523.69/12.38/5.52 to 3.39/8.31/4.903.39/8.31/4.90, all in centimeters. In the DexYCB convergence study, online re-weighting reaches lower MPJPE throughout training and finishes at approximately 12.8012.80 mm versus 12.9512.95 mm for fixed-weight offline sampling.

  11. Knowl 11 — ArtiBoost does not explicitly solve the synthetic-to-real domain gap

    limitation

    The method does not explicitly mitigate the domain gap between rendered and real images. The authors attribute the dominant improvement to the greater diversity of hand-object pose variants rather than to more realistic rendered appearance. In addition, the renderer is non-differentiable, so the current method explores a predefined lookup table such as CCV-space rather than continuously optimizing the data distribution. The proposed future direction is a generative and contrastive model that learns features shared by real and synthetic images.

Coverage note — No other substantial load-bearing contribution was omitted; appendix implementation details and qualitative visualizations were not expanded into separate knowls because they add reproduction detail or illustration rather than a distinct methodological or quantitative contribution.

References

  1. 1.Adnane Boukhayma, Rodrigo de Bem, and Philip HS Torr. 3d hand shape and pose from images in the wild. In CVPR, 2019. 2
  2. 2.Samarth Brahmbhatt, Ankur Handa, James Hays, and Dieter Fox. ContactGrasp: Functional Multi-finger Grasp Synthesis from Contact. In IROS, 2019. 1, 2
  3. 3.Samarth Brahmbhatt, Chengcheng Tang, Christopher D Twigg, Charles C Kemp, and James Hays. ContactPose: A dataset of grasps with object contact and hand pose. In ECCV, 2020. 1
  4. 4.Berk Calli, Arjun Singh, Aaron Walsman, Siddhartha Srinivasa, Pieter Abbeel, and Aaron M. Dollar. The ycb object and model set: Towards common benchmarks for manipulation research. In ICAR, 2015. 7
  5. 5.Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa, and Jitendra Malik. Reconstructing hand-object interactions in the wild. In ICCV, 2021. 1, 2
  6. 6.Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S. Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox. DexYCB: A benchmark for capturing hand grasping of objects. In CVPR, 2021. 1, 2, 6
  7. 7.Wenzheng Chen, Huan Wang, Yangyan Li, Hao Su, Zhenhua Wang, Changhe Tu, Dani Lischinski, Daniel Cohen-Or, and Baoquan Chen. Synthesizing training images for boosting human 3d pose estimation. In 3DV, 2016. 2
  8. 8.Xingyu Chen, Yufeng Liu, Chongyang Ma, Jianlong Chang, Huayan Wang, Tian Chen, Xiaoyan Guo, Pengfei Wan, and Wen Zheng. Camera-space hand mesh recovery via semantic aggregationand adaptive 2d-1d registration. In CVPR, 2021. 1, 2
  9. 9.Xingyu Chen, Yufeng Liu, Dong Yajiao, Xiong Zhang, Chongyang Ma, Yanmin Xiong, Yuan Zhang, and Xiaoyan Guo. MobRecon: Mobile-friendly hand mesh reconstruction from monocular image. In arXiv preprint arXiv:2112.02753, 2021. 1, 2
  10. 10.Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gregory Rogez. Ganhand: Predicting ´ human grasp affordances in multi-object scenes. In CVPR, 2020. 1, 2, 7
  11. 11.Bardia Doosti, Shujon Naha, Majid Mirbagheri, and David J Crandall. HOPE-Net: A graph-based model for hand-object pose estimation. In CVPR, 2020. 1
  12. 12.Carlo Ferrari and John F Canny. Planning optimal grasps. In ICRA, 1992. 2
  13. 13.Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek, and Tae-Kyun Kim. First-person hand action benchmark with rgb-d videos and 3d hand pose annotations. In CVPR, 2018. 1, 6
  14. 14.Liuhao Ge, Zhou Ren, Yuncheng Li, Zehao Xue, Yingying Wang, Jianfei Cai, and Junsong Yuan. 3d hand shape and pose estimation from a single rgb image. In CVPR, 2019. 1, 2
  15. 15.Weifeng Ge, Weilin Huang, Dengke Dong, and Matthew R. Scott. Deep metric learning with hierarchical triplet loss. In ECCV, 2018. 3
  16. 16.Kehong Gong, Jianfeng Zhang, and Jiashi Feng. Poseaug: A differentiable pose augmentation framework for 3d human pose estimation. In CVPR, 2021. 3
  17. 17.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The ”something something” video database for learning and evaluating visual common sense. In ICCV, 2017. 1
  18. 18.Patrick Grady, Chengcheng Tang, Christopher D Twigg, Minh Vo, Samarth Brahmbhatt, and Charles C Kemp. Contactopt: Optimizing contact to improve grasps. In CVPR, 2021. 1, 2
  19. 19.Shreyas Hampali, Markus Oberweger, Mahdi Rad, and V. Lepetit. HO-3D: A multi-user, multi-object dataset for joint 3d hand-object pose estimation. In arXiv preprint arXiv:1907.01481, 2019. 8
  20. 20.Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. Honnotate: A method for 3d annotation of hand and object poses. In CVPR, 2020. 1, 2, 6
  21. 21.Shreyas Hampali, Sayan Deb Sarkar, and Vincent Lepetit. HO-3D v3: Improving the accuracy of hand-object annotations of the ho-3d dataset. In arXiv preprint arXiv:2107.00887, 2021. 2, 6
  22. 22.Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. HandsFormer: Keypoint transformer for monocular 3d pose estimation ofhands and object in interaction. In arXiv preprint arXiv:2104.14639, 2021. 6, 7
  23. 23.Yana Hasson, Bugra Tekin, Federica Bogo, Ivan Laptev, Marc Pollefeys, and Cordelia Schmid. Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction. In CVPR, 2020. 1, 2, 7, 8
  24. 24.Yana Hasson, Gul Varol, Ivan Laptev, and Cordelia Schmid. ¨ Towards unconstrained joint hand-object reconstruction from rgb videos. In 3DV, 2021. 1, 2
  25. 25.Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. In CVPR, 2019. 1, 2, 5
  26. 26.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 6
  27. 27.Alexander Hermans, Lucas Beyer, and Bastian Leibe. In defense of the triplet loss for person re-identification. arXiv preprint arXiv:1703.07737, 2017. 3
  28. 28.Toma´s Hoda ˇ n, Martin Sundermeyer, Bertram Drost, Yann ˇ Labbe, Eric Brachmann, Frank Michel, Carsten Rother, and ´ Jiˇr´ı Matas. Bop challenge 2020 on 6d object localization. In ECCV Workshops, 2020. 7
  29. 29.Lin Huang, Jianchao Tan, Jingjing Meng, Ji Liu, and Junsong Yuan. HOT-Net: Non-autoregressive transformer for 3d hand-object pose estimation. In ACMMM, 2020. 1
  30. 30.Hanwen Jiang, Shaowei Liu, Jiashun Wang, and Xiaolong Wang. Hand-object contact consistency reasoning for human grasps generation. In ICCV, 2021. 3
  31. 31.Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasping field: Learning implicit representations for human grasps. In 3DV, 2020. 1, 3
  32. 32.Wadim Kehl, Fabian Manhardt, Federico Tombari, Slobodan Ilic, and Nassir Navab. SSD-6d: Making rgb-based 3d detection and 6d pose estimation great again. In ICCV, 2017. 2
  33. 33.Mia Kokic, Danica Kragic, and Jeannette Bohg. Learning task-oriented grasping from human activity datasets. IEEE Robotics and Automation Letters, 5(2), 2020. 1, 2
  34. 34.Felix Kuhnke and Jorn Ostermann. Deep head pose estimation using synthetic images and partial adversarial domain adaption for continuous label spaces. ICCV, 2019. 3
  35. 35.Taein Kwon, Bugra Tekin, Jan Stuhmer, Federica Bogo, and Marc Pollefeys. H2o: Two hands manipulating objects for first person interaction recognition. In ICCV, 2021. 1
  36. 36.Jiefeng Li, Chao Xu, Zhicun Chen, Siyuan Bian, Lixin Yang, and Cewu Lu. HybrIK: A hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation. In CVPR, 2021. 1, 2
  37. 37.Xiaolong Li, He Wang, Li Yi, Leonidas J Guibas, A Lynn Abbott, and Shuran Song. Category-level articulated object pose estimation. In CVPR, 2020. 1, 2
  38. 38.John Lin, Ying Wu, and Thomas S Huang. Modeling the constraints of human hand motion. In Proceedings workshop on human motion, 2000. 3
  39. 39.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ICCV, 2017. 3
  40. 40.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 5
  41. 41.Liu Liu, Han Xue, Wenqiang Xu, Haoyuan Fu, and Cewu Lu. Towards real-world category-level articulation pose estimation. IEEE Transactions on Image Processing, 2022. 1, 2
  42. 42.Shaowei Liu, Hanwen Jiang, Jiarui Xu, Sifei Liu, and Xiaolong Wang. Semi-supervised 3d hand-object poses estimation with interactions in time. In CVPR, 2021. 1, 2, 7
  43. 43.Jun Lv, Wenqiang Xu, Lixin Yang, Sucheng Qian, Chongzhao Mao, and Cewu Lu. HandTailor: Towards high-precision monocular 3d hand recovery. In BMVC, 2021. 7
  44. 44.George Marsaglia et al. Choosing a point from the surface of a sphere. The Annals of Mathematical Statistics, 1972. 5
  45. 45.Joanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu, Xiaolong Wang, and Trevor Darrell. Something-else: Compositional action recognition with spatial-temporal interaction networks. In CVPR, 2020. 1
  46. 46.Matthew Matl. PyRender. https://github.com/mmatl/pyrender, 2019. 5
  47. 47.A.T. Miller and P.K. Allen. Graspit! a versatile simulator for robotic grasping. IEEE Robotics Automation Magazine, 2004. 2, 7
  48. 48.Franziska Mueller, Florian Bernard, Oleksandr Sotnychenko, Dushyant Mehta, Srinath Sridhar, Dan Casas, and Christian Theobalt. Ganerated hands for real-time 3d hand tracking from monocular rgb. In CVPR, 2018. 5
  49. 49.Neng Qian, Jiayi Wang, Franziska Mueller, Florian Bernard, Vladislav Golyanik, and Christian Theobalt. Html: A parametric hand texture model for 3d hand reconstruction and personalization. In ECCV, 2020. 5
  50. 50.Gregory Rogez, James S. Supancic, III, and Deva Ramanan. Understanding everyday hands in action from rgb-d images. In ICCV, 2015. 1, 2
  51. 51.Javier Romero, Hedvig Kjellstrom, and Danica Kragic. ¨ Hands in action: real-time 3d reconstruction of hands in interaction with objects. In ICRA, 2010. 1, 2
  52. 52.Javier Romero, Dimitrios Tzionas, and Michael J Black. Embodied hands: Modeling and capturing hands and bodies together. ACM Transactions on Graphics, 2017. 2, 3
  53. 53.Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. In CVPR, 2015. 3
  54. 54.Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick. Training region-based object detectors with online hard example mining. In CVPR, 2016. 3
  55. 55.Kihyuk Sohn. Improved deep metric learning with multi-class n-pair loss objective. In NIPS, 2016. 3
  56. 56.Xiao Sun, Bin Xiao, Fangyin Wei, Shuang Liang, and Yichen Wei. Integral human pose regression. In ECCV, 2018. 2
  57. 57.Martin Sundermeyer, Zoltan-Csaba Marton, Maximilian Durner, Manuel Brucker, and Rudolph Triebel. Implicit 3d orientation learning for 6d object detection from rgb images. In ECCV, 2018. 2
  58. 58.Omid Taheri, Nima Ghorbani, Michael J Black, and Dimitrios Tzionas. GRAB: A dataset of whole-body human grasping of objects. In ECCV, 2020. 3, 4, 5
  59. 59.Xiao Tang, Tianyu Wang, and Chi-Wing Fu. Towards accurate alignment in real-time 3d hand-mesh reconstruction. In ICCV, 2021. 2
  60. 60.Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+O: Unified egocentric recognition of 3d hand-object poses and interactions. In CVPR, 2019. 1, 6
  61. 61.Bugra Tekin, Sudipta N Sinha, and Pascal Fua. Real-time seamless single shot 6d object pose prediction. In CVPR, 2018. 2
  62. 62.Gul Varol, Ivan Laptev, Cordelia Schmid, and Andrew Zisserman. Synthetic humans for action recognition from unseen viewpoints. International Journal of Computer Vision, 2021. 1, 2
  63. 63.Chengde Wan, Thomas Probst, Luc Van Gool, and Angela Yao. Self-supervised 3d hand pose estimation through training by fitting. In CVPR, 2019. 2
  64. 64.Haonan Yan, Jiaqi Chen, Xujie Zhang, Shengkai Zhang, Nianhong Jiao, Xiaodan Liang, and Tianxiang Zheng. Ultrapose: Synthesizing dense pose with 1 billion points by human-body decoupling 3d model. In ICCV, 2021. 1, 2
  65. 65.Linlin Yang, Shicheng Chen, and Angela Yao. Semihand: Semi-supervised hand pose estimation with consistency. In ICCV, 2021. 2
  66. 66.Lixin Yang, Jiasen Li, Wenqiang Xu, Yiqun Diao, and Cewu Lu. BiHand: Recovering hand mesh with multi-stage bisected hourglass networks. In BMVC, 2020. 2
  67. 67.Linlin Yang and Angela Yao. Disentangling latent hands for image synthesis and pose estimation. In CVPR, 2019. 2
  68. 68.Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu. CPF: Learning a contact potential field to model the hand-object interaction. In ICCV, 2021. 1, 2, 4
  69. 69.Yuxiao Zhou, Marc Habermann, Weipeng Xu, Ikhsanul Habibie, Christian Theobalt, and Feng Xu. Monocular real-time hand shape and motion capture using multi-modal data. In CVPR, 2020. 2
  70. 70.Christian Zimmermann and Thomas Brox. Learning to estimate 3d hand pose from single rgb images. In ICCV, 2017. 1, 2, 5

Citation

MLA
Li, K., et al. “ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis”. arXiv, 2021, http://arxiv.org/abs/2109.05488v2.
APA
Li, K., Yang, L., Zhan, X., Lv, J., Xu, W., Li, J., & Lu, C. (2021). ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis. arXiv. http://arxiv.org/abs/2109.05488v2
Chicago
Li, K., L. Yang, X. Zhan, et al. 2021. “ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis”. arXiv. http://arxiv.org/abs/2109.05488v2.
Harvard
Li, K. et al. (2021) “ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2109.05488v2.
Vancouver
1. Li K, Yang L, Zhan X, Lv J, Xu W, Li J, Lu C (2021) ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis. arXiv

BibTeX

@article{li2021artiboost,
  title = {ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis},
  author = {Li, Kailin and Yang, Lixin and Zhan, Xinyu and Lv, Jun and Xu, Wenqiang and Li, Jiefeng and Lu, Cewu},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2109.05488v2},
  eprint = {2109.05488}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE