ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis
Lixin YangKailin LiXinyu ZhanJun LvWenqiang XuJiefeng LiCewu Lu
Proposes an online synthetic data generation framework that adaptively samples and renders diverse, physically valid hand-object interactions based on training loss feedback to improve single-image 3D hand-object pose estimation.
Estimating the articulated three-dimensional poses of hands and objects from a single standard camera image is vital for applications in robotics and augmented reality. However, real-world data collection and manual annotation for this task are exceptionally difficult, slow, and expensive. Because human hands have high degrees of freedom and interact closely with objects, existing datasets suffer from limited diversity in hand poses, object configurations, and camera viewpoints.
The article demonstrates an online data enhancement framework called ArtiBoost, which systematically boosts hand-object pose estimation performance by continuously exploring and synthesizing diverse, physically plausible interaction data during model training.
The approach constructs a structured search space encompassing object types, anatomically valid hand grasp configurations, and camera viewpoints. Grasp poses are generated using contact constraints to ensure realism while preventing impossible hand-object intersections. Rather than generating a static synthetic dataset offline, the framework operates in real time alongside model training. It renders synthetic images, mixes them into batches of real images, and uses training error feedback to adaptively re-weight and sample difficult hand-object configurations that the machine learning model struggles to discern.
The findings show substantial improvements across standard benchmarks. Integrating ArtiBoost into standard classification and regression baseline models allowed them to outperform previous state-of-the-art architectures, improving hand pose error by approximately 10% to 28% on the HO3D benchmark. The contact-guided grasp synthesis significantly outperformed conventional offline grasp datasets. Crucially, a baseline model trained on only 10% of real annotated data combined with synthetic data achieved better accuracy than the same model trained on 100% of real annotated data.
These results demonstrate that online, targeted synthetic data generation can drastically lower the cost and operational bottlenecks associated with collecting massive real-world training datasets. By prioritizing harder examples and ensuring anatomical plausibility, machine learning systems can achieve higher precision and faster convergence without requiring overly complex neural network architectures.
Organizations developing vision-based manipulation or tracking systems should consider adopting dynamic, contact-aware data synthesis to augment scarce labeled data and reduce annotation costs. For deployment, engineering teams should evaluate integrating adaptive sampling into existing training pipelines.
The primary limitations include a persistent visual domain gap between synthetic renderings and real images, as well as dependence on a predefined lookup space rather than a fully differentiable rendering pipeline. Nonetheless, the experimental evidence strongly supports that increasing pose diversity via targeted online synthesis is a highly effective, reliable strategy for improving 3D hand-object pose estimation.
- Paper: Expressive Body Capture: 3D Hands, Face, and Body From a Single Image, Georgios Pavlakos et al. (2019). Presents expressive whole-body and articulated hand parametric modeling (SMPL-X) that underlies modern 3D hand representation and pose estimation from monocular images.
- Paper: PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes, Yu Xiang et al. (2017). Establishes foundational CNN-based frameworks and loss formulations for 6D object pose estimation from RGB images in occluded settings.
- Paper: Learning from Simulated and Unsupervised Images through Adversarial Training, Ashish Shrivastava et al. (2016). Demonstrates adversarial refinement methods to bridge the domain gap between synthetic and real data for 3D hand pose estimation.
- Paper: Domain randomization for transferring deep neural networks from simulation to the real world, Josh Tobin et al. (2017). Introduces domain randomization as a core technique to train deep vision models on synthetic renderings for robust real-world transfer.
- Paper: A Brief Introduction to Boosting, R. Schapire (1999). Provides the foundational algorithmic principles of boosting and adaptive sample reweighting that inspire ArtiBoost's loss-feedback data exploration pipeline.
- Paper: ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation, Zicong Fan et al. (2023). Provides a comprehensive real-world dataset of bimanual dexterous hand-object manipulation with accurate 3D contacts, extending the evaluation of articulated hand-object interaction beyond monocular synthetic exploration.
- Paper: HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos, Prithviraj Banerjee et al. (2025). Expands 3D hand-object pose and tracking research into egocentric, multi-view video settings across diverse real-world daily interactions.
- Paper: Objaverse: A Universe of Annotated 3D Objects, Matt Deitke et al. (2022). Scales 3D object repositories to massive collections of annotated and articulable 3D assets that can directly empower synthetic data synthesis pipelines like ArtiBoost.
