ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation
Zicong FanOmid TaheriDimitrios TzionasMuhammed KocabasManuel KaufmannMichael J. BlackOtmar Hilliges
Presents the first large-scale multi-view dataset featuring accurate 3D meshes and dynamic contact annotations for bimanual manipulation of articulated objects, establishing new benchmarks and baselines for spatio-temporally consistent 3D motion reconstruction and interaction field estimation.
Enabling artificial intelligence systems to understand physical interaction with the surrounding world remains a fundamental challenge in computer vision and robotics. While humans intuitively manipulate complex items—such as opening a laptop or cutting with scissors—machine learning models struggle to interpret these actions accurately. Existing research datasets have predominantly focused on simple, single-handed grasping of static, rigid objects, lacking the realistic data needed to model how two hands dynamically interact with moving, multi-part items.
The article aims to address this capability gap by introducing ARCTIC, a large-scale multimodal dataset designed to evaluate and benchmark physically consistent, two-handed manipulation of articulated objects. It establishes two core evaluation tasks: reconstructing synchronized three-dimensional motion from standard video and estimating dense spatial interaction fields between hands and objects.
To achieve this, the authors recorded 10 participants performing unconstrained manipulation and grasping across 11 articulated items, generating 2.1 million video frames. The experimental setup synchronized eight static allocentric video cameras and one head-mounted egocentric camera with an array of 54 high-resolution infrared motion capture cameras. Minimal markers and pre-scanned digital models allowed the precise recovery of three-dimensional full-body, hand, and articulated object meshes alongside dynamic contact data without interfering with natural movements. The authors also developed baseline neural network models—ArcticNet for motion reconstruction and InterField for distance estimation—testing both single-frame and recurrent temporal versions.
The findings show that ARCTIC captures a significantly broader range of hand postures and contact regions than prior benchmarks, showing extensive palm engagement beyond simple fingertip contact. In reconstruction evaluations, temporal baseline models consistently outperformed single-frame variants. For allocentric motion reconstruction, the temporal model reduced motion deviation error from 10.4 mm to 9.3 mm and lowered hand acceleration error from 5.7 to 5.0 meters per second squared, while achieving an object reconstruction success rate of approximately 73.5%. Similarly, for interaction field estimation, temporal modeling yielded smoother predictions and reduced distance errors to 8.7 mm from hand to object.
These results demonstrate that temporal context is critical for achieving smooth, physically consistent reconstructions and avoiding unnatural motion artifacts. Bridging this dataset gap provides essential foundation tools for advancing robotic manipulation, augmented and virtual reality, and human behavior analysis. Moving beyond static grasp assumptions reduces the risk of models failing in real-world environments where dynamic coordination is required.
For future development, the authors recommend using the interaction field representation as an explicit constraint to guide and refine pose estimation pipelines. Research efforts should also focus on generative models capable of synthesizing realistic bimanual interactions with multi-part objects and extending depth-based object tracking methods to account for human occlusion.
Key limitations include the controlled studio environment, the evaluation of 11 specific object categories, and the reliance on baseline architectures designed primarily to establish benchmark standards rather than reach peak performance. Nevertheless, given the high capture fidelity and precise marker-based motion alignment, stakeholders can have high confidence in the dataset as an empirical standard for training and benchmarking advanced manipulation systems.
- Paper: Expressive Body Capture: 3D Hands, Face, and Body From a Single Image, Georgios Pavlakos et al. (2019). Introduces SMPL-X and unified parametric representations of hands and bodies that form the foundational 3D mesh representations used by ARCTIC to model dynamic hand-object interactions.
- Paper: AMASS: Archive of Motion Capture As Surface Shapes, Naureen Mahmood et al. (2019). Establishes standard methodologies for fitting parametric meshes (MANO and SMPL) to motion capture marker data, underpinning ARCTIC's ground-truth 3D registration pipeline.
- Paper: End-to-End Recovery of Human Shape and Pose, Angjoo Kanazawa et al. (2017). Pioneers modern end-to-end regression frameworks for estimating parametric 3D body and mesh surfaces from monocular images, directly informing the baseline architectures tested on ARCTIC.
- Paper: Ego4D: Around the World in 3,000 Hours of Egocentric Video, Kristen Grauman et al. (2021). Provides the broader benchmark context for observing hands interacting with objects in natural visual streams that ARCTIC extends to precise 3D articulated meshes.
- Paper: EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations, Ahmad Darkhalil et al. (2022). Formulates foundational benchmarks for segmenting hands and active objects undergoing state transformations, highlighting the need for metric 3D dynamic contact datasets like ARCTIC.
- Paper: HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos, Prithviraj Banerjee et al. (2025). Extends the capture of 3D hand-object interactions to multi-view egocentric video using wearable headsets, building on ARCTIC's foundational principles of metric 3D hand-object tracking.
- Paper: NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions, Juze Zhang et al. (2023). Builds on multi-view capture paradigms for human-object interactions by deploying neural modeling pipelines to handle complex contact and occlusion dynamics.
- Paper: RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation, Songming Liu et al. (2025). Applies principles of bimanual coordination and dexterous interaction modeling to train large diffusion foundation models for robotic dual-arm manipulation.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). Translates dynamic hand-object interaction and state evolution concepts into closed-loop vision-language-action policies for real-time dynamic manipulation.
