ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation

Zicong FanOmid TaheriDimitrios TzionasMuhammed KocabasManuel KaufmannMichael J. BlackOtmar Hilliges

article2023CVPR404 citations

Presents the first large-scale multi-view dataset featuring accurate 3D meshes and dynamic contact annotations for bimanual manipulation of articulated objects, establishing new benchmarks and baselines for spatio-temporally consistent 3D motion reconstruction and interaction field estimation.

Listen

Enabling artificial intelligence systems to understand physical interaction with the surrounding world remains a fundamental challenge in computer vision and robotics. While humans intuitively manipulate complex items—such as opening a laptop or cutting with scissors—machine learning models struggle to interpret these actions accurately. Existing research datasets have predominantly focused on simple, single-handed grasping of static, rigid objects, lacking the realistic data needed to model how two hands dynamically interact with moving, multi-part items.

The article aims to address this capability gap by introducing ARCTIC, a large-scale multimodal dataset designed to evaluate and benchmark physically consistent, two-handed manipulation of articulated objects. It establishes two core evaluation tasks: reconstructing synchronized three-dimensional motion from standard video and estimating dense spatial interaction fields between hands and objects.

To achieve this, the authors recorded 10 participants performing unconstrained manipulation and grasping across 11 articulated items, generating 2.1 million video frames. The experimental setup synchronized eight static allocentric video cameras and one head-mounted egocentric camera with an array of 54 high-resolution infrared motion capture cameras. Minimal markers and pre-scanned digital models allowed the precise recovery of three-dimensional full-body, hand, and articulated object meshes alongside dynamic contact data without interfering with natural movements. The authors also developed baseline neural network models—ArcticNet for motion reconstruction and InterField for distance estimation—testing both single-frame and recurrent temporal versions.

The findings show that ARCTIC captures a significantly broader range of hand postures and contact regions than prior benchmarks, showing extensive palm engagement beyond simple fingertip contact. In reconstruction evaluations, temporal baseline models consistently outperformed single-frame variants. For allocentric motion reconstruction, the temporal model reduced motion deviation error from 10.4 mm to 9.3 mm and lowered hand acceleration error from 5.7 to 5.0 meters per second squared, while achieving an object reconstruction success rate of approximately 73.5%. Similarly, for interaction field estimation, temporal modeling yielded smoother predictions and reduced distance errors to 8.7 mm from hand to object.

These results demonstrate that temporal context is critical for achieving smooth, physically consistent reconstructions and avoiding unnatural motion artifacts. Bridging this dataset gap provides essential foundation tools for advancing robotic manipulation, augmented and virtual reality, and human behavior analysis. Moving beyond static grasp assumptions reduces the risk of models failing in real-world environments where dynamic coordination is required.

For future development, the authors recommend using the interaction field representation as an explicit constraint to guide and refine pose estimation pipelines. Research efforts should also focus on generative models capable of synthesizing realistic bimanual interactions with multi-part objects and extending depth-based object tracking methods to account for human occlusion.

Key limitations include the controlled studio environment, the evaluation of 11 specific object categories, and the reliance on baseline architectures designed primarily to establish benchmark standards rather than reach peak performance. Nevertheless, given the high capture fidelity and precise marker-based motion alignment, stakeholders can have high confidence in the dataset as an empirical standard for training and benchmarking advanced manipulation systems.

Cover for ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation

Abstract

Humans intuitively understand that inanimate objects do not move by themselves, but that state changes are typically caused by human manipulation (e.g., the opening of a book). This is not yet the case for machines. In part this is because there exist no datasets with ground-truth 3D annotations for the study of physically consistent and synchronised motion of hands and articulated objects. To this end, we introduce ARCTIC – a dataset of two hands that dexterously manipulate objects, containing 2.1M video frames paired with accurate 3D hand and object meshes and detailed, dynamic contact information. It contains bi-manual articulation of objects such as scissors or laptops, where hand poses and object states evolve jointly in time. We propose two novel articulated hand-object interaction tasks: (1) Consistent motion reconstruction: Given a monocular video, the goal is to reconstruct two hands and articulated objects in 3D, so that their motions are spatio-temporally consistent. (2) Interaction field estimation: Dense relative hand-object distances must be estimated from images. We introduce two baselines ArcticNet and InterField, respectively and evaluate them qualitatively and quantitatively on ARCTIC. Our code and data are available at https://arctic.is.tue.mpg.de.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. ARCTIC Dataset
  • 3.1. Data Characteristics
  • 3.2. Acquisition Setup
  • 4. Evaluation Protocol
  • 5. Baselines and Experiments
  • 5.1. Consistent motion reconstruction
  • 5.2. Interaction field estimation
  • 6. Conclusions
  • References

Knowls

  1. Knowl 1 — ARCTIC dataset composition and scope

    data/table

    ARCTIC (ARticulated objeCTs in InteraCtion) is a dataset of dexterous bimanual manipulation of articulated objects. It contains 339 sequences recorded from 10 subjects, including 5 females, manipulating 11 articulated objects. The data comprise approximately 2.1 million RGB frames: 1.7 million frames labeled as subjects “using” objects and 457,000 frames labeled as subjects “grasping” objects.

    Each frame is paired with accurate 3D meshes for both hands and the articulated object, as well as full-body pose represented with SMPL-X. The RGB data come from eight synchronized static allocentric views and one moving egocentric view; the cameras capture 2800 × 2000 images at 30 FPS. Depth images of the hands, body, and objects can also be rendered from the 3D annotations.

    The objects consist of two rigid parts that rotate around a shared axis. Compared with the existing datasets considered by the authors, ARCTIC uniquely combines articulated objects, both hands, full-body annotations, dexterous manipulation, calibrated allocentric and egocentric views, and motion-capture-based annotations. The unconstrained bimanual interactions produce substantially greater hand-pose diversity and more spatially distributed contact regions than the grasping-focused datasets used for comparison.

  2. Knowl 2 — Motion-capture acquisition and 3D annotation pipeline

    model/method

    ARCTIC synchronizes calibrated multi-view RGB video with a Vicon motion-capture system containing 54 high-resolution infrared cameras. Small hemispherical markers with 1.5 mm radius are placed mainly on the dorsal sides of the hands and on the objects so that they minimally interfere with natural interaction while remaining trackable during fast motion and occlusion.

    The annotation pipeline has five main stages: (1) build personalized canonical subject templates and scanned object geometries; (2) estimate the rotation axis of each two-part articulated object; (3) capture synchronized RGB video and marker-based motion; (4) fit body, hand, and object poses to the marker observations; and (5) compute hand-object contact from mesh proximity. Subject templates are obtained by fitting SMPL-X to 3D scans of each subject in a canonical T-pose and additional poses. Objects are scanned with a handheld 3D scanner, separated into two articulated parts, and assigned a canonical rest pose.

    Marker locations are associated with subject or object mesh vertices and refined using MoSh++. SMPL-X pose is optimized so that the personalized body and hand surfaces explain the observed marker positions. Each object is represented by a 6D rigid pose for its base part plus a 1D articulation angle relative to the canonical pose. The base pose is recovered from the rigid transformation between the current object markers and their canonical mesh correspondences; the articulation angle is computed from the estimated rotation axis and rest pose.

  3. Knowl 3 — Two articulated hand-object interaction tasks

    definition

    ARCTIC defines two benchmark tasks.

    Consistent motion reconstruction takes a monocular RGB video as input and reconstructs, for every frame, the 3D motion of the left hand, right hand, and articulated object. A successful reconstruction should jointly preserve accurate hand and object poses, hand-object contact, object articulation, and temporally smooth motion. In particular, vertices that remain in contact while a hand moves or an object articulates should move consistently together.

    Interaction field estimation takes RGB images from a video and predicts dense relative spatial relationships rather than only binary contact. For every vertex on either hand, the task estimates its shortest distance to the object mesh; for every object vertex, it estimates its shortest distance to each hand mesh. The four target fields are left-hand-to-object, right-hand-to-object, object-to-left-hand, and object-to-right-hand distances.

  4. Knowl 4 — ArcticNet articulated motion reconstruction baseline

    model/method

    ArcticNet reconstructs two hands and one articulated object from RGB images. Its hand representation is MANO: each hand has pose parameters b8 in b2^{48}, including global orientation, and shape parameters b2 in b2^{10}. MANO maps these parameters to a hand mesh H(b8,b2) in b2^{778 7 3}; a fixed learned linear regressor converts the mesh vertices into 3D joint locations. Each hand also has a 3D translation.

    The scanned articulated object is represented by a function O(a9) that outputs a mesh with VV vertices. Its pose a9 in b2^7 contains one articulation rotation c9 in b2, a 3D object rotation R_o in b2^3 represented with axis-angle coordinates, and a 3D translation T_o in b2^3.

    The single-frame baseline, ArcticNet-SF, uses a CNN image encoder followed by separate decoders for the left hand, right hand, and object. The hand decoders regress MANO parameters and hand translations, while the object decoder regresses the articulated object pose. Weak-perspective projection is used for image-space translation estimation. ArcticNet-LSTM has the same regression structure but first aggregates image features over a temporal window with an LSTM, allowing hand and object predictions to reason jointly over time. The baselines are trained using ground-truth 3D keypoints, projected 2D keypoints, and hand and object model parameters.

  5. Knowl 5 — Interaction field representation and InterField baseline

    model/method

    For two meshes MaM_a and MbM_b, let VaV_a and VbV_b be their vertex counts, and let v_i^a,v_j^b in b2^3 denote their vertices. The interaction field from mesh MaM_a to mesh MbM_b is the vector of shortest vertex-to-vertex distances:

    Fia7b=min⁡1≤j≤Vb∥via−vjb∥2,1≤i≤Va.F_i^{a 7 b}=\min_{1\leq j\leq V_b}\left\lVert v_i^a-v_j^b\right\rVert_2,\qquad 1\leq i\leq V_a.

    Here F^{a 7 b}in b2^{V_a} assigns one nonnegative distance to every vertex of MaM_a. ARCTIC requires prediction of Fl7oF^{l 7 o}, Fr7oF^{r 7 o}, Fo7lF^{o 7 l}, and Fo7rF^{o 7 r}, where ll, rr, and oo denote the left hand, right hand, and object meshes.

    InterField estimates these fields from RGB images. A CNN first produces an image feature vector xin b2^d. For a target mesh, each subsampled canonical-pose vertex v_iin b2^3 is concatenated with the image feature to form p_i=[x;v_i]in b2^{d+3}. A PointNet processes the set of concatenated vectors, and a regression head predicts one distance per subsampled vertex; the predictions are then upsampled to the full mesh. The four fields share the CNN and PointNet but use different regression heads. InterField-SF processes one frame at a time, whereas InterField-LSTM aggregates image features over a temporal window before field regression. Ground-truth meshes are used only to visualize predicted fields, not as network inputs.

  6. Knowl 6 — Subject-disjoint evaluation protocols

    experimental setup

    ARCTIC is split by subject into eight training subjects, one male validation subject, and one female test subject. The same split supports two camera protocols.

    In the allocentric protocol, models are trained and evaluated only with images from the eight static third-person views. In the egocentric protocol, training can use images from all available views for the training subjects, but evaluation uses only images from the moving first-person view. This design evaluates both conventional third-person reconstruction and a mixed-reality-like egocentric setting while keeping the subjects disjoint across training, validation, and testing.

  7. Knowl 7 — Metrics for consistent motion reconstruction

    equation

    ARCTIC evaluates reconstruction using contact, motion, pose, articulation, object, and relative-position metrics. All distances are measured in millimeters unless stated otherwise.

    For a frame, let (hi,oi)(h_i,o_i) be one of CC ground-truth hand-object vertex pairs whose ground-truth distance is below 3 mm, and let (h^i,o^i)(\hat h_i,\hat o_i) be the corresponding predicted vertices. Contact deviation is

    CDev=1C∑i=1C∥h^i−o^i∥2.\mathrm{CDev}=\frac{1}{C}\sum_{i=1}^{C}\left\lVert\hat h_i-\hat o_i\right\rVert_2.

    For motion deviation, let hith_i^t and ojto_j^t be hand and object vertices at frame tt. A ground-truth stable-contact window (m,n)(m,n) satisfies ∥hit−ojt∥2≤α\lVert h_i^t-o_j^t\rVert_2\leq\alpha for every t∈{m,…,n}t\in\{m,\ldots,n\}, with α=3\alpha=3 mm; only windows of at least 15 frames, corresponding to at least 0.5 seconds, are retained. For predicted vertices, define δh^it=h^it−h^it−1\delta\hat h_i^t=\hat h_i^t-\hat h_i^{t-1} and δo^jt=o^jt−o^jt−1\delta\hat o_j^t=\hat o_j^t-\hat o_j^{t-1}. Motion deviation is

    MDev=1n−m∑t=m+1n∥δh^it−δo^jt∥2,\mathrm{MDev}=\frac{1}{n-m}\sum_{t=m+1}^{n}\left\lVert\delta\hat h_i^t-\delta\hat o_j^t\right\rVert_2,

    averaged over all detected stable-contact windows. Acceleration error (ACC) is the difference between predicted and ground-truth vertex accelerations, reported in m/s2\mathrm{m/s^2} after subtracting the root trajectory of each entity; the object root is the center of its base.

    Mean per-joint position error (MPJPE) is the root-relative Euclidean error over the 21 hand joints. Average articulation error (AAE) is the absolute error in the predicted object articulation angle, reported in degrees. For an object with diameter DD, VoV_o vertices, ground-truth vertices oio_i, predicted vertices o^i\hat o_i, and object roots removed, the object success rate is

    SuccessRate=1Vo∑i=1Vo1 ⁣(∥oi−o^i∥2<0.05D)×100%.\mathrm{SuccessRate}=\frac{1}{V_o}\sum_{i=1}^{V_o}\mathbf{1}\!\left(\left\lVert o_i-\hat o_i\right\rVert_2<0.05D\right)\times100\%.

    For entities a,b∈{l,r,o}a,b\in\{l,r,o\}, where ll and rr are the two hands and oo is the object, and where J0a,J0b∈R3J_0^a,J_0^b\in\mathbb{R}^3 are ground-truth root locations while J^0a,J^0b\hat J_0^a,\hat J_0^b are predicted roots, the mean relative-root position error is

    MRRPEa→b=∥(J0a−J0b)−(J^0a−J^0b)∥2.\mathrm{MRRPE}_{a\rightarrow b}=\left\lVert (J_0^a-J_0^b)-(\hat J_0^a-\hat J_0^b)\right\rVert_2.
  8. Knowl 8 — Metrics for interaction field estimation

    equation

    For a predicted interaction field F^a→b\hat F^{a\rightarrow b} and ground-truth field Fa→bF^{a\rightarrow b}, the average distance error is the mean absolute error over the VaV_a vertices of source mesh MaM_a:

    AverageDistanceError=1Va∑i=1Va∣Fia→b−F^ia→b∣.\mathrm{AverageDistanceError}=\frac{1}{V_a}\sum_{i=1}^{V_a}\left|F_i^{a\rightarrow b}-\hat F_i^{a\rightarrow b}\right|.

    Each field value is a distance in millimeters. The evaluation considers both hand-to-object and object-to-hand fields and averages the corresponding metrics over the two hands when reporting a single summary value.

    To measure temporal smoothness, the field is predicted for every frame of a sequence. The acceleration error is the mean absolute difference between the predicted and ground-truth acceleration sequences of the field values, reported in m/s2\mathrm{m/s^2}. Lower average distance error indicates more accurate relative spatial distances, while lower acceleration error indicates smoother temporal field predictions.

  9. Knowl 9 — ArcticNet reconstruction results

    data/table

    The table compares the single-frame ArcticNet-SF and temporal ArcticNet-LSTM baselines on the subject-disjoint ARCTIC validation and test splits. CDev measures contact deviation, MRRPE measures relative root position error, MDev measures disagreement in the motion of stable-contact vertices, ACC measures acceleration error, MPJPE measures root-relative hand-joint error, AAE measures object articulation error, and Success Rate measures object-vertex accuracy. For compactness, hh averages the two hands; r/lr/l and r/or/o denote right-hand-to-left-hand and right-hand-to-object relative-root errors.

    Could not parse LaTeX table

    Temporal modeling generally improves contact consistency, stable-contact motion, and smoothness: ArcticNet-LSTM lowers MDev and ACC on every split and lowers CDev on both allocentric splits and the egocentric test split. It also improves allocentric object success from 71.8% to 74.9% on validation and from 71.4% to 73.5% on testing. The egocentric results remain substantially harder, with test success rates of 53.9% for ArcticNet-SF and 53.5% for ArcticNet-LSTM.

  10. Knowl 10 — InterField interaction-distance results

    data/table

    The table compares single-frame InterField-SF with temporal InterField-LSTM. Each entry reports hand-to-object/object-to-hand values in that order; average distance error is in millimeters and ACC is in m/s2\mathrm{m/s^2}. Lower values are better for both metrics.

    Could not parse LaTeX table

    InterField-LSTM has lower average distance error and lower acceleration error than InterField-SF for both directions on all validation and test splits. The strongest temporal improvement occurs on the allocentric validation split, where hand-to-object/object-to-hand distance error decreases from 9.6/9.9 mm to 9.0/8.9 mm and acceleration error decreases from 3.0/2.9 to 2.1/2.0 m/s2\mathrm{m/s^2}. The qualitative predictions form heatmaps that correlate with ground-truth proximity patterns, including regions that are near the object without being in binary contact.

Coverage note — Detailed per-object contact heatmaps, qualitative reconstruction figures, and the full comparison table of prior datasets were omitted because they primarily visualize or corroborate the dataset properties and baseline behavior already captured here, rather than adding separate load-bearing methods or results.

References

  1. 1.Luca Ballan, Aparna Taneja, Jürgen Gall, Luc Van Gool, and Marc Pollefeys. Motion capture of hands in action using discriminative salient points. In European Conference on Computer Vision (ECCV), pages 640–653, 2012.
  2. 2.Keni Bernardin, Koichi Ogawara, Katsushi Ikeuchi, and Ruediger Dillmann. A sensor fusion approach for recognizing continuous human grasping sequences using hidden markov models. Transactions on Robotics, 21(1):47–57, 2005.
  3. 3.Bharat Lal Bhatnagar, Xianghui Xie, Ilya A Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. BEHAVE: Dataset and method for tracking human object interactions. In Computer Vision and Pattern Recognition (CVPR), pages 15935–15946, 2022.
  4. 4.Michael J. Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang. BEDLAM: A dataset of bodies exhibiting detailed lifelike animated motion. In Computer Vision and Pattern Recognition (CVPR), June 2023.
  5. 5.Adnane Boukhayma, Rodrigo de Bem, and Philip H. S. Torr. 3D hand shape and pose from images in the wild. In Computer Vision and Pattern Recognition (CVPR), pages 10843–10852, 2019.
  6. 6.Samarth Brahmbhatt, Chengcheng Tang, Christopher D. Twigg, Charles C. Kemp, and James Hays. ContactPose: A dataset of grasps with object contact and hand pose. In European Conference on Computer Vision (ECCV), volume 12358, pages 361–378, 2020.
  7. 7.Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa, and Jitendra Malik. Reconstructing hand-object interactions in the wild. In International Conference on Computer Vision (ICCV), pages 12417–12426, 2021.
  8. 8.Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S. Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox. DexYCB: A benchmark for capturing hand grasping of objects. In Computer Vision and Pattern Recognition (CVPR), pages 9044–9053, 2021.
  9. 9.Yixin Chen, Sai Kumar Dwivedi, Michael J. Black, and Dimitrios Tzionas. Detecting human-object contact in images. In Computer Vision and Pattern Recognition (CVPR), June 2023.
  10. 10.Yuanpei Chen, Yaodong Yang, Tianhao Wu, Shengjie Wang, Xidong Feng, Jiechuang Jiang, Stephen Marcus McAleer, Hao Dong, Zongqing Lu, and Song-Chun Zhu. Towards human-level bimanual dexterous manipulation with reinforcement learning. arXiv preprint arXiv:2206.08686, 2022.
  11. 11.Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. D-Grasp: Physically plausible dynamic grasp synthesis for hand-object interactions. In Computer Vision and Pattern Recognition (CVPR), pages 20545–20554, 2022.
  12. 12.Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gregory Rogez. GanHand: Predicting human grasp affordances in multi-object scenes. In Computer Vision and Pattern Recognition (CVPR), pages 5030–5040, 2020.
  13. 13.Zicong Fan, Adrian Spurr, Muhammed Kocabas, Siyu Tang, Michael J. Black, and Otmar Hilliges. Learning to disambiguate strongly interacting hands via probabilistic per-pixel part segmentation. In International Conference on 3D Vision (3DV), pages 1–10, 2021.
  14. 14.Thomas Feix, Javier Romero, Heinz-Bodo Schmiedmayer, Aaron M. Dollar, and Danica Kragic. The grasp taxonomy of human grasp types. Transactions on Human-Machine Systems (THMS), 46(1):66–77, 2016.
  15. 15.Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek, and Tae-Kyun Kim. First-person hand action benchmark with RGB-D videos and 3D hand pose annotations. In Computer Vision and Pattern Recognition (CVPR), pages 409–419, 2018.
  16. 16.Patrick Grady, Chengcheng Tang, Samarth Brahmbhatt, Christopher D. Twigg, Chengde Wan, James Hays, and Charles C. Kemp. PressureVision: Estimating hand pressure from a single RGB image. European Conference on Computer Vision (ECCV), 13666:328–345, 2022.
  17. 17.Patrick Grady, Chengcheng Tang, Christopher D. Twigg, Minh Vo, Samarth Brahmbhatt, and Charles C. Kemp. ContactOpt: Optimizing contact to improve grasps. In Computer Vision and Pattern Recognition (CVPR), pages 1471–1481, 2021.
  18. 18.Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D annotation of hand and object poses. In Computer Vision and Pattern Recognition (CVPR), pages 3193–3203, 2020.
  19. 19.Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint transformer: Solving joint identification in challenging hands and object interactions for accurate 3D pose estimation. In Computer Vision and Pattern Recognition (CVPR), pages 11090–11100, 2022.
  20. 20.Yana Hasson, Bugra Tekin, Federica Bogo, Ivan Laptev, Marc Pollefeys, and Cordelia Schmid. Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction. In Computer Vision and Pattern Recognition (CVPR), pages 568–577, 2020.
  21. 21.Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J. Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. In Computer Vision and Pattern Recognition (CVPR), pages 11807–11816, 2019.
  22. 22.Chun-Hao P. Huang, Hongwei Yi, Markus Hoschle, Matvey Safroshkin, Tsvetelina Alexiadis, Senya Polikovsky, Daniel Scharstein, and Michael J. Black. Capturing and inferring dense full-body human-scene contact. In Computer Vision and Pattern Recognition (CVPR), pages 13274–13285, June 2022.
  23. 23.Yinghao Huang, Omid Taheri, Michael J. Black, and Dimitrios Tzionas. InterCap: Joint markerless 3D tracking of humans and objects in interaction. In German Conference on Pattern Recognition (GCPR), volume 13485, pages 281–299, 2022.
  24. 24.Umar Iqbal, Pavlo Molchanov, Thomas Breuel Juergen Gall, and Jan Kautz. Hand pose estimation via latent 2.5D heatmap regression. In European Conference on Computer Vision (ECCV), pages 118–134, 2018.
  25. 25.Noriko Kamakura, Michiko Matsuo, Harumi Ishii, Fumiko Mitsuboshi, and Yoriko Miura. Patterns of static prehension in normal hands. American Journal of Occupational Therapy, 34(7):437–445, 1980.
  26. 26.Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In Computer Vision and Pattern Recognition (CVPR), pages 7122–7131, 2018.
  27. 27.Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J. Black, Krikamol Muandet, and Siyu Tang. Grasping Field: Learning implicit representations for human grasps. In International Conference on 3D Vision (3DV), pages 333–344, 2020.
  28. 28.Muhammed Kocabas, Nikos Athanasiou, and Michael J. Black. VIBE: Video inference for human body pose and shape estimation. In Computer Vision and Pattern Recognition (CVPR), pages 5253–5263, 2020.
  29. 29.Muhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, and Michael J. Black. PARE: Part attention regressor for 3D human body estimation. In International Conference on Computer Vision (ICCV), pages 11127–11137, 2021.
  30. 30.Taein Kwon, Bugra Tekin, Jan Stuhmer, Federica Bogo, and Marc Pollefeys. H2O: Two hands manipulating objects for first person interaction recognition. In International Conference on Computer Vision (ICCV), pages 10138–10148, 2021.
  31. 31.Mengcheng Li, Liang An, Hongwen Zhang, Lianpeng Wu, Feng Chen, Tao Yu, and Yebin Liu. Interacting attention graph for single image two-hand reconstruction. In Computer Vision and Pattern Recognition (CVPR), pages 2761–2770, 2022.
  32. 32.Xiaolong Li, He Wang, Li Yi, Leonidas J. Guibas, A. Lynn Abbott, and Shuran Song. Category-level articulated object pose estimation. In Computer Vision and Pattern Recognition (CVPR), pages 3703–3712, 2020.
  33. 33.Shaowei Liu, Hanwen Jiang, Jiarui Xu, Sifei Liu, and Xiaolong Wang. Semi-supervised 3D hand-object poses estimation with interactions in time. In Computer Vision and Pattern Recognition (CVPR), pages 14687–14697, 2021.
  34. 34.Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. In Computer Vision and Pattern Recognition (CVPR), pages 21013–21022, 2022.
  35. 35.Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: Archive of motion capture as surface shapes. In International Conference on Computer Vision (ICCV), pages 5441–5450, 2019.
  36. 36.Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Kyoung Mu Lee. InterHand2.6M: A dataset and baseline for 3D interacting hand pose estimation from a single RGB image. In European Conference on Computer Vision (ECCV), volume 12365, pages 548–564, 2020.
  37. 37.Franziska Mueller, Florian Bernard, Oleksandr Sotnychenko, Dushyant Mehta, Srinath Sridhar, Dan Casas, and Christian Theobalt. GANerated hands for real-time 3D hand tracking from monocular RGB. In Computer Vision and Pattern Recognition (CVPR), pages 49–59, 2018.
  38. 38.Franziska Mueller, Dushyant Mehta, Oleksandr Sotnychenko, Srinath Sridhar, Dan Casas, and Christian Theobalt. Real-time hand tracking under occlusion from an egocentric RGB-D sensor. In International Conference on Computer Vision (ICCV), pages 1163–1172, 2017.
  39. 39.Supreeth Narasimhaswamy, Trung Nguyen, and Minh Hoai Nguyen. Detecting hands and recognizing physical contact in the wild. In Conference on Neural Information Processing Systems (NeurIPS), volume 33, 2020.
  40. 40.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Computer Vision and Pattern Recognition (CVPR), pages 10975–10985, 2019.
  41. 41.Tu-Hoa Pham, Nikolaos Kyriazis, Antonis A. Argyros, and Abderrahmane Kheddar. Hand-object contact force estimation from markerless visual tracking. Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 40(12):2883–2896, 2018.
  42. 42.Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. PointNet: Deep learning on point sets for 3D classification and segmentation. In Computer Vision and Pattern Recognition (CVPR), pages 77–85, 2017.
  43. 43.James M. Rehg and Takeo Kanade. Visual tracking of high DOF articulated structures: An application to human hand tracking. In European Conference on Computer Vision (ECCV), volume 801, pages 35–46, 1994.
  44. 44.Gregory Rogez, James Steven Supancić III, and Deva Ramanan. Understanding everyday hands in action from RGB-D images. In International Conference on Computer Vision (ICCV), pages 3889–3897, 2015.
  45. 45.Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bodies together. Transactions on Graphics (TOG), 36(6):245:1–245:17, 2017.
  46. 46.István Sárándi, Timm Linder, Kai O. Arras, and Bastian Leibe. Metric-scale truncation-robust heatmaps for 3D human pose estimation. In International Conference on Automatic Face & Gesture Recognition (FG), pages 407–414, 2020.
  47. 47.Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In Computer Vision and Pattern Recognition (CVPR), pages 21064–21074, 2022.
  48. 48.Dandan Shan, Jiaqi Geng, Michelle Shu, and David F. Fouhey. Understanding human hands in contact at internet scale. In Computer Vision and Pattern Recognition (CVPR), pages 9866–9875, 2020.
  49. 49.Tomas Simon, Hanbyul Joo, Iain Matthews, and Yaser Sheikh. Hand keypoint detection in single images using multiview bootstrapping. In Computer Vision and Pattern Recognition (CVPR), pages 4645–4653, 2017.
  50. 50.Adrian Spurr, Aneesh Dahiya, Xi Wang, Xucong Zhang, and Otmar Hilliges. Self-supervised 3D hand pose estimation from monocular RGB via contrastive learning. In International Conference on Computer Vision (ICCV), pages 11210–11219, 2021.
  51. 51.Adrian Spurr, Umar Iqbal, Pavlo Molchanov, Otmar Hilliges, and Jan Kautz. Weakly supervised 3D hand pose estimation via biomechanical constraints. In European Conference on Computer Vision (ECCV), volume 12362, pages 211–228, 2020.
  52. 52.Adrian Spurr, Jie Song, Seonwook Park, and Otmar Hilliges. Cross-modal deep variational hand pose estimation. In Computer Vision and Pattern Recognition (CVPR), pages 89–98, 2018.
  53. 53.Srinath Sridhar, Franziska Mueller, Michael Zollhoefer, Dan Casas, Antti Oulasvirta, and Christian Theobalt. Real-time joint tracking of a hand manipulating an object from RGB-D input. In European Conference on Computer Vision (ECCV), volume 9906, pages 294–310, 2016.
  54. 54.Stefan Stevsiˇc and Otmar Hilliges. Spatial attention improves iterative 6D object pose estimation. In International Conference on 3D Vision (3DV), pages 1070–1078, 2020.
  55. 55.Omid Taheri, Vassileios Choutas, Michael J. Black, and Dimitrios Tzionas. GOAL: Generating 4D whole-body motion for hand-object grasping. In Computer Vision and Pattern Recognition (CVPR), pages 13253–13263, 2022.
  56. 56.Omid Taheri, Nima Ghorbani, Michael J. Black, and Dimitrios Tzionas. GRAB: A dataset of whole-body human grasping of objects. In European Conference on Computer Vision (ECCV), volume 12349, pages 581–600, 2020.
  57. 57.Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+O: Unified egocentric recognition of 3D hand-object poses and interactions. In Computer Vision and Pattern Recognition (CVPR), pages 4511–4520, 2019.
  58. 58.3dMDhand system series. https : / / 3dmd . com / products/.
  59. 59.Shashank Tripathi, Lea Muller, Chun-Hao P. Huang, Taheri Omid, Michael J. Black, and Dimitrios Tzionas. 3D human pose estimation via intuitive physics. In Computer Vision and Pattern Recognition (CVPR), June 2023.
  60. 60.Aggeliki Tsoli and Antonis A. Argyros. Joint 3D tracking of a deformable object in interaction with a hand. In European Conference on Computer Vision (ECCV), volume 11218, pages 504–520, 2018.
  61. 61.Dimitrios Tzionas, Luca Ballan, Abhilash Srikantha, Pablo Aponte, Marc Pollefeys, and Juergen Gall. Capturing hands in action using discriminative salient points and physics simulation. International Journal of Computer Vision (IJCV), 118(2):172–193, 2016.
  62. 62.Dimitrios Tzionas and Juergen Gall. A comparison of directional distances for hand pose estimation. In German Conference on Pattern Recognition (GCPR), volume 8142, pages 131–141, 2013.
  63. 63.Dimitrios Tzionas and Juergen Gall. 3D object reconstruction from hand-object interactions. In International Conference on Computer Vision (ICCV), pages 729–737, 2015.
  64. 64.Dimitrios Tzionas and Juergen Gall. Reconstructing articulated rigged models from RGB-D videos. In European Conference on Computer Vision Workshops (ECCVw), volume 9915, pages 620–633, 2016.
  65. 65.Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-SNE. Journal of Machine Learning Research (JMLR), 9(86):2579–2605, 2008.
  66. 66.Vicon Vantage: Cutting edge, flagship camera with intelligent feedback and resolution. https://www.vicon.com/hardware/cameras/vantage.
  67. 67.Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu. CPF: Learning a contact potential field to model the hand-object interaction. In International Conference on Computer Vision (ICCV), pages 11097–11106, 2021.
  68. 68.Yifei Yin, Chen Guo, Manuel Kaufmann, Juan Zarate, Jie Song, and Otmar Hilliges. Hi4D: 4D instance segmentation of close human interaction. In Computer Vision and Pattern Recognition (CVPR), 2023.
  69. 69.Sergey Zakharov, Ivan Shugurov, and Slobodan Ilic. DPOD: 6D pose object detector and refiner. In International Conference on Computer Vision (ICCV), pages 1941–1950, 2019.
  70. 70.Baowen Zhang, Yangang Wang, Xiaoming Deng, Yinda Zhang, Ping Tan, Cuixia Ma, and Hongan Wang. Interacting two-hand 3D pose and shape reconstruction from single color image. In International Conference on Computer Vision (ICCV), pages 11354–11363, 2021.
  71. 71.He Zhang, Yuting Ye, Takaaki Shiratori, and Taku Komura. ManipNet: Neural manipulation synthesis with a hand-object spatial representation. Transactions on Graphics (TOG), 40(4):1–14, 2021.
  72. 72.Hao Zhang, Yuxiao Zhou, Yifei Tian, Jun-Hai Yong, and Feng Xu. Single depth view based real-time reconstruction of hand-object interactions. Transactions on Graphics (TOG), 40(3):29:1–29:12, 2021.
  73. 73.Xiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang, and Wen Zheng. End-to-end hand mesh recovery from a monocular RGB image. In International Conference on Computer Vision (ICCV), pages 2354–2364, 2019.
  74. 74.Keyang Zhou, Bharat Lal Bhatnagar, Jan Eric Lenssen, and Gerard Pons-Moll. TOCH: Spatio-temporal object-to-hand correspondence for motion refinement. In European Conference on Computer Vision (ECCV), volume 13663, pages 1–19, 2022.
  75. 75.Yuxiao Zhou, Marc Habermann, Weipeng Xu, Ikhsanul Habibie, Christian Theobalt, and Feng Xu. Monocular real-time hand shape and motion capture using multi-modal data. In Computer Vision and Pattern Recognition (CVPR), pages 5345–5354, 2020.
  76. 76.Andrea Ziani, Zicong Fan, Muhammed Kocabas, Sammy Christen, and Otmar Hilliges. TempCLR: Reconstructing hands via time-coherent contrastive learning. In International Conference on 3D Vision (3DV), pages 627–636, 2022.
  77. 77.Christian Zimmermann and Thomas Brox. Learning to estimate 3D hand pose from single RGB images. In International Conference on Computer Vision (ICCV), pages 4913–4921, 2017.
  78. 78.Christian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan Russell, Max Argus, and Thomas Brox. FreiHAND: A dataset for markerless capture of hand pose and shape from single RGB images. In International Conference on Computer Vision (ICCV), pages 813–822, 2019.

Citation

MLA
Fan, Z., et al. “ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation”. arXiv, 2022, http://arxiv.org/abs/2204.13662v3.
APA
Fan, Z., Taheri, O., Tzionas, D., Kocabas, M., Kaufmann, M., Black, M. J., & Hilliges, O. (2022). ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation. arXiv. http://arxiv.org/abs/2204.13662v3
Chicago
Fan, Z., O. Taheri, D. Tzionas, et al. 2022. “ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation”. arXiv. http://arxiv.org/abs/2204.13662v3.
Harvard
Fan, Z. et al. (2022) “ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.13662v3.
Vancouver
1. Fan Z, Taheri O, Tzionas D, Kocabas M, Kaufmann M, Black MJ, Hilliges O (2022) ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation. arXiv

BibTeX

@article{fan2022arctic,
  title = {ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation},
  author = {Fan, Zicong and Taheri, Omid and Tzionas, Dimitrios and Kocabas, Muhammed and Kaufmann, Manuel and Black, Michael J. and Hilliges, Otmar},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.13662v3},
  eprint = {2204.13662}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE