AMASS: Archive of Motion Capture As Surface Shapes
Naureen MahmoodNima GhorbaniNikolaus F. TrojeGerard Pons-MollMichael J. Black
Presents AMASS, a unified database that standardizes 15 distinct motion capture datasets into rigged 3D human surface meshes across 40 hours of movement, establishing a massive training resource for deep learning and computer animation.
Data-driven artificial intelligence and computer animation require vast amounts of high-quality human motion data to train realistic models. While numerous optical marker-based motion capture collections exist, they are historically fragmented across incompatible skeletal frameworks, varying body dimensions, and disparate marker configurations. Standard methods typically discard realistic body shape and soft-tissue movements by treating non-rigid skin deformations as measurement noise, or alter original movements through artificial skeleton retargeting. This fragmentation has severely restricted the scale and visual fidelity of training data available for modern machine learning workflows.
The article establishes a method to reconstruct metrically accurate 3D human body shape, pose, hand articulation, and dynamic soft-tissue motion directly from sparse motion capture markers. Leveraging this approach, the authors demonstrate the creation of the Archive of Mocap as Surface Shapes (AMASS), a unified, large-scale meta-dataset designed to standardize previously incompatible archival motion capture collections without sacrificing individual body morphology or movement dynamics.
The researchers developed an optimization framework called MoSh++ that fits a standardized 3D body mesh to sparse marker data in two distinct stages. The pipeline integrates the Skinned Multi-Person Linear (SMPL) model, MANO hand representations, and DMPL dynamic soft-tissue deformations. To tune algorithm parameters and evaluate geometric accuracy, the team recorded a new Synchronized Scans and Markers benchmark combining optical marker tracking with high-resolution four-dimensional body scanning across 30 dynamic motion trials. Using this calibrated framework, the authors unified 15 distinct international motion capture datasets encompassing diverse marker layouts ranging from 37 to 91 markers.
The evaluation demonstrates significant technical and practical advancements over previous surface-fitting methods. On the benchmark dataset, the new method reduced average 3D body shape reconstruction error from 12.1 millimeters down to 7.4 millimeters using a standard 46-marker layout, representing an approximate 39% improvement in geometric accuracy. When tracking dynamic motion with soft-tissue estimation, reconstruction error dropped from 10.24 millimeters to 7.3 millimeters, an improvement of roughly 29%. Furthermore, the framework operates with high parameter efficiency, achieving superior surface recovery using only 16 shape and 8 dynamic components compared to older systems requiring 100 components. The resulting AMASS meta-dataset consolidates more than 40 hours of motion data spanning 346 unique subjects and over 11,000 distinct motion sequences.
These findings provide an essential infrastructure for computer vision, robotics, and digital graphics. By offering consistent parameters that seamlessly plug into standard game engines and graphics packages, the archive eliminates the need to normalize subjects to identical body proportions. Practitioners can directly render textured, realistic virtual characters or extract custom skeletons, significantly lowering the cost and complexity of generating synthetic training data for deep learning algorithms.
Organizations developing motion models can immediately leverage the publicly available dataset and adopt the conversion framework to standardize existing proprietary motion libraries. To maximize practical utility, future efforts should prioritize transitioning the optimization pipeline from central processing units to parallel graphics processing unit frameworks to achieve real-time tracking performance. The authors also outline the expansion of the benchmark to include dense hand ground-truth and facial capture integration via compatible statistical head models.
While the reconstructed dataset is robust and comprehensive, users must account for operational limitations. The current fitting process runs offline at roughly two seconds per frame for full dynamic optimization, and extreme multi-marker occlusions require automated regularizer weighting that may yield minor smoothing of fast limb dynamics. Nevertheless, given the rigorous validation against synchronized four-dimensional ground-truth scans, there is high confidence in the geometric accuracy and physiological realism of the unified motion library.
- Paper: SMPL, M. Loper et al. (2015). Introduces the SMPL parametric body model that serves as the core rigged 3D surface mesh representation underpinning AMASS and MoSh++.
- Paper: Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments, Catalin Ionescu et al. (2014). Establishes a foundational multi-camera 3D human sensing dataset, demonstrating the limitations of isolated capture benchmarks that AMASS unifies into a common mesh framework.
- Paper: Expressive Body Capture: 3D Hands, Face, and Body From a Single Image, Georgios Pavlakos et al. (2019). Extends SMPL parametric body fitting to include expressive hands and facial details (SMPL-X), leveraging motion capture priors aligned with AMASS-style unified 3D representations.
