AMASS: Archive of Motion Capture As Surface Shapes

Naureen MahmoodNima GhorbaniNikolaus F. TrojeGerard Pons-MollMichael J. Black

article2019ICCV2,008 citations

Presents AMASS, a unified database that standardizes 15 distinct motion capture datasets into rigged 3D human surface meshes across 40 hours of movement, establishing a massive training resource for deep learning and computer animation.

Listen

Data-driven artificial intelligence and computer animation require vast amounts of high-quality human motion data to train realistic models. While numerous optical marker-based motion capture collections exist, they are historically fragmented across incompatible skeletal frameworks, varying body dimensions, and disparate marker configurations. Standard methods typically discard realistic body shape and soft-tissue movements by treating non-rigid skin deformations as measurement noise, or alter original movements through artificial skeleton retargeting. This fragmentation has severely restricted the scale and visual fidelity of training data available for modern machine learning workflows.

The article establishes a method to reconstruct metrically accurate 3D human body shape, pose, hand articulation, and dynamic soft-tissue motion directly from sparse motion capture markers. Leveraging this approach, the authors demonstrate the creation of the Archive of Mocap as Surface Shapes (AMASS), a unified, large-scale meta-dataset designed to standardize previously incompatible archival motion capture collections without sacrificing individual body morphology or movement dynamics.

The researchers developed an optimization framework called MoSh++ that fits a standardized 3D body mesh to sparse marker data in two distinct stages. The pipeline integrates the Skinned Multi-Person Linear (SMPL) model, MANO hand representations, and DMPL dynamic soft-tissue deformations. To tune algorithm parameters and evaluate geometric accuracy, the team recorded a new Synchronized Scans and Markers benchmark combining optical marker tracking with high-resolution four-dimensional body scanning across 30 dynamic motion trials. Using this calibrated framework, the authors unified 15 distinct international motion capture datasets encompassing diverse marker layouts ranging from 37 to 91 markers.

The evaluation demonstrates significant technical and practical advancements over previous surface-fitting methods. On the benchmark dataset, the new method reduced average 3D body shape reconstruction error from 12.1 millimeters down to 7.4 millimeters using a standard 46-marker layout, representing an approximate 39% improvement in geometric accuracy. When tracking dynamic motion with soft-tissue estimation, reconstruction error dropped from 10.24 millimeters to 7.3 millimeters, an improvement of roughly 29%. Furthermore, the framework operates with high parameter efficiency, achieving superior surface recovery using only 16 shape and 8 dynamic components compared to older systems requiring 100 components. The resulting AMASS meta-dataset consolidates more than 40 hours of motion data spanning 346 unique subjects and over 11,000 distinct motion sequences.

These findings provide an essential infrastructure for computer vision, robotics, and digital graphics. By offering consistent parameters that seamlessly plug into standard game engines and graphics packages, the archive eliminates the need to normalize subjects to identical body proportions. Practitioners can directly render textured, realistic virtual characters or extract custom skeletons, significantly lowering the cost and complexity of generating synthetic training data for deep learning algorithms.

Organizations developing motion models can immediately leverage the publicly available dataset and adopt the conversion framework to standardize existing proprietary motion libraries. To maximize practical utility, future efforts should prioritize transitioning the optimization pipeline from central processing units to parallel graphics processing unit frameworks to achieve real-time tracking performance. The authors also outline the expansion of the benchmark to include dense hand ground-truth and facial capture integration via compatible statistical head models.

While the reconstructed dataset is robust and comprehensive, users must account for operational limitations. The current fitting process runs offline at roughly two seconds per frame for full dynamic optimization, and extreme multi-marker occlusions require automated regularizer weighting that may yield minor smoothing of fast limb dynamics. Nevertheless, given the rigorous validation against synchronized four-dimensional ground-truth scans, there is high confidence in the geometric accuracy and physiological realism of the unified motion library.

Cover for AMASS: Archive of Motion Capture As Surface Shapes

Abstract

Large datasets are the cornerstone of recent advances in computer vision using deep learning. In contrast, existing human motion capture (mocap) datasets are small and the motions limited, hampering progress on learning models of human motion. While there are many different datasets available, they each use a different parameterization of the body, making it difficult to integrate them into a single meta dataset. To address this, we introduce AMASS, a large and varied database of human motion that unifies 15 different optical marker-based mocap datasets by representing them within a common framework and parameterization. We achieve this using a new method, MoSh++, that converts mocap data into realistic 3D human meshes represented by a rigged body model; here we use SMPL [doi:https://doi.org/10.1145/2816795.2818013], which is widely used and provides a standard skeletal representation as well as a fully rigged surface mesh. The method works for arbitrary marker sets, while recovering soft-tissue dynamics and realistic hand motion. We evaluate MoSh++ and tune its hyperparameters using a new dataset of 4D body scans that are jointly recorded with marker-based mocap. The consistent representation of AMASS makes it readily useful for animation, visualization, and generating training data for deep learning. Our dataset is significantly richer than previous human motion collections, having more than 40 hours of motion data, spanning over 300 subjects, more than 11,000 motions, and will be publicly available to the research community.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Technical Approach
  • 3.1 The Body Model
  • 3.2 Model Fitting
  • 3.3 Optimization and Runtime
  • 4 Evaluation
  • 4.1 Synchronized Scans and Markers (SSM)
  • 4.2 Hyper-parameter Search using SSM
  • 4.3 Shape Estimation Evaluation
  • 4.4 Pose and Soft-tissue Estimation Evaluation
  • 4.5 Hand Articulation
  • 5 AMASS Dataset
  • 6 Future Work and Conclusions
  • References
  • 1 Acknowledgements
  • 2 Optimization and Runtime
  • 3 Data Collection
  • 4 Model Size
  • 5 Diversity and Quality

Knowls

  1. Knowl 1 — Unified Body Surface Model (SMPL-H + DMPL) in MoSh++

    model/method

    MoSh++ represents the moving human body by unifying the SMPL-H articulated body and hand model with the DMPL dynamic soft-tissue deformation model into a single generative surface mesh formulation. The posed 3D mesh surface S(β,θ,ϕ)S(\beta, \theta, \phi) containing N=6890N = 6890 vertices is computed via linear blend skinning:

    S(β,θ,ϕ)=G(T(β,θ,ϕ),J(β),θ,W)S(\beta, \theta, \phi) = G(T(\beta, \theta, \phi), J(\beta), \theta, W)

    where G(T,J,θ,W):R3N×R3K×R∣θ∣×R4×3N→R3NG(T, J, \theta, W) : \mathbb{R}^{3N} \times \mathbb{R}^{3K} \times \mathbb{R}^{|\theta|} \times \mathbb{R}^{4 \times 3N} \to \mathbb{R}^{3N} applies linear blend skinning using skeletal joint locations J(β)∈R3KJ(\beta) \in \mathbb{R}^{3K} for K=52K = 52 joints, blend weights WW, and pose parameters θ\theta. The rest-pose template mesh T(β,θ,ϕ)T(\beta, \theta, \phi) is defined additively as:

    T(β,θ,ϕ)=Tμ+Bs(β)+Bp(θ)+Bd(ϕ)T(\beta, \theta, \phi) = T_\mu + B_s(\beta) + B_p(\theta) + B_d(\phi)

    where Tμ∈R3NT_\mu \in \mathbb{R}^{3N} is the base mean template mesh in a canonical rest pose. The blendshape functions modify vertex offsets relative to TμT_\mu:

    • Bs(β)B_s(\beta) represents identity-dependent shape deformations parameterized by linear shape coefficients β∈R16\beta \in \mathbb{R}^{16}.
    • Bp(θ)B_p(\theta) represents pose-dependent blend deformations parameterized by skeletal pose θ\theta.
    • Bd(ϕ)B_d(\phi) represents dynamic soft-tissue deformations parameterized by linear coefficients ϕ∈R8\phi \in \mathbb{R}^{8} learned from dynamic 4D scans.

    The articulated structure consists of K=52K = 52 joints (n=24n = 24 body joints and 2828 hand joints, with 1414 joints per hand). The overall pose vector θ∈R159\theta \in \mathbb{R}^{159} consists of 3 rotational degrees of freedom in exponential coordinates for each of the 52 joints plus a 3-dimensional root translation vector γ∈R3\gamma \in \mathbb{R}^3 (3×52+3=1593 \times 52 + 3 = 159 parameters).

  2. Knowl 2 — MoSh++ Stage I: Subject Shape and Latent Marker Optimization

    algorithm

    The first stage of MoSh++ estimates subject-specific identity shape parameters β∈R16\beta \in \mathbb{R}^{16}, latent pose-invariant marker positions M~={m~i}\tilde{M} = \{\tilde{m}_i\}, and frame poses Θ={θt}t=1F\Theta = \{\theta_t\}_{t=1}^F across a subset of F=12F = 12 randomly chosen frames from a mocap sequence by minimizing an unconstrained objective function:

    E(M~,β,ΘB,ΘH)=λDED(M~,β,ΘB,ΘH)+λβEβ(β)+λθBEθB(θB)+λθHEθH(θH)+λRER(M~,β)+λIEI(M~,β)E(\tilde{M}, \beta, \Theta_B, \Theta_H) = \lambda_D E_D(\tilde{M}, \beta, \Theta_B, \Theta_H) + \lambda_\beta E_\beta(\beta) + \lambda_{\theta B} E_{\theta B}(\theta_B) + \lambda_{\theta H} E_{\theta H}(\theta_H) + \lambda_R E_R(\tilde{M}, \beta) + \lambda_I E_I(\tilde{M}, \beta)

    Input: Observed markers M={mi,t}t=1…F,i=1…nM = \{m_{i,t}\}_{t=1\dots F, i=1\dots n}, initial marker attachments M~\tilde{M}
    Output: Subject shape β\beta, latent marker offsets M~\tilde{M}, frame poses ΘB,ΘH\Theta_B, \Theta_H
    b←46/nb \leftarrow 46 / n
    Initialize λD←600b/8\lambda_D \leftarrow 600 b / 8, λβ←1.25×8\lambda_\beta \leftarrow 1.25 \times 8, λθB←0.375×8\lambda_{\theta B} \leftarrow 0.375 \times 8, λθH←0.125×8\lambda_{\theta H} \leftarrow 0.125 \times 8, λI←37.5×8\lambda_I \leftarrow 37.5 \times 8, λR←104\lambda_R \leftarrow 10^4
    for k=1k = 1 to 44 do
        if k≥3k \ge 3 then
            Include 24-D MANO hand pose parameters into optimization
        else
            Set hand poses to rest configuration
        end if
        Optimize E(M~,β,ΘB,ΘH)E(\tilde{M}, \beta, \Theta_B, \Theta_H) using Powell's dogleg minimization
        λD←λD×2\lambda_D \leftarrow \lambda_D \times 2
        λβ←λβ/2\lambda_\beta \leftarrow \lambda_\beta / 2
        λθB←λθB/2\lambda_{\theta B} \leftarrow \lambda_{\theta B} / 2
        λθH←λθH/2\lambda_{\theta H} \leftarrow \lambda_{\theta H} / 2
        λI←λI/2\lambda_I \leftarrow \lambda_I / 2
    end for
    return β,M~,ΘB,ΘH\beta, \tilde{M}, \Theta_B, \Theta_H

    Definitions of energy terms and parameters:

    • EDE_D: squared Euclidean distance between simulated surface marker positions m(m~i,β,θt)m(\tilde{m}_i, \beta, \theta_t) and observed 3D marker coordinates mi,tm_{i,t}.
    • b=46/nb = 46/n: scaling factor normalizing marker density relative to a baseline 46-markerset, where nn is the count of active markers in the sequence.
    • Eβ(β)E_\beta(\beta): Mahalanobis distance regularizer over SMPL shape space.
    • EθB(θB)E_{\theta B}(\theta_B): prior on body pose parameters.
    • EθH(θH)=θ^HTΣθH−1θ^HE_{\theta H}(\theta_H) = \hat{\theta}_H^T \Sigma_{\theta H}^{-1} \hat{\theta}_H: Mahalanobis distance on the 24-dimensional MANO hand PCA pose subspace projection θ^H\hat{\theta}_H, with diagonal covariance matrix ΣθH\Sigma_{\theta H}.
    • ER(M~,β)E_R(\tilde{M}, \beta): penalty constraining latent markers to stay at prescribed distance d=9.5 mmd = 9.5\text{ mm} above the mesh surface.
    • EI(M~,β)E_I(\tilde{M}, \beta): regularizer penalizing marker deviations from their initial designated template surface positions.
  3. Knowl 3 — MoSh++ Stage II: Per-Frame Pose and Soft-Tissue Dynamics Optimization

    algorithm

    In Stage II of MoSh++, the body shape β\beta and latent marker locations M~\tilde{M} obtained from Stage I are held constant, and MoSh++ solves for body pose θB\theta_B, hand pose θH\theta_H, and dynamic soft-tissue coefficients ϕ∈R8\phi \in \mathbb{R}^8 on every frame tt of the sequence by optimizing:

    E(θB,θH,ϕ)=λDED(θB,θH,ϕ)+λθBEθB(θB)+λθHEθH(θH)+λuEu(θB,θH)+λϕEϕ(ϕ)+λvEv(ϕ)E(\theta_B, \theta_H, \phi) = \lambda_D E_D(\theta_B, \theta_H, \phi) + \lambda_{\theta B} E_{\theta B}(\theta_B) + \lambda_{\theta H} E_{\theta H}(\theta_H) + \lambda_u E_u(\theta_B, \theta_H) + \lambda_\phi E_\phi(\phi) + \lambda_v E_v(\phi)

    Input: Mocap marker sequence {Mt}t=1T\{M_t\}_{t=1}^T, fixed shape β\beta, fixed marker positions M~\tilde{M}
    Output: Frame poses {(θB)t,(θH)t}t=1T\{(\theta_B)_t, (\theta_H)_t\}_{t=1}^T, dynamic deformations {ϕt}t=1T\{\phi_t\}_{t=1}^T
    for frame t=1t = 1 to TT do
        xt←x_t \leftarrow number of missing markers in frame tt
        qt←1+2.5×(xt/∣M∣)q_t \leftarrow 1 + 2.5 \times (x_t / |M|)
        λD←400×(46/∣M∣)\lambda_D \leftarrow 400 \times (46 / |M|)
        λθB←1.6×qt\lambda_{\theta B} \leftarrow 1.6 \times q_t, λθH←1.0×qt\lambda_{\theta H} \leftarrow 1.0 \times q_t
        λu←2.5\lambda_u \leftarrow 2.5, λϕ←1.0\lambda_\phi \leftarrow 1.0, λv←6.0\lambda_v \leftarrow 6.0
        if t=1t = 1 then
            Initialize model by rigid alignment between simulated and observed markers
            Run graduated optimization of EE without dynamics, varying λθB∈[16qt,8qt,1.6qt]\lambda_{\theta B} \in [16 q_t, 8 q_t, 1.6 q_t]
        else
            Initialize (θB)t,(θH)t,ϕt(\theta_B)_t, (\theta_H)_t, \phi_t with solution from frame t−1t-1
            Step 1: Optimize EE over θB,θH\theta_B, \theta_H while fixing ϕ=0\phi = 0 (omitting Eϕ,EvE_\phi, E_v)
            Step 2: Optimize full EE over θB,θH,ϕ\theta_B, \theta_H, \phi together
        end if
    end for
    return {(θB)t,(θH)t,ϕt}t=1T\{(\theta_B)_t, (\theta_H)_t, \phi_t\}_{t=1}^T

    Key terms and priors:

    • Eu(θB,θH)E_u(\theta_B, \theta_H): temporal velocity smoothness penalty on pose changes between adjacent frames.
    • Eϕ(ϕ)=ϕtTΣϕ−1ϕtE_\phi(\phi) = \phi_t^T \Sigma_\phi^{-1} \phi_t: Mahalanobis distance prior on DMPL dynamics coefficients using covariance matrix Σϕ\Sigma_\phi derived from the DYNA dataset.
    • Ev(ϕ)E_v(\phi): temporal velocity smoothness penalty on dynamic coefficients ϕ\phi.
    • When no hand markers are present in the markerset, hand parameters are fixed to the neutral mean pose of the MANO model.
  4. Knowl 4 — AMASS Dataset Composition and Corpus Statistics

    data/table

    AMASS unifies 15 optical marker-based motion capture datasets into a single SMPL parameterization with consistent surface topology, skeletal kinematics, hand articulations, and soft-tissue dynamics. The source captures contain varying marker layouts ranging from 37 to 91 markers.

    Dataset Markers Subjects Motions Minutes
    ACCAD 82 20 258 27.22
    BioMotion 41 111 3130 541.82
    CMU 41 97 2030 559.18
    EKUT 46 4 349 30.74
    Eyes Japan 37 12 795 385.42
    HumanEva 39 3 28 8.48
    KIT 50 55 4233 662.04
    MPI HDM05 41 4 219 147.63
    MPI Limits 53 3 40 24.14
    MPI MoSh 87 20 78 16.65
    SFU 53 7 44 15.23
    SSM 86 3 30 1.87
    TCD Hand 91 1 62 8.05
    TotalCapture 53 5 40 43.71
    Transitions 53 1 115 15.84
    Total - 346 11451 2488.01

    Across the 15 component datasets, AMASS contains 346 distinct subjects, 11,451 individual motion sequences, and 2,488.01 minutes (over 41.4 hours) of motion capture data represented as 16-dimensional shape parameters β\beta, 8-dimensional soft-tissue parameters ϕ\phi, and 159-dimensional pose vectors θ\theta per frame.

  5. Knowl 5 — Synchronized Scans and Markers (SSM) Benchmark

    experimental setup

    The Synchronized Scans and Markers (SSM) dataset provides ground-truth dense 3D dynamic surface geometry captured synchronously with sparse optical mocap marker tracking for evaluation and hyperparameter optimization.

    Hardware and capture protocol:

    • Optical mocap system: 24 OptiTrack Prime 17W optical cameras tracking subjects fitted with 67 skin-mounted reflective markers (based on the MoSh optimized marker layout) or a standardized 46-marker subset.
    • 4D surface scanner: 3dMD body scanning system running synchronously at 60 frames per second, comprising 22 stereo camera pairs, 22 color cameras, 34 speckle projectors, and white-light LED illumination panels.
    • Subjects and data volume: 3 subjects exhibiting diverse body shapes performing a total of 30 distinct motion sequences.

    Evaluation metric: For each frame, 10,000 points are sampled uniformly at random from the 3D surface scan mesh, and the Euclidean distance from each sample point to the closest point on the reconstructed body model mesh is calculated. The metric reported is the mean scan-to-model distance in millimeters (mm\text{mm}).

  6. Knowl 6 — Reconstruction Accuracy of MoSh++ vs MoSh on 3D/4D Scans

    empirical result

    MoSh++ evaluated on the held-out test split of the SSM dataset achieves superior surface reconstruction accuracy compared to the original MoSh (BlendSCAPE) method across all evaluated tracking modes when tested using a standard 46-markerset:

    • Shape estimation error: MoSh achieves an average scan-to-mesh error of 12.1 mm12.1\text{ mm}, whereas MoSh++ achieves 7.4 mm7.4\text{ mm}.
    • Pose estimation error (without soft-tissue dynamics): MoSh yields 10.5 mm10.5\text{ mm} scan-to-mesh error, whereas MoSh++ yields 8.1 mm8.1\text{ mm}.
    • Pose estimation error with dynamic soft-tissue modeling: MoSh yields 10.24 mm10.24\text{ mm} error, whereas MoSh++ using DMPL soft-tissue dynamics achieves 7.3 mm7.3\text{ mm} error.

    For reference, direct 3D scan alignment produces a baseline scan-to-mesh distance of 0.5 mm0.5\text{ mm}. Additionally, MoSh++ tracking using 46 markers matches or outperforms MoSh tracking using the denser 67-marker setup (which yields ≈7.0 mm\approx 7.0\text{ mm} shape error and ≈8.7 mm\approx 8.7\text{ mm} pose error for MoSh).

  7. Knowl 7 — Missing-Marker Adaptive Pose Regularization

    equation

    To prevent unnatural joint poses when markers are occluded or missing during capture, MoSh++ adaptively scales the body and hand pose prior weights (λθB\lambda_{\theta B} and λθH\lambda_{\theta H}) via a variable occlusion factor qq:

    q=1+2.5×(x∣M∣)q = 1 + 2.5 \times \left(\frac{x}{|M|}\right)

    λθB=1.6×q,λθH=1.0×q\lambda_{\theta B} = 1.6 \times q, \quad \lambda_{\theta H} = 1.0 \times q

    where x∈[0,∣M∣]x \in [0, |M|] is the number of missing/unobserved markers in the current frame and ∣M∣|M| is the total count of session markers.

    Properties:

    • When all markers are visible (x=0x = 0), q=1.0q = 1.0, applying nominal regularization.
    • As marker occlusions increase, qq increases monotonically up to a maximum value of q=3.5q = 3.5 when all markers are lost (x=∣M∣x = |M|), which increases the pose prior penalty by a factor of 3.5 to hold the skeleton within valid kinematic distributions.
  8. Knowl 8 — Dimensionality Selection for Body Shape and Soft-Tissue Dynamics

    empirical result

    Cross-validation line search on the SSM validation set establishes that ∣β∣=16|\beta| = 16 shape components and ∣ϕ∣=8|\phi| = 8 DMPL dynamic soft-tissue components minimize scan-to-model reconstruction error when fitting to sparse marker sets.

    While increasing the number of shape components beyond 16 (e.g., to 32, 64, 128, or 256) decreases the residual distance between markers and model surface, it causes overfitting to sparse marker points, increasing test scan-to-mesh reconstruction error from 7.4 mm7.4\text{ mm} (∣β∣=16|\beta|=16) to >8.1 mm> 8.1\text{ mm} (∣β∣=128|\beta|=128). Similarly, dynamic components beyond ∣ϕ∣=8|\phi| = 8 lead to overfitting on sparse marker motions.

  9. Knowl 9 — Computational Performance and Runtime Limitations of MoSh++

    limitation

    MoSh++ is implemented using Powell's dogleg numerical optimization in the CPU-based Chumpy auto-differentiation framework and does not operate in real time. On a single quad-core Intel Core i7 CPU (3.1 GHz, 16 GB RAM), average per-sequence and per-frame runtimes on the SSM dataset are:

    • Stage I (Subject Shape and Marker Optimization): approximately 25 minutes per motion capture sequence (over F=12F=12 frames).
    • Stage II without soft-tissue dynamics: approximately 0.5 seconds per frame.
    • Stage II with dynamic soft-tissue optimization: approximately 2.0 seconds per frame.

Coverage note — None was omitted; all key technical contributions (model formulation, two-stage fitting algorithms, occlusion handling, SSM benchmark, hyperparameter validation, quantitative results, AMASS table, and runtime limitations) are fully covered.

References

  1. 1.4D Scan. http://www.3dmd.com/. 6, 12
  2. 2.SFU Motion Capture Database. http://mocap.cs.sfu.ca/. 8
  3. 3.MocapClub. http://www.mocapclub.com/, 2007. 2, 3
  4. 4.CMUKitchen. http://kitchen.cs.cmu.edu/, 2009. 2
  5. 5.ACCAD. https://accad.osu.edu/research/motion-lab/system-data, 2018. 1, 2, 8
  6. 6.Eyes Japan. http://mocapdata.com, 2018. 8
  7. 7.M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mane, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viegas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. arXiv:1603.04467 [cs], 2016. 7
  8. 8.I. Akhter and M. J. Black. Pose-Conditioned Joint Angle Limits for 3D Human Pose Reconstruction. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR 2015), 2015. 1, 2, 3, 8
  9. 9.S. Alexanderson, C. O'Sullivan, and J. Beskow. Robust online motion capture labeling of finger markers. In Proceedings of the 9th International Conference on Motion in Games. ACM, 2016. 3
  10. 10.B. Allen, B. Curless, and Z. Popovic. The space of human body shapes: Reconstruction and parameterization from range scans. ACM Transactions on Graphics (TOG), 2003. 3
  11. 11.B. Allen, B. Curless, Z. Popovic, and A. Hertzmann. Learning a Correlated Model of Identity and Pose-dependent Body Shape Variation for Real-time Synthesis. In Proceedings of the 2006 ACM SIGGRAPH/Eurographics Symposium on Computer Animation, SCA '06, 2006. 3
  12. 12.T. P. Andriacchi and E. J. Alexander. Studies of human locomotion: Past, present and future. Journal of Biomechanics, 2000. 3
  13. 13.D. Anguelov, P. Srinivasan, D. Koller, S. Thrun, J. Rodgers, and J. Davis. SCAPE: Shape Completion and Animation of PEople. ACM Transactions on Graphics, 2005. 2, 3
  14. 14.F. De la Torre, J. Hodgins, A. Bargteil, X. Martin, J. Macey, A. Collado, and P. Beltran. Guide to the carnegie mellon university multimodal activity (cmu-mmac) database. Robotics Institute, 2008. 1, 2, 3, 8
  15. 15.G. Dueck and T. Scheuer. Threshold accepting: A general purpose optimization algorithm appearing superior to simulated annealing. Journal of computational physics, 1990. 5
  16. 16.G. D. Forney. The viterbi algorithm. Proceedings of the IEEE, 1973. 3
  17. 17.G. E. Gorton, D. A. Hebert, and M. E. Gannotti. Assessment of the kinematic variability among 12 motion analysis laboratories. Gait & posture, 2009. 2, 3
  18. 18.S. Han, B. Liu, R. Wang, Y. Ye, C. D. Twigg, and K. Kin. Online Optical Marker-based Hand Tracking with Deep Labels. ACM Trans. Graph., 2018. 3
  19. 19.D. A. Hirshberg, M. Loper, E. Rachlin, and M. J. Black. Coregistration: Simultaneous alignment and modeling of articulated 3D shape. In A. Fitzgibbon, S. Lazebnik, P. Perona, Y. Sato, and C. Schmid, editors, European Conference on Computer Vision, Lecture Notes in Computer Science. Springer, Springer Berlin Heidelberg, 2012. 3
  20. 20.D. Holden. Robust Solving of Optical Motion Capture Data by Denoising. ACM Transactions on Graphics, 2018. 7
  21. 21.D. Holden, J. Saito, and T. Komura. A deep learning framework for character motion synthesis and editing. ACM Transactions on Graphics (TOG), 2016. 2, 3
  22. 22.L. Hoyet, K. Ryall, R. McDonnell, and C. O'Sullivan. Sleight of Hand: Perception of Finger Motion from Reduced Marker Sets. In Proceedings of the ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games, I3D '12. ACM, 2012. 1, 2, 8
  23. 23.M. Lafortune, P. Cavanagh, H. Sommer, and A. Kalenak. Three-dimensional kinematics of the human knee during walking. Journal of biomechanics, 1992. 3
  24. 24.A. Leardini, L. Chiari, U. D. Croce, and A. Cappozzo. Human movement analysis using stereophotogrammetry: Part 3. Soft tissue artifact assessment and compensation. Gait &/ Posture, 2005. 3
  25. 25.T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero. Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 2017. 7
  26. 26.M. Loper. Chumpy. 2013. 6
  27. 27.M. Loper, N. Mahmood, and M. J. Black. MoSh: Motion and shape capture from sparse markers. ACM Transactions on Graphics (TOG), 2014. 2, 3, 4, 7, 8, 12
  28. 28.M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black. SMPL: A Skinned Multi-Person Linear Model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 2015. 1, 2, 3, 4, 7
  29. 29.C. Mandery, O. Terlemez, M. Do, N. Vahrenkamp, and T. Asfour. The KIT Whole-Body Human Motion Database. In International Conference on Advanced Robotics (ICAR), 2015. 1, 2, 3, 8
  30. 30.J. Maycock, T. Rohlig, M. Schroder, M. Botsch, and H. Ritter. Fully automatic optical motion tracking using an inverse kinematics approach. In Humanoid Robots (Humanoids), 2015 IEEE-RAS 15th International Conference On. IEEE, 2015. 3
  31. 31.M. Muller, A. Baak, and H.-P. Seidel. Efficient and Robust Annotation of Motion Capture Data. In Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation (SCA), 2009. 1, 3
  32. 32.M. Muller, T. Roder, M. Clausen, B. Eberhardt, B. Kruger, and A. Weber. Documentation Mocap Database HDM05. Technical report, Universitat Bonn, 2007. 1, 2, 3, 8
  33. 33.I. NaturalPoint. Motion Capture Systems. 6, 12
  34. 34.J. Nocedal and S. J. Wright. Numerical Optimization. Springer, New York, 2nd edition, 2006. 6
  35. 35.G. Pons-Moll, J. Romero, N. Mahmood, and M. J. Black. Dyna: A Model of Dynamic Human Shape in Motion. ACM Transactions on Graphics, (Proc. SIGGRAPH), 2015. 3, 4, 5
  36. 36.J. Romero, D. Tzionas, and M. J. Black. Embodied Hands: Modeling and Capturing Hands and Bodies Together. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 2017. () Two first authors contributed equally. 2, 4, 5, 8
  37. 37.M. Schroder, J. Maycock, and M. Botsch. Reduced marker layouts for optical motion capture of hands. In Proceedings of the 8th ACM SIGGRAPH Conference on Motion in Games. ACM, 2015. 3
  38. 38.L. Sigal, A. Balan, and M. J. Black. HumanEva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. International Journal of Computer Vision, 2010. 2, 3, 8
  39. 39.J. Taylor, L. Bordeaux, T. Cashman, B. Corish, C. Keskin, T. Sharp, E. Soto, D. Sweeney, J. Valentin, B. Luff, et al. Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondences. ACM Transactions on Graphics (TOG), 2016. 3
  40. 40.A. Tkach, A. Tagliasacchi, E. Remelli, M. Pauly, and A. Fitzgibbon. Online generative model personalization for hand tracking. ACM Transactions on Graphics (TOG), 2017. 3
  41. 41.N. F. Troje. Decomposing biological motion: A framework for analysis and synthesis of human gait patterns. Journal of Vision, 2002. 1, 2, 3, 8
  42. 42.M. Trumble, A. Gilbert, C. Malleson, A. Hilton, and J. Collomosse. Total Capture: 3D Human Pose Estimation Fusing Video and Inertial Sensors. In BMVC17, 2017. 2, 3, 8
  43. 43.G. Varol, J. Romero, X. Martin, N. Mahmood, M. J. Black, I. Laptev, and C. Schmid. Learning from Synthetic Humans. In CVPR, 2017. 4
  44. 44.C. von Laßberg, W. Rapp, B. Mohler, and J. Krug. Neuromuscular onset succession of high level gymnasts during dynamic leg acceleration phases on high bar. Journal of Electromyography and Kinesiology, 2013. 3

Citation

MLA
Mahmood, N., et al. “AMASS: Archive of Motion Capture as Surface Shapes”. arXiv, 2019, http://arxiv.org/abs/1904.03278v1.
APA
Mahmood, N., Ghorbani, N., Troje, N. F., Pons-Moll, G., & Black, M. J. (2019). AMASS: Archive of Motion Capture as Surface Shapes. arXiv. http://arxiv.org/abs/1904.03278v1
Chicago
Mahmood, N., N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black. 2019. “AMASS: Archive of Motion Capture as Surface Shapes”. arXiv. http://arxiv.org/abs/1904.03278v1.
Harvard
Mahmood, N. et al. (2019) “AMASS: Archive of Motion Capture as Surface Shapes”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1904.03278v1.
Vancouver
1. Mahmood N, Ghorbani N, Troje NF, Pons-Moll G, Black MJ (2019) AMASS: Archive of Motion Capture as Surface Shapes. arXiv

BibTeX

@article{mahmood2019amass,
  title = {AMASS: Archive of Motion Capture as Surface Shapes},
  author = {Mahmood, Naureen and Ghorbani, Nima and Troje, Nikolaus F. and Pons-Moll, Gerard and Black, Michael J.},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1904.03278v1},
  eprint = {1904.03278}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE