Expressive Body Capture: 3D Hands, Face, and Body From a Single Image

Georgios PavlakosVasileios ChoutasNima GhorbaniTimo BolkartAhmed A. A. OsmanDimitrios TzionasMichael J. Black

article2019CVPR2,587 citationsBest Paper Finalist

Introduces SMPL-X and SMPLify-X, providing a unified 3D parametric body model with articulated hands and expressive facial geometry alongside an optimization framework to reconstruct complete 3D human shape and pose from a single monocular image.

Listen

The article addresses the challenge of capturing detailed 3D representations of human bodies, hands, and faces from everyday single images, which is essential for understanding actions, interactions, and emotions but remains difficult because existing models lack sufficient expressivity and paired training data is scarce.

The work set out to create a unified 3D body model called SMPL-X that jointly represents the full body, articulated hands, and expressive face, along with an optimization-based fitting method called SMPLify-X that recovers model parameters from one RGB image.

Researchers built SMPL-X by combining and refining existing body, hand, and head models, then training its shape and pose spaces on thousands of curated 3D scans. SMPLify-X detects 2D joints with OpenPose, fits the model using an improved variational pose prior learned from motion-capture data, a precise collision penalty, automatic gender classification, and a faster PyTorch implementation. Accuracy was measured on a new curated set of 100 images with pseudo ground-truth meshes.

The richer SMPL-X model produced lower vertex-to-vertex errors of 52.9 mm compared with 54–65 mm for simpler body-only or body-plus-hands variants. Replacing the prior or removing the collision term raised errors, while the gender-specific model outperformed the neutral version. Qualitative fits on in-the-wild images showed natural hand gestures and facial expressions, and the method proved more robust to noisy detections than hand-only baselines.

These results matter because they demonstrate that a single expressive model can deliver holistic 3D capture from ordinary photos, supporting applications in scene understanding, animation, and human–computer interaction without specialized multi-camera setups.

The authors recommend curating larger in-the-wild datasets of SMPL-X fits to train direct regression networks. They note that performance still depends on the quality of 2D joint detections and can fail under heavy occlusion or depth ambiguity, so further gains will require visibility-aware terms and additional data.

No sufficiently relevant recommendations were found.

Cover for Expressive Body Capture: 3D Hands, Face, and Body From a Single Image

Abstract

To facilitate the analysis of human actions, interactions and emotions, we compute a 3D model of human body pose, hand pose, and facial expression from a single monocular image. To achieve this, we use thousands of 3D scans to train a new, unified, 3D model of the human body, SMPL-X, that extends SMPL with fully articulated hands and an expressive face. Learning to regress the parameters of SMPL-X directly from images is challenging without paired images and 3D ground truth. Consequently, we follow the approach of SMPLify, which estimates 2D features and then optimizes model parameters to fit the features. We improve on SMPLify in several significant ways: (1) we detect 2D features corresponding to the face, hands, and feet and fit the full SMPL-X model to these; (2) we train a new neural network pose prior using a large MoCap dataset; (3) we define a new interpenetration penalty that is both fast and accurate; (4) we automatically detect gender and the appropriate body models (male, female, or neutral); (5) our PyTorch implementation achieves a speedup of more than 8x over Chumpy. We use the new method, SMPLify-X, to fit SMPL-X to both controlled images and images in the wild. We evaluate 3D accuracy on a new curated dataset comprising 100 images with pseudo ground-truth. This is a step towards automatic expressive human capture from monocular RGB data. The models, code, and data are available for research purposes at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 2.1 Modeling the body
  • 2.2 Inferring the body
  • 3 Technical approach
  • 3.1 Unified model: SMPL-X
  • 3.2 SMPLify-X: SMPL-X from a single image
  • 3.3 Variational Human Body Pose Prior
  • 3.4 Collision penalizer
  • 3.5 Deep Gender Classifier
  • 3.6 Optimization
  • 4 Experiments
  • 4.1 Evaluation datasets
  • 4.2 Qualitative & Quantitative evaluations
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — SMPL-X Expressive Parametric Human Body Model

    model/method

    SMPL-X (SMPL eXpressive) is a unified statistical 3D model of the human body, face, and hands. The mesh has N=10,475N = 10,475 vertices and K=54K = 54 joints, incorporating joints for the neck, jaw, eyeballs, and fingers. The model function M(β,θ,ψ):R∣β∣×∣θ∣×∣ψ∣→R3NM(\beta, \theta, \psi) : \mathbb{R}^{|\beta| \times |\theta| \times |\psi|} \to \mathbb{R}^{3N} maps shape parameters β\beta, pose parameters θ\theta, and facial expression parameters ψ\psi to vertex positions via vertex-based linear blend skinning:

    M(β,θ,ψ)=W(TP(β,θ,ψ),J(β),θ,W)M(\beta, \theta, \psi) = W(T_P(\beta, \theta, \psi), J(\beta), \theta, \mathcal{W})

    TP(β,θ,ψ)=Tˉ+BS(β;S)+BE(ψ;E)+BP(θ;P)T_P(\beta, \theta, \psi) = \bar{T} + B_S(\beta; \mathcal{S}) + B_E(\psi; \mathcal{E}) + B_P(\theta; \mathcal{P})

    where:

    • Tˉ∈R3N\bar{T} \in \mathbb{R}^{3N} is the rest-pose template mesh.
    • BS(β;S)=∑n=1∣β∣βnSnB_S(\beta; \mathcal{S}) = \sum_{n=1}^{|\beta|} \beta_n S_n is the shape blend shape function with identity PCA basis vectors Sn∈R3NS_n \in \mathbb{R}^{3N} and coefficients β∈R∣β∣\beta \in \mathbb{R}^{|\beta|} (∣β∣=10|\beta| = 10), trained on 3,800 A-pose 3D scans from CAESAR.
    • BE(ψ;E)=∑n=1∣ψ∣ψnEnB_E(\psi; \mathcal{E}) = \sum_{n=1}^{|\psi|} \psi_n E_n is the facial expression blend shape function with PCA basis vectors En∈R3NE_n \in \mathbb{R}^{3N} and coefficients ψ∈R∣ψ∣\psi \in \mathbb{R}^{|\psi|} (∣ψ∣=10|\psi| = 10), adopted from the FLAME head model.
    • BP(θ;P)=∑n=19K(Rn(θ)−Rn(θ∗))PnB_P(\theta; \mathcal{P}) = \sum_{n=1}^{9K} (R_n(\theta) - R_n(\theta^*)) P_n is the pose blend shape function with corrective displacements Pn∈R3NP_n \in \mathbb{R}^{3N}, where R(θ)∈R9KR(\theta) \in \mathbb{R}^{9K} maps pose θ\theta to joint relative rotation matrices and θ∗\theta^* denotes the rest pose.
    • J(β)=J(Tˉ+BS(β;S))J(\beta) = \mathcal{J}(\bar{T} + B_S(\beta; \mathcal{S})) predicts 3D joint locations from the shape-deformed rest template via a learned sparse linear regressor J\mathcal{J}.
    • W(T,J,θ,W)W(T, J, \theta, \mathcal{W}) represents standard linear blend skinning parameterized by blend weights W∈RN×K\mathcal{W} \in \mathbb{R}^{N \times K}.

    The pose vector θ∈R3(K+1)\theta \in \mathbb{R}^{3(K+1)} decomposes into body joints θb\theta_b, jaw joint θf\theta_f, and finger joints θh\theta_h. Hand pose is parameterized by a lower-dimensional PCA space from MANO: θh=∑n=1∣mh∣mhnMn\theta_h = \sum_{n=1}^{|m_h|} m_{hn} M_n, where mh∈R24m_h \in \mathbb{R}^{24} (12 coefficients per hand). The complete model has 119 parameters: 75 for global body rotation and body/jaw/eye joints, 24 for hand PCA coefficients, 10 for body shape, and 10 for facial expressions.

  2. Knowl 2 — SMPLify-X Single-Image Fitting Objective

    model/method

    Fitting SMPL-X to 2D keypoints detected in a single RGB image (SMPLify-X) is formulated as continuous unconstrained energy minimization over body shape β\beta, pose parameters θ=(θb(Z),θf,mh)\theta = (\theta_b(Z), \theta_f, m_h), and facial expressions ψ\psi:

    E(β,θ,ψ)=EJ(β,θ,K,Jest)+λθbEθb(θb)+λθfEθf(θf)+λmhEmh(mh)+λαEα(θb)+λβEβ(β)+λEEE(ψ)+λCEC(θ,β)E(\beta, \theta, \psi) = E_J(\beta, \theta, K, J_{\text{est}}) + \lambda_{\theta_b} E_{\theta_b}(\theta_b) + \lambda_{\theta_f} E_{\theta_f}(\theta_f) + \lambda_{m_h} E_{m_h}(m_h) + \lambda_\alpha E_\alpha(\theta_b) + \lambda_\beta E_\beta(\beta) + \lambda_E E_E(\psi) + \lambda_C E_C(\theta, \beta)

    where:

    • EJ(β,θ,K,Jest)=∑joint iγiωiρ(ΠK(Rθ(J(β))i)−Jest,i)E_J(\beta, \theta, K, J_{\text{est}}) = \sum_{\text{joint } i} \gamma_i \omega_i \rho(\Pi_K(R_\theta(J(\beta))_i) - J_{\text{est}, i}) is the 2D reprojection loss, where ΠK\Pi_K denotes 3D-to-2D perspective projection with camera intrinsics KK, Rθ(J(β))iR_\theta(J(\beta))_i is the posed 3D joint position, Jest,iJ_{\text{est}, i} is the detected 2D keypoint, ωi∈[0,1]\omega_i \in [0, 1] is the detection confidence score from OpenPose, γi\gamma_i is an annealing stage weight, and ρ(e)=e2e2+σ2\rho(e) = \frac{e^2}{e^2 + \sigma^2} is the robust Geman-McClure error function.
    • Eβ(β)=∥β∥22E_\beta(\beta) = \|\beta\|_2^2 is the Mahalanobis distance regularizing body shape under unit-variance scaling.
    • Eθf(θf)=∥θf∥22E_{\theta_f}(\theta_f) = \|\theta_f\|_2^2, Emh(mh)=∥mh∥22E_{m_h}(m_h) = \|m_h\|_2^2, and EE(ψ)=∥ψ∥22E_E(\psi) = \|\psi\|_2^2 are L2L_2 priors penalizing deviation from the neutral state for facial pose, hand pose PCA coefficients, and facial expressions.
    • Eα(θb)=∑i∈{elbows,knees}exp⁡(θi)E_\alpha(\theta_b) = \sum_{i \in \{\text{elbows}, \text{knees}\}} \exp(\theta_i) penalizes extreme joint bending angles at the elbows and knees.
    • Eθb(θb)=∥Z∥22E_{\theta_b}(\theta_b) = \|Z\|_2^2 is the VPoser latent body pose prior on Z∈R32Z \in \mathbb{R}^{32}.
    • EC(θ,β)E_C(\theta, \beta) is the 3D mesh self-collision penalty.
    • λ∙\lambda_{\bullet} are term-specific scalar weighting hyperparameters.
  3. Knowl 3 — VPoser Variational Human Body Pose Prior

    model/method

    VPoser is a variational autoencoder (VAE) trained on human motion capture data to model a continuous latent space of physically plausible 3D body poses Z∈R32Z \in \mathbb{R}^{32}. The input to the network is a concatenation of 3×33 \times 3 rotation matrices R∈SO(3)23R \in SO(3)^{23} representing the relative rotations of 23 body joints (R∈[−1,1]207R \in [-1, 1]^{207}), and the output R^\hat{R} reconstructs these rotation matrices.

    The training loss is defined as:

    Ltotal=c1LKL+c2Lrec+c3Lorth+c4Ldet1+c5Lreg\mathcal{L}_{\text{total}} = c_1 \mathcal{L}_{KL} + c_2 \mathcal{L}_{\text{rec}} + c_3 \mathcal{L}_{\text{orth}} + c_4 \mathcal{L}_{\text{det1}} + c_5 \mathcal{L}_{\text{reg}}

    LKL=DKL(q(Z∣R)∥N(0,I))\mathcal{L}_{KL} = D_{KL}(q(Z|R) \parallel \mathcal{N}(0, I))

    Lrec=∥R−R^∥22\mathcal{L}_{\text{rec}} = \|R - \hat{R}\|_{2}^2

    Lorth=∥R^R^T−I∥22\mathcal{L}_{\text{orth}} = \|\hat{R}\hat{R}^T - I\|_{2}^2

    Ldet1=∣det⁡(R^)−1∣\mathcal{L}_{\text{det1}} = |\det(\hat{R}) - 1|

    Lreg=∥ϕ∥22\mathcal{L}_{\text{reg}} = \|\phi\|_2^2

    where ϕ\phi denotes the network weights, and the hyperparameter coefficients are c1=0.005c_1 = 0.005, c2=1.0−c1=0.995c_2 = 1.0 - c_1 = 0.995, c3=1.0c_3 = 1.0, c4=1.0c_4 = 1.0, and c5=0.0005c_5 = 0.0005.

    Both encoder and decoder are symmetric, fully connected networks using LeakyReLU activations (negative slope 0.20.2) and dropout (0.250.25):

    • Encoder: 207→512→512→32207 \to 512 \to 512 \to 32 (yielding posterior parameters μ(R)\mu(R) and Σ(R)\Sigma(R)).
    • Decoder: 32→512→512→20732 \to 512 \to 512 \to 207 with a tanh⁡\tanh output activation.

    During SMPLify-X fitting, optimization directly searches over the latent vector Z∈R32Z \in \mathbb{R}^{32} with a quadratic penalty ∥Z∥22\|Z\|_2^2. The decoder maps ZZ to R^\hat{R}, which is converted to axis-angle joint angles θb(Z)\theta_b(Z) via the inverse Rodrigues formula.

  4. Knowl 4 — Mesh Interpenetration Penalization via Conic Distance Fields

    equation

    Self-collisions between mesh triangles are penalized using local conic 3D distance fields computed over colliding triangle pairs detected via Bounding Volume Hierarchies (BVH). For a set of colliding triangle pairs C\mathcal{C}, the collision energy EC(θ)E_C(\theta) is:

    EC(θ)=∑(fs(θ),ft(θ))∈C(∑vs∈fs∥−Ψft(vs)ns∥22+∑vt∈ft∥−Ψfs(vt)nt∥22)E_C(\theta) = \sum_{(f_s(\theta), f_t(\theta)) \in \mathcal{C}} \left( \sum_{v_s \in f_s} \| -\Psi_{f_t}(v_s) \mathbf{n}_s \|_2^2 + \sum_{v_t \in f_t} \| -\Psi_{f_s}(v_t) \mathbf{n}_t \|_2^2 \right)

    where vs,vt∈R3v_s, v_t \in \mathbb{R}^3 are vertex coordinates, and ns,nt∈R3\mathbf{n}_s, \mathbf{n}_t \in \mathbb{R}^3 are triangle unit normals. The local conic distance field Ψfs(vt):R3→R+\Psi_{f_s}(v_t) : \mathbb{R}^3 \to \mathbb{R}^+ for intruder vertex vtv_t and receiver face fsf_s with circumcenter ofs∈R3o_{f_s} \in \mathbb{R}^3 and circumradius rfs>0r_{f_s} > 0 is:

    Ψfs(vt)={∣(1−Φ(vt))Υ(nfs⋅(vt−ofs))∣2if Φ(vt)<10if Φ(vt)≥1\Psi_{f_s}(v_t) = \begin{cases} |(1 - \Phi(v_t)) \Upsilon(\mathbf{n}_{f_s} \cdot (v_t - o_{f_s}))|^2 & \text{if } \Phi(v_t) < 1 \\ 0 & \text{if } \Phi(v_t) \ge 1 \end{cases}

    Φ(vt)=∥(vt−ofs)−(nfs⋅(vt−ofs))nfs∥2−rfs−σrfs(nfs⋅(vt−ofs))+rfs\Phi(v_t) = \frac{\|(v_t - o_{f_s}) - (\mathbf{n}_{f_s} \cdot (v_t - o_{f_s})) \mathbf{n}_{f_s}\|_2 - r_{f_s}}{\frac{-\sigma}{r_{f_s}}(\mathbf{n}_{f_s} \cdot (v_t - o_{f_s})) + r_{f_s}}

    Υ(x)={−x+1−σx≤−σ−1−2σ4σ2x2−12σx+14(3−2σ)x∈(−σ,+σ)0x≥+σ\Upsilon(x) = \begin{cases} -x + 1 - \sigma & x \le -\sigma \\ \frac{-1-2\sigma}{4\sigma^2} x^2 - \frac{1}{2\sigma} x + \frac{1}{4}(3 - 2\sigma) & x \in (-\sigma, +\sigma) \\ 0 & x \ge +\sigma \end{cases}

    where coordinates are expressed in meters and the cone field-of-view parameter is σ=0.0001\sigma = 0.0001. Collisions are explicitly ignored for body regions in permanent or frequent self-contact (eyes, toes, armpits, crotch, and immediately adjacent links along the kinematic chain).

  5. Knowl 5 — SMPLify-X Multi-Stage Annealed Fitting Algorithm

    algorithm

    SMPLify-X estimates camera parameters and SMPL-X model parameters (β,Z,θf,mh,ψ)(\beta, Z, \theta_f, m_h, \psi) from 2D OpenPose detections using L-BFGS optimization with strong Wolfe line search across three annealed stages.

    Input: 2D joint coordinates and confidences (J_est, omega), camera intrinsics K, initial body model M
    Output: Optimized SMPL-X parameters (beta, Z, theta_f, m_h, psi)
    Estimate camera translation and global body orientation while keeping body shape beta and pose fixed
    Fix camera parameters
    Initialize weights: lambda_alpha, lambda_beta, lambda_E to high regularization; lambda_C to low
    Stage 1 (Coarse Body Fitting):
        Set joint weights: gamma_b = 1.0, gamma_h = 0.0, gamma_f = 0.0
        Optimize beta and latent body pose Z via L-BFGS (learning rate = 1.0, max_iter = 30)
    Stage 2 (Arm and Hand Pose Refinement):
        Set joint weights: gamma_b = 1.0, gamma_h = 0.1, gamma_f = 0.0
        Reduce lambda_alpha, lambda_beta, lambda_E; increase lambda_C
        Optimize beta, Z, and hand PCA pose m_h via L-BFGS
    Stage 3 (Full Expressive Capture):
        Set joint weights: gamma_b = 1.0, gamma_h = 2.0, gamma_f = 2.0
        Further reduce regularization lambda_alpha, lambda_beta, lambda_E; maximize lambda_C
        Optimize beta, Z, m_h, jaw pose theta_f, and facial expressions psi via L-BFGS
    Decode final body pose: theta_b = VPoserDecoder(Z)
    return (beta, theta_b, theta_f, m_h, psi)

    The optimization is implemented in PyTorch using custom CUDA kernels for the BVH collision operator, achieving an 8×8\times speedup over the Chumpy implementation in SMPLify.

  6. Knowl 6 — Deep Gender Classification for Body Model Selection

    model/method

    To select the gender-appropriate SMPL-X model (male, female, or gender-neutral), a ResNet18 convolutional neural network is trained for binary gender classification on cropped human images containing detected OpenPose keypoints. Training data was curated from LSP, LSP-extended, MPII, MS-COCO, and LIP datasets by filtering crops with size ≥200×200\ge 200 \times 200 pixels and at least one high-confidence visible joint in the head, torso, and each limb, followed by consensus Amazon Mechanical Turk annotations (50,216 training and 16,170 test samples).

    At inference time:

    • If the predicted class probability for male or female is ≥0.90\ge 0.90, the corresponding gender-specific SMPL-X body model is fitted.
    • If the predicted probability is <0.90< 0.90, the gender-neutral SMPL-X body model is fitted.

    On a class-equalized validation set, this 0.900.90 threshold yields 62.38%62.38\% confident correct predictions, 7.54%7.54\% incorrect predictions, and discards uncertain cases to the neutral model.

  7. Knowl 7 — Expressive Hands and Faces (EHF) Dataset

    experimental setup

    The Expressive Hands and Faces (EHF) dataset is a curated benchmark for evaluating 3D capture of the human body, hands, and face together from a single RGB image. It consists of 100 color frames selected from multi-view 4D scan sequences in the SMPL+H dataset. For each frame, pseudo ground-truth 3D surface meshes are obtained by fitting SMPL-X directly to the high-resolution 4D scan point clouds and manually curating alignments for high accuracy across complex hand gestures and facial expressions.

    Accuracy is quantified using Procrustes-aligned vertex-to-vertex (v2v) mean absolute error (in mm) between the reconstructed mesh and the pseudo ground-truth mesh, in contrast to standard 3D joint error which is blind to bone rotations and facial/hand surface deformations.

  8. Knowl 8 — Comparative Reconstruction Accuracy Across Model Expressivity on EHF

    data/table

    Quantitative comparison of body models on the 100-image EHF evaluation benchmark, fitted using SMPLify-X from a single RGB image and 2D OpenPose keypoints. Evaluations report mean vertex-to-vertex (v2v) surface error and standard 3D body-only joint error (in mm) after Procrustes alignment with ground-truth meshes and joints:

    Model Keypoints v2v error (mm) Joint error (mm)
    “SMPL” Body 57.6 63.5
    “SMPL” Body+Hands+Face 64.5 71.7
    “SMPL+H” Body+Hands 54.2 63.9
    SMPL-X Body+Hands+Face 52.9 62.6

    Standard 3D body joint error fails to reflect differences in hand and face expressivity (SMPL body joint error is 63.5 mm vs. SMPL-X 62.6 mm). The v2v surface error captures expressivity gains, showing progressive error reduction from SMPL (57.6 mm) to SMPL+H (54.2 mm) and SMPL-X (52.9 mm). Attempting to fit SMPL to hand and face keypoints without underlying expressive blendshapes causes severe distortion, increasing v2v error to 64.5 mm.

  9. Knowl 9 — Ablation Study of SMPLify-X Components on EHF

    data/table

    An ablation study isolating the contribution of individual components of SMPLify-X evaluated on the EHF dataset by reporting mean Procrustes-aligned vertex-to-vertex (v2v) error in mm:

    Version v2v error (mm)
    SMPLify-X (Full, gender-specific) 52.9
    Gender neutral model 58.0
    Replace VPoser with GMM 56.4
    No collision term 53.5

    Using the gender-neutral model instead of the gender-specific model increases error by 5.1 mm (58.0 mm vs. 52.9 mm). Replacing the neural VPoser pose prior with the Gaussian Mixture Model (GMM) prior from SMPLify increases error by 3.5 mm (56.4 mm vs. 52.9 mm). Removing the conic distance field collision penalizer increases surface error to 53.5 mm while allowing physically impossible mesh self-interpenetrations.

  10. Knowl 10 — 3D Joint Reconstruction Accuracy on Human3.6M and THF Datasets

    empirical result

    When evaluated on the standard Human3.6M benchmark using 3D joint error after Procrustes alignment following the SMPLify evaluation protocol, SMPLify-X outperforms SMPLify, achieving a mean 3D joint error of 75.9 mm and median error of 60.8 mm, compared to SMPLify's mean error of 82.3 mm and median error of 69.3 mm.

    On the Total Hands and Faces (THF) dataset (200 curated frames from CMU Panoptic Studio with multi-view triangulated 3D keypoints), SMPLify-X achieves 19.8 mm mean 3D joint error on hand joints, outperforming the dedicated RGB hand-pose estimation baseline of Panteleris et al. which achieves 26.5 mm, demonstrating that full-body context regularizes hand fitting.

  11. Knowl 11 — Failure Modes of SMPLify-X Fitting

    limitation

    SMPLify-X suffers from several characteristic failure modes:

    1. Depth ambiguities inherent in 2D reprojection loss can result in incorrect torso rotation or inaccurate ordinal depth ordering of limbs and feet.
    2. Keypoint occlusions leave occluded limbs unconstrained by the data term EJE_J, allowing them to drift into unnatural poses.
    3. Complex self-contact poses (e.g., tightly crossed arms or hands resting on knees) can cause optimization to get trapped in local minima due to the collision penalty, as SMPL-X does not model soft-tissue contact deformation.

Coverage note — None was omitted; all key contributions—the SMPL-X parametric formulation, VPoser prior, collision penalty, optimization schedule, gender classifier, EHF benchmark, ablation studies, and comparative evaluations—are fully covered.

References

  1. 1.Ijaz Akhter and Michael J. Black. Pose-conditioned joint angle limits for 3D human pose reconstruction. In CVPR, 2015. 5
  2. 2.Brett Allen, Brian Curless, and Zoran Popović. The space of human body shapes: Reconstruction and parameterization from range scans. ACM Transactions on Graphics, (Proc. SIGGRAPH), 22(3):587–594, 2003. 2, 3
  3. 3.Brett Allen, Brian Curless, Zoran Popović, and Aaron Hertzmann. Learning a correlated model of identity and pose-dependent body shape variation for real-time synthesis. In ACM SIGGRAPH/Eurographics Symposium on Computer Animation, SCA ’06, pages 147–156. Eurographics Association, 2006. 2, 3
  4. 4.Brian Amberg, Reinhard Knothe, and Thomas Vetter. Expression invariant 3D face recognition with a morphable model. In International Conference on Automatic Face Gesture Recognition, 2008. 2
  5. 5.Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele. 2D human pose estimation: New benchmark and state of the art analysis. In CVPR, 2014. 6, 8
  6. 6.Dragomir Anguelov, Praveen Srinivasan, Daphne Koller, Sebastian Thrun, Jim Rodgers, and James Davis. SCAPE: Shape Completion and Animation of PEople. ACM Transactions on Graphics, (Proc. SIGGRAPH), 24(3):408–416, 2005. 2, 3
  7. 7.Luca Ballan and Guido Maria Cortelazzo. Marker-less motion capture of skinned models in a four camera set-up using optical flow and silhouettes. In International Symposium on 3D Data Processing, Visualization and Transmission (3DPVT), 2008. 3
  8. 8.Luca Ballan, Aparna Taneja, Juergen Gall, Luc Van Gool, and Marc Pollefeys. Motion capture of hands in action using discriminative salient points. In ECCV, 2012. 3, 5, 6
  9. 9.Volker Blanz and Thomas Vetter. A morphable model for the synthesis of 3D faces. In SIGGRAPH, pages 187–194, 1999. 2, 3
  10. 10.Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In ECCV, 2016. 1, 3, 4, 5, 6, 7
  11. 11.James Booth, Anastasios Roussos, Allan Ponniah, David Dunaway, and Stefanos Zafeiriou. Large scale 3D morphable models. IJCV, 126(2-4):233–254, 2018. 2
  12. 12.Christoph Bregler, Jitendra Malik, and Katherine Pullen. Twist based acquisition and tracking of animal and human kinematics. International Journal of Computer Vision (IJCV), 56(3):179–194, 2004. 4
  13. 13.Alan Brunton, Augusto Salazar, Timo Bolkart, and Stefanie Wuhrer. Review of statistical shape spaces for 3D data with comparative analysis for human faces. CVIU, 128(0):1–17, 2014. 2, 3
  14. 14.Chen Cao, Yanlin Weng, Shun Zhou, Yiying Tong, and Kun Zhou. Facewarehouse: A 3D facial expression database for visual computing. IEEE Transactions on Visualization and Computer Graphics, 20(3):413–425, 2014. 2, 3
  15. 15.Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. Realtime multi-person 2D pose estimation using part affinity fields. In CVPR, 2017. 1, 2, 5
  16. 16.Yinpeng Chen, Zicheng Liu, and Zhengyou Zhang. Tensor-based human body modeling. In CVPR, 2013. 3
  17. 17.CMU. CMU MoCap dataset. 5
  18. 18.Total Capture Dataset. http://domedb.perception.cs.cmu.edu. 7
  19. 19.Martin De La Gorce, David J. Fleet, and Nikos Paragios. Model-based 3D hand pose estimation from monocular video. PAMI, 33(9):1793–1805, 2011. 3
  20. 20.Quentin Delamarre and Olivier D. Faugeras. 3D articulated models and multiview tracking with physical forces. CVIU, 81(3):328–357, 2001. 3
  21. 21.P. Ekman and W. Friesen. Facial Action Coding System: A Technique for the Measurement of Facial Movement. Consulting Psychologists Press, 1978. 3
  22. 22.models FLAME website: dataset and code. http://flame.is.tue.mpg.de. 3
  23. 23.Oren Freifeld and Michael J. Black. Lie bodies: A manifold representation of 3D human shape. In ECCV, 2012. 3
  24. 24.Juergen Gall, Carsten Stoll, Edilson De Aguiar, Christian Theobalt, Bodo Rosenhahn, and Hans-Peter Seidel. Motion capture using joint skeleton tracking and surface estimation. In CVPR, 2009. 3
  25. 25.Stuart Geman and Donald E. McClure. Statistical methods for tomographic image reconstruction. In Proceedings of the 46th Session of the International Statistical Institute, Bulletin of the ISI, volume 52, 1987. 5
  26. 26.Nils Hasler, Carsten Stoll, Martin Sunkel, Bodo Rosenhahn, and Hans-Peter Seidel. A statistical model of human pose and body shape. Computer Graphics Forum, 28(2):337–346, 2009. 2, 3
  27. 27.Nils Hasler, Thorsten Thormählen, Bodo Rosenhahn, and Hans-Peter Seidel. Learning skeletons for shape and pose. In Proceedings of the 2010 ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games, I3D ’10, pages 23–30, New York, NY, USA, 2010. ACM. 3
  28. 28.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 6
  29. 29.David A. Hirshberg, Matthew Loper, Eric Rachlin, and Michael J. Black. Coregistration: Simultaneous alignment and modeling of articulated 3D shape. In ECCV, 2012. 3
  30. 30.Yinghao Huang, Federica Bogo, Christoph Lassner, Angjoo Kanazawa, Peter V. Gehler, Javier Romero, Ijaz Akhter, and Michael J. Black. Towards accurate marker-less human shape and pose estimation over time. In 3DV, 2017. 3
  31. 31.Eldar Insafutdinov, Leonid Pishchulin, Bjoern Andres, Mykhaylo Andriluka, and Bernt Schiele. Deepercut: A deeper, stronger, and faster multi-person pose estimation model. In ECCV, 2016. 1
  32. 32.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments. PAMI, 36(7):1325–1339, 2014. 5
  33. 33.Sam Johnson and Mark Everingham. Clustered pose and nonlinear appearance models for human pose estimation. In BMVC, 2010. 2, 6, 8
  34. 34.Sam Johnson and Mark Everingham. Learning effective human pose estimation from inaccurate annotation. In CVPR, 2011. 6, 8
  35. 35.Hanbyul Joo, Hao Liu, Lei Tan, Lin Gui, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh. Panoptic studio: A massively multiview system for social motion capture. In ICCV, 2015. 3, 4
  36. 36.Hanbyul Joo, Tomas Simon, and Yaser Sheikh. Total capture: A 3D deformation model for tracking faces, hands, and bodies. In CVPR, 2018. 2, 3, 4, 6, 7, 8
  37. 37.Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In CVPR, 2018. 1, 2, 3
  38. 38.Tero Karras. Maximizing parallelism in the construction of BVHs, Octrees, and K-d trees. In Proceedings of the Fourth ACM SIGGRAPH / Eurographics Conference on High-Performance Graphics, pages 33–37, 2012. 6
  39. 39.Sameh Khamis, Jonathan Taylor, Jamie Shotton, Cem Keskin, Shahram Izadi, and Andrew Fitzgibbon. Learning an efficient model of hand shape variation from depth images. In CVPR, 2015. 2, 3
  40. 40.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. In ICLR, 2014. 5
  41. 41.Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J Black, and Peter V Gehler. Unite the people: Closing the loop between 3D and 2D human representations. In CVPR, 2017. 3
  42. 42.John P. Lewis, Matt Cordner, and Nickson Fong. Pose space deformation: A unified approach to shape interpolation and skeleton-driven deformation. In ACM Transactions on Graphics (SIGGRAPH), pages 165–172, 2000. 4
  43. 43.Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero. Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphics (TOG), 36(6):194, 2017. 2, 3, 4
  44. 44.Xiaodan Liang, Chunyan Xu, Xiaohui Shen, Jianchao Yang, Si Liu, Jinhui Tang, Liang Lin, and Shuicheng Yan. Human parsing with contextualized convolutional neural network. In ICCV, 2015. 6
  45. 45.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014. 6
  46. 46.Yebin Liu, Juergen Gall, Carsten Stoll, Qionghai Dai, Hans-Peter Seidel, and Christian Theobalt. Markerless motion capture of multiple characters using multiview image segmentation. PAMI, 35(11):2720–2735, 2013. 3
  47. 47.Matthew Loper, Naureen Mahmood, and Michael J Black. MoSh: Motion and shape capture from sparse markers. ACM Transactions on Graphics (TOG), 33(6):220, 2014. 2, 4, 5
  48. 48.Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, Oct. 2015. 2, 3, 6
  49. 49.Matthew M Loper and Michael J Black. OpenDR: An approximate differentiable renderer. In ECCV, 2014. 6
  50. 50.Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: Archive of motion capture as surface shapes. arXiv:1904.03278, 2019. 2, 5
  51. 51.MANO, models SMPL+H website: dataset, and code. http://mano.is.tue.mpg.de. 3, 6
  52. 52.Stan Melax, Leonid Keselman, and Sterling Orsten. Dynamics based 3D skeletal hand tracking. In Graphics Interface, pages 63–70, 2013. 2, 3
  53. 53.Thomas B. Moeslund, Adrian Hilton, and Volker Krüger. A survey of advances in vision-based human motion capture and analysis. CVIU, 104(2):90–126, 2006. 3
  54. 54.Richard M. Murray, Li Zexiang, and S. Shankar Sastry. A Mathematical Introduction to Robotic Manipulation. CRC press, 1994. 4
  55. 55.Jorge Nocedal and Stephen J Wright. Nonlinear Equations. Springer, 2006. 6
  56. 56.Markus Oberweger, Paul Wohlhart, and Vincent Lepetit. Training a feedback loop for hand pose estimation. In ICCV, 2015. 2
  57. 57.Iason Oikonomidis, Nikolaos Kyriazis, and Antonis A. Argyros. Efficient model-based 3D tracking of hand articulations using Kinect. In BMVC, 2011. 2, 3
  58. 58.Mohamed Omran, Christoph Lassner, Gerard Pons-Moll, Peter V Gehler, and Bernt Schiele. Neural body fitting: Unifying deep learning and model-based human pose and shape estimation. In 3DV, 2018. 1, 2, 3
  59. 59.OpenPose. https://github.com/CMU-Perceptual-Computing-Lab/openpose. 2
  60. 60.Paschalis Panteleris, Iason Oikonomidis, and Antonis Argyros. Using a single RGB frame for real time 3D hand pose estimation in the wild. In WACV, 2018. 8
  61. 61.Georgios Pavlakos, Luyang Zhu, Xiaowei Zhou, and Kostas Daniilidis. Learning to estimate 3D human pose and shape from a single color image. In CVPR, 2018. 1, 2, 3, 6
  62. 62.Pascal Paysan, Reinhard Knothe, Brian Amberg, Sami Romdhani, and Thomas Vetter. A 3D face model for pose and illumination invariant face recognition. In 2009 Sixth IEEE International Conference on Advanced Video and Signal Based Surveillance, pages 296–301, 2009. 2
  63. 63.Gerard Pons-Moll, Javier Romero, Naureen Mahmood, and Michael J. Black. Dyna: A model of dynamic human shape in motion. ACM Transactions on Graphics, (Proc. SIGGRAPH), 34(4):120:1–120:14, July 2015. 3
  64. 64.Gerard Pons-Moll and Bodo Rosenhahn. Model-Based Pose Estimation, chapter 9, pages 139–170. Springer, 2011. 4
  65. 65.Helge Rhodin, Jörg Spörri, Isinsu Katircioglu, Victor Constantin, Frédéric Meyer, Erich Müller, Mathieu Salzmann, and Pascal Fua. Learning monocular 3D human pose estimation from multi-view images. In CVPR, 2018. 3
  66. 66.Kathleen M. Robinette, Sherri Blackwell, Hein Daanen, Mark Boehmer, Scott Fleming, Tina Brill, David Hoeferlin, and Dennis Burnsides. Civilian American and European Surface Anthropometry Resource (CAESAR) final report. Technical Report AFRL-HE-WP-TR-2002-0169, US Air Force Research Laboratory, 2002. 3, 4
  67. 67.Javier Romero, Dimitrios Tzionas, and Michael J Black. Embodied hands: Modeling and capturing hands and bodies together. ACM Transactions on Graphics (TOG), 2017. 2, 3, 4, 5, 6
  68. 68.Tanner Schmidt, Richard Newcombe, and Dieter Fox. DART: Dense articulated real-time tracking. In RSS, 2014. 2, 3
  69. 69.Tomas Simon, Hanbyul Joo, Iain Matthews, and Yaser Sheikh. Hand keypoint detection in single images using multiview bootstrapping. In CVPR, 2017. 1, 2, 5
  70. 70.Srinath Sridhar, Antti Oulasvirta, and Christian Theobalt. Interactive markerless articulated hand motion tracking using RGB and depth data. In ICCV, 2013. 2, 3
  71. 71.Jonathan Starck and Adrian Hilton. Surface capture for performance-based animation. IEEE computer graphics and applications, 27(3), 2007. 3
  72. 72.Matthias Teschner, Stefan Kimmerle, Bruno Heidelberger, Gabriel Zachmann, Laks Raghupathi, Arnulph Fuhrmann, Marie-Paule Cani, François Faure, Nadia Magnenat-Thalmann, Wolfgang Strasser, and Pascal Volino. Collision detection for deformable objects. In Eurographics, 2004. 5
  73. 73.Anastasia Tkach, Mark Pauly, and Andrea Tagliasacchi. Sphere-meshes for real-time hand modeling and tracking. ACM Transactions on Graphics (TOG), 35(6), 2016. 2, 3
  74. 74.Dimitrios Tzionas, Luca Ballan, Abhilash Srikantha, Pablo Aponte, Marc Pollefeys, and Juergen Gall. Capturing hands in action using discriminative salient points and physics simulation. IJCV, 118(2):172–193, 2016. 2, 3, 5, 6
  75. 75.Daniel Vlasic, Matthew Brand, Hanspeter Pfister, and Jovan Popović. Face transfer with multilinear models. ACM transactions on graphics (TOG), 24(3):426–433, 2005. 2
  76. 76.Shih-En Wei, Varun Ramakrishna, Takeo Kanade, and Yaser Sheikh. Convolutional pose machines. In CVPR, 2016. 2, 5
  77. 77.Weipeng Xu, Avishek Chatterjee, Michael Zollhöfer, Helge Rhodin, Dushyant Mehta, Hans-Peter Seidel, and Christian Theobalt. Monoperfcap: Human performance capture from monocular video. ACM Transactions on Graphics (TOG), 37(2):27, 2018. 3
  78. 78.Fei Yang, Jue Wang, Eli Shechtman, Lubomir Bourdev, and Dimitri Metaxas. Expression flow for 3D-aware face component transfer. ACM Transactions on Graphics (TOG), 30(4):60, 2011. 2
  79. 79.Shanxin Yuan, Guillermo Garcia-Hernando, Björn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov, Jan Kautz, Sina Honari, Liuhao Ge, et al. Depth-based 3D hand pose estimation: From current achievements to future goals. In CVPR, 2018. 3
  80. 80.Michael Zollhöfer, Justus Thies, Pablo Garrido, Derek Bradley, Thabo Beeler, Patrick Pérez, Marc Stamminger, Matthias Nießner, and Christian Theobalt. State of the art on monocular 3D face reconstruction, tracking, and applications. Computer Graphics Forum, 37(2):523–550, 2018. 3
  81. 81.Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. Tensorflow: A system for large-scale machine learning. In OSDI, volume 16, pages 265–283, 2016. 6, 7
  82. 82.Ijaz Akhter and Michael J. Black. Pose-conditioned joint angle limits for 3D human pose reconstruction. In CVPR, 2015. 6
  83. 83.Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele. 2D human pose estimation: New benchmark and state of the art analysis. In CVPR, 2014. 4, 7, 9
  84. 84.Luca Ballan, Aparna Taneja, Juergen Gall, Luc Van Gool, and Marc Pollefeys. Motion capture of hands in action using discriminative salient points. In ECCV, 2012. 1, 2
  85. 85.Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In ECCV, 2016. 4, 10
  86. 86.François Chollet et al. Keras. https://keras.io, 2015. 7
  87. 87.CMU. CMU MoCap dataset. 6
  88. 88.Total Capture Dataset. http://domedb.perception.cs.cmu.edu. 3, 5
  89. 89.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 7
  90. 90.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments. PAMI, 36(7):1325–1339, 2014. 4, 6
  91. 91.Sam Johnson and Mark Everingham. Clustered pose and nonlinear appearance models for human pose estimation. In BMVC, 2010. 7
  92. 92.Sam Johnson and Mark Everingham. Learning effective human pose estimation from inaccurate annotation. In CVPR, 2011. 7
  93. 93.Hanbyul Joo, Tomas Simon, and Yaser Sheikh. Total capture: A 3D deformation model for tracking faces, hands, and bodies. In CVPR, 2018. 3
  94. 94.Tero Karras. Maximizing parallelism in the construction of BVHs, Octrees, and K-d trees. In Proceedings of the Fourth ACM SIGGRAPH / Eurographics Conference on High-Performance Graphics, pages 33–37, 2012. 3
  95. 95.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 7
  96. 96.Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero. Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphics (TOG), 36(6):194, 2017. 1
  97. 97.Xiaodan Liang, Chunyan Xu, Xiaohui Shen, Jianchao Yang, Si Liu, Jinhui Tang, Liang Lin, and Shuicheng Yan. Human parsing with contextualized convolutional neural network. In ICCV, 2015. 7
  98. 98.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014. 7
  99. 99.Matthew Loper, Naureen Mahmood, and Michael J Black. MoSh: Motion and shape capture from sparse markers. ACM Transactions on Graphics (TOG), 33(6):220, 2014. 6
  100. 100.Andrew L Maas, Awni Y Hannun, and Andrew Y Ng. Rectifier nonlinearities improve neural network acoustic models. In ICML Workshops, 2013. 6
  101. 101.Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: Archive of motion capture as surface shapes. arXiv:1904.03278, 2019. 6
  102. 102.Jorge Nocedal and Stephen Wright. Numerical Optimization. Springer, New York, 2nd edition, 2006. 3
  103. 103.OpenPose. https://github.com/CMU-Perceptual-Computing-Lab/openpose. 2, 4, 5
  104. 104.Paschalis Panteleris, Iason Oikonomidis, and Antonis Argyros. Using a single RGB frame for real time 3D hand pose estimation in the wild. In WACV, 2018. 1, 2
  105. 105.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pyTorch. In NIPS-W, 2017. 6
  106. 106.Kathleen M. Robinette, Sherri Blackwell, Hein Daanen, Mark Boehmer, Scott Fleming, Tina Brill, David Hoeferlin, and Dennis Burnsides. Civilian American and European Surface Anthropometry Resource (CAESAR) final report. Technical Report AFRL-HE-WP-TR-2002-0169, US Air Force Research Laboratory, 2002. 4
  107. 107.Matthias Teschner, Stefan Kimmerle, Bruno Heidelberger, Gabriel Zachmann, Laks Raghupathi, Arnulph Fuhrmann, Marie-Paule Cani, François Faure, Nadia Magnenat-Thalmann, Wolfgang Strasser, and Pascal Volino. Collision detection for deformable objects. In Eurographics, 2004. 1
  108. 108.Dimitrios Tzionas, Luca Ballan, Abhilash Srikantha, Pablo Aponte, Marc Pollefeys, and Juergen Gall. Capturing hands in action using discriminative salient points and physics simulation. IJCV, 118(2):172–193, 2016. 1, 2

Citation

MLA
Pavlakos, G., et al. “Expressive Body Capture: 3D Hands, Face, and Body From a Single Image”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 10967–77, https://doi.org/10.1109/CVPR.2019.01123.
APA
Pavlakos, G., Choutas, V., Ghorbani, N., Bolkart, T., Osman, A. A., Tzionas, D., & Black, M. J. (2019). Expressive Body Capture: 3D Hands, Face, and Body From a Single Image. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10967–10977. https://doi.org/10.1109/CVPR.2019.01123
Chicago
Pavlakos, G., V. Choutas, N. Ghorbani, et al. 2019. “Expressive Body Capture: 3D Hands, Face, and Body From a Single Image”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10967–77. https://doi.org/10.1109/CVPR.2019.01123.
Harvard
Pavlakos, G. et al. (2019) “Expressive Body Capture: 3D Hands, Face, and Body From a Single Image”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 10967–10977. Available at: https://doi.org/10.1109/CVPR.2019.01123.
Vancouver
1. Pavlakos G, Choutas V, Ghorbani N, Bolkart T, Osman AA, Tzionas D, Black MJ (2019) Expressive Body Capture: 3D Hands, Face, and Body From a Single Image. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 10967–10977

BibTeX

@inproceedings{Pavlakos_2019, title={Expressive Body Capture: 3D Hands, Face, and Body From a Single Image}, url={http://dx.doi.org/10.1109/CVPR.2019.01123}, DOI={10.1109/cvpr.2019.01123}, booktitle={2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Pavlakos, Georgios and Choutas, Vasileios and Ghorbani, Nima and Bolkart, Timo and Osman, Ahmed A. and Tzionas, Dimitrios and Black, Michael J.}, year={2019}, month=June, pages={10967–10977} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE