Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image

Federica BogoAngjoo KanazawaChristoph LassnerPeter GehlerJavier RomeroMichael J. Black

article2016ECCV1,839 citations

Introduces SMPLify, the first automated method to reconstruct full 3D human body pose and shape from a single unconstrained image by fitting the SMPL statistical model directly to detected 2D joints.

Listen

Estimating the three-dimensional pose and shape of the human body from a single, unconstrained two-dimensional photograph is a fundamental challenge in computer vision. Previous approaches have generally reconstructed simplified 3D stick figures or relied heavily on manual user input, known silhouettes, or multi-camera setups. These simplified approaches frequently produce anatomically impossible body configurations and fail to capture true 3D surface geometry, limiting their value for downstream applications such as animation, virtual avatars, and automated scene analysis.

The article introduces and evaluates "SMPLify," the first fully automated method to reconstruct both 3D human pose and full 3D body shape directly from a single 2D image. The researchers set out to demonstrate that standard 2D joint locations contain sufficient information to infer plausible 3D body surface meshes without requiring manual intervention.

The approach operates in two main stages: a bottom-up detection step followed by top-down statistical model fitting. First, a convolutional neural network predicts 2D body joint locations and associated confidence values from the image. Second, the system fits a statistical 3D body model (SMPL) to these detected joints by minimizing an objective function. This objective balances joint alignment errors against learned body shape statistics, unnatural joint bending limits, and an efficient collision penalty based on body capsules to prevent self-intersection. The authors evaluated the framework across synthetic benchmarks and three real-world datasets: HumanEva-I, Human3.6M, and the Leeds Sports Pose dataset.

The findings confirm that SMPLify outperforms existing state-of-the-art methods in 3D joint accuracy while uniquely reconstructing a complete 3D surface mesh. On the HumanEva-I benchmark, SMPLify achieved a mean 3D joint error of 79.9 mm, representing an approximate 27% error reduction compared to the previous leading method (110.0 mm). On the Human3.6M benchmark, SMPLify attained a mean joint error of 82.3 mm compared to 106.7 mm for the next best method, reflecting a roughly 23% performance improvement. Synthetic tests demonstrated that 2D joint positions provide robust signals for 3D body shape estimation, outperforming baseline average-shape predictions even under pixel noise. Furthermore, ablation experiments confirmed that multi-modal pose priors and self-intersection penalties successfully rule out physically impossible poses in complex imagery.

These results establish that incorporating rich 3D statistical body models significantly reduces the geometric ambiguity inherent in single-image 3D reconstruction. SMPLify delivers ready-to-animate, rigged 3D meshes using standard desktop hardware in under one minute per image, lowering operational overhead and avoiding the high costs of manual segmentation or specialized capture environments.

For future development, the authors recommend extending the framework to multi-camera and video inputs, integrating facial pose and automatic gender classification, incorporating silhouette cues, and scaling the system to multi-person scenes where 3D meshes can help resolve occlusions. Limitations include reliance on estimated camera focal length, susceptibility to 2D detection errors when subjects overlap, and depth ambiguity under severe occlusions. Nonetheless, the framework provides highly reliable and robust baseline reconstructions across diverse real-world conditions.

Cover for Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image

Abstract

We describe the first method to automatically estimate the 3D pose of the human body as well as its 3D shape from a single unconstrained image. We estimate a full 3D mesh and show that 2D joints alone carry a surprising amount of information about body shape. The problem is challenging because of the complexity of the human body, articulation, occlusion, clothing, lighting, and the inherent ambiguity in inferring 3D from 2D. To solve this, we first use a recently published CNN-based method, DeepCut, to predict (bottom-up) the 2D body joint locations. We then fit (top-down) a recently published statistical body shape model, called SMPL, to the 2D joints. We do so by minimizing an objective function that penalizes the error between the projected 3D model joints and detected 2D joints. Because SMPL captures correlations in human shape across the population, we are able to robustly fit it to very little data. We further leverage the 3D model to prevent solutions that cause interpenetration. We evaluate our method, SMPLify, on the Leeds Sports, HumanEva, and Human3.6M datasets, showing superior pose accuracy with respect to the state of the art.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Approximating Bodies with Capsules
  • 3.2 Objective Function
  • 3.3 Optimization
  • 4 Evaluation
  • 4.1 Quantitative Evaluation: Synthetic Data
  • 4.2 Quantitative Evaluation: Real Data
  • 4.3 Qualitative Evaluation
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — SMPLify Objective Function for 3D Pose and Shape Fitting

    model/method

    SMPLify fits a 3D parametric human body model (SMPL) to 2D joints detected in a single image by minimizing a composite objective function over body shape coefficients β∈R10\beta \in \mathbb{R}^{10} and pose parameters θ∈R72\theta \in \mathbb{R}^{72} (consisting of global orientation and 23 relative joint rotations in axis-angle representation):

    E(β,θ)=EJ(β,θ;K,Jest)+λθEθ(θ)+λaEa(θ)+λspEsp(θ;β)+λβEβ(β)E(\beta, \theta) = E_J(\beta, \theta; K, J_{est}) + \lambda_\theta E_\theta(\theta) + \lambda_a E_a(\theta) + \lambda_{sp} E_{sp}(\theta; \beta) + \lambda_\beta E_\beta(\beta)

    where:

    • EJ(β,θ;K,Jest)E_J(\beta, \theta; K, J_{est}) is the joint-based data term penalizing 2D reprojection error against CNN-detected 2D joint positions Jest,i∈R2J_{est,i} \in \mathbb{R}^2 weighted by detection confidences wi∈[0,1]w_i \in [0, 1] under camera intrinsic parameters KK: EJ(β,θ;K,Jest)=∑joint iwiρ(ΠK(Rθ(J(β)i))−Jest,i)E_J(\beta, \theta; K, J_{est}) = \sum_{\text{joint } i} w_i \rho\left(\Pi_K(R_\theta(J(\beta)_i)) - J_{est,i}\right) where ΠK\Pi_K denotes perspective projection, J(β)iJ(\beta)_i is the unposed 3D position of joint ii regressed from body shape β\beta, RθR_\theta applies kinematic transformations, and ρ(x)=x2x2+σρ2\rho(x) = \frac{x^2}{x^2 + \sigma_\rho^2} is the robust Geman-McClure penalty function to handle noisy detections and occlusions.
    • Eθ(θ)E_\theta(\theta) is a multi-modal Gaussian mixture pose prior learned from motion capture data.
    • Ea(θ)=∑iexp⁡(θi)E_a(\theta) = \sum_{i} \exp(\theta_i) is an asymmetric limit penalty summing over knee and elbow bending angles θi\theta_i to heavily penalize unnatural hyperextension.
    • Esp(θ;β)E_{sp}(\theta; \beta) is a differentiable interpenetration penalty enforcing body non-self-intersection using proxy capsule geometries.
    • Eβ(β)=βTΣβ−1βE_\beta(\beta) = \beta^T \Sigma_\beta^{-1} \beta is a Mahalanobis shape prior where Σβ−1\Sigma_\beta^{-1} is a diagonal matrix of squared singular values obtained from Principal Component Analysis (PCA) on training body scans.
    • λθ,λa,λsp,λβ\lambda_\theta, \lambda_a, \lambda_{sp}, \lambda_\beta are positive scalar regularization weights.
  2. Knowl 2 — Differentiable Interpenetration Penalty via Linear Capsule Parameter Regression

    model/method

    To prevent physically impossible self-intersecting body configurations in an efficient and differentiable manner, the body surface is approximated by 20 volumetric capsules (one per body part, excluding fingers and toes):

    1. Capsule Parameter Regression: For any shape β∈R10\beta \in \mathbb{R}^{10}, capsule radii r(β)r(\beta) and axis lengths are predicted directly from shape coefficients using a linear ridge regressor trained on unposed body scans. Capsules are transformed into 3D space by kinematic rotations RθR_\theta.

    2. Discretization into 3D Gaussians: Each capsule is discretized into a chain of spheres centered at positions C(θ,β)C(\theta, \beta) along its axis with radius r(β)r(\beta). Each sphere is represented as an isotropic 3D Gaussian with standard deviation σ(β)=r(β)3\sigma(\beta) = \frac{r(\beta)}{3}.

    3. Collision Energy: The interpenetration penalty Esp(θ;β)E_{sp}(\theta; \beta) computes the scaled integral of Gaussian overlaps over pairs of incompatible parts I(i)I(i) (body parts that do not intersect in natural poses):

    Esp(θ;β)=∑i∑j∈I(i)exp⁡(−∥Ci(θ,β)−Cj(θ,β)∥2σi2(β)+σj2(β))E_{sp}(\theta; \beta) = \sum_i \sum_{j \in I(i)} \exp\left( - \frac{\|C_i(\theta, \beta) - C_j(\theta, \beta)\|^2}{\sigma_i^2(\beta) + \sigma_j^2(\beta)} \right)

    This term is differentiable with respect to both pose θ\theta and shape β\beta. To prevent the optimizer from unnaturally shrinking the body to avoid collision, EspE_{sp} is included when optimizing pose θ\theta but excluded when updating shape coefficients β\beta.

  3. Knowl 3 — Multi-Modal Pose Prior with Max-Mixture Approximation

    model/method

    To constrain human pose optimization to plausible articulated configurations, a Gaussian Mixture Model (GMM) with N=8N = 8 components is trained on approximately 1 million poses from 100 subjects in the CMU motion capture dataset, parameterized as relative part rotations in axis-angle format:

    p(θ)=∑j=1NgjN(θ;μθ,j,Σθ,j)p(\theta) = \sum_{j=1}^N g_j \mathcal{N}(\theta; \mu_{\theta,j}, \Sigma_{\theta,j})

    where gjg_j are mixture weights, μθ,j\mu_{\theta,j} are mean poses, and Σθ,j\Sigma_{\theta,j} are covariance matrices.

    Because computing the gradient of the negative log-likelihood of a sum is computationally cumbersome and numerically difficult in non-linear solvers, the mixture is approximated using a max operator:

    Eθ(θ)≡−log⁡∑j=1N(gjN(θ;μθ,j,Σθ,j))≈−log⁡(max⁡j(cgjN(θ;μθ,j,Σθ,j)))=min⁡j(−log⁡(cgjN(θ;μθ,j,Σθ,j)))E_\theta(\theta) \equiv -\log \sum_{j=1}^N \left( g_j \mathcal{N}(\theta; \mu_{\theta,j}, \Sigma_{\theta,j}) \right) \approx -\log \left( \max_j \left( c g_j \mathcal{N}(\theta; \mu_{\theta,j}, \Sigma_{\theta,j}) \right) \right) = \min_j \left( -\log \left( c g_j \mathcal{N}(\theta; \mu_{\theta,j}, \Sigma_{\theta,j}) \right) \right)

    where c>0c > 0 is a constant scaling factor. At each optimization step, the Jacobian of Eθ(θ)E_\theta(\theta) is approximated by the Jacobian of the specific Gaussian mode j∗j^* that minimizes the energy for the current pose θ\theta.

  4. Knowl 4 — SMPLify Fitting Algorithm and Camera Initialization

    algorithm

    The complete SMPLify pipeline reconstructs 3D pose θ\theta, shape β\beta, and camera translation γ\gamma from detected 2D joint positions JestJ_{est} and confidence scores ww using a perspective camera with known or estimated focal length.

    Input: 2D joint detections Jest∈RP×2J_{est} \in \mathbb{R}^{P \times 2}, joint confidences w∈RPw \in \mathbb{R}^P, camera focal length ff
    Output: Optimized 3D pose θ∗\theta^*, 3D shape β∗\beta^*, camera translation γ∗\gamma^*
    // Stage 1: Camera Depth Initialization
    Set shape β←0\beta \leftarrow 0 (mean SMPL shape)
    Estimate depth γz\gamma_z via ratio of similar triangles between mean SMPL 3D torso length and detected 2D torso joints
    Initialize translation γ=[γx,γy,γz]T\gamma = [\gamma_x, \gamma_y, \gamma_z]^T assuming body plane is parallel to image plane
    // Stage 2: Torso Alignment
    Optimize γ\gamma and global body orientation θroot\theta_{\text{root}} by minimizing EJE_J over torso joints only, keeping β=0\beta = 0
    // Stage 3: Viewpoint Ambiguity Check
    if 2D distance between detected shoulder joints < threshold then
        Set hypothesis 1 with orientation θroot\theta_{\text{root}}
        Set hypothesis 2 with orientation θroot+180∘\theta_{\text{root}} + 180^\circ
        Optimize both hypotheses and select the one yielding lower EJE_J
    end if
    // Stage 4: Staged Model Optimization
    Initialize weights λθ,λβ\lambda_\theta, \lambda_\beta to high values
    for each optimization stage do
        Update (θ,β)(\theta, \beta) by minimizing E(β,θ)E(\beta, \theta) using Powell's dogleg solver
        Exclude ∇βEsp\nabla_\beta E_{sp} during shape updates
        Decrease regularization weights λθ\lambda_\theta and λβ\lambda_\beta for subsequent stages
    end for
    return θ∗,β∗,γ∗\theta^*, \beta^*, \gamma^*

    The optimization is implemented using Chumpy and OpenDR and converges in under 1 minute per image on a standard desktop machine.

  5. Knowl 5 — SMPL Human Model Parameterization and Gender-Neutral Formulation

    model/method

    SMPLify uses the SMPL body model M(β,θ,γ)M(\beta, \theta, \gamma), which maps shape, pose, and translation parameters to a triangulated 3D mesh containing 6890 vertices.

    • Shape Parameters (β\beta): Linear coefficients β∈R10\beta \in \mathbb{R}^{10} modulate the first 10 principal shape components learned from thousands of aligned 3D human body scans.
    • Pose Parameters (θ\theta): Pose is defined by θ∈R72\theta \in \mathbb{R}^{72}, consisting of relative axis-angle 3D rotations for 23 skeleton joints plus 1 global orientation parameter.
    • Joint Function J(β)J(\beta): 3D joint locations are predicted directly from mesh vertices via a sparse linear combination matrix, directly coupling bone lengths and joint centers with shape β\beta.
    • Gender-Neutral Model: Standard SMPL provides distinct female and male models. To enable fully automatic processing without known gender, a gender-neutral model is trained on approximately 2000 male and 2000 female body shapes with female pose-dependent deformations, maintaining pose accuracy regardless of subject gender.
  6. Knowl 6 — Quantitative Evaluation of 3D Pose on HumanEva-I and Human3.6M

    data/table

    3D pose estimation performance evaluated on single frames from HumanEva-I (validation set: Walking and Boxing sequences for subjects S1, S2, S3) and Human3.6M (subjects S9 and S11 across 15 action categories, trial 1, camera 3). 2D joint inputs are predicted by the DeepCut CNN. Predicted 3D joints are aligned to ground truth via Procrustes analysis over 14 common joints, reporting mean 3D Euclidean joint error in millimeters (mm).

    HumanEva-I 3D joint error (mm) across subjects S1, S2, S3 for Walking and Boxing sequences.

    HumanEva-I 3D joint error (mm) across subjects S1, S2, S3 for Walking and Boxing sequences.
    Method Walking Boxing Mean Median
    S1 S2 S3 S1 S2 S3
    Akhter and Black 186.1 197.8 209.4 165.5 196.5 208.4 194.4 171.2
    Ramakrishna et al. 161.8 182.0 188.6 151.0 170.4 158.3 168.4 145.9
    Zhou et al. 100.0 98.89 123.1 112.5 118.6 110.0 110.0 98.9
    SMPLify 73.3 59.0 99.4 82.1 79.2 87.2 79.9 61.9

    Human3.6M 3D joint error (mm) across 15 actions for subjects S9 and S11.

    Human3.6M 3D joint error (mm) across 15 actions for subjects S9 and S11.
    Method Directions Discussion Eating Greeting Phoning Photo Posing Purchases
    Akhter and Black 199.2 177.6 161.8 197.8 176.2 186.5 195.4 167.3
    Ramakrishna et al. 137.4 149.3 141.6 154.3 157.7 158.9 141.8 158.1
    Zhou et al. 99.7 95.8 87.9 116.8 108.3 107.3 93.5 95.3
    SMPLify 62.0 60.2 67.8 76.5 92.1 77.0 73.0 75.3
    Method Sit SitDown Smoking Waiting WalkDog Walk WalkTogether Mean (Median)
    Akhter and Black 160.7 173.7 177.8 181.9 176.2 198.6 192.7 181.1 (158.1)
    Ramakrishna et al. 168.6 175.6 160.4 161.7 150.0 174.8 150.2 157.3 (136.8)
    Zhou et al. 109.1 137.5 106.0 102.2 106.5 110.4 115.2 106.7 (90.0)
    SMPLify 100.3 137.3 83.4 77.3 79.7 86.8 81.7 82.3 (69.3)

    SMPLify consistently outperforms state-of-the-art 2D-to-3D skeleton lifting methods on both datasets, demonstrating the constraining power of a statistical 3D body shape and pose model.

  7. Knowl 7 — Ablation of Pose Priors and Interpenetration on HumanEva-I

    data/table

    An ablation study on HumanEva-I evaluates the quantitative impact of the mixture-of-Gaussians pose prior EθE_\theta versus a uni-modal Gaussian pose prior Eθ′E_{\theta'}, and the effect of the interpenetration penalty EspE_{sp}. Errors represent 3D Euclidean joint error in millimeters (mm) following Procrustes alignment.

    HumanEva-I ablation study comparing energy configurations (errors in mm).

    HumanEva-I ablation study comparing energy configurations (errors in mm).
    Objective Configuration Walking Boxing Mean Median
    S1 S2 S3 S1 S2 S3
    Eβ+EJ+Eθ′E_\beta + E_J + E_{\theta'} (Uni-modal Gaussian) 98.4 79.6 117.8 105.9 98.5 122.5 104.1 82.3
    Eβ+EJ+Eθ′+EspE_\beta + E_J + E_{\theta'} + E_{sp} 97.9 79.4 116.0 105.8 98.5 122.3 103.7 82.3
    SMPLify (Eβ+EJ+Eθ+Ea+EspE_\beta + E_J + E_\theta + E_a + E_{sp}) 73.3 59.0 99.4 82.1 79.2 87.2 79.9 61.9

    Replacing the 8-component Gaussian mixture pose prior EθE_\theta with a single Gaussian prior Eθ′E_{\theta'} causes a substantial increase in mean joint error (from 79.9 mm to 104.1 mm). The interpenetration penalty EspE_{sp} shows negligible numerical impact on simple laboratory poses (104.1 mm vs 103.7 mm), but prevents unnatural self-intersection artifacts in complex articulated poses.

  8. Knowl 8 — Recoverability of 3D Human Shape from 2D Joint Positions

    empirical result

    Quantitative experiments on 2000 synthetic images (1000 male, 1000 female at 640×480640 \times 480 resolution) demonstrate that 2D joint coordinates provide sufficient geometric signal to estimate 3D body shape:

    1. Robustness to 2D Joint Noise: When zero-mean Gaussian noise with standard deviation from 1 to 5 pixels is added to 2D joint coordinates, fitting SMPL shape coefficients β\beta achieves lower mean vertex-to-vertex Euclidean error (in a canonical pose) than predicting the average dataset shape across all noise levels.
    2. Dependence on Joint Count: When body pose is known, 3D shape reconstruction error decreases monotonically as more 2D joints are incorporated into the optimization: fitting to all 23 SMPL joints yields significantly higher shape accuracy than fitting to 12 joints (torso and limbs), which in turn surpasses fitting to only 4 torso joints.
  9. Knowl 9 — Failure Modes of SMPLify on Unconstrained Images

    limitation

    When applied to unconstrained real-world images such as those in the Leeds Sports Pose (LSP) dataset, SMPLify is subject to three primary failure modes:

    • 2D Joint Estimation Errors: False positives or mislocalizations by the bottom-up CNN detector (e.g., swapped limb labels, or assigning joints from another person in crowded scenes) propagate directly into erroneous 3D mesh configurations.
    • Monocular Depth and Viewing Ambiguities: Extreme foreshortening and projective 2D-to-3D ambiguities can cause the non-convex optimization to converge to incorrect local minima (e.g., flipped limb orientations along the optical axis).
    • Under-Constrained Shape Inference: Because 2D skeleton joints provide only sparse geometric constraints without silhouette boundaries, subtle surface variations (such as weight differences among individuals with identical bone lengths) cannot be precisely recovered.

Coverage note — None. All major contributions—including the full SMPLify optimization framework, capsule-based differentiable interpenetration penalty, max-mixture pose prior, staged optimization algorithm, gender-neutral SMPL model, synthetic shape recovery experiments, quantitative benchmarks on HumanEva-I and Human3.6M, ablation analyses, and qualitative failure modes—are covered.

References

  1. 1.http://smplify.is.tue.mpg.de
  2. 2.http://chumpy.org
  3. 3.http://mocap.cs.cmu.edu
  4. 4.Akhter, I., Black, M.J.: Pose-conditioned joint angle limits for 3D human pose reconstruction. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 1446–1455 (2015)
  5. 5.Anguelov, D., Srinivasan, P., Koller, D., Thrun, S., Rodgers, J., Davis, J.: SCAPE: shape completion and animation of people. ACM Trans. Graph. (TOG) - Proc. ACM SIGGRAPH 24(3), 408–416 (2005)
  6. 6.Balan, A.O., Sigal, L., Black, M.J., Davis, J.E., Haussecker, H.W.: Detailed human shape and pose from images. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 1–8 (2007)
  7. 7.Barron, C., Kakadiaris, I.: Estimating anthropometry and pose from a single uncalibrated image. Comput. Vis. Image Underst. CVIU 81(3), 269–284 (2001)
  8. 8.Bo, L., Sminchisescu, C.: Twin Gaussian processes for structured prediction. Int. J. Comput. Vis. IJCV 87(1–2), 28–52 (2010)
  9. 9.Chen, Y., Kim, T.-K., Cipolla, R.: Inferring 3D shapes and deformations from single views. In: Daniilidis, K., Maragos, P., Paragios, N. (eds.) ECCV 2010. LNCS, vol. 6313, pp. 300–313. Springer, Heidelberg (2010). doi:10.1007/978-3-642-15558-1 22
  10. 10.Ericson, C.: Real-Time Collision Detection. The Morgan Kaufmann Series in Interactive 3-D Technology (2004)
  11. 11.Fan, X., Zheng, K., Zhou, Y., Wang, S.: Pose locality constrained representation for 3D human pose reconstruction. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8689, pp. 174–188. Springer, Heidelberg (2014). doi:10.1007/978-3-319-10590-1 12
  12. 12.Geman, S., McClure, D.: Statistical methods for tomographic image reconstruction. Bull. Int. Stat. Inst. 52(4), 5–21 (1987)
  13. 13.Grest, D., Koch, R.: Human model fitting from monocular posture images. In: Proceedings of VMV, pp. 665–1344 (2005)
  14. 14.Guan, P., Weiss, A., Balan, A., Black, M.J.: Estimating human shape and pose from a single image. In: IEEE International Conference on Computer Vision, ICCV, pp. 1381–1388 (2009)
  15. 15.Guan, P.: Virtual human bodies with clothing and hair: From images to animation. Ph.D. thesis, Brown University, Department of Computer Science, December 2012
  16. 16.Hasler, N., Ackermann, H., Rosenhahn, B., Thormhlen, T., Seidel, H.P.: Multilinear pose and body shape estimation of dressed subjects from image sets. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 1823–1830 (2010)
  17. 17.Ionescu, C., Carreira, J., Sminchisescu, C.: Iterated second-order label sensitive pooling for 3D human pose estimation. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 1661–1668 (2014)
  18. 18.Ionescu, C., Papava, D., Olaru, V., Sminchisescu, C.: Human3.6M: large scale datasets and predictive methods for 3D human sensing in natural environments. IEEE Trans. Pattern Anal. Mach. Intell. TPAMI 36(7), 1325–1339 (2014)
  19. 19.Jain, A., Thorm¨ahlen, T., Seidel, H.P., Theobalt, C.: MovieReshape: tracking and reshaping of humans in videos. ACM Trans. Graph. (TOG) - Proc. ACM SIGGRAPH 29(5), 148:1–148:10 (2010)
  20. 20.Jain, A., Tompson, J., LeCun, Y., Bregler, C.: MoDeep: a deep learning framework using motion features for human pose estimation. In: Cremers, D., Reid, I., Saito, H., Yang, M.-H. (eds.) ACCV 2014. LNCS, vol. 9004, pp. 302–315. Springer, Heidelberg (2015). doi:10.1007/978-3-319-16808-1 21
  21. 21.Jiang, H.: 3D human pose reconstruction using millions of exemplars. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 1674–1677 (2010)
  22. 22.Johnson, S., Everingham, M.: Clustered pose and nonlinear appearance models for human pose estimation. In: Proceedings of the British Machine Vision Conference, pp. 12.1-12.11 (2010)
  23. 23.Kiefel, M., Gehler, P.V.: Human pose estimation with fields of parts. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8693, pp. 331–346. Springer, Heidelberg (2014). doi:10.1007/978-3-319-10602-1 22
  24. 24.Kostrikov, I., Gall, J.: Depth sweep regression forests for estimating 3D human pose from images. In: Proceedings of the British Machine Vision Conference (2014)
  25. 25.Kulkarni, T.D., Kohli, P., Tenenbaum, J.B., Mansinghka, V.: Picture: a probabilistic programming language for scene perception. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 4390–4399 (2015)
  26. 26.Lee, H., Chen, Z.: Determination of 3D human body postures from a single view. Comput. Vis. Graph. Image Process. 30(2), 148–168 (1985)
  27. 27.Li, S., Chan, A.B.: 3D human pose estimation from monocular images with deep convolutional neural network. In: Cremers, D., Reid, I., Saito, H., Yang, M.-H. (eds.) ACCV 2014. LNCS, vol. 9004, pp. 332–347. Springer, Heidelberg (2015). doi:10.1007/978-3-319-16808-1 23
  28. 28.Loper, M.M., Black, M.J.: OpenDR: an approximate differentiable renderer. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8695, pp. 154–169. Springer, Heidelberg (2014). doi:10.1007/978-3-319-10584-0 11
  29. 29.Loper, M., Mahmood, N., Black, M.J.: MoSh: motion and shape capture from sparse markers. ACM Trans. Graph. (TOG) - Proc. ACM SIGGRAPH Asia 33(6), 220:1–220:13 (2014)
  30. 30.Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., Black, M.J.: SMPL: a skinned multi-person linear model. ACM Trans. Graph. (TOG) - Proc. ACM SIGGRAPH Asia 34(6), 248: 1–248: 16 (2015)
  31. 31.Nocedal, J., Wright, S.: Numerical Optimization. Springer, New York (2006)
  32. 32.Olson, E., Agarwal, P.: Inference on networks of mixtures for robust robot mapping. Int. J. Robot. Res. 32(7), 826–840 (2013)
  33. 33.Parameswaran, V., Chellappa, R.: View independent human body pose estimation from a single perspective image. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 16–22 (2004)
  34. 34.Pfister, T., Charles, J., Zisserman, A.: Flowing convnets for human pose estimation in videos. In: IEEE International Conference on Computer Vision, ICCV, pp. 1913–1921 (2015)
  35. 35.Pfister, T., Simonyan, K., Charles, J., Zisserman, A.: Deep convolutional neural networks for efficient pose estimation in gesture videos. In: Cremers, D., Reid, I., Saito, H., Yang, M.-H. (eds.) ACCV 2014. LNCS, vol. 9003, pp. 538–552. Springer, Heidelberg (2015). doi:10.1007/978-3-319-16865-4 35
  36. 36.Pishchulin, L., Insafutdinov, E., Tang, S., Andres, B., Andriluka, M., Gehler, P., Schiele, B.: DeepCut: joint subset partition and labeling for multi person pose estimation. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 4929–4937 (2016)
  37. 37.Pons-Moll, G., Fleet, D., Rosenhahn, B.: Posebits for monocular human pose estimation. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 2345–2352 (2014)
  38. 38.Pons-Moll, G., Taylor, J., Shotton, J., Hertzmann, A., Fitzgibbon, A.: Metric regression forests for correspondence estimation. Int. J. Comput. Vis. IJCV 113(3), 1–13 (2015)
  39. 39.Ramakrishna, V., Kanade, T., Sheikh, Y.: Reconstructing 3D human pose from 2D image landmarks. In: Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., Schmid, C. (eds.) ECCV 2012. LNCS, vol. 7575, pp. 573–586. Springer, Heidelberg (2012). doi:10.1007/978-3-642-33765-9 41
  40. 40.Rother, C., Kolmogorov, V., Blake, A.: Grabcut: interactive foreground extraction using iterated graph cuts. ACM Trans. Graph. (TOG) - Proc. ACM SIGGRAPH 23(3), 309–314 (2004)
  41. 41.Sigal, L., Balan, A., Black, M.J.: HumanEva: synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. Int. J. Comput. Vis. IJCV 87(1), 4–27 (2010)
  42. 42.Sigal, L., Balan, A., Black, M.J.: Combined discriminative and generative articulated pose and non-rigid shape estimation. In: Advances in Neural Information Processing Systems (NIPS), vol. 20, pp. 1337–1344 (2008)
  43. 43.Simo-Serra, E., Quattoni, A., Torras, C., Moreno-Noguer, F.: A joint model for 2D and 3D pose estimation from a single image. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 3634–3641 (2013)
  44. 44.Simo-Serra, E., Ramisa, A., Alenya, G., Torras, C., Moreno-Noguer, F.: Single image 3D human pose estimation from noisy observations. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 2673–2680 (2012)
  45. 45.Sminchisescu, C., Telea, A.: Human pose estimation from silhouettes, a consistent approach using distance level sets. In: WSCG International Conference for Computer Graphics, Visualization and Computer Vision, pp. 413–420 (2002)
  46. 46.Sminchisescu, C., Triggs, B.: Covariance scaled sampling for monocular 3D body tracking. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 447–454 (2001)
  47. 47.Sridhar, S., Mueller, F., Oulasvirta, A., Theobalt, C.: Fast and robust hand tracking using detection-guided optimization. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 3121–3221 (2015)
  48. 48.Taylor, C.: Reconstruction of articulated objects from point correspondences in single uncalibrated image. Comput. Vis. Image Underst. CVIU 80(10), 349–363 (2000)
  49. 49.Tekin, B., Rozantsev, A., Lepetit, V., Fua, P.: Direct prediction of 3D body poses from motion compensated sequences. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 991–1000 (2016)
  50. 50.Thiery, J.M., Guy, E., Boubekeur, T.: Sphere-meshes: shape approximation using spherical quadric error metrics. ACM Trans. Graph. (TOG) - Proc. ACM SIGGRAPH Asia 32(6), 178:1–178:12 (2013)
  51. 51.Toshev, A., Szegedy, C.: DeepPose: human pose estimation via deep neural networks. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 1653–1660 (2014)
  52. 52.Wang, C., Wang, Y., Lin, Z., Yuille, A., Gao, W.: Robust estimation of 3D human poses from a single image. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 2369–2376 (2014)
  53. 53.Wei, S.E., Ramakrishna, V., Kanade, T., Sheikh, Y.: Convolutional pose machines. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 4724–4732 (2016)
  54. 54.Yang, Y., Ramanan, D.: Articulated pose estimation using flexible mixtures of parts. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 3546–3553 (2011)
  55. 55.Yasin, H., Iqbal, U., Kr¨uger, B., Weber, A., Gall, J.: A dual-source approach for 3D pose estimation from a single image. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 4948–4956 (2016)
  56. 56.Zhou, F., Torre, F.: Spatio-temporal matching for human detection in video. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8694, pp. 62–77. Springer, Heidelberg (2014). doi:10.1007/978-3-319-10599-4 5
  57. 57.Zhou, S., Fu, H., Liu, L., Cohen-Or, D., Han, X.: Parametric reshaping of human bodies in images. ACM Trans. Graph. (TOG) - Proc. ACM SIGGRAPH 29(4), 126:1–126:10 (2010)
  58. 58.Zhou, X., Zhu, M., Leonardos, S., Derpanis, K., Daniilidis, K.: Sparse representation for 3D shape estimation: a convex relaxation approach. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 4447–4455 (2015)
  59. 59.Zhou, X., Zhu, M., Leonardos, S., Derpanis, K., Daniilidis, K.: Sparseness meets deepness: 3D human pose estimation from monocular video. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 4447–4455 (2016)

Citation

MLA
Bogo, F., et al. “Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image”. arXiv, 2016, http://arxiv.org/abs/1607.08128v1.
APA
Bogo, F., Kanazawa, A., Lassner, C., Gehler, P., Romero, J., & Black, M. J. (2016). Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image. arXiv. http://arxiv.org/abs/1607.08128v1
Chicago
Bogo, F., A. Kanazawa, C. Lassner, P. Gehler, J. Romero, and M. J. Black. 2016. “Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image”. arXiv. http://arxiv.org/abs/1607.08128v1.
Harvard
Bogo, F. et al. (2016) “Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1607.08128v1.
Vancouver
1. Bogo F, Kanazawa A, Lassner C, Gehler P, Romero J, Black MJ (2016) Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image. arXiv

BibTeX

@article{bogo2016keep,
  title = {Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image},
  author = {Bogo, Federica and Kanazawa, Angjoo and Lassner, Christoph and Gehler, Peter and Romero, Javier and Black, Michael J.},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1607.08128v1},
  eprint = {1607.08128}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF