End-to-End Recovery of Human Shape and Pose

Angjoo KanazawaMichael J. BlackDavid W. JacobsJitendra Malik

article2017CVPR2,129 citations

Introduces an end-to-end deep learning framework that reconstructs full 3D human body meshes directly from single RGB images in real time, using adversarial training to learn valid body shapes and poses without requiring paired 3D annotations.

Listen

Reconstructing full three-dimensional (3D) human body meshes from single two-dimensional (2D) images is a crucial capability for animation, virtual reality, and computer vision. However, standard methods primarily predict sparse 3D joint skeletons rather than full surfaces, or they rely on multi-stage optimization pipelines that require up to a minute per image and struggle to generalize to natural, unconstrained environments where ground-truth 3D data is scarce.

The article demonstrates an end-to-end framework, called Human Mesh Recovery (HMR), that directly infers full 3D human pose, body shape, and camera parameters in real time from a single RGB image. It aims to eliminate multi-step pipelines and enable accurate 3D mesh reconstruction even when trained without paired 3D annotations.

To accomplish this, the authors designed a deep learning system that directly maps image pixels to parameters of a standard generative human body model (SMPL) using an iterative feedback loop. The framework is trained to ensure the projected 3D keypoints match annotated 2D image keypoints. To prevent the model from generating anatomically impossible bodies that happen to fit 2D projections, the architecture integrates a factorized adversarial discriminator. This discriminator is trained on an unpaired dataset of 3D motion-capture body scans (over 700,000 samples) to ensure the inferred joint angles and body shapes conform to real human anatomy without requiring hand-crafted geometric constraints.

Key findings show that HMR substantially outperforms prior mesh-recovery methods while running in real time at approximately 0.04 seconds per image—roughly 1,500 times faster than earlier optimization methods that took 20 to 60 seconds. On standard benchmark datasets such as Human3.6M, HMR achieved a reconstruction error of 56.8 mm, outperforming previous mesh estimation approaches (which had errors around 80.7 to 93.9 mm) and competing closely with specialized 3D skeleton-only models. Crucially, when trained entirely without paired 3D supervision, the system still achieved strong performance (66.5 mm reconstruction error) and produced plausible body shapes, whereas removing the adversarial component produced severely distorted, unnatural shapes. The model also matched top-performing benchmarks on auxiliary tasks like body part segmentation.

These results show that high-fidelity 3D human analysis can be scaled efficiently using readily available 2D in-the-wild images alongside unpaired 3D scan data. This lowers the operational and data collection costs associated with expensive motion-capture labs and unlocks practical, low-latency deployment for real-time video analytics and interactive graphics.

Organizations developing 3D human capture applications should adopt direct feedforward regression models rather than slow multi-stage optimization pipelines. Stakeholders should also consider expanding training datasets with in-the-wild 2D labeled images to further refine real-world robustness. Future work should focus on testing whether scaling up 2D data alone can fully close the accuracy gap with supervised models, as well as evaluating performance across wider demographics and video-based temporal sequences.

The findings carry high confidence for monocular images within standard bounding boxes, though some caution is warranted regarding extreme occlusions, boundary ambiguities, and noisy ground-truth labels in existing evaluation benchmarks.

Cover for End-to-End Recovery of Human Shape and Pose

Abstract

We describe Human Mesh Recovery (HMR), an end-to-end framework for reconstructing a full 3D mesh of a human body from a single RGB image. In contrast to most current methods that compute 2D or 3D joint locations, we produce a richer and more useful mesh representation that is parameterized by shape and 3D joint angles. The main objective is to minimize the reprojection loss of keypoints, which allow our model to be trained using images in-the-wild that only have ground truth 2D annotations. However, the reprojection loss alone leaves the model highly under constrained. In this work we address this problem by introducing an adversary trained to tell whether a human body parameter is real or not using a large database of 3D human meshes. We show that HMR can be trained with and without using any paired 2D-to-3D supervision. We do not rely on intermediate 2D keypoint detections and infer 3D pose and shape parameters directly from image pixels. Our model runs in real-time given a bounding box containing the person. We demonstrate our approach on various images in-the-wild and out-perform previous optimization based methods that output 3D meshes and show competitive results on tasks such as 3D joint location estimation and part segmentation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Model
  • 3.1 3D Body Representation
  • 3.2 Iterative 3D Regression with Feedback
  • 3.3 Factorized Adversarial Prior
  • 3.4 Implementation Details
  • 4 Experimental Results
  • 4.1 3D Joint Location Estimation
  • 4.2 Human Body Segmentation
  • 4.3 Without Paired 3D Supervision
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — HMR 3D Human Body and Camera Parameterization

    model/method

    Human Mesh Recovery (HMR) reconstructs a 3D mesh of a human body and a weak-perspective camera directly from a single RGB image II using the Skinned Multi-Person Linear (SMPL) model.

    The 3D human body and camera are parameterized by an 85-dimensional vector Θ={θ,β,R,t,s}\Theta = \{\theta, \beta, R, t, s\}:

    • Body shape β∈R10\beta \in \mathbb{R}^{10} represents the first 10 coefficients of the SMPL PCA shape space, controlling height, weight, and individual body proportions.
    • Body pose θ∈R3K\theta \in \mathbb{R}^{3K} represents the relative 3D rotations of K=23K = 23 kinematic joints in axis-angle format.
    • Global body rotation R∈R3×3R \in \mathbb{R}^{3 \times 3} is parameterized in axis-angle format.
    • Camera translation t∈R2t \in \mathbb{R}^2 and camera scale s∈Rs \in \mathbb{R} define the weak-perspective projection.

    The SMPL function M(θ,β)∈R3×NM(\theta, \beta) \in \mathbb{R}^{3 \times N} is a differentiable mapping that produces a triangulated surface mesh of N=6980N = 6980 vertices by shaping template vertices according to β\beta and θ\theta, articulating bones via forward kinematics, and applying linear blend skinning. A linear regression matrix maps the resulting mesh vertices to P=19P = 19 3D keypoint positions X(θ,β)∈R3×PX(\theta, \beta) \in \mathbb{R}^{3 \times P} (14 standard body joints plus 5 facial landmark vertices: nose (vertex 333), left eye (vertex 2801), right eye (vertex 6261), left ear (vertex 584), and right ear (vertex 4072)).

    The 2D projection x^∈R2×P\hat{x} \in \mathbb{R}^{2 \times P} of the 3D keypoints is computed via orthographic projection Π\Pi: x^=sΠ(RX(θ,β))+t\hat{x} = s \Pi(R X(\theta, \beta)) + t

  2. Knowl 2 — Iterative 3D Regression with Latent Feature Feedback

    model/method

    Instead of regressing the 85-dimensional parameter vector Θ\Theta in a single feedforward step or using image-space reprojections, HMR uses an Iterative Error Feedback (IEF) loop operating entirely within a latent feature space.

    An input image II is encoded by a ResNet-50 network and average-pooled into an image feature vector ϕ∈R2048\phi \in \mathbb{R}^{2048}. At iteration tt, the 3D regression module receives the concatenated vector [ϕ,Θt][\phi, \Theta_t] and predicts a parameter residual ΔΘt∈R85\Delta \Theta_t \in \mathbb{R}^{85}: Θt+1=Θt+ΔΘt\Theta_{t+1} = \Theta_t + \Delta \Theta_t The parameters are initialized to the mean human parameter vector Θ0=Θˉ\Theta_0 = \bar{\Theta}. The regression module consists of two fully connected layers of 1024 units each with dropout, followed by an 85-dimensional linear output layer. The iteration loop runs for T=3T = 3 steps.

    To prevent the regressor from overshooting and getting trapped in local minima during iterative updates, the 2D reprojection loss Lreproj\mathcal{L}_{\text{reproj}} and direct 3D losses L3D\mathcal{L}_{\text{3D}} are applied solely to the final parameter estimate ΘT\Theta_T. In contrast, the adversarial loss Ladv\mathcal{L}_{\text{adv}} is applied at every intermediate iteration Θt\Theta_t for t∈{1,…,T}t \in \{1, \dots, T\}, ensuring that each corrective step remains on the manifold of plausible 3D human bodies.

  3. Knowl 3 — Factorized Adversarial Prior Architecture and Objective

    model/method

    To prevent 2D-to-3D lifting from predicting anatomically implausible poses or extreme shapes that satisfy 2D reprojection, HMR incorporates a factorized adversarial prior trained on an unpaired database of 3D human meshes.

    The discriminator system decomposes into K+2=25K + 2 = 25 distinct discriminators (for K=23K = 23 joints):

    1. Shape Discriminator (DβD_\beta): A 2-layer multilayer perceptron (layer sizes 10→10→5→110 \to 10 \to 5 \to 1 with ReLU activations) that classifies the 10-dimensional PCA shape coefficient vector β\beta.
    2. Joint Rotation Discriminators (Dθ,kD_{\theta, k}): Each joint's 3D axis-angle rotation vector in θ\theta is converted into a 3×33 \times 3 rotation matrix (9-D) via Rodrigues' formula and passed to a shared embedding network of two fully connected layers (32 hidden units). Individual 1-D linear classifier heads Dθ,kD_{\theta, k} (k∈{1,…,23}k \in \{1, \dots, 23\}) determine if each joint angle is valid.
    3. Kinematic Pose Discriminator (DposeD_{\text{pose}}): Concatenates all 23×3223 \times 32 joint embeddings (736736-D) and passes them through two fully connected layers with 1024 units each to output a single scalar, modeling the joint distribution across the entire kinematic tree.

    The encoder EE and discriminators DiD_i are trained using the Least Squares GAN (LSGAN) objective: min⁡ELadv(E)=∑i=1K+2EΘ∼pE[(Di(E(I))−1)2]\min_E \mathcal{L}_{\text{adv}}(E) = \sum_{i=1}^{K+2} \mathbb{E}_{\Theta \sim p_E} \left[ (D_i(E(I)) - 1)^2 \right] min⁡DiL(Di)=EΘ∼pdata[(Di(Θ)−1)2]+EΘ∼pE[Di(E(I))2]\min_{D_i} \mathcal{L}(D_i) = \mathbb{E}_{\Theta \sim p_{\text{data}}} \left[ (D_i(\Theta) - 1)^2 \right] + \mathbb{E}_{\Theta \sim p_E} \left[ D_i(E(I))^2 \right]

  4. Knowl 4 — HMR Multi-Task Training Objective

    equation

    The complete objective function for training the Human Mesh Recovery (HMR) encoder network is: L=λ(Lreproj+13DL3D)+Ladv\mathcal{L} = \lambda (\mathcal{L}_{\text{reproj}} + \mathbf{1}_{\text{3D}} \mathcal{L}_{\text{3D}}) + \mathcal{L}_{\text{adv}} where λ\lambda balances objective weighting and 13D\mathbf{1}_{\text{3D}} is an indicator function that equals 1 if paired 3D ground truth is available for the training sample and 0 otherwise.

    The constituent loss terms are defined as follows:

    • 2D Reprojection Loss: Lreproj=∑i=1P∥vi(xi−x^i)∥1\mathcal{L}_{\text{reproj}} = \sum_{i=1}^P \|v_i (x_i - \hat{x}_i)\|_1 where xi∈R2x_i \in \mathbb{R}^2 is the ground truth 2D location for keypoint ii, x^i∈R2\hat{x}_i \in \mathbb{R}^2 is the projected keypoint from the predicted 3D mesh, and vi∈{0,1}v_i \in \{0, 1\} is keypoint visibility (11 if visible, 00 otherwise) over P=19P = 19 keypoints.
    • 3D Supervision Loss: L3D=L3D joints+L3D smpl=∑i∥Xi−X^i∥22+∥[βi,θi]−[β^i,θ^i]∥22\mathcal{L}_{\text{3D}} = \mathcal{L}_{\text{3D joints}} + \mathcal{L}_{\text{3D smpl}} = \sum_{i} \|X_i - \hat{X}_i\|_2^2 + \|[\beta_i, \theta_i] - [\hat{\beta}_i, \hat{\theta}_i]\|_2^2 where XiX_i and X^i\hat{X}_i denote ground truth and predicted 3D joint coordinates, and [βi,θi][\beta_i, \theta_i] and [β^i,θ^i][\hat{\beta}_i, \hat{\theta}_i] denote ground truth and predicted SMPL shape and pose parameters.
    • Adversarial Loss: Ladv\mathcal{L}_{\text{adv}} is the sum of least-squares adversarial errors across the K+2K + 2 shape and pose discriminators.
  5. Knowl 5 — Unpaired 2D-to-3D Weakly Supervised Training Setup

    experimental setup

    HMR is trained using a combination of in-the-wild 2D annotated images, indoor 3D annotated datasets, and an unpaired database of 3D human meshes.

    • 2D Keypoint Datasets: LSP (1k images), LSP-extended (10k images), MPII (20k images), and MS COCO (80k images), filtered to retain samples with at least 6 visible keypoints.
    • Paired 3D Datasets: Human3.6M (150k training frames with 3D joint annotations and MoSh-derived ground truth SMPL parameters) and MPI-INF-3DHP (150k training frames with 3D joint annotations; Subject 8 is reserved for hyperparameter validation).
    • Unpaired 3D Mesh Prior Data: 3D meshes obtained by fitting SMPL models via MoSh to MoCap databases: CMU MoCap (~390k samples), Human3.6M training mocap (~150k samples), and PosePrior dataset (~180k samples).
    • Training Details: Input images are cropped and scaled to 224×224224 \times 224 pixels (diagonal of tight bounding box ≈150\approx 150 px) with random scaling, translation, and horizontal flipping. Mini-batch size is 64 (50% 2D in-the-wild and 50% 3D paired when 3D supervision is enabled). Optimization uses Adam with learning rate 10−510^{-5} for the encoder and 10−410^{-4} for the discriminators across 55 epochs on an NVIDIA Titan 1080Ti GPU (~5 days).
  6. Knowl 6 — 3D Joint Estimation Performance on Human3.6M

    data/table

    The tables below compare HMR against existing single-view 3D human pose and shape estimation methods on the Human3.6M dataset under Protocol 1 (tested on subjects S9 and S11 across all cameras) and Protocol 2 (tested on frontal camera 3). Methods marked with * produce full 3D body meshes or kinematic parameters beyond sparse 3D joints.

    Method Reconst. Error (mm)
    Rogez et al. [35] 87.3
    Pavlakos et al. [33] 51.9
    Martinez et al. [26] 47.7
    *Regression Forest from 91 kps [20] 93.9
    *SMPLify [5] 82.3
    *SMPLify from 91 kps [20] 80.7
    *HMR (proposed) 56.8
    *HMR unpaired (proposed) 66.5
    Method MPJPE (mm) Reconst. Error (mm)
    Tome et al. [44] 88.39 -
    Rogez et al. [36] 87.7 71.6
    VNect [28] 80.5 -
    Pavlakos et al. [33] 71.9 51.23
    Mehta et al. [27] 68.6 -
    Sun et al. [40] 59.1 -
    *Deep Kinematic Pose [52] 107.26 -
    *HMR (proposed) 87.97 58.1
    *HMR unpaired (proposed) 106.84 67.45

    Reconstruction error denotes Mean Per Joint Position Error (MPJPE) after rigid Procrustes alignment. Under Protocol 2, HMR reduces reconstruction error by 25.525.5 mm compared to SMPLify (56.856.8 mm vs. 82.382.3 mm). Even without any paired 3D supervision (HMR unpaired), the reconstruction error (66.566.5 mm) outperforms all previous SMPL optimization and regression baselines.

  7. Knowl 7 — Evaluation on MPI-INF-3DHP Benchmark

    data/table

    Quantitative results on the 2929 test frames of the MPI-INF-3DHP dataset (covering 6 subjects across 7 indoor and outdoor actions), evaluated before and after rigid Procrustes alignment. Methods marked with * output 3D meshes.

    Absolute After Rigid Alignment
    Method PCK (%) AUC MPJPE (mm) PCK (%) AUC MPJPE (mm)
    Mehta et al. [27] 75.7 39.3 117.6 - - -
    VNect [28] 76.6 40.4 124.7 83.9 47.3 98.0
    *HMR (proposed) 72.9 36.5 124.2 86.3 47.8 89.8
    *HMR unpaired (proposed) 59.6 27.9 169.5 77.1 40.7 113.2

    PCK is thresholded at 150 mm and AUC is evaluated across multiple PCK thresholds. After rigid alignment, fully supervised HMR achieves a PCK of 86.3% and MPJPE of 89.8 mm, outperforming VNect (83.9% PCK, 98.0 mm MPJPE).

  8. Knowl 8 — Human Body Part Segmentation on LSP

    data/table

    Evaluation of human body foreground and part segmentation on the 1000 test images of the Leeds Sports Pose (LSP) dataset. The task tests foreground-versus-background (Fg vs Bg) and 6 body parts plus background segmentation accuracy (Acc) and average F1 score.

    Fg vs Bg Parts
    Method Acc (%) F1 Acc (%) F1 Run Time
    SMPLify oracle [20] 92.17 0.88 88.82 0.67 -
    SMPLify [5] 91.89 0.88 87.71 0.64 ∼\sim1 min
    Decision Forests [20] 86.60 0.80 82.32 0.51 0.13 sec
    HMR (proposed) 91.67 0.87 87.12 0.60 0.04 sec
    HMR unpaired (proposed) 91.30 0.86 87.00 0.59 0.04 sec

    HMR reaches performance comparable to the optimization-based SMPLify oracle (which uses ground truth segmentations to guide fitting) while performing inference in feedforward real time (0.040.04 seconds per person vs. ∼\sim1 minute for SMPLify).

  9. Knowl 9 — Necessity of Adversarial Regularization in Unpaired 3D Mesh Recovery

    empirical result

    When trained without paired 3D ground truth (setting 13D=0\mathbf{1}_{\text{3D}} = 0), relying solely on the 2D keypoint reprojection loss Lreproj\mathcal{L}_{\text{reproj}} produces extreme failure modes. Due to single-view depth and scale ambiguities, the network minimizes 2D reprojection error by outputting severely distorted, anthropometrically invalid 3D meshes ("monster" geometries with unnatural limb lengths, gross self-intersections, and impossible joint angles) whose 2D projected keypoints nevertheless match ground truth 2D annotations accurately.

    Introducing the factorized adversarial prior regularizes the network to restrict its parameter updates to the manifold of real human shapes and poses. This enables successful weakly-supervised 3D mesh reconstruction from in-the-wild 2D annotations without any paired 3D supervision, yielding a Protocol 2 reconstruction error of 66.5 mm on Human3.6M.

Coverage note — None was omitted; all key architectural components, objectives, datasets, quantitative benchmark tables (Human3.6M Protocols 1 & 2, MPI-INF-3DHP, and LSP Part Segmentation), and empirical ablation insights are included.

References

  1. 1.M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al. Tensorflow: A system for large-scale machine learning. In Operating Systems Design and Implementation, 2016. 6
  2. 2.I. Akhter and M. J. Black. Pose-conditioned joint angle limits for 3D human pose reconstruction. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, June 2015. 3, 6
  3. 3.M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele. 2d human pose estimation: New benchmark and state of the art analysis. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, June 2014. 6
  4. 4.C. Barron and I. Kakadiaris. Estimating anthropometry and pose from a single uncalibrated image. Computer Vision and Image Understanding, CVIU, 81(3):269–284, 2001. 3
  5. 5.F. Bogo, A. Kanazawa, C. Lassner, P. Gehler, J. Romero, and M. J. Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In European Conference on Computer Vision, ECCV, Lecture Notes in Computer Science. Springer International Publishing, Oct. 2016. 2, 3, 4, 5, 6, 7, 8
  6. 6.F. by NSF EIA-0196217. Cmu graphics lab - motion capture library. http://mocap.cs.cmu.edu/. 6
  7. 7.J. Carreira, P. Agrawal, K. Fragkiadaki, and J. Malik. Human pose estimation with iterative error feedback. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2016. 5
  8. 8.Y. Chen, T.-K. Kim, and R. Cipolla. Inferring 3D shapes and deformations from single views. In European Conference on Computer Vision, ECCV, pages 300–313, 2010. 3
  9. 9.P. Dollár, P. Welinder, and P. Perona. Cascaded pose regression. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pages 1078–1085. IEEE, 2010. 5
  10. 10.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in neural information processing systems, pages 2672–2680, 2014. 5
  11. 11.J. C. Gower. Generalized procrustes analysis. Psychometrika, 40(1):33–51, Mar 1975. 7
  12. 12.P. Guan, A. Weiss, A. Balan, and M. J. Black. Estimating human shape and pose from a single image. In IEEE International Conference on Computer Vision, ICCV, pages 1381–1388, 2009. 3
  13. 13.R. A. Güler, G. Trigeorgis, E. Antonakos, P. Snape, S. Zafeiriou, and I. Kokkinos. Densereg: Fully convolutional dense shape regression in-the-wild. IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2017. 4
  14. 14.N. Hasler, H. Ackermann, B. Rosenhahn, T. Thormählen, and H. P. Seidel. Multilinear pose and body shape estimation of dressed subjects from image sets. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pages 1823–1830, 2010. 3
  15. 15.K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. In European Conference on Computer Vision, ECCV, 2016. 6
  16. 16.C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu. Human3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments. pami, 36(7):1325–1339, 2014. 3, 6
  17. 17.S. Johnson and M. Everingham. Clustered pose and nonlinear appearance models for human pose estimation. In bmvc, pages 12.1–12.11, 2010. 6, 8
  18. 18.D. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 6
  19. 19.T. D. Kulkarni, P. Kohli, J. B. Tenenbaum, and V. Mansinghka. Picture: A probabilistic programming language for scene perception. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pages 4390–4399, 2015. 4
  20. 20.C. Lassner, J. Romero, M. Kiefel, F. Bogo, M. J. Black, and P. V. Gehler. Unite the people: Closing the loop between 3d and 2d human representations. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, July 2017. 2, 3, 4, 6, 7, 8
  21. 21.H. Lee and Z. Chen. Determination of 3D human body postures from a single view. Computer Vision Graphics and Image Processing, 30(2):148–168, 1985. 3
  22. 22.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision (ECCV), Zürich, 2014. Oral. 4, 6
  23. 23.M. Loper, N. Mahmood, and M. J. Black. MoSh: Motion and shape capture from sparse markers. ACM Transactions on Graphics (TOG) - Proceedings of ACM SIGGRAPH Asia, 33(6):220:1–220:13, 2014. 5, 6
  24. 24.M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, Oct. 2015. 1, 4
  25. 25.X. Mao, Q. Li, H. Xie, R. Y. K. Lau, Z. Wang, and S. P. Smolley. Least squares generative adversarial networks, 2016. 5
  26. 26.J. Martinez, R. Hossain, J. Romero, and J. J. Little. A simple yet effective baseline for 3d human pose estimation. In IEEE International Conference on Computer Vision, ICCV, 2017. 3, 7
  27. 27.D. Mehta, H. Rhodin, D. Casas, P. Fua, O. Sotnychenko, W. Xu, and C. Theobalt. Monocular 3d human pose estimation in the wild using improved cnn supervision. In Proc. of International Conference on 3D Vision (3DV), 2017. 3, 6, 7, 8
  28. 28.D. Mehta, S. Sridhar, O. Sotnychenko, H. Rhodin, M. Shafiei, H.-P. Seidel, W. Xu, D. Casas, and C. Theobalt. Vnect: Real-time 3d human pose estimation with a single rgb camera. ACM Transactions on Graphics (TOG) - Proceedings of ACM SIGGRAPH, 36, July 2017. 3, 4, 6, 7, 8
  29. 29.F. Moreno-Noguer. 3d human pose estimation from a single image via distance matrix regression. arXiv preprint arXiv:1611.09010, 2016. 3
  30. 30.A. Newell, K. Yang, and J. Deng. Stacked hourglass networks for human pose estimation. In European Conference on Computer Vision, ECCV, pages 483–499, 2016. 3
  31. 31.M. Oberweger, P. Wohlhart, and V. Lepetit. Training a feedback loop for hand pose estimation. In Proceedings of the IEEE International Conference on Computer Vision, pages 3316–3324, 2015. 5
  32. 32.V. Parameswaran and R. Chellappa. View independent human body pose estimation from a single perspective image. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pages 16–22, 2004. 3
  33. 33.G. Pavlakos, X. Zhou, K. G. Derpanis, and K. Daniilidis. Coarse-to-fine volumetric prediction for single-image 3D human pose. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2017. 3, 7, 8
  34. 34.V. Ramakrishna, T. Kanade, and Y. Sheikh. Reconstructing 3d Human Pose from 2d Image Landmarks. Computer Vision–ECCV 2012, pages 573–586, 2012. 3, 4
  35. 35.G. Rogez and C. Schmid. Mocap-guided data augmentation for 3d pose estimation in the wild. In Advances in Neural Information Processing Systems, pages 3108–3116, 2016. 3, 7
  36. 36.G. Rogez, P. Weinzaepfel, and C. Schmid. LCR-Net: Localization-Classification-Regression for Human Pose. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, July 2017. 3, 7
  37. 37.M. Sanzari, V. Ntouskos, and F. Pirri. Bayesian image based 3d pose estimation. In European Conference on Computer Vision, ECCV, pages 566–582, 2016. 3
  38. 38.L. Sigal, A. Balan, and M. J. Black. HumanEva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. International Journal of Computer Vision, IJCV, 87(1):4–27, 2010. 3
  39. 39.N. Silberman and S. Guadarrama. Tensorflow-slim image classification model library. https://github.com/tensorflow/models/tree/master/research/slim. 6
  40. 40.X. Sun, J. Shang, S. Liang, and Y. Wei. Compositional human pose regression. In IEEE International Conference on Computer Vision, ICCV, 2017. 3, 4, 7
  41. 41.J. K. V. Tan, I. Budvytis, and R. Cipolla. Indirect deep structured learning for 3d human shape and pose prediction. In Proceedings of the British Machine Vision Conference, 2017. 4
  42. 42.C. Taylor. Reconstruction of articulated objects from point correspondences in single uncalibrated image. Computer Vision and Image Understanding, CVIU, 80(10):349–363, 2000. 2, 3
  43. 43.B. Tekin, P. Marquez Neila, M. Salzmann, and P. Fua. Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation. In IEEE International Conference on Computer Vision, ICCV, 2017. 3
  44. 44.D. Tome, C. Russell, and L. Agapito. Lifting from the deep: Convolutional 3d pose estimation from a single image. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017. 3, 7
  45. 45.S. Tulsiani and J. Malik. Viewpoints and keypoints. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1510–1519, 2015. 3
  46. 46.H.-Y. Tung, H.-W. Tung, E. Yumer, and K. Fragkiadaki. Self-supervised learning of motion capture. In Advances in Neural Information Processing Systems, pages 5242–5252, 2017. 4
  47. 47.H.-Y. F. Tung, A. W. Harley, W. Seto, and K. Fragkiadaki. Adversarial inverse graphics networks: Learning 2d-to-3d lifting and image-to-image translation from unpaired supervision. In IEEE International Conference on Computer Vision, ICCV, 2017. 3, 5, 8
  48. 48.G. Varol, J. Romero, X. Martin, N. Mahmood, M. J. Black, I. Laptev, and C. Schmid. Learning from Synthetic Humans. In CVPR, 2017. 4, 5
  49. 49.S.-E. Wei, V. Ramakrishna, T. Kanade, and Y. Sheikh. Convolutional pose machines. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pages 4724–4732, 2016. 3
  50. 50.J. Wu, T. Xue, J. J. Lim, Y. Tian, J. B. Tenenbaum, A. Torralba, and W. T. Freeman. Single image 3d interpreter network. In European Conference on Computer Vision, ECCV, 2016. 3, 8
  51. 51.X. Zhou, Q. Huang, X. Sun, X. Xue, and Y. Wei. Weakly-supervised transfer for 3d human pose estimation in the wild. In IEEE International Conference on Computer Vision, ICCV, 2017. 3, 4
  52. 52.X. Zhou, X. Sun, W. Zhang, S. Liang, and Y. Wei. Deep kinematic pose regression. In ECCV Workshop on Geometry Meets Deep Learning, pages 186–201, 2016. 3, 4, 5, 7
  53. 53.X. Zhou, M. Zhu, S. Leonardos, K. Derpanis, and K. Daniilidis. Sparse representation for 3D shape estimation: A convex relaxation approach. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pages 4447–4455, 2015. 3
  54. 54.X. Zhou, M. Zhu, S. Leonardos, K. Derpanis, and K. Daniilidis. Sparseness meets deepness: 3D human pose estimation from monocular video. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pages 4966–4975, 2016. 3
  55. 55.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. ICCV, 2017. 4

Citation

MLA
Kanazawa, A., et al. “End-to-end Recovery of Human Shape and Pose”. arXiv, 2017, http://arxiv.org/abs/1712.06584v2.
APA
Kanazawa, A., Black, M. J., Jacobs, D. W., & Malik, J. (2017). End-to-end Recovery of Human Shape and Pose. arXiv. http://arxiv.org/abs/1712.06584v2
Chicago
Kanazawa, A., M. J. Black, D. W. Jacobs, and J. Malik. 2017. “End-to-end Recovery of Human Shape and Pose”. arXiv. http://arxiv.org/abs/1712.06584v2.
Harvard
Kanazawa, A. et al. (2017) “End-to-end Recovery of Human Shape and Pose”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1712.06584v2.
Vancouver
1. Kanazawa A, Black MJ, Jacobs DW, Malik J (2017) End-to-end Recovery of Human Shape and Pose. arXiv

BibTeX

@article{kanazawa2017end,
  title = {End-to-end Recovery of Human Shape and Pose},
  author = {Kanazawa, Angjoo and Black, Michael J. and Jacobs, David W. and Malik, Jitendra},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1712.06584v2},
  eprint = {1712.06584}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE