Monocular 3D Human Pose Estimation in the Wild Using Improved CNN Supervision

Dushyant MehtaHelge RhodinDan CasasPascal FuaOleksandr SotnychenkoWeipeng XuChristian Theobalt

article20163DV1,355 citations

Develops a monocular 3D human pose estimation method that integrates 2D-to-3D transfer learning with a diverse marker-less motion capture dataset, significantly improving model generalization across unconstrained in-the-wild scenes.

Listen

Estimating three-dimensional (3D) human body pose from a single standard camera image is a critical capability for applications in robotics, augmented reality, biomechanics, and surveillance. While computer vision models perform well on two-dimensional (2D) joint detection across diverse environments, 3D pose estimation has historically struggled to operate outside controlled laboratory settings. This limitation stems from the scarcity of diverse 3D training data, as traditional motion-capture systems require restrictive marker suits, studio environments, or specialized sensors that fail to reflect complex, real-world appearances, clothing, and camera perspectives.

The article demonstrates a feedforward deep learning approach that significantly improves the accuracy and generalizability of single-camera 3D human pose estimation in uncontrolled, in-the-wild environments. The core objective is to overcome data scarcity through transfer learning from abundant 2D datasets, enhanced neural network training mechanisms, and a newly created multi-view markerless dataset.

The authors developed a three-stage methodology. First, a 2D pose network identifies the subject's location and joints in the frame. Second, a 3D pose network directly predicts body joint coordinates relative to the pelvis using multi-modal joint representations and a training technique called corrective skip connections. Third, the localized crop is projected back into full camera coordinates using closed-form perspective correction and 3D alignment without iterative optimization. To train and evaluate the framework, the authors introduced the MPI-INF-3DHP dataset—containing over 1.3 million frames of eight actors captured via markerless multi-camera motion capture—featuring diverse casual clothing, complex motions, background compositing, and a dedicated outdoor test benchmark.

The findings show that transferring mid- and high-level representations from 2D pose models into the 3D network yields massive performance gains, achieving 64.7% correct 3D keypoints on the new in-the-wild benchmark using existing laboratory data alone, compared to 41.4% with standard domain adaptation and 26.0% without transfer learning. When combining this transfer learning scheme with the new augmented dataset, the model achieved a state-of-the-art accuracy of 76.5% correct keypoints on in-the-wild tests and reduced error on standard laboratory benchmarks to 72.88 millimeters. Furthermore, multi-modal pose fusion markedly improved performance on difficult non-upright body positions, reducing joint error by 3.5 millimeters on sitting poses and 5.5 millimeters on crouching poses, while closed-form perspective correction added a 3 percentage point boost in keypoint accuracy.

These results demonstrate that high-quality 3D human tracking in unconstrained environments does not require expensive multi-sensor setups, specialized suits, or slow optimization routines. Because the feedforward model processes images in under 250 milliseconds per frame, it offers a viable, low-cost path for deploying 3D human sensing at scale using ordinary single-lens cameras.

Organizations developing monocular computer vision systems should adopt transfer learning from 2D pose networks as a standard architectural baseline, rather than relying solely on synthetic avatars or standard domain adaptation. For production deployment, teams should integrate the proposed perspective correction step into cropped image pipelines. Next steps should focus on optimizing network architectures to achieve real-time throughput below 30 milliseconds per frame and integrating temporal filtering to smooth out frame-to-frame tracking jitter in continuous video streams.

Confidence in the reported laboratory and outdoor benchmark accuracy is high, as the methods were rigorously tested against established datasets. However, practitioners should note that the current training corpus still predominantly uses chest-height camera angles, meaning accuracy may degrade when processing images taken from severe top-down or low-angle viewpoints until multi-elevation training sets are integrated.

arXiv: 1611.09813
Cover for Monocular 3D Human Pose Estimation in the Wild Using Improved CNN Supervision

Abstract

We propose a CNN-based approach for 3D human body pose estimation from single RGB images that addresses the issue of limited generalizability of models trained solely on the starkly limited publicly available 3D pose data. Using only the existing 3D pose data and 2D pose data, we show state-of-the-art performance on established benchmarks through transfer of learned features, while also generalizing to in-the-wild scenes. We further introduce a new training set for human body pose estimation from monocular images of real humans that has the ground truth captured with a multi-camera marker-less motion capture system. It complements existing corpora with greater diversity in pose, human appearance, clothing, occlusion, and viewpoints, and enables an increased scope of augmentation. We also contribute a new benchmark that covers outdoor and indoor scenes, and demonstrate that our 3D pose dataset shows better in-the-wild performance than existing annotated data, which is further improved in conjunction with transfer learning from 2D pose data. All in all, we argue that the use of transfer learning of representations in tandem with algorithmic and data contributions is crucial for general 3D body pose estimation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 CNN-based 3D Pose Estimation
  • 3.1 Bounding Box and 2D Pose Computation
  • 3.2 3D Pose Regression
  • 3.2.1 Multi-level Corrective Skip Connections
  • 3.2.2 Multi-modal Pose Fusion
  • 3.3 Global Pose Computation
  • 4 Transfer Learning
  • 5 MPI-INF-3DHP: Human Pose Dataset
  • 6 Experiments and Evaluation
  • 6.1 Impact of Supervision Methods
  • 6.2 Transfer Learning
  • 6.3 Benefit of MPI-INF-3DHP
  • 6.4 Other Components
  • 6.5 Quantitative Comparison
  • 7 Discussion
  • 8 Conclusion
  • 1 Further Discussion of Design Choices Regarding Multi-modal Fusion
  • 2 Further Discussion of Multi-level Corrective Skip
  • 3 Global Pose Computation
  • 3.1 3D localization
  • 3.2 Perspective correction
  • 4 CNN Architecture and Training Specifics
  • 4.1 2DPoseNet
  • 4.2 3DPoseNet
  • 4.2.1 3D Pose Training Data
  • 4.2.2 Domain Adaptation To In The Wild 2D Pose Data
  • 5 MPI-INF-3DHP Dataset
  • 5.1 The Challenge of Learning Invariance to Viewpoint Elevation
  • References

Knowls

  1. Knowl 1 — Multi-Modal Kinematic 3D Human Pose Fusion

    model/method

    Rather than regressing 3D joint locations relative only to the root pelvis joint or only to immediate kinematic parent joints, the 3DPoseNet architecture predicts joint positions in three complementary kinematic representations and fuses them:

    1. Root-relative pose (PP): 3D joint positions relative to the pelvis root (joint #15).
    2. Order-1 parent-relative pose (O1O_1): 3D joint positions relative to each joint's direct parent in the kinematic tree hierarchy.
    3. Order-2 parent-relative pose (O2O_2): 3D joint positions relative to each joint's grandparent (Order-2 parent) in the kinematic tree hierarchy.

    All three modes ensure at least one reference joint lies in the relatively low-entropy torso. Attached after layer res5a of a ResNet-101 backbone, three separate prediction stubs—each consisting of a convolution layer (5×55 \times 5 kernel, stride 2, 128 feature channels) followed by a fully connected layer—separately output PP, O1O_1, and O2O_2. The distinct convolutional branches provide decorrelation among the predictions.

    These three output vectors are concatenated and fed into a fusion multi-layer perceptron comprising three fully connected layers (with 2000, 1000, and 3N3N nodes, where NN is the number of joints, e.g., 51 outputs for 17 joints) to output the final fused root-centered 3D pose PfusedP_{\text{fused}}. This enables the network to dynamically weight the most reliable kinematic constraints based on limb visibility and pose complexity.

  2. Knowl 2 — Multi-Level Corrective Skip Connections for Training Regularization

    model/method

    Multi-level corrective skip connections act as a training-time architectural regularization mechanism that boosts gradient propagation without incurring any computational overhead during inference.

    Auxiliary prediction stubs with 96 convolutional features followed by a fully connected layer are attached to intermediate layers (specifically res3b3 and res4b20 of the ResNet backbone). The intermediate outputs Pskip,lP_{\text{skip}, l} are summed with the deep output PdeepP_{\text{deep}} produced by the main head at res5a:

    Psum=Pdeep+∑lPskip,lP_{\text{sum}} = P_{\text{deep}} + \sum_{l} P_{\text{skip}, l}

    During training, ground truth supervision is applied simultaneously to both the composite prediction PsumP_{\text{sum}} and the deep branch prediction PdeepP_{\text{deep}} using Euclidean losses:

    L=wsum∥Psum−PGT∥2+wdeep∥Pdeep−PGT∥2\mathcal{L} = w_{\text{sum}} \|P_{\text{sum}} - P_{\text{GT}}\|^2 + w_{\text{deep}} \|P_{\text{deep}} - P_{\text{GT}}\|^2

    At test time, the skip branches are discarded completely, and predictions are generated solely from PdeepP_{\text{deep}} (or downstream fusion layers attached to PdeepP_{\text{deep}}). This forces the core network layers to learn representations that are accurate independently, outperforming conventional identity skip connections.

  3. Knowl 3 — Feature Transfer from 2D Pose Estimation via Layer-Wise Learning Rate Discrepancy

    model/method

    To transfer generalized mid- and high-level representations learned on large-scale in-the-wild 2D pose datasets (e.g., MPII and LSP) to monocular 3D pose regression without catastrophic forgetting:

    • The 3D pose estimation network (3DPoseNet) backbone is initialized with the weights of a fully convolutional 2D pose estimation network (2DPoseNet, ResNet-101 base).
    • An intentional learning rate discrepancy is established: the learning rate for all backbone layers up to res4b22 is scaled down by a factor of 1/10001/1000 (0.001×0.001\times the base learning rate).
    • The higher-level layer res5a and the randomly initialized 3D prediction stubs are trained at the full base learning rate (multiplier 1.01.0).

    This discrepancy prevents the uncorrupted 2D visual features in lower and middle layers from being overwritten by 3D regression gradients, outperforming both standard end-to-end fine-tuning and unsupervised gradient-reversal domain adaptation.

  4. Knowl 4 — Closed-Form Monocular 3D Localization and Perspective Correction

    algorithm

    Given a cropped bounding box RGB image, camera focal length ff, predicted 2D keypoints K∈R2×NK \in \mathbb{R}^{2 \times N}, and predicted root-centered 3D pose P∈R3×NP \in \mathbb{R}^{3 \times N}, this algorithm computes the global 3D pose P[G]∈R3×NP^{[G]} \in \mathbb{R}^{3 \times N} in closed form without iterative non-linear optimization.

    Input: Camera focal length ff, 2D keypoint locations K∈R2×NK \in \mathbb{R}^{2 \times N}, root-centered 3D pose P∈R3×NP \in \mathbb{R}^{3 \times N}
    Output: Global 3D pose P[G]∈R3×NP^{[G]} \in \mathbb{R}^{3 \times N}
    // 1. Perspective Correction for Virtual Camera Rotation
    Compute 2D centroid Kˉ=1N∑i=1NKi\bar{K} = \frac{1}{N} \sum_{i=1}^N K^i
    Compute rotation matrix RR around camera vertical axis aligning the virtual view vector to the original optical axis
    Prot←R⋅PP_{\text{rot}} \leftarrow R \cdot P
    // 2. Closed-Form Translation via Projective Alignment
    Compute 3D (x,y)(x,y) centroid Pˉ[xy]=1N∑i=1NProt,[xy]i\bar{P}_{[xy]} = \frac{1}{N} \sum_{i=1}^N P_{\text{rot}, [xy]}^i
    z←f⋅∑i=1N∥Prot,[xy]i−Pˉ[xy]∥2∑i=1N∥Ki−Kˉ∥2z \leftarrow f \cdot \frac{\sqrt{\sum_{i=1}^N \|P_{\text{rot}, [xy]}^i - \bar{P}_{[xy]}\|^2}}{\sqrt{\sum_{i=1}^N \|K^i - \bar{K}\|^2}}
    x←Kˉ[x]⋅zf−Pˉ[x]x \leftarrow \bar{K}_{[x]} \cdot \frac{z}{f} - \bar{P}_{[x]}
    y←Kˉ[y]⋅zf−Pˉ[y]y \leftarrow \bar{K}_{[y]} \cdot \frac{z}{f} - \bar{P}_{[y]}
    T←(x,y,z)⊤T \leftarrow (x, y, z)^\top
    // 3. Assemble Global Reconstruction
    for i=1i = 1 to NN do
        P[G],i←Proti+TP^{[G], i} \leftarrow P_{\text{rot}}^i + T
    end for
    return P[G]P^{[G]}
  5. Knowl 5 — MPI-INF-3DHP Dataset and Multi-Region Texture Augmentation

    definition

    The MPI-INF-3DHP dataset is a 3D human pose dataset recorded in a multi-camera green-screen studio with ground truth obtained from a marker-less motion capture system (14 synchronized cameras). It contains recordings of 8 actors (4 male, 4 female) performing 8 activity sets each (~1 minute per activity) covering walking, sitting, exercise, dynamic sports, floor poses, and dancing. Over 1.3 million frames are captured, including >500k>500\text{k} frames from 5 chest-height cameras (15∘15^\circ elevation variation), alongside top-down, high angled (45∘45^\circ), and low cameras.

    Each subject performs activities in two apparel configurations: casual everyday clothing and plain-colored clothing. For plain-colored sequences, automatic chroma-key segmentation provides masks for the background, chair/sofa furniture, upper body, and lower body. Appearance augmentation uses a simplified intrinsic image decomposition: because plain clothing intensity variation is dominated by shading, average pixel intensity serves as a proxy for the shading component, allowing independent compositing of random cloth textures onto upper body, lower body, furniture, and sampled internet backgrounds to synthesize photorealistic in-the-wild variation.

  6. Knowl 6 — 3D Percentage of Correct Keypoints (3DPCK) and Area Under Curve (AUC) Metrics

    definition

    To evaluate 3D human pose estimation with greater robustness to outlier predictions than Mean Per Joint Position Error (MPJPE):

    • 3DPCK (3D Percentage of Correct Keypoints): The percentage of predicted joint positions whose Euclidean distance to the ground truth joint position (after root centering) is below a fixed threshold of 150 mm150\text{ mm} (approximately half the diameter of a human head, matching the standard PCKh metric in 2D pose estimation):

    3DPCK=1N∑i=1NI(∥Pi−PGTi∥≤150 mm)\text{3DPCK} = \frac{1}{N} \sum_{i=1}^N \mathbb{I}\left(\|P^i - P_{\text{GT}}^i\| \le 150\text{ mm}\right)

    • AUC (Area Under Curve): The normalized area under the 3DPCK curve evaluated over a range of threshold values from 0 mm0\text{ mm} to 150 mm150\text{ mm}.

    Evaluations are standardized on the 14 common articulated skeleton joints (head top, neck, shoulders, elbows, wrists, hips, knees, and ankles) grouped bilaterally.

  7. Knowl 7 — Ablation Study of CNN Supervision and Multi-Modal Fusion on Human3.6m

    data/table

    The table below measures the effect of adding architectural supervision components, 2D-to-3D feature transfer, and additional training data to a ResNet-101 baseline. The evaluation is conducted on Human3.6m (Subjects S9 and S11, every 64th frame, 17 joints, without test skeleton scaling) in terms of MPJPE in mm across all 15 activity classes.

    Method Direct Discuss Eating Greet Phone Posing Purch. Sitting Sit Down Smoke Take Photo Wait Walk Walk Dog Walk Pair Total
    Base + Regular Skip 113.34 112.26 97.40 110.50 108.63 112.09 105.67 125.97 173.41 109.34 120.87 107.75 97.30 126.05 117.45 115.29
    Base 98.98 100.14 86.07 101.83 101.34 96.74 94.89 125.28 158.31 100.21 112.49 99.57 83.39 109.61 95.79 104.32
    + Corr. Skip 92.57 99.08 85.46 95.43 96.93 89.56 95.67 123.54 160.98 97.13 107.56 93.86 76.99 110.93 88.73 101.09
    + Fusion 93.80 99.17 84.73 95.60 94.48 89.40 93.15 119.94 154.61 95.94 106.09 94.13 77.25 108.82 87.38 99.79
    + Transfer 2DPoseNet 59.69 69.74 60.55 68.77 76.36 59.05 75.04 96.19 122.92 70.82 85.42 68.45 54.41 82.03 59.79 74.14
    + MPI-INF-3DHP 57.51 68.58 59.56 67.34 78.06 56.86 69.13 99.98 117.53 69.44 82.40 67.96 55.24 76.50 61.40 72.88

    The results show that standard skip connections degrade performance (115.29 mm115.29\text{ mm} vs 104.32 mm104.32\text{ mm}), whereas corrective skip connections improve it to 101.09 mm101.09\text{ mm}. Multi-modal kinematic fusion reduces total error to 99.79 mm99.79\text{ mm}. Incorporating 2D feature transfer from 2DPoseNet yields a large reduction to 74.14 mm74.14\text{ mm}, which is further improved to 72.88 mm72.88\text{ mm} when co-training with the augmented MPI-INF-3DHP dataset.

  8. Knowl 8 — Performance Benchmark on MPI-INF-3DHP Test Set

    data/table

    The table below evaluates the generalizability of models across three distinct scene settings in the MPI-INF-3DHP test set: Studio with Green Screen (Studio GS), Studio without Green Screen (Studio no GS), and unconstrained Outdoors. Metrics reported are 3DPCK (%) and AUC with weights transferred from 2DPoseNet.

    3D Training Dataset Method Studio GS (3DPCK) Studio no GS (3DPCK) Outdoor (3DPCK) All (3DPCK) All (AUC)
    Human3.6m Domain adaptation 44.1 42.6 35.2 41.4 17.7
    Human3.6m Ours (full model) 70.8 62.3 58.5 64.7 31.7
    MPI-INF-3DHP Aug. Ours (full model) 82.6 66.7 62.0 71.7 36.4
    MPI-INF-3DHP Unaug. Ours (full model) 84.1 68.9 59.6 72.5 36.9
    MPI-INF-3DHP Aug. + H3.6m Ours, w/o persp. corr. 81.9 68.6 67.4 73.5 37.6
    MPI-INF-3DHP Aug. + H3.6m Ours, w/o GT BB 80.4 71.2 69.8 74.4 39.6
    MPI-INF-3DHP Aug. + H3.6m Ours (full model) 84.6 72.4 69.7 76.5 40.8

    The data shows that our feature transfer approach (64.7%64.7\% 3DPCK on Human3.6m data) substantially outperforms domain adaptation (41.4%41.4\% 3DPCK). Combining Human3.6m with MPI-INF-3DHP augmented data yields the highest overall accuracy (76.5%76.5\% 3DPCK, 40.840.8 AUC). Omitting perspective correction drops performance from 76.5%76.5\% to 73.5%73.5\%, and using 2D-detected bounding boxes instead of ground truth boxes achieves 74.4%74.4\% 3DPCK.

  9. Knowl 9 — Benchmark Comparison with State of the Art on Human3.6m

    data/table

    The table below compares the 3DPoseNet model against state-of-the-art monocular 3D pose estimation methods on Human3.6m (trained on Subjects 1, 5, 6, 7, 8; evaluated on Subjects 9, 11). Error is reported as total MPJPE in mm.

    Method Total MPJPE (mm)
    Deep Kinematic Pose (Zhou et al., 2016) 107.26
    Sparse Deep (Zhou et al., 2015) 113.01
    Motion Comp. Seq. (Tekin et al., 2016) 124.97
    LinKDE (Ionescu et al., 2014) 162.14
    Du et al. (2016) 126.47
    Rogez et al. (2016) 121.20
    SMPLify (Bogo et al., 2016) 82.3
    3D=2D+Matching (Chen Ramanan, 2017) 114.18
    Distance Matrix (Moreno-Noguer, 2017) 87.30
    Volumetric Coarse-Fine (Pavlakos et al., 2017) 71.90
    LCR-Net (Rogez et al., 2017) 87.7
    Full model (w/o MPI-INF-3DHP) [17 joints, unscaled] 74.11
    Full model (w/o MPI-INF-3DHP) [17 joints, subject-scaled] 68.61
    Full model (w/o MPI-INF-3DHP) [14 joints, Procrustes aligned] 54.59
    Full model (+ MPI-INF-3DHP) [17 joints, unscaled] 72.88

    The full model trained on Human3.6m achieves 74.11 mm74.11\text{ mm} MPJPE without any test-time scaling or Procrustes alignment, outperforming prior monocular methods. When rescaled to subject bone lengths, error drops to 68.61 mm68.61\text{ mm}, and under Procrustes alignment on 14 joints it achieves 54.59 mm54.59\text{ mm}.

  10. Knowl 10 — Learning Rate Discrepancy Multiplier Ablation for 2D-to-3D Transfer Learning

    data/table

    The table below evaluates different learning rate multiplier combinations when transferring weights from 2DPoseNet to 3DPoseNet on the Human3.6m dataset (Subjects 1, 5, 6, 7, 8 for training; every 64th frame of Subjects 9, 11 for testing). Base network without corrective skip connections is used.

    Learning Rate Multiplier
    up to res4b22 res5a 3D Stub S Total MPJPE (mm)
    1 1 1* 118.7
    1/10 1/10 1* 84.6
    1/1000 1/1000 1* 89.2
    1/10 1 1* 90.7
    1/1000 1 1* 80.7

    Note: * indicates randomly initialized weights.

    Uniformly training all layers with a multiplier of 11 leads to catastrophic forgetting (118.7 mm118.7\text{ mm} MPJPE). The optimal configuration scales down layers up to res4b22 by 1/10001/1000 while allowing res5a and the randomly initialized 3D prediction stub to train with full learning rate multiplier 11, achieving 80.7 mm80.7\text{ mm} MPJPE.

  11. Knowl 11 — Pose-Specific Error Reduction via Kinematic Multi-Modal Fusion

    empirical result

    Clustering the Human3.6m test set (Subjects S9 and S11) into three broad pose categories reveals that multi-modal kinematic fusion selectively benefits complex and non-upright body poses:

    • Stand/Walk (constituting 67%67\% of the test set): Fusion has a negligible effect, changing MPJPE from 88.4 mm88.4\text{ mm} to 88.8 mm88.8\text{ mm}.
    • Sit (constituting 25%25\% of the test set): Fusion yields a 3.5 mm3.5\text{ mm} improvement, decreasing MPJPE from 118.9 mm118.9\text{ mm} to 115.4 mm115.4\text{ mm}.
    • Crouch (constituting 8%8\% of the test set): Fusion yields a 5.5 mm5.5\text{ mm} improvement, decreasing MPJPE from 156.0 mm156.0\text{ mm} to 150.5 mm150.5\text{ mm}.

    Because upright poses dominate standard motion capture benchmarks, overall aggregate error metrics dilute the performance gains of kinematic fusion, which specifically resolves limb ambiguities in sitting and crouching configurations.

  12. Knowl 12 — Limitations in Extreme Viewpoint Elevation and Temporal Consistency

    limitation

    The monocular 3D human pose regression framework has two primary limitations:

    1. Viewpoint elevation bias: Because standard 3D pose training sets are captured predominantly at chest height (with minimal camera pitch variation of roughly ±15∘\pm 15^\circ), the network exhibits reduced accuracy when deployed on steep top-down or bottom-up camera angles.
    2. Temporal jitter in monocular video: Operating on single unconstrained frames in a feedforward manner without multi-frame temporal conditioning leads to high-frequency jitter across consecutive video frames, necessitating subsequent model-based temporal tracking or post-filtering for temporally smooth animations.

Coverage note — None was omitted.

References

  1. 1.A. Agarwal and B. Triggs. Recovering 3d human pose from monocular images. IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), 28(1):44–58, 2006. 2
  2. 2.I. Akhter and M. J. Black. Pose-conditioned joint angle limits for 3d human pose reconstruction. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1446–1455, 2015. 2, 11
  3. 3.S. Amin, M. Andriluka, M. Rohrbach, and B. Schiele. Multiview pictorial structures for 3D human pose estimation. In BMVC, 2013. 11
  4. 4.M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele. 2D Human Pose Estimation: New Benchmark and State of the Art Analysis. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2014. 1, 3, 11, 12
  5. 5.A. Baak, M. Müller, G. Bharaj, H.-P. Seidel, and C. Theobalt. A data-driven approach for real-time full body pose reconstruction from a depth camera. In IEEE International Conference on Computer Vision (ICCV), 2011. 1
  6. 6.A. O. Balan, L. Sigal, M. J. Black, J. E. Davis, and H. W. Haussecker. Detailed human shape and pose from images. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1–8, 2007. 1
  7. 7.V. Belagiannis and A. Zisserman. Recurrent human pose estimation. arXiv preprint arXiv:1605.02914, 2016. 1, 11
  8. 8.L. Bo and C. Sminchisescu. Twin gaussian processes for structured prediction. In International Journal of Computer Vision, 2010. 10, 11
  9. 9.F. Bogo, A. Kanazawa, C. Lassner, P. Gehler, J. Romero, and M. J. Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In European Conference on Computer Vision (ECCV), 2016. 1, 2, 7, 10, 11
  10. 10.E. Brau and H. Jiang. 3D Human Pose Estimation via Deep Learning from 2D Annotations. In International Conference on 3D Vision (3DV), 2016. 3
  11. 11.A. Bulat and G. Tzimiropoulos. Human pose estimation via convolutional part heatmap regression. In European Conference on Computer Vision (ECCV), 2016. 1, 11
  12. 12.J. Carreira, P. Agrawal, K. Fragkiadaki, and J. Malik. Human pose estimation with iterative error feedback. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 1, 11
  13. 13.J. Chai and J. K. Hodgins. Performance animation from low-dimensional control signals. ACM Transactions on Graphics (TOG), 24(3):686–696, 2005. 1
  14. 14.C.-H. Chen and D. Ramanan. 3d human pose estimation = 2d pose estimation + matching. In CVPR 2017-IEEE Conference on Computer Vision & Pattern Recognition, 2017. 1, 3, 7
  15. 15.W. Chen, H. Wang, Y. Li, H. Su, Z. Wang, C. Tu, D. Lischinski, D. Cohen-Or, and B. Chen. Synthesizing training images for boosting human 3d pose estimation. In International Conference on 3D Vision (3DV), 2016. 1, 3, 5, 6, 8
  16. 16.X. Chen and A. L. Yuille. Articulated pose estimation by a graphical model with image dependent pairwise relations. In Advances in Neural Information Processing Systems (NIPS), pages 1736–1744, 2014. 1
  17. 17.A. Elhayek, E. Aguiar, A. Jain, J. Tompson, L. Pishchulin, M. Andriluka, C. Bregler, B. Schiele, and C. Theobalt. MARCOnI - ConvNet-based MARker-less Motion Capture in Outdoor and Indoor Scenes. IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), 2016. 2, 6
  18. 18.A. Elhayek, E. de Aguiar, A. Jain, J. Tompson, L. Pishchulin, M. Andriluka, C. Bregler, B. Schiele, and C. Theobalt. Efficient ConvNet-based marker-less motion capture in general scenes with a low number of cameras. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3810–3818, 2015. 11
  19. 19.P. F. Felzenszwalb and D. P. Huttenlocher. Pictorial structures for object recognition. International Journal of Computer Vision (IJCV), 61(1):55–79, 2005. 2
  20. 20.J. Gall, B. Rosenhahn, T. Brox, and H.-P. Seidel. Optimization and filtering for human motion capture. International Journal of Computer Vision (IJCV), 87(1–2):75–92, 2010. 1
  21. 21.Y. Ganin and V. Lempitsky. Unsupervised domain adaptation by backpropagation. In Proceedings of the 32nd International Conference on Machine Learning (ICML-15), pages 1180–1189, 2015. 8, 12
  22. 22.G. Gkioxari, A. Toshev, and N. Jaitly. Chained predictions using convolutional neural networks. In European Conference on Computer Vision (ECCV), 2016. 1, 11
  23. 23.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 3
  24. 24.P. Hu, D. Ramanan, J. Jia, S. Wu, X. Wang, L. Cai, and J. Tang. Bottom-up and top-down reasoning with hierarchical rectified gaussians. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 1, 11
  25. 25.E. Insafutdinov, L. Pishchulin, B. Andres, M. Andriluka, and B. Schiele. Deepercut: A deeper, stronger, and faster multiperson pose estimation model. In European Conference on Computer Vision (ECCV), 2016. 3, 6, 11
  26. 26.C. Ionescu, J. Carreira, and C. Sminchisescu. Iterated second-order label sensitive pooling for 3d human pose estimation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1661–1668, 2014. 2, 12
  27. 27.C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu. Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), 36(7):1325–1339, 2014. 1, 5, 6, 7, 8
  28. 28.A. Jain, J. Tompson, Y. LeCun, and C. Bregler. Modeep: A deep learning framework using motion features for human pose estimation. In Asian Conference on Computer Vision (ACCV), pages 302–315. Springer, 2014. 2
  29. 29.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. In Proceedings of the 22nd ACM International Conference on Multimedia, pages 675–678, 2014. 11
  30. 30.S. Johnson and M. Everingham. Clustered pose and nonlinear appearance models for human pose estimation. In British Machine Vision Conference (BMVC), 2010. doi:10.5244/C.24.12. 1, 3, 6, 11, 12
  31. 31.S. Johnson and M. Everingham. Learning effective human pose estimation from inaccurate annotation. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2011. 1, 3, 12
  32. 32.H. Joo, H. Liu, L. Tan, L. Gui, B. Nabbe, I. Matthews, T. Kanade, S. Nobuhara, and Y. Sheikh. Panoptic studio: A massively multiview system for social motion capture. In ICCV, pages 3334–3342, 2015. 1, 5, 6
  33. 33.A. M. Lehrmann, P. V. Gehler, and S. Nowozin. A Nonparametric Bayesian Network Prior of Human Pose. In IEEE International Conference on Computer Vision (ICCV), 2013. 4
  34. 34.V. Lepetit and P. Fua. Monocular model-based 3D tracking of rigid objects. Now Publishers Inc, 2005. 4
  35. 35.S. Li and A. B. Chan. 3d human pose estimation from monocular images with deep convolutional neural network. In Asian Conference on Computer Vision (ACCV), pages 332–347, 2014. 1, 2, 4
  36. 36.S. Li, W. Zhang, and A. B. Chan. Maximum-margin structured learning with deep networks for 3d human pose estimation. In IEEE International Conference on Computer Vision (ICCV), pages 2848–2856, 2015. 1, 2
  37. 37.I. Lifshitz, E. Fetaya, and S. Ullman. Human pose estimation using deep consensus voting. In European Conference on Computer Vision (ECCV), 2016. 1, 11
  38. 38.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015. 3
  39. 39.A. Meka, M. Zollhöfer, C. Richardt, and C. Theobalt. Live intrinsic video. ACM Trans. Graph. (Proc. SIGGRAPH), 35(4):109:1–14, 2016. 5
  40. 40.F. Moreno-Noguer. 3d human pose estimation from a single image via distance matrix regression. In CVPR 2017-IEEE Conference on Computer Vision & Pattern Recognition, 2017. 1, 7
  41. 41.G. Mori and J. Malik. Recovering 3d human body configurations using shape contexts. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 28(7):1052–1062, 2006. 2
  42. 42.A. Newell, K. Yang, and J. Deng. Stacked hourglass networks for human pose estimation. In European Conference on Computer Vision (ECCV), 2016. 1, 2, 11
  43. 43.S. J. Pan and Q. Yang. A survey on transfer learning. IEEE Transactions on knowledge and data engineering, 22(10):1345–1359, 2010. 3
  44. 44.G. Pavlakos, X. Zhou, K. G. Derpanis, and K. Daniilidis. Coarse-to-fine volumetric prediction for single-image 3D human pose. In CVPR 2017-IEEE Conference on Computer Vision & Pattern Recognition, 2017. 1, 2, 7, 8
  45. 45.L. Pishchulin, E. Insafutdinov, S. Tang, B. Andres, M. Andriluka, P. Gehler, and B. Schiele. Deepcut: Joint subset partition and labeling for multi person pose estimation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 1, 2, 11
  46. 46.L. Pishchulin, A. Jain, M. Andriluka, T. Thormählen, and B. Schiele. Articulated people detection and pose estimation: Reshaping the future. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 3178–3185. IEEE, 2012. 5
  47. 47.G. Pons-Moll, D. J. Fleet, and B. Rosenhahn. Posebits for monocular human pose estimation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2337–2344, 2014. 2
  48. 48.V. Ramakrishna, T. Kanade, and Y. Sheikh. Reconstructing 3d human pose from 2d image landmarks. In European Conference on Computer Vision, pages 573–586. Springer, 2012. 11
  49. 49.H. Rhodin, C. Richardt, D. Casas, E. Insafutdinov, M. Shafiei, H.-P. Seidel, B. Schiele, and C. Theobalt. EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras. ACM Trans. Graph. (Proc. SIGGRAPH Asia), 2016. 5
  50. 50.H. Rhodin, N. Robertini, D. Casas, C. Richardt, H.-P. Seidel, and C. Theobalt. General automatic human shape and motion capture using volumetric contour cues. In European Conference on Computer Vision (ECCV), pages 509–526. Springer, 2016. 2, 11
  51. 51.N. Robertini, D. Casas, H. Rhodin, H.-P. Seidel, and C. Theobalt. Model-based Outdoor Performance Capture. In International Conference on Computer Vision (3DV), 2016. 6
  52. 52.G. Rogez and C. Schmid. Mocap-guided data augmentation for 3d pose estimation in the wild. In Advances in Neural Information Processing Systems, pages 3108–3116, 2016. 3, 7, 8
  53. 53.G. Rogez, P. Weinzaepfel, and C. Schmid. Lcr-net: Localization-classification-regression for human pose. In CVPR 2017-IEEE Conference on Computer Vision & Pattern Recognition, 2017. 3, 7
  54. 54.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015. 3, 4
  55. 55.B. Sapp and B. Taskar. Modec: Multimodal decomposable models for human pose estimation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2013. 1
  56. 56.A. Sharif Razavian, H. Azizpour, J. Sullivan, and S. Carlsson. CNN features off-the-shelf: an astounding baseline for recognition. In Conference on Computer Vision and Pattern Recognition Workshops, pages 806–813, 2014. 3
  57. 57.J. Shotton, T. Sharp, A. Kipman, A. Fitzgibbon, M. Finocchio, A. Blake, M. Cook, and R. Moore. Real-time human pose recognition in parts from single depth images. Communications of the ACM, 56(1):116–124, 2013. 1
  58. 58.L. Sigal, A. O. Balan, and M. J. Black. Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. International Journal of Computer Vision (IJCV), 87(1-2):4–27, 2010. 1, 6, 8, 11
  59. 59.E. Simo-Serra, A. Quattoni, C. Torras, and F. Moreno-Noguer. A joint model for 2d and 3d pose estimation from a single image. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 3634–3641, 2013. 1, 2
  60. 60.E. Simo-Serra, A. Ramisa, G. Alenyà, C. Torras, and F. Moreno-Noguer. Single image 3d human pose estimation from noisy observations. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2673–2680. IEEE, 2012. 1, 2
  61. 61.C. Sminchisescu and B. Triggs. Covariance scaled sampling for monocular 3d body tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), volume 1, pages I–447. IEEE, 2001. 1
  62. 62.J. Starck and A. Hilton. Model-based multiple view reconstruction of people. In IEEE International Conference on Computer Vision (ICCV), pages 915–922, 2003. 1
  63. 63.C. Stoll, N. Hasler, J. Gall, H.-P. Seidel, and C. Theobalt. Fast articulated motion tracking using a sums of Gaussians body model. In IEEE International Conference on Computer Vision (ICCV), pages 951–958, 2011. 1
  64. 64.C. J. Taylor. Reconstruction of articulated objects from point correspondences in a single uncalibrated image. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), volume 1, pages 677–684, 2000. 2
  65. 65.B. Tekin, I. Katircioglu, M. Salzmann, V. Lepetit, and P. Fua. Structured Prediction of 3D Human Pose with Deep Neural Networks. In British Machine Vision Conference (BMVC), 2016. 1, 2
  66. 66.B. Tekin, P. Márquez-Neila, M. Salzmann, and P. Fua. Fusing 2D Uncertainty and 3D Cues for Monocular Body Pose Estimation. arXiv preprint arXiv:1611.05708, 2016. 2, 3
  67. 67.B. Tekin, A. Rozantsev, V. Lepetit, and P. Fua. Direct Prediction of 3D Body Poses from Motion Compensated Sequences. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 1, 2, 7
  68. 68.The Captury. http://www.thecaptury.com/, 2016. 5
  69. 69.D. Tome, C. Russell, and L. Agapito. Lifting from the deep: Convolutional 3d pose estimation from a single image. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 1
  70. 70.J. J. Tompson, A. Jain, Y. LeCun, and C. Bregler. Joint training of a convolutional network and a graphical model for human pose estimation. In Advances in Neural Information Processing Systems (NIPS), pages 1799–1807, 2014. 1, 6
  71. 71.A. Toshev and C. Szegedy. Deeppose: Human pose estimation via deep neural networks. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 1653–1660, 2014. 1, 6
  72. 72.R. Urtasun, D. J. Fleet, and P. Fua. Monocular 3d tracking of the golf swing. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 932–938, 2005. 1
  73. 73.C. Wang, Y. Wang, Z. Lin, A. L. Yuille, and W. Gao. Robust estimation of 3d human poses from a single image. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2361–2368, 2014. 1, 2
  74. 74.S.-E. Wei, V. Ramakrishna, T. Kanade, and Y. Sheikh. Convolutional Pose Machines. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 1, 2, 11
  75. 75.C. R. Wren, A. Azarbayejani, T. Darrell, and A. P. Pentland. Pfinder: real-time tracking of the human body. IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), 19(7):780–785, 1997. 1
  76. 76.H. Yasin, U. Iqbal, B. Krüger, A. Weber, and J. Gall. A Dual-Source Approach for 3D Pose Estimation from a Single Image. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 1, 3, 10, 11
  77. 77.J. Yosinski, J. Clune, Y. Bengio, and H. Lipson. How transferable are features in deep neural networks? In Advances in Neural Information Processing Systems (NIPS), pages 3320–3328, 2014. 3
  78. 78.Y. Yu, F. Yonghao, Z. Yilin, and W. Mohan. Marker-less 3D Human Motion Capture with Monocular Image Sequence and Height-Maps. In European Conference on Computer Vision (ECCV), 2016. 1, 2, 7
  79. 79.F. Zhou and F. De la Torre. Spatio-temporal matching for human detection in video. In European Conference on Computer Vision (ECCV), pages 62–77, 2014. 1
  80. 80.X. Zhou, S. Leonardos, X. Hu, and K. Daniilidis. 3D shape estimation from 2D landmarks: A convex relaxation approach. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4447–4455, 2015. 1, 2, 4
  81. 81.X. Zhou, X. Sun, W. Zhang, S. Liang, and Y. Wei. Deep kinematic pose regression. In ECCV Workshop on Geometry Meets Deep Learning, 2016. 1, 2, 7, 8
  82. 82.X. Zhou, M. Zhu, S. Leonardos, and K. Daniilidis. Sparse representation for 3d shape estimation: A convex relaxation approach. arXiv preprint arXiv:1509.04309, 2015. 2
  83. 83.X. Zhou, M. Zhu, S. Leonardos, K. Derpanis, and K. Daniilidis. Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015. 1, 2, 7, 10, 11

Citation

MLA
Mehta, D., et al. “Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision”. arXiv, 2016, http://arxiv.org/abs/1611.09813v5.
APA
Mehta, D., Rhodin, H., Casas, D., Fua, P., Sotnychenko, O., Xu, W., & Theobalt, C. (2016). Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision. arXiv. http://arxiv.org/abs/1611.09813v5
Chicago
Mehta, D., H. Rhodin, D. Casas, et al. 2016. “Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision”. arXiv. http://arxiv.org/abs/1611.09813v5.
Harvard
Mehta, D. et al. (2016) “Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1611.09813v5.
Vancouver
1. Mehta D, Rhodin H, Casas D, Fua P, Sotnychenko O, Xu W, Theobalt C (2016) Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision. arXiv

BibTeX

@article{mehta2016monocular,
  title = {Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision},
  author = {Mehta, Dushyant and Rhodin, Helge and Casas, Dan and Fua, Pascal and Sotnychenko, Oleksandr and Xu, Weipeng and Theobalt, Christian},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1611.09813v5},
  eprint = {1611.09813}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF