Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation

Jonathan TompsonArjun JainYann LeCunChristoph Bregler

article2014NeurIPS1,619 citations

Introduces an end-to-end trainable hybrid architecture combining deep convolutional networks with Markov random fields to explicitly model geometric spatial constraints between body joints for improved 2D human pose estimation.

Listen

Accurately identifying human body joint positions in single standard images is a major challenge in computer vision. It is critical for applications in surveillance, healthcare, motion capture, and human-computer interaction. The task is complicated by variations in clothing, lighting, body shapes, viewing angles, and frequent limb occlusions. Existing techniques generally rely either on hand-crafted part models, which struggle with visual diversity, or deep learning systems, which often output anatomically implausible poses because they lack structural understanding of human anatomy.

The main objective of the article is to design and evaluate a hybrid computer vision framework that combines a multi-resolution deep convolutional network for part detection with a spatial graphical model that enforces anatomical body constraints. The authors evaluate whether training these two systems together within a single framework improves joint localization accuracy over existing state-of-the-art methods.

The evaluated approach uses a multi-resolution feature detector that processes images across multiple scales to capture both fine details and broader context. The network outputs per-pixel probability heatmaps for joint locations, which are then passed into a spatial model structured as a fully connected graph. This spatial component learns the geometric relationships between body parts and performs message-passing inference. Both components are initially trained independently and then fine-tuned together end-to-end. The system was validated against two standard computer vision benchmarks: the FLIC dataset (comprising movie frames) and the extended-LSP dataset (featuring complex athletic poses).

The article demonstrates several key findings. First, the unified model significantly outperforms prior state-of-the-art methods across all evaluated benchmarks, achieving notable gains in wrist, elbow, knee, and ankle localization. Second, incorporating the spatial model increases joint detection rates by 8 to 12 percentage points for large error thresholds by eliminating anatomically impossible false positives. Third, jointly fine-tuning the detector and spatial model yields an additional 4 to 5 percentage point performance boost. Fourth, expanding the visual context using multiple resolution processing banks substantially increases part detection accuracy over single-resolution baselines. Finally, the combined pipeline operates with an inference speed of approximately 51 milliseconds per image on standard hardware, supporting practical, real-time application.

These findings show that pairing appearance-based deep learning with structural anatomical constraints resolves a key failure mode in computer vision. Instead of forcing a neural network to memorize complex geometric rules, the hybrid design lets the visual detector focus on localized features while the spatial model enforces plausible skeletal arrangements. The resulting efficiency makes high-accuracy human pose estimation viable for real-time commercial and operational deployments without specialized, high-cost computing hardware.

Organizations developing motion analysis or computer vision systems should adopt hybrid architectures that combine visual feature detectors with learnable structural models. Future implementations should focus on increasing the capacity of the spatial model to handle highly articulated and dynamic poses, as well as curating larger, high-diversity training datasets to improve edge-case performance.

The reported results carry high confidence on moderately posed subjects, as in standard upright or front-facing scenes. However, stakeholders should note a performance drop on extreme, highly acrobatic poses—such as those found in sports datasets—where learned spatial relationships are less rigid. Expanding training data and incorporating more expressive spatial priors remain necessary before deploying the system in unconstrained, high-variance environments.

arXiv: 1406.2984
Cover for Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation

Abstract

This paper proposes a new hybrid architecture that consists of a deep Convolutional Network and a Markov Random Field. We show how this architecture is successfully applied to the challenging problem of articulated human pose estimation in monocular images. The architecture can exploit structural domain constraints such as geometric relationships between body joint locations. We show that joint training of these two model paradigms improves performance and allows us to significantly outperform existing state-of-the-art techniques.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Model
  • 3.1 Convolutional Network Part-Detector
  • 3.2 Higher-Level Spatial-Model
  • 3.3 Unified Model
  • 4 Results
  • 5 Conclusion
  • 6 Acknowledgments
  • References

Knowls

  1. Knowl 1 — Differentiable MRF-Inspired Spatial-Model for Joint Consistency

    model/method

    To eliminate anatomically incorrect false-positive predictions from local part detectors, a Markov Random Field (MRF) inspired Spatial-Model is formulated over the 2D spatial positions of all body joints. Let VV denote the set of body joints forming the vertices of a fully connected graph. The local part detector provides unary potentials for each joint location. Pairwise potentials between joint v∈Vv \in V and joint A∈VA \in V are parameterized by learned 2D convolution filters pA∣vp_{A|v}, representing the conditional displacement probability distribution of joint AA given the location of joint vv.

    In standard probability form, the marginal likelihood pˉA\bar{p}_A for joint AA across the image domain is evaluated via a single round of sum-product belief propagation:

    pˉA=1Z∏v∈V(pA∣v∗pv+bv→A)\bar{p}_A = \frac{1}{Z} \prod_{v \in V} \left( p_{A|v} * p_v + b_{v \to A} \right)

    where ∗* denotes 2D spatial convolution, pvp_v is the unary distribution (heat-map) for joint vv, bv→Ab_{v \to A} is a learned scalar background probability bias for the message from vv to AA (preventing false negatives if joint vv is missed), and ZZ is the partition function.

    Rather than hand-crafting tree structures or joint priors, every pair of joints is connected, and the spatial prior kernels pA∣vp_{A|v} and background biases are learned end-to-end via back-propagation. If two joints share no structural relationship, the corresponding pairwise kernel learns a uniform distribution.

  2. Knowl 2 — Energy-Based Message Passing Formulation for Spatial Priors

    equation

    To render the MRF message-passing update numerically stable and fully differentiable for back-propagation with Stochastic Gradient Descent, the probability product formulation is mapped into log-space energy updates without explicitly computing the partition function:

    eˉA=exp⁡(∑v∈Vlog⁡(SoftPlus(eA∣v)∗ReLU(ev)+SoftPlus(bv→A)))\bar{e}_A = \exp \left( \sum_{v \in V} \log \left( \text{SoftPlus}(e_{A|v}) * \text{ReLU}(e_v) + \text{SoftPlus}(b_{v \to A}) \right) \right)

    where:

    • eˉA\bar{e}_A is the final output marginal energy map for joint A∈VA \in V.
    • VV is the complete set of annotated body joints.
    • eve_v is the input unary heat-map energy for joint vv.
    • eA∣ve_{A|v} is the 2D convolutional weight kernel modeling the spatial displacement of joint AA relative to joint vv.
    • bv→Ab_{v \to A} is a scalar bias parameter representing the background energy of the message.
    • ∗* denotes a 2D spatial convolution operator.
    • SoftPlus(x)=1βlog⁡(1+exp⁡(βx))\text{SoftPlus}(x) = \frac{1}{\beta} \log(1 + \exp(\beta x)) with scale parameter 0.5≤β≤20.5 \le \beta \le 2, applied to kernel weights and biases to guarantee strictly non-negative values and non-zero gradients during training.
    • ReLU(x)=max⁡(x,ϵ)\text{ReLU}(x) = \max(x, \epsilon) with threshold 0<ϵ≤0.010 < \epsilon \le 0.01, ensuring strictly positive inputs to the log⁡\log operation to avoid numerical instability.

    The log-space summation decouples convolution output gradients across stages, allowing back-propagation gradients with respect to one pairwise convolution output to be independent of other branches.

  3. Knowl 3 — Multi-Resolution Sliding-Window ConvNet Part-Detector

    model/method

    The Part-Detector is a deep convolutional neural network that produces dense per-pixel 2D heat-maps indicating the likelihood of key body joint locations from an input monocular RGB image. To capture both fine-grained local part appearance and broader contextual information, the architecture incorporates multi-resolution input banks with overlapping receptive fields.

    The input image is processed at multiple spatial resolutions (typically 3 resolution banks) forming an approximate Laplacian pyramid via downsampling and Local Contrast Normalization (LCN), which supplies non-overlapping spectral content to each bank. In a fully convolutional formulation equivalent to a sliding-window detector:

    1. Each resolution bank consists of three stages of 5×55 \times 5 convolutions, Rectified Linear Unit (ReLU) activations, and pooling.
    2. The lower-resolution feature maps are upscaled point-wise to match the spatial dimensions of the higher-resolution feature map.
    3. The multi-resolution feature maps are combined and processed by 9×99 \times 9 and 1×11 \times 1 convolutional layers (equivalent to sliding fully-connected layers) to produce the multi-channel joint heat-maps (e.g., 90×6090 \times 60 pixels for a 320×240320 \times 240 input image).

    The Part-Detector is trained in a supervised manner using batched Stochastic Gradient Descent with Nesterov Momentum minimizing the Mean Squared Error (MSE) against target 2D Gaussian heat-maps centered at ground-truth joint coordinates with a small fixed variance.

  4. Knowl 4 — Unified End-to-End Joint Training Pipeline

    model/method

    The unified human pose estimation pipeline integrates the multi-resolution Part-Detector and the convolutional Spatial-Model into a single end-to-end trainable neural network using a three-phase training procedure:

    1. Independent Part-Detector Training: The multi-resolution ConvNet part detector is trained independently on RGB images using Mean Squared Error (MSE) loss against ground-truth 2D Gaussian heat-maps until convergence, and its heat-map predictions on the training set are saved.
    2. Independent Spatial-Model Training: The convolutional MRF spatial-model is trained on the predicted heat-maps from the first stage, updating the pairwise convolution kernels and message biases.
    3. Unified Joint Fine-Tuning: The trained Part-Detector and Spatial-Model are merged into a single computational graph and fine-tuned end-to-end using back-propagation and Stochastic Gradient Descent.

    Joint fine-tuning improves accuracy because the Spatial-Model constrains the search space of anatomically plausible part combinations, allowing the Part-Detector to dedicate its representational capacity toward precise local joint localization rather than global ambiguity resolution.

  5. Knowl 5 — Empirical Initialization and FFT Implementation for Spatial Convolutions

    model/method

    To model long-range anatomical constraints across an entire body heat-map (e.g., a 90×6090 \times 60 pixel heat-map output), pairwise spatial convolution kernels require large spatial dimensions. A kernel size of 128×128128 \times 128 is used to accommodate joint displacement radii up to 64 pixels, with zero-padding applied to the input heat-maps to prevent boundary pixel loss.

    To manage the computational complexity of these large spatial kernels, 2D convolutions in the Spatial-Model are implemented using GPU-accelerated Fast Fourier Transforms (FFT).

    To improve optimization stability, accelerate convergence, and avoid poor local minima, the pairwise convolution weights eA∣ve_{A|v} are initialized using empirical 2D displacement histograms computed directly from ground-truth joint offset vectors in the training dataset rather than random weight initialization.

  6. Knowl 6 — FLIC-plus Dataset Construction and Scene Deduplication

    experimental setup

    The standard Frames Labeled in Cinema (FLIC) dataset comprises 5,003 labeled frames from Hollywood movies (3,987 training and 1,016 test images). An extended dataset called FLIC-full contains 20,928 training images, but many of its frames share common movie scenes with the 1,016 test images, risking train-test leakage.

    To resolve this, the FLIC-plus dataset was constructed:

    1. Human annotators on Amazon Mechanical Turk generated unique scene identifiers for all images in the FLIC test set and the FLIC-full candidate set.
    2. Candidate training images that shared a scene label with any image in the test set were removed.
    3. The 253 original FLIC training images that overlapped with test scenes were re-included to ensure FLIC-plus strictly forms a clean superset of the standard FLIC training set.

    The resulting FLIC-plus dataset consists of 17,380 training images guaranteed to be scene-independent from the 1,016 FLIC test images.

  7. Knowl 7 — Articulated Human Pose Estimation Performance on FLIC and LSP Benchmarks

    empirical result

    The unified ConvNet and MRF architecture was evaluated on the FLIC (1,016 test images) and Leeds Sports Pose (LSP, 1,000 test images) benchmarks using the Percentage of Correct Parts metric (detection rate within a normalized distance error threshold scaled by torso height):

    • FLIC Dataset: The unified model trained on FLIC and FLIC-plus significantly outperformed previous methods (including DeepPose by Toshev et al., MODEC by Sapp & Taskar, and pictorial structures by Yang & Ramanan). For elbow and wrist localization, the proposed model achieved higher detection rates across all normalized distance thresholds, showing especially large gains in the high-precision radius regime.
    • LSP Dataset: Evaluated under person-centric (non-observer-centric) coordinates, the model outperformed previous methods (Toshev et al., Dantone et al., and Pishchulin et al.) across wrist, elbow, ankle, and knee localizations across distance thresholds from 0 to 20 normalized pixels.
    • Runtime: Forward propagation through both the Part-Detector and Spatial-Model requires 51 ms per image on a 12-CPU workstation equipped with an NVIDIA Titan GPU, supporting near real-time inference.
  8. Knowl 8 — Ablation of Spatial-Model and Joint Unified Training

    empirical result

    Ablation experiments on the FLIC wrist localization task demonstrate the progressive benefit of each structural stage:

    1. Stand-alone Part-Detector: Achieves baseline localization accuracy, but produces anatomically impossible outlier false positives (e.g., face detections placed far from shoulder detections).
    2. Addition of Spatial-Model: Has negligible impact on accuracy at very small distance thresholds (low radii), but increases the detection rate by 8%8\% to 12%12\% at larger distance error thresholds by pruning false-positive outliers.
    3. Unified Joint Training: End-to-end fine-tuning of the combined Part-Detector and Spatial-Model adds a further 4%4\% to 5%5\% improvement in detection rate at large radius thresholds over independently trained cascaded models.
  9. Knowl 9 — Impact of Multi-Resolution Banks on Part Localization

    empirical result

    Evaluating the Part-Detector on the FLIC dataset across varying numbers of input resolution banks demonstrates that multi-resolution context substantially improves localization accuracy:

    • A network using 3 resolution banks consistently outperforms networks using 2 resolution banks or 1 resolution bank across all normalized distance error thresholds (from 0 to 20 pixels) for wrist detection.
    • Increasing the input context allows the network to incorporate broader body structure cues at coarse scales without inflating parameters in the fine-scale convolution stages.
  10. Knowl 10 — Spatial-Model Limitations on Highly Articulated Pose Distributions

    limitation

    The Spatial-Model relies on static 2D convolutional prior kernels pA∣vp_{A|v} to model the displacement distribution between joint pairs. While highly effective for datasets with constrained view distributions and predominantly upright front-facing poses (such as FLIC), the pairwise distributions become diffuse and less constrained on datasets with wide pose variance and extreme articulation (such as athletes in the LSP dataset).

    Consequently, the simple unimodal pairwise convolution prior is less effective at pruning outliers for complex, highly unconstrained poses where relative joint displacements vary extensively.

Coverage note — None was omitted. All primary contributions—including the multi-resolution Part-Detector, the differentiable MRF Spatial-Model and its log-space energy formulation, the unified training framework, dataset deduplication (FLIC-plus), FFT implementation, empirical histogram initialization, ablation studies, benchmark results on FLIC/LSP, and structural limitations—are fully covered.

References

  1. 1.M. Andriluka, S. Roth, and B. Schiele. Pictorial structures revisited: People detection and articulated pose estimation. In CVPR, 2009.
  2. 2.M. Bergtholdt, J. Kappes, S. Schmidt, and C. Schnörr. A study of parts-based object class detection using complete graphs. IJCV, 2010.
  3. 3.L. Bourdev and J. Malik. Poselets: Body part detectors trained using 3d human pose annotations. In ICCV, 2009.
  4. 4.H. Bourlard, Y. Konig, and N. Morgan. Remap: recursive estimation and maximization of a posteriori probabilities in connectionist speech recognition. In EUROSPEECH, 1995.
  5. 5.P. Buehler, A. Zisserman, and M. Everingham. Learning sign language by watching TV (using weakly aligned subtitles). CVPR, 2009.
  6. 6.R. Collobert, K. Kavukcuoglu, and C. Farabet. Torch7: A matlab-like environment for machine learning. In BigLearn, NIPS Workshop, 2011.
  7. 7.N. Dalal and B. Triggs. Histograms of oriented gradients for human detection. In CVPR, 2005.
  8. 8.M. Dantone, J. Gall, C. Leistner, and L. Van Gool. Human pose estimation using body parts dependent joint regressors. In CVPR'13.
  9. 9.M. Eichner and V. Ferrari. Better appearance models for pictorial structures. In BMVC, 2009.
  10. 10.P. Felzenszwalb, D. McAllester, and D. Ramanan. A discriminatively trained, multiscale, deformable part model. In CVPR, 2008.
  11. 11.A. Giusti, D. Ciresan, J. Masci, L. Gambardella, and J. Schmidhuber. Fast image scanning with deep max-pooling convolutional neural networks. In CoRR, 2013.
  12. 12.G. Gkioxari, P. Arbelaez, L. Bourdev, and J. Malik. Articulated pose estimation using discriminative armlet classifiers. In CVPR'13.
  13. 13.K. Grauman, G. Shakhnarovich, and T. Darrell. Inferring 3d structure with a statistical image-based shape model. In ICCV, 2003.
  14. 14.G. Heitz, S. Gould, A. Saxena, and D. Koller. Cascaded classification models: Combining models for holistic scene understanding. 2008.
  15. 15.A. Jain, J. Tompson, M. Andriluka, G. Taylor, and C. Bregler. Learning human pose estimation features with convolutional networks. In ICLR, 2014.
  16. 16.S. Johnson and M. Everingham. Learning Effective Human Pose Estimation from Inaccurate Annotation. In CVPR'11.
  17. 17.S. Johnson and M. Everingham. Clustered pose and nonlinear appearance models for human pose estimation. In BMVC, 2010.
  18. 18.D. Lowe. Object recognition from local scale-invariant features. In ICCV, 1999.
  19. 19.M. Mathieu, M. Henaff, and Y. LeCun. Fast training of convolutional networks through ffts. In CoRR, 2013.
  20. 20.G. Mori and J. Malik. Estimating human body configurations using shape context matching. ECCV, 2002.
  21. 21.F. Morin and Y. Bengio. Hierarchical probabilistic neural network language model. In Proceedings of the Tenth International Workshop on Artificial Intelligence and Statistics, 2005.
  22. 22.F. Ning, D. Delhomme, Y. LeCun, F. Piano, L. Bottou, and P. Barbano. Toward automatic phenotyping of developing embryos from videos. IEEE TIP, 2005.
  23. 23.L. Pishchulin, M. Andriluka, P. Gehler, and B. Schiele. Poselet conditioned pictorial structures. In CVPR'13.
  24. 24.L. Pishchulin, M. Andriluka, P. Gehler, and B. Schiele. Strong appearance and expressive spatial models for human pose estimation. In ICCV'13.
  25. 25.D. Ramanan, D. Forsyth, and A. Zisserman. Strike a pose: Tracking people by finding stylized poses. In CVPR, 2005.
  26. 26.S. Ross, D. Munoz, M. Hebert, and J.A Bagnell. Learning message-passing inference machines for structured prediction. In CVPR, 2011.
  27. 27.B. Sapp and B. Taskar. Modec: Multimodal decomposable models for human pose estimation. In CVPR, 2013.
  28. 28.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. In ICLR, 2014.
  29. 29.J. Tompson, M. Stein, Y. LeCun, and K. Perlin. Real-time continuous pose recovery of human hands using convolutional networks. In TOG, 2014.
  30. 30.A. Toshev and C. Szegedy. Deeppose: Human pose estimation via deep neural networks. In CVPR, 2014.
  31. 31.Yi Yang and Deva Ramanan. Articulated pose estimation with flexible mixtures-of-parts. In CVPR'11.

Citation

MLA
Tompson, J., et al. “Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation”. arXiv, 2014, http://arxiv.org/abs/1406.2984v2.
APA
Tompson, J., Jain, A., LeCun, Y., & Bregler, C. (2014). Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation. arXiv. http://arxiv.org/abs/1406.2984v2
Chicago
Tompson, J., A. Jain, Y. LeCun, and C. Bregler. 2014. “Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation”. arXiv. http://arxiv.org/abs/1406.2984v2.
Harvard
Tompson, J. et al. (2014) “Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1406.2984v2.
Vancouver
1. Tompson J, Jain A, LeCun Y, Bregler C (2014) Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation. arXiv

BibTeX

@article{tompson2014joint,
  title = {Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation},
  author = {Tompson, Jonathan and Jain, Arjun and LeCun, Yann and Bregler, Christoph},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1406.2984v2},
  eprint = {1406.2984}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors