How Far are We from Solving the 2D & 3D Face Alignment Problem? (and a Dataset of 230,000 3D Facial Landmarks)

Adrian BulatGeorgios Tzimiropoulos

article2017ICCV1,698 citations

Presents a high-capacity deep neural network baseline alongside the 230,000-image LS3D-W dataset, demonstrating that modern architectures are approaching saturated performance across standard 2D and 3D face alignment benchmarks.

Listen

Facial landmark localization—identifying key structural points on human faces across images and video—is a foundational component for facial recognition, tracking, and augmented reality. While traditional cascaded regression techniques perform well on controlled, near-frontal faces, they struggle with severe head rotations, partial occlusions, poor image quality, and imprecise initializations. Moreover, the field has suffered from a lack of large-scale, consistent benchmarks for three-dimensional facial landmarking, making it difficult to assess how close modern computer vision methods are to truly solving this task across real-world conditions.

The article investigates how close state-of-the-art deep neural networks are to reaching saturation performance on major two-dimensional and three-dimensional face alignment benchmarks. To achieve this, it constructs a high-capacity architecture called the Face Alignment Network, creates a unified large-scale 3D facial dataset, and systematically evaluates the model against traditional performance bottlenecks.

To build the Face Alignment Network, the authors combined stacked hourglass network architectures with hierarchical, parallel, and multi-scale residual blocks. They addressed the scarcity of three-dimensional data by developing a guided neural network that converts existing 2D landmark annotations into 3D, creating the Large Scale 3D Faces in-the-Wild (LS3D-W) dataset containing roughly 230,000 images. The researchers conducted cross-dataset evaluations across multiple established benchmarks (such as 300-W, 300-VW, Menpo, and AFLW2000-3D) and carried out controlled ablation studies measuring accuracy across variations in facial pose, image resolution, initialization noise, and model parameter size.

The evaluation yielded several critical findings. First, both the 2D and 3D alignment networks achieved near-saturating performance across existing datasets, matching or exceeding prior state-of-the-art methods and producing failure rates under 0.4% on standard test sets. Second, visual inspections revealed that the remaining prediction errors were frequently due to low-quality or inaccurate ground-truth labels rather than model shortcomings. Third, the network exhibited high robustness across extreme yaw rotations (from -90 to 90 degrees), with performance declining only slightly at extreme profile angles. Fourth, the architecture proved highly resilient to real-world degradation, maintaining high accuracy on face resolutions as low as 30 pixels and under bounding box initialization noise of up to 30%. Finally, scaling experiments showed that reducing model size from 24 million parameters down to 12 million caused negligible performance loss while enabling processing speeds up to 150 frames per second on standard hardware.

These results demonstrate that deep landmark localization architectures are effectively capable of solving standard 2D and 3D face alignment under in-the-wild conditions. For technical leaders and product teams, this shifts the operational bottleneck away from algorithmic accuracy toward engineering trade-offs, such as optimizing model footprint and inference speed for mobile or edge deployment. Furthermore, the findings indicate that existing benchmarks have reached saturation, meaning future performance improvements will depend on higher-quality annotations and more challenging test conditions.

Organizations deploying facial alignment solutions should consider adopting stacked heatmap regression architectures and may safely compress model sizes down toward 12 million parameters to reduce compute costs and improve latency without noticeable accuracy loss. Engineering teams should also prioritize cleaning upstream bounding-box detection pipelines and refining label quality rather than attempting to over-engineer landmark localization models. Future development should focus on lightweight network designs (under 6 million parameters) to support resource-constrained hardware and expanding benchmarks to cover rare, extreme head poses.

The primary limitation of the study is its reliance on synthetically expanded datasets to train large-pose models, which introduces minor distortions around facial boundaries such as the ears. Additionally, the manual verification of large-scale datasets remains subject to label noise. Nonetheless, given the extensive evaluation across hundreds of thousands of diverse images, confidence in the primary findings remains high.

Cover for How Far are We from Solving the 2D & 3D Face Alignment Problem? (and a Dataset of 230,000 3D Facial Landmarks)

Abstract

This paper investigates how far a very deep neural network is from attaining close to saturating performance on existing 2D and 3D face alignment datasets. To this end, we make the following 5 contributions: (a) we construct, for the first time, a very strong baseline by combining a state-of-the-art architecture for landmark localization with a state-of-the-art residual block, train it on a very large yet synthetically expanded 2D facial landmark dataset and finally evaluate it on all other 2D facial landmark datasets. (b) We create a guided by 2D landmarks network which converts 2D landmark annotations to 3D and unifies all existing datasets, leading to the creation of LS3D-W, the largest and most challenging 3D facial landmark dataset to date ~230,000 images. (c) Following that, we train a neural network for 3D face alignment and evaluate it on the newly introduced LS3D-W. (d) We further look into the effect of all "traditional" factors affecting face alignment performance like large pose, initialization and resolution, and introduce a "new" one, namely the size of the network. (e) We show that both 2D and 3D face alignment networks achieve performance of remarkable accuracy which is probably close to saturating the datasets used. Training and testing code as well as the dataset can be downloaded from this https URL

Table of Contents

  • 1 Introduction
  • 2 Closely related work
  • 3 Datasets
  • 3.1 Training datasets
  • 3.2 Test datasets
  • 3.2.1 2D datasets
  • 3.2.2 3D datasets
  • 3.3 Metrics
  • 4 Method
  • 4.1 2D and 3D Face Alignment Networks
  • 4.2 2D-to-3D Face Alignment Network
  • 4.3 Training
  • 5 2D face alignment
  • 6 Large Scale 3D Faces in-the-Wild dataset
  • 7 3D face alignment
  • 8 Ablation studies
  • 9 Conclusions
  • 10 Acknowledgments
  • References
  • A1 Additional numeric results
  • A2 Additional visual results
  • A3 Full 3D face alignment

Knowls

  1. Knowl 1 — Face Alignment Network (FAN) Architecture

    model/method

    The Face Alignment Network (FAN) is a convolutional neural network designed for 2D and 3D facial landmark localization via heatmap regression.

    FAN is constructed by stacking four Hourglass (HG) modules. Inside each Hourglass module, the conventional residual bottleneck block is replaced by a hierarchical, parallel, and multi-scale residual block. This modification improves landmark localization precision while maintaining the parameter count.

    The network outputs spatial confidence heatmaps, with one heatmap per target facial landmark (N=68N = 68 landmarks). During training, the network is optimized using RMSprop with a batch size of 10 for 40 epochs. The learning rate is initialized at 10−410^{-4} and decayed by a factor of 10 at epoch 15 (to 10−510^{-5}) and epoch 30 (to 10−610^{-6}). Data augmentation includes horizontal flipping, random in-plane rotation (\\pm 50^\\circ), color jittering, random scale jittering (scaling factors between 0.8 and 1.2), and random occlusion.

    Two variants are trained depending on landmark definition:

    • 2D-FAN: Trained on the 300W-LP-2D dataset (and fine-tuned on the 300-W training set) to predict 2D facial landmarks.
    • 3D-FAN: Trained on the 300W-LP-3D dataset to predict 2D projections of 3D facial landmarks that preserve 3D correspondence across large yaw variations (from -90^\\circ to +90^\\circ).
  2. Knowl 2 — 2D-to-3D Face Alignment Network (2D-to-3D-FAN)

    model/method

    2D-to-3D-FAN is a neural network designed to convert 2D facial landmark annotations into 3D facial landmark coordinates (specifically, their 2D projections that maintain anatomical correspondence across wide pose changes up to \\pm 90^\\circ).

    The network is built on a 4-stack Hourglass architecture with hierarchical, parallel, and multi-scale residual blocks. Its input is a 71-channel tensor consisting of:

    1. The 3-channel RGB face image.
    2. 68 additional spatial heatmap channels, where channel kk contains a 2D Gaussian (sigma=1textpx\\sigma = 1\\text{ px}) centered at the (x,y)(x, y) coordinate of the kk-th 2D landmark.

    The network outputs 68 confidence heatmaps corresponding to the 2D locations of the 3D facial landmarks. It is trained on the 300W-LP dataset (which contains paired 2D and 3D landmark annotations) using the RMSprop optimizer for 40 epochs. The initial learning rate is set to 10−310^{-3} and decreased by a factor of 10 at epoch 15 and epoch 30, with data augmentation consisting of rotations in [-70^\\circ, 70^\\circ] and scaling factors in [0.7,1.3][0.7, 1.3].

  3. Knowl 3 — Large Scale 3D Faces in-the-Wild (LS3D-W) Dataset

    experimental setup

    The Large Scale 3D Faces in-the-Wild (LS3D-W) dataset is a facial landmark benchmark containing approximately 230,000 in-the-wild facial images annotated with 68 3D landmarks spanning yaw angles from -90^\\circ to +90^\\circ.

    LS3D-W is constructed by standardizing and unifying existing in-the-wild 2D facial datasets using a two-stage automated conversion pipeline: 2D-FAN is first used to extract 2D facial landmarks, and 2D-to-3D-FAN subsequently converts these 2D landmarks into 3D landmark annotations. The unified source datasets comprise:

    • 300-W test set: 600 images split into Indoor and Outdoor subsets.
    • 300-VW: 218,595 video frames across 114 videos (64 test videos across Categories A, B, and C, and 50 training videos).
    • Menpo: ~9,000 images from FDDB and AFLW, including both frontal and profile faces.
    • AFLW2000-3D: 2,000 re-annotated images from AFLW covering extreme head poses.

    A balanced evaluation subset, LS3D-W Balanced, consists of 7,200 images sampled from LS3D-W with an equal distribution of 2,400 images across each of three yaw angle intervals: [0^\\circ, 30^\\circ], [30^\\circ, 60^\\circ], and [60^\\circ, 90^\\circ].

  4. Knowl 4 — Full 3D Facial Landmark Localization Network (Full-2D-to-3D-FAN)

    model/method

    Full-2D-to-3D-FAN is a two-stage network that estimates full (x,y,z)(x, y, z) 3D coordinates for N=68N = 68 facial landmarks from a single RGB image guided by 2D landmarks.

    The pipeline operates in two cascaded stages:

    1. (x,y)(x, y) Prediction: A 2D-to-3D-FAN takes an RGB image concatenated with 68 2D landmark Gaussian heatmaps (71 channels total) and produces N=68N = 68 confidence heatmaps representing the 2D projected coordinates (xk,yk)(x_k, y_k).
    2. Depth (zz) Regression Subnetwork: A ResNet-152 backbone is modified to accept a (3+N)(3 + N)-channel input (the 3-channel RGB image stacked together with the N=68N = 68 output heatmaps from the first stage). The output fully connected layer is modified to predict an Ntimes1N \\times 1 continuous vector mathbfz=[z1,z2,dots,zN]T\\mathbf{z} = [z_1, z_2, \\dots, z_N]^T representing the depth coordinates of the NN landmarks.

    The depth subnetwork is trained using an L2L_2 regression loss for 50 epochs using RMSprop with the same learning rate schedule as the 2D-to-3D-FAN.

  5. Knowl 5 — Bounding-Box Normalized Mean Error (NME) for Facial Landmarks

    equation

    To evaluate facial landmark localization across arbitrary poses (including profile views where interocular distance approaches zero), the Normalized Mean Error (NME) is normalized by the geometric mean of the ground-truth 2D bounding box dimensions:

    textNME=frac1Nsumk=1Nfrac∣mathbfxk−mathbfyk∣2d\\text{NME} = \\frac{1}{N} \\sum_{k=1}^N \\frac{\\|\\mathbf{x}_k - \\mathbf{y}_k\\|_2}{d}

    where:

    • NN is the total number of annotated facial landmarks (N=68N = 68).
    • mathbfxkinmathbbR2\\mathbf{x}_k \\in \\mathbb{R}^2 denotes the ground-truth 2D landmark coordinate vector for the kk-th landmark.
    • mathbfykinmathbbR2\\mathbf{y}_k \\in \\mathbb{R}^2 denotes the predicted 2D landmark coordinate vector for the kk-th landmark.
    • ∣cdot∣2\\| \\cdot \\|_2 is the standard Euclidean norm (L2L_2 distance) in pixels.
    • d=sqrtwtextbboxcdothtextbboxd = \\sqrt{w_{\\text{bbox}} \\cdot h_{\\text{bbox}}} is the normalization factor, where wtextbboxw_{\\text{bbox}} and htextbboxh_{\\text{bbox}} are the width and height of the tight bounding box computed from the 2D ground-truth landmarks.
  6. Knowl 6 — Comparative Performance of 2D-FAN on 2D Face Alignment Benchmarks

    data/table

    2D-FAN was evaluated across several major 2D face alignment test benchmarks using Area Under the Curve (AUC) computed on cumulative error distributions up to a maximum NME threshold of 7%.

    Dataset 2D-FAN (Ours) MDM iCCR TCDCN CFSS
    300VW-A 72.1% 70.2% 65.9% - -
    300VW-B 71.2% 67.9% 65.5% - -
    300VW-C 64.1% 54.6% 58.1% - -
    Menpo 67.5% 67.1% - 47.9% 60.5%
    300W 66.9% 58.1% - 41.7% 55.9%

    2D-FAN consistently outperforms existing cascaded regression and deep recurrent methods across all datasets. Its error curves on the in-the-wild datasets match the performance achieved by MDM on the near-frontal, controlled LFPW test set. Across 7,200 test images from 300-W and Menpo, only 18 failure cases (defined as textNME>7\\text{NME} > 7\\%) occurred, representing a failure rate of 0.25%.

  7. Knowl 7 — 3D Face Alignment Performance on the LS3D-W Benchmark

    empirical result

    3D-FAN was trained on the synthetically expanded 300W-LP-3D dataset and evaluated across all subsets of the LS3D-W benchmark (comprising ~230,000 images: 300-W-3D test set, 300-VW-3D Categories A, B, and C, Menpo-3D, and re-annotated AFLW2000-3D).

    3D-FAN substantially outperforms the 3D Morphable Model fitting baseline (3DDFA) across all LS3D-W subsets. It produces consistent cumulative error distributions across all subsets, showing improved localization accuracy compared to 2D-FAN at tight error thresholds (textNME<2\\text{NME} < 2\\%) due to the elimination of annotation mark-up discrepancies between training and evaluation data.

  8. Knowl 8 — Robustness of 3D-FAN to Pose, Initialization Noise, and Resolution

    data/table

    3D-FAN was tested on the LS3D-W Balanced benchmark (7,200 images, 2,400 per yaw range) to assess resilience against head yaw, bounding box initialization perturbations, and reduced image resolution, measured by AUC calculated at an NME threshold of 7%.

    Performance across yaw poses:

    Yaw Range Number of Images 3D-FAN (Ours)
    [0∘−30∘][0^\circ - 30^\circ] 2400 73.5%
    [30∘−60∘][30^\circ - 60^\circ] 2400 74.6%
    [60∘−90∘][60^\circ - 90^\circ] 2400 68.8%

    Performance across initialization bounding box noise levels:

    Noise Level [0∘−30∘][0^\circ - 30^\circ] [30∘−60∘][30^\circ - 60^\circ] [60∘−90∘][60^\circ - 90^\circ]
    0% 74.5% 75.2% 69.8%
    10% 73.5% 74.6% 68.8%
    20% 70.8% 71.7% 66.1%
    30% 63.8% 63.5% 57.2%

    Across face resolutions, 3D-FAN maintains high AUC performance without retraining down to a face bounding box size of 40 pixels, with significant accuracy degradation occurring only when face size falls to 30 pixels.

  9. Knowl 9 — Influence of Parameter Count on 3D-FAN Accuracy and Inference Speed

    data/table

    The parameter capacity of 3D-FAN was evaluated on the LS3D-W Balanced dataset by reducing the number of stacked Hourglass modules from 4 down to 1 and scaling the internal channel capacity of the residual blocks.

    #params [0∘−30∘][0^\circ - 30^\circ] [30∘−60∘][30^\circ - 60^\circ] [60∘−90∘][60^\circ - 90^\circ]
    2M 70.9% 69.9% 55.8%
    4M 71.0% 70.5% 57.0%
    6M 71.5% 71.1% 58.3%
    12M 72.7% 72.7% 67.1%
    18M 73.4% 74.2% 68.3%
    24M 73.5% 74.6% 68.8%

    Between 12M and 24M parameters, localization accuracy remains nearly constant across all pose bins. Accuracy drops significantly below 6M parameters, particularly for profile poses ([60∘,90∘][60^\circ, 90^\circ]). In terms of inference speed on an NVIDIA TitanX GPU, the 24M parameter model operates at 28–30 frames per second (fps), while the 2M parameter model reaches 150 fps.

Coverage note — None omitted; qualitative visualizations of individual facial fittings and background comparisons to classical cascaded regression methods were excluded as they do not constitute standalone contributed findings.

References

  1. 1.M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele. 2d human pose estimation: New benchmark and state of the art analysis. In CVPR, 2014.
  2. 2.P. Belhumeur, D. Jacobs, D. Kriegman, and N. Kumar. Localizing parts of faces using a consensus of exemplars. In CVPR, 2011.
  3. 3.P. N. Belhumeur, D. W. Jacobs, D. J. Kriegman, and N. Kumar. Localizing parts of faces using a consensus of exemplars. TPAMI, 2013.
  4. 4.A. Bulat and G. Tzimiropoulos. Convolutional aggregation of local evidence for large pose face alignment. In BMVC, 2016.
  5. 5.A. Bulat and G. Tzimiropoulos. Human pose estimation via convolutional part heatmap regression. In ECCV, 2016.
  6. 6.A. Bulat and G. Tzimiropoulos. Two-stage convolutional part heatmap regression for the 1st 3d face alignment in the wild (3dfaw) challenge. In ECCV, 2016.
  7. 7.A. Bulat and G. Tzimiropoulos. Binarized convolutional landmark localizers for human pose estimation and face alignment with limited resources. In ICCV, 2017.
  8. 8.X. Cao, Y. Wei, F. Wen, and J. Sun. Face alignment by explicit shape regression. In CVPR, 2012.
  9. 9.R. Collobert, K. Kavukcuoglu, and C. Farabet. Torch7: A matlab-like environment for machine learning. In NIPS-W, 2011.
  10. 10.D. Cristinacce and T. F. Cootes. Feature detection and tracking with constrained local models. In BMVC, 2006.
  11. 11.P. Dollar, P. Welinder, and P. Perona. Cascaded pose regression. In CVPR, 2010.
  12. 12.P. F. Felzenszwalb and D. P. Huttenlocher. Pictorial structures for object recognition. IJCV, 61(1):55–79, 2005.
  13. 13.R. Gross, I. Matthews, J. Cohn, T. Kanade, and S. Baker. Multi-pie. IVC, 2010.
  14. 14.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016.
  15. 15.K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. In ECCV, 2016.
  16. 16.G. B. Huang, M. Ramesh, T. Berg, and E. Learned-Miller. Labeled faces in the wild: A database for studying face recognition in unconstrained environments. Technical report, University of Massachusetts, Amherst, 2007.
  17. 17.E. Insafutdinov, L. Pishchulin, B. Andres, M. Andriluka, and B. Schiele. Deepercut: A deeper, stronger, and faster multi-person pose estimation model. arXiv, 2016.
  18. 18.V. Jain and E. G. Learned-Miller. Fddb: A benchmark for face detection in unconstrained settings. UMass Amherst Technical Report, 2010.
  19. 19.L. A. Jeni, S. Tulyakov, L. Yin, N. Sebe, and J. F. Cohn. The first 3d face alignment in the wild (3dfaw) challenge. In ECCV, 2016.
  20. 20.A. Jourabloo and X. Liu. Large-pose face alignment via cnn-based dense 3d model fitting. In CVPR, 2016.
  21. 21.M. Kostinger, P. Wohlhart, P. M. Roth, and H. Bischof. Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization. In ICCV-W, 2011.
  22. 22.V. Le, J. Brandt, Z. Lin, L. Bourdev, and T. S. Huang. Interactive facial feature localization. In ECCV, 2012.
  23. 23.A. Newell, K. Yang, and J. Deng. Stacked hourglass networks for human pose estimation. In ECCV, 2016.
  24. 24.T. Pfister, J. Charles, and A. Zisserman. Flowing convnets for human pose estimation in videos. In ICCV, 2015.
  25. 25.L. Pishchulin, M. Andriluka, P. Gehler, and B. Schiele. Poselet conditioned pictorial structures. In CVPR, 2013.
  26. 26.L. Pishchulin, M. Andriluka, P. Gehler, and B. Schiele. Strong appearance and expressive spatial models for human pose estimation. In CVPR, 2013.
  27. 27.L. Pishchulin, E. Insafutdinov, S. Tang, B. Andres, M. Andriluka, P. V. Gehler, and B. Schiele. Deepcut: Joint subset partition and labeling for multi person pose estimation. In CVPR, 2016.
  28. 28.C. Sagonas, E. Antonakos, G. Tzimiropoulos, S. Zafeiriou, and M. Pantic. 300 faces in-the-wild challenge: Database and results. IVC, 47:3–18, 2016.
  29. 29.C. Sagonas, G. Tzimiropoulos, S. Zafeiriou, and M. Pantic. 300 faces in-the-wild challenge: The first facial landmark localization challenge. In CVPR, 2013.
  30. 30.C. Sagonas, G. Tzimiropoulos, S. Zafeiriou, and M. Pantic. A semi-automatic methodology for facial landmark annotation. In CVPR, 2013.
  31. 31.E. Sánchez-Lozano, B. Martinez, G. Tzimiropoulos, and M. Valstar. Cascaded continuous regression for real-time incremental face tracking. In ECCV, 2016.
  32. 32.B. Sapp and B. Taskar. Modec: Multimodal decomposable models for human pose estimation. In CVPR, 2013.
  33. 33.J. Shen, S. Zafeiriou, G. G. Chrysos, J. Kossaifi, G. Tzimiropoulos, and M. Pantic. The first facial landmark tracking in-the-wild challenge: Benchmark and results. In ICCVW, 2015.
  34. 34.B. M. Smith and L. Zhang. Collaborative facial landmark localization for transferring annotations across datasets. In ECCV, 2014.
  35. 35.Y. Sun, X. Wang, and X. Tang. Deep convolutional network cascade for facial point detection. In CVPR, 2013.
  36. 36.Y. Tian, C. L. Zitnick, and S. G. Narasimhan. Exploring the spatial hierarchy of mixture models for human pose estimation. In ECCV. 2012.
  37. 37.T. Tieleman and G. Hinton. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning, 4(2), 2012.
  38. 38.J. J. Tompson, A. Jain, Y. LeCun, and C. Bregler. Joint training of a convolutional network and a graphical model for human pose estimation. In NIPS, 2014.
  39. 39.A. Toshev and C. Szegedy. Deeppose: Human pose estimation via deep neural networks. In CVPR, 2014.
  40. 40.G. Trigeorgis, P. Snape, M. A. Nicolaou, E. Antonakos, and S. Zafeiriou. Mnemonic descent method: A recurrent process applied for end-to-end face alignment. In CVPR, 2016.
  41. 41.G. Tzimiropoulos. Project-out cascaded regression with an application to face alignment. In CVPR, 2015.
  42. 42.S.-E. Wei, V. Ramakrishna, T. Kanade, and Y. Sheikh. Convolutional pose machines. In CVPR, 2016.
  43. 43.X. Xiong and F. De la Torre. Supervised descent method and its applications to face alignment. In CVPR, 2013.
  44. 44.Y. Yang and D. Ramanan. Articulated pose estimation with flexible mixtures-of-parts. In CVPR, 2011.
  45. 45.S. Zaferiou. The menpo facial landmark localisation challenge. In CVPR-W, 2017.
  46. 46.J. Zhang, M. Kan, S. Shan, and X. Chen. Leveraging datasets with varying annotations for face alignment via deep regression network. In ICCV, 2015.
  47. 47.Z. Zhang, P. Luo, C. C. Loy, and X. Tang. Facial landmark detection by deep multi-task learning. In ECCV. 2014.
  48. 48.S. Zhu, C. Li, C. Change Loy, and X. Tang. Face alignment by coarse-to-fine shape searching. In CVPR, 2015.
  49. 49.S. Zhu, C. Li, C. C. Loy, and X. Tang. Transferring landmark annotations for cross-dataset face alignment. arXiv, 2014.
  50. 50.X. Zhu, Z. Lei, X. Liu, H. Shi, and S. Z. Li. Face alignment across large poses: A 3d solution. In CVPR, 2016.
  51. 51.X. Zhu and D. Ramanan. Face detection, pose estimation, and landmark localization in the wild. In CVPR. IEEE, 2012.

Citation

MLA
Bulat, A., and G. Tzimiropoulos. “How Far Are We from Solving the 2D & 3D Face Alignment Problem? (and a Dataset of 230,000 3D Facial Landmarks)”. 2017 IEEE International Conference on Computer Vision (ICCV), 2017, pp. 1021–30, https://doi.org/10.1109/ICCV.2017.116.
APA
Bulat, A., & Tzimiropoulos, G. (2017). How Far are We from Solving the 2D & 3D Face Alignment Problem? (and a Dataset of 230,000 3D Facial Landmarks). 2017 IEEE International Conference on Computer Vision (ICCV), 1021–1030. https://doi.org/10.1109/ICCV.2017.116
Chicago
Bulat, A., and G. Tzimiropoulos. 2017. “How Far Are We from Solving the 2D & 3D Face Alignment Problem? (and a Dataset of 230,000 3D Facial Landmarks)”. 2017 IEEE International Conference on Computer Vision (ICCV), 1021–30. https://doi.org/10.1109/ICCV.2017.116.
Harvard
Bulat, A. and Tzimiropoulos, G. (2017) “How Far are We from Solving the 2D & 3D Face Alignment Problem? (and a Dataset of 230,000 3D Facial Landmarks)”, 2017 IEEE International Conference on Computer Vision (ICCV). IEEE, pp. 1021–1030. Available at: https://doi.org/10.1109/ICCV.2017.116.
Vancouver
1. Bulat A, Tzimiropoulos G (2017) How Far are We from Solving the 2D & 3D Face Alignment Problem? (and a Dataset of 230,000 3D Facial Landmarks). In: 2017 IEEE International Conference on Computer Vision (ICCV). IEEE, pp 1021–1030

BibTeX

@inproceedings{Bulat_2017, title={How Far are We from Solving the 2D &amp; 3D Face Alignment Problem? (and a Dataset of 230,000 3D Facial Landmarks)}, url={http://dx.doi.org/10.1109/ICCV.2017.116}, DOI={10.1109/iccv.2017.116}, booktitle={2017 IEEE International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Bulat, Adrian and Tzimiropoulos, Georgios}, year={2017}, month=Oct, pages={1021–1030} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE