Facial Landmark Detection by Deep Multi-task Learning

Zhanpeng ZhangPing LuoChen Change LoyXiaoou Tang

article2014ECCV1,545 citations

Proposes a tasks-constrained deep convolutional network with task-wise early stopping that optimizes facial landmark detection alongside auxiliary tasks like pose estimation and attribute inference, significantly reducing model complexity while maintaining high accuracy under severe occlusion.

Listen

Facial landmark detection—identifying key facial features such as the eyes, nose, and mouth corners—is a vital foundation for computer vision applications, including face recognition and demographic analysis. However, real-world deployment frequently encounters major performance drops when dealing with extreme head turns, partial occlusions (such as sunglasses), and changing expressions. Traditional systems treat landmark detection as an isolated problem or rely on heavy, multi-stage neural network pipelines that are computationally expensive and difficult to maintain.

The article evaluates whether joint multi-task learning can improve landmark detection accuracy while simplifying model complexity. Specifically, it demonstrates the Tasks-Constrained Deep Convolutional Network (TCDCN), which optimizes landmark localization simultaneously with related tasks: head pose estimation, gender classification, smiling detection, and glasses detection.

To evaluate this framework, the authors developed a multi-task learning architecture that shares low-level visual representations across tasks while applying task-specific prediction heads. To overcome the practical challenge of tasks learning at different speeds and causing overfitting, the authors introduced an automated task-wise early stopping criterion. The model was trained on a 10,000-image dataset (the Multi-Task Facial Landmark dataset) and evaluated against leading commercial and academic baselines on challenging public benchmark datasets, including AFLW, AFW, and COFW.

The analysis yielded three critical findings. First, multi-task learning substantially improved detection accuracy, reducing the overall failure rate by more than 10% compared to a single-task baseline, with head pose estimation providing the single largest performance benefit across all facial points. Second, the proposed single-network architecture outperformed existing state-of-the-art cascaded neural networks while operating roughly seven times faster (17 milliseconds versus 120 milliseconds per face on a standard central processing unit), eliminating the need for a complex 23-network pipeline. Third, using the network's five-point output as an initial baseline for dense landmark algorithms consistently reduced localization errors on heavily occluded faces.

These results demonstrate that auxiliary tasks act as natural regularizers, allowing shared visual representations to become more robust against pose and expression shifts without adding run-time complexity. For decision-makers, this enables significant cost and computational savings: high-accuracy facial analytics can now run in real time on standard consumer hardware without requiring specialized graphical processing units or bulky cascaded models.

Organizations developing or deploying face-analysis technologies should adopt unified multi-task architectures in place of complex cascaded pipelines to lower computational latency and improve edge-device capability. In addition, teams should utilize automated task-wise early stopping during model training to ensure multi-task convergence without manual hyperparameter tuning. Future technical initiatives should expand this multi-task framework to dense facial landmark maps and explore similar shared-task architectures in other visual domain challenges.

The reported findings carry high confidence across standard benchmark settings, though performance was evaluated specifically on low-resolution inputs (40 by 40 pixels) predicting five primary landmark points. Deployments requiring dense 3D surface mapping or ultra-high-resolution edge details may require further pilot validation before fully replacing specialized downstream alignment pipelines.

  • Paper: Regularized multi--task learning, T. Evgeniou et al. (2004). It establishes the foundational mathematical principles of multi-task learning and task-coupling regularization that motivate learning shared representations across related visual tasks.
  • Paper: Convex multi-task feature learning, Andreas Argyriou et al. (2008). It introduces convex multi-task feature learning for joint representation sharing, providing the foundational theory for using auxiliary objectives as regularizers.
  • Paper: DeepPose: Human Pose Estimation via Deep Neural Networks, Alexander Toshev et al. (2014). It pioneered casting keypoint localization directly as deep neural network coordinate regression, establishing the direct regression paradigm adopted by TCDCN.
  • Paper: Deep Neural Networks for Object Detection, Christian Szegedy et al. (2013). It demonstrates how deep convolutional networks can directly regress geometric coordinates and masks for visual localization tasks.
  • Paper: Sharing visual features for multiclass and multiview object detection, Antonio Torralba et al. (2007). It demonstrates how sharing visual features across multiple viewpoints and object categories improves efficiency and generalizability.
  • Paper: Recognizing Action Units for Facial Expression Analysis, Ying-li Tian et al. (2001). It provides foundational insights into tracking facial geometry and classifying discrete facial actions that underpin multi-attribute facial analysis.
Cover for Facial Landmark Detection by Deep Multi-task Learning

Abstract

Facial landmark detection has long been impeded by the problems of occlusion and pose variation. Instead of treating the detection task as a single and independent problem, we investigate the possibility of improving detection robustness through multi-task learning. Specifically, we wish to optimize facial landmark detection together with heterogeneous but subtly correlated tasks, e.g.head pose estimation and facial attribute inference. This is non-trivial since different tasks have different learning difficulties and convergence rates. To address this problem, we formulate a novel tasks-constrained deep model, with task-wise early stopping to facilitate learning convergence. Extensive evaluations show that the proposed task-constrained learning (i) outperforms existing methods, especially in dealing with faces with severe occlusion and pose variation, and (ii) reduces model complexity drastically compared to the state-of-the-art method based on cascaded deep model [21].

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Tasks-Constrained Facial Landmark Detection
  • 3.1 Problem Formulation
  • 3.2 Learning Tasks-Constrained Deep Convolutional Network
  • 4 Implementation and Experiments
  • 4.1 Evaluating the Effectiveness of Learning with Related Task
  • 4.2 The Benefits of Task-Wise Early Stopping
  • 4.3 Comparison with the Cascaded CNN [21]
  • 4.4 Comparison with Other State-of-the-Art Methods
  • 4.5 TCDCN for Robust Initialization
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — Tasks-Constrained Deep Convolutional Network (TCDCN) Architecture

    model/method

    The Tasks-Constrained Deep Convolutional Network (TCDCN) is a unified convolutional neural network designed to jointly estimate facial landmarks and auxiliary facial attributes from a single 40×4040 \times 40 pixel grayscale face image.

    The feature extraction backbone consists of four locally connected convolutional layers (where filter weights are not spatially shared across locations), three max-pooling layers, and one shared fully connected layer:

    1. Convolutional layer 1 (5×55 \times 5 filters) followed by 2×22 \times 2 non-overlapping max-pooling, outputting an 18×18×1618 \times 18 \times 16 feature map.
    2. Convolutional layer 2 (3×33 \times 3 filters) followed by 2×22 \times 2 non-overlapping max-pooling, outputting an 8×8×488 \times 8 \times 48 feature map.
    3. Convolutional layer 3 (3×33 \times 3 filters) followed by 2×22 \times 2 non-overlapping max-pooling, outputting a 3×3×643 \times 3 \times 64 feature map.
    4. Convolutional layer 4 (2×22 \times 2 filters), producing a 2×2×642 \times 2 \times 64 feature map.
    5. Fully connected layer producing a shared 100-dimensional latent representation vector xl\mathbf{x}^l.

    The absolute tangent function σ(u)=∣tanh⁡(u)∣\sigma(u) = |\tanh(u)| is used as the non-linear activation function across all intermediate layers. The shared 100-dimensional representation feeds into parallel task-specific output heads: a linear regression head predicting the 2D coordinates of five facial landmarks (10 continuous values), and classification heads (softmax/logistic regression) for auxiliary tasks.

  2. Knowl 2 — TCDCN Multi-Task Objective Loss Function

    equation

    The multi-task objective function for training the Tasks-Constrained Deep Convolutional Network combines a least-squares regression loss for landmark localization with cross-entropy classification losses for auxiliary tasks, subject to L2L_2 weight regularization:

    argmin⁡Wr,{Wa}a∈A12∑i=1N∥yir−f(xi;Wr)∥22−∑i=1N∑a∈Aλayialog⁡(p(yia∣xi;Wa))+∑t=1T∥Wt∥22\operatorname{argmin}_{\mathbf{W}^r, \{\mathbf{W}^a\}_{a \in \mathcal{A}}} \frac{1}{2} \sum_{i=1}^N \|\mathbf{y}_i^r - f(\mathbf{x}_i; \mathbf{W}^r)\|_2^2 - \sum_{i=1}^N \sum_{a \in \mathcal{A}} \lambda^a \mathbf{y}_i^a \log\left(p(\mathbf{y}_i^a \mid \mathbf{x}_i; \mathbf{W}^a)\right) + \sum_{t=1}^T \|\mathbf{W}^t\|_2^2

    where:

    • NN is the number of training samples.
    • xi∈R100\mathbf{x}_i \in \mathbb{R}^{100} is the shared high-level feature vector extracted by the convolutional network for the ii-th face image.
    • yir∈R10\mathbf{y}_i^r \in \mathbb{R}^{10} contains the ground-truth 2D coordinates for the five facial landmarks (left eye center, right eye center, nose tip, left mouth corner, right mouth corner) for sample ii, and f(xi;Wr)=(Wr)Txif(\mathbf{x}_i; \mathbf{W}^r) = (\mathbf{W}^r)^T \mathbf{x}_i is the linear prediction function parameterized by weights Wr\mathbf{W}^r.
    • A\mathcal{A} is the set of auxiliary tasks, which include head pose classification (5 classes: 0∘,±30∘,±60∘0^\circ, \pm 30^\circ, \pm 60^\circ), gender classification (binary), wearing glasses (binary), and smiling (binary).
    • λa\lambda^a is the task importance weighting coefficient for auxiliary task aa.
    • p(yia=m∣xi;Wa)=exp⁡((Wma)Txi)∑jexp⁡((Wja)Txi)p(y_i^a = m \mid \mathbf{x}_i; \mathbf{W}^a) = \frac{\exp\left((\mathbf{W}_m^a)^T \mathbf{x}_i\right)}{\sum_j \exp\left((\mathbf{W}_j^a)^T \mathbf{x}_i\right)} is the class posterior probability modeled by softmax regression, with Wma\mathbf{W}_m^a denoting the mm-th column of the weight matrix Wa\mathbf{W}^a.
    • Wt∈{Wr}∪{Wa}a∈A\mathbf{W}^t \in \{\mathbf{W}^r\} \cup \{\mathbf{W}^a\}_{a \in \mathcal{A}} denotes the parameter weight matrix for task tt among the total T=1+∣A∣T = 1 + |\mathcal{A}| tasks.
  3. Knowl 3 — Task-Wise Early Stopping Criterion for Multi-Task Deep Learning

    model/method

    To prevent auxiliary tasks with faster convergence rates or lower learning difficulty from overfitting the shared feature representation and degrading the primary landmark detection task, training of auxiliary task aa is halted when the following condition is met:

    k⋅med⁡j=t−ktEtra(j)∑j=t−ktEtra(j)−k⋅med⁡j=t−ktEtra(j)⋅Evala(t)−min⁡j=1…tEtra(j)λa⋅min⁡j=1…tEtra(j)>ϵ\frac{k \cdot \operatorname{med}_{j=t-k}^t E_{\text{tr}}^a(j)}{\sum_{j=t-k}^t E_{\text{tr}}^a(j) - k \cdot \operatorname{med}_{j=t-k}^t E_{\text{tr}}^a(j)} \cdot \frac{E_{\text{val}}^a(t) - \min_{j=1\dots t} E_{\text{tr}}^a(j)}{\lambda^a \cdot \min_{j=1\dots t} E_{\text{tr}}^a(j)} > \epsilon

    where:

    • tt is the current training iteration.
    • kk is the length of a sliding iteration strip.
    • med⁡(⋅)\operatorname{med}(\cdot) computes the median value over the window.
    • Etra(j)E_{\text{tr}}^a(j) and Evala(t)E_{\text{val}}^a(t) denote the loss of auxiliary task aa on the training set at iteration jj and on the validation set at iteration tt, respectively.
    • λa\lambda^a is the importance weight of auxiliary task aa (updated via gradient descent).
    • ϵ\epsilon is a predefined stopping threshold.

    The first factor captures the training progress over the recent window: a rapid drop in training loss yields a small value, indicating that training should continue. The second factor measures the validation generalization error relative to the minimum historical training loss, scaled inversely by task importance λa\lambda^a so that more critical auxiliary tasks are retained for longer.

  4. Knowl 4 — Error Back-Propagation through the Shared Representation Layer

    equation

    In TCDCN, gradients from the primary landmark regression task and all active auxiliary classification tasks are combined at the shared fully connected feature layer ll and back-propagated to the lower convolutional layers.

    The error vector εl\boldsymbol{\varepsilon}^l at the shared feature representation layer xl\mathbf{x}^l is computed as:

    εl=Wr(yir−(Wr)Txl)+∑a∈AWa(yia−p(yia∣xl;Wa))\boldsymbol{\varepsilon}^l = \mathbf{W}^r \left(\mathbf{y}_i^r - (\mathbf{W}^r)^T \mathbf{x}^l\right) + \sum_{a \in \mathcal{A}} \mathbf{W}^a \left(\mathbf{y}_i^a - p(\mathbf{y}_i^a \mid \mathbf{x}^l; \mathbf{W}^a)\right)

    For any intermediate layer l−1l-1, the error vector is propagated backward via:

    εl−1=(Wsl)Tεl⊙∂σ(ul−1)∂ul−1\boldsymbol{\varepsilon}^{l-1} = (\mathbf{W}^{s_l})^T \boldsymbol{\varepsilon}^l \odot \frac{\partial \sigma(\mathbf{u}^{l-1})}{\partial \mathbf{u}^{l-1}}

    where Wsl\mathbf{W}^{s_l} represents the filter weights of layer ll, ul−1\mathbf{u}^{l-1} is the pre-activation input of layer l−1l-1, σ(⋅)\sigma(\cdot) is the non-linear activation function, and ⊙\odot is element-wise multiplication. The gradient with respect to filter weights Wsl\mathbf{W}^{s_l} over receptive field Ω\Omega is given by:

    ∂E∂Wsl=εl(xΩl−1)T\frac{\partial E}{\partial \mathbf{W}^{s_l}} = \boldsymbol{\varepsilon}^l (\mathbf{x}_{\Omega}^{l-1})^T

  5. Knowl 5 — Multi-Task Facial Landmark (MTFL) Dataset Specification

    definition

    The Multi-Task Facial Landmark (MTFL) dataset contains 10,000 outdoor face images collected from the web. Each face image is annotated with:

    1. A face bounding box.
    2. 2D coordinates for five facial landmarks: left eye center, right eye center, nose tip, left mouth corner, and right mouth corner.
    3. Head pose discretized into 5 bins: 0∘0^\circ (frontal), +30∘+30^\circ (left), −30∘-30^\circ (right), +60∘+60^\circ (left profile), and −60∘-60^\circ (right profile).
    4. Gender attribute: male or female.
    5. Eye attribute: wearing glasses or without glasses.
    6. Expression attribute: smiling or not smiling.

    Data augmentation during training includes translation, in-plane rotation, and zooming.

  6. Knowl 6 — Effect of Auxiliary Task Combinations on Facial Landmark Localization

    empirical result

    On 3,000 randomly selected test images from the Annotated Facial Landmarks in the Wild (AFLW) dataset (where 39% of faces exhibit non-frontal poses), training TCDCN jointly with auxiliary tasks systematically reduces landmark detection failure rates (defined as mean landmark error normalized by inter-ocular distance exceeding 10%):

    • Facial Landmark Detection alone (FLD): 35.62% failure rate.
    • FLD + Gender: 31.86% failure rate.
    • FLD + Glasses: 32.87% failure rate.
    • FLD + Smile: 32.37% failure rate.
    • FLD + Pose: 28.76% failure rate.
    • FLD + All auxiliary tasks: 25.00% failure rate.

    Pose estimation provides the largest single-task gain due to its global constraint on landmark geometry across profile angles. Smiling attribute inference specifically reduces error at the lower face landmarks: Pearson's correlation between the learned weights of the smiling classification head and the landmark regression heads is highest for the right mouth corner (r=0.40r = 0.40), left mouth corner (r=0.22r = 0.22), and nose (r=0.17r = 0.17), compared to the eye centers (r=0.11r = 0.11 and r=0.32r = 0.32).

  7. Knowl 7 — Empirical Impact of Task-Wise Early Stopping on Training and Convergence

    empirical result

    Training the full multi-task model (FLD+all) with task-wise early stopping prevents destructive interference caused by varying task convergence rates:

    • Without task-wise early stopping: Training and validation loss curves exhibit pronounced oscillations and slow convergence. Validation mean landmark errors across all five landmarks range between 13.0% and 14.5%.
    • With task-wise early stopping: Training converges rapidly and stably. Validation mean landmark errors drop to between 7.5% and 9.5% across individual landmarks.

    During training with the early stopping rule, individual auxiliary tasks are automatically terminated sequentially as they reach peak validation utility: 'wearing glasses' terminates at approximately iteration 250, 'gender' at iteration 350, 'smiling' after iteration 400, and 'pose' remains active longest, terminating around iteration 750.

  8. Knowl 8 — Computational Complexity and Runtime: TCDCN vs. Cascaded CNN

    empirical result

    The computational complexity of a CNN with LL layers is O(∑l=1Lsl2qlql−1)=O(s2q2)\mathcal{O}\left(\sum_{l=1}^L s_l^2 q_l q_{l-1}\right) = \mathcal{O}(s^2 q^2), where sl2s_l^2 is the 2D feature map size and qlq_l is the number of filters at layer ll.

    While the cascaded deep convolutional architecture of Sun et al. requires 23 separate CNNs across multiple coarse-to-fine levels, TCDCN uses a single CNN with equivalent input size (40×4040 \times 40):

    • CPU Inference Time (Intel Core i5): TCDCN takes 17 ms per image compared to 120 ms (0.12 s) for the cascaded CNN, achieving an approximate 7×7\times speedup.
    • GPU Inference Time (NVIDIA GTX 760): TCDCN processes an image in 1.5 ms.
    • Accuracy: On AFLW, TCDCN achieves lower mean localization error on four of the five landmarks and a lower failure rate compared to the 23-network cascaded CNN.
  9. Knowl 9 — Facial Landmark Localization Benchmark Comparison on AFLW and AFW

    data/table

    The overall mean normalized landmark localization errors (normalized by inter-ocular distance, in %) of TCDCN compared with existing landmark detection methods on the AFLW and AFW benchmarks are:

    Method AFLW Mean Error (%) AFW Mean Error (%)
    Tree Structured Part Model (TSPM) 15.9 14.3
    Luxand Face SDK 13.1 12.2
    Cascaded Deformable Model (CDM) 13.0 11.1
    Explicit Shape Regression (ESR) 12.4 10.4
    Robust Cascaded Pose Regression (RCPR) 11.6 9.3
    Supervised Descent Method (SDM) 8.5 8.8
    TCDCN (Ours) 8.0 8.2

    TCDCN achieves the lowest overall mean error on both benchmarks, demonstrating robustness against extreme pose variations, partial occlusions, and low input resolution (40×4040 \times 40 pixels).

  10. Knowl 10 — TCDCN Five-Point Landmark Estimation as Robust Initialization for RCPR

    empirical result

    Using TCDCN's 5-landmark estimation as an initial shape estimate for Robust Cascaded Pose Regression (RCPR) replaces random training shape initialization on the Caltech Occluded Faces in the Wild (COFW) dataset (507 test images annotated with 29 landmarks under severe occlusions and head rotations).

    Initializing RCPR with TCDCN predictions yields consistent relative error reductions across all 29 individual landmarks:

    relative improvement=original error−reduced errororiginal error\text{relative improvement} = \frac{\text{original error} - \text{reduced error}}{\text{original error}}

    Relative improvements range from approximately 5% up to 18% across individual landmarks, with the largest gains observed on boundary and jawline landmarks in images with heavy occlusion and rotation.

Coverage note — None was omitted; all contributed methodologies, network architectures, loss formulations, training algorithms, datasets, and experimental findings are represented.

References

  1. 1.Asthana, A., Zafeiriou, S., Cheng, S., Pantic, M.: Robust discriminative response map fitting with constrained local models. In: CVPR, pp. 3444–3451 (2013)
  2. 2.Belhumeur, P.N., Jacobs, D.W., Kriegman, D.J., Kumar, N.: Localizing parts of faces using a consensus of exemplars. In: CVPR, pp. 545–552 (2011)
  3. 3.Burgos-Artizzu, X.P., Perona, P., Dollar, P.: Robust face landmark estimation under occlusion. In: ICCV, pp. 1513–1520 (2013)
  4. 4.Cao, X., Wei, Y., Wen, F., Sun, J.: Face alignment by explicit shape regression. In: CVPR, pp. 2887–2894 (2012)
  5. 5.Caruana, R.: Multitask learning. Machine Learning 28(1), 41–75 (1997)
  6. 6.Chen, K., Gong, S., Xiang, T., Loy, C.C.: Cumulative attribute space for age and crowd density estimation. In: CVPR, pp. 2467–2474 (2013)
  7. 7.Collobert, R., Weston, J.: A unified architecture for natural language processing: Deep neural networks with multitask learning. In: ICML, pp. 160–167 (2008)
  8. 8.Cootes, T.F., Edwards, G.J., Taylor, C.J.: Active appearance models. PAMI 23(6), 681–685 (2001)
  9. 9.Cootes, T.F., Ionita, M.C., Lindner, C., Sauer, P.: Robust and accurate shape model fitting using random forest regression voting. In: Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., Schmid, C. (eds.) ECCV 2012, Part VII. LNCS, vol. 7578, pp. 278–291. Springer, Heidelberg (2012)
  10. 10.Dantone, M., Gall, J., Fanelli, G., Van Gool, L.: Real-time facial feature detection using conditional regression forests. In: CVPR, pp. 2578–2585 (2012)
  11. 11.Kostinger, M., Wohlhart, P., Roth, P.M., Bischof, H.: Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization. In: ICCV Workshops, pp. 2144–2151 (2011)
  12. 12.Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: NIPS (2012)
  13. 13.Li, H., Shen, C., Shi, Q.: Real-time visual tracking using compressive sensing. In: CVPR, pp. 1305–1312 (2011)
  14. 14.Liu, X.: Generic face alignment using boosted appearance model. In: CVPR (2007)
  15. 15.Lu, C., Tang, X.: Surpassing human-level face verification performance on LFW with GaussianFace. Tech. rep., arXiv:1404.3840 (2014)
  16. 16.Luo, P., Wang, X., Tang, X.: Hierarchical face parsing via deep learning. In: CVPR, pp. 2480–2487 (2012)
  17. 17.Luo, P., Wang, X., Tang, X.: A deep sum-product architecture for robust facial attributes analysis. In: CVPR, pp. 2864–2871 (2013)
  18. 18.Luxand Incorporated: Luxand face SDK, http://www.luxand.com/
  19. 19.Nair, V., Hinton, G.E.: Rectified linear units improve restricted boltzmann machines. In: ICML, pp. 807–814 (2010)
  20. 20.Prechelt, L.: Automatic early stopping using cross validation: quantifying the criteria. Neural Networks 11(4), 761–767 (1998)
  21. 21.Sun, Y., Wang, X., Tang, X.: Deep convolutional network cascade for facial point detection. In: CVPR, pp. 3476–3483 (2013)
  22. 22.Sun, Y., Wang, X., Tang, X.: Deep learning face representation by joint identification-verification. Tech. rep., arXiv:1406.4773 (2014)
  23. 23.Sun, Y., Wang, X., Tang, X.: Deep learning face representation from predicting 10,000 classes. In: CVPR (2014)
  24. 24.Valstar, M., Martinez, B., Binefa, X., Pantic, M.: Facial point detection using boosted regression and graph models. In: CVPR, pp. 2729–2736 (2010)
  25. 25.Xiong, X., De La Torre, F.: Supervised descent method and its applications to face alignment. In: CVPR, pp. 532–539 (2013)
  26. 26.Yang, H., Patras, I.: Sieving regression forest votes for facial feature detection in the wild. In: ICCV, pp. 1936–1943 (2013)
  27. 27.Yu, X., Huang, J., Zhang, S., Yan, W., Metaxas, D.N.: Pose-free facial landmark fitting via optimized part mixtures and cascaded deformable shape model. In: ICCV, pp. 1944–1951 (2013)
  28. 28.Yuan, X.T., Liu, X., Yan, S.: Visual classification with multitask joint sparse representation. TIP 21(10), 4349–4360 (2012)
  29. 29.Zhang, T., Ghanem, B., Liu, S., Ahuja, N.: Robust visual tracking via structured multi-task sparse learning. IJCV 101(2), 367–383 (2013)
  30. 30.Zhang, Y., Yeung, D.Y.: A convex formulation for learning task relationships in multi-task learning. In: UAI (2011)
  31. 31.Zhang, Z., Zhang, W., Liu, J., Tang, X.: Facial landmark localization based on hierarchical pose regression with cascaded random ferns. In: ACM Multimedia, pp. 561–564 (2013)
  32. 32.Zhu, X., Ramanan, D.: Face detection, pose estimation, and landmark localization in the wild. In: CVPR, pp. 2879–2886 (2012)
  33. 33.Zhu, Z., Luo, P., Wang, X., Tang, X.: Deep learning identity-preserving face space. In: ICCV, pp. 113–120 (2013)
  34. 34.Zhu, Z., Luo, P., Wang, X., Tang, X.: Deep learning multi-view representation for face recognition. Tech. rep., arXiv:1406.6947 (2014)
  35. 35.Zhu, Z., Luo, P., Wang, X., Tang, X.: Recover canonical-view faces in the wild with deep neural networks. Tech. rep., arXiv:1404.3543 (2014)

Citation

MLA
Zhang, Z., et al. “Facial Landmark Detection by Deep Multi-task Learning”. Lecture Notes in Computer Science, Springer International Publishing, 2014, pp. 94–108, https://doi.org/10.1007/978-3-319-10599-4_7.
APA
Zhang, Z., Luo, P., Loy, C. C., & Tang, X. (2014). Facial Landmark Detection by Deep Multi-task Learning. In Lecture Notes in Computer Science (pp. 94–108). Springer International Publishing. https://doi.org/10.1007/978-3-319-10599-4_7
Chicago
Zhang, Z., P. Luo, C. C. Loy, and X. Tang. 2014. “Facial Landmark Detection by Deep Multi-task Learning”. In Lecture Notes in Computer Science. Springer International Publishing. https://doi.org/10.1007/978-3-319-10599-4_7.
Harvard
Zhang, Z. et al. (2014) “Facial Landmark Detection by Deep Multi-task Learning”, Lecture Notes in Computer Science. Springer International Publishing, pp. 94–108. Available at: https://doi.org/10.1007/978-3-319-10599-4_7.
Vancouver
1. Zhang Z, Luo P, Loy CC, Tang X (2014) Facial Landmark Detection by Deep Multi-task Learning. In: Lecture Notes in Computer Science. Springer International Publishing, pp 94–108

BibTeX

@inbook{Zhang_2014, title={Facial Landmark Detection by Deep Multi-task Learning}, ISBN={9783319105994}, ISSN={1611-3349}, url={http://dx.doi.org/10.1007/978-3-319-10599-4_7}, DOI={10.1007/978-3-319-10599-4_7}, booktitle={Computer Vision – ECCV 2014}, publisher={Springer International Publishing}, author={Zhang, Zhanpeng and Luo, Ping and Loy, Chen Change and Tang, Xiaoou}, year={2014}, pages={94–108} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF