End-To-End Multi-Task Learning With Attention

Shikun LiuEdward JohnsAndrew J. Davison

article2018CVPR1,421 citations

Introduces the Multi-Task Attention Network (MTAN), an architecture that applies task-specific soft-attention masks to a shared feature extractor, achieving state-of-the-art visual performance while reducing sensitivity to loss weighting schemes.

Listen

Modern computer vision systems often require models capable of performing several tasks at once, such as object recognition, depth perception, and surface boundary detection. Deploying independent, dedicated networks for each task creates high computational costs, excessive memory consumption, and slow operational speeds. Conversely, conventional multi-task architectures struggle to balance general feature sharing with specialized task requirements, frequently requiring manual and labor-intensive tuning of loss functions to prevent easier tasks from dominating the training process.

The article demonstrates and evaluates the Multi-Task Attention Network (MTAN), an end-to-end framework designed to efficiently share a compact pool of global features while learning task-specific features via soft attention mechanisms. The primary objective is to prove that MTAN achieves high predictive accuracy across diverse tasks while remaining parameter-efficient and structurally robust to variations in loss-weight balancing schemes.

The researchers evaluated the proposed approach through comparative benchmark experiments across both dense pixel-level prediction and multi-domain image classification. For dense visual tasks—such as semantic segmentation, depth estimation, and surface normal prediction—MTAN was integrated with a standard encoder-decoder architecture and tested on the urban outdoor CityScapes dataset and the complex indoor NYUv2 dataset. For classification, the approach was tested across ten distinct domains within the Visual Decathlon Challenge using a Wide Residual Network backbone. The architecture was benchmarked against single-task networks and multiple competitive multi-task baselines across multiple loss-balancing strategies, including an introduced loss-balancing technique called Dynamic Weight Average (DWA).

The experimental findings show that MTAN achieves state-of-the-art or highly competitive performance while scaling significantly better than existing approaches. On the challenging NYUv2 indoor benchmark, MTAN outperformed all baseline multi-task methods across every task and loss-weighting strategy. Unlike architectures that duplicate network capacity linearly as tasks are added, MTAN requires only about a 10% parameter increase per additional task. Furthermore, the architecture proved substantially less sensitive to the choice of training loss weights, maintaining consistent learning curves across equal, uncertainty-based, and dynamic weighting schemes. The performance advantage over single-task models expanded notably as task complexity increased, with learned attention masks effectively filtering shared representations for task-specific needs.

These results indicate that organizations can deploy unified vision systems that substantially reduce memory footprints and inference latency without sacrificing accuracy. Because MTAN exhibits natural resilience to loss weighting, engineering teams can minimize time-consuming manual hyperparameter tuning during deployment pipelines. This makes the architecture particularly suitable for resource-constrained production settings such as autonomous driving and robotics where multiple concurrent visual perception tasks must operate reliably in real time.

Decision-makers should consider adopting soft attention-based multi-task designs over disjoint model ensembles or heavily duplicated networks when deploying multi-functional computer vision systems. Before full production rollout, engineering teams should conduct pilot implementations on their target hardware platforms to validate operational latency and determine whether simple dynamic weighting methods like DWA are sufficient for their domain. Confidence in these conclusions is high given the consistent empirical performance across varied network backbones and standard datasets, though users should evaluate performance on domain-specific datasets if edge cases fall outside the standard benchmarks evaluated in the article.

Cover for End-To-End Multi-Task Learning With Attention

Abstract

We propose a novel multi-task learning architecture, which allows learning of task-specific feature-level attention. Our design, the Multi-Task Attention Network (MTAN), consists of a single shared network containing a global feature pool, together with a soft-attention module for each task. These modules allow for learning of task-specific features from the global features, whilst simultaneously allowing for features to be shared across different tasks. The architecture can be trained end-to-end and can be built upon any feed-forward neural network, is simple to implement, and is parameter efficient. We evaluate our approach on a variety of datasets, across both image-to-image predictions and image classification tasks. We show that our architecture is state-of-the-art in multi-task learning compared to existing methods, and is also less sensitive to various weighting schemes in the multi-task loss function. Code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Multi-Task Attention Network
  • 3.1 Architecture Design
  • 3.2 Task Specific Attention Module
  • 3.3 The Model Objective
  • 4 Experiments
  • 4.1 Image-to-Image Prediction (One-to-Many)
  • 4.1.1 Datasets
  • 4.1.2 Baselines
  • 4.1.3 Dynamic Weight Average
  • 4.1.4 Results on Image-to-Image Predictions
  • 4.1.5 Effect of Task Complexity
  • 4.1.6 Attention Masks as Feature Selectors
  • 4.2 Visual Decathlon Challenge (Many-to-Many)
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — Multi-Task Attention Network Architecture

    model/method

    The Multi-Task Attention Network (MTAN) is a multi-task learning architecture designed to learn both task-shared and task-specific visual representations end-to-end within feed-forward convolutional neural networks. MTAN consists of two primary components:

    1. A Single Shared Backbone Network: Computes a global feature pool across all layers for all tasks.
    2. KK Task-Specific Attention Networks: Each task i∈{1,…,K}i \in \{1, \dots, K\} possesses its own dedicated stream composed of attention modules attached to each convolutional block of the shared backbone.

    Rather than executing hard parameter sharing (where representations fork only at final task-specific heads) or using duplicate feed-forward networks with cross-stitch units, each MTAN attention module extracts task-specific representations by applying learned soft attention masks directly to feature maps extracted by the shared backbone. By sharing a single global feature pool and routing task-specific activations via attention masks, the parameter overhead scales gracefully with the number of tasks KK (requiring approximately a 10%10\% parameter increase per additional task).

  2. Knowl 2 — Task-Specific Attention Module Formulation in MTAN

    equation

    In MTAN, let p(j)p^{(j)} denote the shared feature tensor generated at block jj of the shared network. For a given task ii, let ai(j)a^{(j)}_i be the soft attention mask computed at block jj. The task-specific attended feature representation a^i(j)\hat{a}^{(j)}_i is obtained via channel-wise element-wise multiplication ⊙\odot:

    a^i(j)=ai(j)⊙p(j)\hat{a}^{(j)}_i = a^{(j)}_i \odot p^{(j)}

    For the first encoder block (j=1j = 1), the attention mask is generated exclusively from the initial shared features. For subsequent blocks (j≥2j \ge 2), the attention mask is recursively generated from both current shared features u(j)u^{(j)} and previous task-specific attended features a^i(j−1)\hat{a}^{(j-1)}_i:

    ai(j)=hi(j)(gi(j)([u(j);f(j)(a^i(j−1))])),j≥2a^{(j)}_i = h^{(j)}_i\Big( g^{(j)}_i\big( [u^{(j)}; f^{(j)}(\hat{a}^{(j-1)}_i)] \big) \Big), \quad j \ge 2

    where:

    • [⋅;⋅][\cdot ; \cdot] denotes feature map concatenation along the channel dimension.
    • f(j)f^{(j)} is a shared feature extractor consisting of 3×33 \times 3 convolutions, Batch Normalization, and ReLU, followed by a spatial pooling layer (in the encoder) or an upsampling layer (in the decoder) to match spatial dimensions.
    • gi(j)g^{(j)}_i and hi(j)h^{(j)}_i are task-specific convolutional layers with 1×11 \times 1 kernels, Batch Normalization, and ReLU activations, with hi(j)h^{(j)}_i ending with a Sigmoid activation function ensuring ai(j)∈[0,1]a^{(j)}_i \in [0, 1].

    If ai(j)→1a^{(j)}_i \to 1, the attention mask approaches an identity operator, allowing complete feature sharing; intermediate values allow the network to dynamically filter features specifically required by task ii.

  3. Knowl 3 — Dynamic Weight Average for Multi-Task Loss Balancing

    model/method

    Dynamic Weight Average (DWA) is an adaptive loss-weighting method that balances learning rates across KK tasks by tracking the relative rate of change of each task's loss over training epochs. Unlike gradient-balancing methods that require inspecting internal layer gradients, DWA operates exclusively on scalar task loss histories.

    Let Lk(t)\mathcal{L}_k(t) be the average training loss value for task kk at epoch/iteration tt. The relative loss descending rate wk(t−1)∈(0,+∞)w_k(t-1) \in (0, +\infty) is calculated as:

    wk(t−1)=Lk(t−1)Lk(t−2)w_k(t-1) = \frac{\mathcal{L}_k(t-1)}{\mathcal{L}_k(t-2)}

    At iteration tt, the dynamic weight λk(t)\lambda_k(t) assigned to task kk is calculated using a temperature-scaled softmax normalization multiplied by the total task count KK:

    λk(t):=Kexp⁡(wk(t−1)/T)∑i=1Kexp⁡(wi(t−1)/T)\lambda_k(t) := \frac{K \exp\big(w_k(t-1) / T\big)}{\sum_{i=1}^K \exp\big(w_i(t-1) / T\big)}

    where T>0T > 0 is a temperature hyperparameter that controls the sharpness of task prioritization. A larger TT smooths the distribution toward equal weighting (as T→∞T \to \infty, λk→1\lambda_k \to 1). The weights satisfy ∑k=1Kλk(t)=K\sum_{k=1}^K \lambda_k(t) = K. For the initial steps t=1,2t=1, 2, weights are initialized uniformly to wk(t)=1w_k(t) = 1.

  4. Knowl 4 — Loss Objectives for Multi-Task Dense Scene Understanding

    equation

    For multi-task dense scene understanding involving semantic segmentation, depth estimation, and surface normal prediction over an input image XX with spatial resolution p×qp \times q, the overall training loss is defined as:

    Ltot(X,Y1:K)=∑i=1KλiLi(X,Yi)\mathcal{L}_{tot}(X, Y_{1:K}) = \sum_{i=1}^K \lambda_i \mathcal{L}_i(X, Y_i)

    where each task loss Li\mathcal{L}_i is defined with respect to ground-truth label map YiY_i and predicted output map Y^i\hat{Y}_i:

    1. Semantic Segmentation Loss: Pixel-wise cross-entropy over spatial positions (p,q)(p, q):

    L1(X,Y1)=−1pq∑p,qY1(p,q)log⁡Y^1(p,q)\mathcal{L}_1(X, Y_1) = -\frac{1}{pq} \sum_{p, q} Y_1(p, q) \log \hat{Y}_1(p, q)

    1. Depth Estimation Loss: Mean absolute error (L1L_1 norm), applied on metric true depth for indoor scenes or inverse depth for outdoor scenes:

    L2(X,Y2)=1pq∑p,q∣Y2(p,q)−Y^2(p,q)∣\mathcal{L}_2(X, Y_2) = \frac{1}{pq} \sum_{p, q} |Y_2(p, q) - \hat{Y}_2(p, q)|

    1. Surface Normal Prediction Loss: Negative dot product across unit-normalized 3D vector predictions:

    L3(X,Y3)=−1pq∑p,qY3(p,q)⋅Y^3(p,q)\mathcal{L}_3(X, Y_3) = -\frac{1}{pq} \sum_{p, q} Y_3(p, q) \cdot \hat{Y}_3(p, q)

  5. Knowl 5 — Multi-Task Dense Prediction Performance on NYUv2

    data/table

    The performance of MTAN built upon a SegNet backbone was evaluated on the 3-task NYUv2 validation set (13-class semantic segmentation, true depth estimation, and surface normal prediction) against single-task and multi-task baselines under Equal Weights, Weight Uncertainty, and Dynamic Weight Average (DWA, T=2T=2). Metric units: Segmentation mIoU (%), Pixel Accuracy (%); Depth Absolute Error (m), Relative Error; Surface Normal Angle Distance Mean/Median (degrees), and percentage within angle thresholds t∘t^\circ.

    Type #P Architecture Weighting Segmentation (↑\uparrow) Depth (↓\downarrow) Angle Dist. (↓\downarrow) Within t∘t^\circ (↑\uparrow)
    mIoU Pix Acc Abs Err Rel Err Mean Median 11.25 22.5 30
    Single 3 One Task n.a. 15.10 51.54 0.7508 0.3266 31.76 25.51 22.12 45.33 57.13
    Single 4.56 STAN n.a. 15.73 52.89 0.6935 0.2891 32.09 26.32 21.49 44.38 56.51
    Multi 1.75 Split, Wide Equal 15.89 51.19 0.6494 0.2804 33.69 28.91 18.54 39.91 52.02
    Multi 1.75 Split, Wide Uncert. 15.86 51.12 0.6040 0.2570 32.33 26.62 21.68 43.59 55.36
    Multi 1.75 Split, Wide DWA 16.92 53.72 0.6125 0.2546 32.34 27.10 20.69 42.73 54.74
    Multi 2 Split, Deep Equal 13.03 41.47 0.7836 0.3326 38.28 36.55 9.50 27.11 39.63
    Multi 2 Split, Deep Uncert. 14.53 43.69 0.7705 0.3340 35.14 32.13 14.69 34.52 46.94
    Multi 2 Split, Deep DWA 13.63 44.41 0.7581 0.3227 36.41 34.12 12.82 31.12 43.48
    Multi 4.95 Dense Equal 16.06 52.73 0.6488 0.2871 33.58 28.01 20.07 41.50 53.35
    Multi 4.95 Dense Uncert. 16.48 54.40 0.6282 0.2761 31.68 25.68 21.73 44.58 56.65
    Multi 4.95 Dense DWA 16.15 54.35 0.6059 0.2593 32.44 27.40 20.53 42.76 54.27
    Multi ≈3\approx 3 Cross-Stitch Equal 14.71 50.23 0.6481 0.2871 33.56 28.58 20.08 40.54 51.97
    Multi ≈3\approx 3 Cross-Stitch Uncert. 15.69 52.60 0.6277 0.2702 32.69 27.26 21.63 42.84 54.45
    Multi ≈3\approx 3 Cross-Stitch DWA 16.11 53.19 0.5922 0.2611 32.34 26.91 21.81 43.14 54.92
    Multi 1.77 MTAN (Ours) Equal 17.72 55.32 0.5906 0.2577 31.44 25.37 23.17 45.65 57.48
    Multi 1.77 MTAN (Ours) Uncert. 17.67 55.61 0.5927 0.2592 31.25 25.57 22.99 45.83 57.67
    Multi 1.77 MTAN (Ours) DWA 17.15 54.97 0.5956 0.2569 31.60 25.46 22.48 44.86 57.24

    MTAN achieves superior performance across all 3 tasks with only 1.77×1.77\times the parameters of a single-task network, outperforming Dense (4.95×4.95\times parameters) and Cross-Stitch (3×3\times parameters).

  6. Knowl 6 — Multi-Task Dense Prediction Performance on CityScapes

    data/table

    Evaluation of MTAN and baseline architectures based on SegNet on the CityScapes validation dataset (7-class semantic segmentation and inverse depth estimation) using Equal Weights, Weight Uncertainty, and Dynamic Weight Average (DWA, T=2T=2). #P denotes parameter count relative to a vanilla single-task network.

    #P Architecture Weighting Segmentation (↑\uparrow) Depth (↓\downarrow)
    mIoU Pix Acc Abs Err Rel Err
    2 One Task n.a. 51.09 90.69 0.0158 34.17
    3.04 STAN n.a. 51.90 90.87 0.0145 27.46
    1.75 Split, Wide Equal Weights 50.17 90.63 0.0167 44.73
    1.75 Split, Wide Uncert. Weights 51.21 90.72 0.0158 44.01
    1.75 Split, Wide DWA, T=2T=2 50.39 90.45 0.0164 43.93
    2 Split, Deep Equal Weights 49.85 88.69 0.0180 43.86
    2 Split, Deep Uncert. Weights 48.12 88.68 0.0169 39.73
    2 Split, Deep DWA, T=2T=2 49.67 88.81 0.0182 46.63
    3.63 Dense Equal Weights 51.91 90.89 0.0138 27.21
    3.63 Dense Uncert. Weights 51.89 91.22 0.0134 25.36
    3.63 Dense DWA, T=2T=2 51.78 90.88 0.0137 26.67
    ≈2\approx 2 Cross-Stitch Equal Weights 50.08 90.33 0.0154 34.49
    ≈2\approx 2 Cross-Stitch Uncert. Weights 50.31 90.43 0.0152 31.36
    ≈2\approx 2 Cross-Stitch DWA, T=2T=2 50.33 90.55 0.0153 33.37
    1.65 MTAN (Ours) Equal Weights 53.04 91.11 0.0144 33.63
    1.65 MTAN (Ours) Uncert. Weights 53.86 91.10 0.0144 35.72
    1.65 MTAN (Ours) DWA, T=2T=2 53.29 91.09 0.0144 34.14

    MTAN yields the highest segmentation performance (53.86 mIoU) and matches Dense depth estimation accuracy while using less than half of Dense's parameter count (1.65 vs 3.63).

  7. Knowl 7 — Visual Decathlon Challenge Classification Performance

    data/table

    Evaluation of MTAN based on a Wide Residual Network (WRN-28-4) on the Visual Decathlon Challenge online test set comprising 10 distinct image classification domains. #P denotes parameter count relative to a single-task network. Accuracy is reported per dataset alongside the overall mean accuracy (%) and total Decathlon score (maximum 10,000; 1,000 points per dataset).

    Method #P ImNet Airc. C100 DPed DTD GTSR Flwr Oglt SVHN UCF Mean Score
    Scratch 10 59.87 57.10 75.73 91.20 37.77 96.55 56.30 88.74 96.63 43.27 70.32 1625
    Finetune 10 59.87 60.34 82.12 92.82 55.53 97.53 81.41 87.69 96.55 51.20 76.51 2500
    Feature 1 59.67 23.31 63.11 80.33 45.37 68.16 73.69 58.79 43.54 26.80 54.28 544
    Res. Adapt. 2 59.67 56.68 81.20 93.88 50.85 97.05 66.24 89.62 96.13 47.45 73.88 2118
    DAN 2.17 57.74 64.12 80.07 91.30 56.54 98.46 86.05 89.67 96.77 49.38 77.01 2851
    Piggyback 1.28 57.69 65.29 79.87 96.99 57.45 97.27 79.09 87.63 97.24 47.48 76.60 2838
    Parallel SVD 1.5 60.32 66.04 81.86 94.23 57.82 99.24 85.74 89.25 96.62 52.50 78.36 3398
    MTAN (Ours) 1.74 63.90 61.81 81.59 91.63 56.44 98.80 81.04 89.83 96.88 50.63 77.25 2941

    MTAN achieves a cumulative score of 2941 and a mean accuracy of 77.25% across the 10 domains with 1.74×1.74\times baseline parameters, without requiring dataset-specific regularization techniques such as DropOut, dataset grouping, or per-dataset adaptive weight decay.

  8. Knowl 8 — Scaling Behavior and Task Complexity in Multi-Task Learning

    empirical result

    When comparing performance gains across varying semantic segmentation complexity levels (2-class, 7-class, and 19-class setups paired with inverse depth estimation on CityScapes):

    1. For low-complexity tasks (2 classes: foreground vs. background), the Single-Task Attention Network (STAN) outperforms multi-task learning baselines, as a single task can utilize network capacity directly without cross-task interference.
    2. As task complexity increases to 7 and 19 classes, multi-task models systematically outperform single-task baselines because sharing visual feature representations provides regularizing benefits and improves parameter utilization.
    3. Across all evaluated configurations, the performance gain of MTAN grows at a higher rate with task complexity than Split, Dense, or Cross-Stitch baselines.
  9. Knowl 9 — Robustness of MTAN to Loss Weighting Strategies

    empirical result

    In multi-task optimization experiments across three loss-weighting regimes (Equal Weights, Weight Uncertainty, and Dynamic Weight Average):

    • Baseline architectures (such as Cross-Stitch Networks) exhibit wide discrepancies in optimization dynamics and validation metrics depending on the chosen weighting scheme.
    • MTAN maintains consistent training and validation trajectories across all three weighting schemes for segmentation, depth estimation, and surface normal estimation on NYUv2.
    • This indicates that task-specific soft attention masks provide structural robustness against suboptimal or un-tuned multi-task loss weightings.

Coverage note — None was omitted; all core architectural definitions, loss formulations, adaptive weighting algorithms, quantitative benchmark tables (NYUv2, CityScapes, Visual Decathlon), and empirical findings on task complexity and loss robustness were captured.

References

  1. 1.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE transactions on pattern analysis and machine intelligence, 39(12):2481–2495, 2017.
  2. 2.Rich Caruana. Multitask learning. In Learning to learn, pages 95–133. Springer, 1998.
  3. 3.Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich. Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In International Conference on Machine Learning, pages 793–802, 2018.
  4. 4.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3213–3223, 2016.
  5. 5.Camille Couprie, Clement Farabet, Laurent Najman, and ´ Yann Lecun. Indoor semantic segmentation using depth information. In International Conference on Learning Representations (ICLR2013), April 2013, 2013.
  6. 6.Carl Doersch and Andrew Zisserman. Multi-task self-supervised visual learning. In The IEEE International Conference on Computer Vision (ICCV), Oct 2017.
  7. 7.David Eigen and Rob Fergus. Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. In Proceedings of the IEEE International Conference on Computer Vision, pages 2650–2658, 2015.
  8. 8.Theodoros Evgeniou and Massimiliano Pontil. Regularized multi–task learning. In Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining, pages 109–117. ACM, 2004.
  9. 9.Georgia Gkioxari, Bharath Hariharan, Ross Girshick, and Jitendra Malik. R-cnns for pose estimation and action detection. arXiv preprint arXiv:1406.5212, 2014.
  10. 10.Michelle Guo, Albert Haque, De-An Huang, Serena Yeung, and Li Fei-Fei. Dynamic task prioritization for multi-task learning. In European Conference on Computer Vision, pages 282–299. Springer, 2018.
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  12. 12.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.
  13. 13.Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European Conference on Computer Vision, pages 694–711. Springer, 2016.
  14. 14.Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7482–7491, 2018.
  15. 15.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  16. 16.Iasonas Kokkinos. Ubernet: Training a universal convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
  17. 17.Abhishek Kumar and Hal Daume III. Learning task grouping ´ and overlap in multi-task learning. In Proceedings of the 29th International Coference on International Conference on Machine Learning, pages 1723–1730. Omnipress, 2012.
  18. 18.Mingsheng Long, Jianmin Wang, Guiguang Ding, Jiaguang Sun, and Philip S Yu. Transfer feature learning with joint distribution adaptation. In Proceedings of the IEEE international conference on computer vision, pages 2200–2207, 2013.
  19. 19.Arun Mallya, Dillon Davis, and Svetlana Lazebnik. Piggyback: Adapting a single network to multiple tasks by learning to mask weights. In Proceedings of the European Conference on Computer Vision (ECCV), pages 67–82, 2018.
  20. 20.Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, and Martial Hebert. Cross-stitch networks for multi-task learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3994–4003, 2016.
  21. 21.Pushmeet Kohli Nathan Silberman, Derek Hoiem and Rob Fergus. Indoor segmentation and support inference from rgbd images. In ECCV, 2012.
  22. 22.Sinno Jialin Pan and Qiang Yang. A survey on transfer learning. IEEE Transactions on knowledge and data engineering, 22(10):1345–1359, 2010.
  23. 23.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. Learning multiple visual domains with residual adapters. In Advances in Neural Information Processing Systems, pages 506–516, 2017.
  24. 24.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. Efficient parametrization of multi-domain deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8119–8127, 2018.
  25. 25.Amir Rosenfeld and John K Tsotsos. Incremental learning through deep adaptation. IEEE transactions on pattern analysis and machine intelligence, 2018.
  26. 26.Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. arXiv preprint arXiv:1606.04671, 2016.
  27. 27.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  28. 28.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research, 15(1):1929–1958, 2014.
  29. 29.Sebastian Thrun and Lorien Pratt. Learning to learn. Springer Science & Business Media, 2012.
  30. 30.Fei Wang, Mengqing Jiang, Chen Qian, Shuo Yang, Cheng Li, Honggang Zhang, Xiaogang Wang, and Xiaoou Tang. Residual attention network for image classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3156–3164, 2017.
  31. 31.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. In Edwin R. Hancock Richard C. Wilson and William A. P. Smith, editors, Proceedings of the British Machine Vision Conference (BMVC), pages 87.1–87.12. BMVA Press, September 2016.

Citation

MLA
Liu, S., et al. “End-to-End Multi-Task Learning with Attention”. arXiv, 2018, http://arxiv.org/abs/1803.10704v2.
APA
Liu, S., Johns, E., & Davison, A. J. (2018). End-to-End Multi-Task Learning with Attention. arXiv. http://arxiv.org/abs/1803.10704v2
Chicago
Liu, S., E. Johns, and A. J. Davison. 2018. “End-to-End Multi-Task Learning with Attention”. arXiv. http://arxiv.org/abs/1803.10704v2.
Harvard
Liu, S., Johns, E. and Davison, A.J. (2018) “End-to-End Multi-Task Learning with Attention”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1803.10704v2.
Vancouver
1. Liu S, Johns E, Davison AJ (2018) End-to-End Multi-Task Learning with Attention. arXiv

BibTeX

@article{liu2018end,
  title = {End-to-End Multi-Task Learning with Attention},
  author = {Liu, Shikun and Johns, Edward and Davison, Andrew J.},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1803.10704v2},
  eprint = {1803.10704}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE