Learning Multiple Dense Prediction Tasks from Partially Annotated Data

Wei-Hong LiXialei LiuHakan Bilen

article2022CVPR51 citationsBest Paper Finalist

Presents a label-efficient framework for multi-task dense prediction that learns joint pairwise task spaces to supervise missing annotations across images through cross-task consistency.

Listen

Deploying computer vision models for complex visual understanding typically requires learning multiple pixel-level tasks simultaneously, such as segmenting objects, estimating depth, and detecting surface orientations. Standard multi-task learning frameworks depend on training images that are fully annotated across all target tasks. In real-world applications, however, acquiring complete annotations for every image is prohibitively expensive, operationally challenging due to multi-sensor synchronization issues, and often impossible when integrating new tasks into existing legacy datasets. Consequently, real-world systems must operate in partially annotated environments where images contain labels for only a subset of tasks.

The article evaluates and demonstrates a multi-task partially supervised learning framework designed to train unified vision models when training images are missing labels for one or more tasks. The core objective is to leverage inherent cross-task relationships to supervise unlabelled tasks without requiring complex, direct label transformations or massive fully labelled datasets.

To address this challenge, the authors designed a multi-task learning approach that maps task pairs into a shared, lower-dimensional joint feature space to enforce consistency between the predictions of unlabelled tasks and the ground-truth labels of available tasks. To ensure computational scalability as the number of tasks grows, the framework uses a single encoder network that dynamically adjusts its internal weights for specific task pairs via conditioning. Additionally, the training process includes a regularization term tied to the core image features to prevent the mapping function from collapsing into meaningless, trivial shortcuts. The authors validated this approach on three standard visual benchmarks—Cityscapes, NYU-v2, and PASCAL-Context—across diverse data regimes, including random task availability, single-label-per-image constraints, and heavily imbalanced task ratios.

The experiments show that the proposed framework consistently outperforms standard supervised and semi-supervised baselines under partial supervision. In partial annotation scenarios on the Cityscapes benchmark, the method achieved a segmentation intersection-over-union score of 74.90%, outperforming the standard supervised baseline trained only on partial labels (69.50%) as well as the fully supervised model trained with 100% complete labels across all tasks (73.36%). On the indoor NYU-v2 dataset, when restricted to only one label per image, the framework significantly outperformed the partial supervised baseline across all metrics, raising segmentation accuracy from 25.75% to 30.36% while reducing depth and surface normal errors. Furthermore, in severely imbalanced conditions where 90% of annotations were missing for a given task, the method maintained high accuracy across all outputs, demonstrating strong data efficiency. Finally, when applied to fully supervised settings and combined with adaptive loss weighting, the model surpassed established multi-task benchmarks.

These findings indicate that learning joint pairwise relationships across tasks allows models to extract valuable cross-task supervision from unlabelled visual data. Operationally, this capability reduces data labelling costs, lowers sensor hardware requirements, and accelerates the development timelines needed to expand vision systems with new capabilities. By removing the requirement that all training images must be synchronized across every sensor modality, organizations can unlock and repurpose partially labelled historical data.

Organizations developing multi-task vision systems should adopt joint pairwise consistency learning rather than relying solely on single-task data augmentation or conventional semi-supervised pipelines. When deploying this method, engineering teams should pair the consistency framework with adaptive loss-weighting mechanisms to maximize performance balance across uneven task distributions. Future work should focus on developing automated mechanisms to identify which specific task pairs share meaningful correlations, avoiding the computational overhead of training mappings between unrelated tasks.

The primary limitation of the article is that cross-task relationships were evaluated across all possible task pairs, even though some pairs may offer minimal complementary information. Furthermore, performance improvements depend on the visual domain and the degree of structural correlation between tasks, showing smaller relative gains on diverse image datasets where cross-task links are weaker. Confidence in the core results remains high across standard dense prediction benchmarks, though pilot testing is recommended before applying the framework to domains with unverified task affinities.

  • Paper: Taskonomy: Disentangling Task Transfer Learning, Amir Zamir et al. (2018). It establishes the foundational computational framework for mapping and exploiting cross-task relationships across dense visual predictions, which directly motivates the source's pairwise multi-task transfer.
  • Paper: End-To-End Multi-Task Learning With Attention, Shikun Liu et al. (2018). It introduces shared-encoder multi-task attention architectures and dynamic loss weighting on dense prediction benchmarks like Cityscapes and NYUv2 that the source builds upon.
  • Paper: Multi-task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics, Alex Kendall et al. (2017). It presents the seminal approach for joint learning and principled homoscedastic uncertainty loss weighting across multi-task geometry and dense semantic predictions.
  • Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). It analyzes multi-task gradient conflict dynamics and provides foundational optimization techniques essential for stabilizing joint multi-task representation learning.
  • Paper: Multi-Task Learning as Multi-Objective Optimization, Ozan Sener et al. (2018). It establishes multi-task dense prediction as a multi-objective optimization problem, providing the mathematical context for balancing competing task objectives.
  • Paper: Cross-Stitch Networks for Multi-task Learning, Ishan Misra et al. (2016). It provides the foundational framework for dynamically learning how to share representations across task pairs in multi-task visual learning.
  • Paper: Convex multi-task feature learning, Andreas Argyriou et al. (2008). It formulates the core theoretical principles of learning low-dimensional shared feature representations across multiple related tasks.
Cover for Learning Multiple Dense Prediction Tasks from Partially Annotated Data

Abstract

Despite the recent advances in multi-task learning of dense prediction problems, most methods rely on expensive labelled datasets. In this paper, we present a label efficient approach and look at jointly learning of multiple dense prediction tasks on partially annotated data (i.e. not all the task labels are available for each image), which we call multi-task partially-supervised learning. We propose a multi-task training procedure that successfully leverages task relations to supervise its multi-task learning when data is partially annotated. In particular, we learn to map each task pair to a joint pairwise task-space which enables sharing information between them in a computationally efficient way through another network conditioned on task pairs, and avoids learning trivial cross-task relations by retaining high-level information about the input image. We rigorously demonstrate that our proposed method effectively exploits the images with unlabelled tasks and outperforms existing semi-supervised learning approaches and related methods on three standard benchmarks.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Problem setting
  • 3.2. Cross-task consistency learning
  • 4. Experiments
  • 4.1. Results
  • 4.2. Further results
  • 4.3. Ablation study
  • 4.4. Qualitative results
  • 5. Conclusion and Limitations
  • References

Knowls

  1. Knowl 1 — Multi-Task Partially Supervised Learning Problem Formulation

    definition

    Multi-Task Partially Supervised Learning (MTPSL) is the problem of training a shared multi-task network on a dataset D={(xn,{ynt}t∈Tn)}n=1N\mathcal{D} = \{(\mathbf{x}_n, \{\mathbf{y}_n^t\}_{t \in \mathcal{T}_n})\}_{n=1}^N of NN images, where each image xn∈R3×H×W\mathbf{x}_n \in \mathbb{R}^{3 \times H \times W} is annotated with ground-truth dense labels ynt∈ROt×H×W\mathbf{y}_n^t \in \mathbb{R}^{O_t \times H \times W} for only a non-empty subset of tasks Tn⊂{1,…,K}\mathcal{T}_n \subset \{1, \dots, K\}, while the remaining tasks Un={1,…,K}∖Tn\mathcal{U}_n = \{1, \dots, K\} \setminus \mathcal{T}_n are unlabelled.

    Here KK denotes the total number of dense prediction tasks, OtO_t is the channel dimensionality of the label for task tt, and each image satisfies ∣Tn∣≥1|\mathcal{T}_n| \ge 1 and ∣Tn∣+∣Un∣=K|\mathcal{T}_n| + |\mathcal{U}_n| = K. A shared feature encoder fϕ:R3×H×W→RC×H′×W′f_\phi: \mathbb{R}^{3 \times H \times W} \to \mathbb{R}^{C \times H' \times W'} maps an input image to feature maps, and task-specific decoders hψt:RC×H′×W′→ROt×H×Wh_{\psi^t}: \mathbb{R}^{C \times H' \times W'} \to \mathbb{R}^{O_t \times H \times W} produce predictions y^t(x)=(hψt∘fϕ)(x)\hat{\mathbf{y}}^t(\mathbf{x}) = (h_{\psi^t} \circ f_\phi)(\mathbf{x}) for each task t∈{1,…,K}t \in \{1, \dots, K\}.

  2. Knowl 2 — Task-Pair Conditional Mapping Network

    model/method

    To project predictions and labels of any task pair (s,t)(s, t) into a shared DD-dimensional joint pairwise task-space without allocating O(K2)O(K^2) distinct neural networks, a single task-agnostic backbone mapping network mˉϑ\bar{m}_\vartheta is modulated via feature-wise linear modulation (FiLM) conditioned on the task pair.

    Let A∈{0,1}K×K\mathbf{A} \in \{0, 1\}^{K \times K} be an asymmetric indicator matrix encoding the directed task interaction, where A[s,t]=1\mathbf{A}[s, t] = 1 (with zeros elsewhere and A[k,k]=0\mathbf{A}[k, k] = 0 for all kk). A light-weight auxiliary network aθa_\theta consisting of single-layer fully connected heads takes A\mathbf{A} as input and produces channel-wise scaling vectors aθ,ic(A)∈RM\mathbf{a}^c_{\theta, i}(\mathbf{A}) \in \mathbb{R}^M and bias vectors aθ,ib(A)∈RM\mathbf{a}^b_{\theta, i}(\mathbf{A}) \in \mathbb{R}^M for each intermediate layer ii of mˉϑ\bar{m}_\vartheta.

    The intermediate feature map hi∈RM×Hi×Wi\mathbf{h}_i \in \mathbb{R}^{M \times H_i \times W_i} at layer ii is transformed according to:

    hi←aθ,ic(A)⊙hi+aθ,ib(A)\mathbf{h}_i \leftarrow \mathbf{a}^c_{\theta, i}(\mathbf{A}) \odot \mathbf{h}_i + \mathbf{a}^b_{\theta, i}(\mathbf{A})

    where ⊙\odot denotes the Hadamard product along the channel dimension. The resulting conditioned mapping function is denoted ms→st(⋅)m^{s \to st}(\cdot) for source prediction y^s\hat{\mathbf{y}}^s and mt→st(⋅)m^{t \to st}(\cdot) for target ground truth yt\mathbf{y}^t, each beginning with a task-specific 1-layer convolutional input adapter to match channel dimension OtO_t before shared processing.

  3. Knowl 3 — Feature-Regularized Multi-Task Partially Supervised Training Objective

    equation

    The parameters ϕ\phi (encoder), ψ={ψt}t=1K\psi = \{\psi^t\}_{t=1}^K (task decoders), ϑ\vartheta (task-agnostic mapping network), and θ\theta (conditioning network) are jointly optimized over NN partially annotated samples by minimizing the total loss:

    min⁡ϕ,ψ,ϑ,θ1N∑n=1N(1∣Tn∣∑t∈TnLt(y^t(xn),ynt)+1∣Un∣∑s∈Un,t∈TnLct(ms→st(y^s(xn)),mt→st(ynt))+R(fϕ(xn),ms→st(y^s(xn)))+R(fϕ(xn),mt→st(ynt)))\min_{\phi, \psi, \vartheta, \theta} \frac{1}{N} \sum_{n=1}^N \Bigg( \frac{1}{|\mathcal{T}_n|} \sum_{t \in \mathcal{T}_n} L^t\big(\hat{\mathbf{y}}^t(\mathbf{x}_n), \mathbf{y}_n^t\big) + \frac{1}{|\mathcal{U}_n|} \sum_{s \in \mathcal{U}_n, t \in \mathcal{T}_n} L_{ct}\big(m^{s \to st}(\hat{\mathbf{y}}^s(\mathbf{x}_n)), m^{t \to st}(\mathbf{y}_n^t)\big) + R\big(f_\phi(\mathbf{x}_n), m^{s \to st}(\hat{\mathbf{y}}^s(\mathbf{x}_n))\big) + R\big(f_\phi(\mathbf{x}_n), m^{t \to st}(\mathbf{y}_n^t)\big) \Bigg)

    where:

    • LtL^t is the supervised task-specific loss (e.g., cross-entropy for segmentation, ℓ1\ell_1 norm for depth, cosine loss for surface normals).
    • Lct(a,b)=1−a⋅b∥a∥∥b∥L_{ct}(\mathbf{a}, \mathbf{b}) = 1 - \frac{\mathbf{a} \cdot \mathbf{b}}{\|\mathbf{a}\| \|\mathbf{b}\|} is the cosine distance penalizing cross-task inconsistency in the joint latent space between the unlabelled task prediction y^s(xn)\hat{\mathbf{y}}^s(\mathbf{x}_n) and the labelled task ground truth ynt\mathbf{y}_n^t.
    • R(u,v)=1−u⋅v∥u∥∥v∥R(\mathbf{u}, \mathbf{v}) = 1 - \frac{\mathbf{u} \cdot \mathbf{v}}{\|\mathbf{u}\| \|\mathbf{v}\|} regularizes the mapping output against the encoder representation fϕ(xn)f_\phi(\mathbf{x}_n) to prevent representation collapse to constant or zero vectors.
  4. Knowl 4 — Multi-Task Learning Performance on Cityscapes and NYU-v2 under Partial Annotations

    data/table

    The multi-task partially supervised learning (MTPSL) framework was evaluated against fully supervised learning (SL full), partially supervised learning (SL partial), semi-supervised learning (SSL with perturbation consistency), Direct-Map, Perceptual-Map, Contrastive-Loss, and Discriminator-Loss on Cityscapes (7-class semantic segmentation and depth estimation) and NYU-v2 (13-class semantic segmentation, depth estimation, and surface normal estimation) using a SegNet backbone.

    Evaluation settings include:

    • one: exactly one randomly chosen task label is available per training image.
    • random: a random subset of 11 to K−1K-1 task labels is available per training image.
    Dataset Method Seg. (mIoU) ↑\uparrow Depth (aErr) ↓\downarrow Norm. (mErr) ↓\downarrow
    Cityscapes (one) Supervised Learning (full) 73.36 0.0165 -
    Supervised Learning 69.50 0.0186 -
    Semi-supervised Learning 71.67 0.0178 -
    Perceptual-Map 72.82 0.0169 -
    Direct-Map 72.33 0.0179 -
    Contrastive-Loss 71.79 0.0183 -
    Discriminator-Loss 68.94 0.0208 -
    MTPSL (Ours) 74.90 0.0161 -
    NYU-v2 (full) Supervised Learning (full) 36.95 0.5510 29.51
    NYU-v2 (random) Supervised Learning 27.05 0.6624 33.58
    Semi-supervised Learning 29.50 0.6224 33.31
    Perceptual-Map 32.20 0.6037 32.07
    Direct-Map 29.17 0.6128 33.63
    MTPSL (Ours) 34.26 0.5787 31.06
    NYU-v2 (one) Supervised Learning 25.75 0.6511 33.73
    Semi-supervised Learning 27.52 0.6499 33.58
    Perceptual-Map 26.94 0.6342 34.30
    Direct-Map 19.98 0.6960 37.56
    MTPSL (Ours) 30.36 0.6088 32.08

    MTPSL outperforms all partial-supervision baselines across all metrics. On Cityscapes (one label), MTPSL achieves 74.90% mIoU and 0.0161 aErr, outperforming even full supervision (73.36% mIoU, 0.0165 aErr).

  5. Knowl 5 — Dense Multi-Task Prediction under Severe Task Label Imbalance

    empirical result

    When evaluated on Cityscapes under imbalanced annotation ratios between semantic segmentation and depth estimation (where one task has annotations for 90% of images and the other for only 10%), the proposed joint pairwise mapping method consistently outperforms supervised learning (SL), semi-supervised learning (SSL), Direct-Map, and Perceptual-Map.

    Specifically, the performance across settings is:

    1. Segmentation : Depth = 1 : 9 (10% segmentation, 90% depth annotations):

      • Supervised Learning: 63.37 mIoU, 0.0161 aErr
      • Semi-Supervised Learning: 64.40 mIoU, 0.0179 aErr
      • Perceptual-Map: 68.84 mIoU, 0.0141 aErr
      • Direct-Map: 67.04 mIoU, 0.0153 aErr
      • MTPSL (Ours): 71.89 mIoU, 0.0131 aErr
    2. Segmentation : Depth = 9 : 1 (90% segmentation, 10% depth annotations):

      • Supervised Learning: 72.77 mIoU, 0.0250 aErr
      • Semi-Supervised Learning: 72.97 mIoU, 0.0395 aErr
      • Perceptual-Map: 73.36 mIoU, 0.0237 aErr
      • Direct-Map: 73.13 mIoU, 0.0288 aErr
      • MTPSL (Ours): 74.23 mIoU, 0.0235 aErr

    While SSL degrades depth error relative to partial SL in both regimes, the proposed joint mapping transfers cross-task supervisory signal effectively without suffering from task imbalance.

  6. Knowl 6 — Partial Supervision on PASCAL-Context Multi-Task Benchmark

    data/table

    On PASCAL-Context across five dense prediction tasks (Semantic Segmentation, Human Part Segmentation, Surface Normals Estimation, Saliency Detection, and Semantic Edge Detection) using a ResNet-18 backbone with Atrous Spatial Pyramid Pooling (ASPP) heads, MTPSL achieves higher performance than standard supervised learning and semi-supervised perturbation consistency under both random and single-label settings.

    # labels Method Seg. (mIoU) ↑\uparrow H. Parts (mIoU) ↑\uparrow Norm. (mErr) ↓\downarrow Sal. (mIoU) ↑\uparrow Edge (odsF) ↑\uparrow
    full Supervised Learning 63.9 58.9 15.1 65.4 69.4
    random Supervised Learning 58.4 55.3 16.0 63.9 67.8
    Semi-supervised Learning 59.0 55.8 15.9 64.0 66.9
    MTPSL (Ours) 59.0 55.6 15.9 64.0 67.8
    one Supervised Learning 48.0 55.6 17.2 61.5 64.6
    Semi-supervised Learning 45.0 54.0 16.9 61.7 62.4
    MTPSL (Ours) 49.5 55.8 17.0 61.7 65.1

    In the low-label regime (one label per image), SSL degrades performance on segmentation (45.0 vs. 48.0 mIoU) and parts (54.0 vs. 55.6 mIoU), whereas MTPSL improves segmentation to 49.5 mIoU, human parts to 55.8 mIoU, and edge detection to 65.1 odsF.

  7. Knowl 7 — Cross-Task Pairwise Consistency for Fully Supervised Multi-Task Learning

    data/table

    Cross-task consistency learning in joint latent pairwise spaces generalizes to fully supervised multi-task learning, where complete ground-truth labels are available for all samples. On the NYU-v2 benchmark (13-class semantic segmentation, depth estimation, and surface normal estimation), the method was compared against single-task learning (STL), standard multi-task learning (MTL), architectural multi-task methods (MTAN, X-task), and adaptive loss balancing strategies (Uncertainty, GradNorm, MGDA, DWA).

    Method Seg. (mIoU) ↑\uparrow Depth (aErr) ↓\downarrow Norm. (mErr) ↓\downarrow
    STL 37.45 0.6079 25.94
    MTL 36.95 0.5510 29.51
    MTAN 39.39 0.5696 28.89
    X-task 38.91 0.5342 29.94
    Uncertainty 36.46 0.5376 27.58
    GradNorm 37.19 0.5775 28.51
    MGDA 38.65 0.5572 28.89
    DWA 36.46 0.5429 29.45
    MTPSL (Ours, uniform weights) 41.00 0.5148 28.58
    MTPSL (Ours + Uncertainty) 41.09 0.5090 26.78

    When trained with uniform loss weights, MTPSL achieves 41.00 mIoU and 0.5148 aErr, surpassing MTAN (39.39 mIoU, 0.5696 aErr) and X-task (38.91 mIoU, 0.5342 aErr). Combining MTPSL with Uncertainty-based loss weighting further reduces surface normal error to 26.78 mErr and depth error to 0.5090 aErr.

  8. Knowl 8 — Ablation of Conditional Task-Pair Conditioning and Backbone Feature Regularization

    empirical result

    An ablation study on NYU-v2 isolates the effects of (1) conditioning the joint task-space mapping via the auxiliary network aθa_\theta (cond) and (2) regularizing the mapped embeddings using the image feature encoder representation fϕ(x)f_\phi(\mathbf{x}) (reg).

    On the random label setting:

    • Partial Supervised Baseline: 27.05 mIoU, 0.6624 aErr, 33.58 mErr
    • Ours without task conditioning (w/o cond): 34.13 mIoU, 0.5968 aErr, 31.65 mErr
    • Ours without feature regularization (w/o reg): 33.87 mIoU, 0.5887 aErr, 31.24 mErr
    • Full Model (Ours): 34.26 mIoU, 0.5787 aErr, 31.06 mErr

    On the one label setting:

    • Partial Supervised Baseline: 25.75 mIoU, 0.6511 aErr, 33.73 mErr
    • Ours without task conditioning (w/o cond): 29.19 mIoU, 0.6181 aErr, 32.62 mErr
    • Ours without feature regularization (w/o reg): 28.36 mIoU, 0.6407 aErr, 32.92 mErr
    • Full Model (Ours): 30.36 mIoU, 0.6088 aErr, 32.08 mErr

    Both components yield improvements across all metrics. Removing feature regularization causes the largest degradation under single-label supervision (mIoU drops from 30.36 to 28.36; depth aErr increases from 0.6088 to 0.6407).

  9. Knowl 9 — Limitation of Exhaustive Pairwise Task Relation Modeling

    limitation

    The method learns cross-task consistency by modeling interactions between all possible pairs of tasks s∈Uns \in \mathcal{U}_n and t∈Tnt \in \mathcal{T}_n. However, not all pairs of dense vision tasks share strong mutual information or meaningful geometric/semantic relations. Forcing consistency across loosely coupled or orthogonal tasks can introduce uninformative constraints, indicating the need for mechanisms that automatically identify and select correlated task pairs rather than regularizing all task pairs uniformly.

Coverage note — None was omitted; all contributed formulations, model components, optimization objectives, benchmark evaluations, ablation studies, and limitations are fully covered.

References

  1. 1.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. PAMI, 39(12):2481–2495, 2017.
  2. 2.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin Raffel. Mixmatch: A holistic approach to semi-supervised learning. NeurIPS, 2019.
  3. 3.Hakan Bilen and Andrea Vedaldi. Integrated perception with recurrent multi-task neural networks. In Advances in Neural Information Processing Systems, pages 235–243, 2016.
  4. 4.Felix JS Bragman, Ryutaro Tanno, Sebastien Ourselin, Daniel C Alexander, and Jorge Cardoso. Stochastic filter groups for multi-task cnns: Learning specialist and generalist convolution kernels. In ICCV, pages 1385–1394, 2019.
  5. 5.David Bruggemann, Menelaos Kanakis, Stamatios Georgoulis, and Luc Van Gool. Automated search for resource-efficient branched multi-task networks. arXiv preprint arXiv:2008.10292, 2020.
  6. 6.David Bruggemann, Menelaos Kanakis, Anton Obukhov, Stamatios Georgoulis, and Luc Van Gool. Exploring relational context for multi-task dense prediction. In ICCV, 2021.
  7. 7.Rich Caruana. Multitask learning. Machine learning, 28(1):41–75, 1997.
  8. 8.Vincent Casser, Soeren Pirk, Reza Mahjourian, and Anelia Angelova. Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos. In AAAI, volume 33, pages 8001–8008, 2019.
  9. 9.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, pages 801–818, 2018.
  10. 10.Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu, Sanja Fidler, Raquel Urtasun, and Alan Yuille. Detect what you can: Detecting and representing objects using holistic models and body parts. In CVPR, pages 1971–1978, 2014.
  11. 11.Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich. Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In ICML, pages 794–803. PMLR, 2018.
  12. 12.Zhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong, Henrik Kretzschmar, Yuning Chai, and Dragomir Anguelov. Just pick a sign: Optimizing deep multitask models with gradient sign dropout. NeurIPS, 2020.
  13. 13.Zhihao Chen, Lei Zhu, Liang Wan, Song Wang, Wei Feng, and Pheng-Ann Heng. A multi-task mean teacher for semi-supervised shadow detection. In CVPR, pages 5611–5620, 2020.
  14. 14.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Computer Vision and Pattern Recognition, pages 3213–3223, 2016.
  15. 15.David Eigen and Rob Fergus. Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. In IEEE International Conference on Computer Vision, pages 2650–2658, 2015.
  16. 16.David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep network. arXiv preprint arXiv:1406.2283, 2014.
  17. 17.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 88(2):303–338, 2010.
  18. 18.Yuan Gao, Jiayi Ma, Mingbo Zhao, Wei Liu, and Alan L Yuille. Nddr-cnn: Layerwise feature fusing in multi-task cnns by neural discriminative dimensionality reduction. In CVPR, pages 3205–3214, 2019.
  19. 19.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR, pages 3354–3361. IEEE, 2012.
  20. 20.Ting Gong, Tyler Lee, Cory Stephenson, Venkata Renduchintala, Suchismita Padhy, Anthony Ndirango, Gokce Keskin, and Oguz H Elibol. A comparison of loss weighting strategies for multi task learning in deep neural networks. IEEE Access, 7:141627–141632, 2019.
  21. 21.Vitor Guizilini, Jie Li, Rares Ambrus, Sudeep Pillai, and Adrien Gaidon. Robust semi-supervised monocular depth estimation with reprojected distances. In Conference on robot learning, pages 503–512. PMLR, 2020.
  22. 22.Michelle Guo, Albert Haque, De-An Huang, Serena Yeung, and Li Fei-Fei. Dynamic task prioritization for multitask learning. In ECCV, pages 270–287, 2018.
  23. 23.Pengsheng Guo, Chen-Yu Lee, and Daniel Ulbricht. Learning to branch for multi-task learning. In ICML, pages 3854–3863. PMLR, 2020.
  24. 24.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In ICCV, pages 2961–2969, 2017.
  25. 25.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  26. 26.Derek Hoiem, Alexei A Efros, and Martial Hebert. Closing the loop in scene interpretation. In CVPR, pages 1–8. IEEE, 2008.
  27. 27.Lukas Hoyer, Dengxin Dai, Qin Wang, Yuhua Chen, and Luc Van Gool. Improving semi-supervised and domain-adaptive semantic segmentation with self-supervised depth estimation. arXiv preprint arXiv:2108.12545, 2021.
  28. 28.Abdullah-Al-Zubaer Imran, Chao Huang, Hui Tang, Wei Fan, Yuan Xiao, Dingjun Hao, Zhen Qian, and Demetri Terzopoulos. Partly supervised multitask learning. arXiv preprint arXiv:2005.02523, 2020.
  29. 29.Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In CVPR, pages 7482–7491, 2018.
  30. 30.Yevhen Kuznietsov, Jorg Stuckler, and Bastian Leibe. Semi-supervised deep learning for monocular depth map prediction. In CVPR, pages 6647–6655, 2017.
  31. 31.Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Julien Epps, and Bjorn W Schuller. Multi-task semi-supervised adversarial autoencoding for speech emotion recognition. arXiv preprint arXiv:1907.06078, 2019.
  32. 32.Wei-Hong Li and Hakan Bilen. Knowledge distillation for multi-task learning. In ECCV Workshop on Imbalance Problems in Computer Vision, pages 163–176. Springer, 2020.
  33. 33.Xi Lin, Hui-Ling Zhen, Zhenhua Li, Qing-Fu Zhang, and Sam Kwong. Pareto multi-task learning. NeurIPS, 32:12060–12070, 2019.
  34. 34.Beyang Liu, Stephen Gould, and Daphne Koller. Single image depth estimation from predicted semantic labels. In CVPR, pages 1253–1260. IEEE, 2010.
  35. 35.Qiuhua Liu, Xuejun Liao, and Lawrence Carin. Semi-supervised multitask learning. In Advances in Neural Information Processing Systems, pages 937–944, 2008.
  36. 36.Shikun Liu, Edward Johns, and Andrew J Davison. End-to-end multi-task learning with attention. In CVPR, pages 1871–1880, 2019.
  37. 37.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, pages 3431–3440, 2015.
  38. 38.Yongxi Lu, Abhishek Kumar, Shuangfei Zhai, Yu Cheng, Tara Javidi, and Rogerio Feris. Fully-adaptive feature sharing in multi-task networks with applications in person attribute classification. In CVPR, pages 5334–5343, 2017.
  39. 39.Yao Lu, Soren Pirk, Jan Dlabal, Anthony Brohan, Ankita Pasad, Zhao Chen, Vincent Casser, Anelia Angelova, and Ariel Gordon. Taskology: Utilizing task relations at scale. In CVPR, pages 8700–8709, 2021.
  40. 40.David R Martin, Charless C Fowlkes, and Jitendra Malik. Learning to detect natural image boundaries using local brightness, color, and texture cues. PAMI, 26(5):530–549, 2004.
  41. 41.Robert Mendel, Luis Antonio De Souza, David Rauber, Joao Paulo Papa, and Christoph Palm. Semi-supervised segmentation based on error-correcting supervision. In European Conference on Computer Vision, pages 141–157. Springer, 2020.
  42. 42.Shervin Minaee, Yuri Y Boykov, Fatih Porikli, Antonio J Plaza, Nasser Kehtarnavaz, and Demetri Terzopoulos. Image segmentation using deep learning: A survey. PAMI, 2021.
  43. 43.Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, and Martial Hebert. Cross-stitch networks for multi-task learning. In CVPR, pages 3994–4003, 2016.
  44. 44.Viktor Olsson, Wilhelm Tranheden, Juliano Pinto, and Lennart Svensson. Classmix: Segmentation-based data augmentation for semi-supervised learning. In WACV, pages 1369–1378, 2021.
  45. 45.Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, 2018.
  46. 46.Matteo Poggi, Filippo Aleotti, Fabio Tosi, and Stefano Mattoccia. On the uncertainty of self-supervised monocular depth estimation. In CVPR, pages 3227–3237, 2020.
  47. 47.Sebastian Ruder. An overview of multi-task learning in deep neural networks. arXiv preprint arXiv:1706.05098, 2017.
  48. 48.Sebastian Ruder, Joachim Bingel, Isabelle Augenstein, and Anders Søgaard. Latent multi-task architecture learning. In AAAI, volume 33, pages 4822–4829, 2019.
  49. 49.Suman Saha, Anton Obukhov, Danda Pani Paudel, Menelaos Kanakis, Yuhua Chen, Stamatios Georgoulis, and Luc Van Gool. Learning to relate depth and semantics for unsupervised domain adaptation. In CVPR, pages 8197–8207, 2021.
  50. 50.Ozan Sener and Vladlen Koltun. Multi-task learning as multi-objective optimization. NeurIPS, 2018.
  51. 51.Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from rgbd images. In European conference on computer vision, pages 746–760. Springer, 2012.
  52. 52.Kihyuk Sohn, David Berthelot, Chun-Liang Li, Zizhao Zhang, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. NeurIPS, 2020.
  53. 53.Zhuo Su, Wenzhe Liu, Zitong Yu, Dewen Hu, Qing Liao, Qi Tian, Matti Pietikainen, and Li Liu. Pixel difference networks for efficient edge detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5117–5127, 2021.
  54. 54.Boyang Sun, Jiaxu Xing, Hermann Blum, Roland Siegwart, and Cesar Cadena. See yourself in others: Attending multiple tasks for own failure detection. arXiv preprint arXiv:2110.02549, 2021.
  55. 55.Demetri Terzopoulos et al. Semi-supervised multi-task learning with chest x-ray images. In International Workshop on Machine Learning in Medical Imaging, pages 151–159. Springer, 2019.
  56. 56.Zhi Tian, Chunhua Shen, and Hao Chen. Conditional convolutions for instance segmentation. In ECCV, pages 282–298. Springer, 2020.
  57. 57.Simon Vandenhende, Stamatios Georgoulis, Bert De Brabandere, and Luc Van Gool. Branched multi-task networks: deciding what layers to share. In BMVC, 2020.
  58. 58.Simon Vandenhende, Stamatios Georgoulis, Wouter Van Gansbeke, Marc Proesmans, Dengxin Dai, and Luc Van Gool. Multi-task learning for dense prediction tasks: A survey. PAMI, 2021.
  59. 59.Simon Vandenhende, Stamatios Georgoulis, and Luc Van Gool. Mti-net: Multi-scale task interaction networks for multi-task learning. In ECCV, pages 527–543. Springer, 2020.
  60. 60.Raphael Voges and Bernardo Wagner. Timestamp offset calibration for an imu-camera system under interval uncertainty. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 377–384. IEEE, 2018.
  61. 61.Fei Wang, Xin Wang, and Tao Li. Semi-supervised multi-task learning with task regularizations. In ICDM, pages 562–568. IEEE, 2009.
  62. 62.Qin Wang, Dengxin Dai, Lukas Hoyer, Luc Van Gool, and Olga Fink. Domain adaptive semantic segmentation with self-supervised depth estimation. In ICCV, pages 8515–8525, 2021.
  63. 63.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. In CVPR, pages 8741–8750, 2021.
  64. 64.Fangting Xia, Peng Wang, Xianjie Chen, and Alan L Yuille. Joint multi-person pose estimation and semantic part segmentation. In CVPR, pages 6769–6778, 2017.
  65. 65.Dan Xu, Wanli Ouyang, Xiaogang Wang, and Nicu Sebe. Pad-net: Multi-tasks guided prediction-and-distillation network for simultaneous depth estimation and scene parsing. In CVPR, pages 675–684, 2018.
  66. 66.Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. Gradient surgery for multi-task learning. NeurIPS, 2020.
  67. 67.Zhiding Yu, Chen Feng, Ming-Yu Liu, and Srikumar Ramalingam. Casenet: Deep category-aware semantic edge detection. In CVPR, pages 5964–5973, 2017.
  68. 68.Amir R Zamir, Alexander Sax, Nikhil Cheerla, Rohan Suri, Zhangjie Cao, Jitendra Malik, and Leonidas J Guibas. Robust learning through cross-task consistency. In CVPR, pages 11197–11206, 2020.
  69. 69.Amir R Zamir, Alexander Sax, William Shen, Leonidas J Guibas, Jitendra Malik, and Silvio Savarese. Taskonomy: Disentangling task transfer learning. In CVPR, pages 3712–3722, 2018.
  70. 70.Amir R Zamir, Tilman Wekel, Pulkit Agrawal, Colin Wei, Jitendra Malik, and Silvio Savarese. Generic 3d representation via pose estimation and matching. In ECCV, pages 535–553. Springer, 2016.
  71. 71.Jing Zhang, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemeh Sadat Saleh, Tong Zhang, and Nick Barnes. Uc-net: Uncertainty inspired rgb-d saliency detection via conditional variational autoencoders. In CVPR, pages 8582–8591, 2020.
  72. 72.Miao Zhang, Weisong Ren, Yongri Piao, Zhengkun Rong, and Huchuan Lu. Select, supplement and focus for rgb-d saliency detection. In CVPR, pages 3472–3481, 2020.
  73. 73.Yu Zhang and Qiang Yang. A survey on multi-task learning. arXiv preprint arXiv:1707.08114, 2017.
  74. 74.Yu Zhang and Dit-Yan Yeung. Semi-supervised multi-task regression. In ECML PKDD, pages 617–631. Springer, 2009.
  75. 75.Zhenyu Zhang, Zhen Cui, Chunyan Xu, Zequn Jie, Xiang Li, and Jian Yang. Joint task-recursive learning for semantic segmentation and depth estimation. In ECCV, pages 235–251, 2018.
  76. 76.Zhenyu Zhang, Zhen Cui, Chunyan Xu, Yan Yan, Nicu Sebe, and Jian Yang. Pattern-affinitive propagation across depth, surface normal and semantic segmentation. In CVPR, pages 4106–4115, 2019.
  77. 77.Ling Zhou, Zhen Cui, Chunyan Xu, Zhenyu Zhang, Chaoqun Wang, Tong Zhang, and Jian Yang. Pattern-structure diffusion for multi-task learning. In CVPR, pages 4514–4523, 2020.
  78. 78.Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe. Unsupervised learning of depth and ego-motion from video. In CVPR, pages 1851–1858, 2017.

Citation

MLA
Li, W.-H., et al. “Learning Multiple Dense Prediction Tasks from Partially Annotated Data”. arXiv, 2021, http://arxiv.org/abs/2111.14893v3.
APA
Li, W.-H., Liu, X., & Bilen, H. (2021). Learning Multiple Dense Prediction Tasks from Partially Annotated Data. arXiv. http://arxiv.org/abs/2111.14893v3
Chicago
Li, W.-H., X. Liu, and H. Bilen. 2021. “Learning Multiple Dense Prediction Tasks from Partially Annotated Data”. arXiv. http://arxiv.org/abs/2111.14893v3.
Harvard
Li, W.-H., Liu, X. and Bilen, H. (2021) “Learning Multiple Dense Prediction Tasks from Partially Annotated Data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.14893v3.
Vancouver
1. Li W-H, Liu X, Bilen H (2021) Learning Multiple Dense Prediction Tasks from Partially Annotated Data. arXiv

BibTeX

@article{li2021learning,
  title = {Learning Multiple Dense Prediction Tasks from Partially Annotated Data},
  author = {Li, Wei-Hong and Liu, Xialei and Bilen, Hakan},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.14893v3},
  eprint = {2111.14893}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE