Learning Multiple Dense Prediction Tasks from Partially Annotated Data
Wei-Hong LiXialei LiuHakan Bilen
Presents a label-efficient framework for multi-task dense prediction that learns joint pairwise task spaces to supervise missing annotations across images through cross-task consistency.
Deploying computer vision models for complex visual understanding typically requires learning multiple pixel-level tasks simultaneously, such as segmenting objects, estimating depth, and detecting surface orientations. Standard multi-task learning frameworks depend on training images that are fully annotated across all target tasks. In real-world applications, however, acquiring complete annotations for every image is prohibitively expensive, operationally challenging due to multi-sensor synchronization issues, and often impossible when integrating new tasks into existing legacy datasets. Consequently, real-world systems must operate in partially annotated environments where images contain labels for only a subset of tasks.
The article evaluates and demonstrates a multi-task partially supervised learning framework designed to train unified vision models when training images are missing labels for one or more tasks. The core objective is to leverage inherent cross-task relationships to supervise unlabelled tasks without requiring complex, direct label transformations or massive fully labelled datasets.
To address this challenge, the authors designed a multi-task learning approach that maps task pairs into a shared, lower-dimensional joint feature space to enforce consistency between the predictions of unlabelled tasks and the ground-truth labels of available tasks. To ensure computational scalability as the number of tasks grows, the framework uses a single encoder network that dynamically adjusts its internal weights for specific task pairs via conditioning. Additionally, the training process includes a regularization term tied to the core image features to prevent the mapping function from collapsing into meaningless, trivial shortcuts. The authors validated this approach on three standard visual benchmarks—Cityscapes, NYU-v2, and PASCAL-Context—across diverse data regimes, including random task availability, single-label-per-image constraints, and heavily imbalanced task ratios.
The experiments show that the proposed framework consistently outperforms standard supervised and semi-supervised baselines under partial supervision. In partial annotation scenarios on the Cityscapes benchmark, the method achieved a segmentation intersection-over-union score of 74.90%, outperforming the standard supervised baseline trained only on partial labels (69.50%) as well as the fully supervised model trained with 100% complete labels across all tasks (73.36%). On the indoor NYU-v2 dataset, when restricted to only one label per image, the framework significantly outperformed the partial supervised baseline across all metrics, raising segmentation accuracy from 25.75% to 30.36% while reducing depth and surface normal errors. Furthermore, in severely imbalanced conditions where 90% of annotations were missing for a given task, the method maintained high accuracy across all outputs, demonstrating strong data efficiency. Finally, when applied to fully supervised settings and combined with adaptive loss weighting, the model surpassed established multi-task benchmarks.
These findings indicate that learning joint pairwise relationships across tasks allows models to extract valuable cross-task supervision from unlabelled visual data. Operationally, this capability reduces data labelling costs, lowers sensor hardware requirements, and accelerates the development timelines needed to expand vision systems with new capabilities. By removing the requirement that all training images must be synchronized across every sensor modality, organizations can unlock and repurpose partially labelled historical data.
Organizations developing multi-task vision systems should adopt joint pairwise consistency learning rather than relying solely on single-task data augmentation or conventional semi-supervised pipelines. When deploying this method, engineering teams should pair the consistency framework with adaptive loss-weighting mechanisms to maximize performance balance across uneven task distributions. Future work should focus on developing automated mechanisms to identify which specific task pairs share meaningful correlations, avoiding the computational overhead of training mappings between unrelated tasks.
The primary limitation of the article is that cross-task relationships were evaluated across all possible task pairs, even though some pairs may offer minimal complementary information. Furthermore, performance improvements depend on the visual domain and the degree of structural correlation between tasks, showing smaller relative gains on diverse image datasets where cross-task links are weaker. Confidence in the core results remains high across standard dense prediction benchmarks, though pilot testing is recommended before applying the framework to domains with unverified task affinities.
- Paper: Taskonomy: Disentangling Task Transfer Learning, Amir Zamir et al. (2018). It establishes the foundational computational framework for mapping and exploiting cross-task relationships across dense visual predictions, which directly motivates the source's pairwise multi-task transfer.
- Paper: End-To-End Multi-Task Learning With Attention, Shikun Liu et al. (2018). It introduces shared-encoder multi-task attention architectures and dynamic loss weighting on dense prediction benchmarks like Cityscapes and NYUv2 that the source builds upon.
- Paper: Multi-task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics, Alex Kendall et al. (2017). It presents the seminal approach for joint learning and principled homoscedastic uncertainty loss weighting across multi-task geometry and dense semantic predictions.
- Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). It analyzes multi-task gradient conflict dynamics and provides foundational optimization techniques essential for stabilizing joint multi-task representation learning.
- Paper: Multi-Task Learning as Multi-Objective Optimization, Ozan Sener et al. (2018). It establishes multi-task dense prediction as a multi-objective optimization problem, providing the mathematical context for balancing competing task objectives.
- Paper: Cross-Stitch Networks for Multi-task Learning, Ishan Misra et al. (2016). It provides the foundational framework for dynamically learning how to share representations across task pairs in multi-task visual learning.
- Paper: Convex multi-task feature learning, Andreas Argyriou et al. (2008). It formulates the core theoretical principles of learning low-dimensional shared feature representations across multiple related tasks.
- Paper: VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding, Yi Xin et al. (2024). It extends multi-task dense scene understanding to parameter-efficient vision transformer transfer learning, building on principles of shared feature representations across dense tasks.
- Paper: Uni3D: A Unified Baseline for Multi-Dataset 3D Object Detection, Bo Zhang et al. (2023). It generalizes multi-task and multi-dataset unified representation learning to 3D object detection across disparate annotations and sensor configurations.
- Paper: Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures, Linshan Wu et al. (2023). It applies adaptive probabilistic feature modeling to tackle extremely sparse pixel-level supervision in dense visual segmentation tasks.
