keyword
mask classification
Mask classification is an image segmentation paradigm in computer vision where a model predicts a set of binary masks representing distinct image regions and assigns a single category label to each mask. Unlike traditional per-pixel classification methods that predict a class label for every individual pixel independently, mask classification decouples the spatial localization of segments from their semantic categorization. By generating region-level mask proposals and classifying each entire region as a whole, this approach provides a unified framework capable of addressing semantic, instance, and panoptic segmentation tasks within a shared architecture.
2 items

Per-Pixel Classification is Not All You Need for Semantic Segmentation
Bowen Cheng, Alexander G. Schwing, Alexander Kirillov
Why you should read this
Proposes MaskFormer, a unified mask classification framework that replaces standard per-pixel approaches to achieve state-of-the-art accuracy across both semantic and panoptic segmentation benchmarks.
Modern approaches typically formulate semantic segmentation as a per-pixel classification task, while instance-level segmentation is handled with an alternative mask classification. Our key insight: mask classification is sufficiently general to solve both semantic- and instance-level segmentation tasks in a unified manner using the exact same model, loss, and training procedure. Following this observation, we propose MaskFormer, a simple mask classification model which predicts a set of binary masks, each associated with a single global class label prediction. Overall, the proposed mask classification-based method simplifies the landscape of effective approaches to semantic and panoptic segmentation tasks and shows excellent empirical results. In particular, we observe that MaskFormer outperforms per-pixel classification baselines when the number of classes is large. Our mask classification-based method outperforms both current state-of-the-art semantic (55.6 mIoU on ADE20K) and panoptic segmentation (52.7 PQ on COCO) models.
Added
2026-09-16

Masked-attention Mask Transformer for Universal Image Segmentation
Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, Rohit Girdhar
Why you should read this
Proposes Mask2Former, a unified segmentation architecture that utilizes masked attention to constrain cross-attention within predicted mask regions, achieving state-of-the-art performance across panoptic, instance, and semantic segmentation tasks while significantly improving training efficiency and convergence.
Image segmentation groups pixels with different semantics, e.g., category or instance membership. Each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing spe-cialized architectures for each task. We present Masked- attention Mask Transformer (Mask2Former), a new archi-tecture capable of addressing any image segmentation task (panoptic, instance or semantic). Its key components in-clude masked attention, which extracts localized features by constraining cross-attention within predicted mask regions. In addition to reducing the research effort by at least three times, it outperforms the best specialized architectures by a significant margin on four popular datasets. Most no-tably, Mask2Former sets a new state-of-the-art for panoptic segmentation (57.8 PQ on COCO), instance segmentation (50.1 AP on COCO) and semantic segmentation (57.7 mIoU onADE20K).
Added
2026-05-18
