keyword
Open-vocabulary universal image segmentation
Open-vocabulary universal image segmentation is a computer vision paradigm that unifies multiple image segmentation tasks, such as semantic, instance, and panoptic segmentation, within a single framework capable of segmenting arbitrary visual categories described by natural language. Unlike traditional segmentation systems that are restricted to a fixed and predefined set of categories seen during training, open-vocabulary universal segmentation leverages multimodal vision-language representations to recognize and partition both familiar and previously unseen foreground objects and background regions during inference. By combining task-agnostic mask prediction with text-driven semantic recognition, this approach enables comprehensive pixel-level scene understanding across varying levels of granularity without requiring architecture modifications or task-specific retraining.
3 items

Open-Vocabulary Universal Image Segmentation with MaskCLIP
Zheng Ding, Jieke Wang, Zhuowen Tu
Why you should read this
Presents MaskCLIP, a Transformer-based method that integrates mask tokens into a pre-trained Vision Transformer CLIP backbone via relative mask attention, enabling open-vocabulary panoptic, instance, and semantic segmentation without requiring complex student-teacher distillation.
In this paper, we tackle an emerging computer vision task, open-vocabulary universal image segmentation, that aims to perform semantic/instance/panoptic segmentation (background semantic labeling + foreground instance segmentation) for arbitrary categories of text-based descriptions in inference time. We first build a baseline method by directly adopting pre-trained CLIP models without finetuning or distillation. We then develop MaskCLIP, a Transformer-based approach with a MaskCLIP Visual Encoder, which is an encoder-only module that seamlessly integrates mask tokens with a pre-trained ViT CLIP model for semantic/instance segmentation and class prediction. MaskCLIP learns to efficiently and effectively utilize pre-trained partial/dense CLIP features within the MaskCLIP Visual Encoder that avoids the time-consuming student-teacher training process. MaskCLIP outperforms previous methods for semantic/instance/panoptic segmentation on ADE20K and PASCAL datasets. We show qualitative illustrations for MaskCLIP with online custom categories. Project website: https://maskclip.github.io.
Added
2026-10-01

Alpha-CLIP: A CLIP Model Focusing on Wherever you Want
Zeyi Sun, Ye Fang, Tong Wu, Pan Zhang, Yuhang Zang, Shu Kong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang
Why you should read this
Introduces Alpha-CLIP, an enhanced CLIP model with an auxiliary alpha channel that enables fine-grained, region-specific focus while preserving contextual awareness across open-world recognition, multimodal language models, and 2D/3D generation tasks.
Contrastive Language-Image Pre-training (CLIP) plays an essential role in extracting valuable content information from images across diverse tasks. It aligns textual and visual modalities to comprehend the entire image, including all the details, even those irrelevant to specific tasks. However, for a finer understanding and controlled editing of images, it becomes crucial to focus on specific regions of interest, which can be indicated as points, masks, or boxes by humans or perception models. To fulfill the requirements, we introduce Alpha-CLIP, an enhanced version of CLIP with an auxiliary alpha channel to suggest attentive regions and fine-tuned with constructed millions of RGBA region-text pairs. Alpha-CLIP not only preserves the visual recognition ability of CLIP but also enables precise control over the emphasis of image contents. It demonstrates effectiveness in various tasks, including but not limited to open-world recognition, multimodal large language models, and conditional 2D / 3D generation. It has a strong potential to serve as a versatile tool for image-related tasks. Our project is with codes and models available is linked to https://aleafy.github.io/alpha-clip/.
Added
2026-09-26

Hierarchical Open-vocabulary Universal Image Segmentation
Xudong Wang, Shufan Li, Konstantinos Kallidromitis, Yusuke Kato, Kazuki Kozuka, Trevor Darrell
Why you should read this
Presents HIPIE, a unified open-vocabulary framework that resolves segmentation ambiguity across multiple granularities by incorporating hierarchical visual representations alongside decoupled text-image fusion mechanisms for stuff and thing categories.
Open-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions. However, complex visual scenes can be naturally decomposed into simpler parts and abstracted at multiple levels of granularity, introducing inherent segmentation ambiguity. Unlike existing methods that typically sidestep this ambiguity and treat it as an external factor, our approach actively incorporates a hierarchical representation encompassing different semantic-levels into the learning process. We also propose a decoupled text-image fusion mechanism and representation learning modules for both “things” and “stuff”.1 Additionally, we systematically examine the differences that exist in the textual and visual features between these types of categories. Our resulting model, named HIPIE, tackles HIerarchical, oPen-vocabulary, and unIvErsal segmentation tasks within a unified framework. Benchmarked on over 40 datasets, e.g., ADE20K, COCO, Pascal-VOC Part, RefCOCO/RefCOCOg, ODinW and SeginW, HIPIE achieves the state-of-the-art results at various levels of image comprehension, including semantic-level (e.g., semantic segmentation), instance-level (e.g., panoptic/referring segmentation and object detection), as well as part-level (e.g., part/subpart segmentation) tasks.
Added
2026-09-26
