Hierarchical Open-vocabulary Universal Image Segmentation

Xudong WangShufan LiKonstantinos KallidromitisYusuke KatoKazuki KozukaTrevor Darrell

article2023NeurIPS75 citations

Presents HIPIE, a unified open-vocabulary framework that resolves segmentation ambiguity across multiple granularities by incorporating hierarchical visual representations alongside decoupled text-image fusion mechanisms for stuff and thing categories.

Listen

Modern computer vision systems increasingly rely on open-vocabulary models to identify and segment visual scenes using arbitrary natural language descriptions. However, real-world scenes are inherently multi-layered, as visual content can be interpreted at the level of entire objects, background surfaces, or detailed parts and subparts. Existing segmentation models typically overlook this structural ambiguity or rely on fragmented, task-specific pipelines that apply uniform processing across different visual categories, leading to degraded performance. The article presents and evaluates HIPIE, a unified framework designed to achieve hierarchical, open-vocabulary image segmentation and object detection across multiple levels of granularity.

To address structural ambiguity, the researchers developed an architecture that explicitly separates the processing of countable foreground objects from background regions. Because background categories share high textual similarity, the system applies early multimodal fusion for foreground objects to capitalize on distinct descriptions while using late fusion for background surfaces to prevent feature confusion. The model was trained using composite text prompts combining scene- and part-level labels, and its performance was evaluated across more than 40 diverse computer vision benchmark datasets covering object detection, semantic, panoptic, part, and referring expression tasks.

The evaluation showed that the unified framework achieved state-of-the-art performance across multiple visual recognition tasks. In part-level segmentation, the model surpassed the previous benchmark by 5.2 points in mean Intersection-Over-Union on the Pascal-Panoptic-Parts dataset. When evaluated in zero-shot transfer settings across 25 diverse segmentation datasets in the Segmentation in the Wild benchmark, the system outperformed existing universal models by an average of 8.9 points in mean average precision. Furthermore, the model improved standard object detection precision on the COCO dataset by up to 3.2 points over competitive baselines. Ablation experiments verified that decoupling foreground and background processing pathways was the primary driver of these performance gains.

These results indicate that unified visual models can effectively handle multiple granularities without sacrificing task-specific accuracy. Consolidating whole-object detection, background understanding, and fine-grained part segmentation into a single architecture offers significant operational efficiencies by reducing model deployment complexity and maintenance costs. Such comprehensive visual parsing is particularly valuable for applications demanding high-precision scene understanding, such as autonomous systems, manufacturing inspection, and automated digital content editing.

Organizations developing or deploying visual AI systems should evaluate decoupled, hierarchical frameworks as a means to replace complex multi-model pipelines with a single unified solution. Before full-scale deployment in production or safety-critical environments, technical teams should conduct focused pilot tests to validate performance on domain-specific edge cases. While the article demonstrates strong evidence and high confidence for static image understanding, the current findings are bounded by static benchmarks and reliance on curated training data; further research is recommended to expand the framework into video tracking and to mitigate potential biases inherited from human-annotated datasets.

  • Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer establishes the unified mask classification paradigm that frames both semantic- and instance-level segmentation under a shared query-based framework.
  • Paper: Panoptic Segmentation, Alexander Kirillov et al. (2018). This foundational work formalizes panoptic segmentation and the core dichotomy between countable 'things' and amorphous 'stuff' that HIPIE explicitly decouples and models.
  • Paper: Unified Perceptual Parsing for Scene Understanding, Tete Xiao et al. (2018). UPerNet introduces unified perceptual parsing to segment scenes across multiple hierarchical levels from stuff and objects down to part boundaries.
  • Paper: Semantic Understanding of Scenes Through the ADE20K Dataset, Bolei Zhou et al. (2016). This paper establishes the ADE20K dataset and its hierarchical annotation scheme across scenes, objects, and parts, serving as a primary benchmark for hierarchical universal segmentation.
  • Paper: COCO-Stuff: Thing and Stuff Classes in Context, Holger Caesar et al. (2016). This paper introduces the COCO-Stuff benchmark and standardizes the distinction and spatial context modeling between thing and stuff categories.
  • Paper: Panoptic Feature Pyramid Networks, Alexander Kirillov et al. (2019). Panoptic FPN demonstrates early unified multi-task architectures integrating instance and semantic segmentation heads over shared feature representations.
Cover for Hierarchical Open-vocabulary Universal Image Segmentation

Abstract

Open-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions. However, complex visual scenes can be naturally decomposed into simpler parts and abstracted at multiple levels of granularity, introducing inherent segmentation ambiguity. Unlike existing methods that typically sidestep this ambiguity and treat it as an external factor, our approach actively incorporates a hierarchical representation encompassing different semantic-levels into the learning process. We also propose a decoupled text-image fusion mechanism and representation learning modules for both “things” and “stuff”.1 Additionally, we systematically examine the differences that exist in the textual and visual features between these types of categories. Our resulting model, named HIPIE, tackles HIerarchical, oPen-vocabulary, and unIvErsal segmentation tasks within a unified framework. Benchmarked on over 40 datasets, e.g., ADE20K, COCO, Pascal-VOC Part, RefCOCO/RefCOCOg, ODinW and SeginW, HIPIE achieves the state-of-the-art results at various levels of image comprehension, including semantic-level (e.g., semantic segmentation), instance-level (e.g., panoptic/referring segmentation and object detection), as well as part-level (e.g., part/subpart segmentation) tasks.

Citation

MLA
Wang, X., et al. “Hierarchical Open-vocabulary Universal Image Segmentation”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 21429–53, https://proceedings.neurips.cc/paper_files/paper/2023/file/43663f64775ae439ec52b64305d219d3-Paper-Conference.pdf.
APA
Wang, X., Li, S., Kallidromitis, K., Kato, Y., Kozuka, K., & Darrell, T. (2023). Hierarchical Open-vocabulary Universal Image Segmentation. Advances in Neural Information Processing Systems, 36, 21429–21453. https://proceedings.neurips.cc/paper_files/paper/2023/file/43663f64775ae439ec52b64305d219d3-Paper-Conference.pdf
Chicago
Wang, X., S. Li, K. Kallidromitis, Y. Kato, K. Kozuka, and T. Darrell. 2023. “Hierarchical Open-vocabulary Universal Image Segmentation”. Advances in Neural Information Processing Systems 36: 21429–53. https://proceedings.neurips.cc/paper_files/paper/2023/file/43663f64775ae439ec52b64305d219d3-Paper-Conference.pdf.
Harvard
Wang, X. et al. (2023) “Hierarchical Open-vocabulary Universal Image Segmentation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 21429–21453. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/43663f64775ae439ec52b64305d219d3-Paper-Conference.pdf.
Vancouver
1. Wang X, Li S, Kallidromitis K, Kato Y, Kozuka K, Darrell T (2023) Hierarchical Open-vocabulary Universal Image Segmentation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 21429–21453

BibTeX

@inproceedings{wang2023hierarchical,
  title = {Hierarchical Open-vocabulary Universal Image Segmentation},
  author = {Wang, Xudong and Li, Shufan and Kallidromitis, Konstantinos and Kato, Yusuke and Kozuka, Kazuki and Darrell, Trevor},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {21429-21453},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/43663f64775ae439ec52b64305d219d3-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors