keyword
hierarchical segmentation
Hierarchical segmentation is a computer vision and image analysis technique that partitions an image into a nested, multi-level structure of regions organized by varying degrees of granularity. Unlike single-level segmentation methods that produce a single partition, hierarchical segmentation generates a continuum of nested segments, often represented as a tree, ranging from fine sub-elements and object parts at the lower levels to larger, complete objects and scene categories at higher levels. This structured multi-scale representation allows visual data to be interpreted across multiple levels of semantic abstraction, effectively capturing part-whole relationships and resolving boundary ambiguities across complex scenes.
3 items

Hierarchical Open-vocabulary Universal Image Segmentation
Xudong Wang, Shufan Li, Konstantinos Kallidromitis, Yusuke Kato, Kazuki Kozuka, Trevor Darrell
Why you should read this
Presents HIPIE, a unified open-vocabulary framework that resolves segmentation ambiguity across multiple granularities by incorporating hierarchical visual representations alongside decoupled text-image fusion mechanisms for stuff and thing categories.
Open-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions. However, complex visual scenes can be naturally decomposed into simpler parts and abstracted at multiple levels of granularity, introducing inherent segmentation ambiguity. Unlike existing methods that typically sidestep this ambiguity and treat it as an external factor, our approach actively incorporates a hierarchical representation encompassing different semantic-levels into the learning process. We also propose a decoupled text-image fusion mechanism and representation learning modules for both “things” and “stuff”.1 Additionally, we systematically examine the differences that exist in the textual and visual features between these types of categories. Our resulting model, named HIPIE, tackles HIerarchical, oPen-vocabulary, and unIvErsal segmentation tasks within a unified framework. Benchmarked on over 40 datasets, e.g., ADE20K, COCO, Pascal-VOC Part, RefCOCO/RefCOCOg, ODinW and SeginW, HIPIE achieves the state-of-the-art results at various levels of image comprehension, including semantic-level (e.g., semantic segmentation), instance-level (e.g., panoptic/referring segmentation and object detection), as well as part-level (e.g., part/subpart segmentation) tasks.
Added
2026-09-26

An Optimal Graph Theoretic Approach to Data Clustering: Theory and Its Application to Image Segmentation
Zhenyu Wu, R. Leahy
Why you should read this
Develops a scalable graph-theoretic clustering method using subgraph condensation and equivalent trees to find globally optimal minimum cuts, guaranteeing closed boundary contours in image segmentation across hundreds of thousands of vertices.
A novel graph theoretic approach for data clustering is presented and its application to the image segmentation problem is demonstrated. The data to be clustered are represented by an undirected adjacency graph G with arc capacities assigned to reflect the similarity between the linked vertices. Clustering is achieved by removing arcs of G to form mutually exclusive subgraphs such that the largest inter-subgraph maximum flow is minimized. For graphs of moderate size (~ 2000 vertices), the optimal solution is obtained through partitioning a flow and cut equivalent tree of G, which can be efficiently constructed using the Gomory-Hu algorithm. However for larger graphs this approach is impractical. New theorems for subgraph condensation are derived and are then used to develop a fast algorithm which hierarchically constructs and partitions a partially equivalent tree of much reduced size. This algorithm results in an optimal solution equivalent to that obtained by partitioning the complete equivalent tree and is able to handle very large graphs with several hundred thousand vertices. The new clustering algorithm is applied to the image segmentation problem. The segmentation is achieved by effectively searching for closed contours of edge elements (equivalent to minimum cuts in G), which consist mostly of strong edges, while rejecting contours containing isolated strong edges. This method is able to accurately locate region boundaries and at the same time guarantees the formation of closed edge contours.
Added
2026-09-25

Indoor Segmentation and Support Inference from RGBD Images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, Rob Fergus
Why you should read this
Introduces a benchmark dataset of 1,449 densely annotated indoor RGB-D images alongside an integer programming formulation that infers physical support relations to improve 3D scene interpretation and object segmentation.
ct. We present an approach to interpret the major surfaces, ob- d support relations of an indoor scene from an RGBD image. sting work ignores physical interactions or is applied only to ns and hallways. Our goal is to parse typical, often messy, in- nes into floor, walls, supporting surfaces, and object regions, and er support relationships. One of our main interests is to better nd how 3D cues can best inform a structured 3D interpreta- also contribute a novel integer programming formulation to sical support relations. We offer a new dataset of 1449 RGBD capturing 464 diverse indoor scenes, with detailed annotations. eriments demonstrate our ability to infer support relations in scenes and verify that our 3D scene cues and inferred support etter object segmentation.
Added
2026-09-09
