Built independently by an author, for readers. Read the story and support ChapterPal

keyword

image segmentation

Image segmentation is a fundamental computer vision and digital image processing task that involves partitioning a digital image into multiple distinct regions, segments, or sets of pixels to simplify its representation and facilitate detailed visual analysis. By assigning a specific class label or object identity to each pixel, the process enables machines to locate boundaries and comprehend the spatial extent of objects and background elements within a scene. Common formulations include semantic segmentation, which classifies pixels into predefined categories; instance segmentation, which delineates individual object instances; and panoptic segmentation, which combines both approaches. Methodologies for performing segmentation range from classical algorithms based on edge detection, thresholding, clustering, and graph cuts to modern deep learning architectures, such as fully convolutional networks, vision transformers, and prompt-driven multimodal models.

32 items

Image Segmentation Using Text and Image Prompts

Image Segmentation Using Text and Image Prompts

Timo Lüddecke, Alexander S. Ecker

OrganizationsMax Planck Institute for Dynamics and Self-OrganizationUniversity of Göttingen

Why you should read this

Proposes CLIPSeg, a unified system that adds a lightweight transformer-based decoder to a frozen CLIP model to perform zero-shot, one-shot, and referring expression image segmentation using arbitrary text or visual prompts at test time.

Image segmentation is usually addressed by training a model for a fixed set of object classes. Incorporating additional classes or more complex queries later is expensive as it requires re-training the model on a dataset that encompasses these expressions. Here we propose a system that can generate image segmentations based on arbitrary prompts at test time. A prompt can be either a text or an image. This approach enables us to create a unified model (trained once) for three common segmentation tasks, which come with distinct challenges: referring expression segmentation, zero-shot segmentation and one-shot segmentation. We build upon the CLIP model as a backbone which we extend with a transformer-based decoder that enables dense prediction. After training on an extended version of the PhraseCut dataset, our system generates a binary segmentation map for an image based on a free-text prompt or on an additional image expressing the query. We analyze different variants of the latter image-based prompts in detail. This novel hybrid input allows for dynamic adaptation not only to the three segmentation tasks mentioned above, but to any binary segmentation task where a text or image query can be formulated. Finally, we find our system to adapt well to generalized queries involving affordances or properties. Code is available at https://eckerlab.org/code/clipseg

Added

2026-10-05

LISA: Reasoning Segmentation via Large Language Model

LISA: Reasoning Segmentation via Large Language Model

Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, Jiaya Jia

OrganizationsHarbin Institute of TechnologyMicrosoftSmartMoreThe Chinese University of Hong Kong

Why you should read this

Introduces the task of reasoning segmentation alongside LISA, an end-to-end multimodal architecture that decodes special text tokens into binary masks to segment targets specified by implicit, knowledge-intensive user instructions.

Although perception systems have made remarkable advancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target objects before executing visual recognition tasks. Such systems cannot actively reason and comprehend implicit user intention. In this work, we propose a new segmentation task — reasoning segmentation. The task is designed to output a segmentation mask given a complex and implicit query text. Furthermore, we establish a benchmark comprising over one thousand image-instruction-mask data samples, incorporating intricate reasoning and world knowledge for evaluation purposes. Finally, we present LISA: large Language Instructed Segmentation Assistant, which inherits the language generation capabilities of multimodal Large Language Models (LLMs) while also possessing the ability to produce segmentation masks. We expand the original vocabulary with a <SEG> token and propose the embedding-as-mask paradigm to unlock the segmentation capability. Remarkably, LISA can handle cases involving complex reasoning and world knowledge. Also, it demonstrates robust zero-shot capability when trained exclusively on reasoning-free datasets. In addition, fine-tuning the model with merely 239 reasoning segmentation data samples results in further performance enhancement. Both quantitative and qualitative experiments show our method effectively unlocks new reasoning segmentation capabilities for multimodal LLMs. Code, models, and data are available at github.com/dvlab-research/LISA.

Added

2026-10-05

Language-driven Semantic Segmentation

Language-driven Semantic Segmentation

Boyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun, René Ranftl

OrganizationsAppleCornell UniversityIntelUniversity of Copenhagen

Why you should read this

Proposes LSeg, a model that aligns per-pixel visual embeddings with text representations through contrastive learning, enabling zero-shot semantic segmentation of arbitrary unseen categories without additional training.

We present LSeg, a novel model for language-driven semantic image segmentation. LSeg uses a text encoder to compute embeddings of descriptive input labels (e.g., "grass" or "building") together with a transformer-based image encoder that computes dense per-pixel embeddings of the input image. The image encoder is trained with a contrastive objective to align pixel embeddings to the text embedding of the corresponding semantic class. The text embeddings provide a flexible label representation in which semantically similar labels map to similar regions in the embedding space (e.g., "cat" and "furry"). This allows LSeg to generalize to previously unseen categories at test time, without retraining or even requiring a single additional training sample. We demonstrate that our approach achieves highly competitive zero-shot performance compared to existing zero- and few-shot semantic segmentation methods, and even matches the accuracy of traditional segmentation algorithms when a fixed label set is provided. Code and demo are available at this https URL.

Added

2026-10-05

GSVA: Generalized Segmentation via Multimodal Large Language Models

GSVA: Generalized Segmentation via Multimodal Large Language Models

Zhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan, Shiji Song, Gao Huang

OrganizationsTsinghua University

Why you should read this

Presents a multimodal framework that enables language models to segment multiple target objects simultaneously and explicitly reject non-existent targets using dedicated tokens, setting a new state of the art on the gRefCOCO benchmark.

Generalized Referring Expression Segmentation (GRES) extends the scope of classic RES to refer to multiple objects in one expression or identify the empty targets absent in the image. GRES poses challenges in modeling the complex spatial relationships of the instances in the image and identifying non-existing referents. Multimodal Large Language Models (MLLMs) have recently shown tremendous progress in these complicated vision-language tasks. Connecting Large Language Models (LLMs) and vision models, MLLMs are proficient in understanding contexts with visual inputs. Among them, LISA, as a representative, adopts a special [SEG] token to prompt a segmentation mask decoder, e.g., SAM, to enable MLLMs in the RES task. However, existing solutions to GRES remain unsatisfactory since current segmentation MLLMs cannot correctly handle the cases where users might reference multiple subjects in a singular prompt or provide descriptions incongruent with any image target. In this paper, we propose Generalized Segmentation Vision Assistant (GSVA) to address this gap. Specifically, GSVA reuses the [SEG] token to prompt the segmentation model towards supporting multiple mask references simultaneously and innovatively learns to generate a [REJ] token to reject the null targets explicitly. Experiments validate GSVA’s efficacy in resolving the GRES issue, marking a notable enhancement and setting a new record on the GRES benchmark gRefCOCO dataset. GSVA also proves effective across various classic referring segmentation and comprehension tasks. Code is available at https://github.com/LeapLabTHU/GSVA.

Added

2026-09-26

An Optimal Graph Theoretic Approach to Data Clustering: Theory and Its Application to Image Segmentation

An Optimal Graph Theoretic Approach to Data Clustering: Theory and Its Application to Image Segmentation

Zhenyu Wu, R. Leahy

OrganizationsUniversity of PennsylvaniaUniversity of Southern California

Why you should read this

Develops a scalable graph-theoretic clustering method using subgraph condensation and equivalent trees to find globally optimal minimum cuts, guaranteeing closed boundary contours in image segmentation across hundreds of thousands of vertices.

A novel graph theoretic approach for data clustering is presented and its application to the image segmentation problem is demonstrated. The data to be clustered are represented by an undirected adjacency graph G with arc capacities assigned to reflect the similarity between the linked vertices. Clustering is achieved by removing arcs of G to form mutually exclusive subgraphs such that the largest inter-subgraph maximum flow is minimized. For graphs of moderate size (~ 2000 vertices), the optimal solution is obtained through partitioning a flow and cut equivalent tree of G, which can be efficiently constructed using the Gomory-Hu algorithm. However for larger graphs this approach is impractical. New theorems for subgraph condensation are derived and are then used to develop a fast algorithm which hierarchically constructs and partitions a partially equivalent tree of much reduced size. This algorithm results in an optimal solution equivalent to that obtained by partitioning the complete equivalent tree and is able to handle very large graphs with several hundred thousand vertices. The new clustering algorithm is applied to the image segmentation problem. The segmentation is achieved by effectively searching for closed contours of edge elements (equivalent to minimum cuts in G), which consist mostly of strong edges, while rejecting contours containing isolated strong edges. This method is able to accurately locate region boundaries and at the same time guarantees the formation of closed edge contours.

Added

2026-09-25

Kernel k-means: spectral clustering and normalized cuts

Kernel k-means: spectral clustering and normalized cuts

Inderjit S. Dhillon, Yuqiang Guan, Brian Kulis

OrganizationsUniversity of Texas at Austin

Why you should read this

Proves a theoretical equivalence between weighted kernel k-means and spectral clustering objectives, enabling graph-based normalized cuts to be minimized through efficient iterative algorithms without relying on computationally expensive eigenvector calculations.

Kernel k-means and spectral clustering have both been used to identify clusters that are non-linearly separable in input space. Despite significant research, these methods have remained only loosely related. In this paper, we give an explicit theoretical connection between them. We show the generality of the weighted kernel k-means objective function, and derive the spectral clustering objective of normalized cut as a special case. Given a positive definite similarity matrix, our results lead to a novel weighted kernel k-means algorithm that monotonically decreases the normalized cut. This has important implications: a) eigenvector-based algorithms, which can be computationally prohibitive, are not essential for minimizing normalized cuts, b) various techniques, such as local search and acceleration schemes, may be used to improve the quality as well as speed of kernel k-means. Finally, we present results on several interesting data sets, including diametrical clustering of large gene-expression matrices and a handwriting recognition data set.

Added

2026-09-25

Learning with Hypergraphs: Clustering, Classification, and Embedding

Learning with Hypergraphs: Clustering, Classification, and Embedding

Dengyong Zhou, Jiayuan Huang, B. Schölkopf

OrganizationsMax Planck Institute for Biological CyberneticsNEC Laboratories America, Inc.University of Waterloo

Why you should read this

Generalizes spectral graph theory to hypergraphs by formulating normalized hypergraph cuts, Laplacians, and random walks to enable higher-order relational clustering, embedding, and transductive classification without losing multi-object structural information.

We usually endow the investigated objects with pairwise relationships, which can be illustrated as graphs. In many real-world problems, however, relationships among the objects of our interest are more complex than pairwise. Naively squeezing the complex relationships into pairwise ones will inevitably lead to loss of information which can be expected valuable for our learning tasks however. There we consider using hypergraphs instead to completely represent complex relationships among the objects of our interest, and thus the problem of learning with hypergraphs arises. Our main contribution in this paper is to generalize the powerful methodology of spectral clustering which originally operates on undirected graphs to hypergraphs, and further develop algorithms for hypergraph embedding and transductive classification on the basis of the spectral hypergraph clustering approach. Our experiments on a number of benchmarks showed the advantages of hypergraphs over usual graphs.

Added

2026-09-24

Blobworld: Image Segmentation Using Expectation-Maximization and Its Application to Image Querying

Blobworld: Image Segmentation Using Expectation-Maximization and Its Application to Image Querying

C. Carson, Serge J. Belongie, H. Greenspan, Jitendra Malik

Why you should read this

Introduces a region-based image retrieval framework that uses Expectation-Maximization clustering across joint color, texture, and spatial features to enable object-level queries that significantly outperform traditional global histogram methods.

Retrieving images from large and varied collections using image content as a key is a challenging and important problem. We present a new image representation which provides a transformation from the raw pixel data to a small set of image regions which are coherent in color and texture. This “Blobworld” representation is created by clustering pixels in a joint color-texture-position feature space. The segmentation algorithm is fully automatic and has been run on a collection of 10,000 natural images. We describe a system that uses the Blobworld representation to retrieve images from this collection. An important aspect of the system is that the user is allowed to view the internal representation of the submitted image and the query results. Similar systems do not offer the user this view into the workings of the system; consequently, query results from these systems can be inexplicable, despite the availability of knobs for adjusting the similarity metrics. By finding image regions which roughly correspond to objects, we allow querying at the level of objects rather than global image properties. We present results indicating that querying for distinctive objects using Blobworld produces significantly higher precision than does querying using color and texture histograms of the entire image.

Added

2026-09-24

Matching Words and Pictures

Matching Words and Pictures

Kobus Barnard, Pinar Duygulu, David Forsyth, Nando de Freitas, David Blei, Michael I. Jordan

OrganizationsMiddle East Technical UniversityUniversity of ArizonaUniversity of British ColumbiaUniversity of California

Why you should read this

Proposes probabilistic and statistical translation models to learn the joint distribution of segmented image regions and words, establishing a foundational framework for automatic image annotation, object recognition, and text-based image retrieval.

We present a new approach for modeling multi-modal data sets, focusing on the specific case of segmented images with associated text. Learning the joint distribution of image regions and words has many applications. We consider in detail predicting words associated with whole images (auto-annotation) and corresponding to particular image regions (region naming). Auto-annotation might help organize and access large collections of images. Region naming is a model of object recognition as a process of translating image regions to words, much as one might translate from one language to another. Learning the relationships between image regions and semantic correlates (words) is an interesting example of multi-modal data mining, particularly because it is typically hard to apply data mining techniques to collections of images. We develop a number of models for the joint distribution of image regions and words, including several which explicitly learn the correspondence between regions and words. We study multi-modal and correspondence extensions to Hofmann’s hierarchical clustering/aspect model, a translation model adapted from statistical machine translation (Brown et al.), and a multi-modal extension to mixture of latent Dirichlet allocation (MoM-LDA). All models are assessed using a large collection of annotated images of real

Added

2026-09-18

A Closed-Form Solution to Natural Image Matting

A Closed-Form Solution to Natural Image Matting

Anat Levin, Dani Lischinski, Yair Weiss

OrganizationsThe Hebrew University of Jerusalem

Why you should read this

Presents a closed-form solution to natural image matting that analytically eliminates unknown foreground and background colors, enabling globally optimal alpha matte extraction from sparse user scribbles by solving a single sparse linear system.

Interactive digital matting, the process of extracting a foreground object from an image based on limited user input, is an important task in image and video editing. From a computer vision perspective, this task is extremely challenging because it is massively ill-posed — at each pixel we must estimate the foreground and the background colors, as well as the foreground opacity (“alpha matte”) from a single color measurement. Current approaches either restrict the estimation to a small part of the image, estimating foreground and background colors based on nearby pixels where they are known, or perform iterative nonlinear estimation by alternating foreground and background color estimation with alpha estimation. In this paper we present a closed form solution to natural image matting. We derive a cost function from local smoothness assumptions on foreground and background colors, and show that in the resulting expression it is possible to analytically eliminate the foreground and background colors to obtain a quadratic cost function in alpha. This allows us to find the globally optimal alpha matte by solving a sparse linear system of equations. Furthermore, the closed form formula allows us to predict the properties of the solution by analyzing the eigenvectors of a sparse matrix, closely related to matrices used in spectral image segmentation algorithms. We show that high quality mattes can be obtained on natural images from a small amount of user input.

Added

2026-09-16

Level set evolution without re-initialization: a new variational formulation

Level set evolution without re-initialization: a new variational formulation

Chunming Li, Chenyang Xu, C. Gui, M. Fox

OrganizationsSiemens Corporate ResearchUniversity of Connecticut

Why you should read this

Proposes a distance-regularizing variational formulation for active contours that maintains the level set function as a signed distance profile throughout evolution, eliminating costly re-initialization steps while enabling faster numerical convergence with simple finite difference schemes.

In this paper, we present a new variational formulation for geometric active contours that forces the level set function to be close to a signed distance function, and therefore completely eliminates the need of the costly re-initialization procedure. Our variational formulation consists of an internal energy term that penalizes the deviation of the level set function from a signed distance function, and an external energy term that drives the motion of the zero level set toward the desired image features, such as object boundaries. The resulting evolution of the level set function is the gradient flow that minimizes the overall energy functional. The proposed variational level set formulation has three main advantages over the traditional level set formulations. First, a significantly larger time step can be used for numerically solving the evolution partial differential equation, and therefore speeds up the curve evolution. Second, the level set function can be initialized with general functions that are more efficient to construct and easier to use in practice than the widely used signed distance function. Third, the level set evolution in our formulation can be easily implemented by simple finite difference scheme and is computationally more efficient. The proposed algorithm has been applied to both simulated and real images with promising results.

Added

2026-09-16

Diffusion Models in Vision: A Survey

Diffusion Models in Vision: A Survey

Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, Mubarak Shah

OrganizationsUniversity of BucharestUniversity of Central Florida

Why you should read this

Systematizes the theoretical foundations of diffusion models across probabilistic, score-based, and stochastic differential equation frameworks while analyzing their vision applications, generative trade-offs, and computational bottlenecks.

Denoising diffusion models represent a recent emerging topic in computer vision, demonstrating remarkable results in the area of generative modeling. A diffusion model is a deep generative model that is based on two stages, a forward diffusion stage and a reverse diffusion stage. In the forward diffusion stage, the input data is gradually perturbed over several steps by adding Gaussian noise. In the reverse stage, a model is tasked at recovering the original input data by learning to gradually reverse the diffusion process, step by step. Diffusion models are widely appreciated for the quality and diversity of the generated samples, despite their known computational burdens, i.e. low speeds due to the high number of steps involved during sampling. In this survey, we provide a comprehensive review of articles on denoising diffusion models applied in vision, comprising both theoretical and practical contributions in the field. First, we identify and present three generic diffusion modeling frameworks, which are based on denoising diffusion probabilistic models, noise conditioned score networks, and stochastic differential equations. We further discuss the relations between diffusion models and other deep generative models, including variational auto-encoders, generative adversarial networks, energy-based models, autoregressive models and normalizing flows. Then, we introduce a multi-perspective categorization of diffusion models applied in computer vision. Finally, we illustrate the current limitations of diffusion models and envision some interesting directions for future research.

Added

2026-09-16

A Multiphase Level Set Framework for Image Segmentation Using the Mumford and Shah Model

A Multiphase Level Set Framework for Image Segmentation Using the Mumford and Shah Model

Luminita A. Vese, Tony F. Chan

OrganizationsUniversity of California, Los Angeles

Why you should read this

Develops an efficient multiphase level set formulation of the Mumford-Shah model that eliminates vacuum and overlap issues by construction, enabling the representation of complex topologies and triple junctions using only log₂ n level set functions.

We propose a new multiphase level set framework for image segmentation using the Mumford and Shah model, for piecewise constant and piecewise smooth optimal approximations. The proposed method is also a generalization of an active contour model without edges based 2-phase segmentation, developed by the authors earlier in T. Chan and L. Vese (1999. In Scale-Space’99, M. Nilsen et al. (Eds.), LNCS, vol. 1682, pp. 141–151) and T. Chan and L. Vese (2001. IEEE-IP, 10(2):266–277). The multiphase level set formulation is new and of interest on its own: by construction, it automatically avoids the problems of vacuum and overlap; it needs only log n level set functions for n phases in the piecewise constant case; it can represent boundaries with complex topologies, including triple junctions; in the piecewise smooth case, only two level set functions formally suffice to represent any partition, based on The Four-Color Theorem. Finally, we validate the proposed models by numerical results for signal and image denoising and segmentation, implemented using the Osher and Sethian level set method.

Added

2026-09-14