keyword
image segmentation
Image segmentation is a fundamental computer vision and digital image processing task that involves partitioning a digital image into multiple distinct regions, segments, or sets of pixels to simplify its representation and facilitate detailed visual analysis. By assigning a specific class label or object identity to each pixel, the process enables machines to locate boundaries and comprehend the spatial extent of objects and background elements within a scene. Common formulations include semantic segmentation, which classifies pixels into predefined categories; instance segmentation, which delineates individual object instances; and panoptic segmentation, which combines both approaches. Methodologies for performing segmentation range from classical algorithms based on edge detection, thresholding, clustering, and graph cuts to modern deep learning architectures, such as fully convolutional networks, vision transformers, and prompt-driven multimodal models.
32 items

Image Segmentation Using Text and Image Prompts
Timo Lüddecke, Alexander S. Ecker
Why you should read this
Proposes CLIPSeg, a unified system that adds a lightweight transformer-based decoder to a frozen CLIP model to perform zero-shot, one-shot, and referring expression image segmentation using arbitrary text or visual prompts at test time.
Image segmentation is usually addressed by training a model for a fixed set of object classes. Incorporating additional classes or more complex queries later is expensive as it requires re-training the model on a dataset that encompasses these expressions. Here we propose a system that can generate image segmentations based on arbitrary prompts at test time. A prompt can be either a text or an image. This approach enables us to create a unified model (trained once) for three common segmentation tasks, which come with distinct challenges: referring expression segmentation, zero-shot segmentation and one-shot segmentation. We build upon the CLIP model as a backbone which we extend with a transformer-based decoder that enables dense prediction. After training on an extended version of the PhraseCut dataset, our system generates a binary segmentation map for an image based on a free-text prompt or on an additional image expressing the query. We analyze different variants of the latter image-based prompts in detail. This novel hybrid input allows for dynamic adaptation not only to the three segmentation tasks mentioned above, but to any binary segmentation task where a text or image query can be formulated. Finally, we find our system to adapt well to generalized queries involving affordances or properties. Code is available at https://eckerlab.org/code/clipseg
Added
2026-10-05

LISA: Reasoning Segmentation via Large Language Model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, Jiaya Jia
Why you should read this
Introduces the task of reasoning segmentation alongside LISA, an end-to-end multimodal architecture that decodes special text tokens into binary masks to segment targets specified by implicit, knowledge-intensive user instructions.
Although perception systems have made remarkable advancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target objects before executing visual recognition tasks. Such systems cannot actively reason and comprehend implicit user intention. In this work, we propose a new segmentation task — reasoning segmentation. The task is designed to output a segmentation mask given a complex and implicit query text. Furthermore, we establish a benchmark comprising over one thousand image-instruction-mask data samples, incorporating intricate reasoning and world knowledge for evaluation purposes. Finally, we present LISA: large Language Instructed Segmentation Assistant, which inherits the language generation capabilities of multimodal Large Language Models (LLMs) while also possessing the ability to produce segmentation masks. We expand the original vocabulary with a <SEG> token and propose the embedding-as-mask paradigm to unlock the segmentation capability. Remarkably, LISA can handle cases involving complex reasoning and world knowledge. Also, it demonstrates robust zero-shot capability when trained exclusively on reasoning-free datasets. In addition, fine-tuning the model with merely 239 reasoning segmentation data samples results in further performance enhancement. Both quantitative and qualitative experiments show our method effectively unlocks new reasoning segmentation capabilities for multimodal LLMs. Code, models, and data are available at github.com/dvlab-research/LISA.
Added
2026-10-05

Language-driven Semantic Segmentation
Boyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun, René Ranftl
Why you should read this
Proposes LSeg, a model that aligns per-pixel visual embeddings with text representations through contrastive learning, enabling zero-shot semantic segmentation of arbitrary unseen categories without additional training.
We present LSeg, a novel model for language-driven semantic image segmentation. LSeg uses a text encoder to compute embeddings of descriptive input labels (e.g., "grass" or "building") together with a transformer-based image encoder that computes dense per-pixel embeddings of the input image. The image encoder is trained with a contrastive objective to align pixel embeddings to the text embedding of the corresponding semantic class. The text embeddings provide a flexible label representation in which semantically similar labels map to similar regions in the embedding space (e.g., "cat" and "furry"). This allows LSeg to generalize to previously unseen categories at test time, without retraining or even requiring a single additional training sample. We demonstrate that our approach achieves highly competitive zero-shot performance compared to existing zero- and few-shot semantic segmentation methods, and even matches the accuracy of traditional segmentation algorithms when a fixed label set is provided. Code and demo are available at this https URL.
Added
2026-10-05

Tuning Computer Vision Models With Task Rewards
André Susano Pinto, Alexander Kolesnikov, Yuge Shi, Lucas Beyer, Xiaohua Zhai
Why you should read this
Demonstrates that fine-tuning pretrained vision models with REINFORCE using task-specific rewards directly optimizes non-differentiable evaluation metrics across diverse tasks including object detection, panoptic segmentation, and image colorization without task-specific architectural modifications.
Misalignment between model predictions and intended usage can be detrimental for the deployment of computer vision models. The issue is exacerbated when the task involves complex structured outputs, as it becomes harder to design procedures which address this misalignment. In natural language processing, this is often addressed using reinforcement learning techniques that align models with a task reward. We adopt this approach and show its surprising effectiveness to improve generic models pretrained to imitate example outputs across multiple computer vision tasks, such as object detection, panoptic segmentation, colorization and image captioning. We believe this approach has the potential to be widely useful for better aligning models with a diverse range of computer vision tasks.
Added
2026-09-26

Adaptive Early-Learning Correction for Segmentation from Noisy Annotations
Sheng Liu, Kangning Liu, Weicheng Zhu, Yiqiu Shen, Carlos Fernandez-Granda
Why you should read this
Proposes an adaptive label correction method that prevents deep segmentation models from memorizing noisy annotations by independently monitoring category-specific early-learning dynamics and enforcing spatial consistency.
Added
2026-09-26

GSVA: Generalized Segmentation via Multimodal Large Language Models
Zhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan, Shiji Song, Gao Huang
Why you should read this
Presents a multimodal framework that enables language models to segment multiple target objects simultaneously and explicitly reject non-existent targets using dedicated tokens, setting a new state of the art on the gRefCOCO benchmark.
Generalized Referring Expression Segmentation (GRES) extends the scope of classic RES to refer to multiple objects in one expression or identify the empty targets absent in the image. GRES poses challenges in modeling the complex spatial relationships of the instances in the image and identifying non-existing referents. Multimodal Large Language Models (MLLMs) have recently shown tremendous progress in these complicated vision-language tasks. Connecting Large Language Models (LLMs) and vision models, MLLMs are proficient in understanding contexts with visual inputs. Among them, LISA, as a representative, adopts a special [SEG] token to prompt a segmentation mask decoder, e.g., SAM, to enable MLLMs in the RES task. However, existing solutions to GRES remain unsatisfactory since current segmentation MLLMs cannot correctly handle the cases where users might reference multiple subjects in a singular prompt or provide descriptions incongruent with any image target. In this paper, we propose Generalized Segmentation Vision Assistant (GSVA) to address this gap. Specifically, GSVA reuses the [SEG] token to prompt the segmentation model towards supporting multiple mask references simultaneously and innovatively learns to generate a [REJ] token to reject the null targets explicitly. Experiments validate GSVA’s efficacy in resolving the GRES issue, marking a notable enhancement and setting a new record on the GRES benchmark gRefCOCO dataset. GSVA also proves effective across various classic referring segmentation and comprehension tasks. Code is available at https://github.com/LeapLabTHU/GSVA.
Added
2026-09-26

Medical Image Analysis: Progress over Two Decades and the Challenges Ahead
J. Duncan, N. Ayache
Why you should read this
Synthesizes twenty years of foundational advances in medical image analysis across segmentation, registration, motion tracking, and image-guided surgery while defining critical open problems for integrating biomechanical models and physiological data.
Added
2026-09-25

An Optimal Graph Theoretic Approach to Data Clustering: Theory and Its Application to Image Segmentation
Zhenyu Wu, R. Leahy
Why you should read this
Develops a scalable graph-theoretic clustering method using subgraph condensation and equivalent trees to find globally optimal minimum cuts, guaranteeing closed boundary contours in image segmentation across hundreds of thousands of vertices.
A novel graph theoretic approach for data clustering is presented and its application to the image segmentation problem is demonstrated. The data to be clustered are represented by an undirected adjacency graph G with arc capacities assigned to reflect the similarity between the linked vertices. Clustering is achieved by removing arcs of G to form mutually exclusive subgraphs such that the largest inter-subgraph maximum flow is minimized. For graphs of moderate size (~ 2000 vertices), the optimal solution is obtained through partitioning a flow and cut equivalent tree of G, which can be efficiently constructed using the Gomory-Hu algorithm. However for larger graphs this approach is impractical. New theorems for subgraph condensation are derived and are then used to develop a fast algorithm which hierarchically constructs and partitions a partially equivalent tree of much reduced size. This algorithm results in an optimal solution equivalent to that obtained by partitioning the complete equivalent tree and is able to handle very large graphs with several hundred thousand vertices. The new clustering algorithm is applied to the image segmentation problem. The segmentation is achieved by effectively searching for closed contours of edge elements (equivalent to minimum cuts in G), which consist mostly of strong edges, while rejecting contours containing isolated strong edges. This method is able to accurately locate region boundaries and at the same time guarantees the formation of closed edge contours.
Added
2026-09-25

Contour and Texture Analysis for Image Segmentation
J. Malik, Serge J. Belongie, Thomas K. Leung, Jianbo Shi
Why you should read this
Proposes a unified image segmentation framework that combines intervening contour cues and texton-based texture analysis through an adaptive gating mechanism within the normalized cuts graph-partitioning algorithm.
This paper provides an algorithm for partitioning grayscale images into disjoint regions of coherent brightness and texture. Natural images contain both textured and untextured regions, so the cues of contour and texture differences are exploited simultaneously. Contours are treated in the intervening contour framework, while texture is analyzed using textons. Each of these cues has a domain of applicability, so to facilitate cue combination we introduce a gating operator based on the texturedness of the neighborhood at a pixel. Having obtained a local measure of how likely two nearby pixels are to belong to the same region, we use the spectral graph theoretic framework of normalized cuts to find partitions of the image into regions of coherent texture and brightness. Experimental results on a wide range of images are shown.
Added
2026-09-25

Kernel k-means: spectral clustering and normalized cuts
Inderjit S. Dhillon, Yuqiang Guan, Brian Kulis
Why you should read this
Proves a theoretical equivalence between weighted kernel k-means and spectral clustering objectives, enabling graph-based normalized cuts to be minimized through efficient iterative algorithms without relying on computationally expensive eigenvector calculations.
Kernel k-means and spectral clustering have both been used to identify clusters that are non-linearly separable in input space. Despite significant research, these methods have remained only loosely related. In this paper, we give an explicit theoretical connection between them. We show the generality of the weighted kernel k-means objective function, and derive the spectral clustering objective of normalized cut as a special case. Given a positive definite similarity matrix, our results lead to a novel weighted kernel k-means algorithm that monotonically decreases the normalized cut. This has important implications: a) eigenvector-based algorithms, which can be computationally prohibitive, are not essential for minimizing normalized cuts, b) various techniques, such as local search and acceleration schemes, may be used to improve the quality as well as speed of kernel k-means. Finally, we present results on several interesting data sets, including diametrical clustering of large gene-expression matrices and a handwriting recognition data set.
Added
2026-09-25

Spectral grouping using the Nystrom method
Charless C. Fowlkes, Serge J. Belongie, Fan Chung, Jitendra Malik
Why you should read this
Proposes using the Nyström method to extrapolate spectral clustering solutions from a small subset of sample points, scaling image and video segmentation linearly with resolution while drastically cutting computational and memory costs.
Spectral graph theoretic methods have recently shown great promise for the problem of image segmentation. However, due to the computational demands of these approaches, applications to large problems such as spatiotemporal data and high resolution imagery have been slow to appear. The contribution of this paper is a method that substantially reduces the computational requirements of grouping algorithms based on spectral partitioning making it feasible to apply them to very large grouping problems. Our approach is based on a technique for the numerical solution of eigenfunction problems known as the Nyström method. This method allows one to extrapolate the complete grouping solution using only a small number of samples. In doing so, we leverage the fact that there are far fewer coherent groups in a scene than pixels.
Added
2026-09-25

Learning with Hypergraphs: Clustering, Classification, and Embedding
Dengyong Zhou, Jiayuan Huang, B. Schölkopf
Why you should read this
Generalizes spectral graph theory to hypergraphs by formulating normalized hypergraph cuts, Laplacians, and random walks to enable higher-order relational clustering, embedding, and transductive classification without losing multi-object structural information.
We usually endow the investigated objects with pairwise relationships, which can be illustrated as graphs. In many real-world problems, however, relationships among the objects of our interest are more complex than pairwise. Naively squeezing the complex relationships into pairwise ones will inevitably lead to loss of information which can be expected valuable for our learning tasks however. There we consider using hypergraphs instead to completely represent complex relationships among the objects of our interest, and thus the problem of learning with hypergraphs arises. Our main contribution in this paper is to generalize the powerful methodology of spectral clustering which originally operates on undirected graphs to hypergraphs, and further develop algorithms for hypergraph embedding and transductive classification on the basis of the spectral hypergraph clustering approach. Our experiments on a number of benchmarks showed the advantages of hypergraphs over usual graphs.
Added
2026-09-24

Blobworld: Image Segmentation Using Expectation-Maximization and Its Application to Image Querying
C. Carson, Serge J. Belongie, H. Greenspan, Jitendra Malik
Why you should read this
Introduces a region-based image retrieval framework that uses Expectation-Maximization clustering across joint color, texture, and spatial features to enable object-level queries that significantly outperform traditional global histogram methods.
Retrieving images from large and varied collections using image content as a key is a challenging and important problem. We present a new image representation which provides a transformation from the raw pixel data to a small set of image regions which are coherent in color and texture. This “Blobworld” representation is created by clustering pixels in a joint color-texture-position feature space. The segmentation algorithm is fully automatic and has been run on a collection of 10,000 natural images. We describe a system that uses the Blobworld representation to retrieve images from this collection. An important aspect of the system is that the user is allowed to view the internal representation of the submitted image and the query results. Similar systems do not offer the user this view into the workings of the system; consequently, query results from these systems can be inexplicable, despite the availability of knobs for adjusting the similarity metrics. By finding image regions which roughly correspond to objects, we allow querying at the level of objects rather than global image properties. We present results indicating that querying for distinctive objects using Blobworld produces significantly higher precision than does querying using color and texture histograms of the entire image.
Added
2026-09-24

Matching Words and Pictures
Kobus Barnard, Pinar Duygulu, David Forsyth, Nando de Freitas, David Blei, Michael I. Jordan
Why you should read this
Proposes probabilistic and statistical translation models to learn the joint distribution of segmented image regions and words, establishing a foundational framework for automatic image annotation, object recognition, and text-based image retrieval.
We present a new approach for modeling multi-modal data sets, focusing on the specific case of segmented images with associated text. Learning the joint distribution of image regions and words has many applications. We consider in detail predicting words associated with whole images (auto-annotation) and corresponding to particular image regions (region naming). Auto-annotation might help organize and access large collections of images. Region naming is a model of object recognition as a process of translating image regions to words, much as one might translate from one language to another. Learning the relationships between image regions and semantic correlates (words) is an interesting example of multi-modal data mining, particularly because it is typically hard to apply data mining techniques to collections of images. We develop a number of models for the joint distribution of image regions and words, including several which explicitly learn the correspondence between regions and words. We study multi-modal and correspondence extensions to Hofmann’s hierarchical clustering/aspect model, a translation model adapted from statistical machine translation (Brown et al.), and a multi-modal extension to mixture of latent Dirichlet allocation (MoM-LDA). All models are assessed using a large collection of annotated images of real
Added
2026-09-18

A Closed-Form Solution to Natural Image Matting
Anat Levin, Dani Lischinski, Yair Weiss
Why you should read this
Presents a closed-form solution to natural image matting that analytically eliminates unknown foreground and background colors, enabling globally optimal alpha matte extraction from sparse user scribbles by solving a single sparse linear system.
Interactive digital matting, the process of extracting a foreground object from an image based on limited user input, is an important task in image and video editing. From a computer vision perspective, this task is extremely challenging because it is massively ill-posed — at each pixel we must estimate the foreground and the background colors, as well as the foreground opacity (“alpha matte”) from a single color measurement. Current approaches either restrict the estimation to a small part of the image, estimating foreground and background colors based on nearby pixels where they are known, or perform iterative nonlinear estimation by alternating foreground and background color estimation with alpha estimation. In this paper we present a closed form solution to natural image matting. We derive a cost function from local smoothness assumptions on foreground and background colors, and show that in the resulting expression it is possible to analytically eliminate the foreground and background colors to obtain a quadratic cost function in alpha. This allows us to find the globally optimal alpha matte by solving a sparse linear system of equations. Furthermore, the closed form formula allows us to predict the properties of the solution by analyzing the eigenvectors of a sparse matrix, closely related to matrices used in spectral image segmentation algorithms. We show that high quality mattes can be obtained on natural images from a small amount of user input.
Added
2026-09-16

Level set evolution without re-initialization: a new variational formulation
Chunming Li, Chenyang Xu, C. Gui, M. Fox
Why you should read this
Proposes a distance-regularizing variational formulation for active contours that maintains the level set function as a signed distance profile throughout evolution, eliminating costly re-initialization steps while enabling faster numerical convergence with simple finite difference schemes.
In this paper, we present a new variational formulation for geometric active contours that forces the level set function to be close to a signed distance function, and therefore completely eliminates the need of the costly re-initialization procedure. Our variational formulation consists of an internal energy term that penalizes the deviation of the level set function from a signed distance function, and an external energy term that drives the motion of the zero level set toward the desired image features, such as object boundaries. The resulting evolution of the level set function is the gradient flow that minimizes the overall energy functional. The proposed variational level set formulation has three main advantages over the traditional level set formulations. First, a significantly larger time step can be used for numerically solving the evolution partial differential equation, and therefore speeds up the curve evolution. Second, the level set function can be initialized with general functions that are more efficient to construct and easier to use in practice than the widely used signed distance function. Third, the level set evolution in our formulation can be easily implemented by simple finite difference scheme and is computationally more efficient. The proposed algorithm has been applied to both simulated and real images with promising results.
Added
2026-09-16

Diffusion Models in Vision: A Survey
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, Mubarak Shah
Why you should read this
Systematizes the theoretical foundations of diffusion models across probabilistic, score-based, and stochastic differential equation frameworks while analyzing their vision applications, generative trade-offs, and computational bottlenecks.
Denoising diffusion models represent a recent emerging topic in computer vision, demonstrating remarkable results in the area of generative modeling. A diffusion model is a deep generative model that is based on two stages, a forward diffusion stage and a reverse diffusion stage. In the forward diffusion stage, the input data is gradually perturbed over several steps by adding Gaussian noise. In the reverse stage, a model is tasked at recovering the original input data by learning to gradually reverse the diffusion process, step by step. Diffusion models are widely appreciated for the quality and diversity of the generated samples, despite their known computational burdens, i.e. low speeds due to the high number of steps involved during sampling. In this survey, we provide a comprehensive review of articles on denoising diffusion models applied in vision, comprising both theoretical and practical contributions in the field. First, we identify and present three generic diffusion modeling frameworks, which are based on denoising diffusion probabilistic models, noise conditioned score networks, and stochastic differential equations. We further discuss the relations between diffusion models and other deep generative models, including variational auto-encoders, generative adversarial networks, energy-based models, autoregressive models and normalizing flows. Then, we introduce a multi-perspective categorization of diffusion models applied in computer vision. Finally, we illustrate the current limitations of diffusion models and envision some interesting directions for future research.
Added
2026-09-16

Self-Tuning Spectral Clustering
Lihi Zelnik-Manor, P. Perona
Why you should read this
Introduces a self-tuning spectral clustering method that computes local scaling for each data point and infers the number of groups directly from eigenvector structure, effectively clustering multi-scale, cluttered datasets without manual parameter selection or randomized k-means initialization.
We study a number of open issues in spectral clustering: (i) Selecting the appropriate scale of analysis, (ii) Handling multi-scale data, (iii) Clustering with irregular background clutter, and, (iv) Finding automatically the number of groups. We first propose that a ‘local’ scale should be used to compute the affinity between each pair of points. This local scaling leads to better clustering especially when the data includes multiple scales and when the clusters are placed within a cluttered background. We further suggest exploiting the structure of the eigenvectors to infer automatically the number of groups. This leads to a new algorithm in which the final randomly initialized k-means stage is eliminated.
Added
2026-09-14

A Multiphase Level Set Framework for Image Segmentation Using the Mumford and Shah Model
Luminita A. Vese, Tony F. Chan
Why you should read this
Develops an efficient multiphase level set formulation of the Mumford-Shah model that eliminates vacuum and overlap issues by construction, enabling the representation of complex topologies and triple junctions using only log₂ n level set functions.
We propose a new multiphase level set framework for image segmentation using the Mumford and Shah model, for piecewise constant and piecewise smooth optimal approximations. The proposed method is also a generalization of an active contour model without edges based 2-phase segmentation, developed by the authors earlier in T. Chan and L. Vese (1999. In Scale-Space’99, M. Nilsen et al. (Eds.), LNCS, vol. 1682, pp. 141–151) and T. Chan and L. Vese (2001. IEEE-IP, 10(2):266–277). The multiphase level set formulation is new and of interest on its own: by construction, it automatically avoids the problems of vacuum and overlap; it needs only log n level set functions for n phases in the piecewise constant case; it can represent boundaries with complex topologies, including triple junctions; in the piecewise smooth case, only two level set functions formally suffice to represent any partition, based on The Four-Color Theorem. Finally, we validate the proposed models by numerical results for signal and image denoising and segmentation, implemented using the Osher and Sethian level set method.
Added
2026-09-14

Random Walks for Image Segmentation
L. Grady
Why you should read this
Introduces an interactive multilabel segmentation algorithm that maps random walker hitting probabilities to the combinatorial Dirichlet problem, enabling exact, rapid boundary localization across arbitrary graphs through standard sparse linear solvers.
A novel method is proposed for performing multilabel, interactive image segmentation. Given a small number of pixels with user-defined (or predefined) labels, one can analytically and quickly determine the probability that a random walker starting at each unlabeled pixel will first reach one of the prelabeled pixels. By assigning each pixel to the label for which the greatest probability is calculated, a high-quality image segmentation may be obtained. Theoretical properties of this algorithm are developed along with the corresponding connections to discrete potential theory and electrical circuits. This algorithm is formulated in discrete space (i.e., on a graph) using combinatorial analogues of standard operators and principles from continuous potential theory, allowing it to be applied in arbitrary dimension on arbitrary graphs.
Added
2026-09-14
