keyword
edge detection
Edge detection is an image processing and computer vision technique used to identify boundaries and significant structural contours within a digital image by locating points of sharp discontinuity in visual properties such as brightness, color, or texture. These localized transitions typically correspond to physical object boundaries, surface orientation changes, shadows, or variations in material properties. Methodologies for detecting edges range from classic gradient-based differential operators, such as the Sobel, Prewitt, and Canny detectors, to multiscale steerable filters and deep convolutional neural networks. By filtering out less informative background data while preserving essential structural features, edge detection significantly reduces data complexity and serves as a fundamental preliminary step for higher-level visual tasks, including image segmentation, object recognition, motion tracking, and three-dimensional reconstruction.
11 items

An Optimal Graph Theoretic Approach to Data Clustering: Theory and Its Application to Image Segmentation
Zhenyu Wu, R. Leahy
Why you should read this
Develops a scalable graph-theoretic clustering method using subgraph condensation and equivalent trees to find globally optimal minimum cuts, guaranteeing closed boundary contours in image segmentation across hundreds of thousands of vertices.
A novel graph theoretic approach for data clustering is presented and its application to the image segmentation problem is demonstrated. The data to be clustered are represented by an undirected adjacency graph G with arc capacities assigned to reflect the similarity between the linked vertices. Clustering is achieved by removing arcs of G to form mutually exclusive subgraphs such that the largest inter-subgraph maximum flow is minimized. For graphs of moderate size (~ 2000 vertices), the optimal solution is obtained through partitioning a flow and cut equivalent tree of G, which can be efficiently constructed using the Gomory-Hu algorithm. However for larger graphs this approach is impractical. New theorems for subgraph condensation are derived and are then used to develop a fast algorithm which hierarchically constructs and partitions a partially equivalent tree of much reduced size. This algorithm results in an optimal solution equivalent to that obtained by partitioning the complete equivalent tree and is able to handle very large graphs with several hundred thousand vertices. The new clustering algorithm is applied to the image segmentation problem. The segmentation is achieved by effectively searching for closed contours of edge elements (equivalent to minimum cuts in G), which consist mostly of strong edges, while rejecting contours containing isolated strong edges. This method is able to accurately locate region boundaries and at the same time guarantees the formation of closed edge contours.
Added
2026-09-25

Contour and Texture Analysis for Image Segmentation
J. Malik, Serge J. Belongie, Thomas K. Leung, Jianbo Shi
Why you should read this
Proposes a unified image segmentation framework that combines intervening contour cues and texton-based texture analysis through an adaptive gating mechanism within the normalized cuts graph-partitioning algorithm.
This paper provides an algorithm for partitioning grayscale images into disjoint regions of coherent brightness and texture. Natural images contain both textured and untextured regions, so the cues of contour and texture differences are exploited simultaneously. Contours are treated in the intervening contour framework, while texture is analyzed using textons. Each of these cues has a domain of applicability, so to facilitate cue combination we introduce a gating operator based on the texturedness of the neighborhood at a pixel. Having obtained a local measure of how likely two nearby pixels are to belong to the same region, we use the spectral graph theoretic framework of normalized cuts to find partitions of the image into regions of coherent texture and brightness. Experimental results on a wide range of images are shown.
Added
2026-09-25

Face Recognition: The Problem of Compensating for Changes in Illumination Direction
Yael Adini, Y. Moses, S. Ullman
Why you should read this
Demonstrates through systematic empirical testing that standard illumination-invariant image representations such as edge maps, intensity derivatives, and 2D Gabor filters fail to overcome lighting direction changes in face recognition, establishing the need for richer three-dimensional or model-based recognition strategies.
A face recognition system must recognize a face from a novel image despite the variations between images of the same face. A common approach to overcoming image variations because of changes in the illumination conditions is to use image representations that are relatively insensitive to these variations. Examples of such representations are edge maps, image intensity derivatives, and images convolved with 2D Gabor-like filters. Here we present an empirical study that evaluates the sensitivity of these representations to changes in illumination, as well as viewpoint and facial expression. Our findings indicated that none of the representations considered is sufficient by itself to overcome image variations because of a change in the direction of illumination. Similar results were obtained for changes due to viewpoint and expression. Image representations that emphasized the horizontal features were found to be less sensitive to changes in the direction of illumination. However, systems based only on such representations failed to recognize up to 20 percent of the faces in our database. Humans performed considerably better under the same conditions. We discuss possible reasons for this superiority and alternative methods for overcoming illumination effects in recognition.
Added
2026-09-25

Deep Layer Aggregation
F. Yu, Dequan Wang, Trevor Darrell
Why you should read this
Introduces iterative and hierarchical deep layer aggregation structures that merge feature hierarchies across stages and resolutions, achieving higher visual recognition accuracy with fewer parameters than standard architectures like ResNet and DenseNet.
Visual recognition requires rich representations that span levels from low to high, scales from small to large, and resolutions from fine to coarse. Even with the depth of features in a convolutional network, a layer in isolation is not enough: compounding and aggregating these representations improves inference of what and where. Architectural efforts are exploring many dimensions for network backbones, designing deeper or wider architectures, but how to best aggregate layers and blocks across a network deserves further attention. Although skip connections have been incorporated to combine layers, these connections have been “shallow” themselves, and only fuse by simple, one-step operations. We augment standard architectures with deeper aggregation to better fuse information across layers. Our deep layer aggregation structures iteratively and hierarchically merge the feature hierarchy to make networks with better accuracy and fewer parameters. Experiments across architectures and tasks show that deep layer aggregation improves recognition and resolution compared to existing branching and merging schemes.
Added
2026-09-24

Finite-Element Methods for Active Contour Models and Balloons for 2-D and 3-D Images
L. Cohen, I. Cohen
Why you should read this
Presents a three-dimensional generalization of the balloon deformable surface model and implements a finite element framework that achieves faster convergence and superior numerical stability for volumetric medical image segmentation.
The use of energy-minimizing curves, known as "snakes" to extract features of interest in images has been introduced by Kass, Witkin and Terzopoulos [23]. A balloon model was introduced in [12] as a way to generalize and solve some of the problems encountered with the original method. We present a 3D generalization of the balloon model as a 3D deformable surface, which evolves in 3D images. It is deformed under the action of internal and external forces attracting the surface toward detected edgels by means of an attraction potential. We also show properties of energy-minimizing surfaces concerning their relationship with 3D edge points. To solve the minimization problem for a surface, two simplified approaches are shown first, defining a 3D surface as a series of 2D planar curves. Then, after comparing Finite Element Method and Finite Difference Method in the 2D problem, we solve the 3D model using the Finite Element Method yielding greater stability and faster convergence. We have applied this model for segmenting magnetic resonance images.
Added
2026-09-24

ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data
Foivos I. Diakogiannis, François Waldner, Peter Caccetta, Chen Wu
Why you should read this
Introduces ResUNet-a, a deep learning architecture that couples conditioned multi-task learning and atrous pyramid pooling with a modified Generalized Dice loss to tackle severe class imbalance in high-resolution aerial image segmentation.
Scene understanding of high resolution aerial images is of great importance for the task of automated monitoring in various remote sensing applications. Due to the large within-class and small between-class variance in pixel values of objects of interest, this remains a challenging task. In recent years, deep convolutional neural networks have started being used in remote sensing applications and demonstrate state of the art performance for pixel level classification of objects. \textcolor{black}{Here we propose a reliable framework for performant results for the task of semantic segmentation of monotemporal very high resolution aerial images. Our framework consists of a novel deep learning architecture, ResUNet-a, and a novel loss function based on the Dice loss. ResUNet-a uses a UNet encoder/decoder backbone, in combination with residual connections, atrous convolutions, pyramid scene parsing pooling and multi-tasking inference. ResUNet-a infers sequentially the boundary of the objects, the distance transform of the segmentation mask, the segmentation mask and a colored reconstruction of the input. Each of the tasks is conditioned on the inference of the previous ones, thus establishing a conditioned relationship between the various tasks, as this is described through the architecture's computation graph. We analyse the performance of several flavours of the Generalized Dice loss for semantic segmentation, and we introduce a novel variant loss function for semantic segmentation of objects that has excellent convergence properties and behaves well even under the presence of highly imbalanced classes.} The performance of our modeling framework is evaluated on the ISPRS 2D Potsdam dataset. Results show state-of-the-art performance with an average F1 score of 92.9\% over all classes for our best model.
Added
2026-09-17

Feature Detection with Automatic Scale Selection
Tony Lindeberg
Why you should read this
Develops a foundational scale-space framework for automatic scale selection by identifying local extrema of gamma-normalized Gaussian derivatives, enabling vision algorithms to adaptively detect and localize image features such as blobs, corners, and edges across varying scales without manual parameter tuning.
The fact that objects in the world appear in different ways depending on the scale of observation has important implications if one aims at describing them. It shows that the notion of scale is of utmost importance when processing unknown measurement data by automatic methods. In their seminal works, Witkin (1983) and Koenderink (1984) proposed to approach this problem by representing image structures at different scales in a so-called scale-space representation. Traditional scale-space theory building on this work, however, does not address the problem of how to select local appropriate scales for further analysis. This article proposes a systematic methodology for dealing with this problem. A framework is proposed for generating hypotheses about interesting scale levels in image data, based on a general principle stating that local extrema over scales of different combinations of γ-normalized derivatives are likely candidates to correspond to interesting structures. Specifically, it is shown how this idea can be used as a major mechanism in algorithms for automatic scale selection, which adapt the local scales of processing to the local image structure. Support for the proposed approach is given in terms of a general theoretical investigation of the behaviour of the scale selection method under rescalings of the input pattern and by experiments on real-world and synthetic data. Support is also given by a detailed analysis of how different types of feature detectors perform when integrated with a scale selection mechanism and then applied to characteristic model patterns. Specifically, it is described in detail how the proposed methodology applies to the problems of blob detection, junction detection, edge detection, ridge detection and local frequency estimation. In many computer vision applications, the poor performance of the low-level vision modules constitutes a major bottle-neck. It will be argued that the inclusion of mechanisms for automatic scale selection is essential if we are to construct vision systems to analyse complex unknown environments.
Added
2026-09-12

The Design and Use of Steerable Filters
W. Freeman, E. Adelson
Why you should read this
Introduces the mathematical theory and implementation of steerable filters, demonstrating how filters at arbitrary orientations can be synthesized from a compact set of basis filters to efficiently analyze local orientation, phase, and scale across diverse computer vision tasks.
Oriented filters are useful in many early vision and image processing tasks. One often needs to apply the same filter, rotated to different angles under adaptive control, or wishes to calculate the filter response at various orientations. We present an efficient architecture to synthesize filters of arbitrary orientations from linear combinations of basis filters, allowing one to adaptively “steer” a filter to any orientation, and to determine analytically the filter output as a function of orientation. Steerable filters may be designed in quadrature pairs to allow adaptive control over phase as well as orientation. We show how to design and steer the filters and present examples of their use in several tasks: the analysis of orientation and phase, angularly adaptive filtering, edge detection, and shape from shading. One can also build a self-similar steerable pyramid representation. The same concepts can be generalized to the design of 3-D steerable filters, which should be useful in the analysis of image sequences and volumetric data.
Added
2026-09-11

Seeded Region Growing
Rolf Adams, L. Bischof
Why you should read this
Proposes seeded region growing, a parameter-free segmentation algorithm that rapidly partitions images by iteratively assimilating adjacent pixels based on statistical similarity to seed-initialized regions.
We present here a new algorithm for segmentation of intensity images which is robust, rapid, and free of tuning parameters. The method, however, requires the input of a number of seeds, either individual pixels or regions, which will control the formation of regions into which the image will be segmented. In this correspondence, we present the algorithm, discuss briefly its properties, and suggest two ways in which it can be employed, namely, by using manual seed selection or by automated procedures.
Added
2026-09-11

Image inpainting
Marcelo Bertalmio, Guillermo Sapiro, Vicent Caselles, Coloma Ballester
Why you should read this
Introduces an automatic digital inpainting algorithm that reconstructs missing or damaged image regions by propagating surrounding isophote lines inward, establishing a mathematical framework for photo restoration and object removal.
Inpainting, the technique of modifying an image in an undetectable form, is as ancient as art itself. The goals and applications of inpainting are numerous, from the restoration of damaged paintings and photographs to the removal/replacement of selected objects. In this paper, we introduce a novel algorithm for digital inpainting of still images that attempts to replicate the basic techniques used by professional restorators. After the user selects the regions to be restored, the algorithm automatically fills-in these regions with information surrounding them. The fill-in is done in such a way that isophote lines arriving at the regions’ boundaries are completed inside. In contrast with previous approaches, the technique here introduced does not require the user to specify where the novel information comes from. This is automatically done (and in a fast way), thereby allowing to simultaneously fill-in numerous regions containing completely different structures and surrounding backgrounds. In addition, no limitations are imposed on the topology of the region to be inpainted. Applications of this technique include the restoration of old photographs and damaged film; removal of superimposed text like dates, subtitles, or publicity; and the removal of entire objects from the image like microphones or wires in special effects.
Added
2026-09-10

Segment Anything
Alexander M. Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan‐Yen Lo, Piotr Dollár, Ross Girshick
Why you should read this
Establishes the first tangible "foundation model" for image segmentation that supports zero-shot prompting via points, boxes, or text.
We introduce the Segment Anything (SA) project: a new task, model, and dataset for image segmentation. Using our efficient model in a data collection loop, we built the largest segmentation dataset to date (by far), with over 1 billion masks on 11M licensed and privacy respecting images. The model is designed and trained to be promptable, so it can transfer zero-shot to new image distributions and tasks. We evaluate its capabilities on numerous tasks and find that its zero-shot performance is impressive – often competitive with or even superior to prior fully supervised results. We are releasing the Segment Anything Model (SAM) and corresponding dataset (SA-1B) of 1B masks and 11M images at segment-anything.com to foster research into foundation models for computer vision. We recommend reading the full paper at: arxiv.org/abs/2304.02643.
Added
2026-01-28
