Built independently by an author, for readers. Read the story and support ChapterPal

keyword

edge detection

Edge detection is an image processing and computer vision technique used to identify boundaries and significant structural contours within a digital image by locating points of sharp discontinuity in visual properties such as brightness, color, or texture. These localized transitions typically correspond to physical object boundaries, surface orientation changes, shadows, or variations in material properties. Methodologies for detecting edges range from classic gradient-based differential operators, such as the Sobel, Prewitt, and Canny detectors, to multiscale steerable filters and deep convolutional neural networks. By filtering out less informative background data while preserving essential structural features, edge detection significantly reduces data complexity and serves as a fundamental preliminary step for higher-level visual tasks, including image segmentation, object recognition, motion tracking, and three-dimensional reconstruction.

11 items

An Optimal Graph Theoretic Approach to Data Clustering: Theory and Its Application to Image Segmentation

An Optimal Graph Theoretic Approach to Data Clustering: Theory and Its Application to Image Segmentation

Zhenyu Wu, R. Leahy

OrganizationsUniversity of PennsylvaniaUniversity of Southern California

Why you should read this

Develops a scalable graph-theoretic clustering method using subgraph condensation and equivalent trees to find globally optimal minimum cuts, guaranteeing closed boundary contours in image segmentation across hundreds of thousands of vertices.

A novel graph theoretic approach for data clustering is presented and its application to the image segmentation problem is demonstrated. The data to be clustered are represented by an undirected adjacency graph G with arc capacities assigned to reflect the similarity between the linked vertices. Clustering is achieved by removing arcs of G to form mutually exclusive subgraphs such that the largest inter-subgraph maximum flow is minimized. For graphs of moderate size (~ 2000 vertices), the optimal solution is obtained through partitioning a flow and cut equivalent tree of G, which can be efficiently constructed using the Gomory-Hu algorithm. However for larger graphs this approach is impractical. New theorems for subgraph condensation are derived and are then used to develop a fast algorithm which hierarchically constructs and partitions a partially equivalent tree of much reduced size. This algorithm results in an optimal solution equivalent to that obtained by partitioning the complete equivalent tree and is able to handle very large graphs with several hundred thousand vertices. The new clustering algorithm is applied to the image segmentation problem. The segmentation is achieved by effectively searching for closed contours of edge elements (equivalent to minimum cuts in G), which consist mostly of strong edges, while rejecting contours containing isolated strong edges. This method is able to accurately locate region boundaries and at the same time guarantees the formation of closed edge contours.

Added

2026-09-25

Face Recognition: The Problem of Compensating for Changes in Illumination Direction

Face Recognition: The Problem of Compensating for Changes in Illumination Direction

Yael Adini, Y. Moses, S. Ullman

OrganizationsWeizmann Institute of Science

Why you should read this

Demonstrates through systematic empirical testing that standard illumination-invariant image representations such as edge maps, intensity derivatives, and 2D Gabor filters fail to overcome lighting direction changes in face recognition, establishing the need for richer three-dimensional or model-based recognition strategies.

A face recognition system must recognize a face from a novel image despite the variations between images of the same face. A common approach to overcoming image variations because of changes in the illumination conditions is to use image representations that are relatively insensitive to these variations. Examples of such representations are edge maps, image intensity derivatives, and images convolved with 2D Gabor-like filters. Here we present an empirical study that evaluates the sensitivity of these representations to changes in illumination, as well as viewpoint and facial expression. Our findings indicated that none of the representations considered is sufficient by itself to overcome image variations because of a change in the direction of illumination. Similar results were obtained for changes due to viewpoint and expression. Image representations that emphasized the horizontal features were found to be less sensitive to changes in the direction of illumination. However, systems based only on such representations failed to recognize up to 20 percent of the faces in our database. Humans performed considerably better under the same conditions. We discuss possible reasons for this superiority and alternative methods for overcoming illumination effects in recognition.

Added

2026-09-25

Deep Layer Aggregation

Deep Layer Aggregation

F. Yu, Dequan Wang, Trevor Darrell

OrganizationsUniversity of California Berkeley

Why you should read this

Introduces iterative and hierarchical deep layer aggregation structures that merge feature hierarchies across stages and resolutions, achieving higher visual recognition accuracy with fewer parameters than standard architectures like ResNet and DenseNet.

Visual recognition requires rich representations that span levels from low to high, scales from small to large, and resolutions from fine to coarse. Even with the depth of features in a convolutional network, a layer in isolation is not enough: compounding and aggregating these representations improves inference of what and where. Architectural efforts are exploring many dimensions for network backbones, designing deeper or wider architectures, but how to best aggregate layers and blocks across a network deserves further attention. Although skip connections have been incorporated to combine layers, these connections have been “shallow” themselves, and only fuse by simple, one-step operations. We augment standard architectures with deeper aggregation to better fuse information across layers. Our deep layer aggregation structures iteratively and hierarchically merge the feature hierarchy to make networks with better accuracy and fewer parameters. Experiments across architectures and tasks show that deep layer aggregation improves recognition and resolution compared to existing branching and merging schemes.

Added

2026-09-24

Finite-Element Methods for Active Contour Models and Balloons for 2-D and 3-D Images

Finite-Element Methods for Active Contour Models and Balloons for 2-D and 3-D Images

L. Cohen, I. Cohen

OrganizationsCEREMADEINRIAParis Dauphine University

Why you should read this

Presents a three-dimensional generalization of the balloon deformable surface model and implements a finite element framework that achieves faster convergence and superior numerical stability for volumetric medical image segmentation.

The use of energy-minimizing curves, known as "snakes" to extract features of interest in images has been introduced by Kass, Witkin and Terzopoulos [23]. A balloon model was introduced in [12] as a way to generalize and solve some of the problems encountered with the original method. We present a 3D generalization of the balloon model as a 3D deformable surface, which evolves in 3D images. It is deformed under the action of internal and external forces attracting the surface toward detected edgels by means of an attraction potential. We also show properties of energy-minimizing surfaces concerning their relationship with 3D edge points. To solve the minimization problem for a surface, two simplified approaches are shown first, defining a 3D surface as a series of 2D planar curves. Then, after comparing Finite Element Method and Finite Difference Method in the 2D problem, we solve the 3D model using the Finite Element Method yielding greater stability and faster convergence. We have applied this model for segmenting magnetic resonance images.

Added

2026-09-24

ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data

ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data

Foivos I. Diakogiannis, François Waldner, Peter Caccetta, Chen Wu

OrganizationsCSIRO Agriculture & FoodCSIRO’s Data61University of Western Australia

Why you should read this

Introduces ResUNet-a, a deep learning architecture that couples conditioned multi-task learning and atrous pyramid pooling with a modified Generalized Dice loss to tackle severe class imbalance in high-resolution aerial image segmentation.

Scene understanding of high resolution aerial images is of great importance for the task of automated monitoring in various remote sensing applications. Due to the large within-class and small between-class variance in pixel values of objects of interest, this remains a challenging task. In recent years, deep convolutional neural networks have started being used in remote sensing applications and demonstrate state of the art performance for pixel level classification of objects. \textcolor{black}{Here we propose a reliable framework for performant results for the task of semantic segmentation of monotemporal very high resolution aerial images. Our framework consists of a novel deep learning architecture, ResUNet-a, and a novel loss function based on the Dice loss. ResUNet-a uses a UNet encoder/decoder backbone, in combination with residual connections, atrous convolutions, pyramid scene parsing pooling and multi-tasking inference. ResUNet-a infers sequentially the boundary of the objects, the distance transform of the segmentation mask, the segmentation mask and a colored reconstruction of the input. Each of the tasks is conditioned on the inference of the previous ones, thus establishing a conditioned relationship between the various tasks, as this is described through the architecture's computation graph. We analyse the performance of several flavours of the Generalized Dice loss for semantic segmentation, and we introduce a novel variant loss function for semantic segmentation of objects that has excellent convergence properties and behaves well even under the presence of highly imbalanced classes.} The performance of our modeling framework is evaluated on the ISPRS 2D Potsdam dataset. Results show state-of-the-art performance with an average F1 score of 92.9\% over all classes for our best model.

Added

2026-09-17

Feature Detection with Automatic Scale Selection

Feature Detection with Automatic Scale Selection

Tony Lindeberg

OrganizationsComputational Vision and Active Perception LaboratoryDepartment of Numerical Analysis and Computing ScienceKTH Royal Institute of Technology

Why you should read this

Develops a foundational scale-space framework for automatic scale selection by identifying local extrema of gamma-normalized Gaussian derivatives, enabling vision algorithms to adaptively detect and localize image features such as blobs, corners, and edges across varying scales without manual parameter tuning.

The fact that objects in the world appear in different ways depending on the scale of observation has important implications if one aims at describing them. It shows that the notion of scale is of utmost importance when processing unknown measurement data by automatic methods. In their seminal works, Witkin (1983) and Koenderink (1984) proposed to approach this problem by representing image structures at different scales in a so-called scale-space representation. Traditional scale-space theory building on this work, however, does not address the problem of how to select local appropriate scales for further analysis. This article proposes a systematic methodology for dealing with this problem. A framework is proposed for generating hypotheses about interesting scale levels in image data, based on a general principle stating that local extrema over scales of different combinations of γ-normalized derivatives are likely candidates to correspond to interesting structures. Specifically, it is shown how this idea can be used as a major mechanism in algorithms for automatic scale selection, which adapt the local scales of processing to the local image structure. Support for the proposed approach is given in terms of a general theoretical investigation of the behaviour of the scale selection method under rescalings of the input pattern and by experiments on real-world and synthetic data. Support is also given by a detailed analysis of how different types of feature detectors perform when integrated with a scale selection mechanism and then applied to characteristic model patterns. Specifically, it is described in detail how the proposed methodology applies to the problems of blob detection, junction detection, edge detection, ridge detection and local frequency estimation. In many computer vision applications, the poor performance of the low-level vision modules constitutes a major bottle-neck. It will be argued that the inclusion of mechanisms for automatic scale selection is essential if we are to construct vision systems to analyse complex unknown environments.

Added

2026-09-12

The Design and Use of Steerable Filters

The Design and Use of Steerable Filters

W. Freeman, E. Adelson

OrganizationsMassachusetts Institute of Technology

Why you should read this

Introduces the mathematical theory and implementation of steerable filters, demonstrating how filters at arbitrary orientations can be synthesized from a compact set of basis filters to efficiently analyze local orientation, phase, and scale across diverse computer vision tasks.

Oriented filters are useful in many early vision and image processing tasks. One often needs to apply the same filter, rotated to different angles under adaptive control, or wishes to calculate the filter response at various orientations. We present an efficient architecture to synthesize filters of arbitrary orientations from linear combinations of basis filters, allowing one to adaptively “steer” a filter to any orientation, and to determine analytically the filter output as a function of orientation. Steerable filters may be designed in quadrature pairs to allow adaptive control over phase as well as orientation. We show how to design and steer the filters and present examples of their use in several tasks: the analysis of orientation and phase, angularly adaptive filtering, edge detection, and shape from shading. One can also build a self-similar steerable pyramid representation. The same concepts can be generalized to the design of 3-D steerable filters, which should be useful in the analysis of image sequences and volumetric data.

Added

2026-09-11

Image inpainting

Image inpainting

Marcelo Bertalmio, Guillermo Sapiro, Vicent Caselles, Coloma Ballester

OrganizationsUniversitat Pompeu FabraUniversity of Minnesota

Why you should read this

Introduces an automatic digital inpainting algorithm that reconstructs missing or damaged image regions by propagating surrounding isophote lines inward, establishing a mathematical framework for photo restoration and object removal.

Inpainting, the technique of modifying an image in an undetectable form, is as ancient as art itself. The goals and applications of inpainting are numerous, from the restoration of damaged paintings and photographs to the removal/replacement of selected objects. In this paper, we introduce a novel algorithm for digital inpainting of still images that attempts to replicate the basic techniques used by professional restorators. After the user selects the regions to be restored, the algorithm automatically fills-in these regions with information surrounding them. The fill-in is done in such a way that isophote lines arriving at the regions’ boundaries are completed inside. In contrast with previous approaches, the technique here introduced does not require the user to specify where the novel information comes from. This is automatically done (and in a fast way), thereby allowing to simultaneously fill-in numerous regions containing completely different structures and surrounding backgrounds. In addition, no limitations are imposed on the topology of the region to be inpainted. Applications of this technique include the restoration of old photographs and damaged film; removal of superimposed text like dates, subtitles, or publicity; and the removal of entire objects from the image like microphones or wires in special effects.

Added

2026-09-10

Segment Anything

Segment Anything

Alexander M. Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan‐Yen Lo, Piotr Dollár, Ross Girshick

OrganizationsMeta

Why you should read this

Establishes the first tangible "foundation model" for image segmentation that supports zero-shot prompting via points, boxes, or text.

We introduce the Segment Anything (SA) project: a new task, model, and dataset for image segmentation. Using our efficient model in a data collection loop, we built the largest segmentation dataset to date (by far), with over 1 billion masks on 11M licensed and privacy respecting images. The model is designed and trained to be promptable, so it can transfer zero-shot to new image distributions and tasks. We evaluate its capabilities on numerous tasks and find that its zero-shot performance is impressive – often competitive with or even superior to prior fully supervised results. We are releasing the Segment Anything Model (SAM) and corresponding dataset (SA-1B) of 1B masks and 11M images at segment-anything.com to foster research into foundation models for computer vision. We recommend reading the full paper at: arxiv.org/abs/2304.02643.

Added

2026-01-28