Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Felzenszwalb superpixels

Felzenszwalb superpixels are perceptually coherent, contiguous clusters of pixels produced by a graph-based over-segmentation algorithm that partitions an image according to local visual features. Developed by Pedro Felzenszwalb and Daniel Huttenlocher, the technique models an image as an undirected graph where pixels serve as nodes and edge weights measure the dissimilarity between adjacent pixels. The algorithm iteratively merges neighboring pixel components whenever the variation across their shared boundary is smaller than the internal variation within either component, scaled by a threshold parameter. By evaluating boundary differences relative to the internal characteristics of the regions being joined, this method operates in near-linear time and effectively preserves fine details in low-contrast areas while preventing over-segmentation in heavily textured regions.

1 item

Learning Hierarchical Features for Scene Labeling

Learning Hierarchical Features for Scene Labeling

C. Farabet, C. Couprie, Laurent Najman, Yann LeCun

OrganizationsNew York UniversityUniversité Paris-Est

Why you should read this

Proposes a multiscale convolutional network that learns hierarchical features directly from raw pixels across image pyramids, enabling near-real-time scene parsing while eliminating the need for hand-engineered features and complex graphical model post-processing.

Scene labeling consists in labeling each pixel in an image with the category of the object it belongs to. We propose a method that uses a multiscale convolutional network trained from raw pixels to extract dense feature vectors that encode regions of multiple sizes centered on each pixel. The method alleviates the need for engineered features, and produces a powerful representation that captures texture, shape and contextual information. We report results using multiple post-processing methods to produce the final labeling. Among those, we propose a technique to automatically retrieve, from a pool of segmentation components, an optimal set of components that best explain the scene; these components are arbitrary, e.g. they can be taken from a segmentation tree, or from any family of over-segmentations. The system yields record accuracies on the Sift Flow Dataset (33 classes) and the Barcelona Dataset (170 classes) and near-record accuracy on Stanford Background Dataset (8 classes), while being an order of magnitude faster than competing approaches, producing a 320 × 240 image labeling in less than a second, including feature extraction.

Added

2026-09-14