Built independently by an author, for readers. Read the story and support ChapterPal

keyword

spatial pooling

Spatial pooling is an operation in computer vision and artificial neural networks that aggregates feature values across local spatial neighborhoods of an image or feature map into single summary values. Commonly performed using operations such as taking the maximum or average value within a defined region, spatial pooling reduces the spatial resolution and dimensionality of intermediate feature representations. This downsampling decreases computational requirements and memory usage while conferring a degree of invariance to small translations, geometric distortions, and scale changes. By summarizing localized activations across uniform grid cells or segmented regions, spatial pooling helps visual processing models capture multiscale contextual information for tasks such as image recognition, object detection, and scene parsing.

1 item

Learning Hierarchical Features for Scene Labeling

Learning Hierarchical Features for Scene Labeling

C. Farabet, C. Couprie, Laurent Najman, Yann LeCun

OrganizationsNew York UniversityUniversité Paris-Est

Why you should read this

Proposes a multiscale convolutional network that learns hierarchical features directly from raw pixels across image pyramids, enabling near-real-time scene parsing while eliminating the need for hand-engineered features and complex graphical model post-processing.

Scene labeling consists in labeling each pixel in an image with the category of the object it belongs to. We propose a method that uses a multiscale convolutional network trained from raw pixels to extract dense feature vectors that encode regions of multiple sizes centered on each pixel. The method alleviates the need for engineered features, and produces a powerful representation that captures texture, shape and contextual information. We report results using multiple post-processing methods to produce the final labeling. Among those, we propose a technique to automatically retrieve, from a pool of segmentation components, an optimal set of components that best explain the scene; these components are arbitrary, e.g. they can be taken from a segmentation tree, or from any family of over-segmentations. The system yields record accuracies on the Sift Flow Dataset (33 classes) and the Barcelona Dataset (170 classes) and near-record accuracy on Stanford Background Dataset (8 classes), while being an order of magnitude faster than competing approaches, producing a 320 × 240 image labeling in less than a second, including feature extraction.

Added

2026-09-14