Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multiscale feature learning

Multiscale feature learning is a machine learning approach in which representations of data, such as images or signals, are automatically extracted and integrated across multiple spatial, temporal, or structural scales. By analyzing inputs at both fine and coarse granularities, models can simultaneously capture high-resolution local details like edges and textures alongside broader contextual information such as overall object shapes and scene layouts. This method is commonly implemented using neural network architectures with varying receptive fields, feature pyramids, or multi-resolution inputs, enabling automated recognition systems to achieve high accuracy and robustness against variations in object size, distance, and perspective without relying on manual feature engineering.

1 item

Learning Hierarchical Features for Scene Labeling

Learning Hierarchical Features for Scene Labeling

C. Farabet, C. Couprie, Laurent Najman, Yann LeCun

OrganizationsNew York UniversityUniversité Paris-Est

Why you should read this

Proposes a multiscale convolutional network that learns hierarchical features directly from raw pixels across image pyramids, enabling near-real-time scene parsing while eliminating the need for hand-engineered features and complex graphical model post-processing.

Scene labeling consists in labeling each pixel in an image with the category of the object it belongs to. We propose a method that uses a multiscale convolutional network trained from raw pixels to extract dense feature vectors that encode regions of multiple sizes centered on each pixel. The method alleviates the need for engineered features, and produces a powerful representation that captures texture, shape and contextual information. We report results using multiple post-processing methods to produce the final labeling. Among those, we propose a technique to automatically retrieve, from a pool of segmentation components, an optimal set of components that best explain the scene; these components are arbitrary, e.g. they can be taken from a segmentation tree, or from any family of over-segmentations. The system yields record accuracies on the Sift Flow Dataset (33 classes) and the Barcelona Dataset (170 classes) and near-record accuracy on Stanford Background Dataset (8 classes), while being an order of magnitude faster than competing approaches, producing a 320 × 240 image labeling in less than a second, including feature extraction.

Added

2026-09-14