Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multiscale convolutional network

A multiscale convolutional network is a deep learning architecture that processes visual input across multiple spatial resolutions or receptive field sizes to learn rich representations of data. By examining imagery at various scales simultaneously or progressively, such as through multi-resolution image pyramids, parallel convolutional streams, or varying kernel dimensions, the network captures fine-grained local textures and object boundaries alongside broad semantic context. The extracted feature maps from different scales are merged or fused to generate dense, scale-aware representations, allowing the model to effectively perform complex spatial perception tasks such as semantic segmentation, scene labeling, and object detection.

1 item

Learning Hierarchical Features for Scene Labeling

Learning Hierarchical Features for Scene Labeling

C. Farabet, C. Couprie, Laurent Najman, Yann LeCun

OrganizationsNew York UniversityUniversité Paris-Est

Why you should read this

Proposes a multiscale convolutional network that learns hierarchical features directly from raw pixels across image pyramids, enabling near-real-time scene parsing while eliminating the need for hand-engineered features and complex graphical model post-processing.

Scene labeling consists in labeling each pixel in an image with the category of the object it belongs to. We propose a method that uses a multiscale convolutional network trained from raw pixels to extract dense feature vectors that encode regions of multiple sizes centered on each pixel. The method alleviates the need for engineered features, and produces a powerful representation that captures texture, shape and contextual information. We report results using multiple post-processing methods to produce the final labeling. Among those, we propose a technique to automatically retrieve, from a pool of segmentation components, an optimal set of components that best explain the scene; these components are arbitrary, e.g. they can be taken from a segmentation tree, or from any family of over-segmentations. The system yields record accuracies on the Sift Flow Dataset (33 classes) and the Barcelona Dataset (170 classes) and near-record accuracy on Stanford Background Dataset (8 classes), while being an order of magnitude faster than competing approaches, producing a 320 × 240 image labeling in less than a second, including feature extraction.

Added

2026-09-14