Built independently by an author, for readers. Read the story and support ChapterPal

keyword

segmentation tree

A segmentation tree is a hierarchical data structure used in computer vision and image processing to represent an image as a nested series of regions across multiple scales. In this structure, the leaf nodes typically correspond to fine-grained spatial primitives, such as individual pixels or superpixels generated from an initial over-segmentation, while the root represents the entire image. Intermediate nodes are formed by progressively merging adjacent, visually or semantically similar regions at increasing levels of coarseness. By capturing structural relationships and visual context at multiple granularities simultaneously, a segmentation tree allows automated algorithms to search across scales and retrieve an optimal combination of regions for tasks such as scene parsing, object detection, and semantic labeling.

1 item

Learning Hierarchical Features for Scene Labeling

Learning Hierarchical Features for Scene Labeling

C. Farabet, C. Couprie, Laurent Najman, Yann LeCun

OrganizationsNew York UniversityUniversité Paris-Est

Why you should read this

Proposes a multiscale convolutional network that learns hierarchical features directly from raw pixels across image pyramids, enabling near-real-time scene parsing while eliminating the need for hand-engineered features and complex graphical model post-processing.

Scene labeling consists in labeling each pixel in an image with the category of the object it belongs to. We propose a method that uses a multiscale convolutional network trained from raw pixels to extract dense feature vectors that encode regions of multiple sizes centered on each pixel. The method alleviates the need for engineered features, and produces a powerful representation that captures texture, shape and contextual information. We report results using multiple post-processing methods to produce the final labeling. Among those, we propose a technique to automatically retrieve, from a pool of segmentation components, an optimal set of components that best explain the scene; these components are arbitrary, e.g. they can be taken from a segmentation tree, or from any family of over-segmentations. The system yields record accuracies on the Sift Flow Dataset (33 classes) and the Barcelona Dataset (170 classes) and near-record accuracy on Stanford Background Dataset (8 classes), while being an order of magnitude faster than competing approaches, producing a 320 × 240 image labeling in less than a second, including feature extraction.

Added

2026-09-14