keyword
Stanford Background dataset
The Stanford Background dataset is a benchmark image dataset in computer vision designed for evaluating methods in semantic scene segmentation, scene parsing, and image labeling. Introduced in 2009, the collection contains 715 outdoor scene images compiled from multiple public computer vision databases. The images were selected based on specific criteria, including having an approximate resolution of 320 by 240 pixels, a horizon location within the image frame, and at least one foreground object. Each image is provided with detailed pixel-level annotations categorized into eight semantic classes, comprising sky, tree, road, grass, water, building, mountain, and foreground objects, as well as geometric surface orientations, making it a standard resource for training and testing scene understanding algorithms.
2 items

Parsing Natural Scenes and Natural Language with Recursive Neural Networks
Richard Socher, Cliff Chiung-Yu Lin, Andrew Y. Ng, Christopher D. Manning
Why you should read this
Presents a unified recursive neural network architecture that jointly predicts hierarchical parse trees and semantic representations across both visual scenes and natural language, setting a new benchmark for scene segmentation and classification while maintaining competitive syntactic parsing accuracy.
Recursive structure is commonly found in the inputs of different modalities such as natural scene images or natural language sentences. Discovering this recursive structure helps us to not only identify the units that an image or sentence contains but also how they interact to form a whole. We introduce a max-margin structure prediction architecture based on recursive neural networks that can successfully recover such structure both in complex scene images as well as sentences. The same algorithm can be used both to provide a competitive syntactic parser for natural language sentences from the Penn Treebank and to outperform alternative approaches for semantic scene segmentation, annotation and classification. For segmentation and annotation our algorithm obtains a new level of state-of-the-art performance on the Stanford background dataset (78.1%). The features from the image parse tree outperform Gist descriptors for scene classification by 4%.
Added
2026-09-25

Learning Hierarchical Features for Scene Labeling
C. Farabet, C. Couprie, Laurent Najman, Yann LeCun
Why you should read this
Proposes a multiscale convolutional network that learns hierarchical features directly from raw pixels across image pyramids, enabling near-real-time scene parsing while eliminating the need for hand-engineered features and complex graphical model post-processing.
Scene labeling consists in labeling each pixel in an image with the category of the object it belongs to. We propose a method that uses a multiscale convolutional network trained from raw pixels to extract dense feature vectors that encode regions of multiple sizes centered on each pixel. The method alleviates the need for engineered features, and produces a powerful representation that captures texture, shape and contextual information. We report results using multiple post-processing methods to produce the final labeling. Among those, we propose a technique to automatically retrieve, from a pool of segmentation components, an optimal set of components that best explain the scene; these components are arbitrary, e.g. they can be taken from a segmentation tree, or from any family of over-segmentations. The system yields record accuracies on the Sift Flow Dataset (33 classes) and the Barcelona Dataset (170 classes) and near-record accuracy on Stanford Background Dataset (8 classes), while being an order of magnitude faster than competing approaches, producing a 320 × 240 image labeling in less than a second, including feature extraction.
Added
2026-09-14
