Built independently by an author, for readers. Read the story and support ChapterPal

keyword

scene parsing

Scene parsing is a computer vision task that involves segmenting an image and labeling every pixel with its corresponding semantic category, such as discrete objects, background elements, or object parts. Unlike standard object detection, which locates items within bounding boxes, scene parsing provides a complete, dense spatial breakdown of an entire visual environment by classifying both countable entities like people or vehicles and continuous background regions like sky, ground, and walls. This comprehensive pixel-level understanding enables automated systems to interpret complex environments, recognize individual components, and analyze how these visual elements spatially and contextually interact to form a coherent scene.

4 items

Parsing Natural Scenes and Natural Language with Recursive Neural Networks

Parsing Natural Scenes and Natural Language with Recursive Neural Networks

Richard Socher, Cliff Chiung-Yu Lin, Andrew Y. Ng, Christopher D. Manning

OrganizationsStanford University

Why you should read this

Presents a unified recursive neural network architecture that jointly predicts hierarchical parse trees and semantic representations across both visual scenes and natural language, setting a new benchmark for scene segmentation and classification while maintaining competitive syntactic parsing accuracy.

Recursive structure is commonly found in the inputs of different modalities such as natural scene images or natural language sentences. Discovering this recursive structure helps us to not only identify the units that an image or sentence contains but also how they interact to form a whole. We introduce a max-margin structure prediction architecture based on recursive neural networks that can successfully recover such structure both in complex scene images as well as sentences. The same algorithm can be used both to provide a competitive syntactic parser for natural language sentences from the Penn Treebank and to outperform alternative approaches for semantic scene segmentation, annotation and classification. For segmentation and annotation our algorithm obtains a new level of state-of-the-art performance on the Stanford background dataset (78.1%). The features from the image parse tree outperform Gist descriptors for scene classification by 4%.

Added

2026-09-25

Semantic Understanding of Scenes Through the ADE20K Dataset

Semantic Understanding of Scenes Through the ADE20K Dataset

Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, Antonio Torralba

OrganizationsMassachusetts Institute of TechnologyPeking UniversityThe Chinese University of Hong KongUniversity of Toronto

Why you should read this

Presents ADE20K, an open-vocabulary dataset of 25,000 densely annotated images spanning over 3,000 object and part classes, establishing standardized benchmarks and baselines for fine-grained scene parsing and instance segmentation.

Semantic understanding of visual scenes is one of the holy grails of computer vision. Despite efforts of the community in data collection, there are still few image datasets covering a wide range of scenes and object categories with pixel-wise annotations for scene understanding. In this work, we present a densely annotated dataset ADE20K, which spans diverse annotations of scenes, objects, parts of objects, and in some cases even parts of parts. Totally there are 25k images of the complex everyday scenes containing a variety of objects in their natural spatial context. On average there are 19.5 instances and 10.5 object classes per image. Based on ADE20K, we construct benchmarks for scene parsing and instance segmentation. We provide baseline performances on both of the benchmarks and re-implement the state-of-the-art models for open source. We further evaluate the effect of synchronized batch normalization and find that a reasonably large batch size is crucial for the semantic segmentation performance. We show that the networks trained on ADE20K are able to segment a wide variety of scenes and objects1.

Added

2026-09-14

Learning Hierarchical Features for Scene Labeling

Learning Hierarchical Features for Scene Labeling

C. Farabet, C. Couprie, Laurent Najman, Yann LeCun

OrganizationsNew York UniversityUniversité Paris-Est

Why you should read this

Proposes a multiscale convolutional network that learns hierarchical features directly from raw pixels across image pyramids, enabling near-real-time scene parsing while eliminating the need for hand-engineered features and complex graphical model post-processing.

Scene labeling consists in labeling each pixel in an image with the category of the object it belongs to. We propose a method that uses a multiscale convolutional network trained from raw pixels to extract dense feature vectors that encode regions of multiple sizes centered on each pixel. The method alleviates the need for engineered features, and produces a powerful representation that captures texture, shape and contextual information. We report results using multiple post-processing methods to produce the final labeling. Among those, we propose a technique to automatically retrieve, from a pool of segmentation components, an optimal set of components that best explain the scene; these components are arbitrary, e.g. they can be taken from a segmentation tree, or from any family of over-segmentations. The system yields record accuracies on the Sift Flow Dataset (33 classes) and the Barcelona Dataset (170 classes) and near-record accuracy on Stanford Background Dataset (8 classes), while being an order of magnitude faster than competing approaches, producing a 320 × 240 image labeling in less than a second, including feature extraction.

Added

2026-09-14

Scene Parsing through ADE20K Dataset

Scene Parsing through ADE20K Dataset

Bolei Zhou, Hang Zhao, Xavier Puig, S. Fidler, Adela Barriuso, A. Torralba

OrganizationsMassachusetts Institute of TechnologyUniversity of Toronto

Why you should read this

Introduces the ADE20K benchmark and a cascade segmentation network to hierarchically parse complex scenes into stuff, objects, and object parts while tackling the long-tail distribution of real-world visual categories.

Scene parsing, or recognizing and segmenting objects and stuff in an image, is one of the key problems in computer vision. Despite the community's efforts in data collection, there are still few image datasets covering a wide range of scenes and object categories with dense and detailed annotations for scene parsing. In this paper, we introduce and analyze the ADE20K dataset, spanning diverse annotations of scenes, objects, parts of objects, and in some cases even parts of parts. A scene parsing benchmark is built upon the ADE20K with 150 object and stuff classes included. Several segmentation baseline models are evaluated on the benchmark. A novel network design called Cascade Segmentation Module is proposed to parse a scene into stuff, objects, and object parts in a cascade and improve over the baselines. We further show that the trained scene parsing networks can lead to applications such as image content removal and scene synthesis¹.

Added

2026-09-11