Built independently by an author, for readers. Read the story and support ChapterPal

keyword

pixel-level prediction

Pixel-level prediction is a computer vision process in which a machine learning model generates an individual output value or classification label for every single pixel in an input image. Unlike image-level classification, which assigns a single category to an entire image, or object detection, which localizes targets within coarse bounding boxes, pixel-level prediction produces dense output maps that match the spatial dimensions of the original visual data. This fine-grained approach is fundamental to dense prediction tasks such as semantic segmentation, instance segmentation, and depth estimation. To achieve accurate results, neural network architectures typically combine encoder-decoder structures or dense upsampling operations that capture high-level semantic context while preserving detailed spatial boundaries across the entire scene.

1 item

Understanding Convolution for Semantic Segmentation

Understanding Convolution for Semantic Segmentation

Panqu Wang, Pengfei Chen, Ye Yuan, Ding Liu, Zehua Huang, Xiaodi Hou, Garrison Cottrell

OrganizationsCarnegie Mellon UniversityTuSimpleUniversity of California, San DiegoUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes Dense Upsampling Convolution to capture fine pixel-level details and Hybrid Dilated Convolution to expand receptive fields without gridding artifacts, achieving leading accuracy across major semantic segmentation benchmarks.

Recent advances in deep learning, especially deep convolutional neural networks (CNNs), have led to significant improvement over previous semantic segmentation systems. Here we show how to improve pixel-wise semantic segmentation by manipulating convolution-related operations that are of both theoretical and practical value. First, we design dense upsampling convolution (DUC) to generate pixel-level prediction, which is able to capture and decode more detailed information that is generally missing in bilinear upsampling. Second, we propose a hybrid dilated convolution (HDC) framework in the encoding phase. This framework 1) effectively enlarges the receptive fields (RF) of the network to aggregate global information; 2) alleviates what we call the "gridding issue" caused by the standard dilated convolution operation. We evaluate our approaches thoroughly on the Cityscapes dataset, and achieve a state-of-art result of 80.1% mIOU in the test set at the time of submission. We also have achieved state-of-the-art overall on the KITTI road estimation benchmark and the PASCAL VOC2012 segmentation task. Our source code can be found at this https URL .

Added

2026-09-18