Built independently by an author, for readers. Read the story and support ChapterPal

keyword

strided convolution

A strided convolution is a convolutional neural network operation in which the filter moves across the input data with a step size, or stride, greater than one unit at a time. Unlike standard convolutions that slide the filter one pixel at a time, this larger shift skips intermediate positions and produces an output feature map with reduced spatial dimensions. By performing feature extraction and spatial downsampling simultaneously, strided convolutions decrease the computational and memory requirements of subsequent layers while progressively expanding the receptive field of the network, often serving as an alternative to separate pooling operations.

2 items

Understanding The Robustness in Vision Transformers

Understanding The Robustness in Vision Transformers

Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Animashree Anandkumar, Jiashi Feng, José M. Álvarez

OrganizationsArizona State UniversityByteDanceCalifornia Institute of TechnologyNational University of SingaporeNVIDIAUniversity of Hong Kong

Why you should read this

Explains how visual grouping in self-attention reduces corruption sensitivity through an information bottleneck perspective, introducing fully attentional networks that integrate dynamic channel selection to achieve superior corruption error on ImageNet-C.

Recent studies show that Vision Transformers (ViTs) exhibit strong robustness against various corruptions. Although this property is partly attributed to the self-attention mechanism, there is still a lack of systematic understanding. In this paper, we examine the role of self-attention in learning robust representations. Our study is motivated by the intriguing properties of the emerging visual grouping in Vision Transformers, which indicates that self-attention may promote robustness through improved mid-level representations. We further propose a family of fully attentional networks (FANs) that strengthen this capability by incorporating an attentional channel processing design. We validate the design comprehensively on various hierarchical backbones. Our model achieves a state-of-the-art 87.1% accuracy and 35.8% mCE on ImageNet-1k and ImageNet-C with 76.8M parameters. We also demonstrate state-of-the-art accuracy and robustness in two downstream tasks: semantic segmentation and object detection. Code will be available at https://github.com/NVlabs/FAN.

Added

2026-09-26

Understanding Convolution for Semantic Segmentation

Understanding Convolution for Semantic Segmentation

Panqu Wang, Pengfei Chen, Ye Yuan, Ding Liu, Zehua Huang, Xiaodi Hou, Garrison Cottrell

OrganizationsCarnegie Mellon UniversityTuSimpleUniversity of California, San DiegoUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes Dense Upsampling Convolution to capture fine pixel-level details and Hybrid Dilated Convolution to expand receptive fields without gridding artifacts, achieving leading accuracy across major semantic segmentation benchmarks.

Recent advances in deep learning, especially deep convolutional neural networks (CNNs), have led to significant improvement over previous semantic segmentation systems. Here we show how to improve pixel-wise semantic segmentation by manipulating convolution-related operations that are of both theoretical and practical value. First, we design dense upsampling convolution (DUC) to generate pixel-level prediction, which is able to capture and decode more detailed information that is generally missing in bilinear upsampling. Second, we propose a hybrid dilated convolution (HDC) framework in the encoding phase. This framework 1) effectively enlarges the receptive fields (RF) of the network to aggregate global information; 2) alleviates what we call the "gridding issue" caused by the standard dilated convolution operation. We evaluate our approaches thoroughly on the Cityscapes dataset, and achieve a state-of-art result of 80.1% mIOU in the test set at the time of submission. We also have achieved state-of-the-art overall on the KITTI road estimation benchmark and the PASCAL VOC2012 segmentation task. Our source code can be found at this https URL .

Added

2026-09-18