Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Fully Attentional Networks

Fully Attentional Networks are a class of vision transformer architectures designed to enhance model robustness and visual feature learning by utilizing attention mechanisms across both spatial and channel dimensions. While standard vision transformers typically employ self-attention to capture spatial relationships among image tokens and rely on conventional feed-forward multi-layer perceptrons for channel transformations, fully attentional networks integrate attentional channel processing throughout the backbone. By extending attention-driven computation across all feature axes, these architectures improve visual grouping capabilities and mid-level representations, enabling machine learning models to maintain high accuracy and stability against natural image corruptions, perturbations, and out-of-distribution variations.

1 item

Understanding The Robustness in Vision Transformers

Understanding The Robustness in Vision Transformers

Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Animashree Anandkumar, Jiashi Feng, José M. Álvarez

OrganizationsArizona State UniversityByteDanceCalifornia Institute of TechnologyNational University of SingaporeNVIDIAUniversity of Hong Kong

Why you should read this

Explains how visual grouping in self-attention reduces corruption sensitivity through an information bottleneck perspective, introducing fully attentional networks that integrate dynamic channel selection to achieve superior corruption error on ImageNet-C.

Recent studies show that Vision Transformers (ViTs) exhibit strong robustness against various corruptions. Although this property is partly attributed to the self-attention mechanism, there is still a lack of systematic understanding. In this paper, we examine the role of self-attention in learning robust representations. Our study is motivated by the intriguing properties of the emerging visual grouping in Vision Transformers, which indicates that self-attention may promote robustness through improved mid-level representations. We further propose a family of fully attentional networks (FANs) that strengthen this capability by incorporating an attentional channel processing design. We validate the design comprehensively on various hierarchical backbones. Our model achieves a state-of-the-art 87.1% accuracy and 35.8% mCE on ImageNet-1k and ImageNet-C with 76.8M parameters. We also demonstrate state-of-the-art accuracy and robustness in two downstream tasks: semantic segmentation and object detection. Code will be available at https://github.com/NVlabs/FAN.

Added

2026-09-26