Understanding The Robustness in Vision Transformers

Daquan ZhouZhiding YuEnze XieChaowei XiaoAnimashree AnandkumarJiashi FengJosé M. Álvarez

article2022ICML264 citations

Explains how visual grouping in self-attention reduces corruption sensitivity through an information bottleneck perspective, introducing fully attentional networks that integrate dynamic channel selection to achieve superior corruption error on ImageNet-C.

Listen

Modern computer vision models increasingly power mission-critical and safety-sensitive applications, such as autonomous driving and mobile visual recognition. While vision transformers frequently handle visual corruptions—such as bad weather, blur, and digital noise—better than standard convolutional neural networks, the structural reasons for this resilience have remained poorly understood. Recent developments, including improved convolutional architectures that match transformer accuracy, have created uncertainty regarding whether attention mechanisms provide fundamental robustness advantages in real-world operating environments.

The article aims to explain why self-attention improves visual robustness against corruptions and to introduce a novel architectural framework that enhances this capability across standard and high-demand vision tasks.

To evaluate this, the authors analyzed how internal representations evolve across network layers using spectral clustering, measuring how models group visual patterns and suppress noise. They established a theoretical link connecting self-attention to the information bottleneck principle, which removes irrelevant image noise while preserving critical target features. Building on these theoretical insights, the authors designed Fully Attentional Networks (FANs), a family of vision architectures that incorporate an efficient channel attention mechanism to perform dynamic, content-dependent feature selection. They evaluated these architectures against standard convolutional and transformer models across multiple corrupted and out-of-distribution benchmarks, including ImageNet-C, Cityscapes-C, and COCO-C, across multiple model scales.

The analysis yielded several key findings regarding model robustness. First, spectral analysis demonstrated that self-attention inherently groups visual tokens into clean object clusters in intermediate layers, effectively squeezing out noise perturbations in a manner that standard convolutional networks fail to replicate. Second, the proposed FAN architecture significantly outperformed existing models in corruption robustness without compromising clean image accuracy; for example, the small FAN variant achieved a 47.7% mean corruption error on ImageNet-C, outperforming comparable convolutional baselines by 29.0% and recent competitive designs by 5.5%. Third, scaling the FAN framework established new state-of-the-art supervised performance, reaching 87.1% clean accuracy on ImageNet-1k alongside an industry-leading 35.8% mean corruption error. Finally, these robustness gains directly transferred to dense downstream applications, improving semantic segmentation on Cityscapes-C by 6.8% mean intersection-over-union over previous leading transformers and improving object detection accuracy on COCO-C by 6.2% mean average precision.

These findings indicate that attentional grouping provides architectural advantages that advanced training techniques and standard convolutional layers cannot fully achieve alone. In deployment contexts, adopting fully attentional backbones can substantially reduce operational risks, perception failures, and safety hazards in variable outdoor conditions. Furthermore, hybrid designs that pair initial convolutional layers with higher-level attentional channel processing offer a practical path for high-resolution vision systems, balancing computational efficiency with superior resilience.

Organizations developing safety-critical perception systems should consider incorporating fully attentional or hybrid transformer backbones rather than relying solely on pure convolutional networks or standard transformer designs. Teams should benchmark existing perception pipelines against standard corrupted datasets to quantify vulnerability to environmental perturbations. When computational overhead is a constraint in dense prediction settings, technical leaders should prioritize hybrid models, which effectively balance memory footprint with robust performance.

The article focuses primarily on natural corruptions, out-of-distribution shifts, and standard architectural benchmarks, rather than deliberate adversarial attacks or extreme hardware-constrained edge environments. Given the consistent experimental results across multiple standard benchmarks, model sizes, and downstream tasks, decision-makers can have strong confidence in fully attentional architectures for general robust vision applications.

Cover for Understanding The Robustness in Vision Transformers

Abstract

Recent studies show that Vision Transformers (ViTs) exhibit strong robustness against various corruptions. Although this property is partly attributed to the self-attention mechanism, there is still a lack of systematic understanding. In this paper, we examine the role of self-attention in learning robust representations. Our study is motivated by the intriguing properties of the emerging visual grouping in Vision Transformers, which indicates that self-attention may promote robustness through improved mid-level representations. We further propose a family of fully attentional networks (FANs) that strengthen this capability by incorporating an attentional channel processing design. We validate the design comprehensively on various hierarchical backbones. Our model achieves a state-of-the-art 87.1% accuracy and 35.8% mCE on ImageNet-1k and ImageNet-C with 76.8M parameters. We also demonstrate state-of-the-art accuracy and robustness in two downstream tasks: semantic segmentation and object detection. Code will be available at https://github.com/NVlabs/FAN.

Table of Contents

  • 1. Introduction
  • 2. Fully Attentional Networks
  • 2.1. Preliminaries on Vision Transformers
  • 2.2. Intriguing Properties of Self-Attention
  • 2.3. An Information Bottleneck Perspective
  • 2.4. Fully Attentional Networks
  • 2.5. Efficient Channel Self-attention
  • 3. Experiment Results & Analysis
  • 3.1. Experiment details
  • 3.2. Analysis
  • 3.3. Fully Attentional Networks
  • 3.4. Comparison to SOTAs on various tasks
  • 4. Related Works
  • 5. Conclusion
  • References
  • A. Supplementary Details
  • A.1. Proof on the relationship between the Information Bottleneck and Self-Attention
  • A.2. Implementation details
  • A.3. Impact of head numbers
  • A.4. Detailed benchmark results on corrupted images on classification, segmentation and detection
  • A.5. Architecture details of FAN-Swin and FAN-Hybrid
  • A.6. Feature clustering and visualizations

Knowls

  1. Knowl 1 — Self-Attention as an Iterative Information Bottleneck Optimization Step

    theoretical result

    The Information Bottleneck (IB) principle seeks a compact mapping f(Z∣X)f(Z|X) from an observed noisy input X∼N(X′,ϵ)X \sim \mathcal{N}(X', \epsilon) that preserves information relevant to predicting a target clean code X′X':

    fIB∗(Z∣X)=arg⁡min⁡f(Z∣X)I(X,Z)−βI(Z,X′)f^*_{\text{IB}}(Z|X) = \arg\min_{f(Z|X)} I(X, Z) - \beta I(Z, X')

    subject to the Markov constraint Z↔X↔X′Z \leftrightarrow X \leftrightarrow X', where I(⋅,⋅)I(\cdot, \cdot) denotes mutual information and β\beta is a compression-prediction trade-off parameter.

    For a self-attention module with input token matrix X=[x1,…,xn]∈Rd×nX = [x_1, \dots, x_n] \in \mathbb{R}^{d \times n} and output token representations Z=[z1,…,zn]∈Rd×nZ = [z_1, \dots, z_n] \in \mathbb{R}^{d \times n}, assuming data conditional distribution p(x∣i)∝exp⁡(−12ϵ2∥x−xi∥2)p(x|i) \propto \exp\left(-\frac{1}{2\epsilon^2} \|x - x_i\|^2\right) for small ϵ\epsilon and modeling cluster distributions as Gaussians N(x∣μc,Σ)\mathcal{N}(x | \mu_c, \Sigma), the iterative IB clustering update for cluster center representation zcz_c is:

    zc=∑i=1nlog⁡(nc/n)ndet⁡Σexp⁡(μc⊤Σ−1xi1/2)∑c′=1nexp⁡(μc′⊤Σ−1xi1/2)xiz_c = \sum_{i=1}^n \frac{\log(n_c/n)}{n \det \Sigma} \frac{\exp\left(\frac{\mu_c^\top \Sigma^{-1} x_i}{1/2}\right)}{\sum_{c'=1}^n \exp\left(\frac{\mu_{c'}^\top \Sigma^{-1} x_i}{1/2}\right)} x_i

    In matrix notation, this optimization step matches standard self-attention:

    Z=Softmax(Q⊤Kd)V⊤Z = \text{Softmax}\left(\frac{Q^\top K}{d}\right) V^\top

    where V=[x1,…,xn]log⁡(nc/n)ndet⁡ΣV = [x_1, \dots, x_n] \frac{\log(n_c/n)}{n \det \Sigma}, K=[μ1,…,μn]=WKXK = [\mu_1, \dots, \mu_n] = W_K X, Q=Σ−1[x1,…,xn]Q = \Sigma^{-1}[x_1, \dots, x_n], d=1/2d = 1/2, and ncn_c, Σ\Sigma, and WKW_K are learnable variables. This shows that self-attention acts as an iterative soft clustering operator that maps noisy inputs xix_i to cluster centroids μc\mu_c stored in the key matrix KK, explaining the emergence of visual grouping and noise suppression in Vision Transformers.

  2. Knowl 2 — Efficient Channel Attention Mechanism for Fully Attentional Networks

    model/method

    Conventional channel self-attention evaluates a full d×dd \times d inter-channel covariance matrix with a Softmax function, incurring quadratic computational complexity O(d2)\mathcal{O}(d^2) with respect to channel dimension dd and diminishing non-dominant channel activations. Efficient Channel Attention (ECA) overcomes these issues by computing attention relative to a 1D spatial token prototype and normalizing with a Sigmoid function.

    Given token features Z∈Rd×nZ \in \mathbb{R}^{d \times n} across dd channels and nn spatial tokens, ECA first computes the spatial prototype Zˉ∈R1×n\bar{Z} \in \mathbb{R}^{1 \times n} by averaging across all channel dimensions:

    Zˉ=1d∑c=1dZc,:\bar{Z} = \frac{1}{d} \sum_{c=1}^d Z_{c, :}

    The channel transformation is then defined as:

    ECA(Z)=Sigmoid((WQ′σ(Z))σ(Zˉ)⊤n)⊙MLP(Z)\text{ECA}(Z) = \text{Sigmoid}\left(\frac{(W'_Q \sigma(Z)) \sigma(\bar{Z})^\top}{\sqrt{n}}\right) \odot \text{MLP}(Z)

    where σ(⋅)\sigma(\cdot) denotes the Softmax operation along the token dimension, WQ′∈Rd×dW'_Q \in \mathbb{R}^{d \times d} is a learnable projection matrix, ⊙\odot denotes elementwise multiplication, and MLP(⋅)\text{MLP}(\cdot) is a two-layer feed-forward network with an intermediate GELU activation.

    ECA evaluates channel correlations with the global spatial prototype with linear complexity O(d)\mathcal{O}(d), while the Sigmoid gating allows each channel to be dynamically reweighted without suppressing independent, informative channels.

  3. Knowl 3 — Fully Attentional Network Family and Macro-Architectures

    model/method

    The Fully Attentional Network (FAN) architecture replaces static multi-layer perceptron (MLP) channel transformations in Vision Transformers with attentional channel processing (Efficient Channel Attention, or ECA), resulting in full attentional processing along both the token and channel dimensions.

    FAN is instantiated in three macro-architectural designs:

    1. FAN-ViT: An isotropic vision transformer stacking standard multi-head self-attention (MHSA) for token mixing followed directly by ECA for dynamic channel selection.
    2. FAN-Swin: A hierarchical vision transformer combining shifted-window self-attention for local spatial aggregation with ECA for dynamic channel feature reweighting.
    3. FAN-Hybrid: A hybrid backbone that uses three ConvNeXt convolutional blocks per stage in the first two stages to extract low-level spatial patterns, followed by FAN blocks in the final two stages where mid-level visual grouping and semantic clustering emerge.

    The configurations across model scales are defined as follows:

    Model Variant Blocks Channel Dim Heads Parameters FLOPs
    FAN-T 12 192 4 7.3M 1.4G
    FAN-S 12 384 8 28.3M 5.3G
    FAN-B 18 448 8 54.0M 10.4G
    FAN-L 24 480 10 80.5M 15.8G

    For classification, two class attention blocks are appended at the top layers to aggregate the global class representation.

  4. Knowl 4 — Token Visual Grouping and Noise Suppression Across Transformer Layers

    empirical result

    Spectral clustering analysis on token affinity matrices Sij=zi⊤zjS_{ij} = z_i^\top z_j across the layers of Vision Transformers (ViT-S and FAN-S) shows that the number of near-zero eigenvalues (eigenvalues below predefined thresholds ϵ∈{10−2,10−3,10−4}\epsilon \in \{10^{-2}, 10^{-3}, 10^{-4}\}) increases progressively with network depth. This indicates that token representations collapse into a small number of distinct, coherent visual object clusters in middle and deep stages.

    When standard Gaussian perturbation noise x∼N(0,1)x \sim \mathcal{N}(0, 1) is injected into the input, the normalized feature perturbation norm decays monotonically across successive self-attention blocks, mirroring the increase in zero eigenvalues.

    In contrast, standard convolutional networks such as ResNet-50 do not exhibit continuous visual grouping: feature noise norms only decrease at explicit pooling/downsampling stages and plateau across standard convolution blocks, leading to significantly lower overall noise attenuation compared to self-attention architectures.

  5. Knowl 5 — Impact of Attention Head Count and Channel Dimension on Robustness

    empirical result

    Interpreting Multi-Head Self-Attention (MHSA) as a mixture of parallel information bottlenecks, varying the number of attention heads under a fixed total channel budget of 384 on DeiT-S produces the following clean and corruption robustness trade-offs on ImageNet-1K and ImageNet-C:

    Attention Heads Channels / Head Clean ImageNet-1K (%) ImageNet-C Retention (%)
    2 192 78.3 68.0
    3 128 79.3 70.7
    8 48 79.9 72.7
    12 32 80.1 73.3
    16 24 79.8 73.4

    Increasing the number of attention heads improves representation diversity and corruption retention rate (rising from 68.0% with 2 heads to 73.4% with 16 heads). However, reducing the channel dimension per head below 32 channels leads to degraded clean accuracy (dropping from 80.1% at 12 heads to 79.8% at 16 heads). The optimal balance between clean generalization and corruption robustness is achieved at 32 channels per head.

  6. Knowl 6 — Inherent Robustness Advantage of Self-Attention over Convolutions under Controlled Training

    empirical result

    To isolate the architectural contribution of self-attention from modern training tricks, ResNet-50 and ViT-S were evaluated under identical training conditions using the DeiT training recipe (including CutMix, Mixup, RandAugment, and Random Erasing):

    Model Configuration Parameters IN-1K Clean Acc (%) IN-C Robust Acc (%) mCE (↓\downarrow)
    ResNet-50 (Baseline) 25M 76.0 38.8 76.7
    + DeiT Recipe 25M 79.0 43.9 69.7
    + Squeeze-and-Excitation 25M 79.8 50.1 63.1
    + Strided Convolutions (ResNet-50∗^*) 25M 80.2 52.1 61.6
    ViT-S (Baseline) 22M 77.9 54.2 63.5
    ViT-S∗^* (12 blocks, DeiT Recipe) 22M 79.9 58.0 56.2

    Although data augmentation, Squeeze-and-Excitation channel attention, and strided downsampling improve ResNet-50 robust accuracy from 38.8% to 52.1%, the parameter-matched ViT-S∗^* trained with the identical recipe achieves 58.0% robust accuracy and 56.2% mCE (compared to 61.6% mCE for ResNet-50∗^*). This confirms that self-attention provides an intrinsic architectural robustness advantage over convolutional operations.

  7. Knowl 7 — Image Classification and Corruption Benchmark Performance of FAN Models

    data/table

    Fully Attentional Networks (FAN) were evaluated on ImageNet-1K and ImageNet-C (15 corruption types across blur, noise, digital, and weather at 5 severity levels) without corruption fine-tuning. The retention rate is defined as Robust Acc./Clean Acc.\text{Robust Acc.} / \text{Clean Acc.}, and mean corruption error (mCE) follows the standard benchmark definition.

    Model Parameters FLOPs IN-1K Clean (%) IN-C Robust (%) Retention (%)
    ResNet-18 11M 1.8G 69.9 32.7 46.8
    MobileNetV2 4M 0.4G 73.0 35.0 47.9
    PVT-V2-B1 13M 2.1G 78.7 51.7 65.7
    FAN-T-ViT 7M 1.3G 79.2 57.5 72.6
    FAN-T-Hybrid 7M 3.5G 80.1 57.4 71.7
    ResNet-50 25M 4.1G 79.0 50.6 64.1
    DeiT-S 22M 4.6G 79.9 58.1 72.7
    Swin-T 28M 4.5G 81.3 55.4 68.1
    ConvNeXt-T 29M 4.5G 82.1 59.1 71.9
    FAN-S-ViT 28M 5.3G 82.9 64.5 77.8
    FAN-S-Hybrid 26M 6.7G 83.5 64.7 77.5
    Swin-S 50M 8.7G 83.0 60.4 72.8
    ConvNeXt-S 50M 8.7G 83.1 61.7 74.2
    FAN-B-ViT 54M 10.4G 83.6 67.0 80.1
    FAN-B-Hybrid 50M 11.3G 83.9 66.4 79.1
    FAN-B-Hybrid†^\dagger 50M 11.3G 85.6 70.5 82.4
    DeiT-B 89M 17.6G 81.8 62.7 76.7
    Swin-B 88M 15.4G 83.5 60.4 72.3
    ConvNeXt-B 89M 15.4G 83.8 61.7 73.6
    FAN-L-ViT 81M 15.8G 83.9 67.7 80.7
    FAN-L-Hybrid 77M 16.9G 84.3 68.3 81.0
    FAN-L-Hybrid†^\dagger 77M 16.9G 86.5 73.6 85.1

    (†^\dagger denotes ImageNet-22K pretraining).

    Across all model sizes, FAN architectures outperform both standard CNNs (ResNet, ConvNeXt) and Vision Transformers (DeiT, Swin) in clean accuracy, robust accuracy, and retention rate, reaching up to 85.1% retention under ImageNet-22K pretraining.

  8. Knowl 8 — Out-of-Distribution Generalization on ImageNet-A, ImageNet-R, and ImageNet-C

    data/table

    Zero-shot out-of-distribution (OOD) performance was evaluated on ImageNet-A (natural adversarial examples), ImageNet-R (artistic and abstract renditions), and ImageNet-C (common corruptions, measured via mean corruption error mCE where lower is better):

    Model Parameters Clean IN-1K (%) IN-A Acc (%) IN-R Acc (%) IN-C mCE (↓\downarrow)
    ImageNet-1K Pretrained
    XCiT-S12 26.3M 81.9 25.0 45.5 51.5
    RVT-S∗^* 23.3M 81.9 25.7 47.7 51.4
    Swin-T 28.3M 81.2 21.6 41.3 59.6
    ConvNeXt-T 28.6M 82.1 24.2 47.2 53.2
    FAN-S-ViT 28.0M 82.5 29.1 50.4 47.7
    FAN-S-Hybrid 26.0M 83.6 33.9 50.7 47.8
    Swin-S 50.0M 83.4 35.8 46.6 52.7
    ConvNeXt-S 50.2M 82.1 31.2 49.5 51.2
    FAN-B-ViT 54.0M 83.6 35.4 51.8 44.4
    FAN-B-Hybrid 50.0M 83.9 39.6 52.9 45.2
    Swin-B 87.8M 83.4 35.8 64.2 54.4
    ConvNeXt-B 88.6M 83.8 36.7 51.3 46.8
    FAN-L-ViT 80.5M 83.9 37.2 53.1 43.3
    FAN-L-Hybrid 76.8M 84.3 41.8 53.2 43.0
    ImageNet-22K Pretrained
    ConvNeXt-B‡^\ddagger 88.6M 86.8 62.3 64.9 43.1
    FAN-L-Hybrid 76.8M 86.5 60.7 64.3 35.8
    FAN-L-Hybrid‡^\ddagger 76.8M 87.1 74.5 71.1 36.0

    (‡^\ddagger indicates fine-tuning at 384×384384 \times 384 resolution).

    FAN models consistently exceed competing CNN and Transformer architectures across natural adversarial examples and domain shifts, achieving 35.8% mCE on ImageNet-C and 74.5% accuracy on ImageNet-A when pretrained on ImageNet-22K.

  9. Knowl 9 — Semantic Segmentation Robustness on Cityscapes-C

    data/table

    Semantic segmentation robustness was evaluated by fine-tuning models on clean Cityscapes and testing on Cityscapes-C across 16 algorithmic corruptions covering noise, blur, digital, and weather categories. Performance is measured using mean Intersection over Union (mIoU) and retention rate (City-C mIoU/City Clean mIoU\text{City-C mIoU} / \text{City Clean mIoU}):

    Model Encoder Params City Clean mIoU (%) City-C mIoU (%) Retention (%)
    DeepLabv3+ (ResNet-50) 25.4M 76.6 36.8 48.0
    DeepLabv3+ (ResNet-101) 47.9M 77.1 39.4 51.1
    DeepLabv3+ (Xception-65) 22.8M 78.4 42.7 54.5
    PSPNet 13.7M 78.8 34.5 43.8
    ConvNeXt-T 29.0M 79.0 54.4 68.9
    SETR (DeiT-S) 22.1M 76.0 55.3 72.8
    SWIN-T 28.4M 78.1 47.3 60.6
    SegFormer-B0 3.4M 76.2 48.8 64.0
    SegFormer-B1 13.1M 78.4 52.7 67.2
    SegFormer-B2 24.2M 81.0 59.6 73.6
    SegFormer-B5 81.4M 82.4 65.8 79.9
    FAN-T-Hybrid 7.4M 81.2 57.1 70.3
    FAN-S-Hybrid 26.3M 81.5 66.4 81.5
    FAN-B-Hybrid 50.4M 82.2 66.9 81.5
    FAN-L-Hybrid 76.8M 82.3 68.7 83.5

    FAN-S-Hybrid achieves 66.4% corrupted mIoU (81.5% retention), surpassing the comparable SegFormer-B2 by 6.8% mIoU and ConvNeXt-T by 12.0% mIoU. FAN-L-Hybrid attains 68.7% mIoU under corruptions with an 83.5% retention rate.

  10. Knowl 10 — Object Detection Robustness on COCO-C

    data/table

    Object detection robustness was evaluated by training detection frameworks on clean COCO and benchmarking zero-shot performance on COCO-C across 16 corruption types. Performance is reported in box mean Average Precision (mAP) on clean COCO and corrupted COCO-C, along with the retention rate (COCO-C mAP/COCO Clean mAP\text{COCO-C mAP} / \text{COCO Clean mAP}):

    Detector Backbone Encoder Params COCO Clean mAP COCO-C mAP Retention (%)
    Mask R-CNN
    ResNet-50 25.4M 39.9 21.3 53.3
    ResNet-101 44.1M 41.8 23.3 55.7
    DeiT-S 22.1M 40.0 26.9 67.3
    Swin-T 28.0M 46.0 29.3 63.7
    FAN-T-Hybrid 7.4M 45.8 29.7 64.8
    FAN-S-Hybrid 26.3M 49.1 35.5 72.3
    Cascade R-CNN
    FAN-T-Hybrid 7.4M 50.2 33.1 65.9
    FAN-S-Hybrid 26.3M 53.3 38.7 72.6
    FAN-L-Hybrid 76.8M 54.1 40.6 75.0
    FAN-L-Hybrid†^\dagger 76.8M 55.1 42.0 76.2

    (†^\dagger denotes ImageNet-22K pretraining).

    Under the Mask R-CNN framework, FAN-S-Hybrid outperforms Swin-T by 6.2% corrupted mAP (35.5% vs 29.3%) and ResNet-50 by 14.2% corrupted mAP. Under Cascade R-CNN, FAN-L-Hybrid reaches 42.0% mAP under corruption with 76.2% retention, confirming the transferability of attentional robustness to object detection.

Coverage note — Detailed per-corruption breakdown sub-tables (Tables 12, 13, and 14 in the appendix) were aggregated into the main benchmark tables (Knowls 7, 8, 9, and 10) to preserve non-redundant and essential quantitative comparisons.

References

  1. 1.Bai, Y., Mei, J., Yuille, A. L., and Xie, C. Are transformers more robust than cnns? In NeurIPS, 2021.
  2. 2.Buhmann, J. M., Malik, J., and Perona, P. Image recognition: Visual grouping, recognition, and learning. Proceedings of the National Academy of Sciences, 96(25): 14203–14204, 1999.
  3. 3.Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. End-to-end object detection with transformers. In ECCV, pp. 213–229. Springer, 2020.
  4. 4.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., ´ Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9650–9660, 2021.
  5. 5.Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Xu, J., et al. Mmdetection: Open mmlab detection toolbox and benchmark. arXiv:1906.07155, 2019.
  6. 6.Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, pp. 801–818, 2018.
  7. 7.Contributors, M. Mmsegmentation: Openmmlab semantic segmentation toolbox and benchmark, 2020.
  8. 8.Dai, Z., Cai, B., Lin, Y., and Chen, J. Up-detr: Unsupervised pre-training for object detection with transformers. In CVPR, pp. 1601–1610, 2021.
  9. 9.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2020.
  10. 10.El-Nouby, A., Touvron, H., Caron, M., Bojanowski, P., Douze, M., Joulin, A., Laptev, I., Neverova, N., Synnaeve, G., Verbeek, J., et al. Xcit: Cross-covariance image transformers. In NeurIPS, 2021.
  11. 11.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, pp. 770–778, 2016.
  12. 12.He, K., Gkioxari, G., Dollar, P., and Girshick, R. Mask ´ r-cnn. In ICCV, pp. 2961–2969, 2017.
  13. 13.He, T., Zhang, Z., Zhang, H., Zhang, Z., Xie, J., and Li, M. Bag of tricks for image classification with convolutional neural networks. In CVPR, pp. 558–567, 2019.
  14. 14.Hendrycks, D. and Dietterich, T. Benchmarking neural network robustness to common corruptions and perturbations. ICLR, 2019.
  15. 15.Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., and Song, D. Natural adversarial examples. In CVPR, pp. 15262–15271, 2021.
  16. 16.Heo, B., Yun, S., Han, D., Chun, S., Choe, J., and Oh, S. J. Rethinking spatial dimensions of vision transformers. In ICCV, pp. 11936–11945, 2021.
  17. 17.Hu, J., Shen, L., and Sun, G. Squeeze-and-excitation networks. In CVPR, pp. 7132–7141, 2018.
  18. 18.Kamann, C. and Rother, C. Benchmarking the robustness of semantic segmentation models. In CVPR, pp. 8828–8838, 2020.
  19. 19.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. In NeurIPS, pp. 1097–1105, 2012.
  20. 20.Kurakin, A., Goodfellow, I., and Bengio, S. Adversarial machine learning at scale. In ICLR, 2016.
  21. 21.LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D. Backpropagation applied to handwritten zip code recognition. Neural computation, 1(4):541–551, 1989.
  22. 22.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, pp. 10012–10022, 2021.
  23. 23.Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. A convnet for the 2020s. CVPR, 2022.
  24. 24.Long, J., Shelhamer, E., and Darrell, T. Fully convolutional networks for semantic segmentation. In CVPR, pp. 3431–3440, 2015.
  25. 25.Mao, X., Qi, G., Chen, Y., Li, X., Duan, R., Ye, S., He, Y., and Xue, H. Towards robust vision transformer. In CVPR, 2021.
  26. 26.Naseer, M., Ranasinghe, K., Khan, S., Hayat, M., Khan, F. S., and Yang, M.-H. Intriguing properties of vision transformers. In NeurIPS, 2021.
  27. 27.Ng, A. Y., Jordan, M. I., and Weiss, Y. On spectral clustering: Analysis and an algorithm. In NIPS, 2002.
  28. 28.Paul, S. and Chen, P.-Y. Vision transformers are robust learners. In AAAI, 2022.
  29. 29.Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., and Dosovitskiy, A. Do vision transformers see like convolutional neural networks? NeurIPS, 34, 2021.
  30. 30.Ren, S., He, K., Girshick, R., and Sun, J. Faster r-cnn: Towards real-time object detection with region proposal networks. NeurIPS, 28:91–99, 2015.
  31. 31.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, pp. 4510–4520, 2018.
  32. 32.Shao, R., Shi, Z., Yi, J., Chen, P.-Y., and Hsieh, C.-J. On the adversarial robustness of visual transformers. arXiv:2103.15670, 2021.
  33. 33.Tan, M. and Le, Q. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, pp. 6105–6114. PMLR, 2019.
  34. 34.Tang, C., Zhao, Y., Wang, G., Luo, C., Xie, W., and Zeng, W. Sparse mlp for image recognition: Is self-attention really necessary? arXiv:2109.05422, 2021.
  35. 35.Tishby, N. and Zaslavsky, N. Deep learning and the information bottleneck principle. In 2015 IEEE Information Theory Workshop (ITW), pp. 1–5. IEEE, 2015.
  36. 36.Tishby, N., Pereira, F. C., and Bialek, W. The information bottleneck method. physics/0004057, 2000.
  37. 37.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jegou, H. Training data-efficient image trans- ´ formers & distillation through attention. In ICML, pp. 10347–10357. PMLR, 2021a.
  38. 38.Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., and Jegou, H. Going deeper with image transformers. In ´ ICCV, pp. 32–42, 2021b.
  39. 39.U.C. Berkeley. Reorganization: Grouping, contour detection, segmentation, ecological statistics. https://www2.eecs.berkeley.edu/Research/Projects/CS/vision/grouping/.
  40. 40.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. NeurIPS, 30, 2017.
  41. 41.Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In CVPR, pp. 568–578, 2021.
  42. 42.Wang, Y., Xu, Z., Wang, X., Shen, C., Cheng, B., Shen, H., and Xia, H. End-to-end video instance segmentation with transformers. CVPR, 2020.
  43. 43.Wightman, R. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019.
  44. 44.Wu, Z., Shen, C., and Van Den Hengel, A. Wider or deeper: Revisiting the resnet model for visual recognition. Pattern Recognition, 90:119–133, 2019.
  45. 45.Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J. M., and Luo, P. Segformer: Simple and efficient design for semantic segmentation with transformers. In NeurIPS, 2021.
  46. 46.Yang, A. Y., Wright, J., Ma, Y., and Sastry, S. S. Unsupervised segmentation of natural images via lossy data compression. Computer Vision and Image Understanding, 110(2):212–225, 2008.
  47. 47.Yu, F. and Koltun, V. Multi-scale context aggregation by dilated convolutions. ICLR, 2016.
  48. 48.Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Tay, F. E., Feng, J., and Yan, S. Tokens-to-token vit: Training vision transformers from scratch on imagenet. ICCV, 2021.
  49. 49.Zelnik-Manor, L. and Perona, P. Self-tuning spectral clustering. In NIPS, 2004.
  50. 50.Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. Pyramid scene parsing network. In CVPR, pp. 2881–2890, 2017.
  51. 51.Zhao, H., Qi, X., Shen, X., Shi, J., and Jia, J. Icnet for real-time semantic segmentation on high-resolution images. In ECCV, pp. 405–420, 2018.
  52. 52.Zheng, M., Gao, P., Wang, X., Li, H., and Dong, H. End-to-end object detection with adaptive clustering transformer. BMVC, 2020.
  53. 53.Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., and Torr, P. H. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In CVPR, pp. 6881–6890, 2021.
  54. 54.Zhou, D., Kang, B., Jin, X., Yang, L., Lian, X., Hou, Q., and Feng, J. Deepvit: Towards deeper vision transformer. arXiv:2103.11886, 2021a.
  55. 55.Zhou, D., Shi, Y., Kang, B., Yu, W., Jiang, Z., Li, Y., Jin, X., Hou, Q., and Feng, J. Refiner: Refining self-attention for vision transformers. arXiv preprint arXiv:2106.03714, 2021b.
  56. 56.Zhu, C., Ping, W., Xiao, C., Shoeybi, M., Goldstein, T., Anandkumar, A., and Catanzaro, B. Long-short transformer: Efficient transformers for language and vision. In NeurIPS, 2021.
  57. 57.Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. Deformable detr: Deformable transformers for end-to-end object detection. ICLR, 2020.

Citation

MLA
Zhou, D., et al. “Understanding The Robustness in Vision Transformers”. International Conference on Machine Learning, vol. 162, 2022, pp. 27378–94, https://proceedings.mlr.press/v162/zhou22m.html.
APA
Zhou, D., Yu, Z., Xie, E., Xiao, C., Anandkumar, A., Feng, J., & Alvarez, J. M. (2022). Understanding The Robustness in Vision Transformers. International Conference on Machine Learning, 162, 27378–27394. https://proceedings.mlr.press/v162/zhou22m.html
Chicago
Zhou, D., Z. Yu, E. Xie, et al. 2022. “Understanding The Robustness in Vision Transformers”. International Conference on Machine Learning 162: 27378–94. https://proceedings.mlr.press/v162/zhou22m.html.
Harvard
Zhou, D. et al. (2022) “Understanding The Robustness in Vision Transformers”, International Conference on Machine Learning. PMLR, pp. 27378–27394. Available at: https://proceedings.mlr.press/v162/zhou22m.html.
Vancouver
1. Zhou D, Yu Z, Xie E, Xiao C, Anandkumar A, Feng J, Alvarez JM (2022) Understanding The Robustness in Vision Transformers. In: International Conference on Machine Learning. PMLR, pp 27378–27394

BibTeX

@InProceedings{pmlr-v162-zhou22m,
  title = 	 {Understanding The Robustness in Vision Transformers},
  author =       {Zhou, Daquan and Yu, Zhiding and Xie, Enze and Xiao, Chaowei and Anandkumar, Animashree and Feng, Jiashi and Alvarez, Jose M.},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {27378--27394},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/zhou22m/zhou22m.pdf},
  url = 	 {https://proceedings.mlr.press/v162/zhou22m.html},
  abstract = 	 {Recent studies show that Vision Transformers (ViTs) exhibit strong robustness against various corruptions. Although this property is partly attributed to the self-attention mechanism, there is still a lack of an explanatory framework towards a more systematic understanding. In this paper, we examine the role of self-attention in learning robust representations. Our study is motivated by the intriguing properties of self-attention in visual grouping which indicate that self-attention could promote improved mid-level representation and robustness. We thus propose a family of fully attentional networks (FANs) that incorporate self-attention in both token mixing and channel processing. We validate the design comprehensively on various hierarchical backbones. Our model with a DeiT architecture achieves a state-of-the-art 47.6% mCE on ImageNet-C with 29M parameters. We also demonstrate significantly improved robustness in two downstream tasks: semantic segmentation and object detection}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/