PVT v2: Improved baselines with Pyramid Vision Transformer

Wenhai WangEnze XieXiang LiDeng-Ping FanKaitao SongDing LiangTong LuPing LuoLing Shao

article2021Computational Visual Media2,624 citations

Introduces PVT v2, an upgraded vision transformer that achieves linear computational complexity through overlapping patch embeddings and convolutional feed-forward networks, rivaling Swin Transformer on dense prediction tasks.

Listen

Deploying transformer-based neural networks in computer vision tasks has grown rapidly, offering an alternative to traditional convolutional neural networks (CNNs). However, early vision transformers face significant operational bottlenecks: their computational requirements scale quadratically with image resolution, non-overlapping image patch processing creates loss of spatial continuity, and rigid position encodings restrict models from handling arbitrary image dimensions.

The article demonstrates that introducing three architectural enhancements to the Pyramid Vision Transformer (PVT v1) resolves these computational and architectural limitations. The researchers set out to evaluate an upgraded framework, called PVT v2, across core computer vision tasks to establish whether it delivers state-of-the-art performance with lower computational overhead.

The authors implemented and benchmarked PVT v2 across six size configurations (B0 through B5) using three core architectural changes: linear spatial reduction attention using average pooling, overlapping patch embeddings via padded convolutions, and a convolutional feed-forward network with zero-padding position encoding. The models were evaluated through empirical experiments on standard industry benchmarks: ImageNet-1K (1.28 million training images) for classification, COCO 2017 (118,000 training images) across various one-stage and two-stage object detectors, and ADE20K for semantic segmentation.

The experimental findings show substantial performance and efficiency gains across all domains. First, on object detection benchmarks, PVT v2 consistently outperformed standard CNNs and competing transformers; when paired with Generalized Focal Loss on COCO, PVT v2 achieved an Average Precision of 50.2, surpassing Swin-T by 2.6 points and ResNet-50 by 5.7 points. Second, in semantic segmentation on ADE20K, PVT v2 configurations achieved more than a 5.3% improvement in mean Intersection over Union over PVT v1 baselines. Third, for ImageNet-1K classification, the largest model variant (B5) achieved an 83.8% top-1 accuracy, outperforming competing models such as Swin-B while requiring fewer parameters and floating-point operations. Finally, ablation studies showed that the linear attention module reduced computational load by 22% while maintaining near-identical accuracy, keeping computational scaling linear rather than quadratic as image size increases.

These findings indicate that vision transformers can now achieve the linear computational scaling of traditional CNNs while providing superior feature extraction accuracy. For organizations deploying vision-based artificial intelligence, this translates directly to reduced computing and infrastructure costs when processing high-resolution imagery, faster model inference times, and greater flexibility across varying image inputs without requiring custom position readjustments.

Organizations and practitioners developing computer vision applications should consider adopting PVT v2 backbones in place of older CNN or early transformer baselines for detection and segmentation pipelines. For resource-constrained or real-time deployment environments, the linear attention variant (PVT v2-Li) offers a practical trade-off, lowering computational load with negligible impact on final accuracy. Future research and development should explore extending these baselines across broader production domains and specialized computer vision workflows.

The presented evaluations are based on controlled benchmark datasets under standardized training and testing configurations. While confidence in these benchmark comparisons is high due to consistent testing protocols across multiple model variants, readers should note that performance may vary across unconventional real-world camera inputs or specialized, out-of-distribution vision environments.

Cover for PVT v2: Improved baselines with Pyramid Vision Transformer

Abstract

Transformer recently has presented encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs, including (1) linear complexity attention layer, (2) overlapping patch embedding, and (3) convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linear and achieves significant improvements on fundamental vision tasks such as classification, detection, and segmentation. Notably, the proposed PVT v2 achieves comparable or better performances than recent works such as Swin Transformer. We hope this work will facilitate state-of-the-art Transformer researches in computer vision. Code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Limitations in PVT v1
  • 3.2 Linear Spatial Reduction Attention
  • 3.3 Overlapping Patch Embedding
  • 3.4 Convolutional Feed-Forward
  • 3.5 Details of PVT v2 Series
  • 3.6 Advantages of PVT v2
  • 4 Experiment
  • 4.1 Image Classification
  • 4.2 Object Detection
  • 4.3 Semantic Segmentation
  • 4.4 Ablation Study
  • 4.4.1 Model Analysis
  • 4.4.2 Computation Overhead Analysis
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Linear Spatial Reduction Attention in PVT v2

    model/method

    Linear Spatial Reduction Attention (Linear SRA) reduces the computational complexity of the attention mechanism in Pyramid Vision Transformers from quadratic to linear with respect to the input resolution. In standard Spatial Reduction Attention (SRA), key and value spatial dimensions are downsampled using a convolutional layer with a fixed spatial reduction ratio RR, leaving the sequence length dependent on the input image resolution h×wh \times w.

    In Linear SRA, an adaptive average pooling layer replaces the convolution, projecting the spatial dimension of key (KK) and value (VV) feature maps from h×wh \times w to a fixed spatial dimension of P×PP \times P (set to 7×77 \times 7) before computing multi-head attention with the query (QQ). Because PP is constant regardless of input dimensions, the key and value token count remains fixed, enabling the attention layer to scale linearly with input image size while maintaining linear computational and memory costs.

  2. Knowl 2 — Computational Complexity of SRA versus Linear SRA

    equation

    For an input feature map of height hh, width ww, and channel dimension cc, the computational complexity of standard Spatial Reduction Attention (SRA) and Linear Spatial Reduction Attention (Linear SRA) are given by:

    Ω(SRA)=2h2w2cR2+hwc2R2\Omega(\text{SRA}) = \frac{2h^2w^2c}{R^2} + hwc^2R^2

    Ω(linear SRA)=2hwP2c\Omega(\text{linear SRA}) = 2hwP^2c

    where RR is the spatial reduction ratio used in standard SRA, and PP is the fixed target pooling grid size used in Linear SRA (P=7P = 7). In Linear SRA, because PP is independent of input dimensions hh and ww, the computational cost scales as O(hwc)O(hwc) rather than O((hw)2)O((hw)^2).

  3. Knowl 3 — Overlapping Patch Embedding for Tokenization

    model/method

    Overlapping patch embedding tokenizes image and feature representations while preserving local continuity across patch boundaries, resolving the boundary information loss found in non-overlapping patch embedding.

    For an input feature map of dimensions h×w×ch \times w \times c, the overlapping patch embedding applies a 2D convolution with stride SS, kernel size 2S−12S - 1, zero-padding size S−1S - 1, and C′C' output channels. This design enlarges the patch receptive field so that adjacent patches overlap by half of their area, generating an output representation of size (h/S)×(w/S)×C′(h/S) \times (w/S) \times C' while preserving spatial resolution consistency.

  4. Knowl 4 — Convolutional Feed-Forward Network and Zero-Padding Position Encoding

    model/method

    The Convolutional Feed-Forward Network (CFFN) introduces implicit positional information and local feature continuity into transformer blocks, removing the need for fixed-size positional embeddings that constrain vision transformers to fixed input resolutions.

    The CFFN places a 3×33 \times 3 depth-wise convolution (DWConv) with zero-padding size of 1 between the first fully-connected (FC) linear layer and the Gaussian Error Linear Unit (GELU) activation:

    Output=FC(GELU(DWConv3×3(FC(Input))))\text{Output} = \text{FC}\big(\text{GELU}(\text{DWConv}_{3\times 3}(\text{FC}(\text{Input})))\big)

    Because zero-padding encodes absolute position cues into convolutional operations, CFFN removes the requirement for fixed positional embeddings, allowing the network to process arbitrary and variable input image resolutions without positional interpolation.

  5. Knowl 5 — Architectural Configurations of the PVT v2 Model Family

    data/table

    The PVT v2 family scales across six capacity levels (B0, B1, B2, B3, B4, B5) along with a linear attention variant (B2-Li). Models follow a 4-stage hierarchical structure where feature map resolutions reduce by strides S1=4,S2=2,S3=2,S4=2S_1=4, S_2=2, S_3=2, S_4=2, while channel counts CiC_i increase. Each stage ii is parameterized by depth LiL_i, attention heads NiN_i, SRA reduction ratio RiR_i (or adaptive pooling size Pi=7P_i=7 for B2-Li), and FFN expansion ratio EiE_i.

    Stage B0 B1 B2 B2-Li B3 B4 B5
    Stage 1 C1=32C_1=32 C1=64C_1=64 C1=64C_1=64 C1=64C_1=64 C1=64C_1=64 C1=64C_1=64 C1=64C_1=64
    (H/4×W/4H/4 \times W/4) L1=2,R1=8L_1=2, R_1=8 L1=2,R1=8L_1=2, R_1=8 L1=3,R1=8L_1=3, R_1=8 L1=3,P1=7L_1=3, P_1=7 L1=3,R1=8L_1=3, R_1=8 L1=3,R1=8L_1=3, R_1=8 L1=3,R1=8L_1=3, R_1=8
    N1=1,E1=8N_1=1, E_1=8 N1=1,E1=8N_1=1, E_1=8 N1=1,E1=8N_1=1, E_1=8 N1=1,E1=8N_1=1, E_1=8 N1=1,E1=8N_1=1, E_1=8 N1=1,E1=8N_1=1, E_1=8 N1=1,E1=4N_1=1, E_1=4
    Stage 2 C2=64C_2=64 C2=128C_2=128 C2=128C_2=128 C2=128C_2=128 C2=128C_2=128 C2=128C_2=128 C2=128C_2=128
    (H/8×W/8H/8 \times W/8) L2=2,R2=4L_2=2, R_2=4 L2=2,R2=4L_2=2, R_2=4 L2=3,R2=4L_2=3, R_2=4 L2=3,P2=7L_2=3, P_2=7 L2=3,R2=4L_2=3, R_2=4 L2=8,R2=4L_2=8, R_2=4 L2=6,R2=4L_2=6, R_2=4
    N2=2,E2=8N_2=2, E_2=8 N2=2,E2=8N_2=2, E_2=8 N2=2,E2=8N_2=2, E_2=8 N2=2,E2=8N_2=2, E_2=8 N2=2,E2=8N_2=2, E_2=8 N2=2,E2=8N_2=2, E_2=8 N2=2,E2=4N_2=2, E_2=4
    Stage 3 C3=160C_3=160 C3=320C_3=320 C3=320C_3=320 C3=320C_3=320 C3=320C_3=320 C3=320C_3=320 C3=320C_3=320
    (H/16×W/16H/16 \times W/16) L3=2,R3=2L_3=2, R_3=2 L3=2,R3=2L_3=2, R_3=2 L3=6,R3=2L_3=6, R_3=2 L3=6,P3=7L_3=6, P_3=7 L3=18,R3=2L_3=18, R_3=2 L3=27,R3=2L_3=27, R_3=2 L3=40,R3=2L_3=40, R_3=2
    N3=5,E3=4N_3=5, E_3=4 N3=5,E3=4N_3=5, E_3=4 N3=5,E3=4N_3=5, E_3=4 N3=5,E3=4N_3=5, E_3=4 N3=5,E3=4N_3=5, E_3=4 N3=5,E3=4N_3=5, E_3=4 N3=5,E3=4N_3=5, E_3=4
    Stage 4 C4=256C_4=256 C4=512C_4=512 C4=512C_4=512 C4=512C_4=512 C4=512C_4=512 C4=512C_4=512 C4=512C_4=512
    (H/32×W/32H/32 \times W/32) L4=2,R4=1L_4=2, R_4=1 L4=2,R4=1L_4=2, R_4=1 L4=3,R4=1L_4=3, R_4=1 L4=3,P4=7L_4=3, P_4=7 L4=3,R4=1L_4=3, R_4=1 L4=3,R4=1L_4=3, R_4=1 L4=3,R4=1L_4=3, R_4=1
    N4=8,E4=4N_4=8, E_4=4 N4=8,E4=4N_4=8, E_4=4 N4=8,E4=4N_4=8, E_4=4 N4=8,E4=4N_4=8, E_4=4 N4=8,E4=4N_4=8, E_4=4 N4=8,E4=4N_4=8, E_4=4 N4=8,E4=4N_4=8, E_4=4
  6. Knowl 6 — ImageNet-1K Image Classification Performance of PVT v2

    data/table

    ImageNet-1K top-1 classification performance for PVT v2 models trained for 300 epochs from scratch with 224×224224 \times 224 input resolution. PVT v2 models show consistent accuracy gains over PVT v1 and contemporary CNN and vision transformer backbones at similar parameter counts and FLOPs.

    Method #Param (M) GFLOPs Top-1 Acc (%)
    PVT v2-B0 3.4 0.6 70.5
    ResNet18 11.7 1.8 69.8
    DeiT-Tiny/16 5.7 1.3 72.2
    PVT v1-Tiny 13.2 1.9 75.1
    PVT v2-B1 13.1 2.1 78.7
    ResNet50 25.6 4.1 76.1
    DeiT-Small/16 22.1 4.6 79.9
    PVT v1-Small 24.5 3.8 79.8
    Swin-T 29.0 4.5 81.3
    Twins-SVT-S 24.0 2.8 81.7
    PVT v2-B2-Li 22.6 3.9 82.1
    PVT v2-B2 25.4 4.0 82.0
    ResNet101 44.7 7.9 77.4
    PVT v1-Medium 44.2 6.7 81.2
    PVT v2-B3 45.2 6.9 83.2
    ResNet152 60.2 11.6 78.3
    PVT v1-Large 61.4 9.8 81.7
    Swin-S 50.0 8.7 83.0
    Twins-SVT-B 56.0 8.3 83.2
    PVT v2-B4 62.6 10.1 83.6
    ViT-Base/16 86.6 17.6 81.8
    DeiT-Base/16 86.6 17.6 81.8
    Swin-B 88.0 15.4 83.3
    Twins-SVT-L 99.2 14.8 83.7
    PVT v2-B5 82.0 11.8 83.8
  7. Knowl 7 — Object Detection and Instance Segmentation Performance on COCO 2017

    data/table

    COCO 2017 validation results using PVT v2 as the backbone for RetinaNet (1×1\times schedule) and Mask R-CNN (1×1\times schedule), demonstrating substantial improvements in bounding box AP (APb\text{AP}^b) and mask AP (APm\text{AP}^m) compared to ResNet and PVT v1 counterparts.

    RetinaNet 1×1\times Mask R-CNN 1×1\times
    Backbone #P (M) AP AP50\text{AP}_{50} AP75\text{AP}_{75} #P (M) APb\text{AP}^b AP50b\text{AP}^b_{50} AP75b\text{AP}^b_{75} APm\text{AP}^m AP50m\text{AP}^m_{50} AP75m\text{AP}^m_{75}
    PVT v2-B0 13.0 37.2 57.2 39.5 23.5 38.2 60.5 40.7 36.2 57.8 38.6
    ResNet18 21.3 31.8 49.6 33.6 31.2 34.0 54.0 36.7 31.2 51.0 32.7
    PVT v1-Tiny 23.0 36.7 56.9 38.9 32.9 36.7 59.2 39.3 35.1 56.7 37.3
    PVT v2-B1 23.8 41.2 61.9 43.9 33.7 41.8 64.3 45.9 38.8 61.2 41.6
    ResNet50 37.7 36.3 55.3 38.6 44.2 38.0 58.6 41.4 34.4 55.1 36.7
    PVT v1-Small 34.2 40.4 61.3 43.0 44.1 40.4 62.9 43.8 37.8 60.1 40.3
    PVT v2-B2-Li 32.3 43.6 64.7 46.8 42.2 44.1 66.3 48.4 40.5 63.2 43.6
    PVT v2-B2 35.1 44.6 65.6 47.6 45.0 45.3 67.1 49.6 41.2 64.2 44.4
    ResNet101 56.7 38.5 57.8 41.2 63.2 40.4 61.1 44.2 36.4 57.7 38.8
    PVT v1-Medium 53.9 41.9 63.1 44.3 63.9 42.0 64.4 45.6 39.0 61.6 42.1
    PVT v2-B3 55.0 45.9 66.8 49.3 64.9 47.0 68.1 51.7 42.5 65.7 45.7
    PVT v1-Large 71.1 42.6 63.7 45.4 81.0 42.9 65.0 46.6 39.5 61.9 42.5
    PVT v2-B4 72.3 46.1 66.9 49.2 82.2 47.5 68.7 52.0 42.7 66.1 46.1
    PVT v2-B5 91.7 46.2 67.1 49.5 101.6 47.4 68.6 51.9 42.5 65.7 46.0
  8. Knowl 8 — Dense Object Detection Comparisons with Swin Transformer

    data/table

    Performance comparison between Swin-T and PVT v2 backbones across four detection frameworks on COCO 2017 val with input size 1280×8001280 \times 800. PVT v2-B2 consistently outperforms Swin-T across all detectors, while PVT v2-B2-Li reduces computational cost (GFLOPs) substantially with minimal AP reduction.

    Backbone Method APb\text{AP}^b AP50b\text{AP}^b_{50} AP75b\text{AP}^b_{75} #P (M) GFLOPs
    ResNet50 Cascade Mask R-CNN 46.3 64.3 50.5 82 739
    Swin-T Cascade Mask R-CNN 50.5 69.3 54.9 86 745
    PVT v2-B2-Li Cascade Mask R-CNN 50.9 69.5 55.2 80 725
    PVT v2-B2 Cascade Mask R-CNN 51.1 69.8 55.3 83 788
    ResNet50 ATSS 43.5 61.9 47.0 32 205
    Swin-T ATSS 47.2 66.5 51.3 36 215
    PVT v2-B2-Li ATSS 48.9 68.1 53.4 30 194
    PVT v2-B2 ATSS 49.9 69.1 54.1 33 258
    ResNet50 GFL 44.5 63.0 48.3 32 208
    Swin-T GFL 47.6 66.8 51.7 36 215
    PVT v2-B2-Li GFL 49.2 68.2 53.7 30 197
    PVT v2-B2 GFL 50.2 69.4 54.7 33 261
    ResNet50 Sparse R-CNN 44.5 63.4 48.2 106 166
    Swin-T Sparse R-CNN 47.9 67.3 52.3 110 172
    PVT v2-B2-Li Sparse R-CNN 48.9 68.3 53.4 104 151
    PVT v2-B2 Sparse R-CNN 50.1 69.5 54.9 107 215
  9. Knowl 9 — ADE20K Semantic Segmentation Performance with Semantic FPN

    data/table

    Semantic segmentation results on the ADE20K validation set using Semantic FPN across various backbone networks at an input crop size of 512×512512 \times 512. PVT v2 backbones outperform PVT v1 counterparts by at least 5.3% mIoU under comparable parameter and FLOP budgets.

    Backbone #P (M) GFLOPs mIoU (%)
    PVT v2-B0 7.6 25.0 37.2
    ResNet18 15.5 32.2 32.9
    PVT v1-Tiny 17.0 33.2 35.7
    PVT v2-B1 17.8 34.2 42.5
    ResNet50 28.5 45.6 36.7
    PVT v1-Small 28.2 44.5 39.8
    PVT v2-B2-Li 26.3 41.0 45.1
    PVT v2-B2 29.1 45.8 45.2
    ResNet101 47.5 65.1 38.8
    ResNeXt101-32x4d 47.1 64.7 39.7
    PVT v1-Medium 48.0 61.0 41.6
    PVT v2-B3 49.0 62.4 47.3
    PVT v1-Large 65.1 79.6 42.1
    PVT v2-B4 66.3 81.3 47.9
    ResNeXt101-64x4d 86.4 103.9 40.2
    PVT v2-B5 85.7 91.1 48.7
  10. Knowl 10 — Ablation of Architectural Components in PVT v2

    data/table

    Ablation study demonstrating the cumulative performance impact of adding Overlapping Patch Embedding (OPE), Convolutional Feed-Forward Network (CFFN), and Linear SRA (LSRA) to the PVT baseline. Evaluated on ImageNet top-1 accuracy and COCO RetinaNet (1×1\times schedule) bounding box AP.

    # Setting ImageNet Acc (%) COCO RetinaNet 1×1\times
    #P (M) GFLOPs AP
    1 PVT v1-Small 79.8 34.2 285.8 40.4
    2 + OPE 81.1 34.9 288.6 42.2
    3 ++ CFFN (PVT v2-B2) 82.0 35.1 290.7 44.6
    4 +++ LSRA (PVT v2-B2-Li) 82.1 32.3 227.4 43.6

    Adding OPE yields a +1.3%+1.3\% accuracy gain on ImageNet and +1.8+1.8 AP on COCO. Adding CFFN further increases accuracy by +0.9%+0.9\% on ImageNet and +2.4+2.4 AP on COCO while eliminating fixed-size position embeddings. Applying LSRA reduces FLOPs on COCO by 21.8%21.8\% while maintaining comparable top-1 ImageNet accuracy (82.1%82.1\% vs. 82.0%82.0\%).

Coverage note — No substantial contributed material was omitted; qualitative detection/segmentation visual examples from Figure 3 were omitted as they duplicate quantitative benchmark tables.

References

  1. 1.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; S. Gelly, S.; et al. An image is worth 16×16 words: Transformers for image recognition at scale. In: Proceedings of the International Conference on Learning Representations, 2021.
  2. 2.Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Jegou, H. Training data-efficient image transformers & distillation through attention. In: Proceedings of the 38th International Conference on Machine Learning, 2021.
  3. 3.Wang, W.; Xie, E.; Li, X.; Fan, D.-P.; Song, K.; Liang, D.; Lu, T.; Luo, P.; Shao, L. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 568–578, 2021.
  4. 4.Wu, H.; Xiao, B.; Codella, N.; Liu, M.; Dai, X.; Yuan, L.; Zhang, L. CvT: Introducing convolutions to vision transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 22–31, 2021.
  5. 5.Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 10012–10022, 2021.
  6. 6.Xu, W.; Xu, Y.; Chang, T.; Tu, Z. Co-scale conv-attentional image transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 9981–9990, 2021.
  7. 7.Graham, B.; El-Nouby, A.; Touvron, H.; Stock, P.; Joulin, A.; Jegou, H. LeViT: A vision transformer in ConvNet’s clothing for faster inference. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 12259–12269, 2021.
  8. 8.Chu, X.; Tian, Z.; Wang, Y.; Zhang, B.; Ren, H.; Wei, X.; Xia, H.; Shen, C. Twins: Revisiting the design of spatial attention in vision transformers. In: Proceedings of the 35th Conference on Neural Information Processing Systems, 2021.
  9. 9.Lin, T. Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollar, P.; Zitnick, C. L. Microsoft COCO: Common objects in context. In: Computer Vision – ECCV 2014. Lecture Notes in Computer Science, Vol. 8693. Fleet, D.; Pajdla, T.; Schiele, B.; Tuytelaars, T. Eds. Springer Cham, 740–755, 2014.
  10. 10.Zhou, B. L.; Zhao, H.; Puig, X.; Fidler, S.; Barriuso, A.; Torralba, A. Scene parsing through ADE20K dataset. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5122–5130, 2017.
  11. 11.Dong, B.; Wang, W.; Fan, D.-P.; Li, J.; Fu, H.; Shao, L. Polyp-PVT: Polyp segmentation with pyramid vision transformers. arXiv preprint arXiv:2108.06932, 2021.
  12. 12.Li, X.; Wang, W.; Wu, L.; Chen, S.; Hu, X.; Li, J.; Tang, J.; Yang, J. Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection. In: Proceedings of the 34th Conference on Neural Information Processing Systems, 2020.
  13. 13.He, K. M.; Zhang, X. Y.; Ren, S. Q.; Sun, J. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In: Proceedings of the IEEE International Conference on Computer Vision, 1026–1034, 2015.
  14. 14.Deng, J.; Dong, W.; Socher, R.; Li, L. J.; Kai, L.; Li, F. F. ImageNet: A large-scale hierarchical image database. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 248–255, 2009.
  15. 15.Yuan, L.; Chen, Y.; Wang, T.; Yu, W.; Shi, Y.; Jiang, Z.; Tay, F. E.; Feng, J.; Yan, S. Tokens-to-token ViT: Training vision transformers from scratch on ImageNet. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 558–567, 2021.
  16. 16.Han, K.; Xiao, A.; Wu, E.; Guo, J.; Xu, C.; Wang, Y. Transformer in transformer. arXiv preprint arXiv:2103.00112, 2021.
  17. 17.Chu, X.; Tian, Z.; Zhang, B.; Wang, X.; Wei, X.; Xia, H.; Shen, C. Conditional positional encodings for vision transformers. arXiv preprint arXiv:2102.10882, 2021.
  18. 18.Chen, C.-F.; Fan, Q.; Panda, R. CrossViT: Cross-attention multi-scale vision transformer for image classification. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 357–366, 2021.
  19. 19.Li, Y.; Zhang, K.; Cao, J.; Timofte, R.; van Gool, L. LocalViT: Bringing locality to vision transformers. arXiv preprint arXiv:2104.05707, 2021.
  20. 20.Islam, M. A.; Jia, S.; Bruce, N. D. B. How much position information do convolutional neural networks encode? In: Proceedings of the International Conference on Learning Representations, 2020.
  21. 21.Howard, A. G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T; Andreetto, M.; Adam, H. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  22. 22.Hendrycks, D.; Gimpel, K. Gaussian error linear units (GELUs). arXiv preprint arXiv:1606.08415, 2016.
  23. 23.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; Polosukhin, I. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems, 6000–6010, 2017.
  24. 24.He, K. M.; Zhang, X. Y.; Ren, S. Q.; Sun, J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–778, 2016.
  25. 25.Xie, S. N.; Girshick, R.; Dollar, P.; Tu, Z. W.; He, K. M. Aggregated residual transformations for deep neural networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5987–5995, 2017.
  26. 26.Radosavovic, I.; Kosaraju, R. P.; Girshick, R.; He, K. M.; Dollar, P. Designing network design spaces. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10425–10433, 2020.
  27. 27.Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S. A.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. ImageNet large scale visual recognition challenge. International Journal of Computer Vision Vol. 115, No. 3, 211–252, 2015.
  28. 28.Szegedy, C.; Liu, W.; Jia, Y. Q.; Sermanet, P.; Reed, S.; Anguelov, D.; Erhan, D.; Vanhoucke, V.; Rabinovich, A. Going deeper with convolutions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1–9, 2015.
  29. 29.Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; Wojna, Z. Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2818–2826, 2016.
  30. 30.Zhang, H.; Cisse, M.; Dauphin, Y. N.; Lopez-Paz, D. mixup: Beyond empirical risk minimization. In: Proceedings of the International Conference on Learning Representations, 2018.
  31. 31.Zhong, Z.; Zheng, L.; Kang, G. L.; Li, S. Z.; Yang, Y. Random erasing data augmentation. Proceedings of the AAAI Conference on Artificial Intelligence Vol. 34, No. 7, 13001–13008, 2020.
  32. 32.Loshchilov, I.; Hutter, F. Decoupled weight decay regularization. In: Proceedings of the International Conference on Learning Representations, 2019.
  33. 33.Loshchilov, I.; Hutter, F. SGDR: Stochastic gradient descent with warm restarts. In: Proceedings of the International Conference on Learning Representations, 2017.
  34. 34.Lin, T. Y.; Goyal, P.; Girshick, R.; He, K. M.; Dollar, P. Focal loss for dense object detection. In: Proceedings of the IEEE International Conference on Computer Vision, 2999–3007, 2017.
  35. 35.He, K. M.; Gkioxari, G.; Dollar, P.; Girshick, R. Mask R-CNN. In: Proceedings of the IEEE International Conference on Computer Vision, 2980–2988, 2017.
  36. 36.Cai, Z. W.; Vasconcelos, N. Cascade R-CNN: Delving into high quality object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6154–6162, 2018.
  37. 37.Zhang, S. F.; Chi, C.; Yao, Y. Q.; Lei, Z.; Li, S. Z. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9756–9765, 2020.
  38. 38.Sun, P. Z.; Zhang, R. F.; Jiang, Y.; Kong, T.; Xu, C. F.; Zhan, W.; Tomizuka, M.; Li, L.; Yuan, Z.; Wang, C.; et al. Sparse R-CNN: End-to-end object detection with learnable proposals. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 14449–14458, 2021.
  39. 39.Glorot, X.; Bengio, Y. Understanding the difficulty of training deep feedforward neural networks. In: Proceedings of the 13th International Conference on Artificial Intelligence and Statistics, 249–256, 2010.
  40. 40.Chen, K.; Wang, J. Q.; Pang, J. M.; Cao, Y. H.; Xiong, Y.; Li, X.; Sun, S.; Feng, W.; Liu, Z.; Xu, J.; et al. MMDetection: Open MMLab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.
  41. 41.Kirillov, A.; Girshick, R.; He, K. M.; Dollar, P. Panoptic feature pyramid networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6392–6401, 2019.
  42. 42.Chen, L. C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A. L. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Transactions on Pattern Analysis and Machine Intelligence Vol. 40, No. 4, 834–848, 2018.

Citation

MLA
Wang, W., et al. “PVT V2: Improved Baselines with Pyramid Vision Transformer”. Computational Visual Media, vol. 8, no. 3, 2022, pp. 415–24, https://doi.org/10.1007/s41095-022-0274-8.
APA
Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., & Shao, L. (2022). PVT v2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3), 415–424. https://doi.org/10.1007/s41095-022-0274-8
Chicago
Wang, W., E. Xie, X. Li, et al. 2022. “PVT V2: Improved Baselines with Pyramid Vision Transformer”. Computational Visual Media 8 (3): 415–24. https://doi.org/10.1007/s41095-022-0274-8.
Harvard
Wang, W. et al. (2022) “PVT v2: Improved baselines with pyramid vision transformer”, Computational Visual Media, 8(3), pp. 415–424. Available at: https://doi.org/10.1007/s41095-022-0274-8.
Vancouver
1. Wang W, Xie E, Li X, Fan D-P, Song K, Liang D, Lu T, Luo P, Shao L (2022) PVT v2: Improved baselines with pyramid vision transformer. Computational Visual Media 8:415–424

BibTeX

@article{Wang_2022, title={PVT v2: Improved baselines with pyramid vision transformer}, volume={8}, ISSN={2096-0433}, url={http://dx.doi.org/10.1007/s41095-022-0274-8}, DOI={10.1007/s41095-022-0274-8}, number={3}, journal={Computational Visual Media}, publisher={Tsinghua University Press}, author={Wang, Wenhai and Xie, Enze and Li, Xiang and Fan, Deng-Ping and Song, Kaitao and Liang, Ding and Lu, Tong and Luo, Ping and Shao, Ling}, year={2022}, month=Sept, pages={415–424} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/