SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation

Meng-Hao GuoChenggang LuQibin HouZheng LiuMing-Ming ChengShiyong Hu

article2022NeurIPS1,379 citations

Demonstrates that convolutional attention can surpass transformer-based models in semantic segmentation through SegNeXt, an architecture that achieves superior accuracy on standard benchmarks like ADE20K and Pascal VOC using drastically fewer parameters and computations.

Listen

Semantic segmentation—the computer vision task of assigning a category label to every pixel in an image—is essential for autonomous driving, robotics, and remote sensing. While recent transformer-based architectures dominate performance leaderboards, their self-attention mechanisms require heavy computational resources that grow quadratically with image resolution. The article demonstrates that a purely convolutional network architecture can achieve superior accuracy and detail preservation while requiring significantly less computation and memory.

To accomplish this, the authors designed SegNeXt, an architecture built around a novel multi-scale convolutional attention module that replaces standard self-attention with lightweight, multi-branch strip convolutions. The overall system pairs this convolutional encoder with a lightweight decoder that aggregates global context from high-level features. The authors evaluated SegNeXt across seven standard vision benchmarks, including ADE20K, Cityscapes, COCO-Stuff, Pascal VOC, and the remote sensing dataset iSAID.

Across all benchmarks, SegNeXt consistently outperformed both established convolutional baselines and state-of-the-art vision transformers. On the ADE20K dataset, it improved accuracy by an average of 2.0% mean Intersection over Union while using equal or fewer computations than competing methods. When processing high-resolution urban scenes in Cityscapes, SegNeXt-S achieved higher accuracy (81.3% versus 81.0%) than SegFormer-B2 while using only one-sixth of the computational operations and half the parameters. Furthermore, the largest model variant achieved 90.6% accuracy on Pascal VOC 2012, matching or exceeding top existing models while requiring one-tenth of the parameters.

These findings demonstrate that convolutional approaches remain highly competitive with transformers when tailored for multi-scale spatial attention and linear computational scaling. In real-world applications, this allows organizations to deploy high-accuracy computer vision models on constrained hardware and edge devices, reducing cloud infrastructure costs, latency, and power consumption without sacrificing segmentation quality.

For engineering and product teams building vision-based systems, the article provides a strong rationale to consider modern convolutional attention architectures like SegNeXt rather than defaulting to transformer backbones. Future work should focus on validating this approach at larger scales—specifically models exceeding 100 million parameters—and assessing whether similar multi-scale convolutional attention mechanisms transfer effectively to other vision tasks and natural language processing.

  • Paper: A ConvNet for the 2020s, Zhuang Liu et al. (2022). ConvNeXt modernized convolutional neural network design to compete directly with vision transformers, providing the core architectural philosophy that SegNeXt adapts into convolutional attention for semantic segmentation.
  • Paper: SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers, Enze Xie et al. (2021). SegFormer established a lightweight transformer benchmark with hierarchical encoding and an MLP decoder, serving as the direct transformer baseline that SegNeXt seeks to outperform using convolutional attention.
  • Paper: Segmenter: Transformer for Semantic Segmentation, Robin Strudel et al. (2021). Segmenter represents the purely attention-driven vision transformer paradigm for semantic segmentation that SegNeXt explicitly re-examines and contrasts against cheap convolutional operations.
  • Paper: Large Kernel Matters — Improve Semantic Segmentation by Global Convolutional Network, Chao Peng et al. (2017). This paper introduced large, factorized convolutional kernels to capture broad contextual information in semantic segmentation, establishing the foundational principle behind SegNeXt's large-kernel convolutional attention.
  • Paper: CCNet: Criss-Cross Attention for Semantic Segmentation, Zilong Huang et al. (2019). CCNet introduced efficient contextual aggregation via criss-cross attention paths, pioneering the search for computationally cheaper spatial attention alternatives in dense prediction.
  • Paper: Dual Attention Network for Scene Segmentation, Jun Fu et al. (2019). DANet formulated spatial and channel attention mechanisms for scene segmentation, providing fundamental context-modeling concepts that SegNeXt re-engineers into streamlined convolutional blocks.
  • Paper: Pyramid Scene Parsing Network, Hengshuang Zhao et al. (2016). PSPNet introduced pyramid spatial pooling for capturing global contextual information in scene parsing, serving as a classical reference point for the contextual characteristics re-examined by SegNeXt.
  • Paper: Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation, Liang-Chieh Chen et al. (2018). DeepLabv3+ combined multi-scale context aggregation with depthwise separable convolutions, setting a standard for efficient dense prediction that SegNeXt advances.
Cover for SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation

Abstract

We present SegNeXt, a simple convolutional network architecture for semantic segmentation. Recent transformer-based models have dominated the field of semantic segmentation due to the efficiency of self-attention in encoding spatial information. In this paper, we show that convolutional attention is a more efficient and effective way to encode contextual information than the self-attention mechanism in transformers. By re-examining the characteristics owned by successful segmentation models, we discover several key components leading to the performance improvement of segmentation models. This motivates us to design a novel convolutional attention network that uses cheap convolutional operations. Without bells and whistles, our SegNeXt significantly improves the performance of previous state-of-the-art methods on popular benchmarks, including ADE20K, Cityscapes, COCO-Stuff, Pascal VOC, Pascal Context, and iSAID. Notably, SegNeXt outperforms EfficientNet-L2 w/ NAS-FPN and achieves 90.6% mIoU on the Pascal VOC 2012 test leaderboard using only 1/10 parameters of it. On average, SegNeXt achieves about 2.0% mIoU improvements compared to the state-of-the-art methods on the ADE20K datasets with the same or fewer computations. Code is available at this https URL (Jittor) and this https URL (Pytorch).

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Semantic Segmentation
  • 2.2 Multi-Scale Networks
  • 2.3 Attention Mechanisms
  • 3 Method
  • 3.1 Convolutional Encoder
  • 3.2 Decoder
  • 4 Experiments
  • 4.1 Encoder Performance on ImageNet
  • 4.2 Ablation study
  • 4.3 Comparison with state-of-the-art methods
  • 5 Conclusions and Discussion
  • References

Knowls

  1. Knowl 1 — Multi-Scale Convolutional Attention Module

    model/method

    The Multi-Scale Convolutional Attention (MSCA) module is an attention mechanism designed to capture spatial dependencies across multiple contextual scales using lightweight convolutional operations rather than standard self-attention.

    Given an input feature map F∈RC×H×WF \in \mathbb{R}^{C \times H \times W}, MSCA operates in three successive processing stages:

    1. Local Information Aggregation: Local spatial patterns are aggregated using a standard depth-wise convolution with kernel size 5×55 \times 5.
    2. Multi-Scale Context Aggregation: The locally aggregated features are processed through four parallel branches:
      • Branch 0 (Identity): An identity mapping preserving the initial features.
      • Branch 1: A pair of depth-wise strip convolutions with kernel sizes 1×71 \times 7 and 7×17 \times 1.
      • Branch 2: A pair of depth-wise strip convolutions with kernel sizes 1×111 \times 11 and 11×111 \times 1.
      • Branch 3: A pair of depth-wise strip convolutions with kernel sizes 1×211 \times 21 and 21×121 \times 1. Decomposing a large 2D depth-wise convolution of size k×kk \times k into a succession of 1×k1 \times k and k×1k \times 1 1D strip convolutions drastically reduces parameter and computational count while capturing strip-like elongated object features.
    3. Channel Relationship Modeling and Attention Gating: The sum of the outputs from all four branches is passed through a 1×11 \times 1 pointwise convolution to model inter-channel relationships. The resulting output serves directly as spatial-channel attention weights and is element-wise multiplied with the original input feature FF to produce the final refined representation.
  2. Knowl 2 — Mathematical Formulation of Multi-Scale Convolutional Attention

    equation

    Let F∈RC×H×WF \in \mathbb{R}^{C \times H \times W} denote the input feature tensor, where CC is the channel dimension, and HH and WW denote spatial height and width. The attention map Att∈RC×H×W\text{Att} \in \mathbb{R}^{C \times H \times W} generated by the Multi-Scale Convolutional Attention (MSCA) module is formulated as:

    Att=Conv1×1(∑i=03Scalei(DW-Conv(F)))\text{Att} = \text{Conv}_{1\times 1}\left(\sum_{i=0}^{3} \text{Scale}_i(\text{DW-Conv}(F))\right)

    The final output tensor Out∈RC×H×W\text{Out} \in \mathbb{R}^{C \times H \times W} is computed via element-wise matrix multiplication:

    Out=Att⊗F\text{Out} = \text{Att} \otimes F

    where:

    • DW-Conv(⋅)\text{DW-Conv}(\cdot) denotes a depth-wise convolution with a 5×55 \times 5 kernel used for local context aggregation.
    • Scale0\text{Scale}_0 represents the identity connection: Scale0(X)=X\text{Scale}_0(X) = X.
    • Scalei\text{Scale}_i for i∈{1,2,3}i \in \{1, 2, 3\} denotes the ii-th depth-wise strip convolution branch approximating a large 2D kernel. Specifically, Scale1\text{Scale}_1 uses kernel sizes 7×17\times 1 and 1×71\times 7; Scale2\text{Scale}_2 uses kernel sizes 11×111\times 1 and 1×111\times 11; and Scale3\text{Scale}_3 uses kernel sizes 21×121\times 1 and 1×211\times 21.
    • Conv1×1(⋅)\text{Conv}_{1\times 1}(\cdot) denotes a 1×11 \times 1 pointwise convolution that models inter-channel relationships.
    • ⊗\otimes denotes the Hadamard (element-wise) product.
  3. Knowl 3 — MSCAN Hierarchical Convolutional Encoder Architecture

    model/method

    The Multi-Scale Convolutional Attention Network (MSCAN) encoder follows a 4-stage hierarchical pyramid design with decreasing spatial resolutions: H4×W4\frac{H}{4} \times \frac{W}{4}, H8×W8\frac{H}{8} \times \frac{W}{8}, H16×W16\frac{H}{16} \times \frac{W}{16}, and H32×W32\frac{H}{32} \times \frac{W}{32}.

    Each stage comprises:

    1. Down-sampling Block: A convolution with kernel size 3×33 \times 3 and stride 2, followed by a Batch Normalization (BN) layer. Batch normalization is employed instead of Layer Normalization across the encoder blocks as it yields superior segmentation accuracy.
    2. A Sequence of Building Blocks: Each building block consists of a Batch Normalization layer, an MSCA attention module, a residual addition, a second Batch Normalization layer, and a Feed-Forward Network (FFN). The FFN consists of a 1×11 \times 1 convolution, a 3×33 \times 3 depth-wise convolution, a GELU non-linear activation, and a final 1×11 \times 1 projection convolution.

    Four encoder configurations are defined based on channel capacities (CC), block depths (LL), and FFN expansion ratios (e.r.e.r.):

    Stage Output Size e.r.e.r. SegNeXt-T SegNeXt-S SegNeXt-B SegNeXt-L
    1 H4×W4×C\frac{H}{4} \times \frac{W}{4} \times C 8 C=32,L=3C=32, L=3 C=64,L=2C=64, L=2 C=64,L=3C=64, L=3 C=64,L=3C=64, L=3
    2 H8×W8×C\frac{H}{8} \times \frac{W}{8} \times C 8 C=64,L=3C=64, L=3 C=128,L=2C=128, L=2 C=128,L=3C=128, L=3 C=128,L=5C=128, L=5
    3 H16×W16×C\frac{H}{16} \times \frac{W}{16} \times C 4 C=160,L=5C=160, L=5 C=320,L=4C=320, L=4 C=320,L=12C=320, L=12 C=320,L=27C=320, L=27
    4 H32×W32×C\frac{H}{32} \times \frac{W}{32} \times C 4 C=256,L=2C=256, L=2 C=512,L=2C=512, L=2 C=512,L=3C=512, L=3 C=512,L=3C=512, L=3
    Decoder Dimension 256 256 512 1024
    Parameters (M) 4.3 13.9 27.6 48.9
  4. Knowl 4 — SegNeXt Decoder Architecture with Hamburger Global Context Modeling

    model/method

    The SegNeXt decoder is designed to extract high-level global semantic representations with computational efficiency by fusing feature maps strictly from the final three stages of the MSCAN encoder (Stages 2, 3, and 4), deliberately excluding Stage 1.

    Key components of the decoder include:

    • Exclusion of Stage 1: Because MSCAN is convolutional, features from Stage 1 contain excessive low-level visual noise that degrades semantic segmentation accuracy while introducing heavy computational and memory overhead.
    • Multi-Level Feature Aggregation: Feature maps from Stage 2 (spatial resolution H8×W8\frac{H}{8} \times \frac{W}{8}), Stage 3 (H16×W16\frac{H}{16} \times \frac{W}{16}), and Stage 4 (H32×W32\frac{H}{32} \times \frac{W}{32}) are each passed through an MLP layer to project them to a unified decoder channel dimension (DD). The lower-resolution features from Stages 3 and 4 are bilinearly upsampled to H8×W8\frac{H}{8} \times \frac{W}{8} and concatenated with Stage 2 features along the channel dimension.
    • Global Context Modeling via Hamburger Module: The concatenated multi-level features are processed by a matrix decomposition-based Hamburger (Ham) module, which models global context with linear computational complexity O(n)\mathcal{O}(n) with respect to pixel count nn.
    • Segmentation Head: The global contextual representation from the Hamburger module is passed to an MLP head to compute per-pixel category logits.
  5. Knowl 5 — Semantic Segmentation Performance on ADE20K, Cityscapes, and COCO-Stuff

    data/table

    SegNeXt outperforms state-of-the-art vision transformer and CNN architectures across standard segmentation benchmarks while maintaining lower or comparable parameter counts and FLOPs. FLOPs are measured on inputs of 512×512512 \times 512 for ADE20K and COCO-Stuff, and 2048×10242048 \times 1024 for Cityscapes.

    Model Params ADE20K Cityscapes COCO-Stuff
    (M) GFLOPs mIoU (SS/MS) GFLOPs mIoU (SS/MS) GFLOPs mIoU (SS/MS)
    SegFormer-B0 3.8 8.4 37.4 / 38.0 125.5 76.2 / 78.1 8.4 35.6 / -
    SegNeXt-T 4.3 6.6 41.1 / 42.2 50.5 79.8 / 81.4 6.6 38.7 / 39.1
    SegFormer-B1 13.7 15.9 42.2 / 43.1 243.7 78.5 / 80.0 15.9 40.2 / -
    HRFormer-S 13.5 109.5 44.0 / 45.1 835.7 80.0 / 81.0 109.5 37.9 / 38.9
    SegNeXt-S 13.9 15.9 44.3 / 45.8 124.6 81.3 / 82.7 15.9 42.2 / 42.8
    SegFormer-B2 27.5 62.4 46.5 / 47.5 717.1 81.0 / 82.2 62.4 44.6 / -
    MaskFormer 42.0 55.0 46.7 / 48.8 - - - -
    SegNeXt-B 27.6 34.9 48.5 / 49.9 275.7 82.6 / 83.8 34.9 45.8 / 46.3
    SegFormer-B3 47.3 79.0 49.4 / 50.0 962.9 81.7 / 83.3 79.0 45.5 / -
    Mask2Former 47.0 74.0 47.7 / 49.6 - - - -
    HRFormer-B 56.2 280.0 48.7 / 50.0 2223.8 81.9 / 82.6 280.0 42.4 / 43.3
    MaskFormer 63.0 79.0 49.8 / 51.0 - - - -
    SegNeXt-L 48.9 70.0 51.0 / 52.1 577.5 83.2 / 83.9 70.0 46.5 / 47.2

    SegNeXt-B achieves 48.5%48.5\% single-scale (SS) mIoU on ADE20K compared to 46.5%46.5\% for SegFormer-B2 while requiring 44%44\% fewer FLOPs (34.9G vs 62.4G). On high-resolution Cityscapes images, SegNeXt-S outperforms SegFormer-B2 (81.3%81.3\% vs 81.0%81.0\% mIoU) using approximately 16\frac{1}{6} the FLOPs (124.6G vs 717.1G) and half the parameter count (13.9M vs 27.6M).

  6. Knowl 6 — Pascal VOC 2012 Benchmark Results and Parameter Efficiency

    data/table

    Evaluation on the Pascal VOC 2012 dataset demonstrates that SegNeXt achieves top semantic segmentation accuracy while maintaining extreme parameter efficiency compared to heavy convolutional and neural architecture search models.

    Method Backbone mIoU (%)
    DANet ResNet101 82.6
    OCRNet HRNetV2-W48 84.5
    HamNet ResNet101 85.9
    EncNet (with COCO pretraining) ResNet101 85.9
    EMANet (with COCO pretraining) ResNet101 87.7
    DeepLabV3+ (with COCO pretraining) Xception-71 87.8
    DeepLabV3+ (with JFT-300M pretraining) Xception-JFT 89.0
    NAS-FPN (with 300M extra unlabeled images) EfficientNet-L2 90.5
    SegNeXt-T MSCAN-T 82.7
    SegNeXt-S MSCAN-S 85.3
    SegNeXt-B MSCAN-B 87.5
    SegNeXt-L (with COCO pretraining) MSCAN-L 90.6

    SegNeXt-L (with COCO pretraining) achieves 90.6%90.6\% mIoU on the Pascal VOC 2012 test leaderboard, surpassing EfficientNet-L2 w/ NAS-FPN (90.5%90.5\% mIoU) while utilizing only 48.7M parameters compared to EfficientNet-L2's 485M parameters (roughly 110\frac{1}{10} the parameters) and without requiring 300M additional unlabeled images.

  7. Knowl 7 — Component Ablation of Multi-Scale Convolutional Attention

    data/table

    An ablation study on ImageNet-1K classification (Top-1 accuracy) and ADE20K semantic segmentation (mIoU) verifies the contribution of each component within the MSCA module. The base architecture tested is MSCAN-T.

    7×77 \times 7 Branch 11×1111 \times 11 Branch 21×2121 \times 21 Branch 1×11 \times 1 Conv Attention Gating Top-1 (%) mIoU (%)
    ✓ × × ✓ ✓ 74.7 39.6
    × ✓ × ✓ ✓ 75.2 39.7
    × × ✓ ✓ ✓ 75.3 40.0
    ✓ ✓ ✓ × ✓ 74.8 39.1
    ✓ ✓ ✓ ✓ × 75.5 40.5
    ✓ ✓ ✓ ✓ ✓ 75.9 41.1

    Key observations:

    • Combining all three multi-scale strip convolution branches (7×77 \times 7, 11×1111 \times 11, and 21×2121 \times 21) outperforms any individual branch alone (achieving 41.1%41.1\% mIoU vs. 39.6%39.6\%, 39.7%39.7\%, and 40.0%40.0\%).
    • The 1×11 \times 1 channel-mixing convolution provides a +2.0%+2.0\% mIoU improvement (41.1%41.1\% vs. 39.1%39.1\%).
    • Element-wise attention multiplication provides an adaptive gating mechanism that improves segmentation performance by +0.6%+0.6\% mIoU (41.1%41.1\% vs. 40.5%40.5\%).
  8. Knowl 8 — Decoder Architecture and Context Module Ablation

    data/table

    Ablation experiments evaluate different global context attention mechanisms and multi-level stage aggregation strategies in the decoder on the ADE20K dataset (input size 512×512512 \times 512).

    1. Attention Mechanism in Decoder (MSCAN-B encoder):

    Architecture Params (M) GFLOPs mIoU (SS) mIoU (MS)
    SegNeXt-B w/ CC (Criss-Cross) 27.8 35.7 47.3 48.6
    SegNeXt-B w/ EMA 27.4 32.3 48.0 49.1
    SegNeXt-B w/ NL (Non-Local) 27.6 40.9 48.6 50.0
    SegNeXt-B w/ Ham (Hamburger) 27.6 34.9 48.5 49.9

    The Hamburger (Ham) module matches the high accuracy of Non-Local attention (48.5%48.5\% vs. 48.6%48.6\% SS mIoU) while operating with linear computational complexity O(n)\mathcal{O}(n) rather than quadratic complexity O(n2)\mathcal{O}(n^2) and consuming 15%15\% fewer GFLOPs (34.9 vs. 40.9).

    2. Feature Stage Fusion Layout (MSCAN-T encoder):

    Architecture Params (M) GFLOPs mIoU (SS) mIoU (MS)
    SegNeXt-T (a) [All stages 1–4 MLP] 4.4 10.0 40.3 41.1
    SegNeXt-T (b) [Stage 4 only + Head] 4.2 4.9 30.9 40.6
    SegNeXt-T (c) [Proposed: Stages 2, 3, 4 + Ham] 4.3 6.6 41.1 42.2
    SegNeXt-T (c) w/ Stage 1 [Stages 1, 2, 3, 4 + Ham] 4.3 12.1 40.7 42.2

    Fusing only Stages 2, 3, and 4 yields higher single-scale performance (41.1%41.1\% mIoU) than including Stage 1 (40.7%40.7\% mIoU) while reducing GFLOPs from 12.1 to 6.6, demonstrating that convolutional Stage 1 representations introduce noisy low-level details and unnecessary computation.

  9. Knowl 9 — Real-Time Semantic Segmentation on Cityscapes Benchmark

    data/table

    SegNeXt-T provides efficient real-time inference on the Cityscapes test set without dedicated hardware or software acceleration (tested using a single NVIDIA RTX 3090 GPU and AMD EPYC 7543 32-core processor).

    Method Input Size mIoU (%)
    ESPNet 512×1024512 \times 1024 60.3
    ESPNetv2 512×1024512 \times 1024 66.2
    ICNet 1024×20481024 \times 2048 69.5
    DFANet 1024×10241024 \times 1024 71.3
    BiSeNet 768×1536768 \times 1536 74.6
    BiSeNetv2 512×1024512 \times 1024 75.3
    DF2-Seg 1024×20481024 \times 2048 74.8
    SwiftNet 1024×20481024 \times 2048 75.5
    SFNet 1024×20481024 \times 2048 77.8
    SegNeXt-T 768×1536768 \times 1536 78.0

    SegNeXt-T achieves 78.0%78.0\% mIoU at 25 frames per second (FPS) on 768×1536768 \times 1536 inputs, setting a new state-of-the-art accuracy benchmark for real-time semantic segmentation on Cityscapes.

  10. Knowl 10 — Semantic Segmentation on Remote Sensing Benchmark iSAID

    data/table

    Evaluation of SegNeXt on the large-scale aerial remote sensing benchmark iSAID under single-scale (SS) test mode demonstrates consistent improvements across all model capacities compared to established CNN and vision transformer approaches.

    Method Backbone mIoU (%)
    DenseASPP ResNet50 57.3
    PSPNet ResNet50 60.3
    SemanticFPN ResNet50 62.1
    RefineNet ResNet50 60.2
    HRNet HRNetW-18 61.5
    GSCNN ResNet50 63.4
    SFNet ResNet50 64.3
    RANet ResNet50 62.1
    PointRend ResNet50 62.8
    FarSeg ResNet50 63.7
    UperNet Swin-T 64.6
    PointFlow ResNet50 66.9
    SegNeXt-T MSCAN-T 68.3
    SegNeXt-S MSCAN-S 68.8
    SegNeXt-B MSCAN-B 69.9
    SegNeXt-L MSCAN-L 70.3

    The smallest variant, SegNeXt-T, achieves 68.3%68.3\% mIoU, outperforming PointFlow (66.9%66.9\%) and UperNet with Swin-T (64.6%64.6\%), while the largest variant, SegNeXt-L, attains 70.3%70.3\% mIoU.

  11. Knowl 11 — Limitations of SegNeXt Architecture

    limitation

    The SegNeXt architecture and experimental evaluations present two main limitations identified by the authors:

    1. Model Scaling: The architecture has not been scaled to very large parameter regimes exceeding 100M+ parameters to determine whether purely convolutional attention continues to match or outperform large-scale vision transformers at scale.
    2. Task Breadth: The efficacy of the Multi-Scale Convolutional Attention (MSCA) block has primarily been verified on semantic segmentation and image classification, leaving its generalization to other core vision domains (such as object detection or video processing) and natural language processing (NLP) tasks unexplored.

Coverage note — ImageNet classification pretraining benchmark numbers (Table 3), Pascal Context results (Table 12), and MSCA vs single large kernel comparison (Table 8) were omitted as standalone knowls because their core findings are directly integrated into the MSCA formulation, ablation, and segmentation performance knowls.

References

  1. 1.Badrinarayanan, V., Kendall, A., Cipolla, R.: Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 39(12), 2481–2495 (2017)
  2. 2.Bertasius, G., Shi, J., Torresani, L.: Semantic segmentation with boundary neural fields. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3602–3610 (2016)
  3. 3.Caesar, H., Uijlings, J., Ferrari, V.: Coco-stuff: Thing and stuff classes in context. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1209–1218 (2018)
  4. 4.Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv preprint arXiv:1412.7062 (2014)
  5. 5.Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell. 40(4), 834–848 (2017)
  6. 6.Chen, L.C., Papandreou, G., Schroff, F., Adam, H.: Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587 (2017)
  7. 7.Chen, L.C., Yang, Y., Wang, J., Xu, W., Yuille, A.L.: Attention to scale: Scale-aware semantic image segmentation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3640–3649 (2016)
  8. 8.Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., Adam, H.: Encoder-decoder with atrous separable convolution for semantic image segmentation. In: Eur. Conf. Comput. Vis. pp. 801–818 (2018)
  9. 9.Chen, L., Zhang, H., Xiao, J., Nie, L., Shao, J., Liu, W., Chua, T.S.: Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 5659–5667 (2017)
  10. 10.Cheng, B., Misra, I., Schwing, A.G., Kirillov, A., Girdhar, R.: Masked-attention mask transformer for universal image segmentation (2021)
  11. 11.Cheng, B., Schwing, A.G., Kirillov, A.: Per-pixel classification is not all you need for semantic segmentation. In: NeurIPS (2021)
  12. 12.Contributors, M.: MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark. https://github.com/open-mmlab/mmsegmentation(Apache-2.0) (2020)
  13. 13.Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., Schiele, B.: The cityscapes dataset for semantic urban scene understanding. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3213–3223 (2016)
  14. 14.Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., Wei, Y.: Deformable convolutional networks. In: Int. Conf. Comput. Vis. pp. 764–773 (2017)
  15. 15.Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 248–255. Ieee (2009)
  16. 16.Ding, H., Jiang, X., Liu, A.Q., Thalmann, N.M., Wang, G.: Boundary-aware feature propagation for scene segmentation. In: Int. Conf. Comput. Vis. pp. 6819–6829 (2019)
  17. 17.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. In: Int. Conf. Learn. Represent. (2020)
  18. 18.Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The pascal visual object classes (voc) challenge. Int. J. Comput. Vis. 88(2), 303–338 (2010)
  19. 19.Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., Lu, H.: Dual attention network for scene segmentation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3146–3154 (2019)
  20. 20.Gao, S.H., Cheng, M.M., Zhao, K., Zhang, X.Y., Yang, M.H., Torr, P.: Res2net: A new multi-scale backbone architecture. IEEE Trans. Pattern Anal. Mach. Intell. 43(2), 652–662 (2021)
  21. 21.Geng, Z., Guo, M.H., Chen, H., Li, X., Wei, K., Lin, Z.: Is attention better than matrix decomposition? In: Int. Conf. Learn. Represent. (2021)
  22. 22.Guo, M.H., Cai, J.X., Liu, Z.N., Mu, T.J., Martin, R.R., Hu, S.M.: Pct: Point cloud transformer. Computational Visual Media 7(2), 187–199 (2021)
  23. 23.Guo, M.H., Liu, Z.N., Mu, T.J., Hu, S.M.: Beyond self-attention: External attention using two linear layers for visual tasks. arXiv preprint arXiv:2105.02358 (2021)
  24. 24.Guo, M.H., Lu, C.Z., Liu, Z.N., Cheng, M.M., Hu, S.M.: Visual attention network. arXiv preprint arXiv:2202.09741 (2022)
  25. 25.Guo, M.H., Xu, T.X., Liu, J.J., Liu, Z.N., Jiang, P.T., Mu, T.J., Zhang, S.H., Martin, R.R., Cheng, M.M., Hu, S.M.: Attention mechanisms in computer vision: A survey. arXiv preprint arXiv:2111.07624 (2021)
  26. 26.He, J., Deng, Z., Zhou, L., Wang, Y., Qiao, Y.: Adaptive pyramid context network for semantic segmentation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 7519–7528 (2019)
  27. 27.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 770–778 (2016)
  28. 28.He, T., Zhang, Z., Zhang, H., Zhang, Z., Xie, J., Li, M.: Bag of tricks for image classification with convolutional neural networks. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 558–567 (2019)
  29. 29.Hou, Q., Zhang, L., Cheng, M.M., Feng, J.: Strip pooling: Rethinking spatial pooling for scene parsing. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4003–4012 (2020)
  30. 30.Hu, J., Shen, L., Sun, G.: Squeeze-and-excitation networks. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 7132–7141 (2018)
  31. 31.Hu, S.M., Liang, D., Yang, G.Y., Yang, G.W., Zhou, W.Y.: Jittor: a novel deep learning framework with meta-operators and unified graph execution. Science China Information Sciences 63(12), 1–21 (2020)
  32. 32.Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K.Q.: Densely connected convolutional networks. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 4700–4708 (2017)
  33. 33.Huang, Z., Shi, X., Zhang, C., Wang, Q., Cheung, K.C., Qin, H., Dai, J., Li, H.: Flowformer: A transformer architecture for optical flow. arXiv preprint arXiv:2203.16194 (2022)
  34. 34.Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., Liu, W.: Ccnet: Criss-cross attention for semantic segmentation. In: Int. Conf. Comput. Vis. pp. 603–612 (2019)
  35. 35.Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: Int. Conf. Mach. Learn. pp. 448–456. PMLR (2015)
  36. 36.Kirillov, A., Girshick, R., He, K., Dollár, P.: Panoptic feature pyramid networks. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 6399–6408 (2019)
  37. 37.Kirillov, A., Wu, Y., He, K., Girshick, R.: Pointrend: Image segmentation as rendering. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 9799–9808 (2020)
  38. 38.Lee, Y., Kim, J., Willette, J., Hwang, S.J.: Mpvit: Multi-path vision transformer for dense prediction. In: IEEE Conf. Comput. Vis. Pattern Recog. (2022)
  39. 39.Li, H., Xiong, P., Fan, H., Sun, J.: Dfanet: Deep feature aggregation for real-time semantic segmentation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 9522–9531 (2019)
  40. 40.Li, X., Zhong, Z., Wu, J., Yang, Y., Lin, Z., Liu, H.: Expectation-maximization attention networks for semantic segmentation. In: Int. Conf. Comput. Vis. pp. 9167–9176 (2019)
  41. 41.Li, X., He, H., Li, X., Li, D., Cheng, G., Shi, J., Weng, L., Tong, Y., Lin, Z.: Pointflow: Flowing semantics through points for aerial image segmentation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 4217–4226 (2021)
  42. 42.Li, X., Li, X., Zhang, L., Cheng, G., Shi, J., Lin, Z., Tan, S., Tong, Y.: Improving semantic segmentation via decoupled body and edge supervision. In: European Conference on Computer Vision. pp. 435–452. Springer (2020)
  43. 43.Li, X., You, A., Zhu, Z., Zhao, H., Yang, M., Yang, K., Tan, S., Tong, Y.: Semantic flow for fast and accurate scene parsing. In: European Conference on Computer Vision. pp. 775–793. Springer (2020)
  44. 44.Li, X., Zhang, W., Pang, J., Chen, K., Cheng, G., Tong, Y., Loy, C.C.: Video k-net: A simple, strong, and unified baseline for video segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18847–18857 (2022)
  45. 45.Li, X., Zhao, H., Han, L., Tong, Y., Tan, S., Yang, K.: Gated fully fusion for semantic segmentation. In: Proceedings of the AAAI conference on artificial intelligence. pp. 11418–11425 (2020)
  46. 46.Li, X., Zhou, Y., Pan, Z., Feng, J.: Partial order pruning: for best speed/accuracy trade-off in neural architecture search. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 9145–9153 (2019)
  47. 47.Lin, G., Milan, A., Shen, C., Reid, I.: Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1925–1934 (2017)
  48. 48.Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: Eur. Conf. Comput. Vis. pp. 740–755. Springer (2014)
  49. 49.Liu, R., Deng, H., Huang, Y., Shi, X., Lu, L., Sun, W., Wang, X., Dai, J., Li, H.: Decoupled spatial-temporal transformer for video inpainting. arXiv preprint arXiv:2104.06637 (2021)
  50. 50.Liu, R., Deng, H., Huang, Y., Shi, X., Lu, L., Sun, W., Wang, X., Dai, J., Li, H.: Fuseformer: Fusing fine-grained information in transformers for video inpainting. In: Int. Conf. Comput. Vis. pp. 14040–14049 (2021)
  51. 51.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. In: Int. Conf. Comput. Vis. (2021)
  52. 52.Liu, Z., Mao, H., Wu, C.Y., Feichtenhofer, C., Darrell, T., Xie, S.: A convnet for the 2020s (2022)
  53. 53.Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3431–3440 (2015)
  54. 54.Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017)
  55. 55.Mehta, S., Rastegari, M., Caspi, A., Shapiro, L., Hajishirzi, H.: Espnet: Efficient spatial pyramid of dilated convolutions for semantic segmentation. In: Proceedings of the european conference on computer vision (ECCV). pp. 552–568 (2018)
  56. 56.Mehta, S., Rastegari, M., Shapiro, L., Hajishirzi, H.: Espnetv2: A light-weight, power efficient, and general purpose convolutional neural network. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 9190–9200 (2019)
  57. 57.Mnih, V., Heess, N., Graves, A., et al.: Recurrent models of visual attention. In: Adv. Neural Inform. Process. Syst. pp. 2204–2212 (2014)
  58. 58.Mottaghi, R., Chen, X., Liu, X., Cho, N.G., Lee, S.W., Fidler, S., Urtasun, R., Yuille, A.: The role of context for object detection and semantic segmentation in the wild. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 891–898 (2014)
  59. 59.Mou, L., Hua, Y., Zhu, X.X.: A relation-augmented fully convolutional network for semantic segmentation in aerial scenes. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 12416–12425 (2019)
  60. 60.Orsic, M., Kreso, I., Bevandic, P., Segvic, S.: In defense of pre-trained imagenet architectures for real-time semantic segmentation of road-driving images. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 12607–12616 (2019)
  61. 61.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al.: Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019)
  62. 62.Peng, C., Zhang, X., Yu, G., Luo, G., Sun, J.: Large kernel matters–improve semantic segmentation by global convolutional network. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 4353–4361 (2017)
  63. 63.Ranftl, R., Bochkovskiy, A., Koltun, V.: Vision transformers for dense prediction. In: Int. Conf. Comput. Vis. pp. 12179–12188 (2021)
  64. 64.Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedical image segmentation. In: International Conference on Medical image computing and computer-assisted intervention. pp. 234–241. Springer (2015)
  65. 65.Strudel, R., Garcia, R., Laptev, I., Schmid, C.: Segmenter: Transformer for semantic segmentation. In: Int. Conf. Comput. Vis. pp. 7262–7272 (2021)
  66. 66.Sun, C., Shrivastava, A., Singh, S., Gupta, A.: Revisiting unreasonable effectiveness of data in deep learning era. In: Proceedings of the IEEE international conference on computer vision. pp. 843–852 (2017)
  67. 67.Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: Going deeper with convolutions. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1–9 (2015)
  68. 68.Takikawa, T., Acuna, D., Jampani, V., Fidler, S.: Gated-scnn: Gated shape cnns for semantic segmentation. In: Int. Conf. Comput. Vis. pp. 5229–5238 (2019)
  69. 69.Tan, M., Le, Q.: Efficientnet: Rethinking model scaling for convolutional neural networks. In: International conference on machine learning. pp. 6105–6114. PMLR (2019)
  70. 70.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., Jégou, H.: Training data-efficient image transformers & distillation through attention. In: Int. Conf. Mach. Learn. pp. 10347–10357. PMLR (2021)
  71. 71.Wang, J., Sun, K., Cheng, T., Jiang, B., Deng, C., Zhao, Y., Liu, D., Mu, Y., Tan, M., Wang, X., et al.: Deep high-resolution representation learning for visual recognition. IEEE Trans. Pattern Anal. Mach. Intell. (2020)
  72. 72.Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., Hu, Q.: Eca-net: Efficient channel attention for deep convolutional neural networks (2020)
  73. 73.Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., Shao, L.: Pvtv2: Improved baselines with pyramid vision transformer. arXiv preprint arXiv:2106.13797 (2021)
  74. 74.Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., Shao, L.: Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In: Int. Conf. Comput. Vis. (2021)
  75. 75.Wang, X., Girshick, R., Gupta, A., He, K.: Non-local neural networks. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 7794–7803 (2018)
  76. 76.Waqas Zamir, S., Arora, A., Gupta, A., Khan, S., Sun, G., Shahbaz Khan, F., Zhu, F., Shao, L., Xia, G.S., Bai, X.: isaid: A large-scale dataset for instance segmentation in aerial images. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. pp. 28–37 (2019)
  77. 77.Wightman, R.: Pytorch image models. https://github.com/rwightman/pytorch-image-models(Apache-2.0) (2019). https://doi.org/10.5281/zenodo.4414861
  78. 78.Xia, F., Wang, P., Chen, L.C., Yuille, A.L.: Zoom better to see clearer: Human and object parsing with hierarchical auto-zoom net. In: European Conference on Computer Vision. pp. 648–663. Springer (2016)
  79. 79.Xiao, T., Liu, Y., Zhou, B., Jiang, Y., Sun, J.: Unified perceptual parsing for scene understanding. In: Eur. Conf. Comput. Vis. pp. 418–434 (2018)
  80. 80.Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J.M., Luo, P.: Segformer: Simple and efficient design for semantic segmentation with transformers. Adv. Neural Inform. Process. Syst. 34 (2021)
  81. 81.Xie, S., Girshick, R., Dollár, P., Tu, Z., He, K.: Aggregated residual transformations for deep neural networks. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1492–1500 (2017)
  82. 82.Yang, J., Li, C., Zhang, P., Dai, X., Xiao, B., Yuan, L., Gao, J.: Focal self-attention for local-global interactions in vision transformers (2021)
  83. 83.Yang, M., Yu, K., Zhang, C., Li, Z., Yang, K.: Denseaspp for semantic segmentation in street scenes. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3684–3692 (2018)
  84. 84.Yu, C., Gao, C., Wang, J., Yu, G., Shen, C., Sang, N.: Bisenet v2: Bilateral network with guided aggregation for real-time semantic segmentation. Int. J. Comput. Vis. 129(11), 3051–3068 (2021)
  85. 85.Yu, C., Wang, J., Peng, C., Gao, C., Yu, G., Sang, N.: Bisenet: Bilateral segmentation network for real-time semantic segmentation. In: Eur. Conf. Comput. Vis. pp. 325–341 (2018)
  86. 86.Yu, F., Koltun, V.: Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122 (2015)
  87. 87.Yuan, Y., Chen, X., Wang, J.: Object-contextual representations for semantic segmentation. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16. pp. 173–190. Springer (2020)
  88. 88.Yuan, Y., Fu, R., Huang, L., Lin, W., Zhang, C., Chen, X., Wang, J.: Hrformer: High-resolution vision transformer for dense predict. Adv. Neural Inform. Process. Syst. 34 (2021)
  89. 89.Yuan, Y., Huang, L., Guo, J., Zhang, C., Chen, X., Wang, J.: Ocnet: Object context network for scene parsing. arXiv preprint arXiv:1809.00916 (2018)
  90. 90.Yuan, Y., Xie, J., Chen, X., Wang, J.: Segfix: Model-agnostic boundary refinement for segmentation. In: European Conference on Computer Vision. pp. 489–506. Springer (2020)
  91. 91.Zhang, H., Dana, K., Shi, J., Zhang, Z., Wang, X., Tyagi, A., Agrawal, A.: Context encoding for semantic segmentation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 7151–7160 (2018)
  92. 92.Zhang, H., Wu, C., Zhang, Z., Zhu, Y., Lin, H., Zhang, Z., Sun, Y., He, T., Mueller, J., Manmatha, R., et al.: Resnest: Split-attention networks. arXiv preprint arXiv:2004.08955 (2020)
  93. 93.Zhao, H., Qi, X., Shen, X., Shi, J., Jia, J.: Icnet for real-time semantic segmentation on high-resolution images. In: Eur. Conf. Comput. Vis. pp. 405–420 (2018)
  94. 94.Zhao, H., Shi, J., Qi, X., Wang, X., Jia, J.: Pyramid scene parsing network. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 2881–2890 (2017)
  95. 95.Zhen, M., Wang, J., Zhou, L., Li, S., Shen, T., Shang, J., Fang, T., Quan, L.: Joint semantic segmentation and boundary detection using iterative pyramid contexts. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 13666–13675 (2020)
  96. 96.Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., Torr, P.H., et al.: Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 6881–6890 (2021)
  97. 97.Zheng, Z., Zhong, Y., Wang, J., Ma, A.: Foreground-aware relation network for geospatial object segmentation in high spatial resolution remote sensing imagery. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 4096–4105 (2020)
  98. 98.Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A.: Scene parsing through ade20k dataset. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 633–641 (2017)
  99. 99.Zoph, B., Ghiasi, G., Lin, T.Y., Cui, Y., Liu, H., Cubuk, E.D., Le, Q.: Rethinking pre-training and self-training. Adv. Neural Inform. Process. Syst. 33, 3833–3845 (2020)

Citation

MLA
Guo, M.-H., et al. “SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation”. arXiv, 2022, http://arxiv.org/abs/2209.08575v1.
APA
Guo, M.-H., Lu, C.-Z., Hou, Q., Liu, Z., Cheng, M.-M., & Hu, S.-M. (2022). SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation. arXiv. http://arxiv.org/abs/2209.08575v1
Chicago
Guo, M.-H., C.-Z. Lu, Q. Hou, Z. Liu, M.-M. Cheng, and S.-M. Hu. 2022. “SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation”. arXiv. http://arxiv.org/abs/2209.08575v1.
Harvard
Guo, M.-H. et al. (2022) “SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2209.08575v1.
Vancouver
1. Guo M-H, Lu C-Z, Hou Q, Liu Z, Cheng M-M, Hu S-M (2022) SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation. arXiv

BibTeX

@article{guo2022segnext,
  title = {SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation},
  author = {Guo, Meng-Hao and Lu, Cheng-Ze and Hou, Qibin and Liu, Zhengning and Cheng, Ming-Ming and Hu, Shi-Min},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2209.08575v1},
  eprint = {2209.08575}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors