Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation

Jiaqi GuHyoukjun KwonDilin WangWei YeMeng LiYu-Hsin ChenLiangzhen LaiVikas ChandraDavid Z. Pan

article2022CVPR234 citations

Presents HRViT, a multi-branch vision transformer backbone for semantic segmentation that combines high-resolution feature representations with efficient attention and heterogeneous branch designs to outperform state-of-the-art models on ADE20K and Cityscapes with fewer parameters and FLOPs.

Listen

Dense prediction tasks such as semantic segmentation, which involves classifying every pixel in an image, are vital for modern visual platforms including augmented and virtual reality devices. While Vision Transformers offer strong representational power through attention mechanisms, conventional designs output low-resolution, single-scale features. Existing adaptations typically rely on sequential downsampling architectures that discard fine-grained spatial details and lack sufficient cross-scale interaction, while convolutional high-resolution networks remain limited by small receptive fields.

The article introduces and evaluates HRViT, a novel multi-scale, high-resolution Vision Transformer backbone engineered specifically for semantic segmentation. The objective is to demonstrate that combining parallel high-resolution branches with co-optimized Transformer building blocks can significantly improve segmentation accuracy while reducing computational and hardware costs.

To achieve this, the authors designed a four-stage, multi-branch network that maintains high-resolution representations throughout processing and repeatedly exchanges information across scales. To overcome the prohibitive computational overhead of combining multi-branch topologies with self-attention, the authors introduced an augmented cross-shaped local self-attention mechanism with shared projection matrices, mixed-scale convolutional feedforward networks, lightweight patch embeddings, and a heterogeneous branch allocation strategy that concentrates model depth on medium-resolution paths. The models were pretrained on the standard ImageNet-1K benchmark and comprehensively evaluated against state-of-the-art vision models on the ADE20K and Cityscapes segmentation datasets.

The evaluation produced several key findings: First, HRViT establishes a superior trade-off between performance and efficiency, achieving a mean Intersection over Union of 50.20% on ADE20K and 83.16% on Cityscapes. Second, compared to leading vision backbones CSWin and MiT, HRViT improves segmentation accuracy by an average of 1.78 to 2.16 percentage points while requiring 28% to 30.7% fewer parameters and 21% to 23.1% fewer computation operations. Third, the benefits are particularly pronounced on compact model variants, where the parallel high-resolution design effectively expands representational capacity under strict resource limits. Fourth, ablation experiments confirmed that each co-optimization component—including key-value sharing, dense cross-scale fusion, and auxiliary convolution paths—directly contributed to accuracy gains without incurring material latency penalties.

These findings demonstrate that brute-force integration of high-resolution structures into Vision Transformers is inefficient, but disciplined branch-block co-optimization resolves the scalability barrier. For decision-makers and engineering teams developing vision systems, adopting HRViT provides higher segmentation precision with significantly reduced memory footprint and compute costs. This efficiency directly translates to improved runtime performance, lower hardware requirements, and decreased energy consumption on edge and resource-constrained devices.

Organizations developing dense computer vision applications should consider adopting HRViT as a drop-in backbone for existing segmentation pipelines, particularly when deploying on resource-limited hardware. When implementing the architecture, practitioners should prioritize heterogeneous branch configurations and balanced attention window sizes, as excessively large attention windows add computational overhead without improving accuracy. The primary limitation of the study is its primary focus on semantic segmentation; further evaluation on broader dense prediction tasks, such as object detection and instance segmentation, is recommended to confirm generalizability across all visual recognition workloads.

Cover for Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation

Abstract

Vision Transformers (ViTs) have emerged with superior performance on computer vision tasks compared to the convolutional neural network (CNN)-based models. However, ViTs mainly designed for image classification will generate single-scale low-resolution representations, which makes dense prediction tasks such as semantic segmentation challenging for ViTs. Therefore, we propose HRViT, which enhances ViTs to learn semantically-rich and spatially-precise multi-scale representations by integrating high-resolution multi-branch architectures with ViTs. We balance the model performance and efficiency of HRViT by various branch-block co-optimization techniques. Specifically, we explore heterogeneous branch designs, reduce the redundancy in linear layers, and augment the attention block with enhanced expressiveness. Those approaches enabled HRViT to push the Pareto frontier of performance and efficiency on semantic segmentation to a new level, as our evaluation results on ADE20K and Cityscapes show. HRViT achieves 50.20% mIoU on ADE20K and 83.16% mIoU on Cityscapes, surpassing state-of-the-art MiT and CSWin backbones with an average of +1.78 mIoU improvement, 28% parameter saving, and 21% FLOPs reduction, demonstrating the potential of HRViT as a strong vision backbone for semantic segmentation. Our code is publicly available 1.

Table of Contents

  • 1. Introduction
  • 2. Proposed HRViT Architecture
  • 2.1. Architecture overview
  • 2.2. Efficient HRViT component design
  • 2.3. Heterogeneous HRViT branch design
  • 2.4. Architectural variants
  • 3. Experiments
  • 3.1. Semantic segmentation on Cityscapes and ADE20K
  • 3.2. Ablation studies
  • 4. Related Work
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — HRViT multi-branch high-resolution backbone

    model/method

    HRViT is a pure Vision Transformer backbone designed for semantic segmentation by combining a high-resolution multi-branch topology with Transformer blocks. A two-layer convolutional stem first reduces an input image by 4×4\times using two stride-2 Conv-BN-ReLU blocks and produces CC-channel features. Four successive Transformer stages then contain respectively one, two, three, and four parallel branches, whose spatial resolutions are maintained at approximately 1/41/4, 1/81/8, 1/161/16, and 1/321/32 of the input resolution rather than being collapsed into a single low-resolution stream.

    Each stage is composed of one or more modules. Every module begins with cross-resolution fusion and an efficient patch-embedding block on each branch, followed by repeated augmented local self-attention blocks (HRViTAttn) and mixed-scale convolutional feedforward networks (MixCFN). High- and low-resolution branches interact repeatedly throughout the network, so the highest-resolution branch receives semantic information from deeper branches while lower-resolution branches receive fine spatial information from higher-resolution branches.

  2. Knowl 2 — Augmented cross-shaped local self-attention

    equation

    For an input feature map x∈RH×W×Cx\in\mathbb{R}^{H\times W\times C}, HRViTAttn splits the channels into two equal parts, xHx_H and xVx_V, and partitions them into horizontal windows of size s×Ws\times W and vertical windows of size H×sH\times s, respectively. The two orientations provide local fine-grained aggregation while allowing information to propagate across both spatial axes. With KK attention heads, head dimension dkd_k, output projection WO∈RC×CW^O\in\mathbb{R}^{C\times C}, query/key/value projections WkQ,WkK,WkV∈Rdk×CW_k^Q,W_k^K,W_k^V\in\mathbb{R}^{d_k\times C}, Hardswish activation σ\sigma, and depth-wise convolution DWConv⁡\operatorname{DWConv}, the block is

    HRViTAttn⁡(x)=BN⁡(σ(WO[y1,…,yK])),yk=zk+DWConv⁡(σ(WkVx)),zk={H-Attn⁡k(x),1≤k<K/2,V-Attn⁡k(x),K/2≤k≤K,zkm=MHSA⁡(WkQxm,WkKxm,WkVxm),\begin{aligned} \operatorname{HRViTAttn}(x)&=\operatorname{BN}\left(\sigma\left(W^O[y_1,\ldots,y_K]\right)\right),\\ y_k&=z_k+\operatorname{DWConv}\left(\sigma\left(W_k^Vx\right)\right),\\ z_k&= \begin{cases} \operatorname{H\text{-}Attn}_k(x), & 1\leq k<K/2,\\ \operatorname{V\text{-}Attn}_k(x), & K/2\leq k\leq K, \end{cases}\\ z_k^m&=\operatorname{MHSA}\left(W_k^Qx^m,W_k^Kx^m,W_k^Vx^m\right), \end{aligned}

    where xmx^m is the mm-th horizontal or vertical window and [y1,…,yK][y_1,\ldots,y_K] denotes channel concatenation. The parallel depth-wise-convolution path operates on the complete four-dimensional feature map before window partitioning, adding local inductive bias without requiring a separate positional-encoding path. If HH or WW is not divisible by ss, zero-padding completes the final window and the padded attention logits are masked before softmax. Fixing one window dimension to ss avoids the quadratic dependence on the full image dimensions while retaining fine spatial detail and an approximate global view through the two orthogonal attention orientations.

  3. Knowl 3 — Key-value sharing in HRViTAttn

    model/method

    HRViT reduces the parameter and computation cost of local self-attention by using the same linear projection for keys and values. For a window xmx^m, the conventional three projections are replaced by one query projection WkQW_k^Q and one shared key/value projection WkVW_k^V. If Qkm=WkQxmQ_k^m=W_k^Qx^m and Vkm=WkVxmV_k^m=W_k^Vx^m, the attention operation is

    MHSA⁡(WkQxm,WkVxm,WkVxm)=softmax⁡(Qkm(Vkm)Tdk)Vkm,\operatorname{MHSA}\left(W_k^Qx^m,W_k^Vx^m,W_k^Vx^m\right)=\operatorname{softmax}\left(\frac{Q_k^m(V_k^m)^T}{\sqrt{d_k}}\right)V_k^m,

    where dkd_k is the dimensionality of attention head kk. Sharing the key and value projection removes one learned linear projection per head; the paper compensates for the resulting expressivity reduction with an additional Hardswish nonlinearity and a BatchNorm layer in HRViTAttn. The BatchNorm layer is initialized as an identity transformation to stabilize training.

  4. Knowl 4 — Diversity-enhanced shortcut with Kronecker projection

    equation

    HRViTAttn includes a diversity-enhanced shortcut (DES), an auxiliary channel-mixing path that applies a nonlinear learned projection without hardware-unfriendly Fourier transforms. Let x∈RH×W×Cx\in\mathbb{R}^{H\times W\times C}, let P∈RC×CP\in\mathbb{R}^{C\times C} be the desired channel projector, and approximate it by P=A⊗BP=A\otimes B, where A,B∈RC×CA,B\in\mathbb{R}^{\sqrt{C}\times\sqrt{C}}. Reshape xx into x~∈RHW×C×C\tilde{x}\in\mathbb{R}^{HW\times\sqrt{C}\times\sqrt{C}}. The shortcut is

    DES⁡(x)=A⋅Hardswish⁡(x~BT).\operatorname{DES}(x)=A\cdot\operatorname{Hardswish}\left(\tilde{x}B^T\right).

    The factorized form reduces the cost of applying a C×CC\times C channel projection, while the Hardswish between the two factors increases the shortcut's nonlinearity. DES is added as an auxiliary feature path and has negligible measured hardware overhead because of the factorization.

  5. Knowl 5 — Mixed-scale convolutional feedforward network

    model/method

    HRViT replaces the standard Transformer feedforward network with MixCFN, which extracts local information at multiple spatial scales. For a branch feature map, LayerNorm is followed by a linear expansion from CC channels to rCrC channels, where rr is the expansion ratio. The expanded channels are split into two paths containing depth-wise convolutions with kernel sizes 3×33\times3 and 5×55\times5. Their outputs are concatenated and passed through a GELU activation and a second linear projection back to the branch width.

    The two convolutional scales provide complementary local receptive fields while preserving the lightweight depth-wise structure. The paper uses larger expansion ratios for small HRViT variants and reduces rr to 22 or 33 for medium and large variants, exploiting channel redundancy with only a marginal reported performance loss.

  6. Knowl 6 — Efficient stem, patch embedding, and dense cross-resolution fusion

    model/method

    HRViT uses early convolutions rather than self-attention to process high-resolution input, reducing the input spatial size by 4×4\times while retaining low-level detail. Before Transformer blocks in every branch and module, its efficient patch embedding replaces a conventional convolution-plus-normalization layer with a pointwise convolution followed by a depth-wise convolution and LayerNorm:

    EffPatchEmbed⁡(x)=LN⁡(DWConv⁡(PWConv⁡(x))).\operatorname{EffPatchEmbed}(x)=\operatorname{LN}\left(\operatorname{DWConv}\left(\operatorname{PWConv}(x)\right)\right).

    Here xx is a branch feature map, PWConv⁡\operatorname{PWConv} is a 1×11\times1 pointwise convolution used for channel matching, DWConv⁡\operatorname{DWConv} is a depth-wise convolution, and LN⁡\operatorname{LN} is LayerNorm.

    A dense cross-resolution fusion layer is inserted at the beginning of every module. Branches are indexed from high resolution to low resolution. To send features from branch ii to a lower-resolution branch j>ij>i, HRViT uses a depth-wise separable convolution with stride 2j−i2^{j-i} and kernel size 2j−i+12^{j-i}+1, followed by channel matching. To send features from branch ii to a higher-resolution branch j<ij<i, it first uses a pointwise convolution to match channels and then nearest-neighbor upsampling by 2i−j2^{i-j}. Features with i=ji=j use a skip connection. Thus, low-resolution branches receive detailed high-resolution information and high-resolution branches receive larger-receptive-field semantic information. The paper retains dense fusion rather than a sparse neighboring-branch fusion because sparse fusion saves less than 1% hardware cost but causes measurable accuracy degradation.

  7. Knowl 7 — Heterogeneous branch allocation and complexity-guided design

    theoretical result

    HRViT assigns different attention window sizes and Transformer depths to its branches instead of replicating the same block configuration at every resolution. Let branch i∈{1,2,3,4}i\in\{1,2,3,4\} have spatial size H/2i−1×W/2i−1H/2^{i-1}\times W/2^{i-1}, channel width 2i−1C2^{i-1}C, attention window size sis_i, and MixCFN expansion ratio rir_i, where H,WH,W and CC denote the spatial dimensions and channel width of the highest-resolution branch. The asymptotic parameter and FLOP costs are

    Params⁡HRViTAttn⁡,i=O(4i−1C2+2i−1C),Params⁡MixCFN⁡,i=O(4i−1C2ri+2i−1Cri),FLOPs⁡HRViTAttn⁡,i=O(HWC2+CHW(H+W)si4i−1),FLOPs⁡MixCFN⁡,i=O(riHWC2+riHWC2i−1).\begin{aligned} \operatorname{Params}_{\operatorname{HRViTAttn},i}&=\mathcal{O}\left(4^{i-1}C^2+2^{i-1}C\right),\\ \operatorname{Params}_{\operatorname{MixCFN},i}&=\mathcal{O}\left(4^{i-1}C^2r_i+2^{i-1}Cr_i\right),\\ \operatorname{FLOPs}_{\operatorname{HRViTAttn},i}&=\mathcal{O}\left(HWC^2+\frac{CHW(H+W)s_i}{4^{i-1}}\right),\\ \operatorname{FLOPs}_{\operatorname{MixCFN},i}&=\mathcal{O}\left(r_iHWC^2+\frac{r_iHWC}{2^{i-1}}\right). \end{aligned}

    The two highest-resolution branches are computationally expensive but parameter-efficient and useful for fine spatial detail, so HRViT gives them narrow windows and few blocks. The medium-resolution third branch has a favorable cost-receptive-field trade-off and receives the deepest Transformer stack with a wide window. The lowest-resolution branch contains many parameters and provides global semantic context, but receives only a few wide-window blocks because its spatial downsampling loses detail. When distributing blocks among modules, HRViT uses nearly even assignments, such as 66-66-66-22 for 20 blocks, rather than highly unbalanced assignments such as 1717-11-11-11, to improve the average ensemble depth and preserve input and gradient flow.

  8. Knowl 8 — Scalable HRViT variants and branch settings

    model/method

    HRViT scales both channel width and depth while preserving the heterogeneous branch pattern. All three reported variants use attention windows (1,2,7,7)(1,2,7,7) from the highest- to lowest-resolution branches. The branch channel widths CC, MixCFN expansion ratios rr, and attention head dimensions dkd_k are:

    • HRViT-b1: C=(32,64,128,256)C=(32,64,128,256), r=(4,4,4,4)r=(4,4,4,4), and dk=(16,32,32,32)d_k=(16,32,32,32).
    • HRViT-b2: C=(48,96,240,384)C=(48,96,240,384), r=(2,3,3,3)r=(2,3,3,3), and dk=(24,24,24,24)d_k=(24,24,24,24).
    • HRViT-b3: C=(64,128,256,512)C=(64,128,256,512), r=(2,2,2,2)r=(2,2,2,2), and dk=(32,32,32,32)d_k=(32,32,32,32).

    The high-resolution branches generally receive 5–6 Transformer blocks, the medium-resolution branch 20–24 blocks, and the low-resolution branch 4–6 blocks. The resulting ImageNet-1K models have 19.7M parameters, 2.7 GFLOPs, and 80.5% top-1 accuracy for b1; 32.5M parameters, 5.1 GFLOPs, and 82.3% for b2; and 37.9M parameters, 5.7 GFLOPs, and 82.8% for b3, measured at 224×224224\times224 input resolution.

  9. Knowl 9 — Segmentation evaluation protocol

    experimental setup

    The paper pretrains HRViT-b1, b2, and b3 on ImageNet-1K using the training settings of DeiT and related Vision Transformers, with stochastic depth and a maximum drop rate of 0.10.1. The drop rate increases along the deepest third branch, while the shallower branches within a module follow the corresponding third-branch rate. ImageNet pretraining uses an HRNetV2 classification head.

    For semantic segmentation, the models are evaluated on ADE20K and Cityscapes using a lightweight SegFormer head in the MMSegmentation framework. Training crops are 512×512512\times512 for ADE20K and 1024×10241024\times1024 for Cityscapes. Test images are 512×2048512\times2048 and 1024×20481024\times2048, respectively; Cityscapes inference uses sliding-window crops of 1024×10241024\times1024. Results are compared with Swin, Twins, MiT, CSWin, and HRNet-based backbones using single-scale mean intersection-over-union (mIoU), parameter count, and FLOPs.

  10. Knowl 10 — Semantic segmentation performance and efficiency

    data/table

    With the SegFormer head, HRViT establishes a favorable performance-efficiency trade-off on both evaluation datasets. On ADE20K validation, the reported mIoUs for matched model groups are MiT-B1 42.20, CSWin-Ti 41.43, and HRViT-b1 45.88; MiT-B2 46.50, CSWin-T 47.88, and HRViT-b2 48.76; and MiT-B3 49.40, CSWin-S 49.93, and HRViT-b3 50.20. HRViT-b1 is 3.68 mIoU points above MiT-B1 while using 40% fewer parameters and 8% fewer FLOPs. HRViT-b3 exceeds CSWin-S while saving 23% of parameters and 13% of FLOPs.

    On Cityscapes validation, the detailed comparison is:

    Could not parse LaTeX table

    HRViT-b1 improves over MiT-B1 and CSWin-Ti by 3.13 and 2.47 mIoU points, respectively. HRViT-b3 improves over MiT-B4 by 0.86 mIoU points while using 55.4% fewer parameters and 30.7% fewer FLOPs. Averaged against the MiT and CSWin comparisons, HRViT reports 2.16 higher mIoU, 30.7% fewer parameters, and 23.1% fewer FLOPs on Cityscapes.

  11. Knowl 11 — Ablation evidence for branch-block co-optimization

    empirical result

    The HRViT-b1 ablation evaluates ImageNet-1K top-1 accuracy and Cityscapes validation mIoU while removing one proposed component at a time. The complete model has 8.1M parameters, 14.1 GFLOPs, 80.52% ImageNet top-1 accuracy, and 81.63% Cityscapes mIoU. The measured variants are:

    Could not parse LaTeX table

    Removing all block optimizations reduces ImageNet accuracy by 0.73 percentage points and Cityscapes mIoU by 1.18 points while increasing parameters by 20% and computation by 13%. Individually, key-value sharing lowers Cityscapes mIoU by 0.63 points and increases parameters and computation; replacing efficient patch embedding increases parameters by 22% and FLOPs by 17% without an accuracy benefit; replacing MixCFN with a standard feedforward network loses approximately 0.66 ImageNet points and 0.11 Cityscapes mIoU; and removing the parallel convolution path loses 0.46 ImageNet points and 0.81 Cityscapes mIoU. Removing the extra Hardswish/BatchNorm path loses 0.15 ImageNet points and 0.51 Cityscapes mIoU. Dense fusion is substantially more accurate than sparse fusion despite the latter's less-than-1% hardware saving.

  12. Knowl 12 — HRViT versus direct HRNet–Transformer substitutions

    empirical result

    The paper compares HRViT with vanilla high-resolution backbones formed by directly replacing HRNetV2 residual blocks with MiT or CSWin Transformer blocks. HRViT's heterogeneous branch allocation and optimized blocks achieve higher accuracy at substantially lower cost:

    Could not parse LaTeX table

    The direct HRNet–Transformer substitutions show that a multi-branch topology can help, but their Transformer cost quickly outweighs the performance gain. Relative to these vanilla baselines, HRViT averages 14.4% fewer parameters and 38.2% fewer FLOPs while improving ImageNet top-1 accuracy by 0.92 points and Cityscapes mIoU by 0.89 points.

Coverage note — Only secondary window-size sensitivity results and the appendix-only alternative UperNet-head evaluation were omitted because they do not introduce a new method or alter the primary conclusions.

References

  1. 1.V. Badrinarayanan, A. Kendall, and R. Cipolla. SegNet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, page 2481–2495, 2017. 1, 8
  2. 2.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-End object detection with transformers. In Proc. ECCV, 2020. 1
  3. 3.Chun-Fu Chen, Quanfu Fan, and Rameswar Panda. CrossViT: Cross-attention multi-scale vision transformer for image classification. In Proc. ICCV, 2021. 1, 8
  4. 4.Liang-Chieh Chen, Yukun Zhu, George Papandreou Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proc. ECCV, 2018. 1, 8
  5. 5.Bowen Cheng, Alexander G. Schwing, and Alexander Kirillov. Per-Pixel Classification is Not All You Need for Semantic Segmentation. In Proc. NeurIPS, 2021. 8
  6. 6.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the Design of Spatial Attention in Vision Transformers. In Proc. NeurIPS, 2021. 1, 4, 5, 8, 12
  7. 7.MMSegmentation Contributors. MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark. https : / / github . com / open -mmlab/mmsegmentation, 2020. 6
  8. 8.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proc. CVPR, 2016. 5, 11
  9. 9.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In CVPR Workshop, 2020. 11
  10. 10.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, , and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proc. CVPR, 2009. 5
  11. 11.Mingyu Ding, Xiaochen Lian, Linjie Yang, Peng Wang, Xiaojie Jin, Zhiwu Lu, and Ping Luo. HR-NAS: Searching Efficient High-Resolution Neural Architectures with Lightweight Transformers. In Proc. CVPR, 2021. 1, 4, 7, 8
  12. 12.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows. arXiv preprint arXiv:2107.00652, 2021. 1, 2, 5, 6, 7, 8, 12
  13. 13.A. Dosovitskiy, L. Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, M. Dehghani, Matthias Minderer, G. Heigold, S. Gelly, Jakob Uszkoreit, and N. Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proc. ICLR, 2021. 1, 7, 8
  14. 14.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale Vision Transformers. In Proc. ICCV, 2021. 8
  15. 15.Ben Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herve J ´ egou, and Matthijs Douze. LeViT: a Vision Transformer in ConvNet’s Clothing for Faster Inference. In Proc. ICCV, 2021. 4
  16. 16.Daniel Haase and Manuel Amthor. Rethinking Depthwise Separable Convolutions: How Intra-Kernel Correlations Lead to Improved MobileNets. In Proc. CVPR, pages 14588–14597, 2020. 4
  17. 17.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In Proc. ECCV, 2016. 6
  18. 18.Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool. LocalViT: Bringing locality to vision transformers. arXiv preprint arXiv:2104.05707, 2021. 1
  19. 19.Iasonas Kokkinos Liang-Chieh Chen, George Papandreou, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected CRFs. In Proc. ICLR, 2015. 1
  20. 20.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proc. ICCV, 2021. 1, 4, 5, 6, 12
  21. 21.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proc. CVPR, 2015. 1, 8
  22. 22.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Proc. ICLR, 2019. 11
  23. 23.A. Newell, K. Yang, and J. Deng. Stacked hourglass networks for human pose estimation. In Proc. ECCV, page 483–499, 2016. 8
  24. 24.Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. Do Vision Transformers See Like Convolutional Neural Networks? arXiv preprint arXiv:2108.08810, 2021. 4
  25. 25.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional networks for biomedical image segmentation. In Proc. MICCAI, May 2015. 1, 8
  26. 26.Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for Semantic Segmentation. In Proc. ICCV, 2021. 8
  27. 27.Yehui Tang, Kai Han, Chang Xu, An Xiao, Yiping Deng, Chao Xu, and Yunhe Wang. Augmented Shortcuts for Vision Transformers. In Proc. NeurIPS, 2021. 4
  28. 28.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers: distillation through attention. In Proc. ICML, pages 10347–10357, 2021. 1, 6, 11
  29. 29.Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, Wenyu Liu, and Bin Xiao. Deep High-Resolution Representation Learning for Visual Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(10):3349–3364, 2021. 1, 4, 6, 8
  30. 30.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions. In Proc. ICCV, 2021. 1, 2, 8
  31. 31.Wenxiao Wang, Lu Yao, Long Chen, Deng Cai, Xiaofei He, and Wei Liu. CrossFormer: A Versatile Vision Transformer Based on Cross-scale Attention. arXiv preprint arXiv:2108.00154, 2021. 1
  32. 32.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, , and Huaxia Xia. End-to-end video instance segmentation with transformers. In Proc. CVPR, 2021. 1
  33. 33.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In Proc. ECCV, page 418–434, 2018. 6, 11, 12
  34. 34.Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollar, and Ross Girshick. Early Convolutions Help Transformers See Better. In Proc. NeurIPS, 2021. 4
  35. 35.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers. In Proc. NeurIPS, 2021. 1, 2, 4, 5, 6, 7, 8, 12
  36. 36.Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu. Co-scale conv-attentional image transformers. In Proc. ICCV, 2021. 1
  37. 37.Changqian Yu, Bin Xiao, Changxin Gao, Lu Yuan, Lei Zhang, Nong Sang, and Jingdong Wang. Lite-HRNet: A Lightweight High-Resolution Network. In Proc. CVPR, 2021. 1, 4, 8
  38. 38.Qihang Yu, Yingda Xia, Yutong Bai, Yongyi Lu, Alan Yuille, and Wei Shen. Glance-and-Gaze Vision Transformer. arXiv preprint arXiv:2106.02277, 2021. 1
  39. 39.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token ViT: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986, 2021. 1
  40. 40.Yuhui Yuan, Rao Fu, Lang Huang, Weihong Lin, Chao Zhang, Xilin Chen, and Jingdong Wang. HRFormer: High-Resolution Transformer for Dense Prediction. In Proc. NeurIPS, 2021. 8
  41. 41.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proc. ICCV, 2019. 11
  42. 42.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, , and David Lopez-Paz. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017. 11
  43. 43.Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao. Multi-scale vision longformer: A new vision transformer for high-resolution image encoding. In Proc. ICCV, 2021. 1
  44. 44.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In Proc. AAAI, 2020. 11
  45. 45.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proc. CVPR, 2017. 5, 11

Citation

MLA
Gu, J., et al. “Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation”. arXiv, 2021, http://arxiv.org/abs/2111.01236v2.
APA
Gu, J., Kwon, H., Wang, D., Ye, W., Li, M., Chen, Y.-H., Lai, L., Chandra, V., & Pan, D. Z. (2021). Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation. arXiv. http://arxiv.org/abs/2111.01236v2
Chicago
Gu, J., H. Kwon, D. Wang, et al. 2021. “Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation”. arXiv. http://arxiv.org/abs/2111.01236v2.
Harvard
Gu, J. et al. (2021) “Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.01236v2.
Vancouver
1. Gu J, Kwon H, Wang D, Ye W, Li M, Chen Y-H, Lai L, Chandra V, Pan DZ (2021) Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation. arXiv

BibTeX

@article{gu2021multi,
  title = {Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation},
  author = {Gu, Jiaqi and Kwon, Hyoukjun and Wang, Dilin and Ye, Wei and Li, Meng and Chen, Yu-Hsin and Lai, Liangzhen and Chandra, Vikas and Pan, David Z.},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.01236v2},
  eprint = {2111.01236}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE