Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation

Hu CaoYueyue WangJoy ChenDongsheng JiangXiaopeng ZhangQi TianManning Wang

article2021ECCV5,703 citations

Introduces Swin-Unet, a pure Transformer U-shaped encoder-decoder network with skip connections that learns long-range contextual interactions to outperform conventional convolutional and hybrid baselines on multi-organ and cardiac image segmentation.

Listen

Accurate medical image segmentation is critical for clinical decision-making, computer-assisted surgery, and automated diagnosis. Current standard solutions rely on convolutional neural networks, which process images through localized operations. However, these local operations struggle to capture broad context and long-range relationships across organs and tissues, frequently leading to over-segmentation errors and imperfect boundary delineations.

The article demonstrates that a pure Transformer architecture structured in a standard encoder-decoder format can perform medical image segmentation without relying on traditional convolution operations. To evaluate this approach, the researchers introduced Swin-Unet and tested its performance against established convolutional and hybrid models on standard abdominal computed tomography and cardiac magnetic resonance imaging benchmarks.

The evaluated model processes image patches using shifted attention windows and reconstructs resolution using dedicated patch expanding layers. On multi-organ computed tomography scans, Swin-Unet achieved the highest overall overlap accuracy at 79.13% and substantially reduced boundary errors, lowering the Hausdorff distance metric to 21.55an improvement of roughly 10 points over leading hybrid models. On cardiac magnetic resonance imaging data, it reached an overall accuracy score of 90.00%. Ablation testing confirmed that incorporating multi-scale skip connections and custom patch expanding layers provided superior boundary precision compared to standard interpolation or convolution-based up-sampling.

These findings indicate that pure attention-based networks can capture global anatomical context more effectively than convolutional baselines, significantly improving the precision of organ boundaries. For clinical workflows, enhanced boundary definition reduces the risk of misidentifying healthy tissue as pathology during surgical planning or automated assessment. Furthermore, the model achieved optimal results using a compact configuration, showing that larger network sizes do not yield meaningful performance gains but add substantial computational expense.

Organizations developing medical imaging systems should consider adopting attention-based encoder-decoder frameworks, specifically using compact model configurations to maintain computational efficiency. Practical implementation should prioritize multi-scale feature connections and dedicated patch expansion rather than basic interpolation. Next steps include developing end-to-end medical pre-training strategies to replace general image pre-initialization, as well as extending the current two-dimensional architecture to directly handle three-dimensional volumetric scans.

Confidence in the reported improvements is strong across the evaluated two-dimensional benchmarks. However, stakeholders should note that the current approach relies on pre-trained weights from natural image datasets and operates solely on two-dimensional image slices, whereas many clinical workflows depend on full three-dimensional volumetric analysis.

Cover for Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation

Abstract

In the past few years, convolutional neural networks (CNNs) have achieved milestones in medical image analysis. Especially, the deep neural networks based on U-shaped architecture and skip-connections have been widely applied in a variety of medical image tasks. However, although CNN has achieved excellent performance, it cannot learn global and long-range semantic information interaction well due to the locality of the convolution operation. In this paper, we propose Swin-Unet, which is an Unet-like pure Transformer for medical image segmentation. The tokenized image patches are fed into the Transformer-based U-shaped Encoder-Decoder architecture with skip-connections for local-global semantic feature learning. Specifically, we use hierarchical Swin Transformer with shifted windows as the encoder to extract context features. And a symmetric Swin Transformer-based decoder with patch expanding layer is designed to perform the up-sampling operation to restore the spatial resolution of the feature maps. Under the direct down-sampling and up-sampling of the inputs and outputs by 4x, experiments on multi-organ and cardiac segmentation tasks demonstrate that the pure Transformer-based U-shaped Encoder-Decoder network outperforms those methods with full-convolution or the combination of transformer and convolution. The codes and trained models will be publicly available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • CNN-based methods
  • Vision transformers
  • Self-attention/Transformer to complement CNNs
  • 3 Method
  • 3.1 Architecture overview
  • 3.2 Swin Transformer block
  • 3.3 Encoder
  • Patch merging layer
  • 3.4 Bottleneck
  • 3.5 Decoder
  • Patch expanding layer
  • 3.6 Skip connection
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Implementation details
  • 4.3 Experiment results on Synapse dataset
  • 4.4 Experiment results on ACDC dataset
  • 4.5 Ablation study
  • Effect of up-sampling:
  • Effect of the number of skip connections:
  • Effect of input size:
  • Effect of model scale:
  • 4.6 Discussion
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Swin-Unet Architecture

    model/method

    Swin-Unet is a pure Transformer-based U-shaped encoder-decoder network for 2D medical image segmentation that avoids standard convolutional operations in its feature extraction and restoration paths.

    Given an input medical image XRH×W×CinX \in \mathbb{R}^{H \times W \times C_{\text{in}}} (where CinC_{\text{in}} is the number of input channels, typically 3 or 1), the architecture proceeds as follows:

    1. Patch Partition and Linear Embedding: The image is split into non-overlapping patches of size 4×44 \times 4, resulting in a flattened feature dimension of 4×4×Cin=484 \times 4 \times C_{\text{in}} = 48 for Cin=3C_{\text{in}}=3. A linear embedding layer projects these patches into token embeddings of dimension CC.
    2. Encoder: Composed of hierarchical stages. Each stage contains two consecutive Swin Transformer blocks (one with window-based self-attention and one with shifted window-based self-attention) followed by a Patch Merging layer that reduces spatial resolution by 2×2\times and doubles the channel dimension (C2C4C8CC \to 2C \to 4C \to 8C). The encoder performs three successive downsampling stages, reaching a spatial resolution of H32×W32\frac{H}{32} \times \frac{W}{32}.
    3. Bottleneck: Consists of two consecutive Swin Transformer blocks operating on the lowest resolution representations (H32×W32×8C\frac{H}{32} \times \frac{W}{32} \times 8C) to learn deep contextual feature interactions without altering resolution or channel dimension.
    4. Decoder: A symmetric counterpart to the encoder. Each stage comprises two consecutive Swin Transformer blocks followed by a Patch Expanding layer that doubles spatial resolution (2×2\times upsampling) and reduces channel dimension by half (8C4C2CC8C \to 4C \to 2C \to C).
    5. Multi-Scale Skip Connections: Features from the encoder stages at scales 1/41/4, 1/81/8, and 1/161/16 are concatenated along the channel dimension with the corresponding upsampled decoder features and projected back to the stage channel dimension via a linear layer.
    6. Segmentation Head: A final Patch Expanding layer performs a 4×4\times upsampling to restore the spatial resolution to the original image dimensions H×WH \times W, followed by a linear projection layer that outputs pixel-level class predictions of shape H×W×KH \times W \times K, where KK is the number of target segmentation classes.
  2. Knowl 2 — Consecutive Swin Transformer Blocks Formulation

    equation

    The basic building unit of Swin-Unet consists of two consecutive Swin Transformer blocks utilizing alternating regular window multi-head self-attention (W-MSA) and shifted window multi-head self-attention (SW-MSA). The computation across two consecutive blocks ll and l+1l+1 is formulated as:

    z^l=W-MSA(LN(zl1))+zl1\hat{z}^l = \text{W-MSA}(\text{LN}(z^{l-1})) + z^{l-1}

    zl=MLP(LN(z^l))+z^lz^l = \text{MLP}(\text{LN}(\hat{z}^l)) + \hat{z}^l

    z^l+1=SW-MSA(LN(zl))+zl\hat{z}^{l+1} = \text{SW-MSA}(\text{LN}(z^l)) + z^l

    zl+1=MLP(LN(z^l+1))+z^l+1z^{l+1} = \text{MLP}(\text{LN}(\hat{z}^{l+1})) + \hat{z}^{l+1}

    where z^l\hat{z}^l and zlz^l denote the intermediate output of the attention module and the output of the multilayer perceptron (MLP) module for the ll-th block, respectively. LN\text{LN} denotes Layer Normalization, and MLP\text{MLP} consists of two fully connected layers with GELU non-linear activation.

    Within each local window, self-attention is computed as:

    Attention(Q,K,V)=SoftMax(QKTd+B)V\text{Attention}(Q, K, V) = \text{SoftMax}\left(\frac{QK^T}{\sqrt{d}} + B\right)V

    where Q,K,VRM2×dQ, K, V \in \mathbb{R}^{M^2 \times d} represent query, key, and value matrices; M2M^2 is the number of patches within a local window, dd is the query/key channel dimension, and BRM2×M2B \in \mathbb{R}^{M^2 \times M^2} is a learnable relative position bias matrix whose values are sampled from a continuous parameter table B^R(2M1)×(2M+1)\hat{B} \in \mathbb{R}^{(2M-1) \times (2M+1)}.

  3. Knowl 3 — Patch Expanding Layer for Transformer-Based Up-Sampling

    model/method

    The patch expanding layer is a pure Transformer-compatible up-sampling module that increases spatial resolution and decreases feature dimension without using convolutional operations or bilinear interpolation.

    For an intermediate feature tensor with spatial resolution WS×HS\frac{W}{S} \times \frac{H}{S} and channel dimension CinC_{\text{in}} (for example, bottleneck features where S=32S = 32 and Cin=8CC_{\text{in}} = 8C):

    1. Channel Expansion: A linear layer is first applied to double the feature dimension to 2×Cin2 \times C_{\text{in}} (e.g., from W32×H32×8C\frac{W}{32} \times \frac{H}{32} \times 8C to W32×H32×16C\frac{W}{32} \times \frac{H}{32} \times 16C).
    2. Spatial Rearrangement: A tensor rearrange (unshuffle) operation reorganizes elements across adjacent spatial coordinates, expanding the spatial dimensions by a factor of 2×2\times while reducing the feature dimension to one-quarter of the expanded channel size (e.g., mapping W32×H32×16CW16×H16×4C\frac{W}{32} \times \frac{H}{32} \times 16C \to \frac{W}{16} \times \frac{H}{16} \times 4C).

    In the final stage of the decoder, the patch expanding layer performs a 4×4\times spatial up-sampling directly to restore feature resolution from W4×H4×C\frac{W}{4} \times \frac{H}{4} \times C to the input resolution W×H×CW \times H \times C before linear projection to class logits.

  4. Knowl 4 — Multi-Scale Skip Connection Mechanism in Swin-Unet

    model/method

    Swin-Unet integrates multi-scale skip connections between the symmetric Swin Transformer encoder and decoder stages at resolution scales 1/41/4, 1/81/8, and 1/161/16 of the original image dimensions.

    At each scale s{4,8,16}s \in \{4, 8, 16\}, shallow high-resolution contextual features from the encoder XencRHs×Ws×CsX_{\text{enc}} \in \mathbb{R}^{\frac{H}{s} \times \frac{W}{s} \times C_s} and upsampled deep features from the decoder XdecRHs×Ws×CsX_{\text{dec}} \in \mathbb{R}^{\frac{H}{s} \times \frac{W}{s} \times C_s} are concatenated along the channel dimension:

    Xconcat=[Xenc,Xdec]RHs×Ws×2CsX_{\text{concat}} = [X_{\text{enc}}, X_{\text{dec}}] \in \mathbb{R}^{\frac{H}{s} \times \frac{W}{s} \times 2C_s}

    A linear transformation layer then projects XconcatX_{\text{concat}} back to dimension CsC_s, preserving the channel size required by the succeeding decoder Swin Transformer blocks while mitigating spatial information loss caused by downsampling.

  5. Knowl 5 — Experimental Setup for Medical Image Segmentation Benchmarks

    experimental setup

    Swin-Unet was evaluated on two public medical segmentation benchmarks:

    1. Synapse Multi-Organ CT Dataset: Comprises 30 clinical abdominal CT cases totaling 3,779 axial slices. 18 cases are designated for training and 12 cases for testing. Models segment 8 abdominal organs: aorta, gallbladder, left kidney, right kidney, liver, pancreas, spleen, and stomach. Performance is evaluated using the Dice Similarity Coefficient (DSC, %) and average Hausdorff Distance (HD, in mm).
    2. Automated Cardiac Diagnosis Challenge (ACDC) Dataset: Comprises cine-MRI scans from 100 patients split into 70 training, 10 validation, and 20 testing cases. Models segment three cardiac structures: left ventricle (LV), right ventricle (RV), and myocardium (MYO). Performance is evaluated using average DSC (%).

    Implementation and Training Configuration:

    • Framework: PyTorch 1.7.0 and Python 3.6 on a single NVIDIA V100 GPU (32 GB memory).
    • Input resolution: 224×224224 \times 224 pixels; patch size: 4×44 \times 4.
    • Parameter initialization: Swin Transformer backbone weights pre-trained on ImageNet.
    • Data augmentation: Random flips and rotations.
    • Optimization: SGD optimizer with momentum 0.9, weight decay 1×1041 \times 10^{-4}, and batch size 24.
  6. Knowl 6 — Multi-Organ CT Segmentation Performance on Synapse

    data/table

    On the Synapse multi-organ CT test set (12 testing cases, 8 organs), Swin-Unet outperforms both pure convolutional networks (V-Net, U-Net, Att-UNet) and hybrid CNN-Transformer architectures (TransUnet), achieving the highest overall mean DSC (79.13%) and lowest Hausdorff Distance (21.55 mm).

    Methods DSC\uparrow HD\downarrow Aorta Gallbladder Kidney(L) Kidney(R) Liver Pancreas Spleen Stomach
    V-Net 68.81 - 75.34 51.87 77.10 80.75 87.84 40.05 80.56 56.98
    DARR 69.77 - 74.74 53.77 72.31 73.24 94.08 54.18 89.90 45.96
    R50 U-Net 74.68 36.87 87.74 63.66 80.60 78.19 93.74 56.90 85.87 74.16
    U-Net 76.85 39.70 89.07 69.72 77.77 68.60 93.43 53.98 86.67 75.58
    R50 Att-UNet 75.57 36.97 55.92 63.91 79.20 72.71 93.56 49.37 87.19 74.95
    Att-UNet 77.77 36.02 89.55 68.88 77.98 71.11 93.57 58.04 87.30 75.75
    R50 ViT 71.29 32.87 73.73 55.13 75.80 72.20 91.51 45.99 81.99 73.95
    TransUnet 77.48 31.69 87.23 63.13 81.87 77.02 94.08 55.86 85.08 75.62
    SwinUnet 79.13 21.55 85.47 66.53 83.28 79.61 94.29 56.58 90.66 76.60

    The reduction in Hausdorff Distance from 31.69 mm (TransUnet) and 36.02 mm (Att-UNet) to 21.55 mm demonstrates that Swin-Unet produces substantially more accurate edge and boundary segmentations.

  7. Knowl 7 — Cardiac MRI Segmentation Performance on ACDC

    data/table

    On the Automated Cardiac Diagnosis Challenge (ACDC) MRI dataset, Swin-Unet achieves an average Dice Similarity Coefficient (DSC) of 90.00% across the right ventricle (RV), myocardium (MYO), and left ventricle (LV), surpassing CNN-based and hybrid Transformer baselines.

    Methods DSC RV Myo LV
    R50 U-Net 87.55 87.10 80.63 94.92
    R50 Att-UNet 86.75 87.58 79.20 93.47
    R50 ViT 87.57 86.07 81.88 94.75
    TransUnet 89.71 88.86 84.53 95.73
    SwinUnet 90.00 88.55 85.62 95.83
  8. Knowl 8 — Ablation on Decoder Up-Sampling Strategies

    data/table

    An ablation study on the Synapse multi-organ CT dataset compared three distinct decoder up-sampling mechanisms integrated into the Swin-Unet architecture: bilinear interpolation, transposed convolution, and the proposed patch expanding layer.

    Up-sampling DSC Aorta Gallbladder Kidney(L) Kidney(R) Liver Pancreas Spleen Stomach
    Bilinear interpolation 76.15 81.84 66.33 80.12 73.91 93.64 55.04 86.10 72.20
    Transposed convolution 77.63 84.81 65.96 82.66 74.61 94.39 54.81 89.42 74.41
    Patch expand 79.13 85.47 66.53 83.28 79.61 94.29 56.58 90.66 76.60

    The patch expanding layer achieves a mean DSC of 79.13%, outperforming transposed convolution (77.63%) by +1.50% and bilinear interpolation (76.15%) by +2.98%, showing that rearranging patch tokens preserves contextual feature representations more effectively than standard up-sampling methods.

  9. Knowl 9 — Ablations on Skip Connections, Input Resolution, and Model Scale

    empirical result

    Ablation experiments on the Synapse dataset demonstrate how architectural choices affect Swin-Unet's segmentation accuracy (measured by mean Dice Similarity Coefficient, DSC %):

    1. Number of Skip Connections:

      • 0 skip connections: 72.46% DSC
      • 1 skip connection (at scale 1/41/4): 76.43% DSC
      • 2 skip connections (at scales 1/4,1/81/4, 1/8): 78.93% DSC
      • 3 skip connections (at scales 1/4,1/8,1/161/4, 1/8, 1/16): 79.13% DSC Segmentation accuracy improves monotonically with the inclusion of multi-scale skip connections.
    2. Input Resolution:

      • 224×224224 \times 224 pixels: 79.13% DSC (Aorta 85.47, Gallbladder 66.53, Left Kidney 83.28, Right Kidney 79.61, Liver 94.29, Pancreas 56.58, Spleen 90.66, Stomach 76.60)
      • 384×384384 \times 384 pixels: 81.12% DSC (Aorta 87.07, Gallbladder 70.53, Left Kidney 84.64, Right Kidney 82.87, Liver 94.72, Pancreas 63.73, Spleen 90.14, Stomach 75.29) Increasing the input resolution from 224 to 384 expands the token sequence length and yields a +1.99% DSC improvement, at the expense of higher computational cost.
    3. Model Scale:

      • Tiny backbone: 79.13% DSC
      • Base backbone: 79.25% DSC Scaling the model capacity from Tiny to Base yields a negligible gain (+0.12% DSC) while substantially increasing computational requirements.
  10. Knowl 10 — Pre-Training and Dimensionality Limitations in Swin-Unet

    limitation

    Swin-Unet has two notable structural limitations:

    1. Pre-Training Discrepancy: The model initializes both encoder and decoder using Swin Transformer weights pre-trained on 2D ImageNet classification, which is suboptimal compared to end-to-end medical pre-training or specialized self-supervised representations tailored for medical imagery.
    2. 2D Slice-Based Processing: Swin-Unet processes volumetric medical imaging modalities (such as abdominal CT and cardiac cine-MRI) as independent 2D slices, failing to explicitly capture 3D inter-slice spatial continuity and volumetric context.

Coverage note — None was omitted; all key architectural components, formulations, experimental benchmarks, ablation results, and discussed limitations are fully covered.

References

  1. 1.A. Hatamizadeh, D. Yang, H. Roth, and D. Xu, “Unetr: Transformers for 3d medical image segmentation,” 2021.
  2. 2.J. Chen, Y. Lu, Q. Yu, X. Luo, E. Adeli, Y. Wang, L. Lu, A. L. Yuille, and Y. Zhou, “Transunet: Transformers make strong encoders for medical image segmentation,” CoRR, vol. abs/2102.04306, 2021.
  3. 3.O. Ronneberger, P.Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention (MICCAI), ser. LNCS, vol. 9351. Springer, 2015, pp. 234–241.
  4. 4.K. S. P. J. M.-H. K. Isensee F, Jaeger PF, “nnu-net: a self-configuring method for deep learning-based biomedical image segmentation,” Nat Methods, vol. 18(2):203-211, 2021.
  5. 5.Q. Jin, Z. Meng, C. Sun, H. Cui, and R. Su, “Ra-unet: A hybrid deep attention-aware network to extract liver and tumor in ct scans,” Frontiers in Bioengineering and Biotechnology, vol. 8, p. 1471, 2020.
  6. 6.O. Çiçek, A. Abdulkadir, S. Lienkamp, T. Brox, and O. Ronneberger, “3d u-net: Learning dense volumetric segmentation from sparse annotation,” in Medical Image Computing and Computer-Assisted Intervention (MICCAI), ser. LNCS, vol. 9901. Springer, Oct 2016, pp. 424–432.
  7. 7.X. Xiao, S. Lian, Z. Luo, and S. Li, “Weighted res-unet for high-quality retina vessel segmentation,” 2018 9th International Conference on Information Technology in Medicine and Education (ITME), pp. 327–331, 2018.
  8. 8.Z. Zhou, M. Rahman Siddiquee, N. Tajbakhsh, and J. Liang, “Unet++: A nested u-net architecture for medical image segmentation.” Springer Verlag, 2018, pp. 3–11.
  9. 9.H. Huang, L. Lin, R. Tong, H. Hu, Q. Zhang, Y. Iwamoto, X. Han, Y.-W. Chen, and J. Wu, “Unet 3+: A full-scale connected unet for medical image segmentation,” 2020.
  10. 10.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 4, pp. 834–848, 2018.
  11. 11.Z. Gu, J. Cheng, H. Fu, K. Zhou, H. Hao, Y. Zhao, T. Zhang, S. Gao, and J. Liu, “Ce-net: Context encoder network for 2d medical image segmentation,” IEEE Transactions on Medical Imaging, vol. 38, no. 10, pp. 2281–2292, 2019.
  12. 12.J. Schlemper, O. Oktay, M. Schaap, M. Heinrich, B. Kainz, B. Glocker, and D. Rueckert, “Attention gated networks: Learning to leverage salient regions in medical images,” Medical Image Analysis, vol. 53, pp. 197–207, 2019.
  13. 13.X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 7794–7803.
  14. 14.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 6230–6239.
  15. 15.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30. Curran Associates, Inc., 2017.
  16. 16.N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” CoRR, vol. abs/2005.12872, 2020.
  17. 17.A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations, 2021.
  18. 18.H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” CoRR, vol. abs/2012.12877, 2020.
  19. 19.Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” CoRR, vol. abs/2103.14030, 2021.
  20. 20.A. Tsai, A. Yezzi, W. Wells, C. Tempany, D. Tucker, A. Fan, W. Grimson, and A. Willsky, “A shape-based approach to the segmentation of medical imagery using level sets,” IEEE Transactions on Medical Imaging, vol. 22, no. 2, pp. 137–154, 2003.
  21. 21.K. Held, E. Kops, B. Krause, W. Wells, R. Kikinis, and H.-W. Muller-Gartner, “Markov random field segmentation of brain mr images,” IEEE Transactions on Medical Imaging, vol. 16, no. 6, pp. 878–886, 1997.
  22. 22.X. Li, H. Chen, X. Qi, Q. Dou, C.-W. Fu, and P.-A. Heng, “H-denseunet: Hybrid densely connected unet for liver and tumor segmentation from ct volumes,” IEEE Transactions on Medical Imaging, vol. 37, no. 12, pp. 2663–2674, 2018.
  23. 23.F. Milletari, N. Navab, and S.-A. Ahmadi, “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” 2016 Fourth International Conference on 3D Vision (3DV), pp. 565–571, 2016.
  24. 24.J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4171–4186. [Online]. Available: https://www.aclweb.org/anthology/N19-1423
  25. 25.W. Wang, E. Xie, X. Li, D. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” CoRR, vol. abs/2102.12122, 2021. [Online]. Available: https://arxiv.org/abs/2102.12122
  26. 26.K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang, “Transformer in transformer,” CoRR, vol. abs/2103.00112, 2021. [Online]. Available: https://arxiv.org/abs/2103.00112
  27. 27.J. M. J. Valanarasu, P. Oza, I. Hacihaliloglu, and V. M. Patel, “Medical transformer: Gated axial-attention for medical image segmentation,” CoRR, vol. abs/2102.10662, 2021.
  28. 28.Y. Zhang, H. Liu, and Q. Hu, “Transfuse: Fusing transformers and cnns for medical image segmentation,” CoRR, vol. abs/2102.08005, 2021. [Online]. Available: https://arxiv.org/abs/2102.08005
  29. 29.W. Wang, C. Chen, M. Ding, J. Li, H. Yu, and S. Zha, “Transbts: Multimodal brain tumor segmentation using transformer,” CoRR, vol. abs/2103.04430, 2021. [Online]. Available: https://arxiv.org/abs/2103.04430
  30. 30.Y. Xie, J. Zhang, C. Shen, and Y. Xia, “Cotr: Efficiently bridging CNN and transformer for 3d medical image segmentation,” CoRR, vol. abs/2103.03024, 2021. [Online]. Available: https://arxiv.org/abs/2103.03024
  31. 31.H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei, “Relation networks for object detection,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 3588–3597.
  32. 32.H. Hu, Z. Zhang, Z. Xie, and S. Lin, “Local relation networks for image recognition,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 3463–3472.
  33. 33.H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou, “Going deeper with image transformers,” CoRR, vol. abs/2103.17239, 2021. [Online]. Available: https://arxiv.org/abs/2103.17239
  34. 34.S. Fu, Y. Lu, Y. Wang, Y. Zhou, W. Shen, E. Fishman, and A. Yuille, “Domain adaptive relational reasoning for 3d multi-organ segmentation,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2020, 2020, pp. 656–666.
  35. 35.F. Milletari, N. Navab, and S.-A. Ahmadi, “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” in 2016 Fourth International Conference on 3D Vision (3DV), 2016, pp. 565–571.
  36. 36.S. Fu, Y. Lu, Y. Wang, Y. Zhou, W. Shen, E. Fishman, and A. Yuille, “Domain adaptive relational reasoning for 3d multi-organ segmentation,” Germany, 2020, pp. 656–666.
  37. 37.O. Oktay, J. Schlemper, L. L. Folgoc, M. Lee, M. Heinrich, K. Misawa, K. Mori, S. McDonagh, N. Y. Hammerla, B. Kainz, B. Glocker, and D. Rueckert, “Attention u-net: Learning where to look for the pancreas,” IMIDL Conference, 2018.

Citation

MLA
Cao, H., et al. “Swin-Unet: Unet-Like Pure Transformer for Medical Image Segmentation”. Lecture Notes in Computer Science, Springer Nature Switzerland, 2023, pp. 205–18, https://doi.org/10.1007/978-3-031-25066-8_9.
APA
Cao, H., Wang, Y., Chen, J., Jiang, D., Zhang, X., Tian, Q., & Wang, M. (2023). Swin-Unet: Unet-Like Pure Transformer for Medical Image Segmentation. In Lecture Notes in Computer Science (pp. 205–218). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-25066-8_9
Chicago
Cao, H., Y. Wang, J. Chen, et al. 2023. “Swin-Unet: Unet-Like Pure Transformer for Medical Image Segmentation”. In Lecture Notes in Computer Science. Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-25066-8_9.
Harvard
Cao, H. et al. (2023) “Swin-Unet: Unet-Like Pure Transformer for Medical Image Segmentation”, Lecture Notes in Computer Science. Springer Nature Switzerland, pp. 205–218. Available at: https://doi.org/10.1007/978-3-031-25066-8_9.
Vancouver
1. Cao H, Wang Y, Chen J, Jiang D, Zhang X, Tian Q, Wang M (2023) Swin-Unet: Unet-Like Pure Transformer for Medical Image Segmentation. In: Lecture Notes in Computer Science. Springer Nature Switzerland, pp 205–218

BibTeX

@inbook{Cao_2023, title={Swin-Unet: Unet-Like Pure Transformer for Medical Image Segmentation}, ISBN={9783031250668}, ISSN={1611-3349}, url={http://dx.doi.org/10.1007/978-3-031-25066-8_9}, DOI={10.1007/978-3-031-25066-8_9}, booktitle={Computer Vision – ECCV 2022 Workshops}, publisher={Springer Nature Switzerland}, author={Cao, Hu and Wang, Yueyue and Chen, Joy and Jiang, Dongsheng and Zhang, Xiaopeng and Tian, Qi and Wang, Manning}, year={2023}, pages={205–218} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/