Class-Aware Adversarial Transformers for Medical Image Segmentation

Chenyu YouRuihan ZhaoFenglin LiuSiyuan DongSandeep ChinchaliUfuk TopcuLawrence H. StaibJames S. Duncan

article2022NeurIPS137 citations

Proposes CASTformer, a medical image segmentation framework combining multi-scale pyramid representations, a class-aware transformer module, and adversarial training to overcome standard tokenization limitations and substantially improve segmentation accuracy across multiple benchmarks.

Listen

Accurate segmentation of anatomical structures and abnormalities in medical images is essential for clinical diagnosis, treatment planning, and post-treatment monitoring. While artificial intelligence models based on transformer architectures have shown strong capability in modeling broad visual relationships, existing methods struggle to capture fine anatomical details and multi-scale variations. They typically rely on rigid image-splitting techniques and single-resolution representations, leading to boundary inaccuracies and high data demands that limit their reliability in clinical settings.

The article introduces and evaluates CASTformer, a class-aware adversarial transformer framework designed specifically for two-dimensional medical image segmentation. The core objective is to demonstrate that combining multi-scale visual representations, adaptive anatomical focus, and adversarial training can significantly improve segmentation accuracy while reducing reliance on large annotated medical datasets.

To achieve this, the authors built a hybrid architecture comprising a multi-scale generator and a discriminator paired through adversarial training. The generator incorporates a multi-level feature pyramid and an iterative class-aware module that learns to focus on meaningful anatomical regions rather than irrelevant backgrounds. The discriminator acts as a critic to enforce realistic anatomical textures and global consistency. The framework was evaluated across standard medical benchmarks, including the Synapse multi-organ computed tomography dataset and the Liver Tumor Segmentation benchmark, comparing performance against established deep learning and transformer baselines while also testing the impact of pre-training on natural computer vision datasets.

CASTformer achieved state-of-the-art results across the evaluated benchmarks. On the Synapse multi-organ dataset, the model achieved an overall accuracy score of 82.55% Dice coefficient, representing an absolute improvement of 5.07% over the prior leading transformer model, with substantial accuracy gains on challenging small organs like the pancreas (up to 10.91% higher). On liver tumor segmentation, CASTformer reached 73.82% overall accuracy, boosting tumor-specific segmentation accuracy by more than 9% in absolute terms compared to prior methods. Furthermore, transfer learning experiments revealed that initializing both the generator and discriminator with pre-trained weights from computer vision models improved performance by up to 8.91% compared to training from scratch, allowing the system to train effectively on smaller datasets.

These findings indicate that architectural refinements focusing on multi-scale contexts and region-specific features can overcome the primary limitations of vision transformers in medical analysis. For healthcare organizations and technology developers, using pre-trained computer vision weights offers a cost-effective pathway to high-performing clinical artificial intelligence tools without the expensive bottleneck of hand-labeling massive proprietary medical datasets. Improved boundary detection directly translates to lower clinical risk during diagnostic review and surgical planning.

Organizations developing medical imaging systems should consider adopting hybrid, multi-scale transformer frameworks and prioritizing transfer learning workflows. Moving forward, clinical adoption will require optimizing computational efficiency, building mechanistic explanations to satisfy clinical transparency requirements, and conducting pilot testing on multi-center clinical data to validate performance across broader operational environments.

The reported conclusions are supported by rigorous benchmarking against top baseline models across multiple standard datasets. However, the study focuses primarily on two-dimensional image segmentation tasks and relies on baseline architectures pre-trained on natural images, meaning performance should be verified cautiously when deploying directly to volumetric three-dimensional clinical pipelines or untested imaging modalities.

arXiv: 2201.10737
Cover for Class-Aware Adversarial Transformers for Medical Image Segmentation

Abstract

Transformers have made remarkable progress towards modeling long-range dependencies within the medical image analysis domain. However, current transformer-based models suffer from several disadvantages: (1) existing methods fail to capture the important features of the images due to the naive tokenization scheme; (2) the models suffer from information loss because they only consider single-scale feature representations; and (3) the segmentation label maps generated by the models are not accurate enough without considering rich semantic contexts and anatomical textures. In this work, we present CASTformer, a novel type of adversarial transformers, for 2D medical image segmentation. First, we take advantage of the pyramid structure to construct multi-scale representations and handle multi-scale variations. We then design a novel class-aware transformer module to better learn the discriminative regions of objects with semantic structures. Lastly, we utilize an adversarial training strategy that boosts segmentation accuracy and correspondingly allows a transformer-based discriminator to capture high-level semantically correlated contents and low-level anatomical features. Our experiments demonstrate that CASTformer dramatically outperforms previous state-of-the-art transformer-based approaches on three benchmarks, obtaining 2.54%-5.88% absolute improvements in Dice over previous models. Further qualitative experiments provide a more detailed picture of the model's inner workings, shed light on the challenges in improved transparency, and demonstrate that transfer learning can greatly improve performance and reduce the size of medical image datasets in training, making CASTformer a strong starting point for downstream medical image analysis tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 4 Experimental Setup
  • 5 Results
  • 6 Analysis
  • 7 Conclusion and Discussion of Broader Impact
  • References

Knowls

  1. Knowl 1 — CASTformer architecture for 2D medical image segmentation

    model/method

    CASTformer is a 2D medical-image segmentation system formed by a transformer-based generator, called CATformer, and a transformer-based discriminator. Given an input image x∈RH×W×3x \in \mathbb{R}^{H\times W\times 3}, CATformer predicts a multi-class segmentation mask y′y'. The generator combines a convolutional feature extractor, class-aware transformer modules, transformer encoder modules, and a lightweight decoder. Its four parallel stages process feature representations at multiple resolutions; each stage contains patch embedding, one class-aware module, and a stage-specific number of transformer encoder layers. The discriminator uses a ResNet-50 plus ViT-B/16 hybrid transformer initialized from ImageNet pretraining and classifies real versus generated anatomical demonstrations. The complete CASTformer model therefore combines multi-scale convolutional features, adaptive region sampling, global self-attention, and adversarial regularization.

  2. Knowl 2 — Hierarchical feature pyramid and lightweight segmentation decoder

    model/method

    The CATformer generator uses a CNN stem of approximately 40 convolutional layers to produce a feature pyramid rather than a single-resolution transformer representation. For an input of spatial size H×WH\times W, the reported feature maps are F1∈RH/2×W/2×C1F_1 \in \mathbb{R}^{H/2\times W/2\times C_1}, F2∈RH/2×W/2×(4C1)F_2 \in \mathbb{R}^{H/2\times W/2\times (4C_1)}, F3∈RH/4×W/4×(8C1)F_3 \in \mathbb{R}^{H/4\times W/4\times (8C_1)}, and F4∈RH/8×W/8×(12C1)F_4 \in \mathbb{R}^{H/8\times W/8\times (12C_1)}, where C1C_1 is the channel count of the first feature map. The pyramid supplies high-resolution features for local boundaries and lower-resolution features for broader spatial context. The decoder first maps the channels of all four outputs to a common dimension with MLP layers, upsamples the features to one quarter of the input resolution, concatenates them, applies another MLP for fusion, and predicts the multi-class mask. This design avoids a computationally heavy hand-crafted decoder while retaining multi-scale information.

  3. Knowl 3 — Iterative class-aware transformer sampling

    algorithm

    The class-aware transformer (CAT) module adaptively samples discriminative locations from a feature map instead of using only fixed, regularly partitioned image patches. For a feature map FiF_i of spatial size Hi×WiH_i\times W_i and channel dimension CiC_i, the module samples n2n^2 locations over MM iterations. At iteration one, the location of sample qq, with q∈{0,…,n2−1}q\in\{0,\ldots,n^2-1\}, is initialized on a regular grid: let rq=⌊q/n⌋r_q=\lfloor q/n\rfloor, cq=q−rqnc_q=q-r_qn, τh=Hi/n\tau_h=H_i/n, and τw=Wi/n\tau_w=W_i/n; then s1,q=(rqτh+τh/2, cqτw+τw/2)s_{1,q}=(r_q\tau_h+\tau_h/2,\ c_q\tau_w+\tau_w/2). Each iteration uses differentiable bilinear interpolation to obtain sampled features It′=Fi(st)∈RCi×n2I'_t=F_i(s_t)\in\mathbb{R}^{C_i\times n^2}. The locations are embedded with a learnable matrix Wt∈RCi×2W_t\in\mathbb{R}^{C_i\times 2}, giving St=WtstS_t=W_ts_t, and the current token sequence is updated by a transformer encoder: It=Transformer⁡(It′+St+It−1)I_t=\operatorname{Transformer}(I'_t+S_t+I_{t-1}), with I0I_0 initialized to zero. For t<Mt<M, a learned linear map θt\theta_t predicts offsets ot=θt(It)∈R2×n2o_t=\theta_t(I_t)\in\mathbb{R}^{2\times n^2}, and the next locations are st+1=st+ots_{t+1}=s_t+o_t. The entire sampling process is differentiable and trained end-to-end, allowing later iterations to move samples toward semantically informative anatomical regions. The reported experiments use n=16n=16 samples per feature map and M=4M=4 iterative steps.

  4. Knowl 4 — Transformer encoder for global contextual modeling

    model/method

    Each transformer encoder module (TEM) models long-range relationships among the complete sequence of patch embeddings. Let xpj∈RP2Cx_p^j\in\mathbb{R}^{P^2C} be the flattened pixels of patch jj, let NN be the number of patches, let CC be the input channel count, let DD be the transformer embedding dimension, and let He∈R(P2C)×DH_e\in\mathbb{R}^{(P^2C)\times D} and Hpos∈RN×DH_{pos}\in\mathbb{R}^{N\times D} be learnable projection and positional-embedding matrices. The initial sequence is E0=[xp1He;…;xpNHe]+HposE_0=[x_p^1H_e;\ldots;x_p^NH_e]+H_{pos}. Each encoder layer applies pre-normalized multi-head self-attention followed by a residual connection and then a pre-normalized MLP followed by another residual connection:

    Eℓ′=MSA⁡(LN⁡(Eℓ−1))+Eℓ−1E'_\ell=\operatorname{MSA}(\operatorname{LN}(E_{\ell-1}))+E_{\ell-1} and Eℓ=MLP⁡(LN⁡(Eℓ′))+Eℓ′E_\ell=\operatorname{MLP}(\operatorname{LN}(E'_\ell))+E'_\ell, where ℓ\ell indexes the encoder layer, LN⁡\operatorname{LN} is layer normalization, and MSA is multi-head self-attention. CATformer combines these global contextual representations with the adaptively sampled tokens from the CAT modules.

  5. Knowl 5 — Class-aware adversarial training objective

    equation

    The discriminator is designed to judge anatomical content rather than the entire image background. For an input image xx and a predicted segmentation mask y′y', the class-aware image is formed by pixelwise multiplication, x~=x⊙y′\tilde{x}=x\odot y', where ⊙\odot denotes elementwise multiplication. The discriminator uses a two-layer MLP head on top of the pretrained ResNet-50/ViT-B/16 hybrid and is trained in a Wasserstein GAN with gradient penalty framework to distinguish real from generated class-aware samples. The generator is optimized jointly with cross-entropy, Dice, and WGAN-GP losses:

    LG=λ1LCE+λ2LDICE+λ3LWGAN-GP,L_G=\lambda_1L_{CE}+\lambda_2L_{DICE}+\lambda_3L_{WGAN\text{-}GP},

    where LCEL_{CE} is the pixelwise cross-entropy loss, LDICEL_{DICE} is the segmentation Dice loss, LWGAN-GPL_{WGAN\text{-}GP} is the adversarial Wasserstein loss with gradient penalty, and λ1,λ2,λ3\lambda_1,\lambda_2,\lambda_3 weight the terms. The reported setting is λ1=0.5\lambda_1=0.5, λ2=0.5\lambda_2=0.5, and λ3=0.1\lambda_3=0.1. The adversarial component encourages masks whose selected anatomical regions have both high-level semantic coherence and low-level anatomical fidelity.

  6. Knowl 6 — Training and evaluation configuration

    experimental setup

    The method is evaluated on the Synapse multi-organ CT, LiTS liver-and-tumor CT, and MP-MRI benchmarks. All experiments use AdamW with learning rate 5×10−45\times10^{-4}, batch size 6, and 300 training epochs for both generator and discriminator. Images are resized to 224×224224\times224 and the patch size is set to 14. The CAT module uses n=16n=16 samples per feature map and M=4M=4 iterative sampling steps. Models are implemented in PyTorch 1.7.0 and trained on one NVIDIA GeForce RTX 3090 GPU with 24 GB memory. Segmentation is evaluated using Dice coefficient and Jaccard index, for which higher values are better, and 95% Hausdorff distance and average symmetric surface distance, for which lower values are better.

  7. Knowl 7 — Synapse multi-organ segmentation performance

    empirical result

    On the Synapse multi-organ CT benchmark, CASTformer achieves an average Dice score of 82.55%, Jaccard index of 74.69%, 95% Hausdorff distance of 22.73, and average symmetric surface distance of 5.81. The generator without the discriminator, CATformer, obtains 82.17% Dice, 73.22% Jaccard, 16.20 95HD, and 4.28 ASD. For comparison, the previous TransUNet result is 77.48% Dice, 64.78% Jaccard, 31.69 95HD, and 8.46 ASD. Thus, CASTformer improves over TransUNet by 5.07 percentage points in Dice and 9.91 percentage points in Jaccard. Its per-organ Dice scores are 89.05% for aorta, 67.48% for gallbladder, 86.05% for left kidney, 82.17% for right kidney, 95.61% for liver, 67.49% for pancreas, 91.00% for spleen, and 81.55% for stomach. Relative to TransUNet, the reported gains are especially large for the stomach (+4.95 points), pancreas (+10.91 points), and the other listed organs: aorta +1.82, gallbladder +0.95, left kidney +2.77, right kidney +2.51, liver +1.35, and spleen +5.92 percentage points. The qualitative comparison on page 6 shows that CASTformer produces more complete organ regions and sharper anatomical boundaries than the compared CNN and transformer baselines.

  8. Knowl 8 — LiTS liver-and-tumor segmentation performance

    empirical result

    On the LiTS CT benchmark, CATformer obtains 72.39% average Dice, whereas CASTformer establishes the reported best result with 73.82% Dice and 64.91% Jaccard. Compared with TransUNet, CASTformer improves Dice by 5.88 percentage points and Jaccard by 4.66 percentage points. CASTformer reaches 95.88% Dice on the liver region, a 2.48-point improvement over the comparison result, and increases tumor Dice from 42.49% to 51.76%. The qualitative LiTS comparison on page 8 shows that the model retains more detailed tumor regions while also producing high-quality liver masks, supporting the paper's claim that adaptive focusing and semantically correlated anatomical information are particularly useful for this difficult segmentation setting.

  9. Knowl 9 — Transfer learning from ImageNet-pretrained hybrid transformers

    empirical result

    On Synapse, initializing the generator and discriminator from the ImageNet-pretrained ResNet-50 plus ViT-B/16 hybrid substantially improves performance. CATformer without pretraining obtains 74.84% Dice, 65.61% Jaccard, 31.81 95HD, and 7.23 ASD; with generator pretraining it reaches 82.17% Dice, 73.22% Jaccard, 16.20 95HD, and 4.28 ASD. For CASTformer, training both networks from scratch gives 73.64% Dice and 62.68% Jaccard, pretraining only the discriminator gives 78.87% and 69.36%, pretraining only the generator gives 81.46% and 71.80%, and pretraining both gives 82.55% and 74.69%, respectively. The both-pretrained setting therefore improves over the both-unpretrained setting by 8.91 Dice points and 12.01 Jaccard points. Pretraining the generator is substantially more beneficial than pretraining only the discriminator, suggesting that transferred visual features in the segmentation generator are more important for exploiting anatomical information and adapting to small medical datasets.

  10. Knowl 10 — Ablation of class-aware sampling and global transformer encoding

    empirical result

    An ablation on Synapse separates the contributions of the class-aware transformer module and the transformer encoder module. The baseline, which removes both modules and is analogous to the TransUNet-style configuration, obtains 77.48% Dice, 64.78% Jaccard, 31.69 95HD, and 8.46 ASD. Removing only CAT gives 80.09% Dice, 70.56% Jaccard, 25.62 95HD, and 7.30 ASD; removing only TEM gives 81.35% Dice, 72.66% Jaccard, 24.43 95HD, and 7.17 ASD. Including both modules in CATformer yields 82.17% Dice, 73.22% Jaccard, 16.20 95HD, and 4.28 ASD, while adding adversarial training produces the final CASTformer values of 82.55% Dice, 74.69% Jaccard, 22.73 95HD, and 5.81 ASD. Relative to the baseline, the CAT-only-removal and TEM-only-removal variants improve Dice by 2.61 and 3.87 percentage points, respectively, while the complete CATformer performs best among the generator ablations. The results support complementary roles: CAT improves selective localization of discriminative anatomical regions, whereas TEM supplies global shape and structural context.

Coverage note — Quantitative MP-MRI results and supplementary analyses of iteration count, sampling count, hyperparameters, alternative GAN losses, and additional interpretability experiments were omitted because the supplied paper text refers to appendices whose contents are not included; no other substantial main-text contribution was deliberately omitted.

References

  1. 1.Mehrdad Moghbel, Syamsiah Mashohor, Rozi Mahmud, and M Iqbal Bin Saripan. Review of liver segmentation and computer assisted detection/diagnosis methods in computed tomography. Artificial Intelligence, 2017.
  2. 2.Yuan Xue, Tao Xu, Han Zhang, L Rodney Long, and Xiaolei Huang. Segan: Adversarial network with multi-scale l 1 loss for medical image segmentation. Neuroinformatics, 2018.
  3. 3.Drew A Hudson and Larry Zitnick. Generative adversarial transformers. In International Conference on Machine Learning (ICML), 2021.
  4. 4.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2020.
  5. 5.Daquan Zhou, Bingyi Kang, Xiaojie Jin, Linjie Yang, Xiaochen Lian, Zihang Jiang, Qibin Hou, and Jiashi Feng. Deepvit: Towards deeper vision transformer. arXiv preprint arXiv:2103.11886, 2021.
  6. 6.Aditya Desai, Zhaozhuo Xu, Menal Gupta, Anu Chandran, Antoine Vial-Aussavy, and Anshumali Shrivastava. Raw nav-merge seismic data to subsurface properties with mlp based multi-modal information unscrambler. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  7. 7.Jieneng Chen, Yongyi Lu, Qihang Yu, Xiangde Luo, Ehsan Adeli, Yan Wang, Le Lu, Alan L Yuille, and Yuyin Zhou. Transunet: Transformers make strong encoders for medical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2021.
  8. 8.Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al. Spatial transformer networks. Advances in neural information processing systems, 28, 2015.
  9. 9.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  10. 10.Hu Cao, Yueyue Wang, Joy Chen, Dongsheng Jiang, Xiaopeng Zhang, Qi Tian, and Manning Wang. Swin-unet: Unet-like pure transformer for medical image segmentation. arXiv preprint arXiv:2105.05537, 2021.
  11. 11.Wenxuan Wang, Chen Chen, Meng Ding, Hong Yu, Sen Zha, and Jiangyun Li. Transbts: Multimodal brain tumor segmentation using transformer. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2021.
  12. 12.Jeya Maria Jose Valanarasu, Poojan Oza, Ilker Hacihaliloglu, and Vishal M Patel. Medical transformer: Gated axial-attention for medical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2021.
  13. 13.Yutong Xie, Jianpeng Zhang, Chunhua Shen, and Yong Xia. Cotr: Efficiently bridging cnn and transformer for 3d medical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2021.
  14. 14.Lingke Kong, Chenyu Lian, Detian Huang, Yanle Hu, Qichao Zhou, et al. Breaking the dilemma of medical image-to-image translation. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  15. 15.Nicolae-Catalin Ristea, Andreea-Iuliana Miron, Olivian Savencu, Mariana-Iuliana Georgescu, Nicolae Verga, Fahad Shahbaz Khan, and Radu Tudor Ionescu. Cytran: Cycle-consistent transformers for non-contrast to contrast ct translation. arXiv preprint arXiv:2110.06400, 2021.
  16. 16.Onat Dalmaz, Mahmut Yurt, and Tolga Çukur. Resvit: Residual vision transformers for multi-modal medical image synthesis. arXiv preprint arXiv:2106.16031, 2021.
  17. 17.Yilmaz Korkmaz, Salman UH Dar, Mahmut Yurt, Muzaffer Özbey, and Tolga Çukur. Unsupervised mri reconstruction via zero-shot learned adversarial transformers. arXiv preprint arXiv:2105.08059, 2021.
  18. 18.Zhicheng Zhang, Lequan Yu, Xiaokun Liang, Wei Zhao, and Lei Xing. Transct: Dual-path transformer for low dose computed tomography. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2021.
  19. 19.Qing Lyu, Chenyu You, Hongming Shan, and Ge Wang. Super-resolution mri through deep learning. arXiv preprint arXiv:1810.06776, 2018.
  20. 20.Qing Lyu, Chenyu You, Hongming Shan, Yi Zhang, and Ge Wang. Super-resolution mri and ct through gan-circle. In Developments in X-ray tomography XII, volume 11113, page 111130X. International Society for Optics and Photonics, 2019.
  21. 21.Indranil Guha, Syed Ahmed Nadeem, Chenyu You, Xiaoliu Zhang, Steven M Levy, Ge Wang, James C Torner, and Punam K Saha. Deep learning based high-resolution reconstruction of trabecular bone microstructures from low-resolution ct scans using gan-circle. In Medical Imaging 2020: Biomedical Applications in Molecular, Structural, and Functional Imaging. International Society for Optics and Photonics, 2020.
  22. 22.Chenyu You, Qingsong Yang, Hongming Shan, Lars Gjesteby, Guang Li, Shenghong Ju, Zhuiyang Zhang, Zhen Zhao, Yi Zhang, Wenxiang Cong, et al. Structurally-sensitive multi-scale deep neural network for low-dose ct denoising. IEEE access, 2018.
  23. 23.Chenyu You, Linfeng Yang, Yi Zhang, and Ge Wang. Low-dose ct via deep cnn with skip connection and network-in-network. In Developments in X-Ray tomography XII. International Society for Optics and Photonics, 2019.
  24. 24.Chenyu You, Lianyi Han, Aosong Feng, Ruihan Zhao, Hui Tang, and Wei Fan. Megan: Memory enhanced graph attention network for space-time video super-resolution. In In Proceedings of WACV 2022, 2021.
  25. 25.Achleshwar Luthra, Harsh Sulakhe, Tanish Mittal, Abhishek Iyer, and Santosh Yadav. Eformer: Edge enhancement based transformer for medical image denoising. arXiv preprint arXiv:2109.08044, 2021.
  26. 26.Dayang Wang, Zhan Wu, and Hengyong Yu. Ted-net: Convolution-free t2t vision transformer-based encoder-decoder dilation network for low-dose ct denoising. arXiv preprint arXiv:2106.04650, 2021.
  27. 27.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2015.
  28. 28.Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena. Self-attention generative adversarial networks. In International Conference on Learning Representations (ICLR), 2019.
  29. 29.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision (ECCV), 2020.
  30. 30.Yanhong Zeng, Jianlong Fu, and Hongyang Chao. Learning joint spatial-temporal transformations for video inpainting. In European Conference on Computer Vision (ECCV), 2020.
  31. 31.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  32. 32.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  33. 33.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877, 2020.
  34. 34.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou. Going deeper with image transformers. arXiv preprint arXiv:2103.17239, 2021.
  35. 35.S. Kevin Zhou, Hayit Greenspan, Christos Davatzikos, James S. Duncan, Bram Van Ginneken, Anant Madabhushi, Jerry L. Prince, Daniel Rueckert, and Ronald M. Summers. A review of deep learning in medical imaging: Imaging traits, technology trends, case studies with progress highlights, and future promises. Proceedings of the IEEE, 2021.
  36. 36.Gonglei Shi, Li Xiao, Yang Chen, and S Kevin Zhou. Marginal loss and exclusion loss for partially supervised multi-organ segmentation. Medical Image Analysis, 2021.
  37. 37.Qingsong Yao, Li Xiao, Peihang Liu, and S Kevin Zhou. Label-free segmentation of covid-19 lesions in lung ct. IEEE Transactions on Medical Imaging, 2021.
  38. 38.Jiuwen Zhu, Yuexiang Li, Yifan Hu, Kai Ma, S Kevin Zhou, and Yefeng Zheng. Rubik’s cube+: A self-supervised feature learning framework for 3d medical image analysis. Medical Image Analysis, 2020.
  39. 39.Xiaoyu Yue, Shuyang Sun, Zhanghui Kuang, Meng Wei, Philip Torr, Wayne Zhang, and Dahua Lin. Vision transformer with progressive sampling. In IEEE International Conference on Computer Vision (ICCV), 2021.
  40. 40.Pauline Luc, Camille Couprie, Soumith Chintala, and Jakob Verbeek. Semantic segmentation using adversarial networks. arXiv preprint arXiv:1611.08408, 2016.
  41. 41.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In Advances in Neural Information Processing Systems (NeurIPS), 2016.
  42. 42.Hayit Greenspan, Bram Van Ginneken, and Ronald M Summers. Guest editorial deep learning in medical imaging: Overview and future promise of an exciting new technique. IEEE Transactions on Medical Imaging, 2016.
  43. 43.Geert Litjens, Thijs Kooi, Babak Ehteshami Bejnordi, Arnaud Arindra Adiyoso Setio, Francesco Ciompi, Mohsen Ghafoorian, Jeroen AWM van der Laak, Bram Van Ginneken, and Clara I Sánchez. A survey on deep learning in medical image analysis. Medical Image Analysis, 2017.
  44. 44.Yizhe Zhang, Lin Yang, Jianxu Chen, Maridel Fredericksen, David P Hughes, and Danny Z Chen. Deep adversarial networks for biomedical image segmentation utilizing unannotated images. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2017.
  45. 45.Xiaomeng Li, Lequan Yu, Hao Chen, Chi-Wing Fu, and Pheng-Ann Heng. Semi-supervised skin lesion segmentation via transformation consistent self-ensembling model. arXiv preprint arXiv:1808.03887, 2018.
  46. 46.Dong Nie, Yaozong Gao, Li Wang, and Dinggang Shen. Asdnet: Attention based semi-supervised deep networks for medical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2018.
  47. 47.Yicheng Wu, Yong Xia, Yang Song, Yanning Zhang, and Weidong Cai. Multiscale network followed network model for retinal vessel segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2018.
  48. 48.Lequan Yu, Shujun Wang, Xiaomeng Li, Chi-Wing Fu, and Pheng-Ann Heng. Uncertainty-aware self-ensembling model for semi-supervised 3d left atrium segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2019.
  49. 49.Yu Zeng, Yunzhi Zhuge, Huchuan Lu, and Lihe Zhang. Joint learning of saliency detection and weakly supervised semantic segmentation. In IEEE International Conference on Computer Vision (ICCV), 2019.
  50. 50.Gerda Bortsova, Florian Dubost, Laurens Hogeweg, Ioannis Katramados, and Marleen de Bruijne. Semi-supervised medical image segmentation via learning consistency under transformations. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2019.
  51. 51.Yicheng Wu, Yong Xia, Yang Song, Donghao Zhang, Dongnan Liu, Chaoyi Zhang, and Weidong Cai. Vessel-net: retinal vessel segmentation under multi-path supervision. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2019.
  52. 52.Chenyu You, Junlin Yang, Julius Chapiro, and James S. Duncan. Unsupervised wasserstein distance guided domain adaptation for 3d multi-domain liver segmentation. In Interpretable and Annotation-Efficient Learning for Medical Image Computing, pages 155–163. Springer International Publishing, 2020.
  53. 53.Linfeng Yang, Rajarshi P Ghosh, J Matthew Franklin, Simon Chen, Chenyu You, Raja R Narayan, Marc L Melcher, and Jan T Liphardt. Nuset: A deep learning tool for reliably separating and analyzing crowded cells. PLoS computational biology, 2020.
  54. 54.Shanlin Sun, Kun Han, Deying Kong, Chenyu You, and Xiaohui Xie. Mirnf: Medical image registration via neural fields. arXiv preprint arXiv:2206.03111, 2022.
  55. 55.Xiaoran Zhang, Chenyu You, Shawn Ahn, Juntang Zhuang, Lawrence Staib, and James Duncan. Learning correspondences of cardiac motion from images using biomechanics-informed modeling. arXiv preprint arXiv:2209.00726, 2022.
  56. 56.Chenyu You, Ruihan Zhao, Lawrence Staib, and James S Duncan. Momentum contrastive voxel-wise representation learning for semi-supervised volumetric medical image segmentation. arXiv preprint arXiv:2105.07059, 2021.
  57. 57.Chenyu You, Yuan Zhou, Ruihan Zhao, Lawrence Staib, and James S Duncan. Simcvd: Simple contrastive voxel-wise representation distillation for semi-supervised medical image segmentation. IEEE Transactions on Medical Imaging, 2022.
  58. 58.Chenyu You, Jinlin Xiang, Kun Su, Xiaoran Zhang, Siyuan Dong, John Onofrey, Lawrence Staib, and James S Duncan. Incremental learning meets transfer learning: Application to multi-site prostate mri segmentation. arXiv preprint arXiv:2206.01369, 2022.
  59. 59.Chenyu You, Weicheng Dai, Lawrence Staib, and James S Duncan. Bootstrapping semi-supervised medical image segmentation with anatomical-aware contrastive distillation. arXiv preprint arXiv:2206.02307, 2022.
  60. 60.Chenyu You, Weicheng Dai, Fenglin Liu, Haoran Su, Xiaoran Zhang, Lawrence Staib, and James S Duncan. Mine your own anatomy: Revisiting medical image segmentation with extremely limited labels. arXiv preprint arXiv:2209.13476, 2022.
  61. 61.Krishna Chaitanya, Ertunc Erdil, Neerav Karani, and Ender Konukoglu. Contrastive learning of global and local features for medical image segmentation with limited annotations. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  62. 62.Yuexing Han, Xiaolong Li, Bing Wang, and Lu Wang. Boundary loss-based 2.5 d fully convolutional neural networks approach for segmentation: A case study of the liver and tumor on computed tomography. Algorithms, 2021.
  63. 63.Konstantinos Kamnitsas, Christian Ledig, Virginia FJ Newcombe, Joanna P Simpson, Andrew D Kane, David K Menon, Daniel Rueckert, and Ben Glocker. Efficient multi-scale 3d cnn with fully connected crf for accurate brain lesion segmentation. Medical Image Analysis, 2017.
  64. 64.John Lafferty, Andrew McCallum, and Fernando CN Pereira. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In International Conference on Machine Learning (ICML), 2001.
  65. 65.Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath, Dong Yang, Andriy Myronenko, Bennett Landman, Holger Roth, and Daguang Xu. Unetr: Transformers for 3d medical image segmentation. arXiv preprint arXiv:2103.10504, 2021.
  66. 66.Fabian Isensee, Paul F Jaeger, Simon AA Kohl, Jens Petersen, and Klaus H Maier-Hein. nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 2021.
  67. 67.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. arXiv preprint arXiv:2012.15840, 2020.
  68. 68.Georg Hille, Shubham Agrawal, Christian Wybranski, Maciej Pech, Alexey Surov, and Sylvia Saalfeld. Joint liver and hepatic lesion segmentation using a hybrid cnn with transformer layers. arXiv preprint arXiv:2201.10981, 2022.
  69. 69.Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath, Dong Yang, Andriy Myronenko, Bennett Landman, Holger R Roth, and Daguang Xu. Unetr: Transformers for 3d medical image segmentation. In IEEE Winter Conference on Applications of Computer Vision (WACV), 2022.
  70. 70.Yucheng Tang, Dong Yang, Wenqi Li, Holger Roth, Bennett Landman, Daguang Xu, Vishwesh Nath, and Ali Hatamizadeh. Self-supervised pre-training of swin transformers for 3d medical image analysis. arXiv preprint arXiv:2111.14791, 2021.
  71. 71.Yifan Jiang, Shiyu Chang, and Zhangyang Wang. Transgan: Two transformers can make one strong gan. arXiv preprint arXiv:2102.07074, 2021.
  72. 72.Yu Zeng, Zhe Lin, and Vishal M Patel. Sketchedit: Mask-free local image manipulation with partial sketches. arXiv preprint arXiv:2111.15078, 2021.
  73. 73.Yu Zeng, Zhe Lin, Huchuan Lu, and Vishal M Patel. Cr-fill: Generative image inpainting with auxiliary contextual reconstruction. In IEEE International Conference on Computer Vision (ICCV), 2021.
  74. 74.Dor Arad Hudson and Larry Zitnick. Compositional transformers for scene generation. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  75. 75.Chenyu You, Guang Li, Yi Zhang, Xiaoliu Zhang, Hongming Shan, Mengzhou Li, Shenghong Ju, Zhen Zhao, Zhuiyang Zhang, Wenxiang Cong, et al. CT super-resolution GAN constrained by the identical, residual, and cycle learning ensemble (gan-circle). IEEE Transactions on Medical Imaging, 2019.
  76. 76.Long Zhao, Zizhao Zhang, Ting Chen, Dimitris N Metaxas, and Han Zhang. Improved transformer for high-resolution gans. arXiv preprint arXiv:2106.07631, 2021.
  77. 77.Yanhong Zeng, Huan Yang, Hongyang Chao, Jianbo Wang, and Jianlong Fu. Improving visual quality of image synthesis by a token-based generator with transformers. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  78. 78.Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes. arXiv preprint arXiv:2106.04803, 2021.
  79. 79.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
  80. 80.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In European Conference on Computer Vision (ECCV), 2018.
  81. 81.Chao Peng, Xiangyu Zhang, Gang Yu, Guiming Luo, and Jian Sun. Large kernel matters–improve semantic segmentation by global convolutional network. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  82. 82.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122, 2015.
  83. 83.Tete Xiao, Piotr Dollar, Mannat Singh, Eric Mintun, Trevor Darrell, and Ross Girshick. Early convolutions help transformers see better. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  84. 84.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In IEEE International Conference on Computer Vision (ICCV), 2017.
  85. 85.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  86. 86.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. arXiv preprint arXiv:2105.15203, 2021.
  87. 87.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems (NeurIPS), 2014.
  88. 88.Martin Arjovsky, Soumith Chintala, and Léon Bottou. Wasserstein generative adversarial networks. In International Conference on Machine Learning (ICML), 2017.
  89. 89.Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron Courville. Improved training of wasserstein gans. In Advances in Neural Information Processing Systems (NeurIPS), 2017.
  90. 90.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019.
  91. 91.Jo Schlemper, Ozan Oktay, Michiel Schaap, Mattias Heinrich, Bernhard Kainz, Ben Glocker, and Daniel Rueckert. Attention gated networks: Learning to leverage salient regions in medical images. Medical Image Analysis, 2019.
  92. 92.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.

Citation

MLA
You, C., et al. “Class-Aware Adversarial Transformers for Medical Image Segmentation”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 29582–96, https://proceedings.neurips.cc/paper_files/paper/2022/file/be99227ef4a4de84bb45d7dc7b53f808-Paper-Conference.pdf.
APA
You, C., Zhao, R., Liu, F., Dong, S., Chinchali, S., Topcu, U., Staib, L., & Duncan, J. (2022). Class-Aware Adversarial Transformers for Medical Image Segmentation. Advances in Neural Information Processing Systems, 35, 29582–29596. https://proceedings.neurips.cc/paper_files/paper/2022/file/be99227ef4a4de84bb45d7dc7b53f808-Paper-Conference.pdf
Chicago
You, C., R. Zhao, F. Liu, et al. 2022. “Class-Aware Adversarial Transformers for Medical Image Segmentation”. Advances in Neural Information Processing Systems 35: 29582–96. https://proceedings.neurips.cc/paper_files/paper/2022/file/be99227ef4a4de84bb45d7dc7b53f808-Paper-Conference.pdf.
Harvard
You, C. et al. (2022) “Class-Aware Adversarial Transformers for Medical Image Segmentation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 29582–29596. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/be99227ef4a4de84bb45d7dc7b53f808-Paper-Conference.pdf.
Vancouver
1. You C, Zhao R, Liu F, Dong S, Chinchali S, Topcu U, Staib L, Duncan J (2022) Class-Aware Adversarial Transformers for Medical Image Segmentation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 29582–29596

BibTeX

@inproceedings{you2022class,
  title = {Class-Aware Adversarial Transformers for Medical Image Segmentation},
  author = {You, Chenyu and Zhao, Ruihan and Liu, Fenglin and Dong, Siyuan and Chinchali, Sandeep and Topcu, Ufuk and Staib, Lawrence and Duncan, James},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {29582-29596},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/be99227ef4a4de84bb45d7dc7b53f808-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors