ResUNet++: An Advanced Architecture for Medical Image Segmentation

Debesh JhaPia H. SmedsrudMichael A. RieglerDag JohansenThomas de LangePal HalvorsenHavard D. Johansen

article2019IEEE International Symposium on Multimedia1,384 citations

Introduces ResUNet++, an advanced deep learning architecture that improves pixel-wise polyp segmentation in colonoscopy images and significantly outperforms standard U-Net and ResUNet models on public benchmarks.

Listen

Colorectal cancer is a leading cause of cancer-related mortality worldwide, but early detection and removal of precancerous polyps during colonoscopy examinations significantly reduces patient risk. However, clinicians still miss a notable fraction of polyps due to variations in polyp appearance and visual interference within the bowel. Automated computer-aided detection systems can act as a reliable second observer during procedures, but they require accurate, pixel-level segmentation to clearly delineate lesion boundaries rather than merely flagging anomalous frames.

The article introduces and evaluates ResUNet++, an advanced deep learning architecture designed for automated, pixel-wise polyp segmentation in colonoscopic imagery. The primary objective is to demonstrate that integrating residual units, channel-wise feature calibration, multi-scale contextual pooling, and attention mechanisms provides superior segmentation accuracy and generalizability compared to established baseline models.

To evaluate the system, the authors conducted experiments using two publicly available benchmark datasets: Kvasir-SEG (1,000 expert-annotated images) and CVC-ClinicDB (612 images from 31 colonoscopy sequences). The models were trained and tested using standardized data augmentation and evaluation protocols on high-performance computing hardware. ResUNet++ was directly benchmarked against widely used biomedical segmentation baselines, specifically standard U-Net and standard ResUNet, as well as a modified ResUNet optimized with a dice coefficient loss function.

Across both benchmarks, ResUNet++ established state-of-the-art performance. On the Kvasir-SEG dataset, ResUNet++ achieved a mean Intersection over Union (mIoU) of 79.27% and an overlap dice coefficient of 81.33%, outperforming the baseline U-Net (43.34% mIoU, 71.47% dice) and baseline ResUNet (43.64% mIoU, 51.44% dice). When evaluated on the CVC-ClinicDB dataset to assess generalizability, ResUNet++ maintained high accuracy with a 79.62% mIoU and 79.55% dice coefficient, whereas U-Net and ResUNet attained mIoUs of only 47.11% and 45.70%, respectively. ResUNet++ also demonstrated the highest recall across both datasets (above 70%), while maintaining strong precision (nearly 88%). Qualitative visual assessments confirmed that the predicted segmentations match ground-truth polyp boundaries much more closely than competing architectures.

These results indicate that ResUNet++ effectively captures complex polyp shapes and multi-scale visual details even when training on relatively modest dataset sizes. By substantially improving boundary detection and true positive identification (recall), such a model lowers the clinical risk of missed lesions and provides clearer visual guidance to endoscopists. In addition, the release of the expert-annotated Kvasir-SEG dataset addresses a major industry bottleneck in reproducible medical AI development.

For future development, the article recommends exploring post-processing techniques, expanding training data volume, and testing the architecture on broader medical and natural image segmentation tasks. Leaders should note, however, that ResUNet++ incorporates more model parameters, which increases computational training requirements. Furthermore, because the evaluation relies on fixed image resizing (256x256 pixels) and offline datasets, further clinical validation in live, real-time video streaming environments is necessary before operational deployment.

arXiv: 1911.07067
Cover for ResUNet++: An Advanced Architecture for Medical Image Segmentation

Abstract

Accurate computer-aided polyp detection and segmentation during colonoscopy examinations can help endoscopists resect abnormal tissue and thereby decrease chances of polyps growing into cancer. Towards developing a fully automated model for pixel-wise polyp segmentation, we propose ResUNet++, which is an improved ResUNet architecture for colonoscopic image segmentation. Our experimental evaluations show that the suggested architecture produces good segmentation results on publicly available datasets. Furthermore, ResUNet++ significantly outperforms U-Net and ResUNet, two key state-of-the-art deep learning architectures, by achieving high evaluation scores with a dice coefficient of 81.33%, and a mean Intersection over Union (mIoU) of 79.27% for the Kvasir-SEG dataset and a dice coefficient of 79.55%, and a mIoU of 79.62% with CVC-612 dataset.

Table of Contents

  • I Introduction
  • II Related Work
  • III ResUNet++
  • III-A Residual Units
  • III-B Squeeze and Excitation Units
  • III-C Atrous Spatial Pyramidal Pooling
  • III-D Attention Units
  • IV Experiments
  • IV-A Datasets
  • IV-B Implementation details
  • V Results
  • V-A Results on the Kvasir-SEG dataset
  • V-B Results on the CVC-612 dataset
  • VI Discussion
  • VII Conclusion
  • References

Knowls

  1. Knowl 1 — ResUNet++ Architecture for Medical Image Segmentation

    model/method

    ResUNet++ is an encoder-decoder deep neural network designed for semantic segmentation of medical images, such as colorectal polyps in colonoscopy frames. The architecture builds upon Deep Residual U-Net (ResUNet) by integrating pre-activation residual units, Squeeze-and-Excitation (SE) blocks, Atrous Spatial Pyramidal Pooling (ASPP), and attention blocks.

    The overall computational path proceeds as follows:

    1. Stem Block: An initial convolutional block consisting of a 3×33 \times 3 convolution, Batch Normalization (BN), Rectified Linear Unit (ReLU) activation, a second 3×33 \times 3 convolution, and an identity shortcut addition.

    2. Encoding Path: Three consecutive encoder blocks. Each encoder block contains two 3×33 \times 3 convolutional operations with BN and ReLU, linked by an identity mapping shortcut. Spatial downsampling by a factor of 2 is performed using a strided convolution at the first layer of each encoder block. The residual output of each encoder block is immediately passed through a Squeeze-and-Excitation (SE) unit to adaptively recalibrate channel-wise feature dependencies.

    3. Bridge: An ASPP module acts as a bottleneck bridge between the encoder and decoder. It fuses parallel dilated convolutions at different atrous rates to capture multi-scale contextual information with an expanded receptive field.

    4. Decoding Path: Three decoder stages. In each stage, the skip connection feature maps from the corresponding encoder level are modulated by an Attention Block to suppress irrelevant features and emphasize polyp regions. Simultaneously, the decoder feature maps from the preceding lower level are upsampled via nearest-neighbor interpolation and concatenated with the attention-weighted encoder features. The concatenated representation is then processed through a residual block (two 3×33 \times 3 convolutions with BN and ReLU, combined via identity addition).

    5. Output Head: The output of the final decoder block passes through a second ASPP block, followed by a 1×11 \times 1 convolution with sigmoid activation to generate the final pixel-wise binary segmentation mask.

  2. Knowl 2 — Building Blocks of ResUNet++: Squeeze-and-Excitation, ASPP, and Attention Units

    model/method

    The ResUNet++ architecture integrates three core functional modules into the residual U-Net backbone:

    1. Squeeze-and-Excitation (SE) Unit: Placed after each residual encoder block to explicitly model channel inter-dependencies. The squeeze step aggregates global spatial context into a channel descriptor using Global Average Pooling. The excitation step computes channel-specific modulation weights via a gating mechanism (two fully connected layers with non-linear activations) to reweight feature responses, suppressing background noise and amplifying relevant feature channels.

    2. Atrous Spatial Pyramidal Pooling (ASPP): Utilized both as the bridge between encoder and decoder, and directly before the final output convolution. ASPP applies multiple parallel atrous (dilated) convolutions with varying sampling rates alongside global feature pooling. This multi-scale sampling captures contextual information across different spatial receptive fields without reducing feature map resolution.

    3. Attention Units: Implemented along the skip connections in the decoder path before feature concatenation. The attention mechanism calculates spatial gating signals from the decoder features to selectively weight the encoder feature maps, ensuring that the network highlights target region boundaries (e.g., polyps) and discards redundant background information.

  3. Knowl 3 — Kvasir-SEG Dataset for Polyp Segmentation

    definition

    The Kvasir-SEG dataset is an open-access medical image segmentation benchmark consisting of 1,000 gastrointestinal polyp images extracted from colonoscopy procedures, paired with corresponding pixel-level ground truth binary segmentation masks. Annotations were generated and validated in collaboration with clinical gastroenterologists from Oslo University Hospital (Norway). The dataset includes polyps from multiple diagnostic classes (adenoma, serrated, hyperplastic, and rare mixed polyps), spanning diverse shapes, scales, colors, textures, and degrees of background occlusion (such as mucosal similarity and stool coverage).

  4. Knowl 4 — Experimental Training Pipeline and Hyperparameter Configuration

    experimental setup

    The training and evaluation protocol for ResUNet++ and baseline segmentation models is structured as follows:

    • Data Preprocessing: Input images are cropped with a margin of 320×320320 \times 320 pixels to enhance dataset diversity, then resized to a uniform resolution of 256×256256 \times 256 pixels.
    • Data Augmentation: Techniques applied during training include center cropping, random cropping, horizontal flipping, vertical flipping, scale augmentation, random rotations uniformly sampled from 0∘0^\circ to 90∘90^\circ, cutout, and brightness adjustments.
    • Data Splits: Datasets are partitioned into 80% for training, 10% for validation, and 10% for testing.
    • Optimization and Schedule: Models are optimized using the Adam optimizer with an initial learning rate of 1×10−41 \times 10^{-4} alongside Stochastic Gradient Descent with Restarts (SGDR), trained for 120 epochs with a mini-batch size of 16.
    • Loss Function: Dice coefficient loss is utilized for training as it empirically provides superior Intersection over Union (mIoU) and boundary alignment compared to Mean Squared Error (MSE), Binary Cross-Entropy (BCE), or composite BCE-Dice losses.
  5. Knowl 5 — Polyp Segmentation Results on the Kvasir-SEG Dataset

    data/table

    Performance of ResUNet++ evaluated against U-Net, original ResUNet (trained with MSE loss), and ResUNet-mod (ResUNet with Dice loss and hyperparameter optimization) on the test split of the Kvasir-SEG dataset:

    Method Dice mIoU Recall Precision
    ResUNet++ 0.8133 0.7927 0.7064 0.8774
    ResUNet-mod 0.7909 0.4287 0.6909 0.8713
    ResUNet 0.5144 0.4364 0.5041 0.7292
    U-Net 0.7147 0.4334 0.6306 0.9222

    ResUNet++ achieves the highest Dice coefficient (81.33%), mean Intersection over Union (mIoU of 79.27%), and Recall (70.64%). While standard U-Net attains the highest Precision (92.22%), its mIoU (43.34%) and Dice (71.47%) remain significantly lower. ResUNet++ outperforms all baselines in mIoU by over 35 percentage points.

  6. Knowl 6 — Generalization Evaluation on the CVC-ClinicDB (CVC-612) Dataset

    data/table

    Generalization performance of ResUNet++ compared with baseline architectures on the CVC-ClinicDB (CVC-612) dataset, which contains 612 colonoscopy frames from 31 sequences:

    Method Dice mIoU Recall Precision
    ResUNet++ 0.7955 0.7962 0.7022 0.8785
    ResUNet-mod 0.7788 0.4545 0.6683 0.8877
    ResUNet 0.4510 0.4570 0.5775 0.5614
    U-Net 0.6419 0.4711 0.6756 0.6868

    ResUNet++ achieves the top Dice score (79.55%), mIoU (79.62%), and Recall (70.22%) on CVC-612, demonstrating robust cross-dataset generalizability and significantly exceeding the mIoU of U-Net (47.11%) and ResUNet-mod (45.45%).

  7. Knowl 7 — Influence of Loss Function Choice on Segmentation mIoU and Dice Score

    empirical result

    The original ResUNet implementation utilizes Mean Squared Error (MSE) loss, which yields inadequate segmentation performance on polyp datasets (Dice of 0.5144 on Kvasir-SEG and 0.4510 on CVC-612). Modifying the objective to Dice coefficient loss along with hyperparameter optimization (yielding ResUNet-mod) improves the Dice score substantially to 0.7909 on Kvasir-SEG and 0.7788 on CVC-612.

    However, while Binary Cross-Entropy (BCE), combined BCE-Dice loss, and MSE achieve comparable Dice scores under certain configurations, only the pure Dice coefficient loss combined with the full ResUNet++ architecture (residual units, SE blocks, ASPP, and attention units) produces a high mean Intersection over Union (mIoU ≈0.79\approx 0.79), whereas all other baseline models and configurations remain below 0.480.48 mIoU.

  8. Knowl 8 — Limitations of ResUNet++

    limitation

    The ResUNet++ model has the following identified limitations:

    1. Model Complexity and Training Cost: The inclusion of SE blocks, dual ASPP modules, and attention units increases parameter count relative to standard U-Net and ResUNet, resulting in longer training times.
    2. Information Loss via Image Resizing: Downsampling native colonoscopy frames to 256×256256 \times 256 pixels to maintain tractable GPU memory usage and training speed causes loss of fine-grained spatial and edge details of tiny or flat polyps.
    3. Absence of Post-Processing: The network operates purely end-to-end without auxiliary boundary refinement techniques (such as Conditional Random Fields or morphological filtering) that could further improve mask delineation.

Coverage note — None was omitted; all substantive contributions (architecture components, Kvasir-SEG dataset introduction, training setup, evaluation tables on both datasets, loss function analysis, and limitations) are fully captured.

References

  1. 1.A. G. Zauber, S. J. Winawer, M. J. O’Brien, I. Lansdorp-Vogelaar, M. van Ballegooijen, B. F. Hankey, W. Shi, J. H. Bond, M. Schapiro, J. F. Panish et al., “Colonoscopic polypectomy and long-term prevention of colorectal-cancer deaths,” New England Journal of Medicine, vol. 366, no. 8, pp. 687–696, 2012.
  2. 2.J. C. Van Rijn, J. B. Reitsma, J. Stoker, P. M. Bossuyt, S. J. Van Deventer, and E. Dekker, “Polyp miss rate determined by tandem colonoscopy: a systematic review,” The American journal of gastroenterology, vol. 101, no. 2, p. 343, 2006.
  3. 3.Y. Mori and S.-e. Kudo, “Detecting colorectal polyps via machine learning,” Nature biomedical engineering, vol. 2, no. 10, p. 713, 2018.
  4. 4.F. Milletari, N. Navab, and S.-A. Ahmadi, “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” in Proceeding of International Conference on 3D Vision (3DV). IEEE, 2016, pp. 565–571.
  5. 5.O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Proceedings of International Conference on Medical image computing and computer-assisted intervention. Springer, 2015, pp. 234–241.
  6. 6.Z. Zhang, Q. Liu, and Y. Wang, “Road extraction by deep residual u-net,” IEEE Geoscience and Remote Sensing Letters, vol. 15, no. 5, pp. 749–753, 2018.
  7. 7.K. Pogorelov, K. R. Randel, C. Griwodz, S. L. Eskeland, T. de Lange, D. Johansen, C. Spampinato, D.-T. Dang-Nguyen, M. Lux, P. T. Schmidt, M. Riegler, and P. Halvorsen, “Kvasir: A multi-class image dataset for computer aided gastrointestinal disease detection,” in Proc. of MMSYS, june 2017, pp. 164–169.
  8. 8.D. Jha, P. H. Smedsrud, M. Riegler, P. Halvorsen, T. de Lange, D. Johansen, and H. Johansen, “Kvasir-seg: A segmented polyp dataset,” in International Conference on Multimedia Modeling. Springer, 2020. [Online]. Available: https://datasets.simula.no/kvasir-seg/
  9. 9.Y. Wang, W. Tavanapong, J. Wong, J. Oh, and P. C. De Groen, “Part-based multiderivative edge cross-sectional profiles for polyp detection in colonoscopy,” IEEE Journal of Biomedical and Health Informatics, vol. 18, no. 4, pp. 1379–1389, 2014.
  10. 10.Y. Mori, S.-e. Kudo, T. M. Berzin, M. Misawa, and K. Takeda, “Computer-aided diagnosis for colonoscopy,” Endoscopy, vol. 49, no. 8, pp. 813–819, 2017.
  11. 11.P. Brandao, O. Zisimopoulos, E. Mazomenos, G. Ciuti, J. Bernal, M. Visentini-Scarzanella, A. Menciassi, P. Dario, A. Koulaouzidis, A. Arezzo et al., “Towards a computed-aided diagnosis system in colonoscopy: automatic polyp segmentation using convolution neural networks,” Journal of Medical Robotics Research, vol. 3, no. 2, p. 1840002, 2018.
  12. 12.P. Wang, X. Xiao, J. R. G. Brown, T. M. Berzin, M. Tu, F. Xiong, X. Hu, P. Liu, Y. Song, D. Zhang et al., “Development and validation of a deep-learning algorithm for the detection of polyps during colonoscopy,” Nature biomedical engineering, vol. 2, no. 10, pp. 741–748, 2018.
  13. 13.Y. Wang, W. Tavanapong, J. Wong, J. H. Oh, and P. C. De Groen, “Polyp-alert: Near real-time feedback during colonoscopy,” International Journal of Computer methods and programs in biomedicine, vol. 120, no. 3, pp. 164–179, 2015.
  14. 14.M. Riegler, K. Pogorelov, S. L. Eskeland, P. T. Schmidt, Z. Albisser, D. Johansen, C. Griwodz, P. Halvorsen, and T. D. Lange, “From annotation to computer-aided diagnosis: Detailed evaluation of a medical multimedia system,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 13, no. 3, p. 26, 2017.
  15. 15.S. A. Hicks, S. Eskeland, M. Lux, T. de Lange, K. R. Randel, M. Jeppsson, K. Pogorelov, P. Halvorsen, and M. Riegler, “Mimir: an automatic reporting and reasoning system for deep learning based analysis in the medical domain,” in Proceedings of the ACM Multimedia Systems Conference. ACM, 2018, pp. 369–374.
  16. 16.V. Thambawita, D. Jha, M. Riegler, P. Halvorsen, H. L. Hammer, H. D. Johansen, and D. Johansen, “The medico-task 2018: Disease detection in the gastrointestinal tract using global features and deep learning,” in Working Notes Proceedings of the MediaEval Workshop. CEUR Workshop Proceedings, 2018.
  17. 17.Y. B. Guo and B. Matuszewski, “Giana polyp segmentation with fully convolutional dilation neural networks,” in Proceedings of International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications. SCITEPRESS-Science and Technology Publications, 2019, pp. 632–641.
  18. 18.J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proceedings of IEEE conference on computer vision and pattern recognition (CVPR), 2015, pp. 3431–3440.
  19. 19.O. Çiçek, A. Abdulkadir, S. S. Lienkamp, T. Brox, and O. Ronneberger, “3d u-net: learning dense volumetric segmentation from sparse annotation,” in Proceeding of International conference on medical image computing and computer-assisted intervention. Springer, 2016, pp. 424–432.
  20. 20.M. Drozdzal, E. Vorontsov, G. Chartrand, S. Kadoury, and C. Pal, “The importance of skip connections in biomedical image segmentation,” in Deep Learning and Data Labeling for Medical Applications. Springer, 2016, pp. 179–187.
  21. 21.F. I. Diakogiannis, F. Waldner, P. Caccetta, and C. Wu, “Resunet-a: a deep learning framework for semantic segmentation of remotely sensed data,” arXiv preprint arXiv:1904.00592, 2019.
  22. 22.Z. Zhou, M. M. R. Siddiquee, N. Tajbakhsh, and J. Liang, “Unet++: A nested u-net architecture for medical image segmentation,” in Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support. Springer, 2018, pp. 3–11.
  23. 23.Y. Wang, W. Tavanapong, J. Wong, J. Oh, and P. C. De Groen, “Part-based multiderivative edge cross-sectional profiles for polyp detection in colonoscopy,” IEEE Journal of Biomedical and Health Informatics, vol. 18, no. 4, pp. 1379–1389, 2013.
  24. 24.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of IEEE conference on computer vision and pattern recognition (CVPR), 2016, pp. 770–778.
  25. 25.J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of IEEE conference on computer vision and pattern recognition (CVPR), 2018, pp. 7132–7141.
  26. 26.K. He., X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” IEEE transactions on pattern analysis and machine intelligence, vol. 37, no. 9, pp. 1904–1916, 2015.
  27. 27.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 4, pp. 834–848, 2018.
  28. 28.L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking atrous convolution for semantic image segmentation,” arXiv preprint arXiv:1706.05587, 2017.
  29. 29.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 4, pp. 834–848, 2017.
  30. 30.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems, 2017, pp. 5998–6008.
  31. 31.H. Li, P. Xiong, J. An, and L. Wang, “Pyramid attention network for semantic segmentation,” arXiv preprint arXiv:1805.10180, 2018.
  32. 32.J. Bernal, F. J. Sánchez, G. Fernández-Esparrach, D. Gil, C. Rodríguez, and F. Vilariño, “Wm-dova maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians,” Computerized Medical Imaging and Graphics, vol. 43, pp. 99–111, 2015.
  33. 33.F. Chollet et al., “Keras,” 2015.
  34. 34.M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al., “Tensorflow: A system for large-scale machine learning,” in Proceeding of {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI}), 2016, pp. 265–283.

Citation

MLA
Jha, D., et al. “ResUNet++: An Advanced Architecture for Medical Image Segmentation”. arXiv, 2019, http://arxiv.org/abs/1911.07067v1.
APA
Jha, D., Smedsrud, P. H., Riegler, M. A., Johansen, D., Lange, T. de ., Halvorsen, P., & Johansen, H. D. (2019). ResUNet++: An Advanced Architecture for Medical Image Segmentation. arXiv. http://arxiv.org/abs/1911.07067v1
Chicago
Jha, D., P. H. Smedsrud, M. A. Riegler, et al. 2019. “ResUNet++: An Advanced Architecture for Medical Image Segmentation”. arXiv. http://arxiv.org/abs/1911.07067v1.
Harvard
Jha, D. et al. (2019) “ResUNet++: An Advanced Architecture for Medical Image Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1911.07067v1.
Vancouver
1. Jha D, Smedsrud PH, Riegler MA, Johansen D, Lange T de, Halvorsen P, Johansen HD (2019) ResUNet++: An Advanced Architecture for Medical Image Segmentation. arXiv

BibTeX

@article{jha2019resunet,
  title = {ResUNet++: An Advanced Architecture for Medical Image Segmentation},
  author = {Jha, Debesh and Smedsrud, Pia H. and Riegler, Michael A. and Johansen, Dag and Lange, Thomas de and Halvorsen, Pal and Johansen, Havard D.},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1911.07067v1},
  eprint = {1911.07067}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF