ResUNet++: An Advanced Architecture for Medical Image Segmentation
Debesh JhaPia H. SmedsrudMichael A. RieglerDag JohansenThomas de LangePal HalvorsenHavard D. Johansen
Introduces ResUNet++, an advanced deep learning architecture that improves pixel-wise polyp segmentation in colonoscopy images and significantly outperforms standard U-Net and ResUNet models on public benchmarks.
Colorectal cancer is a leading cause of cancer-related mortality worldwide, but early detection and removal of precancerous polyps during colonoscopy examinations significantly reduces patient risk. However, clinicians still miss a notable fraction of polyps due to variations in polyp appearance and visual interference within the bowel. Automated computer-aided detection systems can act as a reliable second observer during procedures, but they require accurate, pixel-level segmentation to clearly delineate lesion boundaries rather than merely flagging anomalous frames.
The article introduces and evaluates ResUNet++, an advanced deep learning architecture designed for automated, pixel-wise polyp segmentation in colonoscopic imagery. The primary objective is to demonstrate that integrating residual units, channel-wise feature calibration, multi-scale contextual pooling, and attention mechanisms provides superior segmentation accuracy and generalizability compared to established baseline models.
To evaluate the system, the authors conducted experiments using two publicly available benchmark datasets: Kvasir-SEG (1,000 expert-annotated images) and CVC-ClinicDB (612 images from 31 colonoscopy sequences). The models were trained and tested using standardized data augmentation and evaluation protocols on high-performance computing hardware. ResUNet++ was directly benchmarked against widely used biomedical segmentation baselines, specifically standard U-Net and standard ResUNet, as well as a modified ResUNet optimized with a dice coefficient loss function.
Across both benchmarks, ResUNet++ established state-of-the-art performance. On the Kvasir-SEG dataset, ResUNet++ achieved a mean Intersection over Union (mIoU) of 79.27% and an overlap dice coefficient of 81.33%, outperforming the baseline U-Net (43.34% mIoU, 71.47% dice) and baseline ResUNet (43.64% mIoU, 51.44% dice). When evaluated on the CVC-ClinicDB dataset to assess generalizability, ResUNet++ maintained high accuracy with a 79.62% mIoU and 79.55% dice coefficient, whereas U-Net and ResUNet attained mIoUs of only 47.11% and 45.70%, respectively. ResUNet++ also demonstrated the highest recall across both datasets (above 70%), while maintaining strong precision (nearly 88%). Qualitative visual assessments confirmed that the predicted segmentations match ground-truth polyp boundaries much more closely than competing architectures.
These results indicate that ResUNet++ effectively captures complex polyp shapes and multi-scale visual details even when training on relatively modest dataset sizes. By substantially improving boundary detection and true positive identification (recall), such a model lowers the clinical risk of missed lesions and provides clearer visual guidance to endoscopists. In addition, the release of the expert-annotated Kvasir-SEG dataset addresses a major industry bottleneck in reproducible medical AI development.
For future development, the article recommends exploring post-processing techniques, expanding training data volume, and testing the architecture on broader medical and natural image segmentation tasks. Leaders should note, however, that ResUNet++ incorporates more model parameters, which increases computational training requirements. Furthermore, because the evaluation relies on fixed image resizing (256x256 pixels) and offline datasets, further clinical validation in live, real-time video streaming environments is necessary before operational deployment.
- Paper: U-Net: Convolutional Networks for Biomedical Image Segmentation, Olaf Ronneberger et al. (2015). It introduces the foundational U-Net encoder-decoder architecture with skip connections that ResUNet++ directly modifies and enhances.
- Paper: Road Extraction by Deep Residual U-Net, Zhengxin Zhang et al. (2017). It establishes the Deep Residual U-Net (ResUNet) framework that integrates residual units into U-Net, serving as the immediate predecessor and baseline for ResUNet++.
- Paper: Attention U-Net: Learning Where to Look for the Pancreas, Ozan Oktay et al. (2018). It introduces attention gating mechanisms within U-Net skip connections to focus on target anatomical structures, a core concept incorporated into ResUNet++.
- Paper: UNet++: A Nested U-Net Architecture for Medical Image Segmentation, Zongwei Zhou et al. (2018). It explores nested and dense skip connections with deep supervision in medical image segmentation, providing key context for multi-scale feature aggregation techniques used in ResUNet++.
- Paper: Rethinking Atrous Convolution for Semantic Image Segmentation, Liang-Chieh Chen et al. (2017). It formalizes Atrous Spatial Pyramid Pooling (ASPP) to capture multi-scale context without loss of resolution, which is adapted as a core module in ResUNet++.
- Paper: Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images, Jo Schlemper et al. (2018). It details the theory and architecture of soft-attention gated networks for medical imaging that guide the attention mechanism design in ResUNet++.
- Paper: Dilated Residual Networks, Fisher Yu et al. (2017). It details how dilated convolutions can preserve high spatial resolution in deep residual networks, informing the design of receptive field modules in advanced ResUNet architectures.
- Paper: PraNet: Parallel Reverse Attention Network for Polyp Segmentation, Deng-Ping Fan et al. (2020). It advances colonoscopy polyp segmentation by introducing parallel reverse attention and area boundary refinement modules tested on the same benchmark datasets evaluated in ResUNet++.
- Paper: Kvasir-SEG: A Segmented Polyp Dataset, Debesh Jha et al. (2019). It presents the official release, comprehensive benchmark, and ground truth annotations for the Kvasir-SEG polyp dataset used to evaluate ResUNet++.
- Paper: TransFuse: Fusing Transformers and CNNs for Medical Image Segmentation, Yundong Zhang et al. (2021). It extends medical image and polyp segmentation by fusing shallow convolutional networks with transformer branches to capture both local details and global context.
- Paper: UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation, Huimin Huang et al. (2020). It furthers the evolution of medical image segmentation architectures by proposing full-scale skip connections and deep supervision to capture multi-scale feature interactions.
- Paper: Medical Transformer: Gated Axial-Attention for Medical Image Segmentation, Jeya Maria Jose Valanarasu et al. (2021). It introduces gated axial-attention transformers tailored for medical image segmentation on small-sample datasets, advancing beyond purely convolutional segmentation backbones.
- Paper: Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation, Hu Cao et al. (2021). It transitions U-Net-based medical segmentation to a pure vision transformer encoder-decoder design featuring shifted windows and patch expansion.
- Paper: Segment anything in medical images, Jun Ma et al. (2023). It scales medical image segmentation from specialized convolutional models to a foundational vision model capable of prompt-based segmentation across diverse clinical modalities.
