PraNet: Parallel Reverse Attention Network for Polyp Segmentation
Deng-Ping FanGe-Peng JiTao ZhouGeng ChenHuazhu FuJianbing ShenLing Shao
Introduces PraNet, an efficient deep network that resolves vague lesion boundaries in colonoscopy images by coupling high-level context aggregation with reverse attention modules to achieve accurate, real-time colorectal polyp segmentation.
Colorectal cancer is the third most common cancer globally, making early screening and lesion removal a vital public health priority. While colonoscopy is standard for detecting precancerous polyps, accurate automatic polyp segmentation remains difficult. Polyps exhibit significant variations in color, size, and texture, and their boundaries blend into surrounding mucosa, leading to missed detections and inaccurate boundary outlines.
The article evaluates a deep neural network named PraNet (Parallel Reverse Attention Network) to demonstrate how combining high-level contextual area mapping with reverse attention boundary refinement can achieve accurate, real-time polyp segmentation.
The approach mirrors clinical practice by first predicting a coarse polyp area and then progressively refining the boundary. A parallel partial decoder aggregates high-level semantic features to produce a global guidance map, while recurrent reverse attention modules subtract estimated foreground regions from side-output features to capture subtle boundary cues. Evaluated across five benchmark colonoscopy datasets (ETIS, CVC-ClinicDB, CVC-ColonDB, EndoScene, and Kvasir), the model was trained in an end-to-end setting with weighted intersection-over-union and binary cross-entropy loss functions.
The evaluation produced four key findings. First, PraNet demonstrated superior learning accuracy, achieving an 89.8% to 89.9% mean Dice score on seen datasets, outperforming established architectures like U-Net, U-Net++, and selective feature aggregation networks by more than 7%. Second, the architecture showed strong generalization across unseen datasets, achieving a 70.9% mean Dice score on CVC-ColonDB and 62.8% on ETIS, where baseline methods dropped to 29.7%–51.2%. Third, PraNet achieved real-time inference at approximately 50 frames per second on standard hardware. Fourth, training was exceptionally efficient, converging in only 20 epochs (about 30 minutes) compared to over 20 hours required by competing edge-aggregation models.
These findings indicate that PraNet can enhance computer-aided colonoscopy tools by providing real-time, highly accurate guidance during live procedures without requiring complex pre- or post-processing. Because it effectively separates subtle polyp boundaries without overfitting, it reduces computational costs and lowers clinical risk associated with missed polyps. These results also show that using reverse attention to implicitly refine boundaries outperforms heavy, explicit edge-supervision models in both speed and robustness.
For operational deployment, stakeholders should evaluate integrating PraNet into real-time clinical video streams and validate performance in live procedural pilots. Researchers should also extend this reverse attention framework to other complex medical imaging tasks, such as lung infection segmentation.
Confidence in these findings is strong across the evaluated benchmarks; however, limitations remain regarding performance degradation on highly constrained datasets (such as the 62.8% Dice score on the small ETIS dataset), meaning additional diverse clinical data is recommended before widespread adoption.
- Paper: Kvasir-SEG: A Segmented Polyp Dataset, Debesh Jha et al. (2019). It introduces Kvasir-SEG, the primary standardized benchmark dataset and evaluation suite utilized by PraNet for training and testing polyp segmentation models.
- Paper: U-Net: Convolutional Networks for Biomedical Image Segmentation, Olaf Ronneberger et al. (2015). It establishes the foundational encoder-decoder architecture with skip connections that forms the baseline paradigm for medical image segmentation across the field.
- Paper: Res2Net: A New Multi-Scale Backbone Architecture, Shanghua Gao et al. (2019). It introduces the multi-scale hierarchical feature extraction backbone (Res2Net) that PraNet adopts to aggregate granular receptive fields across layers.
- Paper: UNet++: A Nested U-Net Architecture for Medical Image Segmentation, Zongwei Zhou et al. (2018). It redesigns medical image segmentation through nested skip connections and deep supervision, providing key architectural baselines and motivations for PraNet's partial decoder design.
- Paper: Attention U-Net: Learning Where to Look for the Pancreas, Ozan Oktay et al. (2018). It demonstrates how incorporating attention mechanisms into medical segmentation decoders suppresses background noise and highlights salient anatomical regions.
- Paper: MultiResUNet : Rethinking the U-Net Architecture for Multimodal Biomedical Image Segmentation, Nabil Ibtehaz et al. (2019). It addresses scale variation and boundary representation discrepancies in biomedical segmentation, serving as a direct comparative baseline for polyp segmentation.
- Paper: CE-Net: Context Encoder Network for 2D Medical Image Segmentation, Zaiwang Gu et al. (2019). It develops multi-scale context capture modules for 2D medical imaging, establishing techniques for preserving spatial details and boundary cues.
- Paper: Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation, Hu Cao et al. (2021). It advances beyond convolutional reverse-attention pipelines like PraNet by formulating medical image segmentation entirely around pure Vision Transformers.
- Paper: Masked-attention Mask Transformer for Universal Image Segmentation, Bowen Cheng et al. (2022). It unifies image segmentation paradigms via masked-attention transformers, offering a broader structural evolution from task-specific boundary calibration networks.
- Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). It replaces traditional per-pixel classification decoders with universal mask classification architectures that generalize boundary and region prediction across vision domains.
- Paper: Segmenter: Transformer for Semantic Segmentation, Robin Strudel et al. (2021). It demonstrates how pure transformer patch-based models can replace specialized convolutional decoding mechanisms for semantic segmentation.
