Kvasir-SEG: A Segmented Polyp Dataset
Debesh JhaPia H. SmedsrudMichael A. RieglerPål HalvorsenThomas de LangeDag JohansenHåvard D. Johansen
Introduces an open-access dataset of gastrointestinal polyp images paired with expert-verified pixel-level segmentation masks and bounding boxes to provide a standardized benchmark for evaluating automated colonoscopy analysis models.
Colorectal cancer is one of the most common cancers worldwide, making early detection and removal of precursor polyps critical for patient survival. However, colonoscopy screenings suffer from polyp miss rates between 14% and 30%, largely due to human error and visual complexity. Developing automated, computer-assisted diagnostic tools can significantly improve polyp detection and patient outcomes, but progress has been bottlenecked by the lack of publicly accessible, high-quality medical image datasets with pixel-level ground truth annotations.
The article introduces Kvasir-SEG, a newly released open-access dataset designed for pixel-wise polyp segmentation, and establishes baseline performance benchmarks using both traditional and deep-learning computational methods.
To build the dataset, researchers annotated 1,000 gastrointestinal polyp images from the existing Kvasir collection. A medical doctor and an engineer manually outlined polyp boundaries, which were subsequently validated by an experienced gastroenterologist to produce exact segmentation masks and bounding boxes. The authors evaluated two segmentation approaches on an 80/10/10 training, validation, and test split: an unsupervised baseline using Fuzzy C-means clustering with standard image filtering, and a modern deep-learning approach using a Deep Residual U-Net (ResUNet) architecture enhanced with data augmentation.
The experimental results demonstrated a substantial performance difference between the evaluated methods. The ResUNet deep-learning model achieved strong predictive capability on unseen test data, scoring a Dice coefficient of approximately 0.79 and a mean Intersection over Union of 0.78. In contrast, the unsupervised Fuzzy C-means method performed poorly, achieving a Dice coefficient of only 0.24 and an Intersection over Union of 0.31. ResUNet converged effectively within 91 training epochs and showed competitive or superior accuracy compared to benchmarks reported in prior literature.
These findings confirm that traditional color- and threshold-based segmentation algorithms fail because polyps share visual and color characteristics with surrounding healthy mucosal tissue. In contrast, modern convolutional architectures utilizing residual learning and data augmentation can accurately learn complex spatial patterns even from relatively small medical datasets. By providing verified, pixel-level masks alongside standard evaluation metrics, the open-access release enables reliable benchmarking and accelerates the development of cost-effective, real-time diagnostic aids for clinical endoscopy.
To advance toward clinical adoption, research teams should utilize the open-source Kvasir-SEG benchmark to develop and compare next-generation segmentation architectures. Prior to deploying such models in live hospital workflows, stakeholders must expand training on more diverse image sets, conduct prospective clinical trials, and evaluate real-time processing capabilities under operational colonoscopy conditions.
The primary limitation of the study is the dataset scale, which is constrained to 1,000 images from a single primary clinical source and may not capture every anatomical or pathological variation encountered in diverse patient populations. While confidence is high that deep-learning models like ResUNet clearly outperform traditional segmentation methods, the authors caution that current baseline performance must be improved and further validated before these algorithms can be directly deployed for live patient care.
- Paper: U-Net: Convolutional Networks for Biomedical Image Segmentation, Olaf Ronneberger et al. (2015). Introduces the seminal U-Net architecture that serves as the primary modern deep learning benchmark baseline evaluated on the Kvasir-SEG dataset.
- Paper: UNet++: A Nested U-Net Architecture for Medical Image Segmentation, Zongwei Zhou et al. (2018). Presents an advanced nested encoder-decoder architecture for medical image and polyp segmentation that directly motivates dense mask prediction benchmarks in endoscopy.
- Paper: Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?, Nima Tajbakhsh et al. (2016). Investigates deep CNN training and transfer learning specifically for endoscopic polyp analysis, providing foundational context for benchmarking deep models on colonoscopy data.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). Establishes end-to-end fully convolutional networks for semantic segmentation, laying the architectural groundwork for pixel-wise medical image segmentation.
- Paper: Graph Cuts and Efficient N-D Image Segmentation, Yuri Boykov et al. (2006). Details the classic combinatorial graph cut framework that exemplifies the traditional image segmentation baselines compared in the paper.
- Paper: SLIC Superpixels Compared to State-of-the-Art Superpixel Methods, Radhakrishna Achanta et al. (2012). Provides the foundational SLIC superpixel clustering algorithm used heavily in traditional boundary delineation and classical medical image segmentation baselines.
- Paper: Generalised Dice Overlap as a Deep Learning Loss Function for Highly Unbalanced Segmentations, C. Sudre et al. (2017). Formulates Generalized Dice overlap loss functions designed to address foreground-background imbalance in medical segmentation tasks like polyp delineation.
- Paper: Image Segmentation Using Deep Learning: A Survey, Shervin Minaee et al. (2020). Surveys the landscape of deep learning image segmentation architectures, benchmarks, and loss formulations that utilize datasets like Kvasir-SEG.
- Paper: UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation, Huimin Huang et al. (2020). Develops a full-scale connected U-Net variant that extends architectural designs to better capture multi-scale boundaries in medical imaging benchmarks.
- Paper: Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation, Hu Cao et al. (2021). Introduces a pure Vision Transformer U-Net architecture for medical image segmentation, advancing beyond the CNN baselines established in Kvasir-SEG.
- Paper: U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation, Chenxin Li et al. (2025). Applies Kolmogorov–Arnold Networks to medical image segmentation and explicitly evaluates the resulting architecture on endoscopic polyp benchmarks.
- Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Generalizes segmentation to a promptable foundation model framework that can perform zero-shot and interactive mask prediction across varied domains including medical images.
