Kvasir-SEG: A Segmented Polyp Dataset

Debesh JhaPia H. SmedsrudMichael A. RieglerPål HalvorsenThomas de LangeDag JohansenHåvard D. Johansen

article2019Conference on Multimedia Modeling2,034 citations

Introduces an open-access dataset of gastrointestinal polyp images paired with expert-verified pixel-level segmentation masks and bounding boxes to provide a standardized benchmark for evaluating automated colonoscopy analysis models.

Listen

Colorectal cancer is one of the most common cancers worldwide, making early detection and removal of precursor polyps critical for patient survival. However, colonoscopy screenings suffer from polyp miss rates between 14% and 30%, largely due to human error and visual complexity. Developing automated, computer-assisted diagnostic tools can significantly improve polyp detection and patient outcomes, but progress has been bottlenecked by the lack of publicly accessible, high-quality medical image datasets with pixel-level ground truth annotations.

The article introduces Kvasir-SEG, a newly released open-access dataset designed for pixel-wise polyp segmentation, and establishes baseline performance benchmarks using both traditional and deep-learning computational methods.

To build the dataset, researchers annotated 1,000 gastrointestinal polyp images from the existing Kvasir collection. A medical doctor and an engineer manually outlined polyp boundaries, which were subsequently validated by an experienced gastroenterologist to produce exact segmentation masks and bounding boxes. The authors evaluated two segmentation approaches on an 80/10/10 training, validation, and test split: an unsupervised baseline using Fuzzy C-means clustering with standard image filtering, and a modern deep-learning approach using a Deep Residual U-Net (ResUNet) architecture enhanced with data augmentation.

The experimental results demonstrated a substantial performance difference between the evaluated methods. The ResUNet deep-learning model achieved strong predictive capability on unseen test data, scoring a Dice coefficient of approximately 0.79 and a mean Intersection over Union of 0.78. In contrast, the unsupervised Fuzzy C-means method performed poorly, achieving a Dice coefficient of only 0.24 and an Intersection over Union of 0.31. ResUNet converged effectively within 91 training epochs and showed competitive or superior accuracy compared to benchmarks reported in prior literature.

These findings confirm that traditional color- and threshold-based segmentation algorithms fail because polyps share visual and color characteristics with surrounding healthy mucosal tissue. In contrast, modern convolutional architectures utilizing residual learning and data augmentation can accurately learn complex spatial patterns even from relatively small medical datasets. By providing verified, pixel-level masks alongside standard evaluation metrics, the open-access release enables reliable benchmarking and accelerates the development of cost-effective, real-time diagnostic aids for clinical endoscopy.

To advance toward clinical adoption, research teams should utilize the open-source Kvasir-SEG benchmark to develop and compare next-generation segmentation architectures. Prior to deploying such models in live hospital workflows, stakeholders must expand training on more diverse image sets, conduct prospective clinical trials, and evaluate real-time processing capabilities under operational colonoscopy conditions.

The primary limitation of the study is the dataset scale, which is constrained to 1,000 images from a single primary clinical source and may not capture every anatomical or pathological variation encountered in diverse patient populations. While confidence is high that deep-learning models like ResUNet clearly outperform traditional segmentation methods, the authors caution that current baseline performance must be improved and further validated before these algorithms can be directly deployed for live patient care.

Cover for Kvasir-SEG: A Segmented Polyp Dataset

Abstract

Pixel-wise image segmentation is a highly demanding task in medical-image analysis. In practice, it is difficult to find annotated medical images with corresponding segmentation masks. In this paper, we present Kvasir-SEG: an open-access dataset of gastrointestinal polyp images and corresponding segmentation masks, manually annotated by a medical doctor and then verified by an experienced gastroenterologist. Moreover, we also generated the bounding boxes of the polyp regions with the help of segmentation masks. We demonstrate the use of our dataset with a traditional segmentation approach and a modern deep-learning based Convolutional Neural Network (CNN) approach. The dataset will be of value for researchers to reproduce results and compare methods. By adding segmentation masks to the Kvasir dataset, which only provide frame-wise annotations, we enable multimedia and computer vision researchers to contribute in the field of polyp segmentation and automatic analysis of colonoscopy images.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The Kvasir-SEG dataset
  • 3.1 The original Kvasir dataset
  • 3.2 The Kvasir-SEG Dataset Details
  • 3.3 Mask Extraction
  • 4 Suggested Metrics
  • 5 Evaluation
  • 5.1 Baseline Models
  • 5.2 Implementation Details
  • 5.3 Results and Discussions
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Kvasir-SEG Dataset Specification

    definition

    Kvasir-SEG is an open-access benchmark dataset designed for semantic segmentation, localization, and classification of gastrointestinal polyps in colonoscopy images. The dataset comprises:

    • 1,000 polyp images: Color frames in JPEG format extracted from colonoscopy procedures. These images derive from the polyp class of the original Kvasir dataset, in which 13 lower-quality frames were replaced with clearer polyp images collected from Vestre Viken Health Trust in Norway.
    • 1,000 segmentation masks: Corresponding 1-bit binary images in JPEG format matching the dimensions of the original images. Pixels corresponding to polyp tissue (the region of interest) are represented as white foreground (value 1), while non-polyp areas are black background (value 0).
    • Bounding box annotations: Stored in an accompanying JSON metadata file, providing the coordinate rectangles that enclose each polyp region.

    A subset of the images includes visual artifacts from the Olympus ScopeGuide endoscope position marking probe. All polyp annotations were manually created and verified by medical professionals.

  2. Knowl 2 — Quantitative Polyp Segmentation Performance Benchmark on Kvasir-SEG

    data/table

    The benchmark evaluation on Kvasir-SEG establishes baseline performance metrics using the unsupervised Fuzzy C-Means (FCM) clustering algorithm and the deep learning ResUNet architecture on an unseen 10% test split:

    Model / Subset Loss Dice Coefficient Mean IoU
    ResUNet (Train) 0.059389 0.940609 0.920957
    ResUNet (Validation) 0.196520 0.803479 0.792339
    ResUNet (Test) 0.212236 0.787763 0.777771
    FCM (Test) — 0.239002 0.314187

    The deep learning ResUNet model substantially outperforms the unsupervised FCM baseline, demonstrating that supervised learning with convolutional feature representations and data augmentation is necessary for accurate polyp boundary segmentation.

  3. Knowl 3 — ResUNet Baseline Architecture and Training Setup

    experimental setup

    The deep learning segmentation baseline for Kvasir-SEG uses the Deep Residual U-Net (ResUNet) architecture:

    • Network Architecture: An encoder-decoder structure containing 5 residual convolutional blocks in the encoder (contracting path) and 5 residual convolutional blocks in the decoder (expanding path), linked via skip connections.
    • Input Preprocessing & Resolution: Input images and target masks are resized to 320×320320 \times 320 pixels.
    • Data Augmentation: Training samples undergo random horizontal and vertical flipping, random cropping, scaling, rotation, brightness adjustment, cutout, and random erasing.
    • Dataset Split: The 1,000 images are partitioned into 80% training (800 images), 10% validation (100 images), and 10% testing (100 images).
    • Training Hyperparameters: Optimization uses Nadam with learning rate η=10−4\eta = 10^{-4}, first moment decay β1=0.9\beta_1 = 0.9, and second moment decay β2=0.999\beta_2 = 0.999. Activation functions are Rectified Linear Units (ReLU). The loss function is the Dice loss. Models are trained with a batch size of 8 for up to 150 epochs, reaching convergence at epoch 91.
    • Inference Threshold: Continuous model prediction maps are binarized using a decision threshold of t=0.5t = 0.5.
  4. Knowl 4 — Polyp Annotation and Mask Generation Pipeline

    model/method

    The ground truth generation workflow for Kvasir-SEG consists of a multi-stage clinical annotation and mask synthesis pipeline:

    1. Manual Outlining: The 1,000 polyp frames from the Kvasir dataset are imported into the Labelbox annotation platform. An engineer and a medical doctor manually delineate the precise contours of all visible polyp lesions.
    2. Expert Verification: All annotations are independently reviewed and validated by an experienced gastroenterologist.
    3. Binary Mask Generation: Boundary coordinates exported from Labelbox in JSON format are used to draw contours on an empty black (pixel value 0) canvas, which are then filled with white (pixel value 1) to produce 1-bit color depth binary masks.
    4. Bounding Box Extraction: The outer coordinate bounds enclosing the polyp contours are computed and exported into a unified JSON file for localization tasks.
  5. Knowl 5 — Evaluation Metrics for Polyp Semantic Segmentation

    equation

    Pixel-level polyp segmentation on Kvasir-SEG is evaluated using the Dice Coefficient and Intersection over Union (IoU).

    Let A⊂Z2A \subset \mathbb{Z}^2 denote the set of predicted positive polyp pixels and B⊂Z2B \subset \mathbb{Z}^2 denote the ground truth set of positive polyp pixels. Let TPTP be true positive pixels (∣A∩B∣|A \cap B|), FPFP be false positive pixels (∣A∖B∣|A \setminus B|), and FNFN be false negative pixels (∣B∖A∣|B \setminus A|).

    The Dice Coefficient measures the harmonic overlap between predicted and ground truth masks:

    Dice(A,B)=2×∣A∩B∣∣A∣+∣B∣=2×TP2×TP+FP+FN\text{Dice}(A, B) = \frac{2 \times |A \cap B|}{|A| + |B|} = \frac{2 \times TP}{2 \times TP + FP + FN}

    The Intersection over Union (IoU), calculated at a binary decision threshold t∈[0,1]t \in [0, 1] applied to predicted probabilities, measures the area of overlap over the area of union:

    IoU(A,B)=∣A∩B∣∣A∪B∣=TP(t)TP(t)+FP(t)+FN(t)\text{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} = \frac{TP(t)}{TP(t) + FP(t) + FN(t)}

    where TP(t)TP(t), FP(t)FP(t), and FN(t)FN(t) represent true positives, false positives, and false negatives at threshold tt.

  6. Knowl 6 — Fuzzy C-Means Polyp Segmentation Preprocessing Pipeline

    algorithm

    The unsupervised baseline on Kvasir-SEG applies a multi-step image processing pipeline before clustering with Fuzzy C-Means (FCM):

    Input: RGB colonoscopy image II of dimensions H×W×3H \times W \times 3
    Output: Binary segmentation mask MM of dimensions H×WH \times W
    Igray←ConvertToGrayscale(I)I_{\text{gray}} \leftarrow \text{ConvertToGrayscale}(I)
    Iblur←MedianBlur(Igray)I_{\text{blur}} \leftarrow \text{MedianBlur}(I_{\text{gray}})
    ROI←MedianBasedOtsu(Iblur)\text{ROI} \leftarrow \text{MedianBasedOtsu}(I_{\text{blur}})
    Inorm←NormalizeToUnitRange(Igray)I_{\text{norm}} \leftarrow \text{NormalizeToUnitRange}(I_{\text{gray}})
    Idiff←Inorm−NormalizeToUnitRange(Iblur)I_{\text{diff}} \leftarrow I_{\text{norm}} - \text{NormalizeToUnitRange}(I_{\text{blur}})
    Iedge←Threshold(Idiff)I_{\text{edge}} \leftarrow \text{Threshold}(I_{\text{diff}})
    Idilated←Dilate(Iedge)I_{\text{dilated}} \leftarrow \text{Dilate}(I_{\text{edge}})
    Isub←Inorm−IedgeI_{\text{sub}} \leftarrow I_{\text{norm}} - I_{\text{edge}}
    Iclipped←ClipToRange(Isub,0.0,1.0)I_{\text{clipped}} \leftarrow \text{ClipToRange}(I_{\text{sub}}, 0.0, 1.0)
    V1D←FlattenTo1D(Iclipped)V_{\text{1D}} \leftarrow \text{FlattenTo1D}(I_{\text{clipped}})
    C←FuzzyCMeansClustering(V1D)C \leftarrow \text{FuzzyCMeansClustering}(V_{\text{1D}})
    M←ReshapeTo2D(C,H,W)M \leftarrow \text{ReshapeTo2D}(C, H, W)
    return MM
  7. Knowl 7 — Failure Modes of Color-Based Unsupervised Clustering for Polyp Delineation

    limitation

    Unsupervised clustering algorithms such as Fuzzy C-Means (FCM) perform poorly on polyp segmentation in colonoscopy imagery (Dice score of 0.2390 and mean IoU of 0.3142) due to two primary factors:

    1. Color Ambiguity in Colonoscopy Images: FCM relies heavily on color and pixel intensity distributions to partition image regions. However, polyps typically exhibit color and luminance profiles highly similar to normal surrounding mucosa and other benign gastrointestinal features, causing threshold- and color-based partitioning to fail.
    2. Absence of Parameterized Feature Learning and Invariance: FCM does not have a trainable parameter set or convolutional hierarchy capable of learning spatial, structural, or contextual representations, nor can it leverage data augmentation techniques (e.g., spatial transforms, cutout) to achieve invariance to lighting variations and specular reflections.

Coverage note — No substantial contributed material was omitted; all key contributions including dataset specifications, annotation workflows, metrics, baseline models, preprocessing pipelines, and empirical results are included.

References

  1. 1.Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Levenberg, J., Monga, R., Moore, S., Murray, D., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., Wicke, M., Yu, Y., Zheng, X., Brain, G.: Tensorflow: A System for Large-Scale Machine Learning. In: Proceeding of the ACM Symposium on Operating Systems Design and Implementation (SOSP). pp. 265–283 (2016)
  2. 2.Bernal, J., S´anchez, F.J., Fern´andez-Esparrach, G., Gil, D., Rodr´ıguez, C., Vilari˜no, F.: WM-DOVA maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians. Computerized Medical Imaging and Graphics 43, 99–111 (2015)
  3. 3.Bernal, J., S´anchez, J., Vilarino, F.: Towards automatic polyp detection with a polyp appearance model. Pattern Recognition 45(9), 3166–3182 (2012)
  4. 4.Bernal, J., Tajkbaksh, N., S´anchez, F.J., Matuszewski, B.J., Chen, H., Yu, L., Angermann, Q., Romain, O., Rustad, B., Balasingham, I., et al.: Comparative Validation of Polyp Detection Methods in Video Colonoscopy: Results From the MICCAI 2015 Endoscopic Vision Challenge. IEEE transactions on medical imaging 36(6), 1231–1249 (2017)
  5. 5.Boccardi, M., Ganzola, R., Bocchetta, M., Pievani, M., Redolfi, A., Bartzokis, G., Camicioli, R., Csernansky, J., Leon, M., deToledo Morrell, L., Killiany, R., Lehricy, S., Pantel, J., Pruessner, J.C., Soininen, H., Watson, C., Duchesne, S., Jr, C., Frisonia, G.: Survey of protocols for the manual segmentation of the hippocampus: preparatory steps towards a joint EADC-ADNI harmonized protocol. Journal of Alzheimer’s disease 26(s3), 61–75 (2011)
  6. 6.Cai, W., Chen, S., Zhang, D.: Fast and robust fuzzy c-means clustering algorithms incorporating local information for image segmentation. Pattern recognition 40(3), 825–838 (2007)
  7. 7.Chollet, F.: Building powerful image classification models using very little data. Keras Blog (2016)
  8. 8.Chollet, F.: Keras: The Python Deep Learning library. Astrophysics Source Code Library (2018)
  9. 9.Dravid, A.: Employing Deep Networks for Image Processing on Small Research Datasets. Microscopy Today 27(1), 18–23 (2019)
  10. 10.Goldbloom, A., Hamner, B., et al.: Kaggle: Your home for data science. Competition, Kaggle, Inc, https://www.kaggle.com (2019), accessed: 2019-07-12
  11. 11.Haggar, F.A., Boushey, R.P.: Colorectal Cancer Epidemiology: Incidence, Mortality, Survival, and Risk Factors. Clinics in colon and rectal surgery 22(04), 191–197 (2009)
  12. 12.Kaminski, M.F., Wieszczy, P., Rupinski, M., Wojciechowska, U., Didkowska, J., Kraszewska, E., Kobiela, J., Franczyk, R., Rupinska, M., Kocot, B., Chaber-Ciopinska, A., Pachlewski, J., Polkowski, M., Regula, J.: Increased Rate of Adenoma Detection Associates With Reduced Risk of Colorectal Cancer and Death. Gastroenterology 153(1), 98–105 (2017)
  13. 13.Kang, J., Gwak, J.: Ensemble of instance segmentation models for polyp segmentation in colonoscopy images. IEEE Access 7, 26440–26447 (2019)
  14. 14.Otsu, N.: A threshold Selection Method from Gray-Level Histograms. IEEE transactions on systems, man, and cybernetics 9(1), 62–66 (1979)
  15. 15.Pham, D.L., Xu, C., Prince, J.L.: Current methods in medical image segmentation. Annual review of biomedical engineering 2(1), 315–337 (2000)
  16. 16.Pogorelov, K., Randel, K., Griwodz, C., Sigrun, E., Lange, T., Johansen, D., Spampinato, C., Dang-Nguyen, D., Lux, M., Schmidt, P., Riegler, M., Halvorsen, P.: Kvasir: A Multi-Class Image Dataset for Computer Aided Gastrointestinal Disease Detection. In: Proceedings of Multimedia Systems Conference (MMSYS). pp. 164–169. ACM (2017)
  17. 17.Pogorelov, K., Riegler, M., Halvorsen, P., Hicks, S.A., Randel, K.R., Dang-Nguyen, D.T., Lux, M., Ostroukhova, O., de Lange, T.: Medico Multimedia Task at Mediaeval 2018. In: CEUR Workshop Proceedings - Multimedia Benchmark Workshop (MediaEval) (2018)
  18. 18.Pozdeev, A.A., Obukhova, N.A., Motyko, A.A.: Automatic Analysis of Endoscopic Images for Polyps Detection and Segmentation. In: IEEE Conference of Russian Young Researchers in Electrical and Electronic Engineering (EIConRus). pp. 1216–1220. IEEE (2019)
  19. 19.Riegler, M., Lux, M., Griwodz, C., Spampinato, C., de Lange, T., Eskeland, S.L., Pogorelov, K., Tavanapong, W., Schmidt, P.T., Gurrin, C., Johansen, D., Johansen, H., Halvorsen, P.: Multimedia and Medicine: Teammates for Better Disease Detection and Survival. In: Proceedings of ACM Multimedia (ACM MM). pp. 968–977. ACM (2016)
  20. 20.Riegler, M., Pogorelov, K., Halvorsen, P., Ranheim Randel, K., Losada Eskeland, S., Dang-Nguyen, D.T., Lux, M., Griwodz, C., Spampinato, C., de Lange, T.: Multimedia for medicine: the medico task at Mediaeval 2017. CEUR Workshop Proceedings - Multimedia Benchmark Workshop (MediaEval) (2017)
  21. 21.Rundle, A.G., Lebwohl, B., Vogel, R., Levine, S., Neugut, A.I.: Colonoscopic screening in average-risk individuals ages 40 to 49 vs 50 to 59 years. Gastroenterology 134(5), 1311–1315 (2008)
  22. 22.Sharma, M., Rasmuson, D., Rieger, B., Kjelkerud, D., et al.: Labelbox: The best way to create and manage training data. software, LabelBox, Inc, https://www.labelbox.com/ (2019), accessed: 2019-05-21
  23. 23.Silva, J., Histace, A., Romain, O., Dray, X., Granado, B.: Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer. International Journal of Computer Assisted Radiology and Surgery 9(2), 283–293 (2014)
  24. 24.Tajbakhsh, N., Gurudu, S.R., Liang, J.: Automated polyp detection in colonoscopy videos using shape and context information. IEEE transactions on medical imaging 35(2), 630–644 (2015)
  25. 25.Torre, L.A., Bray, F., Siegel, R.L., Ferlay, J., Lortet-Tieulent, J., Jemal, A.: Global cancer statistics, 2012. CA: a cancer journal for clinicians 65(2), 87–108 (2015)
  26. 26.Van Rijn, J.C., Reitsma, J.B., Stoker, J., Bossuyt, P.M., Van Deventer, S.J., Dekker, E.: Polyp miss rate determined by tandem colonoscopy: a systematic review. The American journal of gastroenterology 101(2), 343 (2006)
  27. 27.Visser, M., M¨uller, D., van Duijn, R., Smits, M., Verburg, N., Hendriks, E., Nabuurs, R., Bot, J., Eijgelaar, R., Witte, M., van, M., Barkhof, F., de, P., de, J.: Inter-rater agreement in glioma segmentations on longitudinal MRI. NeuroImage: Clinical 22, 101727 (2019)
  28. 28.Zhang, Z., Liu, Q., Wang, Y.: Road Extraction by Deep Residual U-Net. IEEE Geoscience and Remote Sensing Letters 15(5), 749–753 (2018)

Citation

MLA
Jha, D., et al. “Kvasir-SEG: A Segmented Polyp Dataset”. arXiv, 2019, http://arxiv.org/abs/1911.07069v1.
APA
Jha, D., Smedsrud, P. H., Riegler, M. A., Halvorsen, P., Lange, T. de ., Johansen, D., & Johansen, H. D. (2019). Kvasir-SEG: A Segmented Polyp Dataset. arXiv. http://arxiv.org/abs/1911.07069v1
Chicago
Jha, D., P. H. Smedsrud, M. A. Riegler, et al. 2019. “Kvasir-SEG: A Segmented Polyp Dataset”. arXiv. http://arxiv.org/abs/1911.07069v1.
Harvard
Jha, D. et al. (2019) “Kvasir-SEG: A Segmented Polyp Dataset”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1911.07069v1.
Vancouver
1. Jha D, Smedsrud PH, Riegler MA, Halvorsen P, Lange T de, Johansen D, Johansen HD (2019) Kvasir-SEG: A Segmented Polyp Dataset. arXiv

BibTeX

@article{jha2019kvasir,
  title = {Kvasir-SEG: A Segmented Polyp Dataset},
  author = {Jha, Debesh and Smedsrud, Pia H. and Riegler, Michael A. and Halvorsen, Pål and Lange, Thomas de and Johansen, Dag and Johansen, Håvard D.},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1911.07069v1},
  eprint = {1911.07069}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF