Skin lesion analysis toward melanoma detection: A challenge at the 2017 International symposium on biomedical imaging (ISBI), hosted by the international skin imaging collaboration (ISIC)

David A. GutmanNoel CodellaM. E. CelebiBrian HelbaMichael A. MarchettiNabin K. MishraAllan C. Halpern

article2016ISBI2,800 citations

Establishes a standardized comparative benchmark for automated melanoma diagnosis by analyzing participant algorithms and results across lesion segmentation, dermoscopic feature detection, and disease classification on the ISIC 2017 dataset.

Listen

Skin cancer is the most common cancer in the United States, with melanoma responsible for over 9,000 deaths annually. Visual diagnosis by experts remains limited in accuracy, and shortages of dermatologists have increased interest in automated tools. The article describes the 2017 ISIC-hosted challenge that provided a large public dataset to benchmark algorithms for automated melanoma detection.

The challenge set out to evaluate methods across three tasks using a fixed snapshot of dermoscopic images: lesion segmentation, detection of four specific dermoscopic features, and classification into melanoma, seborrheic keratosis, or benign nevi. Training, validation, and test sets comprised 2,000, 150, and 600 images respectively. Submissions were evaluated with standard metrics including Jaccard index for segmentation and AUC for classification.

The effort attracted 593 registrations and 46 finalized submissions, making it the largest comparative study in the field. Top segmentation entries reached an average Jaccard index of 0.765 using deep learning ensembles, though 1526% of images showed performance below inter-observer agreement. Feature detection achieved average AUC near 0.9 despite low participation. Classification yielded AUC values around 0.870.96, with ensembles of deep networks plus extra training data performing best; simple fusions of all submissions further improved results to an average AUC of 0.926.

These outcomes indicate that collaborative deep learning approaches can approach or exceed dermatologist-level sensitivity at useful specificity thresholds, supporting scalable automated triage. However, dataset bias across diseases, ages, devices, and ethnicities, incomplete feature annotations, and reliance on single metrics limit generalizability. Future challenges should refine evaluation metrics, reformat tasks for broader participation, and emphasize interpretable outputs for clinical integration. Additional diverse data and pilot studies in real workflows are needed before widespread deployment.

Cover for Skin lesion analysis toward melanoma detection: A challenge at the 2017 International symposium on biomedical imaging (ISBI), hosted by the international skin imaging collaboration (ISIC)

Abstract

This article describes the design, implementation, and results of the latest installment of the dermoscopic image analysis benchmark challenge. The goal is to support research and development of algorithms for automated diagnosis of melanoma, the most lethal skin cancer. The challenge was divided into 3 tasks: lesion segmentation, feature detection, and disease classification. Participation involved 593 registrations, 81 pre-submissions, 46 finalized submissions (including a 4-page manuscript), and approximately 50 attendees, making this the largest standardized and comparative study in this field to date. While the official challenge duration and ranking of participants has concluded, the dataset snapshots remain available for further research and development.

Table of Contents

  • 1 Introduction
  • 2 DATASET DESCRIPTIONS & TASKS
  • 3 EVALUATION METRICS
  • 4 RESULTS
  • 5 DISCUSSION & CONCLUSION
  • References

Knowls

  1. Knowl 1 — Comparative Performance and Multi-Model Fusion in ISIC 2017 Disease Classification

    data/table

    The ISIC 2017 disease classification benchmark evaluated automated systems on a holdout test set of 600 dermoscopic images across three diagnostic classes: melanoma (M, n=117n=117), seborrheic keratosis (SK, n=90n=90), and benign nevi (n=393n=393). In addition to the top individual competitor models, three fusion strategies combining prediction vectors from all 23 final test submissions via 3-fold cross-validation were evaluated: simple score averaging (AVGSC), linear support vector machine fusion (L-SVM), and non-linear SVM fusion (NL-SVM) utilizing a histogram intersection kernel with Platt-style probabilistic score calibration.

    Method AVG-AUC M-AUC SK-AUC M-SP82 M-SP89 M-SP95 M-SENS M-SPEC SK-SENS SK-SPEC
    Matsunaga et al. (Top AVG) 0.911 0.868 0.953 0.729 0.588 0.366 0.735 0.851 0.978 0.773
    Diaz (Top SK) 0.910 0.856 0.965 0.727 0.555 0.404 0.103 0.998 0.178 0.998
    Menegola et al. (Top M) 0.908 0.874 0.943 0.747 0.590 0.395 0.547 0.950 0.356 0.990
    AVGSC (Score Average Fusion) 0.913 0.872 0.954 0.778 0.605 0.435 0.214 0.988 0.600 0.975
    L-SVM (Linear SVM Fusion) 0.926 0.892 0.960 0.834 0.692 0.571 0.718 0.901 0.878 0.931
    NL-SVM (Non-Linear SVM Fusion) 0.904 0.853 0.955 0.801 0.449 0.168 0.675 0.909 0.889 0.928

    In the table, AVG-AUC represents the mean area under the receiver operating characteristic curve across melanoma and seborrheic keratosis binary tasks; M-AUC and SK-AUC represent the class-specific AUCs; M-SP82, M-SP89, and M-SP95 represent melanoma specificity fixed at sensitivities of 82%, 89%, and 95%, respectively; and SENS/SPEC indicate sensitivity and specificity at the default decision threshold of 0.5.

    The linear SVM fusion (L-SVM) achieved the highest overall performance across almost all metrics (AVG-AUC of 0.926, M-AUC of 0.892), outperforming every single competitor model as well as non-linear fusion. Classification of seborrheic keratosis consistently achieved higher AUC values (0.943--0.965) than melanoma classification (0.856--0.892).

  2. Knowl 2 — Empirical Trends in Deep Learning and Ensembling for Skin Disease Classification

    empirical result

    Analysis of the 23 submitted systems and ensemble fusion strategies in the ISIC 2017 skin lesion disease classification challenge revealed five consistent empirical patterns:

    1. Deep Learning Ensembles and External Data: All top-performing submissions employed ensembles of deep convolutional neural networks. Furthermore, all top teams leveraged additional training data beyond the core dataset, including external datasets, supplementary ISIC archive images, or in-house dermoscopic feature annotations.
    2. Relative Task Difficulty: Seborrheic keratosis classification proved consistently easier than melanoma classification across all algorithms, yielding higher ROC Area Under the Curve (AUC) scores (ranging up to 0.965 for seborrheic keratosis versus 0.874 for individual melanoma models). Integrating weakly labeled dermoscopic pattern annotations into training yielded the highest individual seborrheic keratosis performance.
    3. Category Trade-offs: The participant with the highest average AUC across categories was not the top performer in any individual disease category.
    4. Fusion Complexity vs. Performance: Collaborative fusion across all participant submissions improved classification over any single algorithm. However, simpler fusion models (Linear SVM and simple score averaging) outperformed the more complex non-linear SVM with a histogram intersection kernel, which degraded performance (dropping average AUC from 0.926 with Linear SVM to 0.904 with Non-Linear SVM).
    5. Threshold Calibration: Raw prediction thresholds of 0.5 often led to extreme imbalances between sensitivity and specificity in uncalibrated individual submissions. Applying probabilistic score normalization within SVM ensembles effectively balanced sensitivity and specificity across operating points.
  3. Knowl 3 — ISIC 2017 Benchmark Dataset Specifications and Task Structure

    experimental setup

    The ISIC 2017 dermoscopic image analysis challenge established a standardized public benchmark divided into three distinct diagnostic tasks over a dataset split into training (n=2000n=2000), validation (n=150n=150), and holdout test (n=600n=600) partitions:

    • Part 1: Lesion Segmentation: Automated extraction of binary masks from dermoscopic images. Ground truth consists of expert manual lesion boundary tracings encoded as binary masks (pixel value 255 for lesion interior, 0 for exterior).
    • Part 2: Dermoscopic Feature Classification: Localization and detection of four clinical dermoscopic patterns: pigment network, negative network, streaks, and milia-like cysts. To reduce spatial annotation dimensionality and variability, images were segmented into superpixels using the Simple Linear Iterative Clustering (SLIC) algorithm, with expert binary annotations provided per superpixel.
    • Part 3: Disease Classification: Multi-class lesion categorization into three diagnostic classes: melanoma, seborrheic keratosis, and benign nevi. The training set contained 374 melanomas, 254 seborrheic keratoses, and 1372 benign nevi; the validation set contained 30 melanomas, 42 seborrheic keratoses, and 78 benign nevi; and the holdout test set contained 117 melanomas, 90 seborrheic keratoses, and 393 benign nevi. Metadata including patient sex and approximate age (in 5-year intervals) were supplied when available.
  4. Knowl 4 — ISIC 2017 Evaluation Metrics and Clinical Operating Points

    experimental setup

    Evaluation of participant algorithms in the ISIC 2017 challenge utilized task-specific quantitative metrics and threshold definitions:

    • Segmentation Evaluation: Binary masks were generated by thresholding predicted probability maps at a pixel intensity of 128 (on a 0--255 scale). Submissions were evaluated using the Jaccard Index (Intersection over Union, the primary metric for participant ranking), the Dice coefficient (F1F_1-score), and overall pixel-wise accuracy.
    • Classification Evaluation: Predictions were submitted as continuous scores in the interval [0.0,1.0][0.0, 1.0] per disease category, with 0.50.5 serving as the default binary operating threshold. Discriminative power was quantified using the Area Under the Receiver Operating Characteristic Curve (ROC AUC).
    • Clinically Grounded Operating Points for Melanoma: To benchmark against dermatologist capabilities, specificity for melanoma detection was evaluated at three fixed sensitivity targets on the ROC curve:
      • SP82\text{SP82}: Specificity at 82%82\% sensitivity, corresponding to average dermatologist classification performance.
      • SP89\text{SP89}: Specificity at 89%89\% sensitivity, corresponding to dermatologist biopsy/management decision performance.
      • SP95\text{SP95}: Specificity at 95%95\% sensitivity, representing a high-sensitivity screening benchmark.
  5. Knowl 5 — Lesion Segmentation Performance and Inter-Observer Error Discrepancy

    empirical result

    In the ISIC 2017 lesion segmentation benchmark (evaluated on 600600 test images across 21 final submissions), the top-ranking system achieved an average Jaccard Index of 0.7650.765, a Dice coefficient of 0.8490.849, and a pixel-wise accuracy of 93.4%93.4\% using an ensemble of fully convolutional-deconvolutional deep neural networks.

    Analysis of the distribution of individual image Jaccard scores revealed a substantial divergence between aggregate pixel accuracy and acceptable boundary segmentation quality:

    • A Jaccard Index of 0.8\ge 0.8 closely matches visually acceptable segmentations and aligns with the estimated human inter-observer agreement threshold (0.7860.786).
    • 156156 out of 600600 test images (26.0%26.0\%) resulted in a Jaccard Index 0.70\le 0.70, where segmentation correctness becomes highly questionable.
    • 9191 out of 600600 test images (15.2%15.2\%) had a Jaccard Index 0.60\le 0.60.

    This indicates an effective failure rate between 15.2%15.2\% and 26.0%26.0\%, demonstrating that global pixel-wise error rates (6.6%6.6\%) significantly underestimate clinically meaningful segmentation failures.

  6. Knowl 6 — Dermoscopic Feature Classification Benchmark Performance

    data/table

    The ISIC 2017 feature classification benchmark evaluated automated localization and classification of four localized dermoscopic patterns (pigment network, negative network, streaks, and milia-like cysts) mapped across SLIC superpixel partitions on the 600-image holdout test set.

    Method / Rank AVG AUC Network Neg. Network Streaks Milia-Like Cyst
    Kawahara Hamarneh (Rank 1) 0.895 0.945 0.869 0.960 0.807
    Li Shen (Rank 2) 0.833 0.835 0.762 0.896 0.838
    Li Shen (Rank 3) 0.832 0.828 0.762 0.900 0.837

    Across all submitted methods, the area under the ROC curve (AUC) exceeded 0.750.75 for every individual feature class, with the winning fully convolutional architecture reaching an average AUC of 0.8950.895 across categories (peaking at 0.9600.960 for streak detection and 0.9450.945 for pigment networks). These results indicate that localized dermoscopic feature classification is tractable for computer vision models, despite receiving fewer participant submissions than global classification or segmentation tasks.

  7. Knowl 7 — Benchmark and Methodological Limitations of the ISIC 2017 Study

    limitation

    Several limitations were identified in the design, data, and evaluation protocols of the ISIC 2017 challenge:

    1. Dataset and Demographic Bias: The dataset exhibited demographic and acquisition biases, with unequal distributions across patient ages, imaging devices, clinical centers, skin ethnicities, and disease conditions.
    2. Continuous Metric Distortion for Segmentation: Mean Jaccard Index across images failed to reflect the true proportion of catastrophic segmentation failures relative to human inter-observer variability (0.7860.786). A binary threshold-based success/failure error metric based on inter-observer difference tolerance was identified as a necessary replacement for future benchmarks.
    3. Annotation and Formatting Bottlenecks: Structuring the dermoscopic feature detection task around SLIC superpixels rather than standard object bounding boxes or pixel segmentations created an engineering barrier that suppressed community participation. Furthermore, dermoscopic feature ground truth annotations were incomplete across the archive.
    4. Lack of Clinical Interpretability: Nearly all top-performing classification systems operated as opaque deep neural network ensembles, providing no human-interpretable visual or procedural evidence (e.g., ABCD rule features or dermoscopic criteria) to support clinical decision-making.

Coverage note — No substantial contributed material was omitted from the extracted knowls.

References

  1. 1.Rogers HW, Weinstock MA, Feldman SR, Coldiron BM.: “Incidence estimate of nonmelanoma skin cancer (keratinocyte carcinomas) in the US population, 2012” JAMA Dermatol vol. 151, no. 10, pp. 1081-1086. 2015.
  2. 2.“Cancer Facts & Figures 2017”. American Cancer Society, 2017. Available: https://www.cancer.org/research/cancer-factsstatistics/all-cancer-facts-figures/cancer-facts-figures2017.html
  3. 3.Siegel, R.L., Miller, K.D., and Jemal, A.: “Cancer statistics, 2017,” CA: A Cancer Journal for Clinicians, vol. 67, no. 1, pp. 7-30. 2017.
  4. 4.Brady, M.S., Oliveria, S.A., Christos, P.J., Berwick, M., Coit, D.G., Katz, J., Halpern, A.C.: “Patterns of detection in patients with cutaneous melanoma.” Cancer. vol. 89, no. 2, pp. 342-7. 2000.
  5. 5.Kittler, H., Pehamberger, H., Wolff, K., Binder, M.: “Diagnostic accuracy of dermoscopy”. The Lancet Oncology. vol. 3, no. 3, pp. 159-165. 2002.
  6. 6.Carli, P., et al.: “Pattern analysis, not simplified algorithms, is the most reliable method for teaching dermoscopy for melanoma diagnosis to residents in dermatology”. Br J Dermatol. vol. 148, no. 5, pp. 981-4. 2003.
  7. 7.Vestergaard, M.E., Macaskill, P., Holt, P.E., et al.: “Dermoscopy compared with naked eye examination for the diagnosis of primary melanoma: a meta-analysis of studies performed in a clinical setting.” Br J Dermatol. vol. 159, pp. 669-676. 2008.
  8. 8.Argenziano, G. et al.: “Dermoscopy of pigmented skin lesions: Results of a consensus meeting via the Internet” J. American Academy of Dermatology. vol. 48, no. 5, 2003.
  9. 9.Gachon, J., et. al.:“First Prospective Study of the Recognition Process of Melanoma in Dermatological Practice”. Arch Dermatol. vol. 141, no. 4, pp. 434-438, 2005.
  10. 10.Kimball, A.B., Resneck, J.S. Jr.: “The US dermatology workforce: a specialty remains in shortage.” J Am Acad Dermatol. vol. 59, no. 5, pp. 741-5. 2008.
  11. 11.Mishra, N.K., Celebi, M.E.: “An Overview of Melanoma Detection in Dermoscopy Images Using Image Processing and Machine Learning” arxiv.org: 1601.07843. Available: http://arxiv.org/abs/1601.07843
  12. 12.Ali, A.A., Deserno, T.M.: “A Systematic Review of Automated Melanoma Detection in Dermatoscopic Images and its Ground Truth Data” Proc. of SPIE Vol. 8318 83181I-1
  13. 13.Codella NCF, Nguyen B, Pankanti S, Gutman D, Helba B, Halpern A, Smith JR. “Deep learning ensembles for melanoma recognition in dermoscopy images” IBM Journal of Research and Development, vol. 61, no. 4/5, 2017. Available: https://arxiv.org/pdf/1610.04662.pdf
  14. 14.Barata, C., Ruela, M., et al.: “Two Systems for the Detection of Melanomas in Dermoscopy Images using Texture and Color Features”. IEEE Systems Journal, vol. 8, no. 3, pp. 965-979, 2014.
  15. 15.Mendonca, T., Ferreira, P.M., Marques, J.S., Marcal, A.R., Rozeira, J.: “PH2 - a dermoscopic image database for research and benchmarking”. Conf Proc IEEE Eng Med Biol Soc. pp. 5437-40, 2013.
  16. 16.Gutman D, Codella N, Celebi E, Helba B, Marchetti M, Mishra N, Halpern A. “Skin Lesion Analysis toward Melanoma Detection: A Challenge at the International Symposium on Biomedical Imaging (ISBI) 2016, hosted by the International Skin Imaging Collaboration (ISIC)”. eprint arXiv:1605.01397 [cs.CV]. 2016. Available: https://arxiv.org/abs/1605.01397
  17. 17.Marchetti M, et al. “Results of the 2016 International Skin Imaging Collaboration International Symposium on Biomedical Imaging challenge: Comparison of the accuracy of computer algorithms to dermatologists for the diagnosis of melanoma from dermoscopic images”. Journal of the American Academy of Dermatology, 2017. In Press.
  18. 18.Braun, R.P., Rabinovitz, H.S., Oliviero, M., Kopf, A.W., Saurat, J.H.:“Dermoscopy of pigmented skin lesions.”. J Am Acad Dermatol. vol. 52, no. 1, pp. 109-21. 2005.
  19. 19.Rezze, G.G., Soares de S, B.C., Neves, R.I.: “Dermoscopy: the pattern analysis”. An Bras Dermatol., vol. 3, pp. 261-8. 2006.
  20. 20.Achanta, R., Shaji, A., Smith, K., Lucchi, A., Fua, P., and Susstrun,S.: “SLIC Superpixels”, EPFL Technical Report 149300, June 2010.
  21. 21.Yuan Y, Chao M, Lo YC. “Automatic skin lesion segmentation with fully convolutional-deconvolutional networks”. International Skin Imaging Collaboration (ISIC) 2017 Challenge at the International Symposium on Biomedical Imaging (ISBI). Available: https://arxiv.org/pdf/1703.05165.pdf
  22. 22.Kawahara J, Hamarneh G. “Fully Convolutional Networks to Detect Clinical Dermoscopic Features”. International Skin Imaging Collaboration (ISIC) 2017 Challenge at the International Symposium on Biomedical Imaging (ISBI). Available: https://arxiv.org/abs/1703.04559
  23. 23.Li Y, Shen L. “Skin Lesion Analysis Towards Melanoma Detection Using Deep Learning Network”. International Skin Imaging Collaboration (ISIC) 2017 Challenge at the International Symposium on Biomedical Imaging (ISBI). Available: https://arxiv.org/abs/1703.00577
  24. 24.Matsunaga K, Hamada A, Minagawa A, Koga H. “Image Classification of Melanoma, Nevus and Seborrheic Keratosis by Deep Neural Network Ensemble”. International Skin Imaging Collaboration (ISIC) 2017 Challenge at the International Symposium on Biomedical Imaging (ISBI). Available: https://arxiv.org/abs/1703.03108
  25. 25.Daz IG. “Incorporating the Knowledge of Dermatologists to Convolutional Neural Networks for the Diagnosis of Skin Lesions”. International Skin Imaging Collaboration (ISIC) 2017 Challenge at the International Symposium on Biomedical Imaging (ISBI). Available: https://arxiv.org/abs/1703.01976
  26. 26.Menegola A, Tavares J, Fornaciali M, Li LT, Avila S, Valle E. “RECOD Titans at ISIC Challenge 2017”. International Skin Imaging Collaboration (ISIC) 2017 Challenge at the International Symposium on Biomedical Imaging (ISBI). Available: https://arxiv.org/pdf/1703.04819.pdf
  27. 27.Fishbaugh, J, et al. ”Data-Driven Rank Aggregation with Application to Grand Challenges.” International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI). Springer, pp. 754-762. 2017.

Citation

MLA
Codella, N. C. F., et al. “Skin Lesion Analysis Toward Melanoma Detection: A Challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), Hosted by the International Skin Imaging Collaboration (ISIC)”. arXiv, 2017, http://arxiv.org/abs/1710.05006v3.
APA
Codella, N. C. F., Gutman, D., Celebi, M. E., Helba, B., Marchetti, M. A., Dusza, S. W., Kalloo, A., Liopyris, K., Mishra, N., Kittler, H., & Halpern, A. (2017). Skin Lesion Analysis Toward Melanoma Detection: A Challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), Hosted by the International Skin Imaging Collaboration (ISIC). arXiv. http://arxiv.org/abs/1710.05006v3
Chicago
Codella, N. C. F., D. Gutman, M. E. Celebi, et al. 2017. “Skin Lesion Analysis Toward Melanoma Detection: A Challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), Hosted by the International Skin Imaging Collaboration (ISIC)”. arXiv. http://arxiv.org/abs/1710.05006v3.
Harvard
Codella, N.C.F. et al. (2017) “Skin Lesion Analysis Toward Melanoma Detection: A Challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), Hosted by the International Skin Imaging Collaboration (ISIC)”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1710.05006v3.
Vancouver
1. Codella NCF, Gutman D, Celebi ME, et al (2017) Skin Lesion Analysis Toward Melanoma Detection: A Challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), Hosted by the International Skin Imaging Collaboration (ISIC). arXiv

BibTeX

@article{codella2017skin,
  title = {Skin Lesion Analysis Toward Melanoma Detection: A Challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), Hosted by the International Skin Imaging Collaboration (ISIC)},
  author = {Codella, Noel C. F. and Gutman, David and Celebi, M. Emre and Helba, Brian and Marchetti, Michael A. and Dusza, Stephen W. and Kalloo, Aadi and Liopyris, Konstantinos and Mishra, Nabin and Kittler, Harald and Halpern, Allan},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1710.05006v3},
  eprint = {1710.05006}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF