Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)
Noel CodellaVeronica RotembergPhilipp TschandlM. E. CelebiStephen W. DuszaDavid GutmanBrian HelbaAadi KallooKonstantinos LiopyrisMichael A. Marchetti
Establishes standard benchmarks and evaluation protocols for automated melanoma detection across 12,500 dermoscopic images, revealing critical generalization failures among top-performing clinical diagnostic models.
Skin cancer is the most common cancer in the United States, and while early detection of melanoma yields a five-year survival rate of up to 99%, late diagnosis causes survival to plunge to 23%. Automated skin image analysis offers the potential to expand early detection, but machine learning tools require rigorous, clinically realistic evaluation before deployment in healthcare. The article details the results of the 2018 International Skin Imaging Collaboration (ISIC) Challenge, which set out to evaluate automated lesion analysis across three key tasks: boundary segmentation, visual attribute detection, and disease classification.
The benchmark used a dataset of more than 12,500 skin images and attracted 299 submissions across the three tasks from global research teams. To better mirror practical clinical conditions, the organizers introduced three core evaluation protocols: a "Thresholded Jaccard" metric that zeroes out segmentation scores failing to meet human observer consistency, a balanced accuracy metric to prevent diagnostic algorithms from overfitting to artificial disease prevalences, and an external test partition sourced from entirely separate international institutions to test model generalizability.
The evaluation revealed several critical findings. First, while top segmentation models achieved high average overlap scores around 0.80, they still failed outright on nearly 10% of images, with failure rates rising to 20–31% on benign seborrheic keratoses. Second, lesion attribute detection yielded exceptionally poor performance across all submissions, reaching a peak score of only 0.473. Third, disease classification achieved strong overall balanced accuracy, reaching up to 0.885 in the top model. However, many models with comparable test scores showed substantial performance drops on external data, demonstrating that strong performance on familiar datasets does not guarantee generalizability to new clinical environments.
These results highlight substantial implications for clinical safety and artificial intelligence regulation. Standard aggregate metrics can mask critical localized failures, and conventional accuracy metrics encourage systems to exploit class imbalances rather than learn robust diagnostic features. The findings confirm that proprietary training data is not mandatory to achieve strong generalizability, but rigorous validation on multi-institutional data is necessary to prevent safety risks from overfitting in real-world deployments.
The article recommends that future biomedical benchmarking initiatives and healthcare regulatory bodies adopt balanced accuracy, failure-penalizing metrics, and multi-partition external validation datasets when evaluating diagnostic software. Additionally, the low performance in attribute detection suggests that research on automated attribute recognition should be paused until clinical definitions mature, or pivoted toward using machine learning to identify novel diagnostic patterns directly. Stakeholders should note that while diagnostic algorithms show high general capability, residual segmentation failure rates and dataset domain shifts remain significant hurdles requiring ongoing oversight.
- Paper: Skin lesion analysis toward melanoma detection: A challenge at the 2017 International symposium on biomedical imaging (ISBI), hosted by the international skin imaging collaboration (ISIC), D. Gutman et al. (2016). This paper establishes the prior 2017 ISIC challenge structure, tasks, and baseline metrics upon which the 2018 ISIC benchmark directly builds and expands.
- Paper: Dermatologist-level classification of skin cancer with deep neural networks, Andre Esteva et al. (2017). This foundational study demonstrates dermatologist-level classification using deep convolutional neural networks on clinical and dermoscopic skin cancer imagery, motivating ISIC's standardized competitive benchmarks.
- Paper: U-Net: Convolutional Networks for Biomedical Image Segmentation, Olaf Ronneberger et al. (2015). This landmark paper introduces the U-Net architecture, which serves as the primary foundational architecture for biomedical image segmentation models evaluated in the ISIC lesion segmentation task.
- Paper: A survey on deep learning in medical image analysis, Geert Litjens et al. (2017). This comprehensive survey outlines early deep learning paradigms and standard evaluation frameworks across medical image analysis tasks including skin lesion diagnosis.
- Paper: Generalised Dice Overlap as a Deep Learning Loss Function for Highly Unbalanced Segmentations, C. Sudre et al. (2017). This paper details Generalized Dice overlap loss functions designed to stabilize medical image segmentation under severe class imbalance, directly relevant to dermoscopic lesion boundary extraction.
- Paper: MultiResUNet : Rethinking the U-Net Architecture for Multimodal Biomedical Image Segmentation, Nabil Ibtehaz et al. (2019). This work develops MultiResUNet to address multiscale boundary ambiguities in medical image segmentation, directly evaluating improvements on benchmark datasets including dermoscopic skin lesions.
- Paper: WILDS: A Benchmark of in-the-Wild Distribution Shifts, Pang Wei Koh et al. (2020). This paper formalizes and expands the study of out-of-distribution generalization across medical imaging and real-world modalities, extending the generalization challenges highlighted in ISIC 2018.
- Paper: Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift, Stephan Rabanser et al. (2019). This study systematically investigates methods to detect dataset shift and algorithm generalization failures, offering quantitative tooling to address the silent failure modes observed in healthcare benchmarks.
- Paper: Kvasir-SEG: A Segmented Polyp Dataset, Debesh Jha et al. (2019). This paper presents the Kvasir-SEG benchmark dataset and baselines, applying standardized biomedical segmentation evaluation protocols similar to ISIC to gastrointestinal polyp detection.
- Paper: Image Segmentation Using Deep Learning: A Survey, Shervin Minaee et al. (2020). This survey reviews subsequent architectural advancements and deep learning paradigms in semantic and medical image segmentation post-dating earlier challenge baselines.
