Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)

Noel CodellaVeronica RotembergPhilipp TschandlM. E. CelebiStephen W. DuszaDavid GutmanBrian HelbaAadi KallooKonstantinos LiopyrisMichael A. Marchetti

article2019arXiv1,876 citations

Establishes standard benchmarks and evaluation protocols for automated melanoma detection across 12,500 dermoscopic images, revealing critical generalization failures among top-performing clinical diagnostic models.

Listen

Skin cancer is the most common cancer in the United States, and while early detection of melanoma yields a five-year survival rate of up to 99%, late diagnosis causes survival to plunge to 23%. Automated skin image analysis offers the potential to expand early detection, but machine learning tools require rigorous, clinically realistic evaluation before deployment in healthcare. The article details the results of the 2018 International Skin Imaging Collaboration (ISIC) Challenge, which set out to evaluate automated lesion analysis across three key tasks: boundary segmentation, visual attribute detection, and disease classification.

The benchmark used a dataset of more than 12,500 skin images and attracted 299 submissions across the three tasks from global research teams. To better mirror practical clinical conditions, the organizers introduced three core evaluation protocols: a "Thresholded Jaccard" metric that zeroes out segmentation scores failing to meet human observer consistency, a balanced accuracy metric to prevent diagnostic algorithms from overfitting to artificial disease prevalences, and an external test partition sourced from entirely separate international institutions to test model generalizability.

The evaluation revealed several critical findings. First, while top segmentation models achieved high average overlap scores around 0.80, they still failed outright on nearly 10% of images, with failure rates rising to 20–31% on benign seborrheic keratoses. Second, lesion attribute detection yielded exceptionally poor performance across all submissions, reaching a peak score of only 0.473. Third, disease classification achieved strong overall balanced accuracy, reaching up to 0.885 in the top model. However, many models with comparable test scores showed substantial performance drops on external data, demonstrating that strong performance on familiar datasets does not guarantee generalizability to new clinical environments.

These results highlight substantial implications for clinical safety and artificial intelligence regulation. Standard aggregate metrics can mask critical localized failures, and conventional accuracy metrics encourage systems to exploit class imbalances rather than learn robust diagnostic features. The findings confirm that proprietary training data is not mandatory to achieve strong generalizability, but rigorous validation on multi-institutional data is necessary to prevent safety risks from overfitting in real-world deployments.

The article recommends that future biomedical benchmarking initiatives and healthcare regulatory bodies adopt balanced accuracy, failure-penalizing metrics, and multi-partition external validation datasets when evaluating diagnostic software. Additionally, the low performance in attribute detection suggests that research on automated attribute recognition should be paused until clinical definitions mature, or pivoted toward using machine learning to identify novel diagnostic patterns directly. Stakeholders should note that while diagnostic algorithms show high general capability, residual segmentation failure rates and dataset domain shifts remain significant hurdles requiring ongoing oversight.

arXiv: 1902.03368
Cover for Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)

Abstract

This work summarizes the results of the largest skin image analysis challenge in the world, hosted by the International Skin Imaging Collaboration (ISIC), a global partnership that has organized the world's largest public repository of dermoscopic images of skin. The challenge was hosted in 2018 at the Medical Image Computing and Computer Assisted Intervention (MICCAI) conference in Granada, Spain. The dataset included over 12,500 images across 3 tasks. 900 users registered for data download, 115 submitted to the lesion segmentation task, 25 submitted to the lesion attribute detection task, and 159 submitted to the disease classification task. Novel evaluation protocols were established, including a new test for segmentation algorithm performance, and a test for algorithm ability to generalize. Results show that top segmentation algorithms still fail on over 10% of images on average, and algorithms with equal performance on test data can have different abilities to generalize. This is an important consideration for agencies regulating the growing set of machine learning tools in the healthcare domain, and sets a new standard for future public challenges in healthcare.

Table of Contents

  • 1 Introduction
  • 2 Methods
  • 2.1 Part 1: Lesion Segmentation
  • 2.2 Part 2: Lesion Attribute Detection
  • 2.3 Part 3: Lesion Disease Classification
  • 3 Results
  • 3.1 Part 1: Lesion Segmentation
  • 3.2 Part 2: Lesion Attribute Detection
  • 3.3 Part 3: Lesion Disease Classification
  • 4 Discussion & Conclusion
  • References
  • 5 Supplementary Material

Knowls

  1. Knowl 1 — Thresholded Jaccard Metric for Lesion Segmentation

    equation

    The Thresholded Jaccard metric is an evaluation metric designed to measure segmentation quality while severely penalizing segmentations that deviate beyond standard inter-observer and intra-observer variability. For a predicted binary mask PP and ground-truth binary mask GG, the standard Jaccard index is defined as:

    J(P,G)=∣P∩G∣∣P∪G∣J(P, G) = \frac{|P \cap G|}{|P \cup G|}

    Given a quality failure threshold T∈[0,1]T \in [0, 1], the Thresholded Jaccard index JT(P,G)J_T(P, G) for a single image is defined piecewise as:

    JT(P,G)={J(P,G),if J(P,G)≥T0,if J(P,G)<TJ_T(P, G) = \begin{cases} J(P, G), & \text{if } J(P, G) \ge T \\ 0, & \text{if } J(P, G) < T \end{cases}

    The dataset-level score is computed as the arithmetic mean of JT(P,G)J_T(P, G) across all evaluated test images. In the ISIC 2018 challenge, the threshold was set to T=0.65T = 0.65. This threshold was established from empirical human expert studies showing an average inter-observer Jaccard agreement of 0.7860.786 with a range of 0.1180.118 (lowest agreement 0.7540.754); rounding the lowest expert agreement to 0.750.75 and subtracting one rounded range (0.100.10) sets T=0.65T = 0.65 to provide high specificity in identifying true segmentation failures.

  2. Knowl 2 — ISIC 2018 Challenge Tasks and Dataset Organization

    experimental setup

    The 2018 International Skin Imaging Collaboration (ISIC) Challenge comprises three distinct dermoscopic image analysis tasks:

    1. Part 1: Lesion Segmentation: 2,594 dermoscopic images with ground truth binary masks for training, 100 validation images, and 1,000 test images without released masks.
    2. Part 2: Lesion Attribute Detection: 2,594 training images accompanied by 12,970 binary segmentation masks for 5 specific dermoscopic attributes: pigment network, negative network, streaks, milia-like cysts, and globules. The validation and held-out test partitions contain 100 and 1,000 images, respectively.
    3. Part 3: Lesion Disease Classification: 10,015 dermoscopic training images categorized into 7 diagnostic disease classes: melanoma (MEL), melanocytic nevi (NV), basal cell carcinoma (BCC), actinic keratosis / intraepithelial carcinoma (AKIEC), benign keratosis (BKL, including seborrheic keratoses, solar lentigines, and lichen-planus like keratoses), dermatofibroma (DF), and vascular lesions (VASC). The validation set contains 193 images, and the held-out test set contains 1,512 images partitioned into:
      • An Internal Test Set (N=1,196N = 1,196 images) from the two medical centers in Austria and Australia represented in the training data.
      • An External Test Set (N=316N = 316 images) sourced exclusively from institutions in Turkey, New Zealand, Sweden, and Argentina not present in the training set.
  3. Knowl 3 — Dataset-Level Aggregated Jaccard Metric for Attribute Detection

    equation

    In dermoscopic attribute segmentation, certain localized clinical patterns (such as streaks or negative networks) are absent in a subset of images. Standard per-image Jaccard calculation on an image where both the ground truth and predicted attribute masks are empty results in an undefined division by zero (0/00/0).

    To address this, the evaluation metric computes global True Positives (TP), False Positives (FP), and False Negatives (FN) summed pixel-wise across all NN images in the evaluation set for each attribute kk:

    Jglobal(k)=∑i=1NTPi(k)∑i=1NTPi(k)+∑i=1NFPi(k)+∑i=1NFNi(k)J_{\text{global}}^{(k)} = \frac{\sum_{i=1}^{N} \text{TP}_i^{(k)}}{\sum_{i=1}^{N} \text{TP}_i^{(k)} + \sum_{i=1}^{N} \text{FP}_i^{(k)} + \sum_{i=1}^{N} \text{FN}_i^{(k)}}

    The overall benchmark score is the arithmetic mean of Jglobal(k)J_{\text{global}}^{(k)} across the 5 evaluated dermoscopic attributes.

  4. Knowl 4 — Balanced Accuracy and Multi-Center Generalization Protocol for Classification

    model/method

    For multi-class lesion diagnosis, evaluation utilizes Balanced Accuracy (BACC), defined as the macro-averaged recall across all CC diagnostic classes after mutually exclusive multiclass prediction:

    BACC=1C∑c=1CTPcTPc+FNc\text{BACC} = \frac{1}{C} \sum_{c=1}^{C} \frac{\text{TP}_c}{\text{TP}_c + \text{FN}_c}

    where TPc\text{TP}_c and FNc\text{FN}_c denote true positives and false negatives for class cc, respectively (C=7C = 7). BACC removes bias from prevalence disparities (such as over-representation of benign nevi or melanomas in benchmark datasets relative to clinical incidence).

    To evaluate out-of-distribution robustness and domain shift, the test benchmark evaluates models separately on an internal cohort (same clinical institutions as training data) and an external cohort (geographically independent clinics not represented during training), measuring the generalization gap Δgen=BACCinternal−BACCexternal\Delta_{\text{gen}} = \text{BACC}_{\text{internal}} - \text{BACC}_{\text{external}}.

  5. Knowl 5 — Top-Ranked Lesion Segmentation Benchmark Results and Diagnosis Disparities

    data/table

    Evaluating 112 submission entries on the 1,000-image segmentation test set demonstrates that while leading deep learning models achieve average standard Jaccard scores above 0.83 (surpassing average expert agreement of 0.786), they maintain failure rates (Jaccard <0.65< 0.65) between 7.9% and 9.3% overall, with pronounced failure on seborrheic keratoses.

    ALL Melanoma (MEL) Seborrheic Keratoses (SEBK) Benign Nevi (NEVI)
    Rank F TJ J F TJ J F TJ J F TJ J
    1 0.093 0.802 0.838 0.095 0.792 0.832 0.310 0.577 0.698 0.066 0.832 0.856
    2 0.079 0.801 0.838 0.090 0.782 0.830 0.195 0.667 0.743 0.063 0.820 0.851
    3 0.083 0.799 0.834 0.100 0.782 0.826 0.299 0.585 0.706 0.053 0.829 0.852
    4 0.085 0.798 0.838 0.090 0.792 0.839 0.207 0.656 0.740 0.069 0.817 0.848
    5 0.084 0.796 0.837 0.095 0.799 0.849 0.195 0.670 0.738 0.067 0.811 0.845

    Notation: F = Failure rate (fraction of images with Jaccard <0.65< 0.65), TJ = Thresholded Jaccard, J = Standard Jaccard index. ALL denotes the entire 1,000-image test set.

  6. Knowl 6 — Benchmark Results for Lesion Attribute Detection

    empirical result

    Across 26 submitted algorithms evaluated on the 1,000-image held-out test set for Part 2 (detecting globules, milia-like cysts, negative networks, pigment networks, and streaks), algorithm performance was uniformly low. The highest-performing submission achieved an average global Jaccard score of only 0.473 across all 5 dermoscopic attributes.

    This low accuracy reflects known high clinical ambiguity and poor inter-observer agreement among dermatologists on localized dermoscopic criteria, indicating that supervised localization of manually annotated morphological criteria remains a challenge.

  7. Knowl 7 — Top-Ranked Multi-Class Lesion Classification Performance Across Test Partitions

    data/table

    141 entries were evaluated for multi-class classification across 7 disease states on the 1,512-image test set. The top 5 models achieved overall Balanced Accuracy (BACC) between 0.845 and 0.885, with individual category Area Under the ROC Curve (AUC) consistently exceeding 0.94 across classes.

    Partition Rank ACC BACC MEL NV BCC AKIEC BKL DF VASC
    1 0.851 0.885 0.949 0.979 0.997 0.987 0.974 0.992 1.000
    2 0.850 0.882 0.946 0.981 0.997 0.985 0.977 0.990 1.000
    ALL 3 0.827 0.871 0.948 0.978 0.996 0.981 0.971 0.986 0.999
    4 0.896 0.856 0.959 0.983 0.995 0.995 0.990 0.987 0.999
    5 0.884 0.845 0.945 0.974 0.992 0.988 0.969 0.982 0.998
    1 0.841 0.875 0.945 0.978 0.997 0.986 0.969 0.992 1.000
    2 0.842 0.875 0.941 0.980 0.997 0.983 0.972 0.993 0.999
    INT 3 0.820 0.875 0.944 0.977 0.997 0.980 0.965 0.994 0.999
    4 0.907 0.854 0.961 0.982 0.999 0.995 0.988 0.973 0.999
    5 0.894 0.841 0.939 0.973 0.997 0.988 0.963 0.973 0.998
    1 0.886 0.925 0.970 0.984 0.998 1.000 0.984 0.993 1.000
    2 0.880 0.911 0.966 0.986 0.998 1.000 0.987 0.989 1.000
    EXT 3 0.854 0.894 0.970 0.982 0.993 1.000 0.983 0.984 0.999
    4 0.854 0.866 0.984 0.984 0.987 0.997 0.994 0.995 1.000
    5 0.848 0.864 0.976 0.971 0.978 0.998 0.978 0.984 0.999

    Notation: ALL = entire test set (N=1,512N=1,512), INT = internal test partition (N=1,196N=1,196), EXT = external test partition (N=316N=316). ACC = unweighted accuracy, BACC = balanced accuracy. Class AUC columns: MEL = Melanoma, NV = Melanocytic Nevi, BCC = Basal Cell Carcinoma, AKIEC = Actinic Keratosis, BKL = Benign Keratosis, DF = Dermatofibroma, VASC = Vascular Lesion.

  8. Knowl 8 — Generalization Disparity and Metric Discrepancies in Lesion Classification

    empirical result

    Analysis of the 141 classification submissions reveals key dynamics in machine learning evaluation for dermatological diagnosis:

    1. Overfitting to Institution vs Generalization: While the majority of algorithms exhibited overfitting (higher BACC on internal test data than on external test data), top-performing systems maintained or improved performance on external data (Rank 1 achieved 0.925 BACC on external vs 0.875 on internal) without requiring proprietary training datasets.
    2. Divergent Generalization from Equal Overall Performance: Algorithms with virtually identical full test set performance displayed widely varying internal-versus-external performance differences.
    3. Ranking Divergence Across Metrics: Submissions ranked substantially differently under Balanced Accuracy versus standard unweighted Accuracy (R2=0.7423R^2 = 0.7423) and Mean AUC (R2=0.6655R^2 = 0.6655). Unweighted accuracy favored models biased toward high-prevalence classes, while Mean AUC incorporated low-sensitivity regions of ROC curves that do not reflect clinical operating requirements.

Coverage note — None was omitted; all key challenge tasks, novel metrics (Thresholded Jaccard, global Jaccard, BACC), dataset configurations, and empirical benchmark results were extracted.

References

  1. 1.Guy GP, Machlin S, Ekwueme DU, Yabroff KR. Prevalence and costs of skin cancer treatment in the US, 2002—2006 and 2007—2011. Am J Prev Med. 2015;48:183—7.
  2. 2.Cancer Facts and Figures 2018. American Cancer Society. https://www.cancer.org/content/dam/cancer-org/research/cancer-facts-and-statistics/annual-cancer-facts-and-figures/2018/cancer-facts-and-figures-2018.pdf. Accessed May 3, 2018.
  3. 3.Gutman D, Codella NCF, Celebi E, Helba, B, Marchetti M, Mishra N, Halpern A. “Skin Lesion Analysis toward Melanoma Detection: A Challenge at the International Symposium on Biomedical Imaging (ISBI) 2016, hosted by the International Skin Imaging Collaboration (ISIC)”. eprint arXiv:1605.01397. 2016.
  4. 4.Marchetti M, et al. “Results of the 2016 International Skin Imaging Collaboration International Symposium on Biomedical Imaging challenge: Comparison of the accuracy of computer algorithms to dermatologists for the diagnosis of melanoma from dermoscopic images”. J Am Acad Dermatol. 2018 Feb;78(2):270—277
  5. 5.Codella N, et al. “Skin Lesion Analysis toward Melanoma Detection: A Challenge at the International Symposium on Biomedical Imaging (ISBI) 2017, hosted by the International Skin Imaging Collaboration (ISIC)”. IEEE International Symposium of Biomedical Imaging (ISBI) 2018.
  6. 6.Codella NCF, Nguyen B, Pankanti S, Gutman D, Helba B, Halpern A, Smith JR. “Deep learning ensembles for melanoma recognition in dermoscopy images” In: IBM Journal of Research and Development, vol. 61, no. 4/5, 2017.
  7. 7.Esteva A,Kuprel B, Novoa RA, Ko J, Swetter SM, Blau HM, Thrun S. “Dermatologist-level classification of skin cancer with deep neural networks”. Nature, vol 542, pp 115—118. 2017.
  8. 8.Menegola A, Tavares J, Fornaciali M, Li LT, Avila S, Valle E. “RECOD Titans at ISIC Challenge 2017”. 2017 International Symposium on Biomedical Imaging (ISBI) Challenge on Skin Lesion Analysis Towards Melanoma Detection. Available: https://arxiv.org/pdf/1703.04819.pdf
  9. 9.Diaz, I.G. “Incorporating the Knowledge of Dermatologists to Convolutional Neural Networks for the Diagnosis of Skin Lesions. 2017 International Symposium on Biomedical Imaging (ISBI) Challenge on Skin Lesion Analysis Towards Melanoma Detection.” Available: https://arxiv.org/abs/1703.01976
  10. 10.Tschandl P, Sinz C, Kittler H. “Domain-specific classification-pretrained fully convolutional network encoders for skin lesion segmentation.” Computers in Biology and Medicine, vol 104, pp 111—116, 2019
  11. 11.Tschandl P, Rosendahl C, Kittler H. “The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions.” Sci Data. 2018 Aug 14;5:180161.
  12. 12.Carrera C, et al. “Validity and Reliability of Dermoscopic Criteria Used to Differentiate Nevi From Melanoma: A Web-Based International Dermoscopy Society Study.” JAMA Dermatol. 2016 July 01; 152(7): 798—806.

Citation

MLA
Codella, N., et al. “Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)”. arXiv, 2019, http://arxiv.org/abs/1902.03368v2.
APA
Codella, N., Rotemberg, V., Tschandl, P., Celebi, M. E., Dusza, S., Gutman, D., Helba, B., Kalloo, A., Liopyris, K., Marchetti, M., Kittler, H., & Halpern, A. (2019). Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC). arXiv. http://arxiv.org/abs/1902.03368v2
Chicago
Codella, N., V. Rotemberg, P. Tschandl, et al. 2019. “Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)”. arXiv. http://arxiv.org/abs/1902.03368v2.
Harvard
Codella, N. et al. (2019) “Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1902.03368v2.
Vancouver
1. Codella N, Rotemberg V, Tschandl P, et al (2019) Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC). arXiv

BibTeX

@article{codella2019skin,
  title = {Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)},
  author = {Codella, Noel and Rotemberg, Veronica and Tschandl, Philipp and Celebi, M. Emre and Dusza, Stephen and Gutman, David and Helba, Brian and Kalloo, Aadi and Liopyris, Konstantinos and Marchetti, Michael and Kittler, Harald and Halpern, Allan},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1902.03368v2},
  eprint = {1902.03368}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors