MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification
Jiancheng YangRui ShiDonglai WeiZequan LiuLin ZhaoBilian KeHanspeter PfisterBingbing Ni
Presents MedMNIST v2, an accessible suite of 18 standardized 2D and 3D biomedical image datasets paired with extensive baseline benchmarks to facilitate rapid algorithm prototyping and machine learning research without high computational or domain-specific barriers.
Biomedical image analysis is critical for modern healthcare AI, yet developing generalizable machine learning models remains difficult due to high task diversity, varying data scales, and complex imaging protocols. Furthermore, evaluating standard end-to-end medical systems often entangles the core machine learning algorithm with heavy pre-processing and tuning pipelines, obscuring true algorithmic performance. To resolve these challenges, the article develops MedMNIST v2, a standardized, large-scale, and lightweight benchmark designed to isolate and evaluate the generalization capabilities of machine learning algorithms across both 2D and 3D biomedical image classification tasks.
The benchmark consists of 18 datasets comprising 708,069 2D images and 9,998 3D images spanning modalities such as X-ray, computed tomography, ultrasound, electron microscopy, and dermoscopy. Images are pre-processed into low-resolution, fixed-size formats (28×28 for 2D and 28×28×28 for 3D) and categorized under standard train-validation-test splits. The evaluation framework tests multiple standard deep learning residual networks alongside popular automated machine learning tools, including open-source libraries (auto-sklearn and AutoKeras) and commercial software (Google AutoML Vision), using area under the curve and accuracy metrics across all datasets.
The analysis revealed several important performance trends. First, standard deep neural networks (ResNets) demonstrated high robustness across 2D tasks, closely matching the performance of commercial automated tools (Google AutoML achieved an average area under the curve of 0.927 versus 0.925 for ResNet-18) while outperforming commercial AutoML in classification accuracy. Second, traditional statistical machine learning (auto-sklearn) performed poorly on 2D images (0.878 AUC and 0.722 accuracy) but was surprisingly competitive on small-scale 3D datasets, even surpassing 2.5D deep learning configurations. Third, standard 3D convolutions delivered the strongest average results across 3D tasks, outperforming both 2.5D approaches and deep automated machine learning tools. Finally, higher-resolution inputs (224×224) provided slight performance gains over the baseline 28×28 resolution, but the compact 28-pixel format remained effective for broad algorithmic comparisons.
These findings indicate that specialized and expensive commercial AutoML systems do not offer substantial performance advantages over well-tuned standard neural network baselines in biomedical classification. Additionally, the results show that 2.5D architectures fail to capture adequate spatial context compared to full 3D convolutions for volumetric analysis. Consequently, organizations can significantly reduce computational costs, rapid prototyping timelines, and engineering overhead by using standardized, lightweight datasets to evaluate core algorithms before deploying them to resource-intensive pipelines.
Researchers and machine learning teams are recommended to use MedMNIST v2 as a standard initial benchmark for rapid prototyping, architecture search, and algorithm comparison in biomedical classification tasks. However, the article explicitly emphasizes that because the benchmark relies on substantially downsampled images, it is intended strictly for machine learning research and education; it is not validated or designed for direct clinical diagnosis, as low resolutions may miss critical fine-grained pathological features.
- Paper: Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms, Han Xiao et al. (2017). It introduces the standardized 28x28 lightweight benchmark paradigm that MedMNIST directly adapts and extends to biomedical imaging.
- Paper: A survey on deep learning in medical image analysis, Geert Litjens et al. (2017). It provides a foundational overview of deep learning across diverse medical imaging modalities and clinical tasks that MedMNIST standardizes into lightweight benchmarks.
- Paper: CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison, Jeremy Irvin et al. (2019). It establishes large-scale chest radiograph multi-label classification benchmarks that serve as primary source data and motivation for MedMNIST's chest X-ray subsets.
- Paper: ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases, Xiaosong Wang et al. (2017). It introduces hospital-scale thoracic disease classification datasets that form the foundational clinical data standardized in MedMNIST's 2D radiology benchmarks.
- Paper: Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC), Noel Codella et al. (2019). It details the standard multi-class dermoscopic lesion benchmarks that MedMNIST processes into its DermaMNIST collection.
- Paper: Skin lesion analysis toward melanoma detection: A challenge at the 2017 International symposium on biomedical imaging (ISBI), hosted by the international skin imaging collaboration (ISIC), David A. Gutman et al. (2016). It provides the public dermoscopy classification challenge data and protocols that underpin the skin lesion subset of MedMNIST.
- Paper: 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation, Özgün Çiçek et al. (2016). It introduces standard 3D convolutional network architectures that MedMNIST evaluates on its 3D volumetric biomedical classification benchmarks.
- Paper: MobileNetV2: Inverted Residuals and Linear Bottlenecks, Mark Sandler et al. (2018). It defines lightweight convolutional architectures commonly benchmarked as foundational baselines on standardized image datasets like MedMNIST.
- Paper: Segment anything in medical images, Jun Ma et al. (2023). It advances multimodal biomedical AI from standardized classification benchmarks toward universal foundation models across multi-organ and multi-modality clinical tasks.
- Paper: LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day, Chunyuan Li et al. (2023). It builds upon standardized multi-modal biomedical image understanding to train general conversational vision-language assistants for diverse clinical domains.
