MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification

Jiancheng YangRui ShiDonglai WeiZequan LiuLin ZhaoBilian KeHanspeter PfisterBingbing Ni

article2021Scientific Data1,504 citations

Presents MedMNIST v2, an accessible suite of 18 standardized 2D and 3D biomedical image datasets paired with extensive baseline benchmarks to facilitate rapid algorithm prototyping and machine learning research without high computational or domain-specific barriers.

Listen

Biomedical image analysis is critical for modern healthcare AI, yet developing generalizable machine learning models remains difficult due to high task diversity, varying data scales, and complex imaging protocols. Furthermore, evaluating standard end-to-end medical systems often entangles the core machine learning algorithm with heavy pre-processing and tuning pipelines, obscuring true algorithmic performance. To resolve these challenges, the article develops MedMNIST v2, a standardized, large-scale, and lightweight benchmark designed to isolate and evaluate the generalization capabilities of machine learning algorithms across both 2D and 3D biomedical image classification tasks.

The benchmark consists of 18 datasets comprising 708,069 2D images and 9,998 3D images spanning modalities such as X-ray, computed tomography, ultrasound, electron microscopy, and dermoscopy. Images are pre-processed into low-resolution, fixed-size formats (28×28 for 2D and 28×28×28 for 3D) and categorized under standard train-validation-test splits. The evaluation framework tests multiple standard deep learning residual networks alongside popular automated machine learning tools, including open-source libraries (auto-sklearn and AutoKeras) and commercial software (Google AutoML Vision), using area under the curve and accuracy metrics across all datasets.

The analysis revealed several important performance trends. First, standard deep neural networks (ResNets) demonstrated high robustness across 2D tasks, closely matching the performance of commercial automated tools (Google AutoML achieved an average area under the curve of 0.927 versus 0.925 for ResNet-18) while outperforming commercial AutoML in classification accuracy. Second, traditional statistical machine learning (auto-sklearn) performed poorly on 2D images (0.878 AUC and 0.722 accuracy) but was surprisingly competitive on small-scale 3D datasets, even surpassing 2.5D deep learning configurations. Third, standard 3D convolutions delivered the strongest average results across 3D tasks, outperforming both 2.5D approaches and deep automated machine learning tools. Finally, higher-resolution inputs (224×224) provided slight performance gains over the baseline 28×28 resolution, but the compact 28-pixel format remained effective for broad algorithmic comparisons.

These findings indicate that specialized and expensive commercial AutoML systems do not offer substantial performance advantages over well-tuned standard neural network baselines in biomedical classification. Additionally, the results show that 2.5D architectures fail to capture adequate spatial context compared to full 3D convolutions for volumetric analysis. Consequently, organizations can significantly reduce computational costs, rapid prototyping timelines, and engineering overhead by using standardized, lightweight datasets to evaluate core algorithms before deploying them to resource-intensive pipelines.

Researchers and machine learning teams are recommended to use MedMNIST v2 as a standard initial benchmark for rapid prototyping, architecture search, and algorithm comparison in biomedical classification tasks. However, the article explicitly emphasizes that because the benchmark relies on substantially downsampled images, it is intended strictly for machine learning research and education; it is not validated or designed for direct clinical diagnosis, as low resolutions may miss critical fine-grained pathological features.

Cover for MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification

Abstract

We introduce MedMNIST v2, a large-scale MNIST-like dataset collection of standardized biomedical images, including 12 datasets for 2D and 6 datasets for 3D. All images are pre-processed into a small size of 28x28 (2D) or 28x28x28 (3D) with the corresponding classification labels so that no background knowledge is required for users. Covering primary data modalities in biomedical images, MedMNIST v2 is designed to perform classification on lightweight 2D and 3D images with various dataset scales (from 100 to 100,000) and diverse tasks (binary/multi-class, ordinal regression, and multi-label). The resulting dataset, consisting of 708,069 2D images and 10,214 3D images in total, could support numerous research / educational purposes in biomedical image analysis, computer vision, and machine learning. We benchmark several baseline methods on MedMNIST v2, including 2D / 3D neural networks and open-source / commercial AutoML tools. The data and code are publicly available at this https URL.

Table of Contents

  • References

Knowls

  1. Knowl 1 — MedMNIST v2 Benchmark Architecture and Scope

    definition

    MedMNIST v2 is a standardized, lightweight benchmark suite of 18 biomedical image classification datasets covering diverse imaging modalities, scales, and task formats. It encompasses 12 two-dimensional (2D) datasets totaling 708,069 images and 6 three-dimensional (3D) datasets totaling 9,998 volumes.

    All images are preprocessed and standardized to fixed low resolutions: 28×2828 \times 28 pixels for 2D images and 28×28×2828 \times 28 \times 28 voxels for 3D volumes. Across its 18 subsets, MedMNIST v2 encompasses dataset scales spanning 10210^2 to 10510^5 samples and four distinct task categories:

    1. Multi-class classification (MC)
    2. Binary classification (BC)
    3. Multi-label binary classification (ML)
    4. Ordinal regression (OR)

    Primary imaging modalities represented include computed tomography (CT), X-ray, optical coherence tomography (OCT), ultrasound, dermoscopy, electron microscopy, blood cell microscopy, fundus photography, and magnetic resonance angiography (MRA).

  2. Knowl 2 — MedMNIST2D Dataset Specifications

    data/table

    MedMNIST2D comprises 12 standardized 2D biomedical image classification datasets processed into 28×2828 \times 28 resolution.

    Name Data Modality Task Type (Classes/Labels) Total Samples Training/Validation/Test Split
    PathMNIST Colon Pathology Multi-Class (9) 107,180 89,996 / 10,004 / 7,180
    ChestMNIST Chest X-Ray Multi-Label BC (14) 112,120 78,468 / 11,219 / 22,433
    DermaMNIST Dermatoscope Multi-Class (7) 10,015 7,007 / 1,003 / 2,005
    OCTMNIST Retinal OCT Multi-Class (4) 109,309 97,477 / 10,832 / 1,000
    PneumoniaMNIST Chest X-Ray Binary-Class (2) 5,856 4,708 / 524 / 624
    RetinaMNIST Fundus Camera Ordinal Regression (5) 1,600 1,080 / 120 / 400
    BreastMNIST Breast Ultrasound Binary-Class (2) 780 546 / 78 / 156
    BloodMNIST Blood Cell Microscope Multi-Class (8) 17,092 11,959 / 1,712 / 3,421
    TissueMNIST Kidney Cortex Microscope Multi-Class (8) 236,386 165,466 / 23,640 / 47,280
    OrganAMNIST Abdominal CT (Axial) Multi-Class (11) 58,850 34,581 / 6,491 / 17,778
    OrganCMNIST Abdominal CT (Coronal) Multi-Class (11) 23,660 13,000 / 2,392 / 8,268
    OrganSMNIST Abdominal CT (Sagittal) Multi-Class (11) 25,221 13,940 / 2,452 / 8,829

    Grayscale datasets are stored as N×28×28N \times 28 \times 28 arrays and RGB datasets as N×28×28×3N \times 28 \times 28 \times 3 arrays, where NN is the sample count.

  3. Knowl 3 — MedMNIST3D Dataset Specifications

    data/table

    MedMNIST3D consists of 6 volumetric biomedical image classification datasets processed into 28×28×2828 \times 28 \times 28 voxels.

    Name Data Modality Task Type (Classes) Total Samples Training/Validation/Test Split
    OrganMNIST3D Abdominal CT Multi-Class (11) 1,743 972 / 161 / 610
    NoduleMNIST3D Chest CT Binary-Class (2) 1,633 1,158 / 165 / 310
    AdrenalMNIST3D Shape from Abdominal CT Binary-Class (2) 1,584 1,188 / 98 / 298
    FractureMNIST3D Chest CT Multi-Class (3) 1,370 1,027 / 103 / 240
    VesselMNIST3D Shape from Brain MRA Binary-Class (2) 1,909 1,335 / 192 / 382
    SynapseMNIST3D Electron Microscope Binary-Class (2) 1,759 1,230 / 177 / 352

    AdrenalMNIST3D and SynapseMNIST3D are datasets introduced directly in MedMNIST v2. AdrenalMNIST3D classifies 3D segmentations of adrenal glands from abdominal CT as normal vs. adrenal mass. SynapseMNIST3D classifies segmented synaptic sites from multi-beam scanning electron microscopy volumes of rat pyramidal neurons as excitatory vs. inhibitory.

  4. Knowl 4 — MedMNIST Standardization and Data Splitting Protocol

    model/method

    MedMNIST v2 standardizes all source data into lightweight uint8 NumPy arrays (.npz files) containing six fixed keys: train_images, train_labels, val_images, val_labels, test_images, and test_labels. Spatial downscaling uses cubic spline interpolation.

    Data partitions follow a strict three-tier hierarchy to prevent data leakage:

    1. If an official source dataset split exists, that split is retained directly.
    2. If the source provides only official training and validation sets, the source validation set is designated as the MedMNIST test set, and the source training set is partitioned 9:19:1 into training and validation sets.
    3. If no official split exists, images are partitioned randomly at the patient level according to a 7:1:27:1:2 ratio (70% train, 10% validation, 20% test).

    Image tensor dimensions are formatted as:

    • N×28×28N \times 28 \times 28 for 2D single-channel grayscale data,
    • N×28×28×3N \times 28 \times 28 \times 3 for 2D three-channel RGB data,
    • N×28×28×28N \times 28 \times 28 \times 28 for 3D single-channel volumetric data,

    where NN is the number of instances in the split. Label tensors are shaped N×1N \times 1 for single-label tasks and N×LN \times L for multi-label tasks with LL target classes (such as L=14L=14 for ChestMNIST).

  5. Knowl 5 — Baseline Deep Learning and AutoML Benchmarking Protocol

    experimental setup

    Standard machine learning baselines and automated machine learning (AutoML) tools are evaluated across MedMNIST v2:

    • 2D Neural Networks: ResNet-18 and ResNet-50 are trained on 28×2828 \times 28 inputs directly as well as on 224×224224 \times 224 upscaled inputs. Grayscale images are converted to 3-channel tensors. Models are optimized using Adam with an initial learning rate η=0.001\eta = 0.001, reduced by a factor of 0.1 at epochs 50 and 75, for a total of 100 epochs with batch size 128, cross-entropy loss, and early stopping on validation loss.
    • 3D Neural Networks: ResNet-18 and ResNet-50 are converted to 2.5D, standard 3D, and Axial-Coronal-Sagittal (ACS) convolutions. Single-channel inputs are replicated to 3 channels. Models are trained with batch size 32, Adam optimizer (initial learning rate 0.0010.001, decayed by 0.1 at epochs 50 and 75; 100 total epochs). For 3D shape datasets (AdrenalMNIST3D and VesselMNIST3D), training samples are regularized by multiplying by a random scalar drawn uniformly from [0,1][0, 1], and test samples are multiplied by 0.50.5.
    • AutoML Baselines:
      • auto-sklearn: Statistical machine learning pipeline search applied to flattened 1D feature arrays. Search time budgets are set to 2 hours for N<10,000N < 10{,}000, 4 hours for N∈[10,000,50,000]N \in [10{,}000, 50{,}000], and 6 hours for N>50,000N > 50{,}000 (and 4 hours for all 3D datasets).
      • AutoKeras: Deep neural architecture search evaluated over 20 trials with 20 epochs per trial, selecting the model with highest validation Area Under the ROC Curve (AUC).
      • Google AutoML Vision: Cloud-based black-box AutoML generating quantized Edge TensorFlow Lite models, allocated 1 to 4 node-hours based on dataset scale (2D only).

    Evaluation is reported using Area Under the ROC Curve (AUC) and Classification Accuracy (ACC), averaged over at least 3 trials.

  6. Knowl 6 — Performance Benchmark on MedMNIST2D

    data/table

    Average classification performance across all 12 datasets in MedMNIST2D illustrates that standard 2D deep residual networks remain highly competitive against commercial and open-source AutoML frameworks.

    Method Average AUC Average ACC
    ResNet-18 (28×2828 \times 28) 0.922 0.819
    ResNet-18 (224×224224 \times 224) 0.925 0.821
    ResNet-50 (28×2828 \times 28) 0.920 0.816
    ResNet-50 (224×224224 \times 224) 0.923 0.821
    auto-sklearn 0.878 0.722
    AutoKeras 0.917 0.813
    Google AutoML Vision 0.927 0.809

    Google AutoML Vision achieves the highest overall average AUC (0.927), whereas ResNet-18 (224×224224 \times 224) and ResNet-50 (224×224224 \times 224) achieve the highest average accuracy (0.821). ResNet-18 consistently outperforms the deeper ResNet-50 when processing native 28×2828 \times 28 inputs (AUC 0.922 vs. 0.920; ACC 0.819 vs. 0.816). Classical statistical machine learning via auto-sklearn underperforms deep learning approaches across 2D medical images (AUC 0.878, ACC 0.722).

  7. Knowl 7 — Performance Benchmark on MedMNIST3D

    data/table

    Average performance across all 6 volumetric datasets in MedMNIST3D demonstrates that full 3D convolutional models achieve the highest general performance, while 2.5D models underperform.

    Method Average AUC Average ACC
    ResNet-18 + 2.5D 0.750 0.731
    ResNet-18 + 3D 0.849 0.767
    ResNet-18 + ACS 0.842 0.775
    ResNet-50 + 2.5D 0.752 0.732
    ResNet-50 + 3D 0.863 0.780
    ResNet-50 + ACS 0.848 0.762
    auto-sklearn 0.815 0.765
    AutoKeras 0.763 0.737

    ResNet-50 with full 3D convolutions achieves the highest overall performance (average AUC 0.863, average ACC 0.780). ACS (Axial-Coronal-Sagittal) convolutions match or approach 3D convolutional performance (AUC 0.842–0.848). In contrast, 2.5D slice-based convolutional formulations suffer substantial performance degradation (AUC 0.750–0.752), performing worse than auto-sklearn (AUC 0.815, ACC 0.765).

  8. Knowl 8 — Comparative Analysis of 2D Slicing vs. 3D Volumetric Processing on Organ CT

    empirical result

    A controlled comparison on the OrganMNIST3D test set evaluates whether 2D planar slicing captures sufficient diagnostic information relative to native 3D volumetric convolutions.

    Method / Training Configuration Test AUC Test ACC
    2D-Input ResNet-18:
    Trained with OrganAMNIST (Axial) 0.995 0.907
    Trained with axial central 60% slices of OrganMNIST3D 0.995 0.916
    Trained with OrganCMNIST (Coronal) 0.991 0.877
    Trained with coronal central 60% slices of OrganMNIST3D 0.992 0.890
    Trained with OrganSMNIST (Sagittal) 0.959 0.697
    Trained with sagittal central 60% slices of OrganMNIST3D 0.963 0.701
    3D-Input ResNet-18:
    2.5D convolution trained with OrganMNIST3D 0.977 0.788
    3D convolution trained with OrganMNIST3D 0.996 0.907
    ACS convolution trained with OrganMNIST3D 0.994 0.900

    2D ResNet-18 models trained on central axial slices achieve test performance (AUC 0.995, ACC 0.916) comparable to, and slightly higher in accuracy than, full 3D convolutional ResNet-18 (AUC 0.996, ACC 0.907). Coronal slice representations perform slightly worse (AUC 0.992, ACC 0.890), while sagittal slice representations exhibit severe performance loss (AUC 0.963, ACC 0.701).

  9. Knowl 9 — Inapplicability of MedMNIST v2 for Direct Clinical Decision Support

    limitation

    MedMNIST v2 is designed exclusively as a machine learning and computer vision benchmark and for educational purposes; it is not intended for clinical diagnostic application. Downsampling biomedical images and volumetric scans to 28×2828 \times 28 pixels or 28×28×2828 \times 28 \times 28 voxels discards high-frequency spatial features, fine morphological boundaries, and microstructural details (e.g., subtle tissue textures, microcalcifications) necessary for accurate clinical decision making.

Coverage note — None was omitted; all key benchmark datasets, preprocessing rules, model architectures, AutoML configurations, baseline results (2D and 3D), representation analyses, and limitations are fully covered.

References

  1. 1.Shen, D., Wu, G. & Suk, H.-I. Deep learning in medical image analysis. Annual review of biomedical engineering 19, 221–248 (2017).
  2. 2.Litjens, G. et al. A survey on deep learning in medical image analysis. Medical image analysis 42, 60–88 (2017).
  3. 3.Liu, X. et al. A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: a systematic review and meta-analysis. The lancet digital health 1, e271–e297 (2019).
  4. 4.Rebuffi, S.-A., Bilen, H. & Vedaldi, A. Learning multiple visual domains with residual adapters. In Advances in Neural Information Processing Systems, 506–516 (2017).
  5. 5.Simpson, A. L. et al. A large annotated medical image dataset for the development and evaluation of segmentation algorithms. Preprint at https://arxiv.org/abs/1902.09063 (2019).
  6. 6.Antonelli, M. et al. The medical segmentation decathlon. Nature communications 13(1), 1-13 (2022).
  7. 7.Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J. & Maier-Hein, K. H. nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods 18, 203–211 (2021).
  8. 8.LeCun, Y., Cortes, C. & Burges, C. Mnist handwritten digit database. http://yann.lecun.com/exdb/mnist/ (2010).
  9. 9.Yang, J., Shi, R. & Ni, B. Medmnist classification decathlon: A lightweight automl benchmark for medical image analysis. In International Symposium on Biomedical Imaging, 191–195 (2021).
  10. 10.He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition, 770–778 (2016).
  11. 11.Feurer, M. et al. Auto-sklearn: efficient and robust automated machine learning. In Automated Machine Learning, 113–134 (Springer, Cham, 2019).
  12. 12.Jin, H., Song, Q. & Hu, X. Auto-keras: An efficient neural architecture search system. In Conference on Knowledge Discovery and Data Mining, 1946–1956 (ACM, 2019).
  13. 13.Qi, K. & Yang, H. Elastic net nonparallel hyperplane support vector machine and its geometrical rationality. IEEE Transactions on Neural Networks and Learning Systems (2021).
  14. 14.Chen, K. et al. Alleviating data imbalance issue with perturbed input during inference. In Conference on Medical Image Computing and Computer Assisted Intervention, 407–417 (Springer, 2021).
  15. 15.Henn, T. et al. A principled approach to failure analysis and model repairment: Demonstration in medical imaging. In Conference on Medical Image Computing and Computer Assisted Intervention, 509–518 (Springer, 2021).
  16. 16.Kather, J. N. et al. Predicting survival from colorectal cancer histology slides using deep learning: A retrospective multicenter study. PLOS Medicine 16, 1–22, https://doi.org/10.1371/journal.pmed.1002730 (2019).
  17. 17.Kather, J. N., Halama, N. & Marx, A. 100,000 histological images of human colorectal cancer and healthy tissue. Zenodo https://doi.org/10.5281/zenodo.1214456 (2018).
  18. 18.Wang, X. et al. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Conference on Computer Vision and Pattern Recognition, 3462–3471 (2017).
  19. 19.Tschandl, P., Rosendahl, C. & Kittler, H. The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Scientific data 5, 180161 (2018).
  20. 20.Tschandl, P. The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Harvard Dataverse https://doi.org/10.7910/DVN/DBW86T (2018).
  21. 21.Codella, N. et al. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic). Preprint at https://arxiv.org/abs/1902.03368v2 (2019).
  22. 22.Kermany, D. S. et al. Identifying medical diagnoses and treatable diseases by image-based deep learning. Cell 172, 1122–1131.e9, https://doi.org/10.1016/j.cell.2018.02.010 (2018).
  23. 23.Kermany, D. S., Zhang, K. & Goldbaum, M. Large dataset of labeled optical coherence tomography (oct) and chest x-ray images https://doi.org/10.17632/rscbjbr9sj.3 (2018).
  24. 24.DeepDRiD. The 2nd diabetic retinopathy–grading and image quality estimation challenge. https://isbi.deepdr.org/data.html (2020).
  25. 25.Al-Dhabyani, W., Gomaa, M., Khaled, H. & Fahmy, A. Dataset of breast ultrasound images. Data in Brief 28, 104863, https://doi.org/10.1016/j.dib.2019.104863 (2020).
  26. 26.Acevedo, A. et al. A dataset of microscopic peripheral blood cell images for development of automatic recognition systems. Data in Brief 30, 105474, https://doi.org/10.1016/j.dib.2020.105474 (2020).
  27. 27.Acevedo, A. et al. A dataset for microscopic peripheral blood cell images for development of automatic recognition systems. Mendeley Data https://doi.org/10.17632/snkd93bnjr.1 (2020).
  28. 28.Woloshuk, A. et al. In situ classification of cell types in human kidney tissue using 3d nuclear staining. Cytometry Part A (2020).
  29. 29.Ljosa, V., Sokolnicki, K. L. & Carpenter, A. E. Annotated high-throughput microscopy image sets for validation. Nature methods 9, 637–637 (2012).
  30. 30.Bilic, P. et al. The liver tumor segmentation benchmark (lits). Medical Image Analysis 84,102680 (2023).
  31. 31.Xu, X. et al. Efficient multiple organ localization in ct image using 3d region proposal network. IEEE Transactions on Medical Imaging 38, 1885–1898 (2019).
  32. 32.Armato, S. G. III et al. The lung image database consortium (lidc) and image database resource initiative (idri): A completed reference database of lung nodules on ct scans. Medical Physics 38, 915–931, https://doi.org/10.1118/1.3528204 (2011).
  33. 33.Jin, L. et al. Deep-learning-assisted detection and segmentation of rib fractures from ct scans: Development and validation of fracnet. EBioMedicine 62, 103106, https://doi.org/10.1016/j.ebiom.2020.103106 (2020).
  34. 34.Yang, X., Xia, D., Kin, T. & Igarashi, T. Intra: 3d intracranial aneurysm dataset for deep learning. In Conference on Computer Vision and Pattern Recognition (2020).
  35. 35.Attene, M. A lightweight approach to repairing digitized polygon meshes. The Visual Computer 26, 1393–1406 (2010).
  36. 36.Dawson-Haggerty et al. trimesh. https://trimsh.org/ (2019).
  37. 37.Wei, D. et al. Mitoem dataset: Large-scale 3d mitochondria instance segmentation from em images. In Conference on Medical Image Computing and Computer Assisted Intervention, 66–76 (Springer, 2020).
  38. 38.Yang, J. et al. Medmnist v2: A large-scale lightweight benchmark for 2d and 3d biomedical image classification. Zenodo https://doi.org/10.5281/zenodo.5208230 (2021).
  39. 39.Harris, C. R. et al. Array programming with numpy. Nature 585, 357–362 (2020).
  40. 40.Kingma, D. P. & Ba, J. Adam: A method for stochastic optimization. Preprint at https://arxiv.org/abs/1412.6980 (2014).
  41. 41.Yang, J. et al. Reinventing 2d convolutions for 3d images. IEEE Journal of Biomedical and Health Informatics 1–1, https://doi.org/10.1109/JBHI.2021.3049452 (2021).
  42. 42.Pedregosa, F. et al. Scikit-learn: Machine learning in python. the Journal of machine Learning research 12, 2825–2830 (2011).
  43. 43.Chollet, F. et al. Keras. https://keras.io (2015).
  44. 44.Bradley, A. P. The use of the area under the roc curve in the evaluation of machine learning algorithms. Pattern recognition 30, 1145–1159 (1997).

Citation

MLA
Yang, J., et al. “MedMNIST V2 - A Large-scale Lightweight Benchmark for 2D and 3D Biomedical Image Classification”. Scientific Data, vol. 10, no. 1, 2023, https://doi.org/10.1038/s41597-022-01721-8.
APA
Yang, J., Shi, R., Wei, D., Liu, Z., Zhao, L., Ke, B., Pfister, H., & Ni, B. (2023). MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification. Scientific Data, 10(1). https://doi.org/10.1038/s41597-022-01721-8
Chicago
Yang, J., R. Shi, D. Wei, et al. 2023. “MedMNIST V2 - A Large-scale Lightweight Benchmark for 2D and 3D Biomedical Image Classification”. Scientific Data 10 (1). https://doi.org/10.1038/s41597-022-01721-8.
Harvard
Yang, J. et al. (2023) “MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification”, Scientific Data, 10(1). Available at: https://doi.org/10.1038/s41597-022-01721-8.
Vancouver
1. Yang J, Shi R, Wei D, Liu Z, Zhao L, Ke B, Pfister H, Ni B (2023) MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification. Scientific Data. https://doi.org/10.1038/s41597-022-01721-8

BibTeX

@article{Yang_2023, title={MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification}, volume={10}, ISSN={2052-4463}, url={http://dx.doi.org/10.1038/s41597-022-01721-8}, DOI={10.1038/s41597-022-01721-8}, number={1}, journal={Scientific Data}, publisher={Springer Science and Business Media LLC}, author={Yang, Jiancheng and Shi, Rui and Wei, Donglai and Liu, Zequan and Zhao, Lin and Ke, Bilian and Pfister, Hanspeter and Ni, Bingbing}, year={2023}, month=Jan }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/