Deep Learning for Medical Image Analysis

Mina RezaeiHaojin YangChristoph Meinel

article2017AIME2,720 citations

Proposes end-to-end deep learning methods for automated brain abnormality detection, recognition, and segmentation to advance clinical neuroimaging analysis.

Listen

The article addresses the challenge of automating analysis of brain magnetic resonance images to support faster and more reliable diagnosis of conditions such as tumors, stroke, multiple sclerosis, and Alzheimer disease. Manual review of these complex scans is time-consuming and variable, creating demand for computer-aided tools that can handle varying lesion sizes, shapes, and locations while processing multi-modal data efficiently.

The work set out to develop and evaluate end-to-end deep learning methods for classifying brain abnormalities, localizing them within images, and performing instance-level segmentation. The author describes a phased research plan centered on convolutional neural networks, with extensions to generative adversarial networks for data augmentation.

Experiments used five public brain MRI datasets totaling several thousand volumes, including healthy subjects from the IXI collection and tumor cases from the BRATS 2015/2016 benchmarks. Models processed axial, coronal, and sagittal slices across multiple modalities, applied data augmentation, and combined local patch features with global image context. Training relied on GPU-accelerated convolution operations and multi-task loss functions.

The classification network reached 95 percent accuracy across five categories on 1,500 images. Detection improved dice scores by roughly 20 percent on BRATS and 30 percent on ISLES data when multi-modal inputs and contextual features were combined, yielding 94.3 percent accuracy and a whole-tumor dice coefficient of 0.72. Segmentation produced accuracies between 84 and 93 percent for tumor core, enhancing core, non-enhancing core, and edema regions.

These outcomes indicate that tailored deep architectures can deliver clinically useful information on lesion size, location, and type with high reliability, potentially reducing diagnostic delays and supporting more consistent treatment planning. Results compare favorably with earlier wavelet- and SVM-based approaches on the same tasks.

The article recommends completing three-dimensional segmentation work, incorporating generative models for additional training data, and deploying an online platform with GPU parallelism for clinical use. Extension to other body regions is planned once brain-focused methods stabilize.

Limitations include reliance on two-dimensional slices, modest dataset sizes for some conditions, and the preliminary nature of results from an early-stage doctoral project; further validation on larger, multi-site data would increase confidence before routine clinical adoption.

arXiv: 1708.08987
Cover for Deep Learning for Medical Image Analysis

Abstract

This report describes my research activities in the Hasso Plattner Institute and summarizes my Ph.D. plan and several novels, end-to-end trainable approaches for analyzing medical images using deep learning algorithm. In this report, as an example, we explore different novel methods based on deep learning for brain abnormality detection, recognition, and segmentation. This report prepared for the doctoral consortium in the AIME-2017 conference.

Table of Contents

  • 1 Introduction
  • 2 Approach
  • 2.1 Classification
  • 2.2 Detection and Localization
  • 2.3 Semantic Segmentation
  • 3 Data Description
  • 4 Future Work
  • References

Knowls

  1. Knowl 1 — Two-Pathway Convolutional Architecture for Brain Lesion Detection and Localization

    model/method

    The brain abnormality detection and localization framework processes multi-modal 2D MRI slices extracted across the axial, coronal, and sagittal anatomical planes. Input slices contain multiple co-registered channels corresponding to distinct MRI modalities (such as T1, T1-contrast, T2, and FLAIR / DWI). The network processes input data through two concurrent pathways:

    1. Global Pathway: A single convolutional layer operates over the entire 2D slice to capture global spatial and contextual anatomical context.

    2. Local Feature Extraction Pathway: Image patches are fed into a fine-tuned VGG-16 backbone that incorporates L2L_2-norm pooling layers in both forward and backward passes. Following the conv5-3 feature layer, candidate regions of interest (RoIs) are pooled via a Spatial Pyramid Pooling (SPP) layer.

    Features from the global pathway and the SPP-pooled local RoIs are concatenated and passed into fully connected layers. The entire network is trained end-to-end using a multi-task loss combining bounding box regression loss and a multi-class Softmax classification loss. On the BRATS-2015 dataset (220 high-grade glioma and 54 low-grade glioma cases), this architecture achieved a Dice similarity coefficient of 0.72, a sensitivity of 0.89, and an accuracy of 94.3% for whole-tumor detection.

  2. Knowl 2 — Seven-Layer Convolutional Network with SVM for Five-Class Brain MRI Classification

    model/method

    A deep convolutional neural network is designed to perform 5-class brain disease classification (distinguishing healthy brains, Alzheimer's disease, high-grade glioma, low-grade glioma, and multiple sclerosis). The network accepts 3-channel input representations constructed from Volume-of-Interest (VOI) slices aligned along the axial, coronal, and sagittal anatomical planes.

    The feature extractor consists of seven convolutional layers and three pooling stages:

    • Average pooling is applied immediately after the 3rd convolutional layer (conv3).
    • Max pooling layers are placed sequentially after the 5th and 6th convolutional layers (conv5 and conv6) to reduce intermediate feature map spatial dimensions.
    • The convolutional trunk feeds into three fully connected layers containing 4,096 individual neurons each.
    • Regularization is applied after the final fully connected layer to mitigate overfitting.
    • A 5-way Support Vector Machine (SVM) replaces the final linear classifier to assign inputs into the five diagnostic categories.
  3. Knowl 3 — Instance-Level Brain Lesion Segmentation Architecture and Multi-Task Cascade

    model/method

    The semantic and instance-level lesion segmentation pipeline adapts the Faster R-CNN architecture into a single-stage, end-to-end trainable multi-task cascade for multi-modal brain MRI. The model operates across four unified functional stages:

    1. Anchor Detection: Candidate regions are proposed using an adapted anchor generation strategy designed for variable lesion shapes and scales.
    2. Bounding Box Regression: Projections refine candidate bounding boxes surrounding suspicious lesion regions.
    3. Mask-Level Estimation: Pixel-level binary masks are estimated across localized lesion boundaries.
    4. Mask Instance-Level Segmentation: Individual lesion instances are segmented and categorized simultaneously.

    The network is optimized using a joint multi-task loss spanning classification, bounding box coordinates, and pixel mask predictions. Training data is expanded using standard geometric/intensity data augmentations (random horizontal/vertical flips, multi-scaling, contrast adjustments) alongside synthetic MR image patches synthesized via a multi-conditional generative adversarial network (MC-GAN).

  4. Knowl 4 — Modality Ablation on Brain Tumor and Stroke Segmentation (BRATS 2016 and ISLES 2016)

    data/table

    The multi-pathway detection and segmentation architecture was evaluated across various single-modality and multi-modality combinations on the BRATS 2016 brain tumor dataset and the ISLES 2016 ischemic stroke dataset. In the table below, the fourth column corresponds to the FLAIR modality for BRATS 2016 and the DWI modality for ISLES 2016.

    T1 T1c T2 FLAIR / DWI BRATS 2016 DSC ISLES 2016 DSC
    ✓ – – – 61.30% 42.00%
    – ✓ – – 33.46% 27.00%
    – – ✓ – 35.67% 39.98%
    – – – ✓ 72.38% 50.71%
    – ✓ ✓ ✓ 82.53% 54.23%
    ✓ ✓ – ✓ 83.53% 54.87%
    ✓ – ✓ ✓ 82.19% 53.09%
    ✓ ✓ ✓ – 86.73% 56.71%
    ✓ ✓ ✓ ✓ 92.44% 57.03%

    The results demonstrate that combining all four complementary MRI modalities yields the highest segmentation accuracy, improving Dice Similarity Coefficients (DSC) by approximately 20% on BRATS 2016 (reaching 92.44%) and 30% on ISLES 2016 (reaching 57.03%) relative to inferior modality subsets.

  5. Knowl 5 — Segmentation Accuracy across High-Grade Glioma Sub-Regions

    data/table

    The instance-level segmentation model was evaluated on high-grade glioma (HGG) sub-compartments across three sequential structural outputs: bounding box localization, pixel mask estimation, and final instance segmentation.

    Stage Tumor Core Enhancing Core Non-Enhancing Core Edema
    Bounding Box Regression 0.92831 0.92740 0.94070 0.90740
    Mask Estimation 0.88845 0.87132 0.92187 0.90862
    Instance Segmentation 0.89562 0.84086 0.89375 0.89789

    The non-enhancing tumor core achieved the highest accuracy across bounding box regression (94.07%) and mask estimation (92.19%), while overall instance segmentation achieved accuracies ranging between 84.09% (enhancing core) and 89.79% (edema).

  6. Knowl 6 — Five-Class Brain Abnormality Classification Performance Comparison

    data/table

    The 7-layer convolutional neural network with an SVM output classifier was evaluated on a five-class brain MRI dataset containing 1,500 total scans across healthy controls, Alzheimer's disease, high-grade glioma, low-grade glioma, and multiple sclerosis, and compared with conventional feature extraction pipelines.

    Method Classes Total Samples Accuracy Sensitivity Specificity
    7-Layer CNN + 5-Way SVM 5 1500 95.07% 0.91 0.87
    Wavelet (DAUB-4) + PCA + SVM-RBF 2 75 98.70% – –
    Wavelet (Haar) + PCA + KNN 2 1500 98.60% – –

    While shallow baseline models (DAUB-4 / Haar wavelets combined with PCA and SVM-RBF or KNN) achieved 98.6%–98.7% accuracy on restricted 2-class binary tasks, the deep architecture achieved 95.07% accuracy, 0.91 sensitivity, and 0.87 specificity on the significantly more complex 5-class problem.

  7. Knowl 7 — Multi-Source MRI Benchmark Datasets for Brain Abnormality Analysis

    experimental setup

    The experiments utilized five distinct multi-modal brain MRI benchmark datasets:

    1. Healthy Brain Images (IXI Dataset): Approximately 600 MRI scans of healthy subjects collected across three hospitals in London in NIFTI (.nii) format, including T1-, T2-, T1-contrast-, and PD-weighted / Diffusion-weighted images in 15 directions.
    2. Glioma Brain Tumor (BRATS 2015 / 2016): Approximately 300 high-grade glioma (HGG) and low-grade glioma (LGG) training volumes and 200 unlabeled test volumes in .mha format, skull-stripped, co-registered to an anatomical template at 1 mm31\text{ mm}^3 isotropic voxel resolution, comprising T1, T1-contrast (T1c), T2, and FLAIR modalities.
    3. Alzheimer's Disease (OASIS Dataset): A cross-sectional collection of 416 subjects aged 18 to 96 years in .hdr format, containing 3 to 4 individual T1-weighted MRI scans per subject obtained in single scan sessions.
    4. Multiple Sclerosis (ISBI 2008 MS Lesion Challenge): 18 MRI scans in .nhdr format provided by the University of Cyprus e-Health Lab, annotated with expert manual segmentations.
    5. Ischemic Stroke (ISLES 2016): Benchmark dataset for stroke lesion segmentation comprising T1, T1c, FLAIR, and DWI modalities (totaling 28,500 slices per modality).

Coverage note — Prospective doctoral research plans and future directions (such as planned 3D extensions, GPU parallelism, and exploratory GAN image translation) were omitted as non-technical roadmap content.

References

  1. 1.Çiçek, Ö., Abdulkadir, A., Lienkamp, S.S., Brox, T., Ronneberger, O.: 3d u-net: learning dense volumetric segmentation from sparse annotation. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 424–432. Springer (2016)
  2. 2.Dai, J., He, K., Sun, J.: Instance-aware semantic segmentation via multi-task network cascades. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 3150–3158 (2016)
  3. 3.Girshick, R.: Fast r-cnn. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 1440–1448 (2015)
  4. 4.Goodfellow, I., Bengio, Y., Courville, A.: Deep learning. MIT Press (2016)
  5. 5.He, K., Zhang, X., Ren, S., Sun, J.: Spatial pyramid pooling in deep convolutional networks for visual recognition. In: Computer Vision–ECCV 2014 (2014)
  6. 6.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (June 2016)
  7. 7.Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: Advances in neural information processing systems. pp. 1097–1105 (2012)
  8. 8.Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 3431–3440 (2015)
  9. 9.Menze, B.H., Jakab, A., Bauer, S., Kalpathy-Cramer, J., Farahani, K., Kirby, J., Burren, Y., Porz, N., Slotboom, J., Wiest, R., et al.: The multimodal brain tumor image segmentation benchmark (brats). IEEE transactions on medical imaging 34(10), 1993–2024 (2015)
  10. 10.Redmon, J., Divvala, S., Girshick, R., Farhadi, A.: You only look once: Unified, real-time object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 779–788 (2016)
  11. 11.Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. In: Advances in neural information processing systems. pp. 91–99 (2015)
  12. 12.Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)

Citation

MLA
Rezaei, M., et al. “Deep Learning for Medical Image Analysis”. arXiv, 2017, https://doi.org/10.48550/arxiv.1708.08987.
APA
Rezaei, M., Yang, H., & Meinel, C. (2017). Deep Learning for Medical Image Analysis. arXiv. https://doi.org/10.48550/arxiv.1708.08987
Chicago
Rezaei, M., H. Yang, and C. Meinel. 2017. “Deep Learning for Medical Image Analysis”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.1708.08987.
Harvard
Rezaei, M., Yang, H. and Meinel, C. (2017) “Deep Learning for Medical Image Analysis”. arXiv. Available at: https://doi.org/10.48550/arxiv.1708.08987.
Vancouver
1. Rezaei M, Yang H, Meinel C (2017) Deep Learning for Medical Image Analysis. https://doi.org/10.48550/arxiv.1708.08987

BibTeX

@misc{https://doi.org/10.48550/arxiv.1708.08987,
  doi = {10.48550/ARXIV.1708.08987},
  url = {https://arxiv.org/abs/1708.08987},
  author = {Rezaei, Mina and Yang, Haojin and Meinel, Christoph},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Deep Learning for Medical Image Analysis},
  publisher = {arXiv},
  year = {2017},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors