MultiResUNet : Rethinking the U-Net Architecture for Multimodal Biomedical Image Segmentation

Nabil IbtehazMohammad Sohel Rahman

article2019Neural Networks2,123 citations

Proposes MultiResUNet, a modified U-Net architecture that incorporates multiresolution feature extraction to consistently outperform standard models on challenging multimodal biomedical image segmentation benchmarks.

Listen

Automated medical image segmentation is critical for clinical diagnosis and treatment planning, but manual annotation remains labor-intensive and error-prone. While the standard U-Net architecture is widely used as a baseline for deep learning segmentation, it often struggles when dealing with targets of varying scales, faint object boundaries, background noise, and feature discrepancies across network levels.

The article develops and evaluates MultiResUNet, an enhanced deep neural network designed to address these architectural limitations and improve multimodal biomedical image segmentation performance.

To address scale variation without massive memory costs, the model replaces standard convolutional sequences with compact, multi-resolution processing blocks. It also modifies shortcut connections with residual paths to bridge feature representation gaps between early and late network stages. The authors benchmarked the architecture against the standard U-Net across five diverse public medical imaging datasets—spanning endoscopy, dermoscopy, fluorescence microscopy, magnetic resonance imaging (MRI), and electron microscopy—using five-fold cross-validation without specialized pre- or post-processing.

Across all benchmarks, the proposed model consistently outperformed standard U-Net while using slightly fewer model parameters. The most pronounced gains appeared on difficult imaging modalities, showing relative performance improvements of roughly 10.1% on endoscopy polyp images and 5.1% on dermoscopy skin lesion data. The model also achieved relative gains of 2.6% on fluorescence microscopy, 1.4% on 3D MRI brain scans, and 0.6% on electron microscopy. Beyond aggregate metrics, the architecture converged in fewer training epochs, demonstrated smaller performance variance across folds, and showed superior resilience against visual noise, image perturbations, and outlier artifacts.

These findings indicate that integrating multi-scale feature analysis and intermediate processing along shortcut paths significantly boosts segmentation accuracy and diagnostic reliability. By maintaining low parameter counts and high training efficiency, the approach reduces computational burdens while improving safety and consistency when segmenting complex or faint anatomical targets.

Organizations developing medical diagnostic software should consider adopting the proposed architecture as a drop-in enhancement over traditional U-Net foundations. Before deployment in production clinical pipelines, future initiatives should test hyperparameter configurations across additional imaging modalities and integrate the model with domain-specific pre-processing and post-processing tools.

The reported findings are supported by consistent cross-validation across diverse imaging modalities, though downsampling input image resolutions to meet GPU memory limits introduces minor constraints. The overall evidence provides high confidence in the model's reliability for multimodal medical segmentation tasks.

Cover for MultiResUNet : Rethinking the U-Net Architecture for Multimodal Biomedical Image Segmentation

Abstract

In recent years Deep Learning has brought about a breakthrough in Medical Image Segmentation. U-Net is the most prominent deep network in this regard, which has been the most popular architecture in the medical imaging community. Despite outstanding overall performance in segmenting multimodal medical images, from extensive experimentations on challenging datasets, we found out that the classical U-Net architecture seems to be lacking in certain aspects. Therefore, we propose some modifications to improve upon the already state-of-the-art U-Net model. Hence, following the modifications we develop a novel architecture MultiResUNet as the potential successor to the successful U-Net architecture. We have compared our proposed architecture MultiResUNet with the classical U-Net on a vast repertoire of multimodal medical images. Albeit slight improvements in the cases of ideal images, a remarkable gain in performance has been attained for challenging images. We have evaluated our model on five different datasets, each with their own unique challenges, and have obtained a relative improvement in performance of 10.15%, 5.07%, 2.63%, 1.41%, and 0.62% respectively.

Table of Contents

  • 1 Introduction
  • 2 Overview of the UNet Architecture
  • 3 Motivations and High Level Considerations
  • 3.1 Variation of Scale in Medical Images
  • 3.2 Probable Semantic Gap between the Corresponding Levels of Encoder-Decoder
  • 4 Proposed Architecture
  • 5 Datasets
  • 5.1 Fluorescence Microscopy Image
  • 5.2 Electron Microscopy Image
  • 5.3 Dermoscopy Image
  • 5.4 Endoscopy Image
  • 5.5 Magnetic Resonance Image
  • 6 Experiments
  • 6.1 Baseline Model
  • 6.2 Pre-processing / Post-processing
  • 6.3 Training Methodology
  • 6.4 Evaluation Metric
  • 6.5 kk-Fold Cross Validation
  • 7 Results
  • 7.1 MultiResUNet Consistently Outperforms U-Net
  • 7.2 MultiResUNet can Obtain Better Results in Less Number of Epochs
  • 7.3 MultiResUNet Delineates Faint Boundaries Better
  • 7.4 MultiResUNet is More Immune to Perturbations
  • 7.5 MultiResUNet is More Reliable Against Outliers
  • 7.6 Note on Segmenting the Majority Class
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — MultiRes Block

    model/method

    The MultiRes block is a multi-scale convolutional module designed to capture spatial features across different context sizes while maintaining computational and memory efficiency. Instead of computing parallel 3×33 \times 3, 5×55 \times 5, and 7×77 \times 7 convolutions as in standard Inception-like blocks, the MultiRes block factorizes the larger kernels into a sequential cascade of three 3×33 \times 3 convolutional layers. The outputs of these three successive 3×33 \times 3 convolutions correspond to effective receptive fields of 3×33 \times 3, 5×55 \times 5, and 7×77 \times 7 respectively.

    The feature maps produced by all three 3×33 \times 3 convolutional layers are concatenated along the channel dimension. To constrain memory expansion caused by consecutive convolutional layers, the filter count is distributed non-uniformly across the three stages rather than being equal:

    • First 3×33 \times 3 convolution: ⌊W/6⌋\lfloor W / 6 \rfloor filters
    • Second 3×33 \times 3 convolution: ⌊W/3⌋\lfloor W / 3 \rfloor filters
    • Third 3×33 \times 3 convolution: ⌊W/2⌋\lfloor W / 2 \rfloor filters

    where WW is the total channel capacity parameter for the block. In addition, a residual shortcut with a 1×11 \times 1 convolution connects the block input directly to the concatenated output via element-wise addition, where the 1×11 \times 1 convolution projects the input channel dimension to match the total concatenated channel count (⌊W/6⌋+⌊W/3⌋+⌊W/2⌋\lfloor W / 6 \rfloor + \lfloor W / 3 \rfloor + \lfloor W / 2 \rfloor).

  2. Knowl 2 — Res Path Shortcut Connections

    model/method

    In standard U-Net architectures, feature maps from early encoder layers (low-level representations) are directly concatenated with feature maps in corresponding late decoder layers (high-level representations), introducing a semantic discrepancy between merged features. The Res Path replaces direct skip connections with a sequence of non-linear convolutional operations enhanced with residual shortcuts to reconcile encoder and decoder feature representations prior to concatenation.

    Each Res Path passes encoder feature maps through a chain of 3×33 \times 3 convolutional layers, each accompanied by a 1×11 \times 1 convolution shortcut added element-wise to form residual blocks. Because the semantic discrepancy between encoder and decoder representations is largest at outer network levels and decreases toward inner levels, the number of residual convolutional blocks and the channel dimensions along the four Res Paths are structured as follows:

    • Res Path 1 (connecting Level 1 encoder to Level 1 decoder): 4 residual convolutional blocks, 32 filters each
    • Res Path 2 (connecting Level 2 encoder to Level 2 decoder): 3 residual convolutional blocks, 64 filters each
    • Res Path 3 (connecting Level 3 encoder to Level 3 decoder): 2 residual convolutional blocks, 128 filters each
    • Res Path 4 (connecting Level 4 encoder to Level 4 decoder): 1 residual convolutional block, 256 filters each
  3. Knowl 3 — MultiResUNet Architecture and Filter Scaling Formulation

    model/method

    MultiResUNet is an encoder-decoder network for biomedical image segmentation that replaces the standard two-convolution blocks of U-Net with MultiRes blocks and replaces standard direct skip connections with Res Paths.

    The network comprises a 5-level encoder-decoder hierarchy:

    1. Encoder: 4 stages, each containing a MultiRes block followed by a 2×22 \times 2 max pooling layer with stride 2.
    2. Bottleneck: A MultiRes block connecting the lowest encoder stage to the decoder.
    3. Decoder: 4 stages, each beginning with a 2×22 \times 2 transposed convolution that halves the number of feature channels, concatenating the upsampled features with the output of the corresponding Res Path from the encoder, and passing the combined representation through a MultiRes block.
    4. Output: A final 1×11 \times 1 convolution with a Sigmoid activation function to generate binary segmentation probabilities. All other convolutional layers are followed by Rectified Linear Unit (ReLU) activations and Batch Normalization.

    The channel parameter WW for each MultiRes block is derived from the baseline U-Net filter count UU at that level via: W=α×UW = \alpha \times U where U∈{32,64,128,256,512}U \in \{32, 64, 128, 256, 512\} across the five depth levels, and α=1.67\alpha = 1.67 is a scalar coefficient chosen to maintain a total parameter count slightly below that of standard U-Net.

    MultiResUNet extends to 3D volumetric segmentation (MultiResUNet 3D) by replacing all 2D operations (convolutions, max pooling, and transposed convolutions) with their 3D counterparts.

  4. Knowl 4 — Layer-by-Layer Specification of MultiResUNet

    data/table

    The filter counts and kernel sizes for all MultiRes blocks and Res Paths across the MultiResUNet architecture are detailed below:

    Block Layer (Filter Size) # Filters Path Layer (Filter Size) # Filters
    MultiRes Block 1 Conv2D(3,3)\text{Conv2D}(3,3) 8 Res Path 1 Conv2D(3,3)\text{Conv2D}(3,3) 32
    MultiRes Block 9 Conv2D(3,3)\text{Conv2D}(3,3) 17 Conv2D(1,1)\text{Conv2D}(1,1) 32
    Conv2D(3,3)\text{Conv2D}(3,3) 26 Conv2D(3,3)\text{Conv2D}(3,3) 32
    Conv2D(1,1)\text{Conv2D}(1,1) 51 Conv2D(1,1)\text{Conv2D}(1,1) 32
    MultiRes Block 2 Conv2D(3,3)\text{Conv2D}(3,3) 17 Conv2D(3,3)\text{Conv2D}(3,3) 32
    MultiRes Block 8 Conv2D(3,3)\text{Conv2D}(3,3) 35 Conv2D(1,1)\text{Conv2D}(1,1) 32
    Conv2D(3,3)\text{Conv2D}(3,3) 53 Conv2D(3,3)\text{Conv2D}(3,3) 32
    Conv2D(1,1)\text{Conv2D}(1,1) 105 Conv2D(1,1)\text{Conv2D}(1,1) 32
    MultiRes Block 3 Conv2D(3,3)\text{Conv2D}(3,3) 35 Res Path 2 Conv2D(3,3)\text{Conv2D}(3,3) 64
    MultiRes Block 7 Conv2D(3,3)\text{Conv2D}(3,3) 71 Conv2D(1,1)\text{Conv2D}(1,1) 64
    Conv2D(3,3)\text{Conv2D}(3,3) 106 Conv2D(3,3)\text{Conv2D}(3,3) 64
    Conv2D(1,1)\text{Conv2D}(1,1) 212 Conv2D(1,1)\text{Conv2D}(1,1) 64
    MultiRes Block 4 Conv2D(3,3)\text{Conv2D}(3,3) 71 Conv2D(3,3)\text{Conv2D}(3,3) 64
    MultiRes Block 6 Conv2D(3,3)\text{Conv2D}(3,3) 142 Conv2D(1,1)\text{Conv2D}(1,1) 64
    Conv2D(3,3)\text{Conv2D}(3,3) 213
    Conv2D(1,1)\text{Conv2D}(1,1) 426
    MultiRes Block 5 Conv2D(3,3)\text{Conv2D}(3,3) 142 Res Path 3 Conv2D(3,3)\text{Conv2D}(3,3) 128
    Conv2D(3,3)\text{Conv2D}(3,3) 284 Conv2D(1,1)\text{Conv2D}(1,1) 128
    Conv2D(3,3)\text{Conv2D}(3,3) 427 Conv2D(3,3)\text{Conv2D}(3,3) 128
    Conv2D(1,1)\text{Conv2D}(1,1) 853 Conv2D(1,1)\text{Conv2D}(1,1) 128
    Res Path 4 Conv2D(3,3)\text{Conv2D}(3,3) 256
    Conv2D(1,1)\text{Conv2D}(1,1) 256

    MultiRes Blocks 1–4 constitute the encoder pathway, MultiRes Block 5 forms the bottleneck, and MultiRes Blocks 6–9 constitute the decoder pathway. For each MultiRes block, the three 3×33 \times 3 convolutional layers operate sequentially with their outputs concatenated along channels, while the 1×11 \times 1 convolutional layer provides the residual projection shortcut. For each Res Path, pairs of 3×33 \times 3 convolutional layers and parallel 1×11 \times 1 residual shortcuts form consecutive residual units along the shortcut connection.

  5. Knowl 5 — Model Parameter Count Comparison for 2D and 3D Architectures

    data/table

    MultiResUNet is configured to maintain a slightly lower parameter count than standard U-Net architectures:

    2D Models 3D Models
    Model Parameters Model Parameters
    U-Net (baseline) 7,759,521 3D U-Net (baseline) 19,078,593
    MultiResUNet (proposed) 7,262,750 MultiResUNet 3D (proposed) 18,657,689

    In the 2D configuration, MultiResUNet utilizes 7,262,7507{,}262{,}750 parameters, representing a reduction of approximately 6.4%6.4\% compared to the baseline 2D U-Net (7,759,5217{,}759{,}521 parameters). In the 3D configuration, MultiResUNet 3D utilizes 18,657,68918{,}657{,}689 parameters compared to 19,078,59319{,}078{,}593 parameters for baseline 3D U-Net, demonstrating that performance enhancements achieved by MultiResUNet do not depend on expanding model capacity.

  6. Knowl 6 — Segmentation Performance Across Five Multimodal Medical Imaging Benchmarks

    data/table

    MultiResUNet and baseline U-Net were evaluated using 5-fold cross-validation on five distinct biomedical image segmentation datasets across different modalities. Performance is reported as the mean ±\pm standard deviation of the Jaccard Index (expressed as percentages):

    Modality MultiResUNet (%) U-Net (%) Relative Improvement (%)
    Dermoscopy (ISIC-2018) 80.2988±0.371780.2988 \pm 0.3717 76.4277±4.518376.4277 \pm 4.5183 5.06505.0650
    Endoscopy (CVC-ClinicDB) 82.0574±1.595382.0574 \pm 1.5953 74.4984±1.470474.4984 \pm 1.4704 10.146510.1465
    Fluorescence Microscopy (Murphy Lab) 91.6537±0.956391.6537 \pm 0.9563 89.3027±2.195089.3027 \pm 2.1950 2.63262.6326
    Electron Microscopy (ISBI-2012) 87.9477±0.774187.9477 \pm 0.7741 87.4092±0.707187.4092 \pm 0.7071 0.61610.6161
    MRI (BraTS17 3D) 78.1936±0.786878.1936 \pm 0.7868 77.1061±0.776877.1061 \pm 0.7768 1.41041.4104

    MultiResUNet consistently achieves higher Jaccard Index scores than baseline U-Net across all five modalities. The largest performance gains appear in non-uniform datasets characterized by vague boundaries and textural complexity, such as Endoscopy (+10.15% relative improvement) and Dermoscopy (+5.07% relative improvement).

  7. Knowl 7 — Faster Convergence and Reduced Variance Across Training Epochs

    empirical result

    When tracking validation Jaccard Index curves across 150 training epochs under 5-fold cross-validation, MultiResUNet exhibits two systematic training dynamics compared to baseline U-Net:

    1. Accelerated Convergence: MultiResUNet reaches peak validation segmentation accuracy in substantially fewer training epochs across all tested imaging modalities. This accelerated convergence arises from the synergy between the residual connections (in MultiRes blocks and Res Paths) and batch normalization.
    2. Reduced Inter-Fold Variance: The standard deviation band of validation Jaccard Index across the 5 cross-validation folds is consistently narrower for MultiResUNet than for U-Net throughout training, reflecting enhanced stability and robustness against variations in dataset partitioning.
  8. Knowl 8 — Robustness to Ambiguous Boundaries, Perturbations, Outliers, and Majority-Class Segmentation

    empirical result

    Qualitative comparison between MultiResUNet and baseline U-Net demonstrates enhanced robustness across challenging medical imaging conditions:

    • Ambiguous Boundaries: In colonoscopy images (CVC-ClinicDB) where polyps lack clear demarcation from surrounding mucosa, U-Net exhibits substantial under-segmentation or over-segmentation, whereas MultiResUNet accurately delineates subtle polyp boundaries.
    • Background and Foreground Perturbations: In dermoscopy images (ISIC-2018) with textured or irregular backgrounds and varying lesion interiors, U-Net frequently predicts disjoint, fractured foreground fragments, generates false positives on rough backgrounds, or fails entirely. MultiResUNet produces continuous, coherent lesion segmentations under strong perturbations.
    • Outlier Filtering: In fluorescence microscopy images (Murphy Lab), small bright non-nuclear debris particles mimic nuclei. Baseline U-Net misclassifies these artifacts as nuclei, whereas MultiResUNet filters them out.
    • Majority-Class Membrane Delineation: In transmission electron microscopy (ISBI-2012), where foreground regions dominate and are separated only by thin background membranes, U-Net frequently misses narrow separating lines and merges adjacent regions. MultiResUNet preserves these fine membrane boundaries.
  9. Knowl 9 — Datasets, Preprocessing, and Training Protocol for Evaluation

    experimental setup

    The experimental benchmark evaluates semantic segmentation across five datasets:

    1. Fluorescence Microscopy (Murphy Lab): 97 bright-field images containing 4,009 cell nuclei (U2OS and NIH3T3 cells), resized to 256×256256 \times 256.
    2. Electron Microscopy (ISBI-2012): 30 transmission electron microscopy images (512×512512 \times 512 resized to 256×256256 \times 256) of Drosophila ventral nerve cord.
    3. Dermoscopy (ISIC-2018): 2,594 skin lesion images compiled from ISIC-2017 and HAM10000, resized to 256×192256 \times 192.
    4. Endoscopy (CVC-ClinicDB): 612 colonoscopy video frames containing polyps, resized from 384×288384 \times 288 to 256×192256 \times 192.
    5. 3D MRI (BraTS17): 285 multimodal brain tumor scans (210 HGG and 75 LGG; 4 channels: T1, T1Gd, T2, FLAIR), resized from 240×240×155240 \times 240 \times 155 to 80×80×4880 \times 80 \times 48.

    Preprocessing consists solely of image resizing to accommodate GPU memory and normalizing pixel intensities to [0,1][0, 1] via division by 255 (no domain-specific preprocessing is applied). Models are trained for 150 epochs using the Adam optimizer with standard moment decay parameters to minimize the mean binary cross-entropy loss: J=1n∑i=1n∑px∈Xi−[ypxlog⁡(y^px)+(1−ypx)log⁡(1−y^px)]J = \frac{1}{n} \sum_{i=1}^{n} \sum_{p_x \in X_i} -\left[y_{p_x} \log(\hat{y}_{p_x}) + (1 - y_{p_x}) \log(1 - \hat{y}_{p_x})\right] where XiX_i is the ii-th image in a batch of size nn, ypx∈{0,1}y_{p_x} \in \{0, 1\} is the ground-truth binary label at pixel pxp_x, and y^px∈[0,1]\hat{y}_{p_x} \in [0, 1] is the predicted probability. A threshold of 0.5 is applied to generate binary masks. Evaluation is conducted using 5-fold cross-validation with the Jaccard Index: Jaccard Index=∣Y∩Y^∣∣Y∪Y^∣\text{Jaccard Index} = \frac{|Y \cap \hat{Y}|}{|Y \cup \hat{Y}|}

Coverage note — None was omitted; all contributed architectural components, layer specifications, parameter comparisons, empirical benchmarks across 5 modalities, training dynamics, qualitative robustness analyses, and experimental setup details are covered.

References

  1. 1.Johannes Schindelin, Curtis T Rueden, Mark C Hiner, and Kevin W Eliceiri. The imagej ecosystem: an open platform for biomedical image analysis. Molecular reproduction and development, 82(7-8):518–529, 2015.
  2. 2.Kevin McGuinness and Noel E O’connor. A comparative evaluation of interactive segmentation algorithms. Pattern Recognition, 43(2):434–444, 2010.
  3. 3.Noel CF Codella, David Gutman, M Emre Celebi, Brian Helba, Michael A Marchetti, Stephen W Dusza, Aadi Kalloo, Konstantinos Liopyris, Nabin Mishra, Harald Kittler, et al. Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on biomedical imaging (isbi), hosted by the international skin imaging collaboration (isic). In Biomedical Imaging (ISBI 2018), 2018 IEEE 15th International Symposium on, pages 168–172. IEEE, 2018.
  4. 4.Jinzhong Yang, Harini Veeraraghavan, Samuel G Armato III, Keyvan Farahani, Justin S Kirby, Jayashree Kalpathy-Kramer, Wouter van Elmpt, Andre Dekker, Xiao Han, Xue Feng, et al. Autosegmentation for thoracic radiation treatment planning: A grand challenge at aapm 2017. Medical physics, 2018.
  5. 5.Shivang Naik, Scott Doyle, Shannon Agner, Anant Madabhushi, Michael Feldman, and John Tomaszewski. Automated gland and nuclei segmentation for grading of prostate and breast cancer histopathology. In Biomedical Imaging: From Nano to Macro, 2008. ISBI 2008. 5th IEEE International Symposium on, pages 284–287. IEEE, 2008.
  6. 6.Rahimeh Rouhi, Mehdi Jafari, Shohreh Kasaei, and Peiman Keshavarzian. Benign and malignant breast tumors classification based on region growing and cnn segmentation. Expert Systems with Applications, 42(3):990–1002, 2015.
  7. 7.Dzung L Pham, Chenyang Xu, and Jerry L Prince. Current methods in medical image segmentation. Annual review of biomedical engineering, 2(1):315–337, 2000.
  8. 8.Pablo Mesejo, Andrea Valsecchi, Linda Marrakchi-Kacem, Stefano Cagnoni, and Sergio Damas. Biomedical image segmentation using geometric deformable models and metaheuristics. Computerized Medical Imaging and Graphics, 43:167–178, 2015.
  9. 9.Yuhui Zheng, Byeungwoo Jeon, Danhua Xu, QM Wu, and Hui Zhang. Image segmentation by generalized hierarchical fuzzy c-means algorithm. Journal of Intelligent & Fuzzy Systems, 28(2):961–973, 2015.
  10. 10.Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521(7553):436, 2015.
  11. 11.Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  12. 12.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012.
  13. 13.Pierre Sermanet, David Eigen, Xiang Zhang, Michael Mathieu, Rob Fergus, and Yann LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. arXiv preprint arXiv:1312.6229, 2013.
  14. 14.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  15. 15.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1–9, 2015.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  17. 17.Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In AAAI, volume 4, page 12, 2017.
  18. 18.Dan Ciresan, Alessandro Giusti, Luca M Gambardella, and Jurgen Schmidhuber. Deep neural networks segment neuronal membranes in electron microscopy images. In Advances in neural information processing systems, pages 2843–2851, 2012.
  19. 19.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3431–3440, 2015.
  20. 20.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. arXiv preprint arXiv:1511.00561, 2015.
  21. 21.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2018.
  22. 22.Geert Litjens, Thijs Kooi, Babak Ehteshami Bejnordi, Arnaud Arindra Adiyoso Setio, Francesco Ciompi, Mohsen Ghafoorian, Jeroen Awm Van Der Laak, Bram Van Ginneken, and Clara I Sanchez. A survey on deep learning in medical image analysis. Medical image analysis, 42:60–88, 2017.
  23. 23.Syed Muhammad Anwar, Muhammad Majid, Adnan Qayyum, Muhammad Awais, Majdi Alnowami, and Muhammad Khurram Khan. Medical image analysis using convolutional neural networks: a review. Journal of medical systems, 42(11):226, 2018.
  24. 24.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  25. 25.Patrick Ferdinand Christ, Mohamed Ezzeldin A Elshaer, Florian Ettlinger, Sunil Tatavarty, Marc Bickel, Patrick Bilic, Markus Rempfler, Marco Armbruster, Felix Hofmann, Melvin D’Anastasi, et al. Automatic liver and lesion segmentation in ct using cascaded fully convolutional neural networks and 3d conditional random fields. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 415–423. Springer, 2016.
  26. 26.Bill S Lin, Kevin Michael, Shivam Kalra, and Hamid R Tizhoosh. Skin lesion segmentation: U-nets versus clustering. In 2017 IEEE Symposium Series on Computational Intelligence (SSCI), pages 1–7. IEEE, 2017.
  27. 27.Korsuk Sirinukunwattana, Josien PW Pluim, Hao Chen, Xiaojuan Qi, Pheng-Ann Heng, Yun Bo Guo, Li Yang Wang, Bogdan J Matuszewski, Elia Bruni, Urko Sanchez, et al. Gland segmentation in colon histology images: The glas challenge contest. Medical image analysis, 35:489–502, 2017.
  28. 28.Ozgun C icek, Ahmed Abdulkadir, Soeren S Lienkamp, Thomas Brox, and Olaf Ronneberger. 3d u-net: learning dense volumetric segmentation from sparse annotation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 424–432. Springer, 2016.
  29. 29.Jameson Merkow, Alison Marsden, David Kriegman, and Zhuowen Tu. Dense volume-to-volume vascular boundary detection. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 371–379. Springer, 2016.
  30. 30.Arnaud Arindra Adiyoso Setio, Alberto Traverso, Thomas De Bel, Moira SN Berens, Cas van den Bogaard, Piergiorgio Cerello, Hao Chen, Qi Dou, Maria Evelina Fantacci, Bram Geurts, et al. Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: the luna16 challenge. Medical image analysis, 42:1–13, 2017.
  31. 31.Lequan Yu, Xin Yang, Hao Chen, Jing Qin, and Pheng-Ann Heng. Volumetric convnets with mixed residual connections for automated prostate segmentation from 3d mr images. In AAAI, pages 66–72, 2017.
  32. 32.M. D. Zeiler, D. Krishnan, G. W. Taylor, and R. Fergus. Deconvolutional networks. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 2528–2535, June 2010.
  33. 33.Michal Drozdzal, Eugene Vorontsov, Gabriel Chartrand, Samuel Kadoury, and Chris Pal. The importance of skip connections in biomedical image segmentation. In Deep Learning and Data Labeling for Medical Applications, pages 179–187. Springer, 2016.
  34. 34.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016.
  35. 35.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015.
  36. 36.Luıs Pedro Coelho, Aabid Shariff, and Robert F Murphy. Nuclear segmentation in microscope cell images: a hand-segmented dataset and comparison of algorithms. In Biomedical Imaging: From Nano to Macro, 2009. ISBI’09. IEEE International Symposium on, pages 518–521. IEEE, 2009.
  37. 37.Thomas Serre, Lior Wolf, Stanley Bileschi, Maximilian Riesenhuber, and Tomaso Poggio. Robust object recognition with cortex-like mechanisms. IEEE Transactions on Pattern Analysis & Machine Intelligence, (3):411–426, 2007.
  38. 38.Panqu Wang, Pengfei Chen, Ye Yuan, Ding Liu, Zehua Huang, Xiaodi Hou, and Garrison Cottrell. Understanding convolution for semantic segmentation. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 1451–1460. IEEE, 2018.
  39. 39.Xiao-Jiao Mao, Chunhua Shen, and Yu-Bin Yang. Image restoration using convolutional auto-encoders with symmetric skip connections. arXiv preprint arXiv:1606.08921, 2016.
  40. 40.Ignacio Arganda-Carreras, Srinivas C Turaga, Daniel R Berger, Dan Ciresan, Alessandro Giusti, Luca M Gambardella, Jurgen Schmidhuber, Dmitry Laptev, Sarvesh Dwivedi, Joachim M Buhmann, et al. Crowdsourcing the creation of image segmentation algorithms for connectomics. Frontiers in neuroanatomy, 9:142, 2015.
  41. 41.Albert Cardona, Stephan Saalfeld, Stephan Preibisch, Benjamin Schmid, Anchi Cheng, Jim Pulokas, Pavel Tomancak, and Volker Hartenstein. An integrated microand macroarchitectural analysis of the drosophila brain by computer-assisted serial section electron microscopy. PLoS biology, 8(10):e1000502, 2010.
  42. 42.Philipp Tschandl, Cliff Rosendahl, and Harald Kittler. The ham10000 dataset: A large collection of multi-source dermatoscopic images of common pigmented skin lesions. arXiv preprint arXiv:1803.10417, 2018.
  43. 43.Jorge Bernal, F Javier Sanchez, Gloria Fernandez-Esparrach, Debora Gil, Cristina Rodrıguez, and Fernando Vilarino. Wm-dova maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians. Computerized Medical Imaging and Graphics, 43:99–111, 2015.
  44. 44.Bjoern H Menze, Andras Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, et al. The multimodal brain tumor image segmentation benchmark (brats). IEEE transactions on medical imaging, 34(10):1993, 2015.
  45. 45.Spyridon Bakas, Hamed Akbari, Aristeidis Sotiras, Michel Bilello, Martin Rozycki, Justin S Kirby, John B Freymann, Keyvan Farahani, and Christos Davatzikos. Advancing the cancer genome atlas glioma mri collections with expert segmentation labels and radiomic features. Scientific data, 4:170117, 2017.
  46. 46.Guido Van Rossum et al. Python programming language. In USENIX Annual Technical Conference, volume 41, page 36, 2007.
  47. 47.Francois Chollet et al. Keras, 2015.
  48. 48.Martın Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. Tensorflow: a system for large-scale machine learning. In OSDI, volume 16, pages 265–283, 2016.
  49. 49.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  50. 50.John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 12(Jul):2121–2159, 2011.
  51. 51.Tijmen Tieleman and Geoffrey Hinton. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning, 4(2):26–31, 2012.
  52. 52.Sebastian Ruder. An overview of gradient descent optimization algorithms. arXiv preprint arXiv:1609.04747, 2016.
  53. 53.Ron Kohavi et al. A study of cross-validation and bootstrap for accuracy estimation and model selection. In Ijcai, volume 14, pages 1137–1145. Montreal, Canada, 1995.
  54. 54.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.

Citation

MLA
Ibtehaz, N., and M. S. Rahman. “MultiResUNet : Rethinking the U-Net Architecture for Multimodal Biomedical Image Segmentation”. Neural Networks, vol. 121, 2020, pp. 74–87, https://doi.org/10.1016/j.neunet.2019.08.025.
APA
Ibtehaz, N., & Rahman, M. S. (2020). MultiResUNet : Rethinking the U-Net architecture for multimodal biomedical image segmentation. Neural Networks, 121, 74–87. https://doi.org/10.1016/j.neunet.2019.08.025
Chicago
Ibtehaz, N., and M. S. Rahman. 2020. “MultiResUNet : Rethinking the U-Net Architecture for Multimodal Biomedical Image Segmentation”. Neural Networks 121: 74–87. https://doi.org/10.1016/j.neunet.2019.08.025.
Harvard
Ibtehaz, N. and Rahman, M.S. (2020) “MultiResUNet : Rethinking the U-Net architecture for multimodal biomedical image segmentation”, Neural Networks, 121, pp. 74–87. Available at: https://doi.org/10.1016/j.neunet.2019.08.025.
Vancouver
1. Ibtehaz N, Rahman MS (2020) MultiResUNet : Rethinking the U-Net architecture for multimodal biomedical image segmentation. Neural Networks 121:74–87

BibTeX

@article{Ibtehaz_2020, title={MultiResUNet : Rethinking the U-Net architecture for multimodal biomedical image segmentation}, volume={121}, ISSN={0893-6080}, url={http://dx.doi.org/10.1016/j.neunet.2019.08.025}, DOI={10.1016/j.neunet.2019.08.025}, journal={Neural Networks}, publisher={Elsevier BV}, author={Ibtehaz, Nabil and Rahman, M. Sohel}, year={2020}, month=Jan, pages={74–87} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF