Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?

Nima TajbakhshJae Y. ShinSuryakanth R. GuruduR. Todd HurstChristopher B. KendallMichael B. GotwayJianming Liang

article2016IEEE TMI3,002 citations

Establishes that fine-tuning pre-trained convolutional neural networks consistently matches or exceeds the performance of models trained from scratch across diverse medical imaging tasks, providing a practical layer-wise strategy to overcome scarce clinical training data.

Listen

Training deep convolutional neural networks from scratch for medical image tasks requires large labeled datasets and substantial expertise to achieve convergence, yet medical imaging routinely faces scarce expert annotations and class imbalance. This creates practical barriers to deploying CNNs in radiology, cardiology, and gastroenterology despite their strong potential for detection, classification, and segmentation.

The paper therefore set out to determine whether fine-tuning a CNN pre-trained on natural images could match or exceed the performance of a CNN trained entirely from scratch on medical data, and whether the depth of fine-tuning should be adjusted according to the amount of available labeled data.

Four distinct clinical tasks were examined across three imaging modalities: polyp detection in colonoscopy videos, frame informativeness classification in colonoscopy, pulmonary embolism detection in CT pulmonary angiography, and intima-media boundary segmentation in carotid ultrasound. In each case, the same AlexNet architecture was either initialized randomly and trained from scratch or initialized from an ImageNet-pretrained model and fine-tuned layer by layer while holding earlier layers fixed. Performance was measured with FROC, ROC, and localization-error analyses, supported by statistical comparisons at clinically relevant operating points, and the same experiments were repeated after successively reducing the training sets to 50 percent, 25 percent, and, in one case, 1 percent of the original size. Results were also benchmarked against previously published handcrafted detectors and segmenters.

Across all four tasks, fine-tuned networks matched or outperformed networks trained from scratch; the advantage grew markedly as training data were reduced. Deep fine-tuning (updating all convolutional layers) was required for polyp detection and boundary segmentation, whereas moderate or shallow tuning sufficed for frame classification and pulmonary embolism detection. In every application the best fine-tuned model also exceeded the corresponding handcrafted baseline. Convergence was faster with fine-tuning, and the performance gap relative to scratch training widened when fewer than half the original training samples were used.

These outcomes indicate that knowledge transfer from natural-image pretraining is both feasible and advantageous for medical imaging, removing the need to collect impractically large annotated medical datasets or to possess specialized deep-learning tuning skills. Clinically, the approach can shorten development cycles for computer-aided detection and measurement tools while maintaining or improving sensitivity at low false-positive rates.

For any new medical imaging task, practitioners should therefore begin with a pretrained network and apply incremental layer-wise fine-tuning until the desired performance is reached, using the amount of available labeled data to decide how many layers to update. Additional validation on magnetic resonance and histopathology images would strengthen generalizability, and systematic hyperparameter search or deeper architectures could yield further gains once the basic fine-tuning strategy is adopted. The reported results rest on consistent, statistically supported experiments across multiple tasks and data regimes, although they are bounded by the choice of AlexNet and the specific hyperparameter settings explored.

arXiv: 1706.00712
Cover for Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?

Abstract

Training a deep convolutional neural network (CNN) from scratch is difficult because it requires a large amount of labeled training data and a great deal of expertise to ensure proper convergence. A promising alternative is to fine-tune a CNN that has been pre-trained using, for instance, a large set of labeled natural images. However, the substantial differences between natural and medical images may advise against such knowledge transfer. In this paper, we seek to answer the following central question in the context of medical image analysis: \emph{Can the use of pre-trained deep CNNs with sufficient fine-tuning eliminate the need for training a deep CNN from scratch?} To address this question, we considered 4 distinct medical imaging applications in 3 specialties (radiology, cardiology, and gastroenterology) involving classification, detection, and segmentation from 3 different imaging modalities, and investigated how the performance of deep CNNs trained from scratch compared with the pre-trained CNNs fine-tuned in a layer-wise manner. Our experiments consistently demonstrated that (1) the use of a pre-trained CNN with adequate fine-tuning outperformed or, in the worst case, performed as well as a CNN trained from scratch; (2) fine-tuned CNNs were more robust to the size of training sets than CNNs trained from scratch; (3) neither shallow tuning nor deep tuning was the optimal choice for a particular application; and (4) our layer-wise fine-tuning scheme could offer a practical way to reach the best performance for the application at hand based on the amount of available data.

Table of Contents

  • I Introduction
  • II Related Works
  • III Contributions
  • IV Convolutional Neural Networks (CNNs)
  • V Fine-tuning
  • VI Applications and Results
  • VI-A Polyp detection
  • VI-B Pulmonary embolism detection
  • VI-C Colonoscopy frame classification
  • VI-D Intima-media boundary segmentation
  • VII Discussion
  • VIII Conclusion
  • References

Knowls

  1. Knowl 1 — Layer-Wise Fine-Tuning Scheme for Deep Convolutional Neural Networks

    model/method

    In medical image analysis, a pre-trained convolutional neural network (CNN) whose weights were learned on a large-scale natural image dataset (such as ImageNet) is transferred to a target medical imaging task. The network architecture consists of LL layers where the first L3L-3 layers are convolutional and pooling stages, and the final layers are fully connected layers terminated by a softmax classification layer. To adapt the pre-trained CNN to a target application with CC classes, the final fully connected layer (e.g., layer LL, or fc8) is replaced with a new randomly initialized fully connected layer containing CC output units.

    Fine-tuning is conducted in a layer-wise progression by selectively unfreezing layers from the output layer backward toward the input layer. Let αl\alpha_l denote the learning rate assigned to layer l{1,,L}l \in \{1, \dots, L\}. Layer-wise tuning levels are configured as follows:

    • Tuning only the last layer (only fc8): Set αL>0\alpha_L > 0 and αl=0\alpha_l = 0 for all l<Ll < L, equivalent to training a linear classifier on fixed features extracted from layer L1L-1.
    • Shallow tuning (fc7-fc8 or fc6-fc8): Set αl>0\alpha_l > 0 for the top fully connected layers (lL1l \ge L-1 or lL2l \ge L-2) while setting αl=0\alpha_l = 0 for all earlier convolutional layers, equivalent to training a shallow multi-layer neural network on fixed convolutional representations.
    • Moderate to deep tuning (conv5-fc8 down to conv1-fc8): Incrementally set αl>0\alpha_l > 0 for deeper convolutional layers, culminating in deep tuning where all layers from conv1 through fc8 are updated simultaneously.

    The network parameters WW are optimized by minimizing the cross-entropy loss:

    L=1Xi=1Xln(p(yiXi))\mathcal{L} = -\frac{1}{|X|} \sum_{i=1}^{|X|} \ln\left(p(y^i \mid X^i)\right)

    where X|X| is the total number of training samples, XiX^i denotes the ii-th training image with true class label yiy^i, and p(yiXi)p(y^i \mid X^i) is the predicted probability of the correct class. Minimization uses mini-batch stochastic gradient descent with momentum μ\mu, mini-batch size NN, and an epoch-based learning rate decay scheduling rate γ(0,1]\gamma \in (0, 1]:

    γt=γtN/X\gamma^t = \gamma^{\lfloor tN / |X| \rfloor}

    Vlt+1=μVltγtαlL^WlV_l^{t+1} = \mu V_l^t - \gamma^t \alpha_l \frac{\partial \hat{\mathcal{L}}}{\partial W_l}

    Wlt+1=Wlt+Vlt+1W_l^{t+1} = W_l^t + V_l^{t+1}

    where tt denotes the iteration step, WltW_l^t is the weight matrix of layer ll, VltV_l^t is the velocity vector for layer ll, and L^\hat{\mathcal{L}} is the loss computed over the current mini-batch of size NN.

  2. Knowl 2 — AlexNet Architecture and Layer-Wise Fine-Tuning Parameter Settings

    data/table

    The experimental framework utilizes the AlexNet CNN architecture consisting of 5 convolutional layers followed by 3 fully connected layers, accepting an input image patch of size 3×227×2273 \times 227 \times 227. The network specifications are:

    • conv1: 9696 filters of size 11×1111 \times 11, stride 44, padding 00, output 96×55×5596 \times 55 \times 55, followed by pool1 (max pooling, 3×33 \times 3, stride 22, output 96×27×2796 \times 27 \times 27)
    • conv2: 256256 filters of size 5×55 \times 5, stride 11, padding 22, output 256×27×27256 \times 27 \times 27, followed by pool2 (max pooling, 3×33 \times 3, stride 22, output 256×13×13256 \times 13 \times 13)
    • conv3: 384384 filters of size 3×33 \times 3, stride 11, padding 11, output 384×13×13384 \times 13 \times 13
    • conv4: 384384 filters of size 3×33 \times 3, stride 11, padding 11, output 384×13×13384 \times 13 \times 13
    • conv5: 256256 filters of size 3×33 \times 3, stride 11, padding 11, output 256×13×13256 \times 13 \times 13, followed by pool5 (max pooling, 3×33 \times 3, stride 22, output 256×6×6256 \times 6 \times 6)
    • fc6: fully connected layer mapping 256×6×6256 \times 6 \times 6 to 4096×14096 \times 1
    • fc7: fully connected layer mapping 4096×14096 \times 1 to 4096×14096 \times 1
    • fc8: fully connected classification layer mapping 4096×14096 \times 1 to C×1C \times 1, where C=3C=3 for carotid interface segmentation and C=2C=2 for polyp detection, pulmonary embolism detection, and colonoscopy frame classification.

    The hyperparameter configurations for full training from scratch and layer-wise fine-tuning are:

    CNN Configuration μ\mu αconv1\alpha_{\text{conv1}} αconv2\alpha_{\text{conv2}} αconv3\alpha_{\text{conv3}} αconv4\alpha_{\text{conv4}} αconv5\alpha_{\text{conv5}} αfc6\alpha_{\text{fc6}} αfc7\alpha_{\text{fc7}} αfc8\alpha_{\text{fc8}} γ\gamma
    Fine-tuned: conv1-fc8 0.9 0.001 0.001 0.001 0.001 0.001 0.001 0.001 0.01 0.95
    Fine-tuned: conv2-fc8 0.9 0 0.001 0.001 0.001 0.001 0.001 0.001 0.01 0.95
    Fine-tuned: conv3-fc8 0.9 0 0 0.001 0.001 0.001 0.001 0.001 0.01 0.95
    Fine-tuned: conv4-fc8 0.9 0 0 0 0.001 0.001 0.001 0.001 0.01 0.95
    Fine-tuned: conv5-fc8 0.9 0 0 0 0 0.001 0.001 0.001 0.01 0.95
    Fine-tuned: fc6-fc8 0.9 0 0 0 0 0 0.001 0.001 0.01 0.95
    Fine-tuned: fc7-fc8 0.9 0 0 0 0 0 0 0.001 0.01 0.95
    Fine-tuned: only fc8 0.9 0 0 0 0 0 0 0 0.01 0.95
    AlexNet scratch 0.9 0.001 0.001 0.001 0.001 0.001 0.001 0.001 0.001 0.95

    For all models, the learning rate for bias terms is set to twice the learning rate of the corresponding layer weights (2αl2\alpha_l).

  3. Knowl 3 — Dependence of Optimal Fine-Tuning Depth on Domain Distance

    empirical result

    Systematic evaluation across four medical imaging applications demonstrates that knowledge transfer from natural images (ImageNet) to medical imaging tasks is viable and that pre-trained CNNs with adequate fine-tuning consistently match or outperform CNNs trained from scratch. However, the optimal fine-tuning depth is governed by the distance between the source dataset and the target task:

    1. High domain discrepancy (e.g., Polyp Detection in endoscopy, CIMT Segmentation in ultrasound): Shallow tuning (only fc8 or fc7-fc8) yields inferior performance, while deeply fine-tuned networks (conv1-fc8) achieve the highest performance, significantly surpassing CNNs trained from scratch (p<0.05p < 0.05). In these applications, early convolutional filters pre-trained on natural images must be substantially adapted to capture modality-specific low-level textural and boundary cues.
    2. Moderate domain discrepancy (e.g., Colonoscopy Frame Classification): Moderate fine-tuning (conv4-fc8 and conv5-fc8) achieves the highest performance, outperforming deep fine-tuning (conv1-fc8) at 10%10\% and 15%15\% false positive rates. Because the input colonoscopy video frames share high-resolution visual statistics and low-level edge/color primitives with ImageNet images, modifying early convolutional layers degrades generalization.
    3. Volumetric lesion detection (e.g., Pulmonary Embolism Detection in CT): Performance saturates quickly after tuning the fully connected layers, with deep fine-tuning performing on par with a CNN trained from scratch.

    Consequently, neither shallow tuning nor full deep tuning is universally optimal; the required depth of fine-tuning must be chosen based on the domain distance and data availability of the target medical application.

  4. Knowl 4 — Robustness of Fine-Tuned CNNs Under Reduced Training Set Sizes

    empirical result

    Comparing deep CNNs trained from scratch with deeply fine-tuned CNNs (conv1-fc8) across reduced subsets of training data reveals distinct stability properties:

    • Polyp Detection: When training data is reduced to 50%50\% and 25%25\% of unique lesions (at the video level), CNNs trained from scratch suffer severe performance degradation. In contrast, deeply fine-tuned CNNs maintain high sensitivity with minimal loss in free-response ROC (FROC) performance, creating a large, statistically significant performance gap in favor of fine-tuning.
    • Pulmonary Embolism Detection: Reducing training data to 50%50\% and 25%25\% of PE candidates at the patient level causes a significant drop in sensitivity across all false positive rates for CNNs trained from scratch, whereas deeply fine-tuned models maintain superior FROC curves across all operating points.
    • Colonoscopy Frame Classification: Reducing the dataset by factors of 1/101/10, 1/201/20, and 1/1001/100 (10%10\%, 5%5\%, and 1%1\% training data) shows that both scratch and fine-tuned models remain stable at 10%10\%, but scratch models degrade precipitously at 5%5\% and 1%1\%, while deeply fine-tuned CNNs retain strong classification performance (AUC >0.85> 0.85 even at 1%1\% data).

    These results establish that fine-tuned CNNs are substantially more robust to training set size limitations than networks trained entirely from scratch.

  5. Knowl 5 — Convergence Acceleration and Initialization Invariance of Fine-Tuned CNNs

    empirical result

    Tracking the area under the ROC curve (AUC) on validation datasets as a function of the number of mini-batch iterations demonstrates that deeply fine-tuned CNNs (conv1-fc8) converge significantly faster than CNNs trained from scratch.

    Across all four medical applications (polyp detection, PE detection, frame classification, and intima-media boundary segmentation):

    • The fine-tuned CNN reaches peak validation AUC within 10110^1 to 10210^2 mini-batches.
    • Full training from scratch requires between 10310^3 and 10410^4 mini-batches to achieve convergence.
    • Evaluating three distinct weight initialization methods for training from scratch—standard Gaussian random initialization, Xavier initialization, and MSRA (He) initialization—shows that while Xavier and MSRA alter the convergence trajectory, none enables scratch training to reach the rapid convergence rate of pre-trained fine-tuning, nor do they yield statistically significant accuracy improvements over Gaussian initialization after full convergence (with minor variation only in PE detection).
  6. Knowl 6 — Carotid Intima-Media Thickness (CIMT) Interface Segmentation Pipeline

    model/method

    Carotid intima-media thickness (CIMT) measurement in ultrasound B-mode images is formulated as a 3-class pixel classification problem within an expert-defined region of interest (ROI):

    1. Patch Extraction and Classification: For each pixel in the ROI, an image patch is extracted. To match the 3-channel input of AlexNet, the single-channel grayscale patch is replicated across three channels. A 3-way CNN classifies each pixel into one of three classes: lumen-intima (LI) interface, media-adventitia (MA) interface, or non-interface background.
    2. Confidence Map Generation: The trained CNN outputs two continuous probability maps corresponding to the likelihood of a pixel belonging to the LI interface and the MA interface across the ROI.
    3. Boundary Thinning: In each column of the ROI confidence maps, the pixel index corresponding to the maximum probability response is selected for each interface, producing a discrete 1-pixel-thick step-like boundary.
    4. Contour Regularization via Dual Active Contours (Snakes): Two open active contour models (snakes) are initialized along the step-like discrete boundaries of the LI and MA interfaces. The snakes deform under internal elasticity/rigidity and external image gradient forces until convergence, producing smooth continuous boundary curves.
    5. Thickness Computation: CIMT is calculated as the average vertical distance between the two converged open snake contours across the ROI.
  7. Knowl 7 — Segmentation Error of Layer-Wise Fine-Tuned CNNs on Carotid Intima-Media Interfaces

    data/table

    Interface segmentation error was measured on 132 test ROIs from 92 CIMT ultrasound videos as the average absolute distance (in micrometers, μm\mu\text{m}) between system-segmented boundaries and consensus annotations of two medical experts. The means (μ\mu) and standard deviations (σ\sigma) of the localization errors for both the lumen-intima and media-adventitia interfaces across fine-tuning depths, training from scratch, and a handcrafted baseline are summarized below:

    Method Lumen-Intima Interface Media-Adventitia Interface
    Mean Error (μm\mu\text{m}) Std Dev (μm\mu\text{m}) Mean Error (μm\mu\text{m}) Std Dev (μm\mu\text{m})
    FT: only fc8 100.97 132.64 103.26 58.48
    FT: fc7-fc8 34.67 14.77 43.93 21.79
    FT: fc6-fc8 27.87 11.83 34.63 14.49
    FT: conv5-fc8 25.32 11.02 31.94 13.13
    FT: conv4-fc8 25.01 9.94 30.55 11.97
    FT: conv3-fc8 24.37 10.79 31.80 11.98
    FT: conv2-fc8 24.71 10.25 31.74 12.37
    FT: conv1-fc8 24.74 11.42 31.21 12.56
    AlexNet scratch 28.11 13.79 33.45 16.17
    Handcrafted 98.17 16.57 106.62 21.00

    Fine-tuning only fc8 performs poorly, matching the handcrafted method (p>0.5p > 0.5). Unfreezing fc7 and fc6 produces large statistically significant reductions in error (p<0.001p < 0.001). Tuning through conv5 achieves performance parity with deep tuning (conv1-fc8). Deeply fine-tuned CNNs achieve significantly lower localization error than full training from scratch for both the lumen-intima interface (p<0.0001p < 0.0001) and media-adventitia interface (p<0.05p < 0.05).

  8. Knowl 8 — 2-Channel Vessel-Aligned Representation for Pulmonary Embolism Detection

    model/method

    To minimize appearance variability of pulmonary emboli (PE) within CT pulmonary angiography (CTPA) datasets for deep CNN training, a candidate detection and multi-planar reformatting framework is used:

    1. Candidate Generation: A tobogganing algorithm identifies dark filling-defect embolus candidates surrounded by brighter contrast-enhanced blood within the pulmonary arteries.
    2. Vessel-Aligned 2-Channel Extraction: For each candidate voxel, the local vessel centerline and orientation are computed. Two orthogonal 2D planes aligned with the vessel geometry are extracted: a longitudinal view along the vessel trajectory and a cross-sectional view perpendicular to the vessel axis.
    3. Channel Reformatting for CNN Input: The resulting 2-channel patch (longitudinal channel + cross-sectional channel) is converted into a 3-channel input compatible with AlexNet by duplicating the second channel (producing channels: [longitudinal,cross-sectional,cross-sectional][\text{longitudinal}, \text{cross-sectional}, \text{cross-sectional}]).
    4. Data Augmentation: Patches are sampled at 3 physical scales (10 mm10\text{ mm}, 15 mm15\text{ mm}, 20 mm20\text{ mm}), translated along the vessel direction 3 times (up to 20%20\% of patch physical size), and rotated around the vessel axis across 5 angles, yielding 81,000 stratified training patches from 121 CTPA scans.
    5. Inference Aggregation: At test time, patch augmentations are generated for each candidate, and candidate-level PE probabilities are calculated as the mean predicted probability over all augmented patches.
  9. Knowl 9 — Polyp Candidate Generation, Data Augmentation, and FROC Performance

    empirical result

    In colonoscopy video polyp detection evaluated on a dataset of 40 video sequences (training: 3,800 polyp frames, 15,100 non-polyp frames; testing: 5,700 polyp frames, 13,200 non-polyp frames):

    1. Candidate Extraction and Augmentation: Initial candidates are identified via boundary and context detection. Patches are extracted around each candidate bounding box at 3 scaling factors (1.0×,1.2×,1.5×1.0\times, 1.2\times, 1.5\times), jittered with horizontal and vertical translations of 10%10\% of the resized box dimensions, and flipped/mirrored in 8 orientations to generate a stratified training set of 100,000 patches.
    2. Layer-Wise Fine-Tuning Performance: Free-response ROC (FROC) analysis demonstrates strict monotonic performance improvements as more layers are fine-tuned:
      • FT:only fc8 exhibits the lowest sensitivity.
      • FT:fc7-fc8 significantly outperforms FT:only fc8 (p<0.05p < 0.05 at 10210^{-2} and 10310^{-3} false positives per frame).
      • FT:conv5-fc8, FT:conv4-fc8, and FT:conv3-fc8 significantly outperform FT:fc7-fc8 (p<0.05p < 0.05).
      • Deeply fine-tuned CNNs (FT:conv1-fc8 and FT:conv2-fc8) achieve the highest sensitivity, outperforming moderate fine-tuning and matching or exceeding AlexNet trained from scratch at low false positive rates (10310^{-3} to 10110^{-1} false positives/frame).
      • All CNN configurations significantly outperform the handcrafted geometric/context baseline (p<0.05p < 0.05).
  10. Knowl 10 — Colonoscopy Frame Informativeness Classification Setup and ROC Analysis

    empirical result

    Image quality assessment in colonoscopy is evaluated as a binary classification problem (informative vs. non-informative frames) across 6 complete video examinations:

    1. Dataset Construction: Frames are sampled at 1 frame per 5 seconds to reduce temporal redundancy, resulting in a balanced dataset of 4,000 annotated images (500×350500 \times 350 pixels), split at the video level into 2,000 training and 2,000 testing images. Random crops of size 227×227227 \times 227 generate a stratified training set of 40,000 sub-images. During inference, frame-level probability is the average over randomly cropped sub-images.
    2. ROC Comparison: All CNN-based approaches significantly outperform the handcrafted reconstruction-error baseline at 10%10\%, 15%15\%, and 20%20\% false positive rates (p<0.05p < 0.05).
    3. Optimal Tuning Depth: Moderate fine-tuning halfway into the network (FT:conv4-fc8 and FT:conv5-fc8) yields the highest ROC sensitivity, statistically outperforming both shallow fine-tuning (FT:only fc8, FT:fc7-fc8) and deep fine-tuning (FT:conv1-fc8) at 10%10\% and 15%15\% false positive rates. Full training from scratch is significantly outperformed by FT:conv5-fc8 (p<0.05p < 0.05 at 10%10\% and 15%15\% false positive rates).

Coverage note — No substantial contributed material was omitted. The knowls capture the core layer-wise fine-tuning method, network hyperparameters, overarching empirical conclusions on domain distance, robustness to small training sets, convergence properties, task-specific pipelines, and quantitative results for all four medical imaging applications.

References

  1. 1.K. Fukushima, ‘‘Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position,’’ Biological cybernetics, vol. 36, no. 4, pp. 193–202, 1980.
  2. 2.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, ‘‘Gradient-based learning applied to document recognition,’’ Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
  3. 3.Y. LeCun, Y. Bengio, and G. Hinton, ‘‘Deep learning,’’ Nature, vol. 521, no. 7553, pp. 436–444, 2015.
  4. 4.‘‘Online: available at http://www.technologyreview.com/featuredstory/513696/deep-learning/,’’.
  5. 5.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, ‘‘Going deeper with convolutions,’’ arXiv preprint arXiv:1409.4842, 2014.
  6. 6.K. Simonyan and A. Zisserman, ‘‘Very deep convolutional networks for large-scale image recognition,’’ arXiv preprint arXiv:1409.1556, 2014.
  7. 7.M. D. Zeiler and R. Fergus, ‘‘Visualizing and understanding convolutional networks,’’ in Computer Vision–ECCV 2014. Springer, 2014, pp. 818–833.
  8. 8.D. Eigen, J. Rolfe, R. Fergus, and Y. LeCun, ‘‘Understanding deep architectures using a recursive convolutional network,’’ arXiv preprint arXiv:1312.1847, 2013.
  9. 9.D. Erhan, P.-A. Manzagol, Y. Bengio, S. Bengio, and P. Vincent, ‘‘The difficulty of training deep architectures and the effect of unsupervised pre-training,’’ in International Conference on artificial intelligence and statistics, 2009, pp. 153–160.
  10. 10.A. S. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson, ‘‘CNN features off-the-shelf: an astounding baseline for recognition,’’ in Computer Vision and Pattern Recognition Workshops (CVPRW), 2014 IEEE Conference on. IEEE, 2014, pp. 512–519.
  11. 11.H. Azizpour, A. S. Razavian, J. Sullivan, A. Maki, and S. Carlsson, ‘‘From generic to specific deep representations for visual recognition,’’ arXiv preprint arXiv:1406.5774, 2014.
  12. 12.O. Penatti, K. Nogueira, and J. Santos, ‘‘Do deep features generalize from everyday objects to remote sensing and aerial scenes domains?’’ in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2015, pp. 44–51.
  13. 13.W. Zhang, K. Doi, M. L. Giger, Y. Wu, R. M. Nishikawa, and R. A. Schmidt, ‘‘Computerized detection of clustered microcalcifications in digital mammograms using a shift-invariant artificial neural network,’’ Medical Physics, vol. 21, no. 4, pp. 517–524, 1994.
  14. 14.H.-P. Chan, S.-C. B. Lo, B. Sahiner, K. L. Lam, and M. A. Helvie, ‘‘Computer-aided detection of mammographic microcalcifications: Pattern recognition with an artificial neural network,’’ Medical Physics, vol. 22, no. 10, pp. 1555–1567, 1995.
  15. 15.S.-C. B. Lo, S.-L. Lou, J.-S. Lin, M. T. Freedman, M. V. Chien, S. K. Mun et al., ‘‘Artificial convolution neural network techniques and applications for lung nodule detection,’’ Medical Imaging, IEEE Transactions on, vol. 14, no. 4, pp. 711–718, 1995.
  16. 16.N. Tajbakhsh, S. R. Gurudu, and J. Liang, ‘‘A comprehensive computer-aided polyp detection system for colonoscopy videos,’’ in Information Processing in Medical Imaging. Springer, 2015, pp. 327–338.
  17. 17.——, ‘‘Automatic polyp detection in colonoscopy videos using an ensemble of convolutional neural networks,’’ in Biomedical Imaging (ISBI), 2015 IEEE 12th International Symposium on. IEEE, 2015, pp. 79–83.
  18. 18.N. Tajbakhsh and J. Liang, ‘‘Computer-aided pulmonary embolism detection using a novel vessel-aligned multi-planar image representation and convolutional neural networks,’’ in Medical Image Computing and Computer-Assisted Intervention MICCAI 2015, 2015.
  19. 19.D. C. Cireřan, A. Giusti, L. M. Gambardella, and J. Schmidhuber, ‘‘Mitosis detection in breast cancer histology images with deep neural networks,’’ in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2013. Springer, 2013, pp. 411–418.
  20. 20.H. Roth, L. Lu, A. Seff, K. Cherry, J. Hoffman, S. Wang, J. Liu, E. Turkbey, and R. Summers, ‘‘A new 2.5d representation for lymph node detection using random sets of deep convolutional neural network observations,’’ in Medical Image Computing and Computer-Assisted Intervention MICCAI 2014, ser. Lecture Notes in Computer Science, P. Golland, N. Hata, C. Barillot, J. Hornegger, and R. Howe, Eds. Springer International Publishing, 2014, vol. 8673, pp. 520–527.
  21. 21.Y. Zheng, D. Liu, B. Georgescu, H. Nguyen, and D. Comaniciu, ‘‘3d deep learning for efficient and robust landmark detection in volumetric data,’’ in Medical Image Computing and Computer-Assisted Intervention MICCAI 2015, 2015.
  22. 22.J. Y. Shin, N. Tajbakhsh, R. T. Hurst, C. B. Kendall, and J. Liang, ‘‘Automating carotid intima-media thickness video interpretation with convolutional neural networks,’’ to appear in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  23. 23.H. R. Roth, A. Farag, L. Lu, E. B. Turkbey, and R. M. Summers, ‘‘Deep convolutional networks for pancreas segmentation in ct imaging,’’ in SPIE Medical Imaging. International Society for Optics and Photonics, 2015, pp. 94 131G–94 131G.
  24. 24.M. Havaei, A. Davy, D. Warde-Farley, A. Biard, A. Courville, Y. Bengio, C. Pal, P.-M. Jodoin, and H. Larochelle, ‘‘Brain tumor segmentation with deep neural networks,’’ arXiv preprint arXiv:1505.03540, 2015.
  25. 25.W. Zhang, R. Li, H. Deng, L. Wang, W. Lin, S. Ji, and D. Shen, ‘‘Deep convolutional neural networks for multi-modality isointense infant brain image segmentation,’’ NeuroImage, vol. 108, pp. 214–224, 2015.
  26. 26.D. Ciresan, A. Giusti, L. M. Gambardella, and J. Schmidhuber, ‘‘Deep neural networks segment neuronal membranes in electron microscopy images,’’ in Advances in Neural Information Processing Systems 25, F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds. Curran Associates, Inc., 2012, pp. 2843–2851.
  27. 27.A. Prasoon, K. Petersen, C. Igel, F. Lauze, E. Dam, and M. Nielsen, ‘‘Deep feature learning for knee cartilage segmentation using a triplanar convolutional neural network,’’ in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2013. Springer, 2013, pp. 246–253.
  28. 28.Y. Bar, I. Diamant, L. Wolf, and H. Greenspan, ‘‘Deep learning with non-medical training used for chest pathology identification,’’ in SPIE Medical Imaging. International Society for Optics and Photonics, 2015, pp. 94 140V–94 140V.
  29. 29.B. van Ginneken, A. A. Setio, C. Jacobs, and F. Ciompi, ‘‘Off-the-shelf convolutional neural network features for pulmonary nodule detection in computed tomography scans,’’ in Biomedical Imaging (ISBI), 2015 IEEE 12th International Symposium on, April 2015, pp. 286–289.
  30. 30.J. Arevalo, F. Gonzalez, R. Ramos-Pollan, J. Oliveira, and M. Guevara Lopez, ‘‘Convolutional neural networks for mammography mass lesion classification,’’ in Engineering in Medicine and Biology Society (EMBC), 2015 37th Annual International Conference of the IEEE, Aug 2015, pp. 797–800.
  31. 31.T. Schlegl, J. Ofner, and G. Langs, ‘‘Unsupervised pre-training across image domains improves lung tissue classification,’’ in Medical Computer Vision: Algorithms for Big Data. Springer, 2014, pp. 82–93.
  32. 32.H. Chen, D. Ni, J. Qin, S. Li, X. Yang, T. Wang, and P. A. Heng, ‘‘Standard plane localization in fetal ultrasound via domain transferred deep neural networks,’’ Biomedical and Health Informatics, IEEE Journal of, vol. 19, no. 5, pp. 1627–1636, Sept 2015.
  33. 33.G. Carneiro, J. Nascimento, and A. Bradley, ‘‘Unregistered multiview mammogram analysis with pre-trained deep learning models,’’ in Medical Image Computing and Computer-Assisted Intervention MICCAI 2015, ser. Lecture Notes in Computer Science, N. Navab, J. Hornegger, W. M. Wells, and A. F. Frangi, Eds. Springer International Publishing, 2015, vol. 9351, pp. 652–660. [Online]. Available: http://dx.doi.org/10.1007/978-3-319-24574-4 78
  34. 34.H.-C. Shin, L. Lu, L. Kim, A. Seff, J. Yao, and R. M. Summers, ‘‘Interleaved text/image deep mining on a very large-scale radiology database,’’ in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 1090–1099.
  35. 35.M. Gao, U. Bagci, L. Lu, A. Wu, M. Buty, H.-C. Shin, H. Roth, G. Z. Papadakis, A. Depeursinge, R. M. Summers et al., ‘‘Holistic classification of ct attenuation patterns for interstitial lung diseases via deep convolutional neural networks,’’ in the 1st Workshop on Deep Learning in Medical Image Analysis, International Conference on Medical Image Computing and Computer Assisted Intervention, at MICCAI-DLMIA’15, 2015.
  36. 36.J. Margeta, A. Criminisi, R. Cabrera Lozoya, D. C. Lee, and N. Ayache, ‘‘Fine-tuned convolutional neural nets for cardiac mri acquisition plane recognition,’’ Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization, pp. 1–11, 2015.
  37. 37.D. H. Hubel and T. N. Wiesel, ‘‘Receptive fields of single neurones in the cat’s striate cortex,’’ The Journal of physiology, vol. 148, no. 3, pp. 574–591, 1959.
  38. 38.D. C. Edwards, M. A. Kupinski, C. E. Metz, and R. M. Nishikawa, ‘‘Maximum likelihood fitting of FROC curves under an initial-detection-and-candidate-analysis model,’’ Medical physics, vol. 29, no. 12, pp. 2861–2870, 2002.
  39. 39.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, ‘‘Caffe: Convolutional architecture for fast feature embedding,’’ arXiv preprint arXiv:1408.5093, 2014.
  40. 40.X. Glorot and Y. Bengio, ‘‘Understanding the difficulty of training deep feedforward neural networks,’’ in International conference on artificial intelligence and statistics, 2010, pp. 249–256.
  41. 41.K. He, X. Zhang, S. Ren, and J. Sun, ‘‘Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,’’ arXiv preprint arXiv:1502.01852, 2015.
  42. 42.N. Tajbakhsh, S. Gurudu, and J. Liang, ‘‘Automated polyp detection in colonoscopy videos using shape and context information,’’ Medical Imaging, IEEE Transactions on, vol. PP, no. 99, pp. 1–1, 2015.
  43. 43.A. Pabby, R. E. Schoen, J. L. Weissfeld, R. Burt, J. W. Kikendall, P. Lance, M. Shike, E. Lanza, and A. Schatzkin, ‘‘Analysis of colorectal cancer occurrence during surveillance colonoscopy in the dietary polyp prevention trial,’’ Gastrointest Endosc, vol. 61, no. 3, pp. 385–91, 2005.
  44. 44.J. van Rijn, J. Reitsma, J. Stoker, P. Bossuyt, S. van Deventer, and E. Dekker, ‘‘Polyp miss rate determined by tandem colonoscopy: a systematic review,’’ American Journal of Gastroenterology, vol. 101, no. 2, pp. 343–350, 2006.
  45. 45.D. H. Kim, P. J. Pickhardt, A. J. Taylor, W. K. Leung, T. C. Winter, J. L. Hinshaw, D. V. Gopal, M. Reichelderfer, R. H. Hsu, and P. R. Pfau, ‘‘Ct colonography versus colonoscopy for the detection of advanced neoplasia,’’ N Engl J Med, vol. 357, no. 14, pp. 1403–12, 2007.
  46. 46.D. Heresbach, T. Barrioz, M. Lapalus, D. Coumaros, P. Bauret, P. Potier, D. Sautereau, C. Boustiere, J. Grimaud, C. Barth ` el´ em´ y et al., ‘‘Miss rate for colorectal neoplastic polyps: a prospective multicenter study of back-to-back video colonoscopies.’’ Endoscopy, vol. 40, no. 4, pp. 284–290, 2008.
  47. 47.A. Leufkens, M. van Oijen, F. Vleggaar, and P. Siersema, ‘‘Factors influencing the miss rate of polyps in a back-to-back colonoscopy study,’’ Endoscopy, vol. 44, no. 05, pp. 470–475, 2012.
  48. 48.L. Rabeneck, H. El-Serag, J. Davila, and R. Sandler, ‘‘Outcomes of colorectal cancer in the united states: no change in survival (1986-1997).’’ The American journal of gastroenterology, vol. 98, no. 2, p. 471, 2003.
  49. 49.S. A. Karkanis, D. K. Iakovidis, D. E. Maroulis, D. A. Karras, and M. Tzivras, ‘‘Computer-aided tumor detection in endoscopic video using color wavelet features,’’ Information Technology in Biomedicine, IEEE Transactions on, vol. 7, no. 3, pp. 141–152, 2003.
  50. 50.D. K. Iakovidis, D. E. Maroulis, S. A. Karkanis, and A. Brokos, ‘‘A comparative study of texture features for the discrimination of gastric polyps in endoscopic video,’’ in Computer-Based Medical Systems, 2005. Proceedings. 18th IEEE Symposium on. IEEE, 2005, pp. 575–580.
  51. 51.L. A. Alexandre, N. Nobre, and J. Casteleiro, ‘‘Color and position versus texture features for endoscopic polyp detection,’’ in BioMedical Engineering and Informatics, 2008. BMEI 2008. International Conference on, vol. 2. IEEE, 2008, pp. 38–42.
  52. 52.S. Hwang, J. Oh, W. Tavanapong, J. Wong, and P. de Groen, ‘‘Polyp detection in colonoscopy video using elliptical shape feature,’’ in Image Processing, 2007. ICIP 2007. IEEE International Conference on, vol. 2, 2007, pp. II–465–II–468.
  53. 53.J. Bernal, J. Snchez, and F. Vilario, ‘‘Towards automatic polyp detection with a polyp appearance model,’’ Pattern Recognition, vol. 45, no. 9, pp. 3166–3182, 2012.
  54. 54.J. Bernal, J. Sanchez, and F. Vilarino, ‘‘Impact of image preprocessing ´ methods on polyp localization in colonoscopy frames,’’ in Engineering in Medicine and Biology Society (EMBC), 2013 35th Annual International Conference of the IEEE. IEEE, 2013, pp. 7350–7354.
  55. 55.Y. Wang, W. Tavanapong, J. Wong, J. Oh, and P. de Groen, ‘‘Part-based multi-derivative edge cross-section profiles for polyp detection in colonoscopy,’’ Biomedical and Health Informatics, IEEE Journal of, vol. PP, no. 99, pp. 1–1, 2013.
  56. 56.S. Y. Park, D. Sargent, I. Spofford, K. Vosburgh, and Y. A-Rahim, ‘‘A colon video analysis framework for polyp detection,’’ Biomedical Engineering, IEEE Transactions on, vol. 59, no. 5, pp. 1408–1418, 2012.
  57. 57.N. Tajbakhsh, S. Gurudu, and J. Liang, ‘‘A classification-enhanced vote accumulation scheme for detecting colonic polyps,’’ in Abdominal Imaging. Computation and Clinical Applications, ser. Lecture Notes in Computer Science, 2013, vol. 8198, pp. 53–62.
  58. 58.N. Tajbakhsh, C. Chi, S. R. Gurudu, and J. Liang, ‘‘Automatic polyp detection from learned boundaries,’’ in Biomedical Imaging (ISBI), 2014 IEEE 10th International Symposium on, 2014.
  59. 59.N. Tajbakhsh, S. R. Gurudu, and J. Liang, ‘‘Automatic polyp detection using global geometric constraints and local intensity variation patterns,’’ in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2014. Springer, 2014, pp. 179–187.
  60. 60.J. Liang and J. Bi, ‘‘Computer aided detection of pulmonary embolism with tobogganing and multiple instance classification in CT pulmonary angiography,’’ in Information Processing in Medical Imaging. Springer, 2007, pp. 630–641.
  61. 61.K. K. Calder, M. Herbert, and S. O. Henderson, ‘‘The mortality of untreated pulmonary embolism in emergency department patients.’’ Annals of emergency medicine, vol. 45, no. 3, pp. 302–310, 2005. [Online]. Available: http://dx.doi.org/10.1016/j.annemergmed.2004.10.001
  62. 62.G. Sadigh, A. M. Kelly, and P. Cronin, ‘‘Challenges, controversies, and hot topics in pulmonary embolism imaging,’’ American Journal of Roentgenology, vol. 196, no. 3, 2011. [Online]. Available: http://dx.doi.org/10.2214/AJR.10.5830
  63. 63.J. Fairfield, ‘‘Toboggan contrast enhancement for contrast segmentation,’’ in Pattern Recognition, 1990. Proceedings., 10th International Conference on, vol. 1. IEEE, 1990, pp. 712–716.
  64. 64.R. M. Haralick, K. Shanmugam, and I. H. Dinstein, ‘‘Textural features for image classification,’’ Systems, Man and Cybernetics, IEEE Transactions on, no. 6, pp. 610–621, 1973.
  65. 65.N. Tajbakhsh, C. Chi, H. Sharma, Q. Wu, S. R. Gurudu, and J. Liang, ‘‘Automatic assessment of image informativeness in colonoscopy,’’ in Abdominal Imaging. Computational and Clinical Applications. Springer, 2014, pp. 151–158.
  66. 66.M. Arnold, A. Ghosh, G. Lacey, S. Patchett, and H. Mulcahy, ‘‘Indistinct frame detection in colonoscopy videos,’’ in Machine Vision and Image Processing Conference, 2009. IMVIP’09. 13th International. IEEE, 2009, pp. 47–52.
  67. 67.J. Oh, S. Hwang, J. Lee, W. Tavanapong, J. Wong, and P. C. de Groen, ‘‘Informative frame classification for endoscopy video,’’ Medical Image Analysis, vol. 11, no. 2, pp. 110–127, 2007.
  68. 68.R.-M. Menchón-Lara and J.-L. Sancho-G ´ omez, ‘‘Fully automatic seg- ´ mentation of ultrasound common carotid artery images based on machine learning,’’ Neurocomputing, vol. 151, pp. 161–167, 2015.
  69. 69.R.-M. Menchón-Lara, M.-C. Bastida-Jumilla, A. Gonz ´ alez-L ´ opez, and ´ J. L. Sancho-Gomez, ‘‘Automatic evaluation of carotid intima-media ´ thickness in ultrasounds using machine learning,’’ in Natural and Artificial Computation in Engineering and Medical Applications. Springer, 2013, pp. 241–249.
  70. 70.S. Petroudi, C. Loizou, M. Pantziaris, and C. Pattichis, ‘‘Segmentation of the common carotid intima-media complex in ultrasound images using active contours,’’ Biomedical Engineering, IEEE Transactions on , vol. 59, no. 11, pp. 3060–3069, 2012.
  71. 71.X. Xu, Y. Zhou, X. Cheng, E. Song, and G. Li, ‘‘Ultrasound intima–media segmentation using hough transform and dual snake model,’’ Computerized Medical Imaging and Graphics, vol. 36, no. 3, pp. 248–258, 2012.
  72. 72.J. Liang, T. McInerney, and D. Terzopoulos, ‘‘United snakes,’’ Medical image analysis, vol. 10, no. 2, pp. 215–233, 2006.
  73. 73.H. Sharma, R. G. Golla, Y. Zhang, C. B. Kendall, R. T. Hurst, N. Tajbakhsh, and J. Liang, ‘‘Ecg-based frame selection and curvature-based roi detection for measuring carotid intima-media thickness,’’ in SPIE Medical Imaging. International Society for Optics and Photonics, 2014, pp. 904 016–904 016.
  74. 74.H. Lee, R. Grosse, R. Ranganath, and A. Y. Ng, ‘‘Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations,’’ in Proceedings of the 26th Annual International Conference on Machine Learning. ACM, 2009, pp. 609–616.
  75. 75.O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, ‘‘Convolutional neural networks for speech recognition,’’ Audio, Speech, and Language Processing, IEEE/ACM Transactions on, vol. 22, no. 10, pp. 1533–1545, 2014.
  76. 76.D. Wulsin, J. Gupta, R. Mani, J. Blanco, and B. Litt, ‘‘Modeling electroencephalography waveforms with semi-supervised deep belief nets: fast classification and anomaly measurement,’’ Journal of neural engineering, vol. 8, no. 3, p. 036015, 2011.

Citation

MLA
Tajbakhsh, N., et al. “Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?”. IEEE Transactions on Medical Imaging, vol. 35, no. 5, 2016, pp. 1299–312, https://doi.org/10.1109/TMI.2016.2535302.
APA
Tajbakhsh, N., Shin, J. Y., Gurudu, S. R., Hurst, R. T., Kendall, C. B., Gotway, M. B., & Liang, J. (2016). Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?. IEEE Transactions on Medical Imaging, 35(5), 1299–1312. https://doi.org/10.1109/TMI.2016.2535302
Chicago
Tajbakhsh, N., J. Y. Shin, S. R. Gurudu, et al. 2016. “Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?”. IEEE Transactions on Medical Imaging 35 (5): 1299–1312. https://doi.org/10.1109/TMI.2016.2535302.
Harvard
Tajbakhsh, N. et al. (2016) “Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?”, IEEE Transactions on Medical Imaging, 35(5), pp. 1299–1312. Available at: https://doi.org/10.1109/TMI.2016.2535302.
Vancouver
1. Tajbakhsh N, Shin JY, Gurudu SR, Hurst RT, Kendall CB, Gotway MB, Liang J (2016) Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?. IEEE Transactions on Medical Imaging 35:1299–1312

BibTeX

@article{Tajbakhsh_2016, title={Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?}, volume={35}, ISSN={1558-254X}, url={http://dx.doi.org/10.1109/TMI.2016.2535302}, DOI={10.1109/tmi.2016.2535302}, number={5}, journal={IEEE Transactions on Medical Imaging}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Tajbakhsh, Nima and Shin, Jae Y. and Gurudu, Suryakanth R. and Hurst, R. Todd and Kendall, Christopher B. and Gotway, Michael B. and Liang, Jianming}, year={2016}, month=May, pages={1299–1312} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF