A survey of the recent architectures of deep convolutional neural networks

Asifullah KhanAnabia SohailUmme ZahooraAqsa Saeed Qureshi

article2019Artificial Intelligence Review2,796 citations

Classifies recent deep convolutional neural network architectures into seven structural categories—including spatial exploitation, depth, multi-path routing, and attention mechanisms—to explain the design principles driving modern computer vision systems.

Listen

The article surveys the rapid evolution of deep convolutional neural networks, which have achieved strong results on computer vision tasks such as image classification, object detection, and segmentation. Early CNNs struggled with scale and complexity, but hardware advances and large datasets revived interest after 2012, prompting a wave of architectural changes aimed at improving representational power while managing training difficulties like vanishing gradients.

The survey set out to organize the most prominent CNN designs reported between 2012 and 2020 into a clear taxonomy and to explain the basic components, historical development, applications, and remaining challenges. It draws on published architectures, performance benchmarks on datasets such as ImageNet and CIFAR, and comparative tables to identify patterns in how networks are structured.

The analysis shows that the largest gains have come from replacing simple stacked layers with reusable blocks that exploit spatial information at multiple scales, increase depth or width, add shortcut connections, recalibrate feature maps, boost input channels, or apply attention. Notable examples include residual networks that ease training of very deep models, inception-style blocks that capture features at different resolutions, and attention modules that focus computation on relevant regions. These changes often reduce error rates by several percentage points on standard benchmarks while keeping parameter counts manageable.

The findings indicate that architectural innovation, rather than parameter tuning alone, drives most recent progress, yet deeper and wider networks raise computational cost and memory use. This limits deployment on resource-constrained devices and makes models harder to interpret. Applications have expanded beyond vision into speech, text, and video, but success still depends on large labeled datasets and careful hyper-parameter choices.

Leaders should therefore prioritize lightweight or quantized versions of these architectures for practical use, invest in hardware accelerators, and explore ensemble or attention-based refinements to improve robustness. Additional work is needed on automated hyper-parameter search, generative pre-training to reduce labeled-data requirements, and methods that preserve spatial relationships while lowering compute demands. The survey notes that results rest on published benchmarks that may not capture all real-world conditions and that many architectures remain sensitive to implementation details.

Cover for A survey of the recent architectures of deep convolutional neural networks

Abstract

Deep Convolutional Neural Network (CNN) is a special type of Neural Networks, which has shown exemplary performance on several competitions related to Computer Vision and Image Processing. Some of the exciting application areas of CNN include Image Classification and Segmentation, Object Detection, Video Processing, Natural Language Processing, and Speech Recognition. The powerful learning ability of deep CNN is primarily due to the use of multiple feature extraction stages that can automatically learn representations from the data. The availability of a large amount of data and improvement in the hardware technology has accelerated the research in CNNs, and recently interesting deep CNN architectures have been reported. Several inspiring ideas to bring advancements in CNNs have been explored, such as the use of different activation and loss functions, parameter optimization, regularization, and architectural innovations. However, the significant improvement in the representational capacity of the deep CNN is achieved through architectural innovations. Notably, the ideas of exploiting spatial and channel information, depth and width of architecture, and multi-path information processing have gained substantial attention. Similarly, the idea of using a block of layers as a structural unit is also gaining popularity. This survey thus focuses on the intrinsic taxonomy present in the recently reported deep CNN architectures and, consequently, classifies the recent innovations in CNN architectures into seven different categories. These seven categories are based on spatial exploitation, depth, multi-path, width, feature-map exploitation, channel boosting, and attention. Additionally, the elementary understanding of CNN components, current challenges, and applications of CNN are also provided.

Table of Contents

  • A Survey of the Recent Architectures of Deep Convolutional Neural Networks
  • 1 Introduction
  • 2 Basic CNN components
  • 2.1 Convolutional layer
  • 2.2 Pooling layer
  • 2.3 Activation function
  • 2.4 Batch normalization
  • 2.5 Dropout
  • 2.6 Fully connected layer
  • 3 Architectural evolution of deep CNNs
  • 3.1 Origin of CNN: Late 1980s-1999
  • 3.2 Stagnation of CNN: Early 2000
  • 3.3 Revival of CNN: 2006-2011
  • 3.4 Rise of CNN: 2012-2014
  • 3.5 Rapid increase in architectural innovations and applications of CNN: 2015-Present
  • 4 Architectural innovations in CNN
  • 4.1 Spatial Exploitation based CNNs
  • 4.1.1 LeNet
  • 4.1.2 AlexNet
  • 4.1.3 ZfNet
  • 4.1.4 VGG
  • 4.1.5 GoogleNet
  • 4.2 Depth based CNNs
  • 4.2.1 Highway Networks
  • 4.2.2 ResNet
  • 4.2.3 Inception-V3, V4 and Inception-ResNet
  • 4.3 Multi-Path based CNNs
  • 4.3.1 Highway Networks
  • 4.3.2 ResNet
  • 4.3.3 DenseNet
  • **4.4 Width based Multi-Connection CNNs**
  • 4.4.2 Pyramidal Net
  • 4.4.3 Xception
  • 4.4.4 ResNeXt
  • 4.4.5 Inception Family
  • 4.5 Feature-Map (ChannelFMap) Exploitation based CNNs
  • 4.5.1 Squeeze and Excitation Network
  • 4.5.2 Competitive Squeeze and Excitation Networks
  • 4.6 Channel(Input) Exploitation based CNNs
  • 4.6.1 Channel Boosted CNN using TL
  • 4.7 Attention based CNNs
  • 4.7.1 Residual Attention Neural Network
  • 4.7.2 Convolutional Block Attention Module
  • 4.7.3 Concurrent Spatial and Channel Excitation Mechanism
  • 5 Applications of CNNs
  • 5.1 CNN based computer vision and related applications
  • 5.2 CNN based natural language processing
  • 5.3 CNN based object detection and segmentation
  • **5.4 CNN based image classification**
  • **5.5 CNN based speech recognition**
  • 5.6 CNN based video processing
  • 5.7 CNN for low resolution images
  • 5.8 CNN for resource limited systems
  • 5.9 CNN for 1D-Data
  • 6 CNN challenges
  • 7 Future directions
  • 8 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Seven-Category Taxonomy of Deep CNN Architectural Innovations

    model/method

    Modern architectural advancements in deep Convolutional Neural Networks (CNNs) from 2012 to 2020 are structured into a seven-category taxonomy based on their core structural modification principles:

    1. Spatial Exploitation: Adjusting spatial filter dimensions, utilizing stacks of small receptive field kernels (3×33 \times 3) to emulate large receptive fields, and multi-scale kernel branching (e.g., LeNet, AlexNet, ZFNet, VGG, GoogLeNet).
    2. Depth-Based Design: Extending the number of parameterized layers to learn hierarchical feature representations at different abstraction levels, using auxiliary loss classifiers or specialized training schemes to mitigate optimization degradation (e.g., Highway Networks, ResNet, Inception-v3, Inception-v4).
    3. Multi-Path Connectivity: Implementing cross-layer shortcut connections and bypass pathways (such as gating mechanisms, additive skip connections, or concatenation) to enable unimpeded gradient flow and eliminate vanishing gradient problems (e.g., Highway Networks, ResNet, DenseNet, FractalNet, DelugeNet).
    4. Width-Based Multi-Connection: Increasing the channel capacity of layers and introducing cardinality (sets of parallel transformation branches) within blocks to prevent feature reuse stagnation and expand representation space (e.g., Wide ResNet, Pyramidal Net, Xception, ResNeXt).
    5. Feature-Map Exploitation: Recalibrating, weighting, or selecting intermediate channel representations based on inter-channel dependencies and class-discriminative relevance (e.g., Squeeze-and-Excitation Networks, Competitive Squeeze-and-Excitation Networks).
    6. Channel (Input) Boosting: Augmenting the initial input channel space with synthetic, information-rich auxiliary channels generated by pre-trained deep generative models via transfer learning (e.g., Channel Boosted CNN).
    7. Attention Mechanisms: Dynamically generating spatial and channel-wise attention masks via top-down feedback or sequential/concurrent pooling to direct network capacity toward context-relevant image regions (e.g., Residual Attention Network, CBAM, concurrent spatial and channel excitation modules).
  2. Knowl 2 — Squeeze-and-Excitation (SE) Channel Recalibration Mechanism

    model/method

    The Squeeze-and-Excitation (SE) block adaptively recalibrates channel-wise feature responses by explicitly modeling interdependencies between intermediate feature maps. For an intermediate feature tensor containing KK feature maps at layer ll, where each map Flk∈RP×Q\mathbf{F}_l^k \in \mathbb{R}^{P \times Q} has spatial height PP and width QQ, the module executes two operations:

    1. Squeeze Operation (gsqg_{sq}): Generates channel-level descriptors by aggregating spatial dimensions using global average pooling:

    slk=gsq(Flk)=1P×Q∑p=1P∑q=1Qflk(p,q)s_l^k = g_{sq}(\mathbf{F}_l^k) = \frac{1}{P \times Q} \sum_{p=1}^{P} \sum_{q=1}^{Q} f_l^k(p, q)

    where flk(p,q)f_l^k(p, q) denotes the feature activation at spatial position (p,q)(p, q), and slk∈Rs_l^k \in \mathbb{R} is the kk-th element of the squeezed descriptor vector Sl=[sl1,sl2,…,slK]\mathbf{S}_l = [s_l^1, s_l^2, \dots, s_l^K].

    1. Excitation Operation (gexg_{ex}): Employs a two-layer feed-forward bottleneck neural network with reduction ratio rr, parameterized by weight matrices w1∈RKr×K\mathbf{w}_1 \in \mathbb{R}^{\frac{K}{r} \times K} and w2∈RK×Kr\mathbf{w}_2 \in \mathbb{R}^{K \times \frac{K}{r}}, followed by a non-linear activation gtg_t (ReLU) and a gating sigmoid activation gsgg_{s_g}:

    yl+1k=gex(Sl)=gsg(w2⋅gt(w1⋅Sl))y_{l+1}^k = g_{ex}(\mathbf{S}_l) = g_{s_g}(\mathbf{w}_2 \cdot g_t(\mathbf{w}_1 \cdot \mathbf{S}_l))

    where yl+1k∈[0,1]y_{l+1}^k \in [0, 1] serves as the dynamic scalar weight for the kk-th feature map. The output feature map is produced by channel-wise scaling F~lk=yl+1kFlk\tilde{\mathbf{F}}_l^k = y_{l+1}^k \mathbf{F}_l^k.

  3. Knowl 3 — Channel Boosted CNN (CB-CNN) Input Channel Augmentation

    model/method

    Channel Boosted Convolutional Neural Network (CB-CNN) enhances the representational capacity of deep CNN discriminators by expanding the input channel space using artificial auxiliary channels created by pre-trained generative models.

    Let IC\mathbf{I}_C represent the original input channel tensor, and let Am\mathbf{A}_m denote the artificial auxiliary channel generated by the mm-th auxiliary generative learner (such as a deep autoencoder) for m∈{1,2,…,M}m \in \{1, 2, \dots, M\} via inductive transfer learning. The channel-boosted input tensor IB\mathbf{I}_B is formed through channel concatenation gk(⋅)g_k(\cdot):

    IB=gk(IC,[A1,A2,…,AM])\mathbf{I}_B = g_k(\mathbf{I}_C, [\mathbf{A}_1, \mathbf{A}_2, \dots, \mathbf{A}_M])

    The kk-th output feature map Flk\mathbf{F}_l^k at layer ll of the deep discriminative CNN is obtained by convolving the boosted input tensor with the convolutional kernel kl\mathbf{k}_l:

    Flk=gc(IB,kl)\mathbf{F}_l^k = g_c(\mathbf{I}_B, \mathbf{k}_l)

    where gc(⋅)g_c(\cdot) represents the standard convolution operation. This formulation enables the CNN discriminator to access both raw input data and learned explanatory variations.

  4. Knowl 4 — Competitive Squeeze-and-Excitation (CMPE-SE) Inner-Imaging Block

    model/method

    The Competitive Inner-Imaging Squeeze and Excitation (CMPE-SE) module recalibrates feature maps in residual networks by evaluating competition between descriptors from both the identity mapping branch and the residual transformation branch.

    Let Flk\mathbf{F}_l^k represent the identity input feature map at layer ll, and let Fm+1k′\mathbf{F}_{m+1}^{k'} be the corresponding residual mapped feature map after intermediate convolutions l→ml \to m. Descriptors for both streams are extracted via spatial squeeze operations gsq(⋅)g_{sq}(\cdot) using global average pooling:

    Slk=gsq(Flk),Sm+1k′=gsq(Fm+1k′)\mathbf{S}_l^k = g_{sq}(\mathbf{F}_l^k), \quad \mathbf{S}_{m+1}^{k'} = g_{sq}(\mathbf{F}_{m+1}^{k'})

    The identity and residual summary vectors are merged using concatenation gk(⋅)g_k(\cdot) and passed through a joint excitation function gex(⋅)g_{ex}(\cdot):

    Ym+1k=gex(gk(Slk,Sm+1k′))\mathbf{Y}_{m+1}^k = g_{ex}(g_k(\mathbf{S}_l^k, \mathbf{S}_{m+1}^{k'}))

    The recalibrated residual feature map Fm+1k\mathbf{F}_{m+1}^k is obtained by multiplying the excitation mask Ym+1k\mathbf{Y}_{m+1}^k with the residual features:

    Fm+1k=Ym+1k⋅Fm+1k′\mathbf{F}_{m+1}^k = \mathbf{Y}_{m+1}^k \cdot \mathbf{F}_{m+1}^{k'}

    This competition mechanism prevents the redundancy observed when standard SE blocks operate solely on isolated residual branch outputs.

  5. Knowl 5 — Multi-Path Shortcut Routing in Highway Networks and Residual Networks

    model/method

    Multi-path architectures resolve gradient degradation in deep networks by providing parallel information highways that bypass intermediate layers:

    1. Highway Networks: Flow across layers is regulated by parameterized transform gates gtgg_{t_g} and carry gates gcgg_{c_g}:

    Fl+1k=gc(Flk,kl)⋅gtg(Flk,tgkl)+Flk⋅gcg(Flk,cgkl)\mathbf{F}_{l+1}^k = g_c(\mathbf{F}_l^k, \mathbf{k}_l) \cdot g_{t_g}(\mathbf{F}_l^k, {}^{t_g}\mathbf{k}_l) + \mathbf{F}_l^k \cdot g_{c_g}(\mathbf{F}_l^k, {}^{c_g}\mathbf{k}_l)

    where gcg(Flk,cgkl)=1−gtg(Flk,tgkl)g_{c_g}(\mathbf{F}_l^k, {}^{c_g}\mathbf{k}_l) = 1 - g_{t_g}(\mathbf{F}_l^k, {}^{t_g}\mathbf{k}_l), allowing smooth transition between transformed output and identity passthrough.

    1. Deep Residual Networks (ResNet): Gating parameters are eliminated in favor of data-independent, parameter-free additive identity shortcuts:

    Fm+1k′=gc(Fl→mk,kl→m)+Flk\mathbf{F}_{m+1}^{k'} = g_c(\mathbf{F}_{l\to m}^k, \mathbf{k}_{l\to m}) + \mathbf{F}_l^k Fm+1k=ga(Fm+1k′)\mathbf{F}_{m+1}^k = g_a(\mathbf{F}_{m+1}^{k'})

    where Flk\mathbf{F}_l^k is the input feature map, gc(Fl→mk,kl→m)g_c(\mathbf{F}_{l\to m}^k, \mathbf{k}_{l\to m}) represents the residual transformation across layers l→ml \to m, and ga(⋅)g_a(\cdot) denotes a non-linear activation (e.g., ReLU). The residual function to be optimized is:

    gc(Fl→mk,kl→m)=Fm+1k′−Flkg_c(\mathbf{F}_{l\to m}^k, \mathbf{k}_{l\to m}) = \mathbf{F}_{m+1}^{k'} - \mathbf{F}_l^k

    ensuring gradients pass directly back to early layers without diminishing.

  6. Knowl 6 — Dense Cross-Layer Connectivity in DenseNet

    model/method

    DenseNet (Densely Connected Convolutional Networks) establishes feed-forward connections from every layer to all subsequent layers via feature-map concatenation rather than summation.

    For an input image tensor IC\mathbf{I}_C transformed by initial layer kernel k1\mathbf{k}_1 into feature map F2k=gc(IC,k1)\mathbf{F}_2^k = g_c(\mathbf{I}_C, \mathbf{k}_1), the feature input to layer ll combines all preceding feature representations:

    Flk=gk(F1k,F2k,…,Fl−1k)\mathbf{F}_l^k = g_k(\mathbf{F}_1^k, \mathbf{F}_2^k, \dots, \mathbf{F}_{l-1}^k)

    where gk(⋅)g_k(\cdot) denotes channel-wise concatenation. In an ll-layer network, this design creates l(l+1)2\frac{l(l+1)}{2} direct connections compared to ll connections in standard sequential models. This allows every layer direct access to gradients from the loss function, minimizes redundant feature relearning, and preserves low-level and high-level representations across the network hierarchy.

  7. Knowl 7 — Channel Width Expansion in Pyramidal Residual Networks

    model/method

    Pyramidal Residual Networks (PyramidalNet) replace the abrupt doubling of channel dimensions at downsampling stages with a continuous, gradual increase in feature-map channels across all residual units.

    The channel dimension dld_l of the ll-th residual block is given by:

    dl={16if l=1⌊dl−1+λn⌋if 2≤l≤n+1d_l = \begin{cases} 16 & \text{if } l = 1 \\ \left\lfloor d_{l-1} + \frac{\lambda}{n} \right\rfloor & \text{if } 2 \le l \le n + 1 \end{cases}

    where nn is the total number of residual units, λ\lambda is a step-size hyperparameter controlling the final width, and λn\frac{\lambda}{n} represents the per-unit channel growth rate. Residual connections across units with mismatched channel dimensions are implemented using zero-padded identity mapping to avoid parameter overhead. Channel growth is implemented either linearly (additive widening) or geometrically (multiplicative widening).

  8. Knowl 8 — Decoupled Spatial and Channel Convolutions in Xception

    model/method

    Xception decouples spatial correlation learning from cross-channel correlation learning by substituting standard convolution operations with depthwise separable convolutions across nn transformation segments (where cardinality nn equals the channel dimension).

    Spatial convolution is computed independently on each input channel Flk\mathbf{F}_l^k:

    fl+1k(p,q)=∑x,yflk(x,y)⋅elk(u,v)f_{l+1}^k(p, q) = \sum_{x, y} f_l^k(x, y) \cdot e_l^k(u, v)

    where elk(u,v)e_l^k(u, v) is a single 3×33 \times 3 kernel applied to channel kk. This is followed by a pointwise (1×11 \times 1) cross-channel convolution applied across all intermediate channels [Fl+11,…,Fl+1K][\mathbf{F}_{l+1}^1, \dots, \mathbf{F}_{l+1}^K]:

    Fl+2k=gc(Fl+1k,kl+1)\mathbf{F}_{l+2}^k = g_c(\mathbf{F}_{l+1}^k, \mathbf{k}_{l+1})

    where kl+1\mathbf{k}_{l+1} is a 1×11 \times 1 kernel at layer l+1l+1. This decoupling separates spatial feature extraction from channel cross-correlation mapping, improving parameter and computational efficiency.

  9. Knowl 9 — Multi-Scale and Spatial-Channel Attention Mechanisms in CNNs

    model/method

    Attention modules in deep CNNs dynamically reweight intermediate feature representations by modeling spatial context, channel statistics, or their interactions:

    • Residual Attention Network (RAN): Combines a feed-forward trunk processing branch gtm(Flk)g_{tm}(\mathbf{F}_l^k) with an encoder-decoder soft mask branch gsm(Flk)g_{sm}(\mathbf{F}_l^k) that generates object-aware weights through bottom-up top-down processing:

    gam(Flk)=gsm(Flk)⋅gtm(Flk)g_{am}(\mathbf{F}_l^k) = g_{sm}(\mathbf{F}_l^k) \cdot g_{tm}(\mathbf{F}_l^k)

    • Convolutional Block Attention Module (CBAM): Sequentially computes 1D channel attention (combining global average pooling and global max pooling descriptors through an MLP) and 2D spatial attention (combining channel-axis average and max projections via a 7×77 \times 7 convolution) to infer what and where to focus.
    • Concurrent Spatial and Channel Excitation (scSE): Evaluates spatial squeeze and channel excitation (cSE) alongside channel squeeze and spatial excitation (sSE) concurrently, adding both recalibrated representations together to enhance segmentation and localization tasks.
  10. Knowl 10 — Comparative Depth, Complexity, and Error Rates of Core CNN Architectures

    data/table

    The following table summarizes parameters, depth, category classifications, and top-5 error rates on benchmark vision datasets across representative deep CNN architectures from 1998 to 2018.

    Architecture Year Parameters Depth Error Rate Category
    LeNet 1998 0.060 M 5 MNIST: 0.95% Spatial Exploitation
    AlexNet 2012 60 M 8 ImageNet: 16.4% Spatial Exploitation
    ZFNet 2014 60 M 8 ImageNet: 11.7% Spatial Exploitation
    VGG 2014 138 M 19 ImageNet: 7.3% Spatial Exploitation
    GoogLeNet 2015 4 M 22 ImageNet: 6.7% Spatial Exploitation
    Inception-V3 2015 23.6 M 159 ImageNet: 3.5% Depth + Width
    Highway Networks 2015 2.3 M 19 CIFAR-10: 7.76% Depth + Multi-Path
    Inception-V4 2016 35 M 70 ImageNet: 4.01% Depth + Width
    Inception-ResNet 2016 55.8 M 572 ImageNet: 3.52% Depth + Width + Multi-Path
    ResNet 2016 25.6 M 152 ImageNet: 3.6% Depth + Multi-Path
    DelugeNet 2016 20.2 M 146 CIFAR-10: 3.76% Multi-Path
    FractalNet 2016 38.6 M 40 CIFAR-10: 7.27% Multi-Path
    WideResNet 2016 36.5 M 28 CIFAR-10: 3.89% Width
    Xception 2017 22.8 M 126 ImageNet: 0.055 Width
    ResNeXt 2017 68.1 M 101 ImageNet: 4.4% Width
    SE-Net 2017 27.5 M 152 ImageNet: 2.3% Feature-Map Exploitation
    DenseNet 2017 25.6 M 190 CIFAR-10+: 3.46% Multi-Path
    PolyNet 2017 92 M – ImageNet: 4.25% Width
    PyramidalNet 2017 116.4 M 200 ImageNet: 4.7% Width
    CBAM (ResNeXt101) 2018 48.96 M 101 ImageNet: 5.59% Attention
    CMPE-SE-WRN-28 2018 36.92 M 152 CIFAR-10: 3.58% Feature-Map Exploitation

    These results demonstrate the evolution from early shallow spatial models (AlexNet at 16.4% error) toward ultra-deep multi-path architectures (ResNet, DenseNet) and feature/attention recalibration blocks (SE-Net reaching 2.3% error on ImageNet).

Coverage note — Omitted introductory historical timelines (1980s–2011 eras), generic computer vision/NLP application summaries, and general listings of online software libraries/hardware platforms, as these represent secondary background context rather than core architectural taxonomy contributions.

References

  1. 1.Abbas Q, Ibrahim MEA, Jaffar MA (2019) A comprehensive review of recent advances on deep vision systems. Artif Intell Rev 52:39–76. doi: 10.1007/s10462-018-9633-3
  2. 2.Abdel-Hamid O, Deng L, Yu D (2013) Exploring convolutional neural network structures and optimization techniques for speech recognition. In: Interspeech. pp 1173–1175
  3. 3.Abdel-Hamid O, Mohamed AR, Jiang H, Penn G (2012) Applying convolutional neural networks concepts to hybrid NN-HMM model for speech recognition. ICASSP, IEEE Int Conf Acoust Speech Signal Process - Proc 4277–4280. doi: 10.1007/978-3-319-96145-3_2
  4. 4.Abdeljaber O, Avci O, Kiranyaz S, et al (2017) Real-time vibration-based structural damage detection using one-dimensional convolutional neural networks. J Sound Vib. doi: 10.1016/j.jsv.2016.10.043
  5. 5.Abdulkader A (2006) Two-tier approach for Arabic offline handwriting recognition. In: Tenth International Workshop on Frontiers in Handwriting Recognition
  6. 6.Ahmed U, Khan A, Khan SH, et al (2019) Transfer Learning and Meta Classification Based Deep Churn Prediction System for Telecom Industry. 1–10
  7. 7.Akar E, Marques O, Andrews WA, Furht B (2019) Cloud-Based Skin Lesion Diagnosis System Using Convolutional Neural Networks. In: Intelligent Computing-Proceedings of the Computing Conference. pp 982–1000
  8. 8.Amer M, Maul T (2019) A review of modularization techniques in artificial neural networks. Artif Intell Rev 52:527–561. doi: 10.1007/s10462-019-09706-7
  9. 9.Aurisano A, Radovic A, Rocco D, et al (2016) A convolutional neural network neutrino event classifier. J Instrum. doi: 10.1088/1748-0221/11/09/P09001
  10. 10.Aziz A, Sohail A, Fahad L, et al (2020) Channel Boosted Convolutional Neural Network for Classification of Mitotic Nuclei using Histopathological Images. In: 2020 17th International Bhurban Conference on Applied Sciences and Technology (IBCAST). pp 277–284
  11. 11.Badrinarayanan V, Kendall A, Cipolla R (2017) SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation. IEEE Trans Pattern Anal Mach Intell. doi: 10.1109/TPAMI.2016.2644615
  12. 12.Batmaz Z, Yurekli A, Bilge A, Kaleli C (2019) A review on deep learning for recommender systems: challenges and remedies. Artif Intell Rev 52:1–37. doi: 10.1007/s10462-018-9654-y
  13. 13.Bay H, Ess A, Tuytelaars T, Van Gool L (2008) Speeded-Up Robust Features (SURF). Comput Vis Image Underst 110:346–359. doi: 10.1016/j.cviu.2007.09.014
  14. 14.Bengio Y (2009) Learning Deep Architectures for AI. Found Trends® Mach Learn 2:1–127. doi: 10.1561/2200000006
  15. 15.Bengio Y (2013) Deep learning of representations: Looking forward. In: International Conference on Statistical Language and Speech Processing. Springer, pp 1–37
  16. 16.Bengio Y, Courville A, Vincent P (2013) Representation learning: A review and new perspectives. IEEE Trans Pattern Anal Mach Intell 35:1798–1828. doi: 10.1109/TPAMI.2013.50
  17. 17.Bengio Y, Lamblin P, Popovici D, Larochelle H (2007) Greedy layer-wise training of deep networks. In: Advances in neural information processing systems. The MIT Press, pp 153–160
  18. 18.Berg A, Deng J, Fei-Fei L (2010) Large scale visual recognition challenge 2010
  19. 19.Bettoni M, Urgese G, Kobayashi Y, et al (2017) A Convolutional Neural Network Fully Implemented on FPGA for Embedded Platforms. 49–52. doi: 10.1109/NGCAS.2017.16
  20. 20.Bhunia AK, Konwer A, Bhunia AK, et al (2019) Script identification in natural scene image and video frames using an attention based Convolutional-LSTM network. Pattern Recognit 85:172–184
  21. 21.Boureau Y (2009) Icml2010B.Pdf. doi: citeulike-article-id:8496352
  22. 22.Bouvrie J (2006) 1 Introduction Notes on Convolutional Neural Networks. doi: http://dx.doi.org/10.1016/j.protcy.2014.09.007
  23. 23.Bulat A, Tzimiropoulos G (2016) Human Pose Estimation via Convolutional Part Heatmap Regression BT - Computer Vision – ECCV 2016. In: Leibe B, Matas J, Sebe N, Welling M (eds). Springer International Publishing, Cham, pp 717–732
  24. 24.Cai Z, Vasconcelos N (2019) Cascade R-CNN: High Quality Object Detection and Instance Segmentation. IEEE Trans Pattern Anal Mach Intell. doi: 10.1109/tpami.2019.2956516
  25. 25.Chapelle O (1998) Support vector machines for image classification. Stage deuxième année magistère d’informatique l’École Norm Supérieur Lyon 10:1055–1064. doi: 10.1109/72.788646
  26. 26.Chellapilla K, Puri S, Simard P (2006) High performance convolutional neural networks for document processing. In: Tenth International Workshop on Frontiers in Handwriting Recognition
  27. 27.Chen W, Wilson JT, Tyree S, et al (2015) Compressing neural networks with the hashing trick. In: 32nd International Conference on Machine Learning, ICML 2015
  28. 28.Chen Y-N, Han C-C, Wang C-T, et al (2006) The application of a convolution neural network on face and license plate detection. In: Pattern Recognition, 2006. ICPR 2006. 18th International Conference on. pp 552–555
  29. 29.Chevalier M, Thome N, Cord M, et al (2015) LR-CNN for fine-grained classification with varying resolution. In: 2015 IEEE International Conference on Image Processing (ICIP). IEEE, pp 3101–3105
  30. 30.Chollet F (2017) Xception: Deep learning with depthwise separable convolutions. arXiv Prepr 1610–2357
  31. 31.Chouhan N, Khan A (2019) Network anomaly detection using channel boosted and residual learning based deep convolutional neural network. Appl Soft Comput 105612
  32. 32.Ciresan D, Giusti A, Gambardella LM, Schmidhuber J (2012) Deep neural networks segment neuronal membranes in electron microscopy images. In: Advances in neural information processing systems. pp 2843–2851
  33. 33.Cireşan D, Meier U, Masci J, Schmidhuber J (2012) Multi-column deep neural network for traffic sign classification. Neural Networks 32:333–338. doi: 10.1016/j.neunet.2012.02.023
  34. 34.Ciresan DC, Ciresan DC, Meier U, Schmidhuber J (2018) Multi-column deep neural networks for image classification. IEEE Comput Soc Conf Comput Vis Pattern Recognit
  35. 35.Cireşan DC, Giusti A, Gambardella LM, Schmidhuber J (2013) Mitosis Detection in Breast Cancer Histology Images with Deep Neural Networks BT - Medical Image Computing and Computer-Assisted Intervention – MICCAI 2013. In: Proceedings MICCAI. pp 411–418
  36. 36.Ciresan DC, Meier U, Gambardella LM, Schmidhuber J (2010) Deep , Big , Simple Neural Nets for Handwritten. Neural Comput 22:3207–3220
  37. 37.Cireşan DC, Meier U, Masci J, et al (2011) High-Performance Neural Networks for Visual Object Classification. arXiv Prepr arXiv11020183
  38. 38.Collobert R, Weston J (2008) A unified architecture for natural language processing: Deep neural networks with multitask learning. In: Proceedings of the 25th international conference on Machine learning. ACM, pp 160–167
  39. 39.Csáji B (2001) Approximation with artificial neural networks. MSc thesis 45. doi: 10.1.1.101.2647
  40. 40.Dahl G, Mohamed A, Hinton GE (2010) Phone recognition with the mean-covariance restricted Boltzmann machine. In: Advances in neural information processing systems. pp 469–477
  41. 41.Dahl GE, Sainath TN, Hinton GE (2013) Improving deep neural networks for LVCSR using rectified linear units and dropout. In: Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on. IEEE, pp 8609–8613
  42. 42.Dai J, Li Y, He K, Sun J (2016) R-FCN: Object Detection via Region-based Fully Convolutional Networks. doi: 10.1016/j.jpowsour.2007.02.075
  43. 43.Dalal N, Triggs W (2004) Histograms of Oriented Gradients for Human Detection. 2005 IEEE Comput Soc Conf Comput Vis Pattern Recognit CVPR05 1:886–893. doi: 10.1109/CVPR.2005.177
  44. 44.Dauphin YN, De Vries H, Bengio Y (2015) Equilibrated adaptive learning rates for non-convex optimization. Adv Neural Inf Process Syst 2015-Janua:1504–1512
  45. 45.Dauphin YN, Fan A, Auli M, Grangier D (2017) Language modeling with gated convolutional networks. In: Proceedings of the 34th International Conference on Machine Learning-Volume 70. pp 933–941
  46. 46.de Vries H, Memisevic R, Courville A (2016) Deep learning vector quantization. In: European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning
  47. 47.Decoste D, Schölkopf B (2002) Training invariant support vector machines. Mach Learn 46:161–190
  48. 48.Delalleau O, Bengio Y (2011) Shallow vs. deep sum-product networks. In: Advances in Neural Information Processing Systems. pp 666–674
  49. 49.Deng L (2012) The MNIST database of handwritten digit images for machine learning research [best of the web]. IEEE Signal Process Mag 29:141–142
  50. 50.Deng L, Yu D, Delft B— (2013) Deep Learning: Methods and Applications Foundations and Trends R in Signal Processing. Signal Processing 7:3–4. doi: 10.1561/2000000039
  51. 51.Do MN, Vetterli M (2005) The contourlet transform: an efficient directional multiresolution image representation. IEEE Trans image Process 14:2091–2106
  52. 52.Dollár P, Tu Z, Perona P, Belongie S (2009) Integral channel features
  53. 53.Donahue J, Anne Hendricks L, Guadarrama S, et al (2015) Long-term recurrent convolutional networks for visual recognition and description. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp 2625–2634
  54. 54.Dong C, Loy CC, He K, Tang X (2016) Image super-resolution using deep convolutional networks. IEEE Trans Pattern Anal Mach Intell 38:295–307
  55. 55.Erhan D, Bengio Y, Courville A, Vincent P (2009) Visualizing higher-layer features of a deep network. Univ Montr 1341:1
  56. 56.Farfade SS, Saberian MJ, Li L-J (2015) Multi-view Face Detection Using Deep Convolutional Neural Networks. In: Proceedings of the 5th ACM on International Conference on Multimedia Retrieval - ICMR ’15. ACM Press, New York, New York, USA, pp 643–650
  57. 57.Fasel B (2002) Facial expression analysis using shape and motion information extracted by convolutional neural networks. In: Neural Networks for Signal Processing, 2002. Proceedings of the 2002 12th IEEE Workshop on. pp 607–616
  58. 58.Frizzi S, Kaabi R, Bouchouicha M, et al (2016) Convolutional neural network for video fire and smoke detection. In: IECON 2016-42nd Annual Conference of the IEEE Industrial Electronics Society. IEEE, pp 877–882
  59. 59.Frome A, Cheung G, Abdulkader A, et al (2009) Large-scale privacy protection in Google Street View. In: Proceedings of the IEEE International Conference on Computer Vision
  60. 60.Frosst N, Hinton G rey (2018) Distilling a neural network into a soft decision tree. In: CEUR Workshop Proceedings
  61. 61.Fukushima K (1988) Neocognitron: A hierarchical neural network capable of visual pattern recognition. Neural networks 1:119–130
  62. 62.Fukushima K, Miyake S (1982) Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition. In: Competition and cooperation in neural nets. Springer, pp 267–285
  63. 63.Garcia C, Delakis M (2004) Convolutional face finder: A neural architecture for fast and robust face detection. IEEE Trans Pattern Anal Mach Intell. doi: 10.1109/TPAMI.2004.97
  64. 64.Gardner MW, Dorling SR (1998) Artificial neural networks (the multilayer perceptron)—a review of applications in the atmospheric sciences. Atmos Environ 32:2627–2636
  65. 65.Geng X, Lin J, Zhao B, et al (2019) Hardware-Aware Softmax Approximation for Deep Neural Networks. In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics). pp 107–122
  66. 66.Gidaris S, Komodakis N (2015) Object detection via a multi-region and semantic segmentation-aware U model. Proc IEEE Int Conf Comput Vis 2015 Inter:1134–1142. doi: 10.1109/ICCV.2015.135
  67. 67.Girshick R (2015) Fast R-CNN. In: Proceedings of the IEEE International Conference on Computer Vision. pp 1440–1448
  68. 68.Giusti A, Ciresan DC, Masci J, et al (2013) Fast image scanning with deep max-pooling convolutional neural networks. In: 2013 IEEE International Conference on Image Processing. IEEE, pp 4034–4038
  69. 69.Glorot X, Bengio Y (2010) Understanding the difficulty of training deep feedforward neural networks. In: Proceedings of the thirteenth international conference on artificial intelligence and statistics. pp 249–256
  70. 70.Goh H, Thome N, Cord M, Lim J-H (2013) Top-down regularization of deep belief networks. In: Advances in Neural Information Processing Systems (NIPS). pp 1878–1886
  71. 71.Grill-Spector K, Weiner KS, Gomez J, et al (2018) The functional neuroanatomy of face perception: From brain measurements to deep neural networks. Interface Focus 8:20180013. doi: 10.1098/rsfs.2018.0013
  72. 72.Grün F, Rupprecht C, Navab N, Tombari F (2016) A Taxonomy and Library for Visualizing Learned Features in Convolutional Neural Networks. 48:. doi: 10.1080/10962247.2014.948229
  73. 73.Gu J, Wang Z, Kuen J, et al (2018) Recent advances in convolutional neural networks. Pattern Recognit 77:354–377. doi: 10.1016/j.patcog.2017.10.013
  74. 74.Guo Y, Liu Y, Oerlemans A, et al (2016) Deep learning for visual understanding: A review. Neurocomputing 187:27–48. doi: 10.1016/j.neucom.2015.09.116
  75. 75.Hamel P, Eck D (2010) Learning Features from Music Audio with Deep Belief Networks. In: ISMIR. Utrecht, The Netherlands, pp 339–344
  76. 76.Han D, Kim J, Kim J (2017) Deep Pyramidal Residual Networks. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 6307–6315
  77. 77.Han S, Mao H, Dally WJ (2016) Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. In: 4th International Conference on Learning Representations, ICLR 2016 - Conference Track Proceedings
  78. 78.Han W, Feng R, Wang L, Gao L (2018) Adaptive Spatial-Scale-Aware Deep Convolutional Neural Network for High-Resolution Remote Sensing Imagery Scene Classification. IGARSS 2018 - 2018 IEEE Int Geosci Remote Sens Symp 4736–4739. doi: 10.1109/IGARSS.2018.8518290
  79. 79.Hanin B, Sellke M (2017) Approximating Continuous Functions by ReLU Nets of Minimal Width. arXiv Prepr arXiv171011278
  80. 80.He K, Gkioxari G, Dollar P, Girshick R (2017) Mask R-CNN. In: Proceedings of the IEEE International Conference on Computer Vision
  81. 81.He K, Zhang X, Ren S, Sun J (2015a) Deep Residual Learning for Image Recognition. Multimed Tools Appl 77:10437–10453. doi: 10.1007/s11042-017-4440-4
  82. 82.He K, Zhang X, Ren S, Sun J (2015b) Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE Trans Pattern Anal Mach Intell 37:1904–1916
  83. 83.Heikkilä M, Pietikäinen M, Schmid C (2009) Description of interest regions with local binary patterns. Pattern Recognit 42:425–436. doi: 10.1016/j.patcog.2008.08.014
  84. 84.Hinton G, Deng L, Yu D, et al (2012a) Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Process Mag 29:82–97
  85. 85.Hinton G, Sabour S, Frosst N (2018) Matrix capsules with EM routing. In: 6th International Conference on Learning Representations, ICLR 2018 - Conference Track Proceedings
  86. 86.Hinton GE, Krizhevsky A, Wang SD (2011) Transforming auto-encoders. In: International Conference on Artificial Neural Networks. Springer, pp 44–51
  87. 87.Hinton GE, Osindero S, Teh Y-W (2006) A fast learning algorithm for deep belief nets. Neural Comput 18:1527–1554
  88. 88.Hinton GE, Srivastava N, Krizhevsky A, et al (2012b) Improving neural networks by preventing co-adaptation of feature detectors. arXiv Prepr ArXiv 12070580 1–18
  89. 89.Hochreiter S (1998) The vanishing gradient problem during learning recurrent neural nets and problem solutions. Int J Uncertainty, Fuzziness Knowledge-Based Syst 6:107–116
  90. 90.Howard AG, Zhu M, Chen B, et al (2017) MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv Prepr arXiv170404861
  91. 91.Hu B, Lu Z, Li H, Chen Q (2011) Topic modeling for named entity queries. In: Proceedings of the 20th ACM international conference on Information and knowledge management - CIKM ’11. ACM Press, New York, New York, USA, p 2009
  92. 92.Hu J, Shen L, Sun G (2018a) Squeeze-and-Excitation Networks. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, pp 7132–7141
  93. 93.Hu Y, Wen G, Luo M, et al (2018b) Competitive Inner-Imaging Squeeze and Excitation for Residual Network. doi: arXiv:1807.08920v3
  94. 94.Huang G, Liu Z, Van Der Maaten L, Weinberger KQ (2017) Densely connected convolutional networks. Proc - 30th IEEE Conf Comput Vis Pattern Recognition, CVPR 2017 2017-Janua:2261–2269. doi: 10.1109/CVPR.2017.243
  95. 95.Huang G, Sun Y, Liu Z, et al (2016a) Deep networks with stochastic depth. In: European Conference on Computer Vision. Springer, pp 646–661
  96. 96.Huang G, Sun Y, Liu Z, et al (2016b) Deep Networks with Stochastic Depth BT - Computer Vision – ECCV 2016. In: European Conference on Computer Vision. Springer, pp 646–661
  97. 97.Huang KY, Wu CH, Hong QB, et al (2019) Speech Emotion Recognition Using Deep Neural Network Considering Verbal and Nonverbal Speech Sounds. In: ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
  98. 98.Huang Y, Cheng Y, Bapna A, et al (2018) GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. arXiv Prepr arXiv181106965 2014:
  99. 99.Hubel DH, Wiesel TN (1962) Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex. J Physiol 160:106–154. doi: 10.1113/jphysiol.1962.sp006837
  100. 100.Hubel DH, Wiesel TN (1959) Receptive fields of single neurones in the cat’s striate cortex. J Physiol. doi: 10.1113/jphysiol.1959.sp006308
  101. 101.Hubel DH, Wiesel TN (1968) Receptive fields and functional architecture of monkey striate cortex. J Physiol 195:215–243. doi: 10.1113/jphysiol.1968.sp008455
  102. 102.Ian Goodfellow, Bengio Y, Courville A (2017) Deep learning. Nat Methods 13:35. doi: 10.1038/nmeth.3707
  103. 103.Ioffe S, Szegedy C (2015) Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. doi: 10.1016/j.molstruc.2016.12.061
  104. 104.Jaderberg M, Simonyan K, Zisserman A, Kavukcuoglu K (2015) Spatial Transformer Networks. 1–15. doi: 10.1038/nbt.3343
  105. 105.Jarrett K, Kavukcuoglu K, Ranzato M, LeCun Y (2009) What is the best multi-stage architecture for object recognition? BT - Computer Vision, 2009 IEEE 12th International Conference on. Comput Vision, 2009 … 2146–2153
  106. 106.Ji S, Yang M, Yu K, Xu W (2010) 3D convolutional neural networks for human action recognition. ICML, Int Conf Mach Learn 35:221–31. doi: 10.1109/TPAMI.2012.59
  107. 107.Joachims T (1998) Text categorization with support vector machines: Learning with many relevant features. In: European conference on machine learning. pp 137–142
  108. 108.Justus D, Brennan J, Bonner S, McGough AS (2019) Predicting the Computational Cost of Deep Learning Models. In: Proceedings - 2018 IEEE International Conference on Big Data, Big Data 2018
  109. 109.Kafi M, Maleki M, Davoodian N (2015) Functional histology of the ovarian follicles as determined by follicular fluid concentrations of steroids and IGF-1 in Camelus dromedarius. Res Vet Sci 99:37–40. doi: 10.1016/j.rvsc.2015.01.001
  110. 110.Kahng M, Thorat N, Chau DHP, et al (2019) GAN Lab: Understanding Complex Deep Generative Models using Interactive Visual Experimentation. IEEE Trans Vis Comput Graph 25:310–320
  111. 111.Kalchbrenner N, Grefenstette E, Blunsom P (2014) A convolutional neural network for modelling sentences. arXiv Prepr arXiv14042188
  112. 112.Kawaguchi K, Huang J, Kaelbling LP (2019) Effect of depth and width on local minima in deep learning. Neural Comput 31:1462–1498. doi: 10.1162/neco_a_01195
  113. 113.Kawashima T, Kawanishi Y, Ide I, et al (2017) Action recognition from extremely low-resolution thermal image sequence. In: 2017 14th IEEE International Conference on Advanced Video and Signal Based Surveillance, AVSS 2017. IEEE, pp 1–6
  114. 114.Khan A, Qureshi AS, Hussain M, et al (2019) A Recent Survey on the Applications of Genetic Programming in Image Processing. arXiv Prepr arXiv190107387
  115. 115.Khan A, Sohail A, Ali A (2018a) A New Channel Boosted Convolutional Neural Network using Transfer Learning. arXiv Prepr arXiv180408528
  116. 116.Khan A, Zameer A, Jamal T, Raza A (2018b) Deep Belief Networks Based Feature Generation and Regression for Predicting Wind Power. arXiv Prepr arXiv180711682
  117. 117.Krizhevsky A, Sutskever I, Hinton GE (2012) ImageNet Classification with Deep Convolutional Neural Networks. Adv Neural Inf Process Syst 1–9. doi: 10.1061/(ASCE)GT.1943-5606.0001284
  118. 118.Kuen J, Kong X, Wang G, et al (2017) DelugeNets: Deep Networks with Efficient and Flexible Cross-layer Information Inflows. In: Computer Vision Workshop (ICCVW), 2017 IEEE International Conference on. pp 958–966
  119. 119.Kuen J, Kong X, Wang G, Tan YP (2018) DelugeNets: Deep Networks with Efficient and Flexible Cross-Layer Information Inflows. Proc - 2017 IEEE Int Conf Comput Vis Work ICCVW 2017 2018-Janua:958–966. doi: 10.1109/ICCVW.2017.117
  120. 120.Lacey G, Taylor GW, Areibi S (2016) Deep Learning on FPGAs: Past, Present, and Future. arXiv Prepr arXiv160204283
  121. 121.Larsson G, Maire M, Shakhnarovich G (2016) Fractalnet: Ultra-deep neural networks without residuals. arXiv Prepr arXiv160507648 1–11
  122. 122.Laskar MNU, Giraldo LGS, Schwartz O (2018) Correspondence of Deep Neural Networks and the Brain for Visual Textures. 1–17
  123. 123.Le Q V., Ranzato M, Monga R, et al (2011) Building high-level features using large scale unsupervised learning. ICASSP, IEEE Int Conf Acoust Speech Signal Process - Proc 8595–8598. doi: 10.1109/ICASSP.2013.6639343
  124. 124.LeCun Y (2007) Effcient BackPrp. J Exp Psychol Gen 136:23–42
  125. 125.LeCun Y, Bengio Y, Hinton G (2015) Deep learning. Nature 521:436–444. doi: 10.1038/nature14539
  126. 126.LeCun Y, Boser B, Denker JS, et al (1989) Backpropagation applied to handwritten zip code recognition. Neural Comput 1:541–551
  127. 127.LeCun Y, Bottou L, Bengio Y, Haffner P (1998) Gradient-based learning applied to document recognition. Proc IEEE 86:2278–2324
  128. 128.LeCun Y, Jackel LD, Bottou L, et al (1995) Learning algorithms for classification: A comparison on handwritten digit recognition. Neural networks Stat Mech Perspect 261:276
  129. 129.LeCun Y, Kavukcuoglu K, Farabet CC, others (2010) Convolutional networks and applications in vision. In: ISCAS. IEEE, pp 253–256
  130. 130.Lee C-Y, Gallagher PW, Tu Z (2016) Generalizing pooling functions in convolutional neural networks: Mixed, gated, and tree. In: Artificial Intelligence and Statistics. pp 464–472
  131. 131.Lee S, Son K, Kim H, Park J (2017) Car Plate Recognition Based on CNN Using Embedded System with GPU. 239–241
  132. 132.Levi G, Hassner T (2009) Sicherheit und Medien. Sicherheit und Medien. doi: 10.1109/CVPRW.2015.7301352
  133. 133.Li H, Lin Z, Shen X, et al (2015) A convolutional neural network cascade for face detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp 5325–5334
  134. 134.Li S, Liu Z-Q, Chan AB (2014) Heterogeneous Multi-task Learning for Human Pose Estimation with Deep Convolutional Neural Network. In: 2014 IEEE Conference on Computer Vision and Pattern Recognition Workshops. IEEE, pp 488–495
  135. 135.Li X, Bing L, Lam W, Shi B (2018) Transformation Networks for Target-Oriented Sentiment Classification. 946–956
  136. 136.Lin M, Chen Q, Yan S (2013) Network In Network. 1–10. doi: 10.1109/ASRU.2015.7404828
  137. 137.Lin T-Y, Maire M, Belongie S, et al (2014) Microsoft coco: Common objects in context. In: European conference on computer vision. Springer, pp 740–755
  138. 138.Lin TY, Dollár P, Girshick R, et al (2017) Feature pyramid networks for object detection. In: Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017
  139. 139.Lindholm E, Nickolls J, Oberman S, Montrym J (2008) NVIDIA Tesla: A Unified Graphics and Computing Architecture. IEEE Micro 28:39–55. doi: 10.1109/MM.2008.31
  140. 140.Linnainmaa S (1970) The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors. Master’s Thesis (in Finnish), Univ Helsinki 6–7
  141. 141.Liu C-L, Nakashima K, Sako H, Fujisawa H (2003) Handwritten digit recognition: benchmarking of state-of-the-art techniques. Pattern Recognit 36:2271–2285
  142. 142.Liu W, Wang Z, Liu X, et al (2017) A survey of deep neural network architectures and their applications. Neurocomputing 234:11–26. doi: 10.1016/j.neucom.2016.12.038
  143. 143.Liu X, Deng Z, Yang Y (2019) Recent progress in semantic image segmentation. Artif Intell Rev 52:1089–1106. doi: 10.1007/s10462-018-9641-3
  144. 144.Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic segmentation. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 3431–3440
  145. 145.Long ZM, Guo SQ, Chen GJ, Yin BL (2012) Modeling and simulation for the articulated robotic arm test system of the combination drive. 2011 Int Conf Mechatronics Mater Eng ICMME 2011 151:480–483. doi: 10.4028/www.scientific.net/AMM.151.480
  146. 146.Lowe DG (1999) Object recognition from local scale-invariant features. Proc Seventh IEEE Int Conf Comput Vis 1150–1157 vol.2. doi: 10.1109/ICCV.1999.790410
  147. 147.Lowe DG (2004) Distinctive image features from scale-invariant keypoints. Int J Comput Vis 60:91–110
  148. 148.Lu H, Li B, Zhu J, et al (2017a) Wound intensity correction and segmentation with convolutional neural networks. Concurr Comput Pract Exp 29:e3927
  149. 149.Lu Z, Pu H, Wang F, et al (2017b) The expressive power of neural networks: A view from the width. In: Advances in Neural Information Processing Systems. pp 6231–6239
  150. 150.Lv E, Wang X, Cheng Y, Yu Q (2019) Deep ensemble network based on multi-path fusion. Artif Intell Rev 52:151–168. doi: 10.1007/s10462-019-09708-5
  151. 151.Madrazo CF, Heredia I, Lloret L, Marco de Lucas J (2019) Application of a Convolutional Neural Network for image classification for the analysis of collisions in High Energy Physics. EPJ Web Conf. doi: 10.1051/epjconf/201921406017
  152. 152.Mao X, Shen C, Yang Y-B (2016) Image restoration using very deep convolutional encoder-decoder networks with symmetric skip connections. In: Advances in neural information processing systems. pp 2802–2810
  153. 153.Marmanis D, Wegner JD, Galliani S, et al (2016) Semantic segmentation of aerial images with an ensemble of CNNs. ISPRS Ann Photogramm Remote Sens Spat Inf Sci 3:473
  154. 154.Matsugu M, Mori K, Ishii M, Mitarai Y (2002) Convolutional spiking neural network model for robust face detection. In: Neural Information Processing, 2002. ICONIP’02. Proceedings of the 9th International Conference on. pp 660–664
  155. 155.Mikolov T, Karafiát M, Burget L, et al (2010) Recurrent neural network based language model. In: Eleventh Annual Conference of the International Speech Communication Association
  156. 156.Misra D (2019) Mish: A Self Regularized Non-Monotonic Neural Activation Function. arXiv Prepr ArXiv 190808681
  157. 157.Mohamed A, Dahl GE, Hinton G (2012) Acoustic modeling using deep belief networks. IEEE Trans Audio, Speech Lang Process 20:14–22
  158. 158.Montufar GF, Pascanu R, Cho K, Bengio Y (2014) On the number of linear regions of deep neural networks. In: Advances in neural information processing systems. pp 2924–2932
  159. 159.Moons B, Verhelst M (2017) An energy-efficient precision-scalable ConvNet processor in 40-nm CMOS. IEEE J Solid-State Circuits 52:903–914
  160. 160.Morar A, Moldoveanu F, Gröller E (2012) Image segmentation based on active contours without edges. Proc - 2012 IEEE 8th Int Conf Intell Comput Commun Process ICCP 2012 213–220. doi: 10.1109/ICCP.2012.6356188
  161. 161.Nair V, Hinton GE (2010) Rectified linear units improve Restricted Boltzmann machines. In: ICML 2010 - Proceedings, 27th International Conference on Machine Learning
  162. 162.Najafabadi MM, Villanustre F, Khoshgoftaar TM, et al (2015) Deep learning applications and challenges in big data analytics. J Big Data 2:1–21. doi: 10.1186/s40537-014-0007-7
  163. 163.Nguyen G, Dlugolinsky S, Bobák M, et al (2019) Machine Learning and Deep Learning frameworks and libraries for large-scale data mining: a survey. Artif Intell Rev 52:77–124. doi: 10.1007/s10462-018-09679-z
  164. 164.Nguyen Q, Mukkamala M, Hein M (2018) Neural Networks Should Be Wide Enough to Learn Disconnected Decision Regions. arXiv Prepr arXiv180300094
  165. 165.Nickolls J, Buck I, Garland M, Skadron K (2008) Scalable parallel programming with CUDA. In: ACM SIGGRAPH 2008 classes on - SIGGRAPH ’08. ACM Press, New York, New York, USA, p 1
  166. 166.Nwankpa C, Ijomah W, Gachagan A, Marshall S (2018) Activation Functions: Comparison of trends in Practice and Research for Deep Learning. arXiv Prepr arXiv181103378
  167. 167.Oh K-S, Jung K (2004) GPU implementation of neural networks. Pattern Recognit 37:1311–1314
  168. 168.Ojala T, PeitiKainen M, Maenpã T (2002) Multiresolution gray-scale and rotation invariant texture classification with local binary patterns. IEEE Trans Pattern Anal ans Mach Intell
  169. 169.Ojala T, Pietikäinen M, Harwood D (1996) A comparative study of texture measures with classification based on feature distributions. Pattern Recognit 29:51–59. doi: 10.1016/0031-3203(95)00067-4
  170. 170.Oquab M, Bottou L, Laptev I, Sivic J (2014) Learning and transferring mid-level image representations using convolutional neural networks. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. IEEE, pp 1717–1724
  171. 171.Pang J, Chen K, Shi J, et al (2020) Libra R-CNN: Towards Balanced Learning for Object Detection
  172. 172.Pascanu R, Mikolov T, Bengio Y (2012) Understanding the exploding gradient problem. CoRR, abs/12115063
  173. 173.Peng X, Hoffman J, Yu SX, Saenko K (2016) Fine-to-coarse knowledge transfer for low-res image classification. In: 2016 IEEE International Conference on Image Processing (ICIP). IEEE, pp 3683–3687
  174. 174.Potluri S, Fasih A, Vutukuru LK, et al (2011) CNN based high performance computing for real time image processing on GPU. In: Proceedings of the Joint INDS’11 & ISTET’11. pp 1–7
  175. 175.Qiang Yang, Pan SJ, Yang Q, Fellow QY (2008) A Survey on Transfer Learning. IEEE Trans Knowl Data Eng 1:1–15. doi: 10.1109/TKDE.2009.191
  176. 176.Qureshi AS, Khan A (2018) Adaptive Transfer Learning in Deep Neural Networks: Wind Power Prediction using Knowledge Transfer from Region to Region and Between Different Task Domains. arXiv Prepr arXiv181012611
  177. 177.Qureshi AS, Khan A, Zameer A, Usman A (2017) Wind power prediction using deep neural network based meta regression and transfer learning. Appl Soft Comput J 58:742–755. doi: 10.1016/j.asoc.2017.05.031
  178. 178.Ramachandran P, Zoph B, Le Q V. (2017) Swish: a Self-Gated Activation Function. arXiv
  179. 179.Ranjan R, Patel VM, Chellappa R (2015) A deep pyramid deformable part model for face detection. arXiv Prepr arXiv150804389
  180. 180.Ranzato M, Huang FJ, Boureau YL, LeCun Y (2007) Unsupervised learning of invariant feature hierarchies with applications to object recognition. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. IEEE, pp 1–8
  181. 181.Rawat W, Wang Z (2016) Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review. 61:1120–1132. doi: 10.1162/NECO
  182. 182.Ren S, He K, Girshick R, Sun J (2015) Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. 1–9. doi: 10.1109/TPAMI.2016.2577031
  183. 183.Ronneberger O, Fischer P, Brox T (2015) U-net: Convolutional networks for biomedical image segmentation. In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
  184. 184.Roy AG, Navab N, Wachinger C (2018) Concurrent spatial and channel ‘squeeze & excitation’ in fully convolutional networks. Lect Notes Comput Sci (including Subser Lect Notes Artif Intell Lect Notes Bioinformatics) 11070 LNCS:421–429. doi: 10.1007/978-3-030-00928-1_48
  185. 185.Russakovsky O, Deng J, Su H, et al (2015) ImageNet Large Scale Visual Recognition Challenge. Int J Comput Vis. doi: 10.1007/s11263-015-0816-y
  186. 186.Salakhutdinov R, Larochelle H (2010) Efficient learning of deep Boltzmann machines. In: Proceedings of the thirteenth international conference on artificial intelligence and statistics. pp 693–700
  187. 187.Scherer D, Müller A, Behnke S (2010) Evaluation of pooling operations in convolutional architectures for object recognition. In: Artificial Neural Networks--ICANN 2010. Springer, pp 92–101
  188. 188.Schmidhuber J (2007) New millennium AI and the convergence of history. In: Challenges for computational intelligence. Springer, pp 15–35
  189. 189.Sermanet P, Chintala S, Lecun Y (2012) Convolutional neural networks applied to house numbers digit classification. In: Proceedings of the 21st International Conference on Pattern Recognition (ICPR2012), Tsukuba. IEEE, pp 3288–3291
  190. 190.Shakeel MF, Bajwa NA, Anwaar AM, et al (2019) Detecting Driver Drowsiness in Real Time Through Deep Learning Based Object Detection. In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
  191. 191.Sharma A, Muttoo SK (2018) Spatial Image Steganalysis Based on ResNeXt. 2018 IEEE 18th Int Conf Commun Technol 1213–1216. doi: 10.1109/ICCT.2018.8600132
  192. 192.Shi Y, Tian Y, Wang Y, Huang T (2017) Sequential deep trajectory descriptor for action recognition with three-stream CNN. IEEE Trans Multimed 19:1510–1520
  193. 193.Shin H-CC, Roth HR, Gao M, et al (2016) Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning. IEEE Trans Med Imaging 35:1285–1298. doi: 10.1109/TMI.2016.2528162
  194. 194.Simard PY, Steinkraus D, Platt JC (2003) Best practices for convolutional neural networks applied to visual document analysis. In: null. p 958
  195. 195.Simonyan K, Vedaldi A, Zisserman A (2013) Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps. 1–8. doi: 10.1080/00994480.2000.10748487
  196. 196.Simonyan K, Zisserman A (2015) VERY DEEP CONVOLUTIONAL NETWORKS FOR LARGE-SCALE IMAGE RECOGNITION. ICLR 75:398–406. doi: 10.2146/ajhp170251
  197. 197.Simonyan K, Zisserman A (2014) Two-stream convolutional networks for action recognition in videos. In: Advances in neural information processing systems. pp 568–576
  198. 198.Sinha T, Verma B, Haidar A (2018) Optimization of convolutional neural network parameters for image classification. 2017 IEEE Symp Ser Comput Intell SSCI 2017 - Proc 2018-Janua:1–7. doi: 10.1109/SSCI.2017.8285338
  199. 199.Spanhol FA, Oliveira LS, Petitjean C, Heutte L (2016a) A dataset for breast cancer histopathological image classification. IEEE Trans Biomed Eng 63:1455–1462
  200. 200.Spanhol FA, Oliveira LS, Petitjean C, Heutte L (2016b) Breast cancer histopathological image classification using Convolutional Neural Networks. In: 2016 International Joint Conference on Neural Networks (IJCNN). IEEE, pp 2560–2567
  201. 201.Srinivas S, Sarvadevabhatla RK, Mopuri KR, et al (2016) A Taxonomy of Deep Convolutional Neural Nets for Computer Vision. Front Robot AI 2:1–13. doi: 10.3389/frobt.2015.00036
  202. 202.Srivastava N, Hinton G, Krizhevsky A, et al (2014) Dropout: A Simple Way to Prevent Neural Networks from Overfittin. J Mach Learn Res 1:11. doi: 10.1016/j.micromeso.2003.09.025
  203. 203.Srivastava RK, Greff K, Schmidhuber J (2015a) Highway Networks. doi: 10.1002/esp.3417
  204. 204.Srivastava RK, Greff K, Schmidhuber J (2015b) Training very deep networks. In: Advances in Neural Information Processing Systems
  205. 205.Stefanini M, Lancellotti R, Baraldi L, Calderara S (2019) A Deep-learning-based approach to VM behavior Identification in Cloud Systems. In: Proceedings of the 9th International Conference on Cloud Computing and Services Science. SCITEPRESS - Science and Technology Publications, pp 308–315
  206. 206.Strigl D, Kofler K, Podlipnig S (2010) Performance and scalability of GPU-based convolutional neural networks. In: Parallel, Distributed and Network-Based Processing (PDP), 2010 18th Euromicro International Conference on. pp 317–324
  207. 207.Suganuma M, Shirakawa S, Nagao T (2017) A genetic programming approach to designing convolutional neural network architectures. In: Proceedings of the Genetic and Evolutionary Computation Conference. ACM, pp 497–504
  208. 208.Sun L, Jia K, Yeung D-Y, Shi BE (2015) Human action recognition using factorized spatio-temporal convolutional networks. In: Proceedings of the IEEE International Conference on Computer Vision. pp 4597–4605
  209. 209.Sundermeyer M, Schlüter R, Ney H (2012) LSTM neural networks for language modeling. In: Thirteenth annual conference of the international speech communication association
  210. 210.Sze V, Chen YH, Yang TJ, Emer JS (2017) Efficient Processing of Deep Neural Networks: A Tutorial and Survey. Proc. IEEE
  211. 211.Szegedy C, Ioffe S, Vanhoucke V (2016a) Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning. arXiv Prepr arXiv160207261v2 131:262–263. doi: 10.1007/s10236-015-0809-y
  212. 212.Szegedy C, Vanhoucke V, Ioffe S, et al (2016b) Rethinking the Inception Architecture for Computer Vision. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. IEEE, pp 2818–2826
  213. 213.Szegedy C, Wei Liu, Yangqing Jia, et al (2015) Going deeper with convolutions. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 1–9
  214. 214.Szegedy C, Zaremba W, Sutskever I, et al (2014) Intriguing properties of neural networks. In: 2nd International Conference on Learning Representations, ICLR 2014 - Conference Track Proceedings
  215. 215.Targ S, Almeida D, Lyman K (2016) Resnet in Resnet: generalizing residual architectures. arXiv Prepr arXiv160308029
  216. 216.Tong T, Li G, Liu X, Gao Q (2017) Image super-resolution using dense skip connections. In: 2017 IEEE International Conference on Computer Vision (ICCV). pp 4809–4817
  217. 217.Tong W, Song L, Yang X, et al (2015) CNN-based shot boundary detection and video annotation. In: 2015 IEEE international symposium on broadband multimedia systems and broadcasting. IEEE, pp 1–5
  218. 218.Tran D, Bourdev L, Fergus R, et al (2015) Learning spatiotemporal features with 3d convolutional networks. In: Proceedings of the IEEE international conference on computer vision. pp 4489–4497
  219. 219.Ullah A, Ahmad J, Muhammad K, et al (2017) Action recognition in video sequences using deep bi-directional LSTM with CNN features. IEEE Access 6:1155–1166
  220. 220.Vinayakumar R, Soman KP, Poornachandrany P (2017) Applying convolutional neural network for network intrusion detection. In: 2017 International Conference on Advances in Computing, Communications and Informatics, ICACCI 2017
  221. 221.Vincent P, Larochelle H, Bengio Y, Manzagol P-A (2008) Extracting and composing robust features with denoising autoencoders. In: Proceedings of the 25th international conference on Machine learning. ACM, pp 1096–1103
  222. 222.Vinyals O, Toshev A, Bengio S, Erhan D (2017) Show and Tell: Lessons Learned from the 2015 MSCOCO Image Captioning Challenge. IEEE Trans Pattern Anal Mach Intell. doi: 10.1109/TPAMI.2016.2587640
  223. 223.Wahab N, Khan A, Lee YS (2019) Transfer learning based deep CNN for segmentation and detection of mitoses in breast cancer histopathological images. Microscopy 68:216–233. doi: 10.1093/jmicro/dfz002
  224. 224.Wahab N, Khan A, Lee YS (2017) Two-phase deep convolutional neural network for reducing class skewness in histopathological images based breast cancer detection. Comput Biol Med 85:86–97. doi: 10.1016/j.compbiomed.2017.04.012
  225. 225.Wang F, Jiang M, Qian C, et al (2017a) Residual Attention Network for Image Classification. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 6450–6458
  226. 226.Wang H, Raj B (2017) On the Origin of Deep Learning. 1–72. doi: 10.1016/0014-5793(91)81229-2
  227. 227.Wang H, Schmid C (2013) Action recognition with improved trajectories. In: Proceedings of the IEEE international conference on computer vision. pp 3551–3558
  228. 228.Wang T, Wu DJDJ, Coates A, Ng AY (2012) End-to-end text recognition with convolutional neural networks. ICPR, Int Conf Pattern Recognit 3304–3308
  229. 229.Wang X, Gao L, Song J, Shen H (2017b) Beyond Frame-level CNN: Saliency-Aware 3-D CNN With LSTM for Video Action Recognition. IEEE Signal Process Lett 24:510–514. doi: 10.1109/LSP.2016.2611485
  230. 230.Wang Y, Wang L, Wang H, Li P (2019) End-to-End Image Super-Resolution via Deep and Shallow Convolutional Networks. IEEE Access 7:31959–31970. doi: 10.1109/ACCESS.2019.2903582
  231. 231.Woo S, Park J, Lee JY, Kweon IS (2018) CBAM: Convolutional block attention module. Lect Notes Comput Sci (including Subser Lect Notes Artif Intell Lect Notes Bioinformatics) 11211 LNCS:3–19. doi: 10.1007/978-3-030-01234-2_1
  232. 232.Wu J, Leng C, Wang Y, et al (2016) Quantized convolutional neural networks for mobile devices. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
  233. 233.Xie S, Girshick R, Dollar P, et al (2017) Aggregated Residual Transformations for Deep Neural Networks. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 5987–5995
  234. 234.Xie W, Zhang C, Zhang Y, et al (2018) An Energy-Efficient FPGA-Based Embedded System for CNN Application. In: 2018 IEEE International Conference on Electron Devices and Solid State Circuits (EDSSC). IEEE, pp 1–2
  235. 235.Xiong Y, Kim HJ, Hedau V (2019) ANTNets: Mobile Convolutional Neural Networks for Resource Efficient Image Classification. arXiv Prepr ArXiv 190403775
  236. 236.Xu B, Wang N, Chen T, Li M (2015a) Empirical Evaluation of Rectified Activations in Convolutional Network. J Foot Ankle Res 1:O22. doi: 10.1186/1757-1146-1-S1-O22
  237. 237.Xu K, Ba J, Kiros R, et al (2015b) Show, attend and tell: Neural image caption generation with visual attention. In: International conference on machine learning. pp 2048–2057
  238. 238.Yamada Y, Iwamura M, Kise K (2016) Deep pyramidal residual networks with separated stochastic depth. arXiv Prepr arXiv161201230
  239. 239.Yang J, Xiong W, Li S, Xu C (2019) Learning structured and non-redundant representations with deep neural networks. Pattern Recognit 86:224–235
  240. 240.Yang S, Luo P, Loy C-C, Tang X (2015) From facial parts responses to face detection: A deep learning approach. In: Proceedings of the IEEE International Conference on Computer Vision. pp 3676–3684
  241. 241.Yıldırım Ö, Pławiak P, Tan RS, Acharya UR (2018) Arrhythmia detection using deep convolutional neural network with long duration ECG signals. Comput Biol Med. doi: 10.1016/j.compbiomed.2018.09.009
  242. 242.Young SR, Rose DC, Karnowski TP, et al (2015) Optimizing deep learning hyper-parameters through an evolutionary algorithm. In: Proceedings of the Workshop on Machine Learning in High-Performance Computing Environments. ACM, p 4
  243. 243.Zagoruyko S, Komodakis N (2016) Wide Residual Networks. Procedings Br Mach Vis Conf 2016 87.1-87.12. doi: 10.5244/C.30.87
  244. 244.Zeiler MD, Fergus R (2013) Visualizing and Understanding Convolutional Networks. arXiv Prepr arXiv13112901v3 30:225–231. doi: 10.1111/j.1475-4932.1954.tb03086.x
  245. 245.Zhang K, Zhang Z, Li Z, et al (2016) Joint face detection and alignment using multitask cascaded convolutional networks. IeeexploreIeeeOrg 23:1499–1503
  246. 246.Zhang Q, Zhang M, Chen T, et al (2019) Recent advances in convolutional neural network acceleration. Neurocomputing 323:37–51. doi: 10.1016/j.neucom.2018.09.038
  247. 247.Zhang X, LeCun Y (2015) Text understanding from scratch. arXiv Prepr arXiv150201710
  248. 248.Zhang X, Li Z, Loy CC, Lin D (2017) PolyNet: A Pursuit of Structural Diversity in Very Deep Networks. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 3900–3908
  249. 249.Zhang X, Zhou X, Lin M, Sun J (2018a) ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
  250. 250.Zhang Y, Qiu Z, Yao T, et al (2018b) Fully Convolutional Adaptation Networks for Semantic Segmentation. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
  251. 251.Zheng H, Fu J, Mei T, Luo J (2017) Learning multi-attention convolutional neural network for fine-grained image recognition. In: 2017 IEEE International Conference on Computer Vision (ICCV). pp 5219–5227
  252. 252.Zhou B, Khosla A, Lapedriza A, et al (2016) Learning deep features for discriminative localization. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp 2921–2929

Citation

MLA
Khan, A., et al. “A Survey of the Recent Architectures of Deep Convolutional Neural Networks”. Artificial Intelligence Review, vol. 53, no. 8, 2020, pp. 5455–516, https://doi.org/10.1007/s10462-020-09825-6.
APA
Khan, A., Sohail, A., Zahoora, U., & Qureshi, A. S. (2020). A survey of the recent architectures of deep convolutional neural networks. Artificial Intelligence Review, 53(8), 5455–5516. https://doi.org/10.1007/s10462-020-09825-6
Chicago
Khan, A., A. Sohail, U. Zahoora, and A. S. Qureshi. 2020. “A Survey of the Recent Architectures of Deep Convolutional Neural Networks”. Artificial Intelligence Review 53 (8): 5455–5516. https://doi.org/10.1007/s10462-020-09825-6.
Harvard
Khan, A. et al. (2020) “A survey of the recent architectures of deep convolutional neural networks”, Artificial Intelligence Review, 53(8), pp. 5455–5516. Available at: https://doi.org/10.1007/s10462-020-09825-6.
Vancouver
1. Khan A, Sohail A, Zahoora U, Qureshi AS (2020) A survey of the recent architectures of deep convolutional neural networks. Artificial Intelligence Review 53:5455–5516

BibTeX

@article{Khan_2020, title={A survey of the recent architectures of deep convolutional neural networks}, volume={53}, ISSN={1573-7462}, url={http://dx.doi.org/10.1007/s10462-020-09825-6}, DOI={10.1007/s10462-020-09825-6}, number={8}, journal={Artificial Intelligence Review}, publisher={Springer Science and Business Media LLC}, author={Khan, Asifullah and Sohail, Anabia and Zahoora, Umme and Qureshi, Aqsa Saeed}, year={2020}, month=Apr, pages={5455–5516} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF