A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects

Zewen LiFan LiuWenjie YangShou-Heng PengJun Zhou

article2020IEEE TNNLS4,453 citations

Synthesizes structural advancements across one-dimensional, two-dimensional, and multi-dimensional convolutional neural networks, offering experimentally grounded rules of thumb for model design and highlighting critical directions for future research.

Listen

Convolutional neural networks have driven major advances in deep learning for computer vision, speech, and related tasks, yet prior surveys have focused narrowly on applications or older architectures and have overlooked many recent innovations in one-dimensional and multi-dimensional convolutions. This survey addresses the need for a unified, up-to-date treatment by tracing the evolution of CNNs from early perceptrons through modern lightweight and generative models, while also examining practical design choices and emerging challenges.

The authors set out to deliver a general overview of CNN building blocks, representative models, activation and loss functions, optimizers, applications across dimensions, and forward-looking research directions. They achieve this through a structured literature synthesis of models from LeNet-5 to GhostNet, targeted experiments on benchmark datasets, and analysis of open problems.

Experiments compared seven activation functions on LeNet-5 and pre-trained VGG-16 across MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100; ten optimizers on CIFAR-10; and reviewed loss functions for regression and classification. Key results show that ReLU and Leaky ReLU deliver the best balance of accuracy, stability, and training speed for hidden layers, while sigmoid and tanh frequently slow convergence or fail to train deeper models. Cross-entropy remains the default for classification, yet center loss and triplet loss improve intra-class compactness when needed. Deeper and wider networks consistently raise accuracy, residual connections prevent degradation beyond roughly 50 layers, and depth-wise separable convolutions cut parameters dramatically for mobile use. One-dimensional CNNs excel at fixed-length signal tasks such as ECG analysis and traffic prediction, two-dimensional networks dominate image classification and detection, and three-dimensional networks capture spatio-temporal patterns in video and volumetric data.

These findings indicate that practitioners can safely adopt ReLU-family activations and mini-batch gradient methods with modest learning-rate tuning, while model compression and neural architecture search offer practical routes to deployment on resource-limited hardware. At the same time, the survey underscores vulnerabilities to adversarial and backdoor attacks, the loss of spatial relationships inherent in pooling, and the continuing difficulty of selecting architectures without exhaustive search.

The authors therefore recommend systematic use of the reported rules of thumb for function selection, wider adoption of model-compression techniques, and exploration of capsule networks and agentless architecture search to address current limitations. Experiments were conducted on only four image-classification datasets and two base architectures, so results may not generalize to every domain or scale; the field continues to evolve rapidly, and some cited models will soon be superseded. Overall, the synthesis and experimental comparisons provide a reliable foundation for design decisions while clearly delineating the next technical hurdles.

  • Paper: A Survey on Vision Transformer, Kai Han et al. (2020). This survey naturally extends the source by tracking the subsequent paradigm shift from convolutional neural networks to transformer-based vision architectures.
  • Paper: On the Relationship between Self-Attention and Convolutional Layers, Jean-Baptiste Cordonnier et al. (2020). This study builds directly upon standard CNN foundations by theoretically and empirically comparing convolutional layers with emerging self-attention mechanisms.
Cover for A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects

Abstract

Convolutional Neural Network (CNN) is one of the most significant networks in the deep learning field. Since CNN made impressive achievements in many areas, including but not limited to computer vision and natural language processing, it attracted much attention both of industry and academia in the past few years. The existing reviews mainly focus on the applications of CNN in different scenarios without considering CNN from a general perspective, and some novel ideas proposed recently are not covered. In this review, we aim to provide novel ideas and prospects in this fast-growing field as much as possible. Besides, not only two-dimensional convolution but also one-dimensional and multi-dimensional ones are involved. First, this review starts with a brief introduction to the history of CNN. Second, we provide an overview of CNN. Third, classic and advanced CNN models are introduced, especially those key points making them reach state-of-the-art results. Fourth, through experimental analysis, we draw some conclusions and provide several rules of thumb for function selection. Fifth, the applications of one-dimensional, two-dimensional, and multi-dimensional convolution are covered. Finally, some open issues and promising directions for CNN are discussed to serve as guidelines for future work.

Table of Contents

  • A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects
  • I. INTRODUCTION
  • II. BRIEF OVERVIEW OF CNN
  • III. CLASSIC CNN MODELS
  • A. LeNet-5
  • B. AlexNet
  • C. VGGNets
  • D. GoogLeNet
  • 1) Inception v1
  • 2) Inception v2
  • 3) Inception v3
  • 4) Inception v4 and Inception-ResNet
  • E. ResNet
  • F. DCGAN
  • G. MobileNets
  • 1) MobileNet v1
  • 2) MobileNet v2
  • 3) MobileNet v3
  • H. ShuffleNets
  • 1) ShuffleNet v1
  • 2) ShuffleNet v2
  • I. GhostNet
  • IV. DISCUSSION AND EXPERIMENTAL ANALYSIS
  • A. Activation function
  • 1) Discussion of Activation Function
  • 2) Experimental Evaluation
  • B. Loss function
  • 1) Loss Function for Regression
  • 2) Loss Function for Classification
  • 3) Rules of Thumb for Selection
  • C. Optimizer
  • 1) Gradient Descent
  • 2) Gradient Descent Optimization Algorithms
  • 3) Experimental Evaluation
  • 4) Rules of Thumb for Selection
  • V. APPLICATIONS OF CNN
  • A. Applications of one-dimensional CNN
  • 1) Time Series Prediction
  • 2) Signal Identification
  • B. Applications of two-dimensional CNN
  • 1) Image Classification
  • 2) Object Detection
  • 3) Image Segmentation
  • 4) Face Recognition
  • C. Applications of multi-dimensional CNN
  • 1) Human Action Recognition
  • 2) Object Recognition/Detection
  • VI. PROSPECTS FOR CNN
  • A. Model Compression
  • B. Security of CNN
  • C. Network Architecture Search
  • D. Capsule Neural Network
  • VII. CONCLUSION
  • REFERENCES

Knowls

  1. Knowl 1 — Benchmark Comparison of Activation Functions across CNN Models and Datasets

    data/table

    An empirical benchmark assesses the performance and execution time of seven activation functions (Linear, Sigmoid, Tanh, ReLU, Leaky ReLU, PReLU, and ELU) across two convolutional architectures: LeNet-5 (trained from scratch) and VGG-16 (pre-trained on ImageNet, excluding the top three fully connected layers). All evaluations use cross-entropy loss and the Adam optimizer on an Intel Xeon E5-2640 v4 CPU with an NVIDIA TITAN RTX (24GB) GPU.

    Model Dataset Activation Function Batch Size Epochs Validation Accuracy (%) Training Time (s)
    LeNet-5 MNIST Linear 256 50 98.56 92.89
    LeNet-5 MNIST Sigmoid 256 50 98.94 95.51
    LeNet-5 MNIST Tanh 256 50 99.03 92.92
    LeNet-5 MNIST ReLU 256 50 99.18 95.04
    LeNet-5 MNIST Leaky ReLU 256 50 99.10 99.82
    LeNet-5 MNIST PReLU 256 50 99.20 113.42
    LeNet-5 MNIST ELU 256 50 99.20 103.84
    LeNet-5 Fashion-MNIST Linear 256 100 88.10 169.95
    LeNet-5 Fashion-MNIST Sigmoid 256 100 89.84 174.26
    LeNet-5 Fashion-MNIST Tanh 256 100 89.83 181.99
    LeNet-5 Fashion-MNIST ReLU 256 100 90.17 191.77
    LeNet-5 Fashion-MNIST Leaky ReLU 256 100 90.36 190.02
    LeNet-5 Fashion-MNIST PReLU 256 100 90.36 217.20
    LeNet-5 Fashion-MNIST ELU 256 100 90.37 204.64
    LeNet-5 CIFAR-10 Linear 256 200 62.56 614.89
    LeNet-5 CIFAR-10 Sigmoid 256 200 62.65 569.25
    LeNet-5 CIFAR-10 Tanh 256 200 62.94 575.07
    LeNet-5 CIFAR-10 ReLU 256 200 64.40 550.35
    LeNet-5 CIFAR-10 Leaky ReLU 256 200 64.27 582.08
    LeNet-5 CIFAR-10 PReLU 256 200 63.61 650.75
    LeNet-5 CIFAR-10 ELU 256 200 65.70 626.51
    LeNet-5 CIFAR-100 Linear 512 1000 31.24 2381.76
    LeNet-5 CIFAR-100 Sigmoid 512 1000 32.32 2355.64
    LeNet-5 CIFAR-100 Tanh 512 1000 32.69 2376.35
    LeNet-5 CIFAR-100 ReLU 512 1000 32.69 2418.18
    LeNet-5 CIFAR-100 Leaky ReLU 512 1000 33.81 2443.72
    LeNet-5 CIFAR-100 PReLU 512 1000 31.84 2615.02
    LeNet-5 CIFAR-100 ELU 512 1000 35.10 2475.30
    VGG-16 MNIST Linear 512 30 9.82 598.14
    VGG-16 MNIST Sigmoid 512 30 11.35 600.27
    VGG-16 MNIST Tanh 512 30 11.35 596.32
    VGG-16 MNIST ReLU 512 30 99.55 606.83
    VGG-16 MNIST Leaky ReLU 512 30 99.48 608.91
    VGG-16 MNIST PReLU 512 30 99.45 607.27
    VGG-16 MNIST ELU 512 30 11.35 614.81
    VGG-16 Fashion-MNIST Linear 512 30 10.00 599.18
    VGG-16 Fashion-MNIST Sigmoid 512 30 15.66 595.11
    VGG-16 Fashion-MNIST Tanh 512 30 11.01 596.48
    VGG-16 Fashion-MNIST ReLU 512 30 93.16 608.19
    VGG-16 Fashion-MNIST Leaky ReLU 512 30 92.81 610.82
    VGG-16 Fashion-MNIST PReLU 512 30 10.00 612.75
    VGG-16 Fashion-MNIST ELU 512 30 93.87 613.32
    VGG-16 CIFAR-10 Linear 512 100 83.25 958.74
    VGG-16 CIFAR-10 Sigmoid 512 100 10.00 957.62
    VGG-16 CIFAR-10 Tanh 512 100 10.00 957.97
    VGG-16 CIFAR-10 ReLU 512 100 83.22 957.74
    VGG-16 CIFAR-10 Leaky ReLU 512 100 83.37 958.39
    VGG-16 CIFAR-10 PReLU 512 100 82.17 963.67
    VGG-16 CIFAR-10 ELU 512 100 83.14 968.60
    VGG-16 CIFAR-100 Linear 512 200 1.00 1897.14
    VGG-16 CIFAR-100 Sigmoid 512 200 1.00 1868.02
    VGG-16 CIFAR-100 Tanh 512 200 1.00 1897.76
    VGG-16 CIFAR-100 ReLU 512 200 44.77 1901.56
    VGG-16 CIFAR-100 Leaky ReLU 512 200 48.22 1916.38
    VGG-16 CIFAR-100 PReLU 512 200 47.46 1922.29
    VGG-16 CIFAR-100 ELU 512 200 1.00 1917.75

    The empirical results show that:

    1. Non-rectified saturating activations (Sigmoid and Tanh) consistently fail to converge when fine-tuning deep pre-trained networks (yielding chance-level accuracies of 10.00%10.00\% on CIFAR-10 and 1.00%1.00\% on CIFAR-100 for VGG-16).
    2. Linear activations result in poor performance across all benchmarks because a multi-layer linear network collapses into a single linear map.
    3. Rectified variants (ReLU, Leaky ReLU, PReLU, ELU) achieve strong classification accuracy on deep networks, with Leaky ReLU and ELU providing top accuracies across datasets, while Leaky ReLU exhibits superior training stability and lower computational cost than ELU and PReLU.
  2. Knowl 2 — Comparative Benchmark of Optimization Algorithms on CIFAR-10

    data/table

    Ten gradient descent optimization algorithms were evaluated for training a pre-trained VGG-16 backbone (with top classification layer sizes of 512, 256, and 10) on the CIFAR-10 dataset using cross-entropy loss, ReLU activation, batch size 512, and 100 training epochs.

    Optimizer Validation Set Accuracy (%) Training Time (s)
    Mini-Batch Gradient Descent (MBGD) 85.70 926.24
    Momentum 86.37 947.18
    Nesterov Accelerated Gradient 84.32 945.92
    Adagrad 84.68 950.72
    Adadelta 86.06 965.90
    RMSprop 87.32 959.33
    Adam 83.09 953.46
    Adamax 86.18 960.83
    Nadam 86.26 968.72
    AMSgrad 83.25 951.28

    RMSprop achieved the highest validation accuracy (87.32%87.32\%), followed by Momentum (86.37%86.37\%) and Nadam (86.26%86.26\%). Mini-Batch Gradient Descent required the lowest training runtime per 100 epochs (926.24 s926.24\text{ s}) but exhibited the slowest epoch-over-epoch convergence rate compared to adaptive moment-based methods.

  3. Knowl 3 — Empirical Dynamics and Convergence Properties of CNN Activation Functions

    empirical result

    Empirical evaluations on MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100 across shallow (LeNet-5) and deep (VGG-16) convolutional networks demonstrate several activation behaviors:

    1. Sigmoid and Tanh saturation: In deep pre-trained CNNs, using Sigmoid or Tanh causes gradient vanishing in non-convex optimization, preventing the network from converging away from initial chance performance (e.g., 10.00%10.00\% on CIFAR-10 and 1.00%1.00\% on CIFAR-100 for VGG-16).
    2. Linear activation limits: Training without a non-linear activation function yields poor validation accuracy across all datasets because hidden layers collapse into a single affine transformation.
    3. Convergence and computational overhead: ELU achieves the highest validation accuracy on LeNet-5 (65.70%65.70\% on CIFAR-10 and 35.10%35.10\% on CIFAR-100) and VGG-16 (93.87%93.87\% on Fashion-MNIST), but requires longer training runtimes due to the exponential derivative calculations for negative inputs. In some deep configurations (such as VGG-16 on MNIST and CIFAR-100), ELU fails to learn without specialized initialization.
    4. Stability of Leaky ReLU: Leaky ReLU maintains better training stability and avoids dying neurons without incurring the parameter overhead of PReLU or the exponential computational cost of ELU, yielding 48.22%48.22\% on CIFAR-100 with VGG-16 compared to 44.77%44.77\% for standard ReLU.
  4. Knowl 4 — Empirical Convergence and Oscillation Characteristics of CNN Optimizers

    empirical result

    Empirical evaluation of ten optimization algorithms (MBGD, Momentum, Nesterov, Adagrad, Adadelta, RMSprop, Adam, Adamax, Nadam, AMSgrad) on CIFAR-10 with VGG-16 reveals the following dynamics:

    1. Mini-Batch Gradient Descent (MBGD) exhibits the slowest convergence rate across training epochs despite reaching 85.70%85.70\% validation accuracy by epoch 100.
    2. Nesterov, RMSprop, Adamax, and Nadam display pronounced loss oscillations during the middle and late stages of training. While high variance and oscillatory behavior can help parameter updates escape suboptimal local minima, they increase sensitivity to the learning rate and can cause divergence if the learning rate is set too high.
    3. Adaptive learning rate algorithms (RMSprop, Momentum, Nadam, and Adamax) consistently outperform standard Adam (83.09%83.09\%) and AMSgrad (83.25%83.25\%) on CIFAR-10 when using a fixed batch size of 512.
  5. Knowl 5 — Activation Function Selection Guidelines for Convolutional Neural Networks

    model/method

    Based on empirical comparisons across CNN architectures and benchmark datasets, the following rules of thumb guide activation function selection:

    1. Output layers: Employ the Sigmoid function σ(x)=11+ex\sigma(x) = \frac{1}{1 + e^{-x}} for binary classification problems and channel attention weight normalization (mapping activations to (0,1)(0, 1)). Employ the Softmax function for mutually exclusive multi-class classification.
    2. Intermediate hidden layers: ReLU (f(x)=max(0,x)f(x) = \max(0, x)) or Leaky ReLU serves as the standard default choice to avoid gradient vanishing while minimizing computation cost.
    3. Mitigating dead neurons: When a significant proportion of neurons become inactive (zero gradient in ReLU when x<0x < 0), use Leaky ReLU (f(x)=xf(x) = x for x0x \ge 0, f(x)=x/af(x) = x/a for x<0x < 0) with a fixed negative slope parameter (recommended α=1/a=0.02\alpha = 1/a = 0.02) to accelerate convergence, or adopt PReLU where the negative slope is learned.
    4. Saturation avoidance: Avoid Sigmoid and Tanh functions in deep CNN hidden layers due to gradient vanishing caused by near-zero derivatives in saturation regions.
  6. Knowl 6 — Loss Function Taxonomy and Selection Guidelines for CNN Applications

    model/method

    Loss functions in CNNs are categorized according to target problem structures:

    1. Regression problems:
    • Mean Absolute Error (MAE / L1L_1 loss): MAE=1ni=1nyiy^i\text{MAE} = \frac{1}{n} \sum_{i=1}^n |y_i - \hat{y}_i|. Robust to extreme outliers in the dataset, but its derivative is non-smooth at 0, preventing dynamic update-rate decay near the optimum.
    • Mean Squared Error (MSE / L2L_2 loss): MSE=1ni=1n(yiy^i)2\text{MSE} = \frac{1}{n} \sum_{i=1}^n (y_i - \hat{y}_i)^2. Differentiable everywhere with gradient scaling proportional to error, but highly sensitive to dataset outliers.
    1. Classification and representation learning:
    • Cross-Entropy (Softmax) Loss: Evaluates divergence between the predicted probability distribution and ground-truth one-hot distribution. Maximizes classification margin correctness but does not explicitly constrain intra-class compactness.
    • Contrastive Loss: Minimizes distance between pairs of features from the same class while enforcing a minimum margin distance between pairs from different classes.
    • Triplet Loss: Minimizes the distance between an anchor xax_a and a positive sample xpx_p while maximizing the distance between xax_a and a negative sample xnx_n subject to a margin α\alpha: f(xa)f(xp)22+αf(xa)f(xn)22\|f(x_a) - f(x_p)\|_2^2 + \alpha \le \|f(x_a) - f(x_n)\|_2^2
    • Center Loss: Minimizes intra-class variance by penalizing the L2L_2 distance between deep feature embeddings and their corresponding class centroids.
    • Large-Margin Softmax (and angular variants such as SphereFace, CosFace, ArcFace): Introduces angular and additive margin constraints into the weight matrix to simultaneously enforce intra-class feature compactness and inter-class boundary separation.
  7. Knowl 7 — Optimizer Selection and Configuration Guidelines for CNN Training

    model/method

    Practical rules of thumb for configuring optimization algorithms in CNN workflows:

    1. Batch sampling: Always use Mini-Batch Gradient Descent (MBGD) or mini-batch-based adaptive optimizers rather than full-batch BGD (excessive memory cost and slow updates) or pure single-sample SGD (extreme variance preventing stable convergence).
    2. Adaptive moment optimizers: For general computer vision tasks, adaptive algorithms (such as RMSprop, Adam, Nadam, or Momentum) should be tested against dataset distributions to identify optimal convergence properties.
    3. Instability mitigation: When training experiences severe loss oscillation, erratic validation accuracy, or divergence under adaptive optimizers (such as Nesterov, RMSprop, or Nadam), decrement the initial base learning rate.
  8. Knowl 8 — Hardware-Aware Network Design Guidelines for Efficient CNN Architectures

    theoretical result

    Network operational runtime depends heavily on Memory Access Cost (MAC) rather than FLOP count alone. Four core design principles govern the construction of high-speed, hardware-efficient CNN architectures:

    1. Equal Channel Ratio: For 1×11 \times 1 pointwise convolutional layers, MAC is minimized when the number of input channels c1c_1 equals the number of output channels c2c_2 (c1=c2c_1 = c_2) under a fixed compute budget (FLOPs).
    2. Convolution Group Cost: Increasing the number of groups gg in group convolutions reduces FLOPs but increases memory access overhead (MAC), which degrades hardware utilization and reduces training/inference speed on GPUs.
    3. Network Fragmentation Reduction: High degrees of structural fragmentation (e.g., executing multiple small parallel convolution branches within a single block, as in Inception-style structures) reduce GPU compute parallelism and increase kernel launch overhead.
    4. Element-Wise Operation Minimization: Element-wise operations (including ReLU, bias addition, shortcut tensor additions, and depthwise operations) consume substantial execution time relative to their FLOP contribution due to high memory bandwidth requirements; their frequency should be minimized.
  9. Knowl 9 — Structural Limitations of Convolutional Neural Networks and the Capsule Network Paradigm

    limitation

    Conventional CNNs exhibit two fundamental structural limitations:

    1. Loss of Spatial Part-Whole Relationships: Pooling operations (max and average pooling) enforce spatial downsampling and translational invariance by discarding local coordinate information, but ignore the spatial transformation and relative geometric relationships between object subcomponents and the whole object.
    2. Sensitivity to Affine Transformations: Standard CNNs lack equivariance and fail to generalize to transformed variants of objects (including rotations, scaling, and reflections) unless trained with extensive data augmentation.

    Capsule Neural Networks (CapsNet) and Stacked Capsule Autoencoders (SCAE) address these limitations by grouping neurons into multidimensional vector capsules. The length of a capsule output vector represents the probability of entity presence, while the vector orientation explicitly encodes geometric pose, orientation, scale, and deformation parameters.

  10. Knowl 10 — Security Vulnerabilities and Defense Mechanisms in Convolutional Neural Networks

    model/method

    CNN security vulnerabilities fall into two primary threat vectors along with corresponding defensive countermeasures:

    1. Data Poisoning and Backdoor Injection: Imperceptible perturbation patterns are injected into training samples during training. The network behaves normally on clean inputs but consistently misclassifies inputs containing the specific backdoor trigger. Countermeasures include fine-pruning, which combines structural pruning of dormant backdoor-associated neurons with network fine-tuning on clean samples.
    2. Adversarial Attacks: Small, calculated input perturbations (e.g., generated via Fast Gradient Sign Methods) exploit the piecewise-linear nature of CNN activations (such as ReLU and Maxout) to force erroneous classifications with high model confidence. Countermeasures include:
    • Training example augmentation: Training the backbone directly on generated adversarial examples.
    • Architectural modification: Designing denoising layers or sub-structures that filter high-frequency adversarial perturbations.
    • Auxiliary networks: Employing supplementary detection networks attached to the primary backbone to flag perturbed inputs.

Coverage note — Bibliographical surveys of prior CNN architectures (e.g., AlexNet, VGG, ResNet, MobileNets, ShuffleNets) and historical background were omitted, as they represent reviews of existing literature rather than original contributed models or findings.

References

  1. 1.W. S. Mcculloch, and W. H. Pitts, “A logical Calculus of Ideas Immanent in Nervous Activity,” The Bulletin of Mathematical Biophysics, vol. 5, pp. 115-133, 1942.
  2. 2.F. Rosenblatt, “The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain,” Psychological Review, pp. 368-408, 1958.
  3. 3.C. V. D. Malsburg, “Frank Rosenblatt: Principles of Neurodynamics: Perceptrons and the Theory of Brain Mechanisms.”
  4. 4.Davd. Rumhar, Geoffrey. Hinton, and RonadJ. Wams, “Learning representations by back-propagating errors.”
  5. 5.A. Waibel, T. Hanazawa, G. E. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,” IEEE Transactions on Acoustics Speech & Signal Processing, vol. 37, no. 3, pp. 328-339, 1989.
  6. 6.W. Zhang, “Shift-invariant pattern recognition neural network and its optical architecture,” in Proceedings of annual conference of the Japan Society of Applied Physics, 1988.
  7. 7.Y. Lecun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, “Backpropagation Applied to Handwritten Zip Code Recognition,” Neural Computation, vol. 1, no. 4, pp. 541-551.
  8. 8.K. Aihara, T. Takabe, and M. Toyoda, “Chaotic neural networks,” Physics Letters A, vol. 144, no. 6-7, pp. 333-340.
  9. 9.Specht, and D.F., “A general regression neural network,” IEEE Transactions on Neural Networks, vol. 2, no. 6, pp. 568-576.
  10. 10.B. L. Lecun Y , Bengio Y , et al., “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278-2324, 1998.
  11. 11.A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” Advances in neural information processing systems, vol. 25, no. 2, 2012.
  12. 12.N. Aloysius, and M. Geetha, "A review on deep convolutional neural networks," Proceedings of the 2017 IEEE International Conference on Communication and Signal Processing, ICCSP 2017. pp. 588-592.
  13. 13.A. Dhillon, and G. K. Verma, “Convolutional neural network: a review of models, methodologies and applications to object detection,” Progress in Artificial Intelligence, 2019/12/20, 2019.
  14. 14.W. Rawat, and Z. Wang, “Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review,” Neural Computation, pp. 1-98.
  15. 15.Q. Liu, N. Zhang, W. Yang, S. Wang, Z. Cui, X. Chen, and L. Chen, “A Review of Image Recognition with Deep Convolutional Neural Network.”
  16. 16.S. Rehman, H. Ajmal, U. Farooq, Q. U. Ain, and A. Hassan, "Convolutional neural network based image segmentation: a review."
  17. 17.T. Lindeberg, “Scale invariant feature transform,” 2012.
  18. 18.N. Dalal, and B. Triggs, "Histograms of oriented gradients for human detection." pp. 886-893.
  19. 19.T. Ahonen, A. Hadid, and M. Pietikainen, “Face description with local binary patterns: Application to face recognition,” IEEE transactions on pattern analysis and machine intelligence, vol. 28, no. 12, pp. 2037-2041, 2006.
  20. 20.W. T. N. Hubel D H “Receptive fields, binocular interaction and functional architecture in the cat"s visual cortex,” The Journal of Physiology, vol. 160, no. 1, pp. 106-154, 1962.
  21. 21.D. M. Hawkins, “The problem of overfitting,” Journal of chemical information and computer sciences, vol. 44, no. 1, pp. 1-12, 2004.
  22. 22.K. Fukushima, “Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position,” Biological Cybernetics, vol. 36, no. 4, pp. 193-202.
  23. 23.J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable Convolutional Networks.”
  24. 24.L. Sifre, and S. Mallat, “Rigid-Motion Scattering for Texture Classification,” 03/07, 2014.
  25. 25.F. Mamalet, and C. Garcia, Simplifying ConvNets for Fast Learning, 2012.
  26. 26.F. Chollet, “Xception: Deep Learning with Depthwise Separable Convolutions.”
  27. 27.W. Min, B. Liu, and H. Foroosh, “Factorized Convolutional Neural Networks.”
  28. 28.D. Li, A. Zhou, and A. Yao, “HBONet: Harmonious Bottleneck on Two Orthogonal Dimensions.”
  29. 29.S. Xie, R. Girshick, P. Dollar, Z. Tu, and K. He, "Aggregated Residual Transformations for Deep Neural Networks."
  30. 30.T. K. Lee, W. J. Baddar, S. T. Kim, and Y. M. Ro, “Convolution with Logarithmic Filter Groups for Efficient Shallow CNN.”
  31. 31.Y. Ioannou, D. Robertson, R. Cipolla, and A. Criminisi, “Deep Roots: Improving CNN Efficiency with Hierarchical Filter Groups.”
  32. 32.K. Simonyan, and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” Computer Science, 2014.
  33. 33.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, "Going deeper with convolutions." pp. 1-9.
  34. 34.S. Ioffe, and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” arXiv preprint arXiv:1502.03167, 2015.
  35. 35.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, "Rethinking the inception architecture for computer vision." pp. 2818-2826.
  36. 36.C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi, "Inception-v4, inception-resnet and the impact of residual connections on learning."
  37. 37.K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition." pp. 770-778.
  38. 38.K. He, X. Zhang, S. Ren, and S. Jian, "Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification."
  39. 39.S. Zagoruyko, and N. Komodakis, "Wide Residual Networks."
  40. 40.G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Weinberger, “Deep Networks with Stochastic Depth.”
  41. 41.S. Targ, D. Almeida, and K. Lyman, “Resnet in Resnet: Generalizing Residual Architectures.”
  42. 42.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, "Generative adversarial nets." pp. 2672-2680.
  43. 43.A. Radford, L. Metz, and S. Chintala, “Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks,” Computer Science, 2015.
  44. 44.A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017.
  45. 45.M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, "Mobilenetv2: Inverted residuals and linear bottlenecks." pp. 4510-4520.
  46. 46.M. S. Andrew Howard, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V. Le, Hartwig Adam, “Searching for MobileNetV3,” arXiv:1905.02244 [cs.CV], 2019.
  47. 47.T. J. Yang, A. Howard, B. Chen, X. Zhang, A. Go, M. Sandler, V. Sze, and H. Adam, “NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications.”
  48. 48.J. Hu, L. Shen, S. Albanie, G. Sun, and E. Wu, “Squeeze-and-Excitation Networks.”
  49. 49.X. Zhang, X. Zhou, M. Lin, and J. Sun, “ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices.”
  50. 50.N. Ma, X. Zhang, H. T. Zheng, and J. Sun, “ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design.”
  51. 51.K. Han, Y. Wang, Q. Tian, J. Guo, C. Xu, and C. Xu, “GhostNet: More Features from Cheap Operations,” arXiv preprint arXiv:1911.11907, 2019.
  52. 52.V. Nair, and G. E. Hinton, "Rectified Linear Units Improve Restricted Boltzmann Machines Vinod Nair."
  53. 53.M. T. Hagan, H. B. Demuth, and M. H. Beale, Neural network design, 2002.
  54. 54.A. Krizhevsky, “Convolutional Deep Belief Networks on CIFAR-10,” 2010.
  55. 55.D.-A. Clevert, T. Unterthiner, and S. Hochreiter, Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs), 2016.
  56. 56.H. Xiao, K. Rasul, and R. Vollgraf, “Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms.”
  57. 57.A. Krizhevsky, “Learning Multiple Layers of Features from Tiny Images,” University of Toronto, 05/08, 2012.
  58. 58.J. Deng, W. Dong, R. Socher, L. J. Li, K. Li, and F. F. Li, “ImageNet: A large-scale hierarchical image database,” Proc of IEEE Computer Vision & Pattern Recognition, pp. 248-255, 2009.
  59. 59.R. Hadsell, S. Chopra, and Y. LeCun, "Dimensionality reduction by learning an invariant mapping." pp. 1735-1742.
  60. 60.S. Chopra, R. Hadsell, and Y. LeCun, "Learning a similarity metric discriminatively, with application to face verification." pp. 539-546.
  61. 61.Y. Sun, Y. Chen, X. Wang, and X. Tang, "Deep learning face representation by joint identification-verification." pp. 1988-1996.
  62. 62.Y. Sun, X. Wang, and X. Tang, "Deeply learned face representations are sparse, selective, and robust." pp. 2892-2900.
  63. 63.Y. Sun, D. Liang, X. Wang, and X. Tang, “Deepid3: Face recognition with very deep neural networks,” arXiv preprint arXiv:1502.00873, 2015.
  64. 64.F. Schroff, D. Kalenichenko, and J. Philbin, "Facenet: A unified embedding for face recognition and clustering." pp. 815-823.
  65. 65.O. M. Parkhi, A. Vedaldi, and A. Zisserman, “Deep face recognition,” 2015.
  66. 66.B. Amos, B. Ludwiczuk, and M. Satyanarayanan, “Openface: A general-purpose face recognition library with mobile applications,” CMU School of Computer Science, vol. 6, pp. 2, 2016.
  67. 67.D. Cheng, Y. Gong, S. Zhou, J. Wang, and N. Zheng, "Person re-identification by multi-channel parts-based cnn with improved triplet loss function." pp. 1335-1344.
  68. 68.A. Hermans, L. Beyer, and B. Leibe, “In defense of the triplet loss for person re-identification,” arXiv preprint arXiv:1703.07737, 2017.
  69. 69.R. Kuma, E. Weill, F. Aghdasi, and P. Sriram, "Vehicle re-identification: an efficient baseline using triplet embedding." pp. 1-9.
  70. 70.Y. Wen, K. Zhang, Z. Li, and Y. Qiao, "A discriminative feature learning approach for deep face recognition." pp. 499-515.
  71. 71.J. Yao, Y. Yu, Y. Deng, and C. Sun, "A feature learning approach for image retrieval." pp. 405-412.
  72. 72.H. Jin, X. Wang, S. Liao, and S. Z. Li, "Deep person re-identification with improved embedding and efficient training." pp. 261-267.
  73. 73.G. Wisniewksi, H. Bredin, G. Gelly, and C. Barras, "Combining speaker turn embedding and incremental structure prediction for low-latency speaker diarization."
  74. 74.W. Liu, Y. Wen, Z. Yu, and M. Yang, "Large-margin softmax loss for convolutional neural networks." p. 7.
  75. 75.L. Tan, K. Zhang, K. Wang, X. Zeng, X. Peng, and Y. Qiao, "Group emotion recognition with individual facial emotion CNNs and global image based CNNs." pp. 549-552.
  76. 76.Y. Liu, L. He, and J. Liu, “Large margin softmax loss for speaker verification,” arXiv preprint arXiv:1904.03479, 2019.
  77. 77.D. Saad, On-line learning in neural networks: Cambridge University Press, 2009.
  78. 78.G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, "Densely connected convolutional networks." pp. 4700-4708.
  79. 79.Y. Sun, X. Wang, and X. Tang, "Deep learning face representation from predicting 10,000 classes." pp. 1891-1898.
  80. 80.N. Qian, “On the momentum term in gradient descent learning algorithms,” Neural networks, vol. 12, no. 1, pp. 145-151, 1999.
  81. 81.Y. Nesterov, "A method for unconstrained convex minimization problem with the rate of convergence O (1/k^ 2)." pp. 543-547.
  82. 82.W. Su, L. Chen, M. Wu, M. Zhou, Z. Liu, and W. Cao, "Nesterov accelerated gradient descent-based convolution neural network with dropout for facial expression recognition." pp. 1063-1068.
  83. 83.A. L. Maas, P. Qi, Z. Xie, A. Y. Hannun, C. T. Lengerich, D. Jurafsky, and A. Y. Ng, “Building DNN acoustic models for large vocabulary speech recognition,” Computer Speech & Language, vol. 41, pp. 195-213, 2017.
  84. 84.P. Molchanov, S. Gupta, K. Kim, and J. Kautz, "Hand gesture recognition with 3D convolutional neural networks." pp. 1-7.
  85. 85.J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of machine learning research, vol. 12, no. Jul, pp. 2121-2159, 2011.
  86. 86.M. D. Zeiler, “Adadelta: an adaptive learning rate method,” arXiv preprint arXiv:1212.5701, 2012.
  87. 87.J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, "Attention-based models for speech recognition." pp. 577-585.
  88. 88.T. Sercu, C. Puhrsch, B. Kingsbury, and Y. LeCun, "Very deep multilingual convolutional neural networks for LVCSR." pp. 4955-4959.
  89. 89.Y. Kim, “Convolutional neural networks for sentence classification,” arXiv preprint arXiv:1408.5882, 2014.
  90. 90.G. Hinton, N. Srivastava, and K. Swersky, “Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,” Cited on, vol. 14, no. 8, 2012.
  91. 91.D. P. Kingma, and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980, 2014.
  92. 92.S. Sharma, R. Kiros, and R. Salakhutdinov, “Action recognition using visual attention,” arXiv preprint arXiv:1511.04119, 2015.
  93. 93.F. Korzeniowski, and G. Widmer, "A fully convolutional deep auditory model for musical chord recognition." pp. 1-6.
  94. 94.M. J. Van Putten, S. Olbrich, and M. Arns, “Predicting sex from brain rhythms with deep learning,” Scientific reports, vol. 8, no. 1, pp. 1-7, 2018.
  95. 95.S. Niklaus, L. Mai, and F. Liu, "Video frame interpolation via adaptive separable convolution." pp. 261-270.
  96. 96.T. Dozat, “Incorporating nesterov momentum into adam,” 2016.
  97. 97.D. Q. Nguyen, and K. Verspoor, “Convolutional neural networks for chemical-disease relation extraction are improved with character-based word embeddings,” arXiv preprint arXiv:1805.10586, 2018.
  98. 98.S. Maetschke, B. Antony, H. Ishikawa, G. Wollstein, J. Schuman, and R. Garnavi, “A feature agnostic approach for glaucoma detection in OCT volumes,” PloS one, vol. 14, no. 7, 2019.
  99. 99.A. Schindler, T. Lidy, and A. Rauber, “Multi-temporal resolution convolutional neural networks for acoustic scene classification,” arXiv preprint arXiv:1811.04419, 2018.
  100. 100.S. J. Reddi, S. Kale, and S. Kumar, “On the convergence of adam and beyond,” arXiv preprint arXiv:1904.09237, 2019.
  101. 101.M. Jahanifar, N. Z. Tajeddin, N. A. Koohbanani, A. Gooya, and N. Rajpoot, “Segmentation of skin lesions and their attributes using multi-scale convolutional neural networks and domain specific augmentations,” arXiv preprint arXiv:1809.10243, 2018.
  102. 102.F. Monti, F. Frasca, D. Eynard, D. Mannion, and M. M. Bronstein, “Fake news detection on social media using geometric deep learning,” arXiv preprint arXiv:1902.06673, 2019.
  103. 103.S. Liu, E. Gibson, S. Grbic, Z. Xu, A. A. A. Setio, J. Yang, B. Georgescu, and D. Comaniciu, “Decompose to manipulate: manipulable object synthesis in 3D medical images with structured image decomposition,” arXiv preprint arXiv:1812.01737, 2018.
  104. 104.Urtnasan, Erdenebayar, Hyeonggon, Kim, Jong-Uk, Park, Dongwon, Kang, Kyoung-Joung, and Lee, “Automatic Prediction of Atrial Fibrillation Based on Convolutional Neural Network Using a Short-term Normal Electrocardiogram Signal.”
  105. 105.S. Harbola, and V. Coors, “One dimensional convolutional neural network architectures for wind prediction,” Energy Conversion and Management, vol. 195, pp. 70-75, 2019.
  106. 106.D. Han, J. Chen, and J. Sun, “A parallel spatiotemporal deep learning network for highway traffic flow forecasting,” International Journal of Distributed Sensor Networks, vol. 15, no. 2.
  107. 107.Q. Zhang, D. Zhou, and X. Zeng, “HeartID: A Multiresolution Convolutional Neural Network for ECG-based Biometric Human Identification in Smart Health Applications,” IEEE Access, pp. 1-1.
  108. 108.O. Abdeljaber, O. Avci, S. Kiranyaz, M. Gabbouj, and D. J. Inman, “Real-Time Vibration-Based Structural Damage Detection Using One-Dimensional Convolutional Neural Networks,” Journal of Sound & Vibration, vol. 388, pp. 154-170, 2017.
  109. 109.O. Abdeljaber, S. Sassi, O. Avci, S. Kiranyaz, A. A. Ibrahim, and M. Gabbouj, “Fault Detection and Severity Identification of Ball Bearings by Online Condition Monitoring,” IEEE Transactions on Industrial Electronics, pp. 1-1.
  110. 110.K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” IEEE transactions on pattern analysis and machine intelligence, vol. 37, no. 9, pp. 1904-1916, 2015.
  111. 111.Y. Chen, J. Li, H. Xiao, X. Jin, S. Yan, and J. Feng, "Dual path networks." pp. 4467-4475.
  112. 112.F. Iandola, M. Moskewicz, S. Karayev, R. Girshick, T. Darrell, and K. Keutzer, “Densenet: Implementing efficient convnet descriptor pyramids,” arXiv preprint arXiv:1404.1869, 2014.
  113. 113.S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, "Aggregated residual transformations for deep neural networks." pp. 1492-1500.
  114. 114.Q. Li, W. Cai, X. Wang, Y. Zhou, D. D. Feng, and M. Chen, "Medical image classification with convolutional neural network." pp. 844-848.
  115. 115.Y. Jiang, L. Chen, H. Zhang, and X. Xiao, “Breast cancer histopathological image classification using convolutional neural networks with small SE-ResNet module,” PloS one, vol. 14, no. 3, 2019.
  116. 116.D. R. Bruno, and F. S. Osório, "Image classification system based on deep learning applied to the recognition of traffic signs for intelligent robotic vehicle navigation purposes." pp. 1-6.
  117. 117.R. Madan, D. Agrawal, S. Kowshik, H. Maheshwari, S. Agarwal, and D. Chakravarty, “Traffic Sign Classification using Hybrid HOG-SURF Features and Convolutional Neural Networks,” 2019.
  118. 118.M. Zhang, W. Li, and Q. Du, “Diverse region-based CNN for hyperspectral image classification,” IEEE Transactions on Image Processing, vol. 27, no. 6, pp. 2623-2634, 2018.
  119. 119.A. Sharma, X. Liu, X. Yang, and D. Shi, “A patch-based convolutional neural network for remote sensing image classification,” Neural Networks, vol. 95, pp. 19-28, 2017.
  120. 120.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, "You only look once: Unified, real-time object detection." pp. 779-788.
  121. 121.J. Redmon, and A. Farhadi, "YOLO9000: better, faster, stronger." pp. 7263-7271.
  122. 122.J. Redmon, and A. Farhadi, “Yolov3: An incremental improvement,” arXiv preprint arXiv:1804.02767, 2018.
  123. 123.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, "Ssd: Single shot multibox detector." pp. 21-37.
  124. 124.H. Law, and J. Deng, "Cornernet: Detecting objects as paired keypoints." pp. 734-750.
  125. 125.H. Law, Y. Teng, O. Russakovsky, and J. Deng, “Cornernet-lite: Efficient keypoint based object detection,” arXiv preprint arXiv:1904.08900, 2019.
  126. 126.R. Girshick, J. Donahue, T. Darrell, and J. Malik, "Rich feature hierarchies for accurate object detection and semantic segmentation." pp. 580-587.
  127. 127.R. Girshick, "Fast r-cnn." pp. 1440-1448.
  128. 128.S. Ren, K. He, R. Girshick, and J. Sun, "Faster r-cnn: Towards real-time object detection with region proposal networks." pp. 91-99.
  129. 129.T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, "Feature pyramid networks for object detection." pp. 2117-2125.
  130. 130.E. Shelhamer, J. Long, and T. Darrell, “Fully Convolutional Networks for Semantic Segmentation.”
  131. 131.O. Ronneberger, P. Fischer, and T. Brox, "U-Net: Convolutional Networks for Biomedical Image Segmentation."
  132. 132.A. Paszke, A. Chaurasia, S. Kim, and E. Culurciello, “ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation.”
  133. 133.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid Scene Parsing Network.”
  134. 134.L. C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs,” IEEE Transactions on Pattern Analysis & Machine Intelligence, vol. 40, no. 4, pp. 834, 2018.
  135. 135.A. Pal, S. Jaiswal, S. Ghosh, N. Das, and M. Nasipuri, "Segfast: A faster squeezenet based semantic image segmentation technique using depth-wise separable convolutions."
  136. 136.K. He, G. Georgia, D. Piotr, and G. Ross, “Mask R-CNN,” IEEE Transactions on Pattern Analysis & Machine Intelligence, pp. 1-1.
  137. 137.D. Bolya, C. Zhou, F. Xiao, and Y. J. Lee, “YOLACT: Real-time Instance Segmentation.”
  138. 138.T. Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal Loss for Dense Object Detection,” IEEE Transactions on Pattern Analysis & Machine Intelligence, vol. PP, no. 99, pp. 2999-3007, 2017.
  139. 139.A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár, “Panoptic Segmentation.”
  140. 140.A. Kirillov, R. Girshick, K. He, and P. Dollár, “Panoptic Feature Pyramid Networks.”
  141. 141.H. Liu, C. Peng, C. Yu, J. Wang, X. Liu, G. Yu, and W. Jiang, “An End-to-End Network for Panoptic Segmentation.”
  142. 142.Y. Taigman, M. Yang, M. A. Ranzato, and L. Wolf, "Deepface: Closing the gap to human-level performance in face verification." pp. 1701-1708.
  143. 143.W. Liu, Y. Wen, Z. Yu, M. Li, B. Raj, and L. Song, “SphereFace: Deep Hypersphere Embedding for Face Recognition.”
  144. 144.J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “ArcFace: Additive Angular Margin Loss for Deep Face Recognition.”
  145. 145.H. Wang, Y. Wang, Z. Zhou, X. Ji, D. Gong, J. Zhou, Z. Li, and W. Liu, “CosFace: Large Margin Cosine Loss for Deep Face Recognition.”
  146. 146.C. Cao, Y. Zhang, C. Zhang, and H. Lu, "Action Recognition with Joints-Pooled 3D Deep Convolutional Descriptors."
  147. 147.A. Stergiou, and R. Poppe, “Spatio-Temporal FAST 3D Convolutions for Human Action Recognition.”
  148. 148.J. Huang, W. Zhou, H. Li, and W. Li, “Attention-based 3D-CNNs for large-vocabulary sign language recognition,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 29, no. 9, pp. 2822-2832, 2018.
  149. 149.Y. Huang, S.-H. Lai, and S.-H. Tai, “Human Action Recognition Based on Temporal Pose CNN and Multi-dimensional Fusion.”
  150. 150.Z. Wu, S. Song, A. Khosla, F. Yu, L. Zhang, X. Tang, and J. Xiao, “3D ShapeNets: A Deep Representation for Volumetric Shapes.”
  151. 151.D. Maturana, and S. Scherer, "Voxnet: A 3d convolutional neural network for real-time object recognition." pp. 922-928.
  152. 152.S. Song, and J. Xiao, "Deep Sliding Shapes for Amodal 3D Object Detection in RGB-D Images."
  153. 153.Y. Zhou, and O. Tuzel, “VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection.”
  154. 154.F. Pastor, J. M. Gandarias, A. J. García-Cerezo, and J. M. Gómez-de-Gabriel, “Using 3D Convolutional Neural Networks for Tactile Object Recognition with Robotic Palpation,” Sensors, vol. 19, no. 24, pp. 5356, 2019.
  155. 155.K. Jnawali, M. R. Arbabshirani, N. Rao, and A. A. Patel, "Deep 3D convolution neural network for CT brain hemorrhage classification." p. 105751C.
  156. 156.S. Hamidian, B. Sahiner, N. Petrick, and A. Pezeshk, "3D convolutional neural network for automatic detection of lung nodules in chest CT." p. 1013409.
  157. 157.M. Jaderberg, A. Vedaldi, and A. Zisserman, “Speeding up convolutional neural networks with low rank expansions,” arXiv preprint arXiv:1405.3866, 2014.
  158. 158.V. Sindhwani, T. Sainath, and S. Kumar, "Structured transforms for small-footprint deep learning." pp. 3088-3096.
  159. 159.S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” arXiv preprint arXiv:1510.00149, 2015.
  160. 160.H. Mao, S. Han, J. Pool, W. Li, X. Liu, Y. Wang, and W. J. Dally, “Exploring the regularity of sparse structure in convolutional neural networks,” arXiv preprint arXiv:1705.08922, 2017.
  161. 161.M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks.”
  162. 162.X. Lin, C. Zhao, and W. Pan, “Towards Accurate Binary Convolutional Neural Network.”
  163. 163.C. Zhu, S. Han, H. Mao, and W. J. Dally, “Trained Ternary Quantization.”
  164. 164.Y. Choi, M. El-Khamy, and J. Lee, “Towards the limit of network quantization,” arXiv preprint arXiv:1612.01543, 2016.
  165. 165.P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, and G. Venkatesh, “Mixed Precision Training.”
  166. 166.M. Sajjad, S. Khan, T. Hussain, K. Muhammad, A. K. Sangaiah, A. Castiglione, C. Esposito, and S. W. Baik, “CNN-based anti-spoofing two-tier multi-factor authentication system,” Pattern Recognition Letters, vol. 126, pp. 123-131, 2019.
  167. 167.K. Itqan, A. Syafeeza, F. Gong, N. Mustafa, Y. Wong, and M. Ibrahim, “User identification system based on finger-vein patterns using convolutional neural network,” ARPN Journal of Engineering and Applied Sciences, vol. 11, no. 5, pp. 3316-3319, 2016.
  168. 168.H. Ke, D. Chen, X. Li, Y. Tang, T. Shah, and R. Ranjan, “Towards brain big data classification: Epileptic EEG identification with a lightweight VGGNet on global MIC,” IEEE Access, vol. 6, pp. 14722-14733, 2018.
  169. 169.A. Shustanov, and P. Yakimov, “CNN design for real-time traffic sign recognition,” Procedia engineering, vol. 201, pp. 718-725, 2017.
  170. 170.J. Špaňhel, J. Sochor, R. Juránek, A. Herout, L. Maršík, and P. Zemčík, "Holistic recognition of low quality license plates by CNN using track annotated data." pp. 1-6.
  171. 171.T. Xie, and Y. Li, "A Gradient-Based Algorithm to Deceive Deep Neural Networks." pp. 57-65.
  172. 172.C. Liao, H. Zhong, A. Squicciarini, S. Zhu, and D. Miller, “Backdoor embedding in convolutional neural network models via invisible perturbation,” arXiv preprint arXiv:1808.10307, 2018.
  173. 173.A. N. Bhagoji, S. Chakraborty, P. Mittal, and S. Calo, “Analyzing federated learning through an adversarial lens,” arXiv preprint arXiv:1811.12470, 2018.
  174. 174.I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples. arXiv,” preprint, 2018.
  175. 175.K. Liu, B. Dolan-Gavitt, and S. Garg, "Fine-pruning: Defending against backdooring attacks on deep neural networks." pp. 273-294.
  176. 176.N. Akhtar, and A. Mian, “Threat of adversarial attacks on deep learning in computer vision: A survey,” IEEE Access, vol. 6, pp. 14410-14430, 2018.
  177. 177.B. Zoph, and Q. V. Le, “Neural architecture search with reinforcement learning,” arXiv preprint arXiv:1611.01578, 2016.
  178. 178.H. Pham, M. Y. Guan, B. Zoph, Q. V. Le, and J. Dean, “Efficient neural architecture search via parameter sharing,” arXiv preprint arXiv:1802.03268, 2018.
  179. 179.H. Cai, L. Zhu, and S. Han, “Proxylessnas: Direct neural architecture search on target task and hardware,” arXiv preprint arXiv:1812.00332, 2018.
  180. 180.M. Tan, B. Chen, R. Pang, V. Vasudevan, M. Sandler, A. Howard, and Q. V. Le, "Mnasnet: Platform-aware neural architecture search for mobile." pp. 2820-2828.
  181. 181.G. Ghiasi, T.-Y. Lin, and Q. V. Le, "Nas-fpn: Learning scalable feature pyramid architecture for object detection." pp. 7036-7045.
  182. 182.T. Jajodia, and P. Garg, “Image Classification–Cat and Dog Images,” Image, vol. 6, no. 12, 2019.
  183. 183.P. Drews, G. Williams, B. Goldfain, E. A. Theodorou, and J. M. Rehg, “Aggressive deep driving: Model predictive control with a cnn cost model,” arXiv preprint arXiv:1707.05303, 2017.
  184. 184.H. Gao, B. Cheng, J. Wang, K. Li, J. Zhao, and D. Li, “Object classification using CNN-based fusion of vision and LIDAR in autonomous vehicle environment,” IEEE Transactions on Industrial Informatics, vol. 14, no. 9, pp. 4224-4231, 2018.
  185. 185.A. Azulay, and Y. Weiss, “Why do deep convolutional networks generalize so poorly to small image transformations?,” 2018.
  186. 186.S. Sabour, N. Frosst, and G. E. Hinton, "Dynamic routing between capsules." pp. 3856-3866.
  187. 187.A. Kosiorek, S. Sabour, Y. W. Teh, and G. E. Hinton, "Stacked capsule autoencoders." pp. 15486-15496.

Citation

MLA
Li, Z., et al. “A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects”. IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 12, 2022, pp. 6999–7019, https://doi.org/10.1109/TNNLS.2021.3084827.
APA
Li, Z., Liu, F., Yang, W., Peng, S., & Zhou, J. (2022). A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects. IEEE Transactions on Neural Networks and Learning Systems, 33(12), 6999–7019. https://doi.org/10.1109/TNNLS.2021.3084827
Chicago
Li, Z., F. Liu, W. Yang, S. Peng, and J. Zhou. 2022. “A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects”. IEEE Transactions on Neural Networks and Learning Systems 33 (12): 6999–7019. https://doi.org/10.1109/TNNLS.2021.3084827.
Harvard
Li, Z. et al. (2022) “A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects”, IEEE Transactions on Neural Networks and Learning Systems, 33(12), pp. 6999–7019. Available at: https://doi.org/10.1109/TNNLS.2021.3084827.
Vancouver
1. Li Z, Liu F, Yang W, Peng S, Zhou J (2022) A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects. IEEE Transactions on Neural Networks and Learning Systems 33:6999–7019

BibTeX

@article{Li_2022, title={A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects}, volume={33}, ISSN={2162-2388}, url={http://dx.doi.org/10.1109/TNNLS.2021.3084827}, DOI={10.1109/tnnls.2021.3084827}, number={12}, journal={IEEE Transactions on Neural Networks and Learning Systems}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Li, Zewen and Liu, Fan and Yang, Wenjie and Peng, Shouheng and Zhou, Jun}, year={2022}, month=Dec, pages={6999–7019} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF