A survey of the recent architectures of deep convolutional neural networks
Asifullah KhanAnabia SohailUmme ZahooraAqsa Saeed Qureshi
Classifies recent deep convolutional neural network architectures into seven structural categories—including spatial exploitation, depth, multi-path routing, and attention mechanisms—to explain the design principles driving modern computer vision systems.
The article surveys the rapid evolution of deep convolutional neural networks, which have achieved strong results on computer vision tasks such as image classification, object detection, and segmentation. Early CNNs struggled with scale and complexity, but hardware advances and large datasets revived interest after 2012, prompting a wave of architectural changes aimed at improving representational power while managing training difficulties like vanishing gradients.
The survey set out to organize the most prominent CNN designs reported between 2012 and 2020 into a clear taxonomy and to explain the basic components, historical development, applications, and remaining challenges. It draws on published architectures, performance benchmarks on datasets such as ImageNet and CIFAR, and comparative tables to identify patterns in how networks are structured.
The analysis shows that the largest gains have come from replacing simple stacked layers with reusable blocks that exploit spatial information at multiple scales, increase depth or width, add shortcut connections, recalibrate feature maps, boost input channels, or apply attention. Notable examples include residual networks that ease training of very deep models, inception-style blocks that capture features at different resolutions, and attention modules that focus computation on relevant regions. These changes often reduce error rates by several percentage points on standard benchmarks while keeping parameter counts manageable.
The findings indicate that architectural innovation, rather than parameter tuning alone, drives most recent progress, yet deeper and wider networks raise computational cost and memory use. This limits deployment on resource-constrained devices and makes models harder to interpret. Applications have expanded beyond vision into speech, text, and video, but success still depends on large labeled datasets and careful hyper-parameter choices.
Leaders should therefore prioritize lightweight or quantized versions of these architectures for practical use, invest in hardware accelerators, and explore ensemble or attention-based refinements to improve robustness. Additional work is needed on automated hyper-parameter search, generative pre-training to reduce labeled-data requirements, and methods that preserve spatial relationships while lowering compute demands. The survey notes that results rest on published benchmarks that may not capture all real-world conditions and that many architectures remain sensitive to implementation details.
- Paper: ImageNet Classification with Deep Convolutional Neural Networks, Alex Krizhevsky et al. (2012). Read AlexNet first to understand the 2012 breakthrough that launched the wave of CNN architecture changes surveyed here.
- Paper: Deep Residual Learning for Image Recognition, Kaiming He et al. (2016). ResNet’s shortcut connections explain the central strategy the survey describes for making very deep CNNs trainable.
- Paper: Going Deeper with Convolutions, Christian Szegedy et al. (2015). The original Inception design introduces the multi-scale blocks that the survey treats as a major architectural advance.
- Paper: Very Deep Convolutional Networks for Large-Scale Image Recognition, Karen Simonyan et al. (2015). VGG provides a clear depth-and-small-filter baseline for understanding later architectural departures discussed in the survey.
- Paper: Densely Connected Convolutional Networks, Gao Huang et al. (2017). DenseNet’s layer-to-layer feature reuse supplies an important example of the reusable connectivity patterns covered by the survey.
- Paper: Recent advances in convolutional neural networks, Jiuxiang Gu et al. (2015). This earlier CNN survey establishes the architectural and training developments that the source revisits and updates.
- Paper: A ConvNet for the 2020s, Zhuang Liu et al. (2022). ConvNeXt carries the survey’s CNN design story into the transformer era by modernizing ResNet and testing whether convolutional models can remain competitive.
- Paper: ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders, Sanghyun Woo et al. (2023). ConvNeXt V2 continues that modernization by co-designing convolutional architecture with masked-autoencoder pretraining.
- Paper: InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions, Wenhai Wang et al. (2023). InternImage extends CNN scaling to foundation-model sizes by using deformable convolutions to adapt feature sampling.
- Paper: SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation, Meng-Hao Guo et al. (2022). SegNeXt advances the survey’s discussion of attention by replacing costly self-attention with multi-scale convolutional attention for segmentation.
- Paper: CoAtNet: Marrying Convolution and Attention for All Data Sizes, Zihang Dai et al. (2021). CoAtNet continues the CNN architecture trajectory by combining convolutional blocks with self-attention to balance data efficiency and scale.
