Built independently by an author, for readers. Read the story and support ChapterPal

keyword

deep convolutional neural networks

A deep convolutional neural network is a class of artificial neural networks composed of multiple processing layers designed to automatically and adaptively learn hierarchical feature representations directly from grid-structured data such as images, audio spectrograms, and text. Unlike standard fully connected networks, these architectures utilize convolutional layers that apply learnable filters across local receptive fields, allowing them to capture spatial and temporal correlations while dramatically reducing the number of parameters through weight sharing and local connectivity. As data flows through successive convolutional, nonlinear activation, and pooling or subsampling layers, the network progressively transforms low-level patterns like edges into rich, high-level semantic representations before delivering predictions through classification or regression layers. This ability to capture translation-invariant features and scale effectively with large datasets makes deep convolutional neural networks a foundational technology for computer vision, natural language processing, and automated pattern recognition.

13 items

Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts

Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts

Cícero Nogueira dos Santos, Maíra Gatti

OrganizationsBrazilian Research LabIBM

Why you should read this

Proposes CharSCNN, a deep convolutional neural network that jointly extracts character- and sentence-level representations to improve sentiment classification performance on short texts across movie review and Twitter benchmarks.

Sentiment analysis of short texts such as single sentences and Twitter messages is challenging because of the limited contextual information that they normally contain. Effectively solving this task requires strategies that combine the small text content with prior knowledge and use more than just bag-of-words. In this work we propose a new deep convolutional neural network that exploits from character- to sentence-level information to perform sentiment analysis of short texts. We apply our approach for two corpora of two different domains: the Stanford Sentiment Treebank (SSTb), which contains sentences from movie reviews; and the Stanford Twitter Sentiment corpus (STS), which contains Twitter messages. For the SSTb corpus, our approach achieves state-of-the-art results for single sentence sentiment prediction in both binary positive/negative classification, with 85.7% accuracy, and fine-grained classification, with 48.3% accuracy. For the STS corpus, our approach achieves a sentiment prediction accuracy of 86.4%.

Added

2026-09-25

Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification

Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification

Justin Salamon, Juan Pablo Bello

OrganizationsNew York University

Why you should read this

Demonstrates that combining deep convolutional neural networks with systematic audio data augmentation overcomes data scarcity to achieve state-of-the-art accuracy in environmental sound classification.

The ability of deep convolutional neural networks (CNN) to learn discriminative spectro-temporal patterns makes them well suited to environmental sound classification. However, the relative scarcity of labeled data has impeded the exploitation of this family of high-capacity models. This study has two primary contributions: first, we propose a deep convolutional neural network architecture for environmental sound classification. Second, we propose the use of audio data augmentation for overcoming the problem of data scarcity and explore the influence of different augmentations on the performance of the proposed CNN architecture. Combined with data augmentation, the proposed model produces state-of-the-art results for environmental sound classification. We show that the improved performance stems from the combination of a deep, high-capacity model and an augmented training set: this combination outperforms both the proposed CNN without augmentation and a "shallow" dictionary learning model with augmentation. Finally, we examine the influence of each augmentation on the model's classification accuracy for each class, and observe that the accuracy for each class is influenced differently by each augmentation, suggesting that the performance of the model could be improved further by applying class-conditional data augmentation.

Added

2026-09-25

Automatic detection of coronavirus disease (COVID-19) using X-ray images and deep convolutional neural networks

Automatic detection of coronavirus disease (COVID-19) using X-ray images and deep convolutional neural networks

Ali Narin, Ceren Kaya, Ziynet Pamuk

OrganizationsZonguldak Bülent Ecevit University

Why you should read this

Establishes a highly accurate automated COVID-19 screening approach by evaluating five pre-trained convolutional neural network architectures on chest X-rays, identifying ResNet50 as the top performer with up to 99.7% accuracy in differentiating coronavirus from other forms of pneumonia.

The 2019 novel coronavirus disease (COVID-19), with a starting point in China, has spread rapidly among people living in other countries, and is approaching approximately 34,986,502 cases worldwide according to the statistics of European Centre for Disease Prevention and Control. There are a limited number of COVID-19 test kits available in hospitals due to the increasing cases daily. Therefore, it is necessary to implement an automatic detection system as a quick alternative diagnosis option to prevent COVID-19 spreading among people. In this study, five pre-trained convolutional neural network based models (ResNet50, ResNet101, ResNet152, InceptionV3 and Inception-ResNetV2) have been proposed for the detection of coronavirus pneumonia infected patient using chest X-ray radiographs. We have implemented three different binary classifications with four classes (COVID-19, normal (healthy), viral pneumonia and bacterial pneumonia) by using 5-fold cross validation. Considering the performance results obtained, it has seen that the pre-trained ResNet50 model provides the highest classification performance (96.1% accuracy for Dataset-1, 99.5% accuracy for Dataset-2 and 99.7% accuracy for Dataset-3) among other four used models.

Added

2026-09-18

Deep Learning Face Representation by Joint Identification-Verification

Deep Learning Face Representation by Joint Identification-Verification

Yi Sun, Yuheng Chen, Xiaogang Wang, Xiaoou Tang

OrganizationsChinese Academy of SciencesShenzhen Institute of Advanced Technology, Chinese Academy of SciencesThe Chinese University of Hong Kong

Why you should read this

Introduces DeepID2, a deep learning approach that jointly optimizes identification and verification supervisory signals to simultaneously minimize intra-identity variation and enlarge inter-identity differences, reducing error rates on the LFW benchmark by 67%.

The key challenge of face recognition is to develop effective feature representations for reducing intra-personal variations while enlarging inter-personal differences. In this paper, we show that it can be well solved with deep learning and using both face identification and verification signals as supervision. The Deep IDentification-verification features (DeepID2) are learned with carefully designed deep convolutional networks. The face identification task increases the inter-personal variations by drawing DeepID2 extracted from different identities apart, while the face verification task reduces the intra-personal variations by pulling DeepID2 extracted from the same identity together, both of which are essential to face recognition. The learned DeepID2 features can be well generalized to new identities unseen in the training data. On the challenging LFW dataset, 99.15% face verification accuracy is achieved. Compared with the best deep learning result on LFW, the error rate has been significantly reduced by 67%.

Added

2026-09-15

A survey of the recent architectures of deep convolutional neural networks

A survey of the recent architectures of deep convolutional neural networks

Asifullah Khan, Anabia Sohail, Umme Zahoora, Aqsa Saeed Qureshi

OrganizationsCenter for Mathematical SciencesDCISDeep Learning LabPakistan Institute of Engineering and Applied SciencesPattern Recognition Lab

Why you should read this

Classifies recent deep convolutional neural network architectures into seven structural categories—including spatial exploitation, depth, multi-path routing, and attention mechanisms—to explain the design principles driving modern computer vision systems.

Deep Convolutional Neural Network (CNN) is a special type of Neural Networks, which has shown exemplary performance on several competitions related to Computer Vision and Image Processing. Some of the exciting application areas of CNN include Image Classification and Segmentation, Object Detection, Video Processing, Natural Language Processing, and Speech Recognition. The powerful learning ability of deep CNN is primarily due to the use of multiple feature extraction stages that can automatically learn representations from the data. The availability of a large amount of data and improvement in the hardware technology has accelerated the research in CNNs, and recently interesting deep CNN architectures have been reported. Several inspiring ideas to bring advancements in CNNs have been explored, such as the use of different activation and loss functions, parameter optimization, regularization, and architectural innovations. However, the significant improvement in the representational capacity of the deep CNN is achieved through architectural innovations. Notably, the ideas of exploiting spatial and channel information, depth and width of architecture, and multi-path information processing have gained substantial attention. Similarly, the idea of using a block of layers as a structural unit is also gaining popularity. This survey thus focuses on the intrinsic taxonomy present in the recently reported deep CNN architectures and, consequently, classifies the recent innovations in CNN architectures into seven different categories. These seven categories are based on spatial exploitation, depth, multi-path, width, feature-map exploitation, channel boosting, and attention. Additionally, the elementary understanding of CNN components, current challenges, and applications of CNN are also provided.

Added

2026-09-14

Learning Deep Features for Scene Recognition using Places Database

Learning Deep Features for Scene Recognition using Places Database

Bolei Zhou, Àgata Lapedriza, Jianxiong Xiao, A. Torralba, A. Oliva

OrganizationsMassachusetts Institute of TechnologyPrinceton UniversityUniversitat Oberta de Catalunya

Why you should read this

Introduces the 7-million-image Places database and demonstrates that convolutional neural networks trained on scene-centric data develop distinct visual representations that significantly outperform object-trained models on scene recognition tasks.

Scene recognition is one of the hallmark tasks of computer vision, allowing definition of a context for object recognition. Whereas the tremendous recent progress in object recognition tasks is due to the availability of large datasets like ImageNet and the rise of Convolutional Neural Networks (CNNs) for learning high-level features, performance at scene recognition has not attained the same level of success. This may be because current deep features trained from ImageNet are not competitive enough for such tasks. Here, we introduce a new scene-centric database called Places with over 7 million labeled pictures of scenes. We propose new methods to compare the density and diversity of image datasets and show that Places is as dense as other scene datasets and has more diversity. Using CNN, we learn deep features for scene recognition tasks, and establish new state-of-the-art results on several scene-centric datasets. A visualization of the CNN layers' responses allows us to show differences in the internal representations of object-centric and scene-centric networks.

Added

2026-09-12

Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning

Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning

Hoo-Chang Shin, Holger R. Roth, Mingchen Gao, Le Lu, Ziyue Xu, Isabella Nogues, Jianhua Yao, Daniel Mollura, Ronald M. Summers

OrganizationsCenter for Infectious Disease ImagingClinical Image Processing ServiceImaging Biomarkers and Computer-Aided Diagnosis LaboratoryNational Institutes of HealthRadiology and Imaging Sciences Department

Why you should read this

Demonstrates the practical impact of network architecture, dataset scale, and ImageNet transfer learning on medical computer-aided detection, establishing clear performance benchmarks across CT diagnostic tasks.

Remarkable progress has been made in image recognition, primarily due to the availability of large-scale annotated datasets and the revival of deep CNN. CNNs enable learning data-driven, highly representative, layered hierarchical image features from sufficient training data. However, obtaining datasets as comprehensively annotated as ImageNet in the medical imaging domain remains a challenge. There are currently three major techniques that successfully employ CNNs to medical image classification: training the CNN from scratch, using off-the-shelf pre-trained CNN features, and conducting unsupervised CNN pre-training with supervised fine-tuning. Another effective method is transfer learning, i.e., fine-tuning CNN models pre-trained from natural image dataset to medical image tasks. In this paper, we exploit three important, but previously understudied factors of employing deep convolutional neural networks to computer-aided detection problems. We first explore and evaluate different CNN architectures. The studied models contain 5 thousand to 160 million parameters, and vary in numbers of layers. We then evaluate the influence of dataset scale and spatial image context on performance. Finally, we examine when and why transfer learning from pre-trained ImageNet (via fine-tuning) can be useful. We study two specific computer-aided detection (CADe) problems, namely thoraco-abdominal lymph node (LN) detection and interstitial lung disease (ILD) classification. We achieve the state-of-the-art performance on the mediastinal LN detection, with 85% sensitivity at 3 false positive per patient, and report the first five-fold cross-validation classification results on predicting axial CT slices with ILD categories. Our extensive empirical evaluation, CNN model analysis and valuable insights can be extended to the design of high performance CAD systems for other medical imaging tasks.

Added

2026-09-10

Learning Transferable Features with Deep Adaptation Networks

Learning Transferable Features with Deep Adaptation Networks

Mingsheng Long, Yue Cao, Jianmin Wang, Michael I. Jordan

OrganizationsTsinghua UniversityUniversity of California Berkeley

Why you should read this

Introduces Deep Adaptation Networks, an architecture that aligns task-specific feature distributions across domains using multi-kernel mean embedding matching to reduce dataset bias with linear computational scaling.

Recent studies reveal that a deep neural network can learn transferable features which generalize well to novel tasks for domain adaptation. However, as deep features eventually transition from general to specific along the network, the feature transferability drops significantly in higher layers with increasing domain discrepancy. Hence, it is important to formally reduce the dataset bias and enhance the transferability in task-specific layers. In this paper, we propose a new Deep Adaptation Network (DAN) architecture, which generalizes deep convolutional neural network to the domain adaptation scenario. In DAN, hidden representations of all task-specific layers are embedded in a reproducing kernel Hilbert space where the mean embeddings of different domain distributions can be explicitly matched. The domain discrepancy is further reduced using an optimal multi-kernel selection method for mean embedding matching. DAN can learn transferable features with statistical guarantees, and can scale linearly by unbiased estimate of kernel embedding. Extensive empirical evidence shows that the proposed architecture yields state-of-the-art image classification error rates on standard domain adaptation benchmarks.

Added

2026-09-09

ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks

ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks

Qilong Wang, Banggu Wu, Pengfei Zhu, Peihua Li, Wangmeng Zuo, Qinghua Hu

OrganizationsDalian University of TechnologyHarbin Institute of TechnologyTianjin University

Why you should read this

Introduces an ultra-lightweight channel attention module that avoids dimensionality reduction through adaptive 1D convolutions, boosting deep CNN accuracy across vision benchmarks with negligible computational overhead.

Recently, channel attention mechanism has demonstrated to offer great potential in improving the performance of deep convolutional neural networks (CNNs). However, most existing methods dedicate to developing more sophisticated attention modules for achieving better performance, which inevitably increase model complexity. To overcome the paradox of performance and complexity trade-off, this paper proposes an Efficient Channel Attention (ECA) module, which only involves a handful of parameters while bringing clear performance gain. By dissecting the channel attention module in SENet, we empirically show avoiding dimensionality reduction is important for learning channel attention, and appropriate cross-channel interaction can preserve performance while significantly decreasing model complexity. Therefore, we propose a local cross-channel interaction strategy without dimensionality reduction, which can be efficiently implemented via 1D1D convolution. Furthermore, we develop a method to adaptively select kernel size of 1D1D convolution, determining coverage of local cross-channel interaction. The proposed ECA module is efficient yet effective, e.g., the parameters and computations of our modules against backbone of ResNet50 are 80 vs. 24.37M and 4.7e-4 GFLOPs vs. 3.86 GFLOPs, respectively, and the performance boost is more than 2% in terms of Top-1 accuracy. We extensively evaluate our ECA module on image classification, object detection and instance segmentation with backbones of ResNets and MobileNetV2. The experimental results show our module is more efficient while performing favorably against its counterparts.

Added

2026-09-09

ImageNet Classification with Deep Convolutional Neural Networks

ImageNet Classification with Deep Convolutional Neural Networks

Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton

OrganizationsUniversity of Toronto

Why you should read this

The Big Bang of modern computer vision. This paper proved deep CNNs could dominate large-scale image recognition and provided the foundational blueprint that launched the deep learning revolution in vision. Start here to understand the fundamental shift from traditional to deep learning approaches.

We trained a large, deep convolutional neural network to classify the 1.3 million high-resolution images in the LSVRC-2010 ImageNet training set into the 1000 different classes. On the test data, we achieved top-1 and top-5 error rates of 39.7\% and 18.9\% which is considerably better than the previous state-of-the-art results. The neural network, which has 60 million parameters and 500,000 neurons, consists of five convolutional layers, some of which are followed by max-pooling layers, and two globally connected layers with a final 1000-way softmax. To make training faster, we used non-saturating neurons and a very efficient GPU implementation of convolutional nets. To reduce overfitting in the globally connected layers we employed a new regularization method that proved to be very effective.

Added

2025-08-28

License

Published with permission