Built independently by an author, for readers. Read the story and support ChapterPal

keyword

deep convolutional neural network

A deep convolutional neural network is a specialized class of artificial neural network designed to analyze grid-structured data, such as images and spatial or spatiotemporal arrays, through multiple stacked layers of mathematical convolutions, non-linear activations, and pooling operations. Unlike traditional fully connected neural networks, it applies learnable localized filters that slide across input dimensions, enforcing parameter sharing and translation invariance while drastically reducing the number of parameters to train. By organizing these operations into a deep hierarchy, the network automatically learns representations ranging from low-level details such as edges and textures in early layers to complex, high-level semantic features in deeper layers. These architectures typically conclude with fully connected or pooling layers to output predictions for tasks such as classification, object detection, regression, and image reconstruction.

6 items

Learning Traffic as Images: A Deep Convolutional Neural Network for Large-Scale Transportation Network Speed Prediction

Learning Traffic as Images: A Deep Convolutional Neural Network for Large-Scale Transportation Network Speed Prediction

Xiaolei Ma, Zhuang Dai, Zhengbing He, Jihui Na, Yong Wang, Yunpeng Wang

OrganizationsBeihang UniversityBeijing Jiaotong UniversityChongqing Jiaotong University

Why you should read this

Demonstrates that modeling spatiotemporal traffic dynamics as two-dimensional images enables convolutional neural networks to predict large-scale network speeds with a 42.9% accuracy improvement over traditional and recurrent deep learning models.

This paper proposes a convolutional neural network (CNN)-based method that learns traffic as images and predicts large-scale, network-wide traffic speed with a high accuracy. Spatiotemporal traffic dynamics are converted to images describing the time and space relations of traffic flow via a two-dimensional time-space matrix. A CNN is applied to the image following two consecutive steps: abstract traffic feature extraction and network-wide traffic speed prediction. The effectiveness of the proposed method is evaluated by taking two real-world transportation networks, the second ring road and north-east transportation network in Beijing, as examples, and comparing the method with four prevailing algorithms, namely, ordinary least squares, k-nearest neighbors, artificial neural network, and random forest, and three deep learning architectures, namely, stacked autoencoder, recurrent neural network, and long-short-term memory network. The results show that the proposed method outperforms other algorithms by an average accuracy improvement of 42.91% within an acceptable execution time. The CNN can train the model in a reasonable time and, thus, is suitable for large-scale transportation networks.

Added

2026-09-25

Deep Learning for Hyperspectral Image Classification: An Overview

Deep Learning for Hyperspectral Image Classification: An Overview

Shutao Li, Weiwei Song, Leyuan Fang, Yushi Chen, Pedram Ghamisi, Jón Atli Benediktsson

OrganizationsHarbin Institute of TechnologyHelmholtz Institute Freiberg for Resource TechnologyHelmholtz-Zentrum Dresden–RossendorfHunan UniversityKey Laboratory of Visual Perception and Artificial Intelligence of Hunan ProvinceUniversity of Iceland

Why you should read this

Presents a systematic framework categorizing deep learning approaches for hyperspectral image classification into spectral, spatial, and joint spectral-spatial networks while evaluating practical strategies to overcome limited training data constraints in remote sensing.

Hyperspectral image (HSI) classification has become a hot topic in the field of remote sensing. In general, the complex characteristics of hyperspectral data make the accurate classification of such data challenging for traditional machine learning methods. In addition, hyperspectral imaging often deals with an inherently nonlinear relation between the captured spectral information and the corresponding materials. In recent years, deep learning has been recognized as a powerful feature-extraction tool to effectively address nonlinear problems and widely used in a number of image processing tasks. Motivated by those successful applications, deep learning has also been introduced to classify HSIs and demonstrated good performance. This survey paper presents a systematic review of deep learning-based HSI classification literatures and compares several strategies for this topic. Specifically, we first summarize the main challenges of HSI classification which cannot be effectively overcome by traditional machine learning methods, and also introduce the advantages of deep learning to handle these problems. Then, we build a framework which divides the corresponding works into spectral-feature networks, spatial-feature networks, and spectral-spatial-feature networks to systematically review the recent achievements in deep learning-based HSI classification. In addition, considering the fact that available training samples in the remote sensing field are usually very limited and training deep networks require a large number of samples, we include some strategies to improve classification performance, which can provide some guidelines for future studies on this topic. Finally, several representative deep learning-based classification methods are conducted on real HSIs in our experiments.

Added

2026-09-24

DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving

DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving

Chenyi Chen, Ari Seff, Alain Kornhauser, Jianxiong Xiao

OrganizationsPrinceton University

Why you should read this

Proposes a direct perception framework for autonomous driving that bridges the gap between full scene parsing and end-to-end control by learning compact driving affordances that enable simple controllers to drive across diverse simulated and real-world environments.

Today, there are two major paradigms for vision-based autonomous driving systems: mediated perception approaches that parse an entire scene to make a driving decision, and behavior reflex approaches that directly map an input image to a driving action by a regressor. In this paper, we propose a third paradigm: a direct perception approach to estimate the affordance for driving. We propose to map an input image to a small number of key perception indicators that directly relate to the affordance of a road/traffic state for driving. Our representation provides a set of compact yet complete descriptions of the scene to enable a simple controller to drive autonomously. Falling in between the two extremes of mediated perception and behavior reflex, we argue that our direct perception representation provides the right level of abstraction. To demonstrate this, we train a deep Convolutional Neural Network using recording from 12 hours of human driving in a video game and show that our model can work well to drive a car in a very diverse set of virtual environments. We also train a model for car distance estimation on the KITTI dataset. Results show that our direct perception approach can generalize well to real driving images. Source code and data are available on our project website.

Added

2026-09-18

Deep Convolutional Neural Network for Inverse Problems in Imaging

Deep Convolutional Neural Network for Inverse Problems in Imaging

Kyong Hwan Jin, Michael T. McCann, Emmanuel Froustey, Michael Unser

OrganizationsCenter for Biomedical ImagingDassault AviationÉcole Polytechnique Fédérale de Lausanne

Why you should read this

Proposes a framework combining direct physical inversion with a residual convolutional neural network to solve ill-posed imaging inverse problems, achieving superior quality over standard iterative reconstruction while recovering sparse-view computed tomography images in sub-second speeds.

In this paper, we propose a novel deep convolutional neural network (CNN)-based algorithm for solving ill-posed inverse problems. Regularized iterative algorithms have emerged as the standard approach to ill-posed inverse problems in the past few decades. These methods produce excellent results, but can be challenging to deploy in practice due to factors including the high computational cost of the forward and adjoint operators and the difficulty of hyper parameter selection. The starting point of our work is the observation that unrolled iterative methods have the form of a CNN (filtering followed by point-wise non-linearity) when the normal operator (H*H, the adjoint of H times H) of the forward model is a convolution. Based on this observation, we propose using direct inversion followed by a CNN to solve normal-convolutional inverse problems. The direct inversion encapsulates the physical model of the system, but leads to artifacts when the problem is ill-posed; the CNN combines multiresolution decomposition and residual learning in order to learn to remove these artifacts while preserving image structure. We demonstrate the performance of the proposed network in sparse-view reconstruction (down to 50 views) on parallel beam X-ray computed tomography in synthetic phantoms as well as in real experimental sinograms. The proposed network outperforms total variation-regularized iterative reconstruction for the more realistic phantoms and requires less than a second to reconstruct a 512 x 512 image on GPU.

Added

2026-09-14

COVID-Net: a tailored deep convolutional neural network design for detection of COVID-19 cases from chest X-ray images

COVID-Net: a tailored deep convolutional neural network design for detection of COVID-19 cases from chest X-ray images

Linda Wang, Zhong Qiu Lin, Alexander Wong

OrganizationsDarwinAIUniversity of Waterloo

Why you should read this

Introduces COVID-Net, an open-source deep convolutional neural network architecture tailored for COVID-19 detection from chest radiographs, accompanied by the COVIDx benchmark dataset and an explainability framework to audit model predictions.

The Coronavirus Disease 2019 (COVID-19) pandemic continues to have a devastating effect on the health and well-being of the global population. A critical step in the fight against COVID-19 is effective screening of infected patients, with one of the key screening approaches being radiology examination using chest radiography. It was found in early studies that patients present abnormalities in chest radiography images that are characteristic of those infected with COVID-19. Motivated by this and inspired by the open source efforts of the research community, in this study we introduce COVID-Net, a deep convolutional neural network design tailored for the detection of COVID-19 cases from chest X-ray (CXR) images that is open source and available to the general public. To the best of the authors’ knowledge, COVID-Net is one of the first open source network designs for COVID-19 detection from CXR images at the time of initial release. We also introduce COVIDx, an open access benchmark dataset that we generated comprising of 13,975 CXR images across 13,870 patient patient cases, with the largest number of publicly available COVID-19 positive cases to the best of the authors’ knowledge. Furthermore, we investigate how COVID-Net makes predictions using an explainability method in an attempt to not only gain deeper insights into critical factors associated with COVID cases, which can aid clinicians in improved screening, but also audit COVID-Net in a responsible and transparent manner to validate that it is making decisions based on relevant information from the CXR images. By no means a production-ready solution, the hope is that the open access COVID-Net, along with the description on constructing the open source COVIDx dataset, will be leveraged and build upon by both researchers and citizen data scientists alike to accelerate the development of highly accurate yet practical deep learning solutions for detecting COVID-19 cases and accelerate treatment of those who need it the most.

Added

2026-09-14

DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition

DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition

Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, Trevor Darrell

OrganizationsInternational Computer Science InstituteUniversity of California Berkeley

Why you should read this

Demonstrates that deep convolutional activations trained on standard object classification generalize effectively across diverse visual recognition tasks, establishing pre-trained network representations as an effective off-the-shelf alternative to task-specific feature engineering.

We evaluate whether features extracted from the activation of a deep convolutional network trained in a fully supervised fashion on a large, fixed set of object recognition tasks can be re-purposed to novel generic tasks. Our generic tasks may differ significantly from the originally trained tasks and there may be insufficient labeled or unlabeled data to conventionally train or adapt a deep architecture to the new tasks. We investigate and visualize the semantic clustering of deep convolutional features with respect to a variety of such tasks, including scene recognition, domain adaptation, and fine-grained recognition challenges. We compare the efficacy of relying on various network levels to define a fixed feature, and report novel results that significantly outperform the state-of-the-art on several important vision challenges. We are releasing DeCAF, an open-source implementation of these deep convolutional activation features, along with all associated network parameters to enable vision researchers to be able to conduct experimentation with deep representations across a range of visual concept learning paradigms.

Added

2026-09-10