Built independently by an author, for readers. Read the story and support ChapterPal

keyword

convolutional network

A convolutional network is a class of deep artificial neural networks designed primarily to process and analyze structured grid-like data, such as images, video, and audio waveforms. The architecture is defined by convolutional layers that apply learnable filters over local receptive fields of the input, allowing the network to detect spatial or temporal features while sharing parameters across the entire input domain. These layers are commonly combined with non-linear activation functions, pooling or strided operations to reduce spatial dimensions, and fully connected layers to generate final predictions or representations. By exploiting translation invariance and hierarchical feature learning, convolutional networks automatically extract low-level patterns such as edges and textures in early layers and compose them into complex, high-level features in deeper layers, making them widely used in computer vision, audio processing, and pattern recognition.

8 items

Evaluating the Visualization of What a Deep Neural Network Has Learned

Evaluating the Visualization of What a Deep Neural Network Has Learned

Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Bach, Klaus-Robert Müller

OrganizationsFraunhofer Heinrich Hertz InstituteKorea UniversitySingapore University of Technology and DesignTechnische Universität Berlin

Why you should read this

Introduces a region-perturbation framework for objectively evaluating neural network explanation heatmaps and demonstrates that Layer-wise Relevance Propagation outperforms sensitivity analysis and deconvolution methods across standard image benchmarks.

Deep Neural Networks (DNNs) have demonstrated impressive performance in complex machine learning tasks such as image classification or speech recognition. However, due to their multi-layer nonlinear structure, they are not transparent, i.e., it is hard to grasp what makes them arrive at a particular classification or recognition decision given a new unseen data sample. Recently, several approaches have been proposed enabling one to understand and interpret the reasoning embodied in a DNN for a single test image. These methods quantify the ''importance'' of individual pixels wrt the classification decision and allow a visualization in terms of a heatmap in pixel/input space. While the usefulness of heatmaps can be judged subjectively by a human, an objective quality measure is missing. In this paper we present a general methodology based on region perturbation for evaluating ordered collections of pixels such as heatmaps. We compare heatmaps computed by three different methods on the SUN397, ILSVRC2012 and MIT Places data sets. Our main result is that the recently proposed Layer-wise Relevance Propagation (LRP) algorithm qualitatively and quantitatively provides a better explanation of what made a DNN arrive at a particular classification decision than the sensitivity-based approach or the deconvolution method. We provide theoretical arguments to explain this result and discuss its practical implications. Finally, we investigate the use of heatmaps for unsupervised assessment of neural network performance.

Added

2026-09-25

Striving for Simplicity: The All Convolutional Net

Striving for Simplicity: The All Convolutional Net

Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, Martin Riedmiller

OrganizationsUniversity of Freiburg

Why you should read this

Demonstrates that max-pooling can be completely replaced by strided convolutions to build high-performing all-convolutional networks, while introducing guided backpropagation for visualizing learned features.

Most modern convolutional neural networks (CNNs) used for object recognition are built using the same principles: Alternating convolution and max-pooling layers followed by a small number of fully connected layers. We re-evaluate the state of the art for object recognition from small images with convolutional networks, questioning the necessity of different components in the pipeline. We find that max-pooling can simply be replaced by a convolutional layer with increased stride without loss in accuracy on several image recognition benchmarks. Following this finding -- and building on other recent work for finding simple network structures -- we propose a new architecture that consists solely of convolutional layers and yields competitive or state of the art performance on several object recognition datasets (CIFAR-10, CIFAR-100, ImageNet). To analyze the network we introduce a new variant of the "deconvolution approach" for visualizing features learned by CNNs, which can be applied to a broader range of network structures than existing approaches.

Added

2026-09-10

Dimensionality Reduction by Learning an Invariant Mapping

Dimensionality Reduction by Learning an Invariant Mapping

Raia Hadsell, Sumit Chopra, Yann LeCun

OrganizationsNew York University

Why you should read this

Proposes Dimensionality Reduction by Learning an Invariant Mapping (DrLIM), a contrastive learning framework that maps high-dimensional data into low-dimensional spaces while generalizing to unseen samples and remaining invariant to transformations without requiring predefined distance metrics.

Dimensionality reduction involves mapping a set of high dimensional input points onto a low dimensional manifold so that “similar” points in input space are mapped to nearby points on the manifold. Most existing techniques for solving the problem suffer from two drawbacks. First, most of them depend on a meaningful and computable distance metric in input space. Second, they do not compute a “function” that can accurately map new input samples whose relationship to the training data is unknown. We present a method - called Dimensionality Reduction by Learning an Invariant Mapping (DrLIM) - for learning a globally coherent non-linear function that maps the data evenly to the output manifold. The learning relies solely on neighborhood relationships and does not require any distance measure in the input space. The method can learn mappings that are invariant to certain transformations of the inputs, as is demonstrated with a number of experiments. Comparisons are made to other techniques, in particular LLE.

Added

2026-09-09

WaveNet: A Generative Model for Raw Audio

WaveNet: A Generative Model for Raw Audio

Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alexander Graves, Nal Kalchbrenner, Andrew Senior, Koray Kavukcuoglu

OrganizationsGoogle

Why you should read this

Demonstrates that autoregressive principles extend perfectly to high-frequency continuous temporal data like raw audio through dilated causal convolutions.

This paper introduces WaveNet, a deep neural network for generating raw audio waveforms. The model is fully probabilistic and autoregressive, with the predictive distribution for each audio sample conditioned on all previous ones; nonetheless we show that it can be efficiently trained on data with tens of thousands of samples per second of audio. When applied to text-to-speech, it yields state-of-the-art performance, with human listeners rating it as significantly more natural sounding than the best parametric and concatenative systems for both English and Mandarin. A single WaveNet can capture the characteristics of many different speakers with equal fidelity, and can switch between them by conditioning on the speaker identity. When trained to model music, we find that it generates novel and often highly realistic musical fragments. We also show that it can be employed as a discriminative model, returning promising results for phoneme recognition.

Added

2026-02-21