keyword
convolutional neural network
A convolutional neural network is a class of deep feedforward artificial neural networks designed to process and analyze grid-structured data such as images, audio spectrograms, and text sequences. Unlike standard fully connected neural networks, it applies learnable filters across local receptive fields in convolutional layers, enabling the automated extraction of hierarchical and translation-invariant features through parameter sharing. Architectures typically combine these convolutional layers with nonlinear activation functions, pooling layers that downsample spatial dimensions, and dense layers that generate outputs for tasks such as pattern recognition, classification, object detection, and feature representation learning.
8 items

End-to-End Representation Learning for Correlation Filter Based Tracking
Jack Valmadre, Luca Bertinetto, João F. Henriques, Andrea Vedaldi, Philip H. S. Torr
Why you should read this
Develops a differentiable correlation filter layer for deep networks, enabling end-to-end representation learning that allows lightweight models to achieve high-speed visual tracking with state-of-the-art accuracy.
The Correlation Filter is an algorithm that trains a linear template to discriminate between images and their translations. It is well suited to object tracking because its formulation in the Fourier domain provides a fast solution, enabling the detector to be re-trained once per frame. Previous works that use the Correlation Filter, however, have adopted features that were either manually designed or trained for a different task. This work is the first to overcome this limitation by interpreting the Correlation Filter learner, which has a closed-form solution, as a differentiable layer in a deep neural network. This enables learning deep features that are tightly coupled to the Correlation Filter. Experiments illustrate that our method has the important practical benefit of allowing lightweight architectures to achieve state-of-the-art performance at high framerates.
Added
2026-09-25

Stereo Matching by Training a Convolutional Neural Network to Compare Image Patches
Jure Zbontar, Yann LeCun
Why you should read this
Presents a deep learning approach that learns image patch similarity with convolutional neural networks to compute stereo matching costs, setting state-of-the-art depth estimation performance on the KITTI and Middlebury benchmarks.
We present a method for extracting depth information from a rectified image pair. Our approach focuses on the first stage of many stereo algorithms: the matching cost computation. We approach the problem by learning a similarity measure on small image patches using a convolutional neural network. Training is carried out in a supervised manner by constructing a binary classification data set with examples of similar and dissimilar pairs of patches. We examine two network architectures for this task: one tuned for speed, the other for accuracy. The output of the convolutional neural network is used to initialize the stereo matching cost. A series of post-processing steps follow: cross-based cost aggregation, semiglobal matching, a left-right consistency check, subpixel enhancement, a median filter, and a bilateral filter. We evaluate our method on the KITTI 2012, KITTI 2015, and Middlebury stereo data sets and show that it outperforms other approaches on all three data sets.
Added
2026-09-25

Unsupervised Pixel-Level Domain Adaptation with Generative Adversarial Networks
Konstantinos Bousmalis, Nathan Silberman, David Dohan, Dumitru Erhan, Dilip Krishnan
Why you should read this
Proposes an unsupervised generative adversarial framework that translates synthetic images into realistic target-domain imagery at the pixel level, significantly improving classifier generalization across domains and to unseen object classes without requiring target annotations.
Collecting well-annotated image datasets to train modern machine learning algorithms is prohibitively expensive for many tasks. One appealing alternative is rendering synthetic data where ground-truth annotations are generated automatically. Unfortunately, models trained purely on rendered images often fail to generalize to real images. To address this shortcoming, prior work introduced unsupervised domain adaptation algorithms that attempt to map representations between the two domains or learn to extract features that are domain-invariant. In this work, we present a new approach that learns, in an unsupervised manner, a transformation in the pixel space from one domain to the other. Our generative adversarial network (GAN)-based method adapts source-domain images to appear as if drawn from the target domain. Our approach not only produces plausible samples, but also outperforms the state-of-the-art on a number of unsupervised domain adaptation scenarios by large margins. Finally, we demonstrate that the adaptation process generalizes to object classes unseen during training.
Added
2026-09-24

Deep learning in remote sensing: a review
Xiao Xiang Zhu, Devis Tuia, Lichao Mou, Gui-Song Xia, Liangpei Zhang, Feng Xu, Friedrich Fraundorfer
Why you should read this
Surveys key advances and open challenges in applying deep learning to remote sensing data, providing practical resources and strategies to integrate Earth observation domain knowledge for tackling large-scale environmental problems.
Standing at the paradigm shift towards data-intensive science, machine learning techniques are becoming increasingly important. In particular, as a major breakthrough in the field, deep learning has proven as an extremely powerful tool in many fields. Shall we embrace deep learning as the key to all? Or, should we resist a 'black-box' solution? There are controversial opinions in the remote sensing community. In this article, we analyze the challenges of using deep learning for remote sensing data analysis, review the recent advances, and provide resources to make deep learning in remote sensing ridiculously simple to start with. More importantly, we advocate remote sensing scientists to bring their expertise into deep learning, and use it as an implicit general model to tackle unprecedented large-scale influential challenges, such as climate change and urbanization.
Added
2026-09-18

PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes
Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, Dieter Fox
Why you should read this
Presents PoseCNN, a convolutional network that estimates full 6D object poses from color images by decoupling 3D translation and quaternion rotation to handle symmetric objects and severe occlusions, while introducing the standard YCB-Video benchmark.
Estimating the 6D pose of known objects is important for robots to interact with the real world. The problem is challenging due to the variety of objects as well as the complexity of a scene caused by clutter and occlusions between objects. In this work, we introduce PoseCNN, a new Convolutional Neural Network for 6D object pose estimation. PoseCNN estimates the 3D translation of an object by localizing its center in the image and predicting its distance from the camera. The 3D rotation of the object is estimated by regressing to a quaternion representation. We also introduce a novel loss function that enables PoseCNN to handle symmetric objects. In addition, we contribute a large scale video dataset for 6D object pose estimation named the YCB-Video dataset. Our dataset provides accurate 6D poses of 21 objects from the YCB dataset observed in 92 videos with 133,827 frames. We conduct extensive experiments on our YCB-Video dataset and the OccludedLINEMOD dataset to show that PoseCNN is highly robust to occlusions, can handle symmetric objects, and provide accurate pose estimation using only color images as input. When using depth data to further refine the poses, our approach achieves state-of-the-art results on the challenging OccludedLINEMOD dataset. Our code and dataset are available at this https URL.
Added
2026-09-15

VoxCeleb2: Deep Speaker Recognition
Joon Son Chung, Arsha Nagrani, Andrew Zisserman
Why you should read this
Introduces the massive VoxCeleb2 dataset containing over one million utterances across six thousand speakers alongside deep convolutional neural network architectures that dramatically improve speaker recognition in unconstrained, noisy environments.
The objective of this paper is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media. Using a fully automated pipeline, we curate VoxCeleb2 which contains over a million utterances from over 6,000 speakers. This is several times larger than any publicly available speaker recognition dataset. Second, we develop and compare Convolutional Neural Network (CNN) models and training strategies that can effectively recognise identities from voice under various conditions. The models trained on the VoxCeleb2 dataset surpass the performance of previous works on a benchmark dataset by a significant margin.
Added
2026-09-14

A Convolutional Neural Network for Modelling Sentences
Nal Kalchbrenner, Edward Grefenstette, Phil Blunsom
Why you should read this
Introduces the Dynamic Convolutional Neural Network and dynamic k-max pooling, establishing a parser-free framework that captures both local and long-range semantic relationships across variable-length sentences in any language.
The ability to accurately represent sentences is central to language understanding. We describe a convolutional architecture dubbed the Dynamic Convolutional Neural Network (DCNN) that we adopt for the semantic modelling of sentences. The network uses Dynamic k-Max Pooling, a global pooling operation over linear sequences. The network handles input sentences of varying length and induces a feature graph over the sentence that is capable of explicitly capturing short and long-range relations. The network does not rely on a parse tree and is easily applicable to any language. We test the DCNN in four experiments: small scale binary and multi-class sentiment prediction, six-way question classification and Twitter sentiment prediction by distant supervision. The network achieves excellent performance in the first three tasks and a greater than 25% error reduction in the last task with respect to the strongest baseline.
Added
2026-09-11

Dive into Deep Learning (with PyTorch)
Aston Zhang, Zachary Lipton, Mu Li, Alexander Smola
Why you should read this
Combines rigorous mathematical foundations with hands-on, executable code examples in a free, interactive format that takes you from deep learning fundamentals to state-of-the-art techniques, making it ideal whether you're a student, researcher, or practitioner looking to truly understand and implement neural networks.
Source
https://d2l.ai/Added
2025-09-22

