Built independently by an author, for readers. Read the story and support ChapterPal

keyword

convolutional neural network

A convolutional neural network is a class of deep feedforward artificial neural networks designed to process and analyze grid-structured data such as images, audio spectrograms, and text sequences. Unlike standard fully connected neural networks, it applies learnable filters across local receptive fields in convolutional layers, enabling the automated extraction of hierarchical and translation-invariant features through parameter sharing. Architectures typically combine these convolutional layers with nonlinear activation functions, pooling layers that downsample spatial dimensions, and dense layers that generate outputs for tasks such as pattern recognition, classification, object detection, and feature representation learning.

8 items

Stereo Matching by Training a Convolutional Neural Network to Compare Image Patches

Stereo Matching by Training a Convolutional Neural Network to Compare Image Patches

Jure Zbontar, Yann LeCun

OrganizationsFaculty of Computer and Information ScienceMetaNew York UniversityUniversity of Ljubljana

Why you should read this

Presents a deep learning approach that learns image patch similarity with convolutional neural networks to compute stereo matching costs, setting state-of-the-art depth estimation performance on the KITTI and Middlebury benchmarks.

We present a method for extracting depth information from a rectified image pair. Our approach focuses on the first stage of many stereo algorithms: the matching cost computation. We approach the problem by learning a similarity measure on small image patches using a convolutional neural network. Training is carried out in a supervised manner by constructing a binary classification data set with examples of similar and dissimilar pairs of patches. We examine two network architectures for this task: one tuned for speed, the other for accuracy. The output of the convolutional neural network is used to initialize the stereo matching cost. A series of post-processing steps follow: cross-based cost aggregation, semiglobal matching, a left-right consistency check, subpixel enhancement, a median filter, and a bilateral filter. We evaluate our method on the KITTI 2012, KITTI 2015, and Middlebury stereo data sets and show that it outperforms other approaches on all three data sets.

Added

2026-09-25

Unsupervised Pixel-Level Domain Adaptation with Generative Adversarial Networks

Unsupervised Pixel-Level Domain Adaptation with Generative Adversarial Networks

Konstantinos Bousmalis, Nathan Silberman, David Dohan, Dumitru Erhan, Dilip Krishnan

OrganizationsGoogle

Why you should read this

Proposes an unsupervised generative adversarial framework that translates synthetic images into realistic target-domain imagery at the pixel level, significantly improving classifier generalization across domains and to unseen object classes without requiring target annotations.

Collecting well-annotated image datasets to train modern machine learning algorithms is prohibitively expensive for many tasks. One appealing alternative is rendering synthetic data where ground-truth annotations are generated automatically. Unfortunately, models trained purely on rendered images often fail to generalize to real images. To address this shortcoming, prior work introduced unsupervised domain adaptation algorithms that attempt to map representations between the two domains or learn to extract features that are domain-invariant. In this work, we present a new approach that learns, in an unsupervised manner, a transformation in the pixel space from one domain to the other. Our generative adversarial network (GAN)-based method adapts source-domain images to appear as if drawn from the target domain. Our approach not only produces plausible samples, but also outperforms the state-of-the-art on a number of unsupervised domain adaptation scenarios by large margins. Finally, we demonstrate that the adaptation process generalizes to object classes unseen during training.

Added

2026-09-24

Deep learning in remote sensing: a review

Deep learning in remote sensing: a review

Xiao Xiang Zhu, Devis Tuia, Lichao Mou, Gui-Song Xia, Liangpei Zhang, Feng Xu, Friedrich Fraundorfer

OrganizationsFudan UniversityGerman Aerospace Center (DLR)Graz University of TechnologyTechnical University of MunichUniversity of ZurichWageningen University & ResearchWuhan University

Why you should read this

Surveys key advances and open challenges in applying deep learning to remote sensing data, providing practical resources and strategies to integrate Earth observation domain knowledge for tackling large-scale environmental problems.

Standing at the paradigm shift towards data-intensive science, machine learning techniques are becoming increasingly important. In particular, as a major breakthrough in the field, deep learning has proven as an extremely powerful tool in many fields. Shall we embrace deep learning as the key to all? Or, should we resist a 'black-box' solution? There are controversial opinions in the remote sensing community. In this article, we analyze the challenges of using deep learning for remote sensing data analysis, review the recent advances, and provide resources to make deep learning in remote sensing ridiculously simple to start with. More importantly, we advocate remote sensing scientists to bring their expertise into deep learning, and use it as an implicit general model to tackle unprecedented large-scale influential challenges, such as climate change and urbanization.

Added

2026-09-18

PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes

PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes

Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, Dieter Fox

OrganizationsCarnegie Mellon UniversityNVIDIAUniversity of Washington

Why you should read this

Presents PoseCNN, a convolutional network that estimates full 6D object poses from color images by decoupling 3D translation and quaternion rotation to handle symmetric objects and severe occlusions, while introducing the standard YCB-Video benchmark.

Estimating the 6D pose of known objects is important for robots to interact with the real world. The problem is challenging due to the variety of objects as well as the complexity of a scene caused by clutter and occlusions between objects. In this work, we introduce PoseCNN, a new Convolutional Neural Network for 6D object pose estimation. PoseCNN estimates the 3D translation of an object by localizing its center in the image and predicting its distance from the camera. The 3D rotation of the object is estimated by regressing to a quaternion representation. We also introduce a novel loss function that enables PoseCNN to handle symmetric objects. In addition, we contribute a large scale video dataset for 6D object pose estimation named the YCB-Video dataset. Our dataset provides accurate 6D poses of 21 objects from the YCB dataset observed in 92 videos with 133,827 frames. We conduct extensive experiments on our YCB-Video dataset and the OccludedLINEMOD dataset to show that PoseCNN is highly robust to occlusions, can handle symmetric objects, and provide accurate pose estimation using only color images as input. When using depth data to further refine the poses, our approach achieves state-of-the-art results on the challenging OccludedLINEMOD dataset. Our code and dataset are available at this https URL.

Added

2026-09-15