Built independently by an author, for readers. Read the story and support ChapterPal

keyword

deep convolutional networks

Deep convolutional networks, also known as deep convolutional neural networks or ConvNets, are a class of artificial neural networks characterized by multiple processing layers designed to analyze structured grid data such as images. These architectures utilize learnable convolutional filters that slide across input data to capture local spatial and temporal patterns, followed by non-linear activation functions and pooling or normalization layers that reduce dimensionality and provide invariance to spatial shifts. By cascading many parameterized layers, deep convolutional networks automatically learn hierarchical representations, transitioning from low-level visual features such as edges and textures in early layers to complex, task-specific semantic concepts in deeper layers. They serve as foundational models across computer vision and pattern recognition for applications such as image classification, object detection, semantic segmentation, and super-resolution.

14 items

Deep Visual Geo-localization Benchmark

Deep Visual Geo-localization Benchmark

Gabriele Moreno Berton, Riccardo Mereu, Gabriele Trivigno, Carlo Masone, Gabriela Csurka, Torsten Sattler, Barbara Caputo

OrganizationsConsorzio Interuniversitario Nazionale per l'InformaticaCzech Technical University in PragueNaver Labs EuropePolitecnico di Torino

Why you should read this

Establishes a modular visual geo-localization benchmarking framework and conducts extensive evaluations across pipeline components, offering practical guidelines that balance retrieval accuracy with real-world computational and memory costs.

In this paper, we propose a new open-source benchmark-ing framework for Visual Geo-localization (VG) that allows to build, train, and test a wide range of commonly used ar-chitectures, with the flexibility to change individual compo-nents of a geo-localization pipeline. The purpose of this framework is twofold: i) gaining insights into how differ-ent components and design choices in a VG pipeline im-pact the final results, both in terms of performance (re-call@N metric) and system requirements (such as execu-tion time and memory consumption); ii) establish a system-atic evaluation protocol for comparing different methods. Using the proposed framework, we perform a large suite of experiments which provide criteria for choosing back-bone, aggregation and negative mining depending on the use-case and requirements. We also assess the impact of engineering techniques like pre/post-processing, data aug-mentation and image resizing, showing that better perfor-mance can be obtained through somewhat simple proce-dures: for example, downscaling the images' resolution to 80% can lead to similar results with a 36% savings in ex-traction time and dataset storage requirement. Code and trained models are available at dataset storage require-ment. https://deep-vg-bench.herokuapp.com/.

Added

2026-09-26

Sketching without Worrying: Noise-Tolerant Sketch-Based Image Retrieval

Sketching without Worrying: Noise-Tolerant Sketch-Based Image Retrieval

Ayan Kumar Bhunia, Subhadeep Koley, Abdullah Faiz Ur Rahman Khilji, Aneeshan Sain, Pinaki Nath Chowdhury, Tao Xiang, Yi-Zhe Song

OrganizationsiFlyTek-Surrey Joint Research Centre on Artificial IntelligenceUniversity of Surrey

Why you should read this

Proposes a reinforcement learning-based stroke subset selector that filters out noisy or detrimental strokes from amateur drawings to boost fine-grained sketch-based image retrieval accuracy without retraining underlying retrieval models.

Sketching enables many exciting applications, notably, image retrieval. The fear-to-sketch problem (i.e., “I can’t sketch”) has however proven to be fatal for its widespread adoption. This paper tackles this “fear” head on, and for the first time, proposes an auxiliary module for existing retrieval models that predominantly lets the users sketch without having to worry. We first conducted a pilot study that revealed the secret lies in the existence of noisy strokes, but not so much of the “I can’t sketch”. We consequently design a stroke subset selector that detects noisy strokes, leaving only those which make a positive contribution towards successful retrieval. Our Reinforcement Learning based formulation quantifies the importance of each stroke present in a given subset, based on the extent to which that stroke contributes to retrieval. When combined with pre-trained retrieval models as a pre-processing module, we achieve a significant gain of 8%-10% over standard baselines and in turn report new state-of-the-art performance. Last but not least, we demonstrate the selector once trained, can also be used in a plug-and-play manner to empower various sketch applications in ways that were not previously possible.

Added

2026-09-26

Do Better ImageNet Models Transfer Better?

Do Better ImageNet Models Transfer Better?

Simon Kornblith, Jonathon Shlens, Quoc V. Le

OrganizationsGoogle

Why you should read this

Establishes a strong correlation between ImageNet classification accuracy and downstream transfer performance across sixteen neural architectures, while showing that common regularization techniques can unexpectedly degrade learned feature representations.

Transfer learning is a cornerstone of computer vision, yet little work has been done to evaluate the relationship between architecture and transfer. An implicit hypothesis in modern computer vision research is that models that perform better on ImageNet necessarily perform better on other vision tasks. However, this hypothesis has never been systematically tested. Here, we compare the performance of 16 classification networks on 12 image classification datasets. We find that, when networks are used as fixed feature extractors or fine-tuned, there is a strong correlation between ImageNet accuracy and transfer accuracy (r=0.99r = 0.99 and 0.960.96, respectively). In the former setting, we find that this relationship is very sensitive to the way in which networks are trained on ImageNet; many common forms of regularization slightly improve ImageNet accuracy but yield penultimate layer features that are much worse for transfer learning. Additionally, we find that, on two small fine-grained image classification datasets, pretraining on ImageNet provides minimal benefits, indicating the learned features from ImageNet do not transfer well to fine-grained tasks. Together, our results show that ImageNet architectures generalize well across datasets, but ImageNet features are less general than previously suggested.

Added

2026-09-24

RepVGG: Making VGG-style ConvNets Great Again

RepVGG: Making VGG-style ConvNets Great Again

Xiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han, Guiguang Ding, Jian Sun

OrganizationsAberystwyth UniversityMegvii TechnologyThe Hong Kong University of Science and TechnologyTsinghua University

Why you should read this

Proposes a structural re-parameterization technique that trains a multi-branch network and converts it into a simple, high-speed 3x3 convolutional backbone for inference, achieving over 80% ImageNet accuracy while running significantly faster than ResNet architectures.

We present a simple but powerful architecture of convolutional neural network, which has a VGG-like inference-time body composed of nothing but a stack of 3x3 convolution and ReLU, while the training-time model has a multi-branch topology. Such decoupling of the training-time and inference-time architecture is realized by a structural re-parameterization technique so that the model is named RepVGG. On ImageNet, RepVGG reaches over 80% top-1 accuracy, which is the first time for a plain model, to the best of our knowledge. On NVIDIA 1080Ti GPU, RepVGG models run 83% faster than ResNet-50 or 101% faster than ResNet-101 with higher accuracy and show favorable accuracy-speed trade-off compared to the state-of-the-art models like EfficientNet and RegNet. The code and trained models are available at this https URL.

Added

2026-09-14

Res2Net: A New Multi-Scale Backbone Architecture

Res2Net: A New Multi-Scale Backbone Architecture

Shanghua Gao, Ming-Ming Cheng, Kai Zhao, Xinyu Zhang, Ming-Hsuan Yang, Philip H. S. Torr

OrganizationsNankai UniversityTKLNDSTUniversity of California, MercedUniversity of Oxford

Why you should read this

Introduces Res2Net, a plug-and-play convolutional block with internal hierarchical residual connections that captures granular multi-scale features and improves backbone network performance across image classification, object detection, and salient object detection.

Representing features at multiple scales is of great importance for numerous vision tasks. Recent advances in backbone convolutional neural networks (CNNs) continually demonstrate stronger multi-scale representation ability, leading to consistent performance gains on a wide range of applications. However, most existing methods represent the multi-scale features in a layer-wise manner. In this paper, we propose a novel building block for CNNs, namely Res2Net, by constructing hierarchical residual-like connections within one single residual block. The Res2Net represents multi-scale features at a granular level and increases the range of receptive fields for each network layer. The proposed Res2Net block can be plugged into the state-of-the-art backbone CNN models, e.g., ResNet, ResNeXt, and DLA. We evaluate the Res2Net block on all these models and demonstrate consistent performance gains over baseline models on widely-used datasets, e.g., CIFAR-100 and ImageNet. Further ablation studies and experimental results on representative computer vision tasks, i.e., object detection, class activation mapping, and salient object detection, further verify the superiority of the Res2Net over the state-of-the-art baseline methods. The source code and trained models are available on this https URL.

Added

2026-09-12

Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks

Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks

Aditya Chattopadhyay, Anirban Sarkar, Prantik Howlader, Vineeth N Balasubramanian

OrganizationsCiscoIndian Institute of Technology Hyderabad

Why you should read this

Introduces Grad-CAM++, a generalized visual explanation method that improves CNN interpretability by weighting positive partial derivatives to deliver precise object localization and accurately distinguish multiple instances of the same class.

Over the last decade, Convolutional Neural Network (CNN) models have been highly successful in solving complex vision problems. However, these deep models are perceived as "black box" methods considering the lack of understanding of their internal functioning. There has been a significant recent interest in developing explainable deep learning models, and this paper is an effort in this direction. Building on a recently proposed method called Grad-CAM, we propose a generalized method called Grad-CAM++ that can provide better visual explanations of CNN model predictions, in terms of better object localization as well as explaining occurrences of multiple object instances in a single image, when compared to state-of-the-art. We provide a mathematical derivation for the proposed method, which uses a weighted combination of the positive partial derivatives of the last convolutional layer feature maps with respect to a specific class score as weights to generate a visual explanation for the corresponding class label. Our extensive experiments and evaluations, both subjective and objective, on standard datasets showed that Grad-CAM++ provides promising human-interpretable visual explanations for a given CNN architecture across multiple tasks including classification, image caption generation and 3D action recognition; as well as in new settings such as knowledge distillation.

Added

2026-09-11

Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition

Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition

Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun

OrganizationsMicrosoftUniversity of Science and Technology of ChinaXi'an Jiaotong University

Why you should read this

Introduces spatial pyramid pooling to eliminate fixed-size input constraints in convolutional neural networks, enabling arbitrary-scale image classification and speeding up object detection by up to a hundredfold over R-CNN.

Existing deep convolutional neural networks (CNNs) require a fixed-size (e.g., 224x224) input image. This requirement is "artificial" and may reduce the recognition accuracy for the images or sub-images of an arbitrary size/scale. In this work, we equip the networks with another pooling strategy, "spatial pyramid pooling", to eliminate the above requirement. The new network structure, called SPP-net, can generate a fixed-length representation regardless of image size/scale. Pyramid pooling is also robust to object deformations. With these advantages, SPP-net should in general improve all CNN-based image classification methods. On the ImageNet 2012 dataset, we demonstrate that SPP-net boosts the accuracy of a variety of CNN architectures despite their different designs. On the Pascal VOC 2007 and Caltech101 datasets, SPP-net achieves state-of-the-art classification results using a single full-image representation and no fine-tuning. The power of SPP-net is also significant in object detection. Using SPP-net, we compute the feature maps from the entire image only once, and then pool features in arbitrary regions (sub-images) to generate fixed-length representations for training the detectors. This method avoids repeatedly computing the convolutional features. In processing test images, our method is 24-102x faster than the R-CNN method, while achieving better or comparable accuracy on Pascal VOC 2007. In ImageNet Large Scale Visual Recognition Challenge (ILSVRC) 2014, our methods rank #2 in object detection and #3 in image classification among all 38 teams. This manuscript also introduces the improvement made for this competition.

Added

2026-09-06

Very Deep Convolutional Networks for Large-Scale Image Recognition

Very Deep Convolutional Networks for Large-Scale Image Recognition

Karen Simonyan, Andrew Zisserman

OrganizationsUniversity of Oxford

Why you should read this

Demonstrates that systematically increasing neural network depth using small 3×3 filters achieves state-of-the-art image recognition performance, establishing a simple yet powerful design principle that became foundational for modern deep learning architectures.

Abstract:In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respectively. We also show that our representations generalise well to other datasets, where they achieve state-of-the-art results. We have made our two best-performing ConvNet models publicly available to facilitate further research on the use of deep visual representations in computer vision.

Added

2026-05-27

License

Published with permission