Built independently by an author, for readers. Read the story and support ChapterPal

keyword

deep residual networks

Deep residual networks, commonly known as ResNets, are a class of deep artificial neural networks that utilize shortcut or skip connections to enable the training of substantially deeper architectures. In conventional deep neural networks, increasing depth can cause optimization difficulties, such as vanishing or exploding gradients and training degradation. Deep residual networks address this issue by allowing inputs to bypass one or more parameterized layers via identity mappings, reforming the network layers to learn residual functions with reference to the layer inputs rather than unreferenced target functions. This structural design facilitates smooth gradient propagation throughout the entire model during backpropagation, resulting in stable optimization, faster convergence, and superior performance across a broad spectrum of computer vision and machine learning applications.

6 items

Activating More Pixels in Image Super-Resolution Transformer

Activating More Pixels in Image Super-Resolution Transformer

Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao, Chao Dong

OrganizationsChinese Academy of SciencesShanghai Artificial Intelligence LaboratoryShenzhen Institute of Advanced Technology, Chinese Academy of SciencesTencentUniversity of Macau

Why you should read this

Proposes a Hybrid Attention Transformer that combines channel and window self-attention with overlapping cross-attention to expand the spatial range of activated pixels, outperforming existing super-resolution methods by over 1 dB.

Transformer-based methods have shown impressive performance in low-level vision tasks, such as image super-resolution. However, we find that these networks can only utilize a limited spatial range of input information through attribution analysis. This implies that the potential of Transformer is still not fully exploited in existing networks. In order to activate more input pixels for better reconstruction, we propose a novel Hybrid Attention Transformer (HAT). It combines both channel attention and window-based self-attention schemes, thus making use of their complementary advantages of being able to utilize global statistics and strong local fitting capability. Moreover, to better aggregate the cross-window information, we introduce an overlapping cross-attention module to enhance the interaction between neighboring window features. In the training stage, we additionally adopt a same-task pre-training strategy to exploit the potential of the model for further improvement. Extensive experiments show the effectiveness of the proposed modules, and we further scale up the model to demonstrate that the performance of this task can be greatly improved. Our overall method significantly outperforms the state-of-the-art methods by more than 1dB.

Added

2026-10-04

Pre-Trained Image Processing Transformer

Pre-Trained Image Processing Transformer

Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, Wen Gao

OrganizationsHuaweiPeking UniversityPeng Cheng LaboratoryUniversity of Sydney

Why you should read this

Introduces the Image Processing Transformer (IPT), a unified pre-trained architecture trained on large-scale corrupted datasets with contrastive learning to outperform specialized models across multiple low-level vision tasks such as super-resolution, denoising, and deraining.

As the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectiveness over conventional methods. The big progress is mainly contributed to the representation ability of transformer and its variant architectures. In this paper, we study the low-level computer vision task (e.g., denoising, super-resolution and deraining) and develop a new pre-trained model, namely, image processing transformer (IPT). To maximally excavate the capability of transformer, we present to utilize the well-known ImageNet benchmark for generating a large amount of corrupted image pairs. The IPT model is trained on these images with multi-heads and multi-tails. In addition, the contrastive learning is introduced for well adapting to different image processing tasks. The pre-trained model can therefore efficiently employed on desired task after fine-tuning. With only one pre-trained model, IPT outperforms the current state-of-the-art methods on various low-level benchmarks. Code is available at this https URL and this https URL

Added

2026-09-15

Learning Structured Sparsity in Deep Neural Networks

Learning Structured Sparsity in Deep Neural Networks

Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, Hai Li

OrganizationsUniversity of Pittsburgh

Why you should read this

Introduces Structured Sparsity Learning, a regularization framework that prunes filters, channels, and entire layers to generate hardware-friendly deep neural networks that accelerate CPU and GPU inference while maintaining or improving accuracy.

High demand for computation resources severely hinders deployment of large-scale Deep Neural Networks (DNN) in resource constrained devices. In this work, we propose a Structured Sparsity Learning (SSL) method to regularize the structures (i.e., filters, channels, filter shapes, and layer depth) of DNNs. SSL can: (1) learn a compact structure from a bigger DNN to reduce computation cost; (2) obtain a hardware-friendly structured sparsity of DNN to efficiently accelerate the DNNs evaluation. Experimental results show that SSL achieves on average 5.1x and 3.1x speedups of convolutional layer computation of AlexNet against CPU and GPU, respectively, with off-the-shelf libraries. These speedups are about twice speedups of non-structured sparsity; (3) regularize the DNN structure to improve classification accuracy. The results show that for CIFAR-10, regularization on layer depth can reduce 20 layers of a Deep Residual Network (ResNet) to 18 layers while improve the accuracy from 91.25% to 92.60%, which is still slightly higher than that of original ResNet with 32 layers. For AlexNet, structure regularization by SSL also reduces the error by around ~1%. Open source code is in this https URL

Added

2026-09-14

Correlated initialization of deep residual networks

Correlated initialization of deep residual networks

Felix Benning, Ivan Nourdin, Giovanni Peccati

Why you should read this

Proves that layer-correlated weight initializations enable deep residual networks in the infinite-depth limit to converge to Young differential equations driven by Hermite processes, establishing correlation decay as a tunable hyperparameter that bridges the gap between deterministic and Brownian scaling regimes.

We study the large-depth behavior of residual networks whose weights are correlated across layers at initialization. Our results confirm and extend a conjecture of Marion et al. [2025], according to which correlated initializations should interpolate continuously between the Brownian stochastic differential equation arising from independent initialization and the ordinary differential equation arising from perfectly correlated initialization. When the initialization is obtained from the application of a feature function to a stationary Gaussian sequence with regularly varying correlation, we prove that there exists a unique critical scaling such that the infinite-depth limit is the solution of a Young differential equation driven by a Hermite process. Hermite processes reduce to the fractional Brownian motion if the feature function generating the initialization has Hermite rank one, which is the case for the identity function, for example. We show that the critical scaling and asymptotic limit are uniquely determined by the decay of correlations together with the Hermite rank of the feature function. Consequently, the correlation structure and Hermite rank of the initialization represent meaningful hyperparameters in the asymptotic regime. By contrast, under finite-variance iid initialization, the asymptotic driver is universally Brownian up to normalization regardless of the choice of distribution. Our proofs rely on a collection of novel results establishing a robust stability theory for Young differential equations in Banach spaces.

Added

2026-09-05

Creative Commons License
Identity Mappings in Deep Residual Networks

Identity Mappings in Deep Residual Networks

Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun

OrganizationsMicrosoft

Why you should read this

Explains why shuffling the order of operations in ResNet blocks—moving batch normalization and activation functions before the convolutions instead of after—creates cleaner pathways for information flow, enabling networks with over 1000 layers to train successfully when previous designs struggled beyond a few hundred layers.

Deep residual networks have emerged as a family of extremely deep architectures showing compelling accuracy and nice convergence behaviors. In this paper, we analyze the propagation formulations behind the residual building blocks, which suggest that the forward and backward signals can be directly propagated from one block to any other block, when using identity mappings as the skip connections and after-addition activation. A series of ablation experiments support the importance of these identity mappings. This motivates us to propose a new residual unit, which makes training easier and improves generalization. We report improved results using a 1001-layer ResNet on CIFAR-10 (4.62% error) and CIFAR-100, and a 200-layer ResNet on ImageNet. Code is available at: this https URL

Added

2026-02-21