Built independently by an author, for readers. Read the story and support ChapterPal

keyword

residual scaling

Residual scaling is a deep learning technique where the output of a residual branch or transformation block is multiplied by a scaling factor before being added back to the identity shortcut connection. In architectures that rely on residual connections, accumulating the full magnitude of transformed features across numerous consecutive blocks can cause signal variance, activation values, or gradients to grow uncontrollably as the network becomes deeper or wider. By attenuating the contribution of each residual block—typically using a constant fraction, a learnable parameter, or a coefficient scaled inversely with network depth—residual scaling stabilizes signal propagation and gradient flow during training. This mechanism prevents numerical instability and optimization collapse, facilitating the successful training of very deep neural networks, wide residual models, and architectures where standard normalization layers are removed.

4 items

ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks

ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks

Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Chen Change Loy, Yu Qiao, Xiaoou Tang

OrganizationsNanyang Technological UniversityShenzhen Institute of Advanced Technology, Chinese Academy of SciencesThe Chinese University of Hong KongUniversity of Chinese Academy of Sciences

Why you should read this

Proposes a generative adversarial network framework using Residual-in-Residual Dense Blocks and relativistic loss to produce realistic textures while eliminating visual artifacts in single-image super-resolution.

The Super-Resolution Generative Adversarial Network (SRGAN) is a seminal work that is capable of generating realistic textures during single image super-resolution. However, the hallucinated details are often accompanied with unpleasant artifacts. To further enhance the visual quality, we thoroughly study three key components of SRGAN - network architecture, adversarial loss and perceptual loss, and improve each of them to derive an Enhanced SRGAN (ESRGAN). In particular, we introduce the Residual-in-Residual Dense Block (RRDB) without batch normalization as the basic network building unit. Moreover, we borrow the idea from relativistic GAN to let the discriminator predict relative realness instead of the absolute value. Finally, we improve the perceptual loss by using the features before activation, which could provide stronger supervision for brightness consistency and texture recovery. Benefiting from these improvements, the proposed ESRGAN achieves consistently better visual quality with more realistic and natural textures than SRGAN and won the first place in the PIRM2018-SR Challenge. The code is available at this https URL .

Added

2026-09-13

Correlated initialization of deep residual networks

Correlated initialization of deep residual networks

Felix Benning, Ivan Nourdin, Giovanni Peccati

Why you should read this

Proves that layer-correlated weight initializations enable deep residual networks in the infinite-depth limit to converge to Young differential equations driven by Hermite processes, establishing correlation decay as a tunable hyperparameter that bridges the gap between deterministic and Brownian scaling regimes.

We study the large-depth behavior of residual networks whose weights are correlated across layers at initialization. Our results confirm and extend a conjecture of Marion et al. [2025], according to which correlated initializations should interpolate continuously between the Brownian stochastic differential equation arising from independent initialization and the ordinary differential equation arising from perfectly correlated initialization. When the initialization is obtained from the application of a feature function to a stationary Gaussian sequence with regularly varying correlation, we prove that there exists a unique critical scaling such that the infinite-depth limit is the solution of a Young differential equation driven by a Hermite process. Hermite processes reduce to the fractional Brownian motion if the feature function generating the initialization has Hermite rank one, which is the case for the identity function, for example. We show that the critical scaling and asymptotic limit are uniquely determined by the decay of correlations together with the Hermite rank of the feature function. Consequently, the correlation structure and Hermite rank of the initialization represent meaningful hyperparameters in the asymptotic regime. By contrast, under finite-variance iid initialization, the asymptotic driver is universally Brownian up to normalization regardless of the choice of distribution. Our proofs rely on a collection of novel results establishing a robust stability theory for Young differential equations in Banach spaces.

Added

2026-09-05

Creative Commons License
DeepLoop: Depth Scaling for Looped Transformers

DeepLoop: Depth Scaling for Looped Transformers

Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang

OrganizationsPrinceton UniversityUniversity of California, Los Angeles

Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transformer, each residual branch receives and applies its own parameter update, whereas in a looped Transformer one shared update aggregates gradients from repeated visits and is read back by those same visits in the next linearized forward pass. We formalize this tied-depth effect through a first-order perturbation bound controlled by a visit-alignment coefficient κR\kappa_R. The bound recovers the DeepNorm exponent when visits decorrelate, but in the conservative aligned regime it requires the exponent to increase from 1/41/4 to 1/21/2 as loop count grows at fixed physical depth. The resulting method, \textbf{DeepLoop}, keeps the Post-LN DeepNorm architecture and sets α=(2N)1/2\alpha=(2N)^{1/2} and β=(8N)−1/2\beta=(8N)^{-1/2} for unrolled depth NN. On GPT-style looped language models at GPT-2 small and GPT-2 medium scale, DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated. These results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count.

Added

2026-07-31