Built independently by an author, for readers. Read the story and support ChapterPal

keyword

residual connections

A residual connection is an architectural mechanism in artificial neural networks that allows information from an earlier layer to bypass one or more intermediate operations and be added directly to a subsequent layer output. Instead of forcing stacked layers to fit an entirely new target transformation, this design reformulates the learning objective so that the layers only need to learn a residual modification relative to the original input. By providing an uninterrupted shortcut for data and gradient signals across the network, residual connections prevent the degradation and vanishing gradient problems typically encountered in very deep architectures. Consequently, they facilitate smoother loss landscape optimization, accelerate model training, and serve as a foundational structural element across modern deep learning architectures, including deep convolutional networks and transformers.

9 items

Dilated Residual Networks

Dilated Residual Networks

Fisher Yu, Vladlen Koltun, Thomas Funkhouser

OrganizationsIntelPrinceton University

Why you should read this

Proposes dilated residual networks that retain high spatial feature resolution without increasing model complexity, introducing a degridding technique that improves performance across image classification, object localization, and semantic segmentation.

Convolutional networks for image classification progressively reduce resolution until the image is represented by tiny feature maps in which the spatial structure of the scene is no longer discernible. Such loss of spatial acuity can limit image classification accuracy and complicate the transfer of the model to downstream applications that require detailed scene understanding. These problems can be alleviated by dilation, which increases the resolution of output feature maps without reducing the receptive field of individual neurons. We show that dilated residual networks (DRNs) outperform their non-dilated counterparts in image classification without increasing the model's depth or complexity. We then study gridding artifacts introduced by dilation, develop an approach to removing these artifacts (`degridding'), and show that this further increases the performance of DRNs. In addition, we show that the accuracy advantage of DRNs is further magnified in downstream applications such as object localization and semantic segmentation.

Added

2026-09-18

Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning

Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning

Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, Alexander A. Alemi

OrganizationsGoogle

Why you should read this

Introduces the Inception-v4 and Inception-ResNet architectures, demonstrating that integrating residual connections accelerates training while activation scaling stabilizes wide networks to achieve top ImageNet classification accuracy.

Very deep convolutional networks have been central to the largest advances in image recognition performance in recent years. One example is the Inception architecture that has been shown to achieve very good performance at relatively low computational cost. Recently, the introduction of residual connections in conjunction with a more traditional architecture has yielded state-of-the-art performance in the 2015 ILSVRC challenge; its performance was similar to the latest generation Inception-v3 network. This raises the question of whether there are any benefit in combining the Inception architecture with residual connections. Here we give clear empirical evidence that training with residual connections accelerates the training of Inception networks significantly. There is also some evidence of residual Inception networks outperforming similarly expensive Inception networks without residual connections by a thin margin. We also present several new streamlined architectures for both residual and non-residual Inception networks. These variations improve the single-frame recognition performance on the ILSVRC 2012 classification task significantly. We further demonstrate how proper activation scaling stabilizes the training of very wide residual Inception networks. With an ensemble of three residual and one Inception-v4, we achieve 3.08 percent top-5 error on the test set of the ImageNet classification (CLS) challenge

Added

2026-09-05