keyword
residual connections
A residual connection is an architectural mechanism in artificial neural networks that allows information from an earlier layer to bypass one or more intermediate operations and be added directly to a subsequent layer output. Instead of forcing stacked layers to fit an entirely new target transformation, this design reformulates the learning objective so that the layers only need to learn a residual modification relative to the original input. By providing an uninterrupted shortcut for data and gradient signals across the network, residual connections prevent the degradation and vanishing gradient problems typically encountered in very deep architectures. Consequently, they facilitate smoother loss landscape optimization, accelerate model training, and serve as a foundational structural element across modern deep learning architectures, including deep convolutional networks and transformers.
9 items

LayerNorm Induces Recency Bias in Transformer Decoders
Junu Kim, Xiao Liu, Zheng-Wen Lin, Lei Ji, Yeyun Gong, Edward Choi
Why you should read this
Demonstrates that LayerNorm turns the early-token attention bias of causal transformers into recency bias, resolving a fundamental contradiction in transformer behavior and informing positional encoding design.
Causal self-attention provides positional information to Transformer decoders. Prior work has shown that stacks of causal self-attention layers alone induce a positional bias in attention scores toward earlier tokens. However, this differs from the bias toward later tokens typically observed in Transformer decoders, known as recency bias. We address this discrepancy by analyzing the interaction between causal self-attention and other architectural components. We show that stacked causal self-attention layers combined with LayerNorm induce recency bias. Furthermore, we examine the effects of residual connections and the distribution of input token embeddings on this bias. Our results provide new theoretical insights into how positional information interacts with architectural components and suggest directions for improving positional encoding strategies.
Added
2026-10-04

Quantifying Attention Flow in Transformers
Samira Abnar, Willem Zuidema
Why you should read this
Proposes attention rollout and attention flow to track information propagation across Transformer layers, providing reliable token-importance explanations that correlate significantly better with gradient and ablation baselines than raw attention weights.
In the Transformer model, "self-attention" combines information from attended embeddings into the representation of the focal embedding in the next layer. Thus, across layers of the Transformer, information originating from different tokens gets increasingly mixed. This makes attention weights unreliable as explanations probes. In this paper, we consider the problem of quantifying this flow of information through self-attention. We propose two methods for approximating the attention to input tokens given attention weights, attention rollout and attention flow, as post hoc methods when we use attention weights as the relative relevance of the input tokens. We show that these methods give complementary views on the flow of information, and compared to raw attention, both yield higher correlations with importance scores of input tokens obtained using an ablation method and input gradients.
Added
2026-09-25

Transformer Feed-Forward Layers Are Key-Value Memories
Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy
Why you should read this
Shows that transformer feed-forward layers function as key-value memories, revealing how the majority of network parameters store human-interpretable textual patterns and directly predict next-token distributions.
Feed-forward layers constitute two-thirds of a transformer model's parameters, yet their role in the network remains under-explored. We show that feed-forward layers in transformer-based language models operate as key-value memories, where each key correlates with textual patterns in the training examples, and each value induces a distribution over the output vocabulary. Our experiments show that the learned patterns are human-interpretable, and that lower layers tend to capture shallow patterns, while upper layers learn more semantic ones. The values complement the keys' input patterns by inducing output distributions that concentrate probability mass on tokens likely to appear immediately after each pattern, particularly in the upper layers. Finally, we demonstrate that the output of a feed-forward layer is a composition of its memories, which is subsequently refined throughout the model's layers via residual connections to produce the final output distribution.
Added
2026-09-24

Dilated Residual Networks
Fisher Yu, Vladlen Koltun, Thomas Funkhouser
Why you should read this
Proposes dilated residual networks that retain high spatial feature resolution without increasing model complexity, introducing a degridding technique that improves performance across image classification, object localization, and semantic segmentation.
Convolutional networks for image classification progressively reduce resolution until the image is represented by tiny feature maps in which the spatial structure of the scene is no longer discernible. Such loss of spatial acuity can limit image classification accuracy and complicate the transfer of the model to downstream applications that require detailed scene understanding. These problems can be alleviated by dilation, which increases the resolution of output feature maps without reducing the receptive field of individual neurons. We show that dilated residual networks (DRNs) outperform their non-dilated counterparts in image classification without increasing the model's depth or complexity. We then study gridding artifacts introduced by dilation, develop an approach to removing these artifacts (`degridding'), and show that this further increases the performance of DRNs. In addition, we show that the accuracy advantage of DRNs is further magnified in downstream applications such as object localization and semantic segmentation.
Added
2026-09-18

The Deep Ritz Method: A Deep Learning-Based Numerical Algorithm for Solving Variational Problems
Weinan E, Bing Yu
Why you should read this
Proposes the Deep Ritz method, a numerical framework that solves variational problems and high-dimensional partial differential equations by training neural networks to minimize energy formulations using stochastic gradient descent.
We propose a deep learning based method, the Deep Ritz Method, for numerically solving variational problems, particularly the ones that arise from partial differential equations. The Deep Ritz method is naturally nonlinear, naturally adaptive and has the potential to work in rather high dimensions. The framework is quite simple and fits well with the stochastic gradient descent method used in deep learning. We illustrate the method on several problems including some eigenvalue problems.
Added
2026-09-18

Visualizing the Loss Landscape of Neural Nets
Hao Li, Zheng Xu, Gavin Taylor, T. Goldstein
Why you should read this
Introduces a filter-normalized visualization technique that reveals how network architectures like skip connections smooth high-dimensional loss surfaces to improve training stability and generalization.
Neural network training relies on our ability to find "good" minimizers of highly non-convex loss functions. It is well-known that certain network architecture designs (e.g., skip connections) produce loss functions that train easier, and well-chosen training parameters (batch size, learning rate, optimizer) produce minimizers that generalize better. However, the reasons for these differences, and their effects on the underlying loss landscape, are not well understood. In this paper, we explore the structure of neural loss functions, and the effect of loss landscapes on generalization, using a range of visualization methods. First, we introduce a simple "filter normalization" method that helps us visualize loss function curvature and make meaningful side-by-side comparisons between loss functions. Then, using a variety of visualizations, we explore how network architecture affects the loss landscape, and how training parameters affect the shape of minimizers.
Added
2026-09-14

Xception: Deep Learning with Depthwise Separable Convolutions
François Chollet
Why you should read this
Introduces the Xception architecture, demonstrating that replacing Inception modules with depthwise separable convolutions achieves superior image classification performance through more efficient parameter utilization.
We present an interpretation of Inception modules in convolutional neural networks as being an intermediate step in-between regular convolution and the depthwise separable convolution operation (a depthwise convolution followed by a pointwise convolution). In this light, a depthwise separable convolution can be understood as an Inception module with a maximally large number of towers. This observation leads us to propose a novel deep convolutional neural network architecture inspired by Inception, where Inception modules have been replaced with depthwise separable convolutions. We show that this architecture, dubbed Xception, slightly outperforms Inception V3 on the ImageNet dataset (which Inception V3 was designed for), and significantly outperforms Inception V3 on a larger image classification dataset comprising 350 million images and 17,000 classes. Since the Xception architecture has the same number of parameters as Inception V3, the performance gains are not due to increased capacity but rather to a more efficient use of model parameters.
Added
2026-09-05

Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, Alexander A. Alemi
Why you should read this
Introduces the Inception-v4 and Inception-ResNet architectures, demonstrating that integrating residual connections accelerates training while activation scaling stabilizes wide networks to achieve top ImageNet classification accuracy.
Very deep convolutional networks have been central to the largest advances in image recognition performance in recent years. One example is the Inception architecture that has been shown to achieve very good performance at relatively low computational cost. Recently, the introduction of residual connections in conjunction with a more traditional architecture has yielded state-of-the-art performance in the 2015 ILSVRC challenge; its performance was similar to the latest generation Inception-v3 network. This raises the question of whether there are any benefit in combining the Inception architecture with residual connections. Here we give clear empirical evidence that training with residual connections accelerates the training of Inception networks significantly. There is also some evidence of residual Inception networks outperforming similarly expensive Inception networks without residual connections by a thin margin. We also present several new streamlined architectures for both residual and non-residual Inception networks. These variations improve the single-frame recognition performance on the ILSVRC 2012 classification task significantly. We further demonstrate how proper activation scaling stabilizes the training of very wide residual Inception networks. With an ensemble of three residual and one Inception-v4, we achieve 3.08 percent top-5 error on the test set of the ImageNet classification (CLS) challenge
Added
2026-09-05

The Annotated Transformer
Alexander M. Rush
Why you should read this
A complete, hands-on tutorial for understanding and building the Transformer - the foundational technology behind modern AI systems like ChatGPT and Google Translate - by walking through the actual code line by line rather than just abstract theory.
An annotated version of the paper "Attention is All You Need" in the form of a line-by-line implementation.
Added
2026-02-21
