keyword
feed-forward network
A feed-forward network is a type of artificial neural network in which information travels in only one forward direction, moving from the input layer through any intermediate hidden layers directly to the output layer without forming cycles or feedback loops. Unlike recurrent architectures, which maintain internal memory through recurring state transitions, feed-forward networks process each given input independently without feedback connections. Typically composed of fully connected linear layers interspersed with non-linear activation functions, these networks enable models to approximate complex mathematical functions and map inputs to higher-dimensional feature spaces. In modern deep learning systems, such as transformer architectures, feed-forward sub-layers operate position-wise on internal representations to transform features and enrich token embeddings after attention operations.
3 items

Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Mor Geva, Avi Caciularu, Kevin Ro Wang, Yoav Goldberg
Why you should read this
Reveals how transformer feed-forward layers construct predictions by promoting human-interpretable concepts directly in the vocabulary space, enabling practical techniques to cut GPT-2 toxicity by half and save twenty percent of inference computation through early exiting.
Transformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood. In this work, we make a substantial step towards unveiling this underlying prediction process, by reverse-engineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models. We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution. Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable. We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average.
Added
2026-09-29

An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention
Yehjin Shin, Jeongwhan Choi, Hyowon Wi, Noseong Park
Why you should read this
Reveals that self-attention in sequential recommendation behaves as a low-pass filter causing representation oversmoothing, and introduces BSARec, a Fourier transform-based architecture that integrates high-frequency signals to capture abrupt short-term user preferences alongside long-term interests.
Sequential recommendation (SR) models based on Transformers have achieved remarkable successes. The self-attention mechanism of Transformers for computer vision and natural language processing suffers from the oversmoothing problem, i.e., hidden representations becoming similar to tokens. In the SR domain, we, for the first time, show that the same problem occurs. We present pioneering investigations that reveal the low-pass filtering nature of self-attention in the SR, which causes oversmoothing. To this end, we propose a novel method called Beyond Self-Attention for Sequential Recommendation (BSARec), which leverages the Fourier transform to i) inject an inductive bias by considering fine-grained sequential patterns and ii) integrate low and high-frequency information to mitigate oversmoothing. Our discovery shows significant advancements in the SR domain and is expected to bridge the gap for existing Transformer-based SR models. We test our proposed approach through extensive experiments on 6 benchmark datasets. The experimental results demonstrate that our model outperforms 7 baseline methods in terms of recommendation performance. Our code is available at https://github.com/yehjin-shin/BSARec.
Added
2026-09-26

The Annotated Transformer
Alexander M. Rush
Why you should read this
A complete, hands-on tutorial for understanding and building the Transformer - the foundational technology behind modern AI systems like ChatGPT and Google Translate - by walking through the actual code line by line rather than just abstract theory.
An annotated version of the paper "Attention is All You Need" in the form of a line-by-line implementation.
Added
2026-02-21
