Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Linear Transformers

Linear Transformers are variants of the Transformer deep learning architecture designed to compute self-attention in linear time and memory relative to input sequence length, overcoming the quadratic scaling bottleneck of standard attention mechanisms. Rather than computing full pairwise relationships across all tokens using standard softmax normalization, linear Transformers approximate or reformulate attention using kernel feature representations, flow networks, or linearized mappings. By utilizing the associative property of matrix multiplication, they reorder computations to combine keys and values before applying queries, which reduces both time and memory complexity from quadratic to linear. This mathematical formulation enables efficient processing of long sequences, allows autoregressive inference to operate recurrently with constant memory per step, and provides fundamental connections to recurrent neural networks and structured state-space models.

6 items

Flowformer: Linearizing Transformers with Conservation Flows

Flowformer: Linearizing Transformers with Conservation Flows

Haixu Wu, Jialong Wu, Jiehui Xu, Jianmin Wang, Mingsheng Long

OrganizationsTsinghua University

Why you should read this

Proposes Flowformer, a linear-complexity Transformer architecture based on flow network conservation theory that prevents attention degeneration without imposing task-specific inductive biases across vision, language, time series, and reinforcement learning domains.

Transformers based on the attention mechanism have achieved impressive success in various areas. However, the attention mechanism has a quadratic complexity, significantly impeding Transformers from dealing with numerous tokens and scaling up to bigger models. Previous methods mainly utilize the similarity decomposition and the associativity of matrix multiplication to devise linear-time attention mechanisms. They avoid degeneration of attention to a trivial distribution by reintroducing inductive biases such as the locality, thereby at the expense of model generality and expressiveness. In this paper, we linearize Transformers free from specific inductive biases based on the flow network theory. We cast attention as the information flow aggregated from the sources (values) to the sinks (results) through the learned flow capacities (attentions). Within this framework, we apply the property of flow conservation into attention and propose the Flow-Attention mechanism of linear complexity. By respectively conserving the incoming flow of sinks for source competition and the outgoing flow of sources for sink allocation, Flow-Attention inherently generates informative attentions without using specific inductive biases. Empowered by the Flow-Attention, Flowformer yields strong performance in linear time for wide areas, including long sequence, time series, vision, natural language, and reinforcement learning. The code and settings are available at this repository: https://github.com/thuml/Flowformer.

Added

2026-10-01

Transformers Learn In-Context by Gradient Descent

Transformers Learn In-Context by Gradient Descent

Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, Max Vladymyrov

OrganizationsETH ZurichGoogle

Why you should read this

Proves that transformer forward passes mechanistically implement gradient descent during in-context learning by establishing a direct equivalence between linear self-attention layers and gradient-based optimization steps.

At present, the mechanisms of in-context learning in Transformers are not well understood and remain mostly an intuition. In this paper, we suggest that training Transformers on auto-regressive objectives is closely related to gradient-based meta-learning formulations. We start by providing a simple weight construction that shows the equivalence of data transformations induced by 1) a single linear self-attention layer and by 2) gradient-descent (GD) on a regression loss. Motivated by that construction, we show empirically that when training self-attention-only Transformers on simple regression tasks either the models learned by GD and Transformers show great similarity or, remarkably, the weights found by optimization match the construction. Thus we show how trained Transformers become mesa-optimizers i.e. learn models by gradient descent in their forward pass. This allows us, at least in the domain of regression problems, to mechanistically understand the inner workings of in-context learning in optimized Transformers. Building on this insight, we furthermore identify how Transformers surpass the performance of plain gradient descent by learning an iterative curvature correction and learn linear models on deep data representations to solve non-linear regression tasks. Finally, we discuss intriguing parallels to a mechanism identified to be crucial for in-context learning termed induction-head (Olsson et al., 2022) and show how it could be understood as a specific case of in-context learning by gradient descent learning within Transformers.

Added

2026-09-28

In-context Convergence of Transformers

In-context Convergence of Transformers

Yu Huang, Yuan Cheng, Yingbin Liang

OrganizationsNational University of SingaporeThe Ohio State UniversityUniversity of Pennsylvania

Why you should read this

Establishes the first finite-time convergence guarantees and characterizes the stage-wise gradient descent dynamics of single-layer softmax transformers performing in-context learning on linear tasks across balanced and imbalanced feature distributions.

Transformers have recently revolutionized many domains in modern machine learning and one salient discovery is their remarkable in-context learning capability, where models can solve an unseen task by utilizing task-specific prompts without further parameters fine-tuning. This also inspired recent theoretical studies aiming to understand the in-context learning mechanism of transformers, which however focused only on linear transformers. In this work, we take the first step toward studying the learning dynamics of a one-layer transformer with softmax attention trained via gradient descent in order to in-context learn linear function classes. We consider a structured data model, where each token is randomly sampled from a set of feature vectors in either balanced or imbalanced fashion. For data with balanced features, we establish the finite-time convergence guarantee with near-zero prediction error by navigating our analysis over two phases of the training dynamics of the attention map. More notably, for data with imbalanced features, we show that the learning dynamics take a stage-wise convergence process, where the transformer first converges to a near-zero prediction error for the query tokens of dominant features, and then converges later to a near-zero error for query tokens of under-represented features, via one and four training phases. Our proof features new techniques for analyzing the competing strengths of two types of attention weights, the change of which determines different training phases.

Added

2026-09-26

Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François Fleuret

OrganizationsÉcole Polytechnique Fédérale de LausanneIdiap Research InstituteUniversity of GenevaUniversity of Washington

Why you should read this

Presents a novel linear attention mechanism that transforms computationally expensive quadratic-complexity Transformers into efficient linear-complexity recurrent neural networks, achieving up to 4000x faster autoregressive prediction for very long sequences without sacrificing performance.

Transformers achieve remarkable performance in several tasks but due to their quadratic complexity, with respect to the input's length, they are prohibitively slow for very long sequences. To address this limitation, we express the self-attention as a linear dot-product of kernel feature maps and make use of the associativity property of matrix products to reduce the complexity from O(N2)\mathcal{O}\left(N^2\right) to O(N)\mathcal{O}\left(N\right), where NN is the sequence length. We show that this formulation permits an iterative implementation that dramatically accelerates autoregressive transformers and reveals their relationship to recurrent neural networks. Our linear transformers achieve similar performance to vanilla transformers and they are up to 4000x faster on autoregressive prediction of very long sequences.

Added

2026-02-13

Creative Commons License