Built independently by an author, for readers. Read the story and support ChapterPal

keyword

dot-product attention

Dot-product attention is an attention mechanism in neural networks that measures the relevance between different elements in a sequence by computing the mathematical dot product of their vector representations. In this mechanism, input representations are projected into query, key, and value vectors. The network calculates similarity scores by taking the dot product of each query with all keys, frequently scaling the resulting values by the square root of the key vector dimension to prevent vanishing gradients during training. Applying a softmax function across these scaled scores yields a normalized probability distribution of attention weights, which is then used to compute a weighted sum of the corresponding value vectors. Serving as the primary foundational operation in Transformer architectures, dot-product attention enables efficient, highly parallelized modeling of long-range dependencies and contextual relationships across text, audio, and visual data.

14 items

Attention as a Guide for Simultaneous Speech Translation

Attention as a Guide for Simultaneous Speech Translation

Sara Papi, Matteo Negri, Marco Turchi

OrganizationsFondazione Bruno KesslerIndependent ResearcherUniversity of Trento

Why you should read this

Introduces EDATT, an adaptive policy that applies cross-attention patterns during inference to enable offline-trained speech translation models to perform simultaneous translation with significantly lower latency and higher BLEU scores without dedicated streaming retraining.

In simultaneous speech translation (SimulST), effective policies that determine when to write partial translations are crucial to reach high output quality with low latency. Towards this objective, we propose EDATT (Encoder-Decoder Attention), an adaptive policy that exploits the attention patterns between audio source and target textual translation to guide an offline-trained ST model during simultaneous inference. EDATT exploits the attention scores modeling the audio-translation relation to decide whether to emit a partial hypothesis or wait for more audio input. This is done under the assumption that, if attention is focused towards the most recently received speech segments, the information they provide can be insufficient to generate the hypothesis (indicating that the system has to wait for additional audio input). Results on en→{de, es} show that EDATT yields better results compared to the SimulST state of the art, with gains respectively up to 7 and 4 BLEU points for the two languages, and with a reduction in computational-aware latency up to 1.4s and 0.7s compared to existing SimulST policies applied to offline-trained models.

Added

2026-10-05

ESPnet: End-to-End Speech Processing Toolkit

ESPnet: End-to-End Speech Processing Toolkit

Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, Tsubasa Ochiai

OrganizationsDoshisha UniversityJohns Hopkins UniversityMitsubishi Electric Research LaboratoriesNagoya UniversityNTT CorporationPaderborn UniversityPreferred Networks, Inc.Retrieva, Inc.Waseda University

Why you should read this

Presents ESPnet, an open-source platform that bridges PyTorch-based dynamic neural models with Kaldi-style data pipelines to enable reproducible, high-accuracy end-to-end speech recognition and processing across standard benchmarks.

This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a main deep learning engine. ESPnet also follows the Kaldi ASR toolkit style for data processing, feature extraction/format, and recipes to provide a complete setup for speech recognition and other speech processing experiments. This paper explains a major architecture of this software platform, several important functionalities, which differentiate ESPnet from other open source ASR toolkits, and experimental results with major ASR benchmarks.

Added

2026-09-24

How Attentive are Graph Attention Networks?

How Attentive are Graph Attention Networks?

Shaked Brody, Uri Alon, Eran Yahav

OrganizationsCarnegie Mellon UniversityTechnion – Israel Institute of Technology

Why you should read this

Reveals a fundamental expressiveness flaw in standard Graph Attention Networks and introduces GATv2, a dynamic attention variant that resolves the limitation to achieve superior accuracy across graph benchmarks without increasing parameter cost.

Graph Attention Networks (GATs) are one of the most popular GNN architectures and are considered as the state-of-the-art architecture for representation learning with graphs. In GAT, every node attends to its neighbors given its own representation as the query. However, in this paper we show that GAT computes a very limited kind of attention: the ranking of the attention scores is unconditioned on the query node. We formally define this restricted kind of attention as static attention and distinguish it from a strictly more expressive dynamic attention. Because GATs use a static attention mechanism, there are simple graph problems that GAT cannot express: in a controlled problem, we show that static attention hinders GAT from even fitting the training data. To remove this limitation, we introduce a simple fix by modifying the order of operations and propose GATv2: a dynamic graph attention variant that is strictly more expressive than GAT. We perform an extensive evaluation and show that GATv2 outperforms GAT across 11 OGB and other benchmarks while we match their parametric costs. Our code is available at this https URL . GATv2 is available as part of the PyTorch Geometric library, the Deep Graph Library, and the TensorFlow GNN library.

Added

2026-09-17

Swin Transformer V2: Scaling Up Capacity and Resolution

Swin Transformer V2: Scaling Up Capacity and Resolution

Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, Baining Guo

OrganizationsMicrosoft

Why you should read this

Introduces stability, resolution-scaling, and self-supervised pre-training techniques that scale Swin Transformers to 3 billion parameters and high-resolution inputs, achieving state-of-the-art performance across major vision benchmarks using significantly less labeled data.

Large-scale NLP models have been shown to significantly improve the performance on language tasks with no signs of saturation. They also demonstrate amazing few-shot capabilities like that of human beings. This paper aims to explore large-scale models in computer vision. We tackle three major issues in training and application of large vision models, including training instability, resolution gaps between pre-training and fine-tuning, and hunger on labelled data. Three main techniques are proposed: 1) a residual-post-norm method combined with cosine attention to improve training stability; 2) A log-spaced continuous position bias method to effectively transfer models pre-trained using low-resolution images to downstream tasks with high-resolution inputs; 3) A self-supervised pre-training method, SimMIM, to reduce the needs of vast labeled images. Through these techniques, this paper successfully trained a 3 billion-parameter Swin Transformer V2 model, which is the largest dense vision model to date, and makes it capable of training with images of up to 1,536×\times1,536 resolution. It set new performance records on 4 representative vision tasks, including ImageNet-V2 image classification, COCO object detection, ADE20K semantic segmentation, and Kinetics-400 video action classification. Also note our training is much more efficient than that in Google's billion-level visual models, which consumes 40 times less labelled data and 40 times less training time. Code is available at \url{this https URL}.

Added

2026-09-13

Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference

Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference

Dhruv Deshmukh, Saurabh Goyal, Nipun Kwatra, Ramachandran Ramjee

OrganizationsMicrosoft

Why you should read this

Achieves up to 4.1x speedup in long-context LLM inference while maintaining high accuracy, Kascade offers a practical, training-free sparse attention method essential for efficient deployment of reasoning models and RAG on modern hardware.

Attention is the dominant source of latency during long-context LLM inference, an increasingly popular workload with reasoning models and RAG. We propose Kascade, a training-free sparse attention method that leverages known observations such as 1) post-softmax attention is intrinsically sparse, and 2) the identity of high-weight keys is stable across nearby layers. Kascade computes exact Top-k indices in a small set of anchor layers, then reuses those indices in intermediate reuse layers. The anchor layers are selected algorithmically, via a dynamic-programming objective that maximizes cross-layer similarity over a development set, allowing easy deployment across models. The method incorporates efficient implementation constraints (e.g. tile-level operations), across both prefill and decode attention. The Top-k selection and reuse in Kascade is head-aware and we show in our experiments that this is critical for high accuracy. Kascade achieves up to 4.1x speedup in decode attention and 2.2x speedup in prefill attention over FlashAttention-3 baseline on H100 GPUs while closely matching dense attention accuracy on long-context benchmarks such as LongBench and AIME-24.

Added

2026-04-04

Creative Commons License
Attention Is All You Need

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin

OrganizationsGoogleUniversity of Toronto

Why you should read this

Proposes the groundbreaking Transformer architecture, which, by solely relying on attention mechanisms without recurrence or convolutions, achieves superior performance, unprecedented parallelization, and significantly reduced training times in sequence transduction tasks like machine translation.

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.

Added

2026-03-19

License

Published with permission

Relational recurrent neural networks

Relational recurrent neural networks

Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Théophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, Timothy Lillicrap

OrganizationsGoogleUniversity College London

Why you should read this

Shows that memory-based neural networks struggle to reason about relationships between things they've remembered across time, then fixes this by letting memory slots attend to each other, enabling the model to compare and relate stored information—leading to breakthrough performance on tasks requiring temporal reasoning like language modeling and partially-observed reinforcement learning.

Memory-based neural networks model temporal data by leveraging an ability to remember information for long periods. It is unclear, however, whether they also have an ability to perform complex relational reasoning with the information they remember. Here, we first confirm our intuitions that standard memory architectures may struggle at tasks that heavily involve an understanding of the ways in which entities are connected -- i.e., tasks involving relational reasoning. We then improve upon these deficits by using a new memory module -- a \textit{Relational Memory Core} (RMC) -- which employs multi-head dot product attention to allow memories to interact. Finally, we test the RMC on a suite of tasks that may profit from more capable relational reasoning across sequential information, and show large gains in RL domains (e.g. Mini PacMan), program evaluation, and language modeling, achieving state-of-the-art results on the WikiText-103, Project Gutenberg, and GigaWord datasets.

Added

2026-02-21

Reformer: The Efficient Transformer

Reformer: The Efficient Transformer

Nikita Kitaev, Łukasz Kaiser, Anselm Levskaya

OrganizationsGoogleUniversity of California Berkeley

Why you should read this

Introduces Locality Sensitive Hashing (LSH) to identify high-attention pairs without exhaustive search, offering an O(N log N) complexity alternative.

Large Transformer models routinely achieve state-of-the-art results on a number of tasks but training these models can be prohibitively costly, especially on long sequences. We introduce two techniques to improve the efficiency of Transformers. For one, we replace dot-product attention by one that uses locality-sensitive hashing, changing its complexity from O(L2L^2) to O(Llog⁡LL\log L), where LL is the length of the sequence. Furthermore, we use reversible residual layers instead of the standard residuals, which allows storing activations only once in the training process instead of NN times, where NN is the number of layers. The resulting model, the Reformer, performs on par with Transformer models while being much more memory-efficient and much faster on long sequences.

Added

2026-02-11

Rethinking Attention with Performers

Rethinking Attention with Performers

Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, Adrian Weller

OrganizationsGoogleThe Alan Turing InstituteUniversity of Cambridge

Why you should read this

Approximates the softmax attention kernel using positive orthogonal random features, allowing linear time and space complexity.

We introduce Performers, Transformer architectures which can estimate regular (softmax) full-rank-attention Transformers with provable accuracy, but using only linear (as opposed to quadratic) space and time complexity, without relying on any priors such as sparsity or low-rankness. To approximate softmax attention-kernels, Performers use a novel Fast Attention Via positive Orthogonal Random features approach (FAVOR+), which may be of independent interest for scalable kernel methods. FAVOR+ can be also used to efficiently model kernelizable attention mechanisms beyond softmax. This representational power is crucial to accurately compare softmax with other kernels for the first time on large-scale tasks, beyond the reach of regular Transformers, and investigate optimal attention-kernels. Performers are linear architectures fully compatible with regular Transformers and with strong theoretical guarantees: unbiased or nearly-unbiased estimation of the attention matrix, uniform convergence and low estimation variance. We tested Performers on a rich set of tasks stretching from pixel-prediction through text models to protein sequence modeling. We demonstrate competitive results with other examined efficient sparse and dense attention methods, showcasing effectiveness of the novel attention-learning paradigm leveraged by Performers.

Added

2026-02-09

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, Dayiheng Liu B, Jingren Zhou, Junyang Lin B

OrganizationsAlibaba GroupMassachusetts Institute of TechnologyStanford UniversityTsinghua UniversityUniversity of Edinburgh

Why you should read this

Introduces a simple, head-specific sigmoid gate applied after Scaled Dot-Product Attention (SDPA) that consistently improves LLM performance, training stability, and mitigates the attention sink issue.

Gating mechanisms have been widely utilized, from early models like LSTMs [ 1 ] and Highway Networks [ 2 ] to recent state space models [ 3 ], linear attention [ 4 ], and also softmax attention [ 5 , 6 ]. Yet, existing literature rarely examines the specific effects of gating. In this work, we conduct comprehensive experiments to system- atically investigate gating-augmented softmax attention variants. Specifically, we perform a comprehensive comparison over 30 variants of 15B Mixture-of-Experts (MoE) models and 1.7B dense models trained on a 3.5 trillion token dataset. Our central finding is that a simple modification— applying an head-specific sigmoid gate after the Scaled Dot-Product Attention (SDPA)—consistently improves perfor- mance. This modification also enhances training stability, tolerates larger learning rates, and improves scaling properties. By comparing various gating positions and computational variants, we attribute this effectiveness to two key factors: (1) introducing non-linearity upon the low-rank mapping in the softmax attention, and (2) applying query-dependent sparse gating scores to modulate the SDPA output. Notably, we find this sparse gating mechanism mitigates ‘massive activation’ [ 7 ], ‘attention sink’ [ 8 ], and enhances long-context extrapolation performance, and we also release related codes and models to facilitate future research. Furthermore, the most effective SDPA output gating is used in the Qwen3-Next models.

Added

2025-12-04

Creative Commons License