Your Transformer May Not be as Powerful as You Expect
Shengjie LuoShanda LiShuxin ZhengTie-Yan LiuLiwei WangDi He
Proves that standard relative positional encoding prevents Transformers from being universal function approximators due to softmax stochasticity, and introduces a mathematically grounded alternative attention module that restores full expressive capacity with minimal parameter overhead.
Modern artificial intelligence architectures, particularly Transformer models, increasingly rely on Relative Positional Encoding (RPE) to track distances between elements in sequences, images, and graph structures. While empirical studies show that RPE enhances generalization over long sequences compared to traditional Absolute Positional Encoding (APE), the theoretical capabilities of RPE-based models have remained largely unexamined. This article investigates whether RPE-equipped Transformers possess the theoretical capacity to approximate any continuous sequence-to-sequence mapping, a property known as universal function approximation.
The main objective of the article is to mathematically analyze the expressive power of RPE-based Transformers and develop an enhanced architecture that guarantees universal function approximation. To evaluate these properties, the authors conduct rigorous mathematical proofs alongside empirical validations across synthetic tasks, large-scale language modeling on the WikiText-103 dataset, and molecular graph property prediction on the ZINC and PCQM4M benchmark datasets.
The investigation yields several critical findings. First, the article proves that standard RPE-based Transformers are not universal function approximators, regardless of their depth or width, because standard RPE operations force the attention mechanism to output a right stochastic matrix that suppresses positional signals. Second, the authors formulate two theoretical conditions—an attentive condition and a position-aware condition—that restore universal expressiveness. Third, building on these conditions, they introduce Universal RPE-based (URPE) Attention, which applies a lightweight learnable Toeplitz matrix multiplication to the attention map. In practical experiments, URPE achieved 100% accuracy on synthetic position-dependent tasks where standard RPE achieved under 60%. Furthermore, URPE reduced validation perplexity from 23.1 to 22.4 on WikiText-103 and decreased test error by more than 40% on ZINC molecular property benchmarks, all while adding minimal parameters (approximately 4,000 in a 151-million-parameter model) and negligible computational overhead.
These findings indicate that existing RPE models carry architectural bottlenecks that restrict model capacity, especially for position-sensitive tasks. The URPE framework resolves this fundamental deficiency with essentially zero risk to computational budgets, inference latency, or memory consumption. Organizations deploying Transformer backbones can adopt URPE by initializing the new parameters as all-one matrices, enabling seamless fine-tuning of existing pre-trained models without training from scratch.
Decision-makers and engineering teams should consider integrating URPE into current Transformer pipelines, particularly for applications in natural language generation, molecular modeling, and graph learning. However, stakeholders should note that the primary mathematical proofs assume simplified architectures without normalization layers and evaluate bounded input domains. Further testing across additional diverse domain benchmarks is recommended to fully establish performance across other complex tasks.
- Paper: Self-Attention with Relative Position Representations, Peter Shaw et al. (2018). This foundational paper introduced Relative Position Representations into self-attention, establishing the standard relative positional encoding formulation whose theoretical expressive limitations are analyzed in the source.
- Paper: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context, Zihang Dai et al. (2019). This work pioneered relative positional encodings for long-context language modeling, providing the practical baseline architecture and benchmarks evaluated in the source.
- Paper: Do Transformers Really Perform Badly for Graph Representation?, Chengxuan Ying et al. (2021). This paper introduced Graphormer and spatial relative encodings on graph benchmarks like PCQM4M and ZINC, which the source directly builds upon when evaluating graph-structured positional attention.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). This seminal work introduced the Transformer self-attention architecture and original absolute positional encoding that relative positional encoding methods seek to replace.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). This paper introduced Attention with Linear Biases (ALiBi) as a relative positional bias mechanism, offering direct context for the right-stochastic attention limitations analyzed in the source.
- Paper: RoFormer: Enhanced Transformer with Rotary Position Embedding, Jianlin Su et al. (2024). This paper introduces Rotary Position Embedding (RoPE) to incorporate relative distances via vector rotations, offering an alternative relative encoding scheme relevant to the expressivity and position-awareness questions explored in the source.
- Paper: Tighter Bounds on the Expressivity of Transformer Encoders, David Chiang et al. (2023). This work provides formal circuit and logic bounds on Transformer expressivity, complementing the source's universal approximation findings on positional encodings.
- Paper: A Length-Extrapolatable Transformer, Yutao Sun et al. (2023). This paper examines how relative positional representations affect attention resolution and length extrapolation, extending considerations of position-aware attention design.
- Paper: Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis, Ta-Chung Chi et al. (2023). This study analyzes how positional embedding designs shape the empirical receptive field and extrapolation capabilities of Transformers.
- Paper: Randomized Positional Encodings Boost Length Generalization of Transformers, Anian Ruoss et al. (2023). This paper explores the failure modes of relative positional encodings under out-of-distribution sequence lengths and proposes randomized encodings as a solution.
- Paper: LeRoPE: Learnable RoPE Frequencies Improve Language Modeling, Petros Karypis et al. (2026). This work investigates learning relative frequency parameters in positional embeddings, connecting closely to the source's design of lightweight learnable positional attention modifications.
- Paper: Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings, Yoav Gelberg et al. (2025). This paper critically re-evaluates the role and necessity of positional embeddings during inference and context scaling in pretrained language models.
