A Survey of Transformers
Tianyang LinYuxin WangXiangyang LiuXipeng Qiu
Establishes a comprehensive taxonomy of Transformer variants by analyzing architectural modifications, pre-training strategies, and cross-domain applications to help researchers select and design attention-based models.
Modern deep learning relies heavily on the Transformer architecture across various operational domains, such as text analysis, visual processing, and speech systems. However, deploying the standard baseline model introduces substantial computational bottlenecks when dealing with long sequences because memory and processing demands scale quadratically with input length. Additionally, because the baseline architecture makes minimal structural assumptions about input data, it requires massive training datasets to generalize effectively and tends to overfit on smaller data collections.
The article provides a systematic overview and taxonomy of architectural modifications, pre-training strategies, and cross-domain applications developed to overcome these operational limitations. The analysis reviews recent innovations across the research landscape, evaluating structural variations at both the individual module level and the broader network level.
The findings synthesize three key architectural insights. First, computational scaling can be reduced from quadratic to linear complexity using sparse attention patterns, kernel approximations, prototype clustering, or memory compression. Second, introducing inductive priors—such as relative positional representations or localized attention biases—stabilizes optimization and boosts sample efficiency. Third, architecture-level techniques, including hierarchical chunking, recurrence mechanisms, and dynamic early-exit computation, allow models to scale efficiently to long contexts and adapt their computational spend based on sample complexity.
For enterprise practitioners and technical leaders, these findings offer practical pathways to reduce hardware overhead, accelerate inference, and lower training costs without sacrificing performance. Organizations seeking to deploy these architectures should match specific modifications to their operational requirements: sparse or linearized mechanisms for long data sequences, and conditional computation or modular expert routing for variable workloads. Future initiatives should focus on developing stronger theoretical foundations for Transformer representations and building unified multimodal frameworks.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Introduces the original vanilla Transformer architecture and multi-head self-attention mechanism that serves as the baseline and foundation for all variants surveyed.
- Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). Establishes a foundational taxonomy and comprehensive analysis of efficient, sub-quadratic Transformer architectures that directly precede this broader architectural survey.
- Paper: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context, Zihang Dai et al. (2019). Introduces recurrence and relative positional encodings to overcome fixed context lengths, representing a landmark architectural modification discussed in the survey.
- Paper: Reformer: The Efficient Transformer, Nikita Kitaev et al. (2020). Provides a seminal efficiency adaptation using locality-sensitive hashing and reversible layers to drastically reduce computational and memory overhead.
- Paper: Longformer: The Long-Document Transformer, Iz Beltagy et al. (2020). Demonstrates how combining local windowed attention with global attention scales Transformers linearly to long documents, forming a key example of attention pattern modification.
- Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Colin Raffel et al. (2020). Systematically explores pre-training strategies and the unified text-to-text paradigm that underpins the survey's discussion of pre-training methods.
- Paper: Improving Language Understanding by Generative Pre-Training, Alec Radford et al. (2018). Pioneers generative self-supervised pre-training followed by discriminative fine-tuning on Transformer architectures.
- Paper: Generating Long Sequences with Sparse Transformers, Rewon Child et al. (2019). Introduces factorized sparse attention mechanisms that enable Transformers to model long sequences efficiently.
- Paper: Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos et al. (2020). Formulates self-attention as a linear kernel feature map, providing a prominent basis for linear-complexity Transformer designs.
- Paper: A Survey on Vision Transformer, Kai Han et al. (2020). Presents an early, dedicated review of adapting attention and Transformer models to computer vision tasks.
- Paper: Transformers in Time Series: A Survey, Qingsong Wen et al. (2022). Extends the taxonomy of Transformer variants into specialized module- and architecture-level adaptations for time-series forecasting and anomaly detection.
- Paper: A Comprehensive Overview of Large Language Models, Humza Naveed et al. (2023). Broadens the scope from general Transformer variants to modern large language models, surveying scaling dynamics, parameter-efficient tuning, and alignment techniques.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). Proposes an inverted Transformer architecture specifically tailored for multivariate time series forecasting by embedding whole series as tokens.
- Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). Applies scalable Transformer backbones directly to latent diffusion models, replacing traditional convolutional U-Nets in generative modeling.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). Provides a unified framework for parameter-efficient adaptation techniques applied to pretrained Transformer architectures.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). Generalizes multi-modal Transformer modeling into a unified sequence-to-sequence framework handling diverse vision, language, and structured output tasks.
- Paper: A ConvNet for the 2020s, Zhuang Liu et al. (2022). Re-evaluates pure convolutional designs by incorporating architectural principles learned from the success of Vision Transformers.
- Paper: ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders, Sanghyun Woo et al. (2023). Explores masked autoencoder self-supervised pre-training techniques adapted from Transformers to modernize convolutional architectures.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). Continues the exploration of efficient architectural modifications by surveying non-quadratic sequence models, state-space architectures, and sparse mixture-of-experts for LLMs.
- Paper: The Topological Trouble With Transformers, Michael C. Mozer et al. (2026). Critically analyzes the structural limitations of feedforward Transformer attention in tracking internal state across extended sequences.
