A Survey of Transformers

Tianyang LinYuxin WangXiangyang LiuXipeng Qiu

article2021AI Open1,577 citations

Establishes a comprehensive taxonomy of Transformer variants by analyzing architectural modifications, pre-training strategies, and cross-domain applications to help researchers select and design attention-based models.

Listen

Modern deep learning relies heavily on the Transformer architecture across various operational domains, such as text analysis, visual processing, and speech systems. However, deploying the standard baseline model introduces substantial computational bottlenecks when dealing with long sequences because memory and processing demands scale quadratically with input length. Additionally, because the baseline architecture makes minimal structural assumptions about input data, it requires massive training datasets to generalize effectively and tends to overfit on smaller data collections.

The article provides a systematic overview and taxonomy of architectural modifications, pre-training strategies, and cross-domain applications developed to overcome these operational limitations. The analysis reviews recent innovations across the research landscape, evaluating structural variations at both the individual module level and the broader network level.

The findings synthesize three key architectural insights. First, computational scaling can be reduced from quadratic to linear complexity using sparse attention patterns, kernel approximations, prototype clustering, or memory compression. Second, introducing inductive priors—such as relative positional representations or localized attention biases—stabilizes optimization and boosts sample efficiency. Third, architecture-level techniques, including hierarchical chunking, recurrence mechanisms, and dynamic early-exit computation, allow models to scale efficiently to long contexts and adapt their computational spend based on sample complexity.

For enterprise practitioners and technical leaders, these findings offer practical pathways to reduce hardware overhead, accelerate inference, and lower training costs without sacrificing performance. Organizations seeking to deploy these architectures should match specific modifications to their operational requirements: sparse or linearized mechanisms for long data sequences, and conditional computation or modular expert routing for variable workloads. Future initiatives should focus on developing stronger theoretical foundations for Transformer representations and building unified multimodal frameworks.

arXiv: 2106.04554
Cover for A Survey of Transformers

Abstract

Transformers have achieved great success in many artificial intelligence fields, such as natural language processing, computer vision, and audio processing. Therefore, it is natural to attract lots of interest from academic and industry researchers. Up to the present, a great variety of Transformer variants (a.k.a. X-formers) have been proposed, however, a systematic and comprehensive literature review on these Transformer variants is still missing. In this survey, we provide a comprehensive review of various X-formers. We first briefly introduce the vanilla Transformer and then propose a new taxonomy of X-formers. Next, we introduce the various X-formers from three perspectives: architectural modification, pre-training, and applications. Finally, we outline some potential directions for future research.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Vanilla Transformer
  • 2.1.1 Attention Modules
  • 2.1.2 Position-wise FFN
  • 2.1.3 Residual Connection and Normalization
  • 2.1.4 Position Encodings
  • 2.2 Model Usage
  • 2.3 Model Analysis
  • 2.4 Comparing Transformer to Other Network Types
  • 2.4.1 Analysis of Self-Attention
  • 2.4.2 In Terms of Inductive Bias
  • 3 Taxonomy of Transformers
  • 4 Attention
  • 4.1 Sparse Attention
  • 4.1.1 Position-based Sparse Attention
  • 4.1.2 Content-based Sparse Attention
  • 4.2 Linearized Attention
  • 4.2.1 Feature Maps
  • 4.2.2 Aggregation Rule
  • 4.3 Query Prototyping and Memory Compression
  • 4.3.1 Attention with Prototype Queries
  • 4.3.2 Attention with Compressed Key-Value Memory
  • 4.4 Low-rank Self-Attention
  • 4.4.1 Low-rank Parameterization
  • 4.4.2 Low-rank Approximation
  • 4.5 Attention with Prior
  • 4.5.1 Prior that Models locality
  • 4.5.2 Prior from Lower Modules
  • 4.5.3 Prior as Multi-task Adapters
  • 4.5.4 Attention with Only Prior
  • 4.6 Improved Multi-Head Mechanism
  • 4.6.1 Head Behavior Modeling
  • 4.6.2 Multi-head with Restricted Spans
  • 4.6.3 Multi-head with Refined Aggregation
  • 4.6.4 Other Modifications
  • 5 Other Module-level Modifications
  • 5.1 Position Representations
  • 5.1.1 Absolute Position Representations
  • 5.1.2 Relative Position Representations
  • 5.1.3 Other Representations
  • 5.1.4 Position Representations without Explicit Encoding
  • 5.1.5 Position Representation on Transformer Decoders
  • 5.2 Layer Normalization
  • 5.2.1 Placement of Layer Normalization
  • 5.2.2 Substitutes of Layer Normalization
  • 5.2.3 Normalization-free Transformer
  • 5.3 Position-wise FFN
  • 5.3.1 Activation Function in FFN
  • 5.3.2 Adapting FFN for Larger Capacity
  • 5.3.3 Dropping FFN Layers
  • 6 Architecture-level Variants
  • 6.1 Adapting Transformer to Be Lightweight
  • 6.2 Strengthening Cross-Block Connectivity
  • 6.3 Adaptive Computation Time
  • 6.4 Transformers with Divide-and-Conquer Strategies
  • 6.4.1 Recurrent Transformers
  • 6.4.2 Hierarchical Transformers
  • 6.5 Exploring Alternative Architecture
  • 7 Pre-trained Transformers
  • 8 Applications of Transformer
  • 9 Conclusion and Future Directions
  • References

Knowls

  1. Knowl 1 — Taxonomy of Transformer Variants (X-formers)

    model/method

    Transformer variants (X-formers) can be organized into a multi-tiered taxonomy across three primary dimensions:

    1. Module-Level Modifications:

      • Attention Module: Sparse attention (position-based, content-based), linearized attention (kernel feature maps), query prototyping, key-value memory compression, low-rank parameterization/approximation, attention with prior distributions, and improved multi-head mechanisms (head diversity, talking heads, span restrictions, capsule routing).
      • Position Representations: Absolute position encodings (sinusoidal, learned, continuous dynamical systems), relative position encodings (bias terms, disentangled embeddings), hybrid representations (Untied Position Encoding, Rotary Position Embedding), and implicit positional representations (via RNNs or zero-padded convolutions).
      • Layer Normalization: Placement strategies (post-LN vs. pre-LN), parameter-free and alternative normalizations (AdaNorm, scaled ℓ2\ell_2 normalization, PowerNorm), and normalization-free architectures (ReZero).
      • Position-wise Feed-Forward Networks (FFN): Non-linear activation functions (Swish, GELU, GLU variants), capacity expansion (Product-Key Memory, Mixture-of-Experts/GShard, Switch Transformer, hash layers), and FFN pruning/dropping.
    2. Architecture-Level Modifications:

      • Lightweight architectures: Depth-wise convolution hybrids (Lite Transformer), multi-stage sequence downsampling/upsampling (Funnel Transformer), and expanded/reduced representations (DeLighT).
      • Cross-block connectivity: Transparent attention (aggregating all encoder layers for decoder cross-attention), feedback memory, and inter-layer attention map reuse.
      • Adaptive Computation Time (ACT): Dynamic per-token halting (Universal Transformer), conditional layer skipping, and multi-layer early exit criteria.
      • Divide-and-conquer: Segment-level recurrence with memory caching (Transformer-XL, Compressive Transformer, Memformer) and hierarchical processing (sentence-to-document and patch-to-pixel models).
      • Alternative block structures: Non-standard block ordering (Sandwich Transformer, Macaron Transformer FFN-attention-FFN) and neural architecture search (Evolved Transformer, DARTSformer).
    3. Pre-training and Task Adaptation: Encoder-only, decoder-only, and encoder-decoder configurations, alongside multi-task conditioning adapters (such as Conditionally Adaptive Multi-Task Learning).

  2. Knowl 2 — Comparative Complexity, Sequential Operations, and Maximum Path Length Across Neural Network Layers

    data/table

    The properties of self-attention are evaluated against standard alternative deep learning layer types in terms of per-layer computational complexity, minimum sequential operations required for signal propagation, and maximum path length between arbitrary positions. Here, TT denotes the input sequence length, DD is the representation dimension, and KK is the convolution kernel size.

    Layer Type Complexity per Layer Sequential Operations Maximum Path Length
    Self-Attention O(T2⋅D)O(T^2 \cdot D) O(1)O(1) O(1)O(1)
    Fully Connected O(T2⋅D2)O(T^2 \cdot D^2) O(1)O(1) O(1)O(1)
    Convolutional O(K⋅T⋅D2)O(K \cdot T \cdot D^2) O(1)O(1) O(log⁡K(T))O(\log_K(T))
    Recurrent O(T⋅D2)O(T \cdot D^2) O(T)O(T) O(T)O(T)

    Self-attention achieves a constant O(1)O(1) maximum path length and O(1)O(1) sequential operations, facilitating parallel computation and direct modeling of long-range dependencies across all positions without stacking deep layers (as required by convolutional layers with O(log⁡K(T))O(\log_K(T)) path length) or unrolling sequential steps (as required by recurrent layers with O(T)O(T) sequential operations). However, its computational and memory complexity scales quadratically (O(T2⋅D)O(T^2 \cdot D)) with sequence length TT.

  3. Knowl 3 — Kernel-Based Linearized Self-Attention via Associative Property

    equation

    Given query matrix Q∈RT×DQ \in \mathbb{R}^{T \times D}, key matrix K∈RT×DK \in \mathbb{R}^{T \times D}, and value matrix V∈RT×DV \in \mathbb{R}^{T \times D}, standard self-attention computes an output matrix Z=D−1A^VZ = D^{-1} \hat{A} V, where A^=exp⁡(QK⊤)\hat{A} = \exp(Q K^\top) and D=diag(A^1T⊤)D = \text{diag}(\hat{A} \mathbf{1}_T^\top), incurring O(T2D)O(T^2 D) complexity.

    Linearized attention replaces the unnormalized kernel sim(qi,kj)=exp⁡(⟨qi,kj⟩)\text{sim}(q_i, k_j) = \exp(\langle q_i, k_j \rangle) with an explicit or randomized feature map ϕ:RD→RD′\phi: \mathbb{R}^D \to \mathbb{R}^{D'}, yielding the kernel inner product K(qi,kj)=ϕ(qi)ϕ(kj)⊤\mathcal{K}(q_i, k_j) = \phi(q_i) \phi(k_j)^\top. By leveraging matrix associativity, the vector output ziz_i at position ii becomes:

    zi=ϕ(qi)∑j=1T(ϕ(kj)⊗vj)ϕ(qi)∑j′=1Tϕ(kj′)⊤z_i = \frac{\phi(q_i) \sum_{j=1}^T \left(\phi(k_j) \otimes v_j\right)}{\phi(q_i) \sum_{j'=1}^T \phi(k_{j'})^\top}

    where ⊗\otimes denotes the vector outer product.

    In autoregressive (causal) settings where position ii only attends to positions j≤ij \le i, the sums can be computed incrementally via running recurrent state accumulators:

    Si=Si−1+ϕ(ki)⊗vi,ui=ui−1+ϕ(ki)S_i = S_{i-1} + \phi(k_i) \otimes v_i, \quad u_i = u_{i-1} + \phi(k_i)

    zi=ϕ(qi)Siϕ(qi)ui⊤z_i = \frac{\phi(q_i) S_i}{\phi(q_i) u_i^\top}

    This enables O(1)O(1) per-token memory updates and inference time during decoding, reducing overall sequence complexity from O(T2D)O(T^2 D) to O(TDD′)O(T D D').

  4. Knowl 4 — Rotary Position Embedding (RoPE)

    equation

    Rotary Position Embedding (RoPE) injects positional information into self-attention by rotating the affine-transformed query and key vectors in the complex plane. Given an input token vector xt∈RDx_t \in \mathbb{R}^D at position tt, projection matrices WQ,WK∈RD×DkW^Q, W^K \in \mathbb{R}^{D \times D_k}, and a base angle set Θ={θj=10000−2(j−1)/Dk,j∈[1,2,…,Dk/2]}\Theta = \{\theta_j = 10000^{-2(j-1)/D_k}, j \in [1, 2, \dots, D_k/2]\}, the position-encoded query qtq_t and key ktk_t are defined as:

    qt=xtWQRΘ,t,kt=xtWKRΘ,tq_t = x_t W^Q R_{\Theta, t}, \quad k_t = x_t W^K R_{\Theta, t}

    where the rotation matrix RΘ,t∈RDk×DkR_{\Theta, t} \in \mathbb{R}^{D_k \times D_k} is the direct sum (orthogonal block-diagonal composition) of 2D rotation matrices:

    RΘ,t=⨁j=1Dk/2M(t,θj),M(t,θj)=(cos⁡(tθj)sin⁡(tθj)−sin⁡(tθj)cos⁡(tθj))R_{\Theta, t} = \bigoplus_{j=1}^{D_k/2} M(t, \theta_j), \quad M(t, \theta_j) = \begin{pmatrix} \cos(t \theta_j) & \sin(t \theta_j) \\ -\sin(t \theta_j) & \cos(t \theta_j) \end{pmatrix}

    Because rotation matrices satisfy RΘ,i⊤RΘ,j=RΘ,j−iR_{\Theta, i}^\top R_{\Theta, j} = R_{\Theta, j - i}, the resulting dot product between query qiq_i and key kjk_j depends purely on the relative distance j−ij - i:

    qikj⊤=(xiWQ)RΘ,j−i(xjWK)⊤q_i k_j^\top = (x_i W^Q) R_{\Theta, j - i} (x_j W^K)^\top

    This formulation produces absolute position encodings that preserve translation invariance and relative positional dependencies while maintaining compatibility with linear attention mechanisms.

  5. Knowl 5 — Taxonomy of Position-Based Sparse Attention Patterns

    model/method

    Position-based sparse attention reduces the O(T2)O(T^2) complexity of full self-attention by enforcing structural sparsity in the unnormalized attention matrix A^∈RT×T\hat{A} \in \mathbb{R}^{T \times T}:

    A^ij={qikj⊤if token i attends to token j−∞otherwise\hat{A}_{ij} = \begin{cases} q_i k_j^\top & \text{if token } i \text{ attends to token } j \\ -\infty & \text{otherwise} \end{cases}

    Complex sparse attention mechanisms are constructed from compositions of five atomic sparse attention patterns:

    1. Global Attention: Dedicated global hub nodes (selected from sequence tokens or external learned parameters) attend to all tokens and are attended to by all tokens in the sequence.
    2. Band (Sliding Window / Local) Attention: Each query at index ii only attends to keys within a local window [i−w,i+w][i - w, i + w].
    3. Dilated (Strided) Attention: Window attention with a fixed dilation step wd≥1w_d \ge 1, skipping intermediate tokens to expand the receptive field without increasing computational cost.
    4. Random Attention: For each query, a small fixed number of key indices are randomly sampled, exploiting the rapid mixing properties of random graphs.
    5. Block Local Attention: The sequence is divided into non-overlapping blocks of queries, each attending only to a corresponding local memory block.

    Representative compound sparse attention models combine these primitives:

    • Star-Transformer: Band attention of width 3 combined with a single central shared global node.
    • Longformer: Band attention combined with internal global tokens (e.g., [CLS] or question tokens) and dilated attention in upper layers.
    • BigBird: Union of band attention, global node attention, and random attention edges.
  6. Knowl 6 — Disentangled Relative Positional Representations in Attention

    equation

    Standard self-attention computes unnormalized attention logits via scalar products between queries and keys. Relative positional attention models express the logit AijA_{ij} between tokens at positions ii and jj as explicit functions of the offset distance i−ji - j.

    In Transformer-XL, the attention score decomposes content and sinusoidal relative position matrices R∈RT×DkR \in \mathbb{R}^{T \times D_k} with learnable bias vectors u1,u2∈RDku^1, u^2 \in \mathbb{R}^{D_k} and projection matrix WK,R∈RD×DkW^{K,R} \in \mathbb{R}^{D \times D_k}:

    Aij=qikj⊤+qi(Ri−jWK,R)⊤+u1kj⊤+u2(Ri−jWK,R)⊤A_{ij} = q_i k_j^\top + q_i (R_{i-j} W^{K,R})^\top + u^1 k_j^\top + u^2 (R_{i-j} W^{K,R})^\top

    In DeBERTa, relative position embeddings rij=Rclip(i−j)∈RDkr_{ij} = R_{\text{clip}(i-j)} \in \mathbb{R}^{D_k} (clipped to maximum span [−K,K][-K, K]) are untied into separate projection spaces via WK,R,WQ,R∈RD×DkW^{K,R}, W^{Q,R} \in \mathbb{R}^{D \times D_k}:

    Aij=qikj⊤+qi(rijWK,R)⊤+kj(rijWQ,R)⊤A_{ij} = q_i k_j^\top + q_i (r_{ij} W^{K,R})^\top + k_j (r_{ij} W^{Q,R})^\top

    where the three terms represent content-to-content, content-to-position, and position-to-content interactions, respectively.

  7. Knowl 7 — Layer Normalization Placement, Degradation Analysis, and Normalization-Free Alternatives

    theoretical result

    In deep Transformers, the placement and functional form of normalization layers determine optimization stability:

    1. Pre-LN vs. Post-LN:

      • Post-LN: Each sub-layer computes H′=LayerNorm(SubLayer(X)+X)H' = \text{LayerNorm}(\text{SubLayer}(X) + X). At initialization, parameter gradients near the output layer are large, and residual branch dependencies create an amplification effect on output variance shifts. This causes optimization instability unless a learning rate warm-up stage is applied, though Post-LN often achieves superior final performance after convergence.
      • Pre-LN: Each sub-layer computes H′=SubLayer(LayerNorm(X))+XH' = \text{SubLayer}(\text{LayerNorm}(X)) + X. Gradients remain well-scaled throughout depth at initialization, allowing the safe removal of the learning rate warm-up stage.
    2. Parameter-Free and Low-Variance LN Substitutes:

      • AdaNorm: Replaces learned affine scale and shift parameters with fixed scaling based on normalized input variance to avoid overfitting: z=C(1−ky)⊙y,y=x−μσz = C(1 - ky) \odot y, \quad y = \frac{x - \mu}{\sigma} where μ,σ\mu, \sigma are the mean and standard deviation of input xx, and C,kC, k are fixed hyperparameters.
      • Scaled ℓ2\ell_2 Normalization: Projects x∈Rdx \in \mathbb{R}^d onto a sphere of learned radius gg: z=gx∥x∥z = g \frac{x}{\|x\|}.
      • PowerNorm (PN): Substitutes batch variance in Batch Normalization with running quadratic mean statistics: y(t)=x(t)/ψ(t−1)y^{(t)} = x^{(t)} / \psi^{(t-1)}, where (ψ(t))2=α(ψ(t−1))2+(1−α)1∣B∣∑i=1∣B∣(xi(t))2(\psi^{(t)})^2 = \alpha (\psi^{(t-1)})^2 + (1 - \alpha) \frac{1}{|B|} \sum_{i=1}^{|B|} (x_i^{(t)})^2.
    3. Normalization-Free Transformer (ReZero): Replaces LayerNorm entirely by scaling residual connections with a learnable scalar α\alpha initialized to zero: H′=H+α⋅F(H)H' = H + \alpha \cdot F(H), establishing dynamic isometry across deep layers and accelerating convergence.

  8. Knowl 8 — Content-Based Dynamic Sparse Attention via Clustering and Hashing

    model/method

    Rather than utilizing static, position-fixed sparse connection masks, content-based sparse attention conditions connection topologies on dynamic query-key similarities to address the Maximum Inner Product Search (MIPS) problem:

    1. Routing Transformer (kk-means Clustering): Queries {qi}i=1T\{q_i\}_{i=1}^T and keys {ki}i=1T\{k_i\}_{i=1}^T are assigned to the closest centroid among a shared set of centroids {μm}m=1k\{\mu_m\}_{m=1}^k. Each query qiq_i attends exclusively to keys in its assigned cluster: Pi={j:μ(qi)=μ(kj)}\mathcal{P}_i = \{j : \mu(q_i) = \mu(k_j)\}. Cluster centroids are updated online during training via exponential moving averages with hyperparameter λ∈(0,1)\lambda \in (0, 1): μ~←λμ~+(1−λ)(∑i:μ(qi)=μqi+∑j:μ(kj)=μkj),cμ←λcμ+(1−λ)∣μ∣,μ←μ~cμ\tilde{\mu} \leftarrow \lambda \tilde{\mu} + (1 - \lambda) \left( \sum_{i: \mu(q_i)=\mu} q_i + \sum_{j: \mu(k_j)=\mu} k_j \right), \quad c_\mu \leftarrow \lambda c_\mu + (1 - \lambda)|\mu|, \quad \mu \leftarrow \frac{\tilde{\mu}}{c_\mu}

    2. Reformer (Locality-Sensitive Hashing Attention): Uses Locality-Sensitive Hashing (LSH) to project vectors into discrete buckets such that similar vectors share buckets with high probability. Given a random projection matrix R∈RDk×b/2R \in \mathbb{R}^{D_k \times b/2} for bb buckets, the hash function is: h(x)=arg⁡max⁡([xR;−xR])h(x) = \arg\max([xR; -xR]) Attention is constrained to keys falling into the identical hash bucket: Pi={j:h(qi)=h(kj)}\mathcal{P}_i = \{j : h(q_i) = h(k_j)\}, reducing self-attention complexity to O(Tlog⁡T)O(T \log T).

  9. Knowl 9 — Segment-Level Recurrence and Memory Compression for Long-Range Sequence Modeling

    equation

    To model sequences beyond fixed context limits without quadratic overhead, recurrent Transformers partition long sequences into segments and pass cached hidden states across segments:

    1. Transformer-XL Recurrence: For segment τ+1\tau + 1 at layer ll, the input representations Hτ+1(l−1)H_{\tau+1}^{(l-1)} are concatenated along the sequence length dimension with cached representations from the previous segment Hτ(l−1)H_\tau^{(l-1)}: H~τ+1(l)=[SG(Hτ(l−1))∘Hτ+1(l−1)],Kτ+1(l)=H~τ+1(l)WK,Vτ+1(l)=H~τ+1(l)WV\tilde{H}_{\tau+1}^{(l)} = [\text{SG}(H_\tau^{(l-1)}) \circ H_{\tau+1}^{(l-1)}], \quad K_{\tau+1}^{(l)} = \tilde{H}_{\tau+1}^{(l)} W^K, \quad V_{\tau+1}^{(l)} = \tilde{H}_{\tau+1}^{(l)} W^V where SG(⋅)\text{SG}(\cdot) is the stop-gradient operator and ∘\circ denotes sequence concatenation. For a network with LL layers and a cache memory of length NmemN_{\text{mem}}, the maximum context dependency length scales as O(L×Nmem)O(L \times N_{\text{mem}}).

    2. Compressive Transformer: Extends the primary memory cache NmemN_{\text{mem}} by applying a compression function (such as convolution or pooling with compression rate cc) to older representations, storing them in a secondary compressed memory NcmN_{\text{cm}}. The theoretical receptive field expands to O(L×(Nmem+c⋅Ncm))O(L \times (N_{\text{mem}} + c \cdot N_{\text{cm}})).

    3. ERNIE-Doc Same-Layer Recurrence: Replaces the (l−1)(l-1)-th layer memory dependency with cached activations from the identical layer ll of the previous segment: H~τ+1(l)=[SG(Hτ(l))∘Hτ+1(l−1)]\tilde{H}_{\tau+1}^{(l)} = [\text{SG}(H_\tau^{(l)}) \circ H_{\tau+1}^{(l-1)}] enabling representations to propagate context across segments without vertical layer-shift constraints.

  10. Knowl 10 — Multi-Head Attention Aggregation Equivalence and Capsule Routing Alternatives

    equation

    In vanilla multi-head attention with HH heads and output projection matrix WO∈RDm×DmW^O \in \mathbb{R}^{D_m \times D_m}, partitioning WOW^O row-wise into HH sub-matrices WO=[W1O;W2O;… ;WHO]W^O = [W_1^O; W_2^O; \dots; W_H^O] (where each WiO∈RDv×DmW_i^O \in \mathbb{R}^{D_v \times D_m}) reveals that the standard concatenation-and-projection operation is mathematically equivalent to the linear summation of re-parameterized head outputs:

    MultiHeadAttn(Q,K,V)=∑i=1HAttention(QWiQ,KWiK,VWiVWiO)\text{MultiHeadAttn}(Q, K, V) = \sum_{i=1}^H \text{Attention}\left(Q W_i^Q, K W_i^K, V W_i^V W_i^O\right)

    To overcome the expressiveness limitations of linear summation without head interaction:

    • Talking-Head Attention: Linearly transforms attention logits across head dimensions prior to softmax (from hkh_k to hh projections) and transforms post-softmax attention weights before value aggregation (from hh to hvh_v projections).
    • Capsule Routing Aggregation: Transforms the outputs of attention heads into input capsules, applies iterative dynamic routing or Expectation-Maximization (EM) routing to compute routing agreements, and concatenates the resultant output capsules to produce the final layer representation.

Coverage note — Domain-specific application surveys (e.g., standard pipelines in NLP, CV, and audio) and standard overviews of foundational pre-trained models (e.g., vanilla BERT, GPT-2) were omitted as they represent contextual references rather than standalone methodological contributions of the survey.

References

  1. 1.Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020. ETC: Encoding Long and Structured Inputs in Transformers. In Proceedings of EMNLP. Online, 268–284. https://doi.org/10.18653/v1/2020.emnlp-main.19
  2. 2.Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones. 2019. Character-Level Language Modeling with Deeper Self-Attention. In Proceedings of AAAI. 3159–3166. https://doi.org/10.1609/aaai.v33i01.33013159
  3. 3.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021. ViViT: A Video Vision Transformer. arXiv:2103.15691 [cs.CV]
  4. 4.Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer Normalization. CoRR abs/1607.06450 (2016). arXiv:1607.06450
  5. 5.Thomas Bachlechner, Bodhisattwa Prasad Majumder, Huanru Henry Mao, Garrison W. Cottrell, and Julian J. McAuley. 2020. ReZero is All You Need: Fast Convergence at Large Depth. CoRR abs/2003.04887 (2020). arXiv:2003.04887
  6. 6.Alexei Baevski and Michael Auli. 2019. Adaptive Input Representations for Neural Language Modeling. In Proceedings of ICLR. https://openreview.net/forum?id=ByxZX20qFQ
  7. 7.Ankur Bapna, Naveen Arivazhagan, and Orhan Firat. 2020. Controlling Computation versus Quality for Neural Sequence Models. arXiv:2002.07106 [cs.LG]
  8. 8.Ankur Bapna, Mia Chen, Orhan Firat, Yuan Cao, and Yonghui Wu. 2018. Training Deeper Neural Machine Translation Models with Transparent Attention. In Proceedings of EMNLP. Brussels, Belgium, 3028–3033. https://doi.org/10.18653/v1/D18-1338
  9. 9.Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andrew Ballard, Justin Gilmer, George Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Langston, Chris Dyer, Nicolas Heess, Daan Wierstra, Pushmeet Kohli, Matt Botvinick, Oriol Vinyals, Yujia Li, and Razvan Pascanu. 2018. Relational inductive biases, deep learning, and graph networks. arXiv:1806.01261 [cs.LG]
  10. 10.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The Long-Document Transformer. arXiv:2004.05150 [cs.CL]
  11. 11.Srinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J. Reddi, and Sanjiv Kumar. 2020. Low-Rank Bottleneck in Multi-head Attention Models. In Proceedings of ICML. 864–873. http://proceedings.mlr.press/v119/bhojanapalli20a.html
  12. 12.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Proceedings of NeurIPS. 1877–1901. https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
  13. 13.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-End Object Detection with Transformers. In Proceedings of ECCV. 213–229. https://doi.org/10.1007/978-3-030-58452-8_13
  14. 14.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. 2020. Generative Pretraining From Pixels. In Proceedings of ICML. 1691–1703. http://proceedings.mlr.press/v119/chen20s.html
  15. 15.Xie Chen, Yu Wu, Zhenghao Wang, Shujie Liu, and Jinyu Li. 2021. Developing Real-time Streaming Transformer Transducer for Speech Recognition on Large-scale Dataset. arXiv:2010.11395 [cs.CL]
  16. 16.Ziye Chen, Mingming Gong, Lingjuan Ge, and Bo Du. 2020. Compressed Self-Attention for Deep Metric Learning with Low-Rank Approximation. In Proceedings of IJCAI. 2058–2064. https://doi.org/10.24963/ijcai.2020/285
  17. 17.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating Long Sequences with Sparse Transformers. arXiv:1904.10509 [cs.LG]
  18. 18.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, David Belanger, Lucy Colwell, and Adrian Weller. 2020. Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers. arXiv:2006.03555 [cs.LG]
  19. 19.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. 2020. Rethinking Attention with Performers. arXiv:2009.14794 [cs.LG]
  20. 20.Xiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. 2021. Conditional Positional Encodings for Vision Transformers. arXiv:2102.10882 [cs.CV]
  21. 21.Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. 2020. Multi-Head Attention: Collaborate Instead of Concatenate. CoRR abs/2006.16362 (2020). arXiv:2006.16362
  22. 22.Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. 2020. Meshed-Memory Transformer for Image Captioning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020. IEEE, 10575–10584. https://doi.org/10.1109/CVPR42600.2020.01059
  23. 23.Zihang Dai, Guokun Lai, Yiming Yang, and Quoc Le. 2020. Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing. In Proceedings of NeurIPS. https://proceedings.neurips.cc/paper/2020/hash/2cd2915e69546904e4e5d4a2ac9e1652-Abstract.html
  24. 24.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019. Transformer-XL: Attentive Language Models beyond a Fixed-Length Context. In Proceedings of ACL. Florence, Italy, 2978–2988. https://doi.org/10.18653/v1/P19-1285
  25. 25.Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017. Language Modeling with Gated Convolutional Networks. In Proceedings of ICML. 933–941. http://proceedings.mlr.press/v70/dauphin17a.html
  26. 26.Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019. Universal Transformers. In Proceedings of ICLR. https://openreview.net/forum?id=HyzdRiR9Y7
  27. 27.Ameet Deshpande and Karthik Narasimhan. 2020. Guiding Attention for Self-Supervised Learning with Transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020. Online, 4676–4686. https://doi.org/10.18653/v1/2020.findings-emnlp.419
  28. 28.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of HLT-NAACL. Minneapolis, Minnesota, 4171–4186. https://doi.org/10.18653/v1/N19-1423
  29. 29.Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. 2021. CogView: Mastering Text-to-Image Generation via Transformers. arXiv:2105.13290 [cs.CV]
  30. 30.Siyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2020. ERNIE-DOC: The Retrospective Long-Document Modeling Transformer. (2020). arXiv:2012.15688 [cs.CL]
  31. 31.Linhao Dong, Shuang Xu, and Bo Xu. 2018. Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition. In Proceedings of ICASSP. 5884–5888. https://doi.org/10.1109/ICASSP.2018.8462506
  32. 32.Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. 2021. Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth. CoRR abs/2103.03404 (2021). arXiv:2103.03404
  33. 33.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv:2010.11929 [cs.CV]
  34. 34.Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, and Sainbayar Sukhbaatar. 2021. Addressing Some Limitations of Transformers with Feedback Memory. https://openreview.net/forum?id=OCm0rwa1lx1
  35. 35.Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, and Xuanjing Huang. 2021. Mask Attention Networks: Rethinking and Strengthen Transformer. In Proceedings of NAACL. 1692–1701. https://www.aclweb.org/anthology/2021.naacl-main.135
  36. 36.William Fedus, Barret Zoph, and Noam Shazeer. 2021. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. CoRR abs/2101.03961 (2021). arXiv:2101.03961
  37. 37.Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017. Convolutional Sequence to Sequence Learning. In Proceedings of ICML. 1243–1252.
  38. 38.Alex Graves. 2016. Adaptive Computation Time for Recurrent Neural Networks. CoRR abs/1603.08983 (2016). arXiv:1603.08983
  39. 39.Jiatao Gu, Qi Liu, and Kyunghyun Cho. 2019. Insertion-based Decoding with Automatically Inferred Generation Order. Trans. Assoc. Comput. Linguistics 7 (2019), 661–676. https://transacl.org/ojs/index.php/tacl/article/view/1732
  40. 40.Shuhao Gu and Yang Feng. 2019. Improving Multi-head Attention with Capsule Networks. In Proceedings of NLPCC. 314–326. https://doi.org/10.1007/978-3-030-32233-5_25
  41. 41.Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang. 2020. Conformer: Convolution-augmented Transformer for Speech Recognition. In Proceedings of Interspeech. 5036–5040. https://doi.org/10.21437/Interspeech.2020-3015
  42. 42.Maosheng Guo, Yu Zhang, and Ting Liu. 2019. Gaussian Transformer: A Lightweight Approach for Natural Language Inference. In Proceedings of AAAI. 6489–6496. https://doi.org/10.1609/aaai.v33i01.33016489
  43. 43.Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang. 2019. Star-Transformer. In Proceedings of HLT-NAACL. 1315–1325. https://www.aclweb.org/anthology/N19-1133
  44. 44.Qipeng Guo, Xipeng Qiu, Pengfei Liu, Xiangyang Xue, and Zheng Zhang. 2020. Multi-Scale Self-Attention for Text Classification. In Proceedings of AAAI. 7847–7854. https://aaai.org/ojs/index.php/AAAI/article/view/6290
  45. 45.Qipeng Guo, Xipeng Qiu, Xiangyang Xue, and Zheng Zhang. 2019. Low-Rank and Locality Constrained Self-Attention for Sequence Modeling. IEEE/ACM Trans. Audio, Speech and Lang. Proc. 27, 12 (2019), 2213–2222. https://doi.org/10.1109/TASLP.2019.2944078
  46. 46.Chi Han, Mingxuan Wang, Heng Ji, and Lei Li. 2021. Learning Shared Semantic Space for Speech-to-Text Translation. arXiv:2105.03095 [cs.CL]
  47. 47.Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, Zhaohui Yang, Yiman Zhang, and Dacheng Tao. 2021. A Survey on Visual Transformer. arXiv:2012.12556 [cs.CV]
  48. 48.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. 2021. Transformer in Transformer. arXiv:2103.00112 [cs.CV]
  49. 49.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In Proceedings CVPR. 770–778. https://doi.org/10.1109/CVPR.2016.90
  50. 50.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. arXiv:2006.03654
  51. 51.Ruining He, Anirudh Ravula, Bhargav Kanagal, and Joshua Ainslie. 2020. RealFormer: Transformer Likes Residual Attention. arXiv:2012.11747 [cs.LG]
  52. 52.Dan Hendrycks and Kevin Gimpel. 2020. Gaussian Error Linear Units (GELUs). arXiv:1606.08415 [cs.LG]
  53. 53.Geoffrey E. Hinton, Sara Sabour, and Nicholas Frosst. 2018. Matrix capsules with EM routing. In Proceedings of ICLR. https://openreview.net/forum?id=HJWLfGWRb
  54. 54.Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. 2019. Axial Attention in Multidimensional Transformers. CoRR abs/1912.12180 (2019). arXiv:1912.12180
  55. 55.Ronghang Hu, Amanpreet Singh, Trevor Darrell, and Marcus Rohrbach. 2020. Iterative Answer Prediction With Pointer-Augmented Multimodal Transformers for TextVQA. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020. 9989–9999. https://doi.org/10.1109/CVPR42600.2020.01001
  56. 56.Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, and Douglas Eck. 2019. Music Transformer. In Proceedings of ICLR. https://openreview.net/forum?id=rJe4ShAcF7
  57. 57.Hyeong Rae Ihm, Joun Yeop Lee, Byoung Jin Choi, Sung Jun Cheon, and Nam Soo Kim. 2020. Reformer-TTS: Neural Speech Synthesis with Reformer Network. In Proceedings of Interspeech, Helen Meng, Bo Xu, and Thomas Fang Zheng (Eds.). 2012–2016. https://doi.org/10.21437/Interspeech.2020-2189
  58. 58.Sergey Ioffe and Christian Szegedy. 2015. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In Proceedings of ICML. 448–456. http://proceedings.mlr.press/v37/ioffe15.html
  59. 59.Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney. 2019. Language Modeling with Deep Transformers. In Proceedings of Interspeech. 3905–3909. https://doi.org/10.21437/Interspeech.2019-2225
  60. 60.Md. Amirul Islam, Sen Jia, and Neil D. B. Bruce. 2020. How much Position Information Do Convolutional Neural Networks Encode?. In Proceedings of ICLR. https://openreview.net/forum?id=rJeB36NKvB
  61. 61.Yifan Jiang, Shiyu Chang, and Zhangyang Wang. 2021. TransGAN: Two Transformers Can Make One Strong GAN. arXiv:2102.07074 [cs.CV]
  62. 62.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. In Proceedings of ICML. 5156–5165. http://proceedings.mlr.press/v119/katharopoulos20a.html
  63. 63.Guolin Ke, Di He, and Tie-Yan Liu. 2020. Rethinking Positional Encoding in Language Pre-training. arXiv:2006.15595 [cs.CL]
  64. 64.Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. 2021. Transformers in Vision: A Survey. arXiv:2101.01169 [cs.CV]
  65. 65.Jaeyoung Kim, Mostafa El-Khamy, and Jungwon Lee. 2020. T-GSA: Transformer with Gaussian-Weighted Self-Attention for Speech Enhancement. In 2020 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8, 2020. IEEE, 6649–6653. https://doi.org/10.1109/ICASSP40776.2020.9053591
  66. 66.Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The Efficient Transformer. In Proceedings of ICLR. https://openreview.net/forum?id=rkgNKkHtvB
  67. 67.Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander Rush. 2017. OpenNMT: Open-Source Toolkit for Neural Machine Translation. In Proceedings of ACL. 67–72. https://www.aclweb.org/anthology/P17-4012
  68. 68.Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019. Revealing the Dark Secrets of BERT. In Proceedings of EMNLP-IJCNLP. 4364–4373. https://doi.org/10.18653/v1/D19-1445
  69. 69.Guillaume Lample, Alexandre Sablayrolles, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2019. Large Memory Layers with Product Keys. In Proceedings of NeurIPS. 8546–8557. https://proceedings.neurips.cc/paper/2019/hash/9d8df73a3cfbf3c5b47bc9b50f214aff-Abstract.html
  70. 70.Juho Lee, Yoonho Lee, Jungtaek Kim, Adam R. Kosiorek, Seungjin Choi, and Yee Whye Teh. 2019. Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks. In Proceedings of ICML. 3744–3753. http://proceedings.mlr.press/v97/lee19d.html
  71. 71.Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. CoRR abs/2006.16668 (2020). arXiv:2006.16668
  72. 72.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of ACL. 7871–7880. https://doi.org/10.18653/v1/2020.acl-main.703
  73. 73.Jian Li, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang. 2018. Multi-Head Attention with Disagreement Regularization. In Proceedings of EMNLP. Brussels, Belgium, 2897–2903. https://doi.org/10.18653/v1/D18-1317
  74. 74.Jian Li, Baosong Yang, Zi-Yi Dou, Xing Wang, Michael R. Lyu, and Zhaopeng Tu. 2019. Information Aggregation for Multi-Head Attention with Routing-by-Agreement. In Proceedings of HLT-NAACL. 3566–3575. https://doi.org/10.18653/v1/N19-1359
  75. 75.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. VisualBERT: A Simple and Performant Baseline for Vision and Language. arXiv:1908.03557 [cs.CV]
  76. 76.Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019. Neural Speech Synthesis with Transformer Network. In Proceedings of AAAI. 6706–6713. https://doi.org/10.1609/aaai.v33i01.33016706
  77. 77.Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. 2020. UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning. arXiv preprint arXiv:2012.15409 (2020).
  78. 78.Xiaoya Li, Yuxian Meng, Mingxin Zhou, Qinghong Han, Fei Wu, and Jiwei Li. 2020. SAC: Accelerating and Structuring Self-Attention via Sparse Adaptive Connection. In Proceedings of NeurIPS. https://proceedings.neurips.cc/paper/2020/hash/c5c1bda1194f9423d744e0ef67df94ee-Abstract.html
  79. 79.Xiaonan Li, Yunfan Shao, Tianxiang Sun, Hang Yan, Xipeng Qiu, and Xuanjing Huang. 2021. Accelerating BERT Inference for Sequence Labeling via Early-Exit. arXiv:2105.13878 [cs.CL]
  80. 80.Xiaonan Li, Hang Yan, Xipeng Qiu, and Xuanjing Huang. 2020. FLAT: Chinese NER Using Flat-Lattice Transformer. In Proceedings of ACL. 6836–6842. https://doi.org/10.18653/v1/2020.acl-main.611
  81. 81.Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, Jie Zhang, Jianwei Zhang, Xu Zou, Zhikang Li, Xiaodong Deng, Jie Liu, Jinbao Xue, Huiling Zhou, Jianxin Ma, Jin Yu, Yong Li, Wei Lin, Jingren Zhou, Jie Tang, and Hongxia Yang. 2021. M6: A Chinese Multimodal Pretrainer. arXiv:2103.00823 [cs.CL]
  82. 82.Hanxiao Liu, Karen Simonyan, and Yiming Yang. 2019. DARTS: Differentiable Architecture Search. In Proceedings of ICLR. https://openreview.net/forum?id=S1eYHoC5FX
  83. 83.Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han. 2020. Understanding the Difficulty of Training Transformers. In Proceedings of EMNLP. 5747–5763. https://doi.org/10.18653/v1/2020.emnlp-main.463
  84. 84.Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018. Generating Wikipedia by Summarizing Long Sequences. In Proceedings of ICLR. https://openreview.net/forum?id=Hyg0vbWC-
  85. 85.Xuanqing Liu, Hsiang-Fu Yu, Inderjit S. Dhillon, and Cho-Jui Hsieh. 2020. Learning to Encode Position for Transformer with Continuous Dynamical Model. In Proceedings of ICML. 6327–6335. http://proceedings.mlr.press/v119/liu20n.html
  86. 86.Yang Liu and Mirella Lapata. 2019. Hierarchical Transformers for Multi-Document Summarization. In Proceedings of ACL. Florence, Italy, 5070–5081. https://doi.org/10.18653/v1/P19-1500
  87. 87.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL]
  88. 88.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. arXiv:2103.14030 [cs.CV]
  89. 89.Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2020. Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View. https://openreview.net/forum?id=SJl1o2NFwS
  90. 90.Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma, and Luke Zettlemoyer. 2021. Luna: Linear Unified Nested Attention. arXiv:2106.01540 [cs.LG]
  91. 91.Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2020. DeLighT: Very Deep and Light-weight Transformer. arXiv:2008.00623 [cs.LG]
  92. 92.Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018. Document-Level Neural Machine Translation with Hierarchical Attention Networks. In Proceedings of EMNLP. Brussels, Belgium, 2947–2954. https://doi.org/10.18653/v1/D18-1325
  93. 93.Toan Q. Nguyen and Julian Salazar. 2019. Transformers without Tears: Improving the Normalization of Self-Attention. CoRR abs/1910.05895 (2019). arXiv:1910.05895
  94. 94.Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018. Image Transformer. In Proceedings of ICML. 4052–4061. http://proceedings.mlr.press/v80/parmar18a.html
  95. 95.Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong. 2021. Random Feature Attention. In Proceedings of ICLR. https://openreview.net/forum?id=QtTKTdVrFBB
  96. 96.Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville. 2018. FiLM: Visual Reasoning with a General Conditioning Layer. In Proceedings of AAAI. 3942–3951. https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/16528
  97. 97.Ngoc-Quan Pham, Thai-Son Nguyen, Jan Niehues, Markus Müller, and Alex Waibel. 2019. Very Deep Self-Attention Networks for End-to-End Speech Recognition. In Proceedings of Interspeech. 66–70. https://doi.org/10.21437/Interspeech.2019-2702
  98. 98.Jonathan Pilault, Amine El hattami, and Christopher Pal. 2021. Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less Data. In Proceedings of ICLR. https://openreview.net/forum?id=de11dbHzAMF
  99. 99.Ofir Press, Noah A. Smith, and Omer Levy. 2020. Improving Transformer Models by Reordering their Sublayers. In Proceedings of ACL. Online, 2996–3005. https://doi.org/10.18653/v1/2020.acl-main.270
  100. 100.Xipeng Qiu, TianXiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. 2020. Pre-trained Models for Natural Language Processing: A Survey. SCIENCE CHINA Technological Sciences 63, 10 (2020), 1872–1897. https://doi.org/10.1007/s11431-020-1647-3
  101. 101.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. (2018).
  102. 102.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. (2019).
  103. 103.Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. 2020. Compressive Transformers for Long-Range Sequence Modelling. In Proceedings of ICLR. https://openreview.net/forum?id=SylKikSYDH
  104. 104.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 [cs.LG]
  105. 105.Ali Rahimi and Benjamin Recht. 2007. Random Features for Large-Scale Kernel Machines. In Proceedings of NeurIPS. 1177–1184. https://proceedings.neurips.cc/paper/2007/hash/013a006f03dbc5392effeb8f18fda755-Abstract.html
  106. 106.Prajit Ramachandran, Barret Zoph, and Quoc V. Le. 2018. Searching for Activation Functions. In Proceedings of ICLR. https://openreview.net/forum?id=Hkuq2EkPf
  107. 107.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-Shot Text-to-Image Generation. arXiv:2102.12092 [cs.CV]
  108. 108.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017. Learning multiple visual domains with residual adapters. In Proceedings of NeurIPS. 506–516. https://proceedings.neurips.cc/paper/2017/hash/e7b24b112a44fdd9ee93bdf998c6ca0e-Abstract.html
  109. 109.Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus. 2021. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences 118, 15 (2021). https://doi.org/10.1073/pnas.2016239118
  110. 110.Stephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, and Jason Weston. 2021. Hash Layers For Large Sparse Models. arXiv:2106.04426 [cs.LG]
  111. 111.Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2020. Efficient Content-Based Sparse Attention with Routing Transformers. arXiv:2003.05997 [cs.LG]
  112. 112.Sara Sabour, Nicholas Frosst, and Geoffrey E. Hinton. 2017. Dynamic Routing Between Capsules. In Proceedings of NeurIPS. 3856–3866. https://proceedings.neurips.cc/paper/2017/hash/2cad8fa47bbef282badbb8de5374b894-Abstract.html
  113. 113.Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. 2021. Linear Transformers Are Secretly Fast Weight Memory Systems. CoRR abs/2102.11174 (2021). arXiv:2102.11174
  114. 114.Philippe Schwaller, Teodoro Laino, Théophile Gaudin, Peter Bolgar, Christopher A. Hunter, Costas Bekas, and Alpha A. Lee. 2019. Molecular Transformer: A Model for Uncertainty-Calibrated Chemical Reaction Prediction. ACS Central Science 5, 9 (2019), 1572–1583. https://doi.org/10.1021/acscentsci.9b00576
  115. 115.Jie Shao, Xin Wen, Bingchen Zhao, and Xiangyang Xue. 2021. Temporal Context Aggregation for Video Retrieval With Contrastive Learning. In Proceedings of WACV. 3268–3278.
  116. 116.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018. Self-Attention with Relative Position Representations. In Proceedings of HLT-NAACL. New Orleans, Louisiana, 464–468. https://doi.org/10.18653/v1/N18-2074
  117. 117.Noam Shazeer. 2019. Fast Transformer Decoding: One Write-Head is All You Need. CoRR abs/1911.02150 (2019). arXiv:1911.02150
  118. 118.Noam Shazeer. 2020. GLU Variants Improve Transformer. arXiv:2002.05202 [cs.LG]
  119. 119.Noam Shazeer, Zhenzhong Lan, Youlong Cheng, Nan Ding, and Le Hou. 2020. Talking-Heads Attention. CoRR abs/2003.02436 (2020). arXiv:2003.02436
  120. 120.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In Proceedings of ICLR. https://openreview.net/forum?id=B1ckMDqlg
  121. 121.Sheng Shen, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2020. PowerNorm: Rethinking Batch Normalization in Transformers. In Proceedings of ICML. 8741–8751. http://proceedings.mlr.press/v119/shen20e.html
  122. 122.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared D Casper, and Bryan Catanzaro. 2020. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv:1909.08053 [cs.CL]
  123. 123.David R. So, Quoc V. Le, and Chen Liang. 2019. The Evolved Transformer. In Proceedings of ICML. 5877–5886. http://proceedings.mlr.press/v97/so19a.html
  124. 124.Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. 2021. RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv:2104.09864
  125. 125.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In Proceedings of ICLR. https://openreview.net/forum?id=SygXPaEYvH
  126. 126.Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. 2019. Adaptive Attention Span in Transformers. In Proceedings of ACL. Florence, Italy, 331–335. https://doi.org/10.18653/v1/P19-1032
  127. 127.Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin. 2019. Augmenting Self-attention with Persistent Memory. arXiv:1907.01470 [cs.LG]
  128. 128.Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019. VideoBERT: A Joint Model for Video and Language Representation Learning. In Proceedings of ICCV. 7463–7472. https://doi.org/10.1109/ICCV.2019.00756
  129. 129.Tianxiang Sun, Yunhua Zhou, Xiangyang Liu, Xinyu Zhang, Hao Jiang, Zhao Cao, Xuanjing Huang, and Xipeng Qiu. 2021. Early Exiting with Ensemble Internal Classifiers. arXiv:2105.13792 [cs.CL]
  130. 130.Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to Sequence Learning with Neural Networks. In Proceedings of NeurIPS. 3104–3112. https://proceedings.neurips.cc/paper/2014/hash/a14ac55a4f27472c5d894ec1c3c743d2-Abstract.html
  131. 131.Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2020. Synthesizer: Rethinking Self-Attention in Transformer Models. CoRR abs/2005.00743 (2020). arXiv:2005.00743
  132. 132.Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan. 2020. Sparse Sinkhorn Attention. In Proceedings of ICML. 9438–9447. http://proceedings.mlr.press/v119/tay20a.html
  133. 133.Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Transformer Dissection: An Unified Understanding for Transformer’s Attention via the Lens of Kernel. In Proceedings of EMNLP-IJCNLP. Hong Kong, China, 4344–4353. https://doi.org/10.18653/v1/D19-1443
  134. 134.Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016. WaveNet: A Generative Model for Raw Audio. In Proceedings of ISCA. 125. http://www.isca-speech.org/archive/SSW_2016/abstracts/ssw9_DS-4_van_den_Oord.html
  135. 135.Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation Learning with Contrastive Predictive Coding. CoRR abs/1807.03748 (2018). arXiv:1807.03748
  136. 136.Ashish Vaswani, Samy Bengio, Eugene Brevdo, Francois Chollet, Aidan Gomez, Stephan Gouws, Llion Jones, Łukasz Kaiser, Nal Kalchbrenner, Niki Parmar, Ryan Sepassi, Noam Shazeer, and Jakob Uszkoreit. 2018. Tensor2Tensor for Neural Machine Translation. In Proceedings of AMTA. 193–199. https://www.aclweb.org/anthology/W18-1819
  137. 137.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Proceedings of NeurIPS. 5998–6008. https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html
  138. 138.Apoorv Vyas, Angelos Katharopoulos, and François Fleuret. 2020. Fast Transformers with Clustered Attention. arXiv:2007.04825 [cs.LG]
  139. 139.Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen. [n.d.]. On Position Embeddings in BERT, url = https://openreview.net/forum?id=onxoVA9FxMw, year = 2021. In Proceedings of ICLR.
  140. 140.Benyou Wang, Donghao Zhao, Christina Lioma, Qiuchi Li, Peng Zhang, and Jakob Grue Simonsen. 2020. Encoding word order in complex embeddings. In Proceedings of ICLR. https://openreview.net/forum?id=Hke-WTVtwr
  141. 141.Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F. Wong, and Lidia S. Chao. 2019. Learning Deep Transformer Models for Machine Translation. In Proceedings of ACL. 1810–1822. https://doi.org/10.18653/v1/p19-1176
  142. 142.Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. Linformer: Self-Attention with Linear Complexity. arXiv:2006.04768 [cs.LG]
  143. 143.Yujing Wang, Yaming Yang, Jiangang Bai, Mingliang Zhang, Jing Bai, Jing Yu, Ce Zhang, and Yunhai Tong. 2021. Predictive Attention Transformer: Improving Transformer with Attention Map Prediction. https://openreview.net/forum?id=YQVjbJPnPc9
  144. 144.Zhiwei Wang, Yao Ma, Zitao Liu, and Jiliang Tang. 2019. R-Transformer: Recurrent Neural Network Enhanced Transformer. CoRR abs/1907.05572 (2019). arXiv:1907.05572
  145. 145.Chuhan Wu, Fangzhao Wu, Tao Qi, and Yongfeng Huang. 2021. Hi-Transformer: Hierarchical Interactive Transformer for Efficient and Effective Long Document Modeling. arXiv:2106.01040 [cs.CL]
  146. 146.Felix Wu, Angela Fan, Alexei Baevski, Yann N. Dauphin, and Michael Auli. 2019. Pay Less Attention with Lightweight and Dynamic Convolutions. In Proceedings of ICLR. https://openreview.net/forum?id=SkVhlh09tX
  147. 147.Qingyang Wu, Zhenzhong Lan, Jing Gu, and Zhou Yu. 2020. Memformer: The Memory-Augmented Transformer. arXiv:2010.06891 [cs.CL]
  148. 148.Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. 2020. Lite Transformer with Long-Short Range Attention. In Proceedings of ICLR. https://openreview.net/forum?id=ByeMPlHKPH
  149. 149.Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and Philip S. Yu. 2021. A Comprehensive Survey on Graph Neural Networks. IEEE Trans. Neural Networks Learn. Syst. 32, 1 (2021), 4–24. https://doi.org/10.1109/TNNLS.2020.2978386
  150. 150.Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020. DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference. In Proceedings of ACL. 2246–2251. https://doi.org/10.18653/v1/2020.acl-main.204
  151. 151.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu. 2020. On Layer Normalization in the Transformer Architecture. In Proceedings of ICML. 10524–10533. http://proceedings.mlr.press/v119/xiong20b.html
  152. 152.Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. 2021. Nyströmformer: A Nyström-based Algorithm for Approximating Self-Attention. (2021).
  153. 153.Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin. 2019. Understanding and Improving Layer Normalization. In Proceedings of NeurIPS. 4383–4393. https://proceedings.neurips.cc/paper/2019/hash/2f4fe03d77724a7217006e5d16728874-Abstract.html
  154. 154.Hang Yan, Bocao Deng, Xiaonan Li, and Xipeng Qiu. 2019. TENER: Adapting transformer encoder for named entity recognition. arXiv preprint arXiv:1911.04474 (2019).
  155. 155.An Yang, Junyang Lin, Rui Men, Chang Zhou, Le Jiang, Xianyan Jia, Ang Wang, Jie Zhang, Jiamang Wang, Yong Li, Di Zhang, Wei Lin, Lin Qu, Jingren Zhou, and Hongxia Yang. 2021. Exploring Sparse Expert Models and Beyond. arXiv:2105.15082 [cs.LG]
  156. 156.Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, and Tong Zhang. 2018. Modeling Localness for Self-Attention Networks. In Proceedings of EMNLP. Brussels, Belgium, 4449–4458. https://doi.org/10.18653/v1/D18-1475
  157. 157.Yilin Yang, Longyue Wang, Shuming Shi, Prasad Tadepalli, Stefan Lee, and Zhaopeng Tu. 2020. On the Sub-layer Functionalities of Transformer Decoder. In Findings of EMNLP. Online, 4799–4811. https://doi.org/10.18653/v1/2020.findings-emnlp.432
  158. 158.Zihao Ye, Qipeng Guo, Quan Gan, Xipeng Qiu, and Zheng Zhang. 2019. BP-Transformer: Modelling Long-Range Context via Binary Partitioning. arXiv:1911.04070 [cs.CL]
  159. 159.Chengxuan Ying, Guolin Ke, Di He, and Tie-Yan Liu. 2021. LazyFormer: Self Attention with Lazy Update. CoRR abs/2102.12702 (2021). arXiv:2102.12702
  160. 160.Davis Yoshida, Allyson Ettinger, and Kevin Gimpel. 2020. Adding Recurrence to Pretrained Transformers for Improved Efficiency and Context Size. CoRR abs/2008.07027 (2020). arXiv:2008.07027
  161. 161.Weiqiu You, Simeng Sun, and Mohit Iyyer. 2020. Hard-Coded Gaussian Attention for Neural Machine Translation. In Proceedings of ACL. Online, 7689–7700. https://doi.org/10.18653/v1/2020.acl-main.687
  162. 162.Weiwei Yu, Jian Zhou, HuaBin Wang, and Liang Tao. 2021. SETransformer: Speech Enhancement Transformer. Cognitive Computation (02 2021). https://doi.org/10.1007/s12559-020-09817-2
  163. 163.Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020. Big Bird: Transformers for Longer Sequences. arXiv:2007.14062 [cs.LG]
  164. 164.Biao Zhang, Deyi Xiong, and Jinsong Su. 2018. Accelerating Neural Transformer via an Average Attention Network. In Proceedings of ACL. Melbourne, Australia, 1789–1798. https://doi.org/10.18653/v1/P18-1166
  165. 165.Hang Zhang, Yeyun Gong, Yelong Shen, Weisheng Li, Jiancheng Lv, Nan Duan, and Weizhu Chen. 2021. Poolingformer: Long Document Modeling with Pooling Attention. arXiv:2105.04371
  166. 166.Xingxing Zhang, Furu Wei, and Ming Zhou. 2019. HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document Summarization. In Proceedings of ACL. Florence, Italy, 5059–5069. https://doi.org/10.18653/v1/P19-1499
  167. 167.Yuekai Zhao, Li Dong, Yelong Shen, Zhihua Zhang, Furu Wei, and Weizhu Chen. 2021. Memory-Efficient Differentiable Transformer Architecture Search. arXiv:2105.14669 [cs.LG]
  168. 168.Minghang Zheng, Peng Gao, Xiaogang Wang, Hongsheng Li, and Hao Dong. 2020. End-to-End Object Detection with Adaptive Clustering Transformer. CoRR abs/2011.09315 (2020). arXiv:2011.09315
  169. 169.Yibin Zheng, Xinhui Li, Fenglong Xie, and Li Lu. 2020. Improving End-to-End Speech Synthesis with Local Recurrent Neural Network Enhanced Transformer. In Proceedings of ICASSP. 6734–6738. https://doi.org/10.1109/ICASSP40776.2020.9054148
  170. 170.Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. In Proceedings of AAAI.
  171. 171.Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. 2020. BERT Loses Patience: Fast and Robust Inference with Early Exit. arXiv:2006.04152
  172. 172.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2020. Deformable DETR: Deformable Transformers for End-to-End Object Detection. CoRR abs/2010.04159 (2020). arXiv:2010.04159

Citation

MLA
Lin, T., et al. “A Survey of Transformers”. arXiv, 2021, http://arxiv.org/abs/2106.04554v2.
APA
Lin, T., Wang, Y., Liu, X., & Qiu, X. (2021). A Survey of Transformers. arXiv. http://arxiv.org/abs/2106.04554v2
Chicago
Lin, T., Y. Wang, X. Liu, and X. Qiu. 2021. “A Survey of Transformers”. arXiv. http://arxiv.org/abs/2106.04554v2.
Harvard
Lin, T. et al. (2021) “A Survey of Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.04554v2.
Vancouver
1. Lin T, Wang Y, Liu X, Qiu X (2021) A Survey of Transformers. arXiv

BibTeX

@article{lin2021survey,
  title = {A Survey of Transformers},
  author = {Lin, Tianyang and Wang, Yuxin and Liu, Xiangyang and Qiu, Xipeng},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.04554v2},
  eprint = {2106.04554}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF