Compressing Transformers: Features Are Low-Rank, but Weights Are Not!

Hao YuJianxin Wu

article2023AAAI57 citations

Reveals that transformer activations are low-rank even when their weights are not, introducing an unsupervised, few-shot feature-mimicking framework that sharply reduces model parameters and increases throughput across vision and language tasks with minimal accuracy loss.

Listen

Modern transformer models drive leading achievements in computer vision and natural language processing, but their enormous parameter counts and high computational demands restrict practical deployment on resource-constrained platforms. Conventional low-rank compression methods focus on factorizing model weight matrices, but they typically suffer sharp accuracy drops and require slow, data-intensive retraining on entire datasets—an obstacle in scenarios demanding rapid deployment or strict data privacy.

The article demonstrates that while transformer weight matrices are nearly full-rank and difficult to compress directly, their internal output features (activations) exhibit strong low-rank characteristics. Based on this insight, the article evaluates a fast, unsupervised compression framework that factorizes layer activations rather than weights using only a tiny fraction of unlabeled data.

The approach introduces Atomic Feature Mimicking to approximate output activations layer-by-layer, an adaptive search mechanism (Adaptive Atomic Feature Mimicking) that allocates compression levels based on each layer's sensitivity, and Global Feature Mimicking to correct accumulated network errors by aligning penultimate representations. The authors evaluated this framework across computer vision tasks—including standard benchmarks (ImageNet-1K), downstream image classification across multiple datasets, and object detection and segmentation (MS COCO2017)—as well as language modeling (WikiText-103), using proxy datasets of only 2,000 to 4,000 unlabeled samples.

The findings confirm substantial performance advantages over traditional weight-factorization techniques. When applied to the DeiT-B vision model, the proposed method eliminated 33% of parameters and increased processing throughput by 18.8% with only a 0.23% loss in classification accuracy; removing 40% of parameters caused just a 0.57% accuracy reduction while improving throughput by 24.5%. Across Swin Transformer variants, the method outperformed standard Singular Value Decomposition by roughly 4 to nearly 7 percentage points in retained accuracy for the same 33% parameter reduction. Furthermore, the compressed models transferred effectively to downstream vision and language tasks without substantial degradation, occasionally outperforming original full-size baselines on small classification datasets.

These results demonstrate that engineering teams can rapidly compress state-of-the-art transformer architectures in roughly one GPU hour without requiring expensive labeled data or risking privacy exposure. Decision-makers can achieve notable reductions in memory and hardware costs while retaining baseline predictive accuracy, solving a persistent trade-off in edge and cloud deployments.

Organizations aiming to reduce deployment footprints should adopt feature-mimicking strategies in place of standard weight decomposition for transformer compression pipelines. When applying this method, teams should avoid using supervised labels during distillation, as the article finds label-based fine-tuning on few-shot data increases overfitting risk relative to unsupervised feature alignment.

While the framework consistently preserves model accuracy, stakeholders should note that decomposing single linear layers into two sequential layers limits inference throughput gains relative to raw parameter reduction. The evaluations rely on randomly sampled proxy data and a greedy layer-allocation heuristic, meaning further throughput optimization, broader model architecture validation (such as convolutional networks), and formal proxy sampling strategies represent necessary areas for ongoing development.

Cover for Compressing Transformers: Features Are Low-Rank, but Weights Are Not!

Abstract

Transformer and its variants achieve excellent results in various computer vision and natural language processing tasks, but high computational costs and reliance on large training datasets restrict their deployment in resource-constrained settings. Low-rank approximation of model weights has been effective in compressing CNN models, but its application to transformers has been less explored and is less effective. Existing methods require the complete dataset to fine-tune compressed models, which are both time-consuming and data-hungry. This paper reveals that the features (i.e., activations) are low-rank, but model weights are surprisingly not low-rank. Hence, AAFM is proposed, which adaptively determines the compressed model structure and locally compresses each linear layer’s output features rather than the model weights. A second stage, GFM, optimizes the entire compressed network holistically. Both AAFM and GFM only use few training samples without labels, that is, they are few-shot, unsupervised, fast and effective. For example, with only 2K images without labels, 33% of the parameters are removed in DeiT-B with 18.8% relative throughput increase, but only a 0.23% accuracy loss for ImageNet recognition. The proposed methods are successfully applied to the language modeling task in NLP, too. Besides, the few-shot compressed models generalize well in downstream tasks.

Table of Contents

  • Introduction
  • Related Works
  • Transformers
  • Low-Rank Approximation for Transformers
  • Feature Mimicking
  • The Proposed Methods
  • Preliminaries
  • Atomic Feature Mimicking (AFM)
  • Adaptive Atomic Feature Mimicking (AAFM)
  • Global Feature Mimicking (GFM)
  • Experiments
  • Datasets and Metrics
  • Compressing DeiT & Swin
  • Transferring Ability
  • Compressing Transformer for NLP Tasks
  • Analyses
  • Discussions and Conclusions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Low-Rankness of Transformer Features Versus Model Weights

    empirical result

    In vision transformer architectures such as DeiT-B (which contains 12 blocks, each composed of a multi-head self-attention module with Query-Key-Value and projection layers, plus a two-layer Feed-Forward Network with FC1\text{FC1} and FC2\text{FC2}), linear layer weight matrices are nearly full-rank, whereas output activation feature maps exhibit strong low-rank properties.

    Evaluating the proportion of singular vectors (for weights via Singular Value Decomposition) or covariance matrix eigenvectors (for input and output activations via Principal Component Analysis) required to retain 90% of the total energy on the ImageNet-1K validation set reveals that:

    • Output activation features consistently require the smallest fraction of dimensions (between 10% and 50% across transformer blocks).
    • Input activation features require higher dimensionality than output features.
    • Weight matrices require the highest proportion of singular vectors across all layers (often requiring 60% to 75% of total dimensions).

    Because output activations reside in a lower-dimensional subspace than the weight matrices that produce them, factorizing transformations to approximate output features preserves representational capacity far more effectively than decomposing model weights directly.

  2. Knowl 2 — Atomic Feature Mimicking Layer Decomposition

    model/method

    Atomic Feature Mimicking (AFM) replaces an original fully connected layer y=Wx+by = Wx + b (with weight W∈Rm×nW \in \mathbb{R}^{m \times n}, bias b∈Rmb \in \mathbb{R}^m, input x∈Rn×cx \in \mathbb{R}^{n \times c}, and output y∈Rm×cy \in \mathbb{R}^{m \times c}, treating columns as cc token instantiations of a random feature vector y∈Rmy \in \mathbb{R}^m) with two sequential linear layers parameterized by a target low rank k<min⁡(m,n)k < \min(m, n) by approximating the output activations.

    The covariance matrix of the layer output activations over an unlabeled proxy dataset D\mathcal{D} is:

    Cov(y)=E[yyT]−E[y]E[y]T\text{Cov}(y) = \mathbb{E}[yy^T] - \mathbb{E}[y]\mathbb{E}[y]^T

    where E[⋅]\mathbb{E}[\cdot] is the expectation operator. Applying eigendecomposition (Principal Component Analysis) yields:

    Cov(y)=USUT\text{Cov}(y) = U S U^T

    where U∈Rm×mU \in \mathbb{R}^{m \times m} is an orthonormal eigenvector matrix and S∈Rm×mS \in \mathbb{R}^{m \times m} is a diagonal matrix containing sorted eigenvalues in descending order.

    Selecting the top kk eigenvectors gives Uk∈Rm×kU_k \in \mathbb{R}^{m \times k} satisfying UkUkT≈IU_k U_k^T \approx I. Approximating centered features as y−E[y]≈UkUkT(y−E[y])y - \mathbb{E}[y] \approx U_k U_k^T(y - \mathbb{E}[y]) transforms the single linear layer into two sequential operations:

    y≈Uk(UkTWx+UkTb)+E[y]−UkUkTE[y]y \approx U_k\left(U_k^T W x + U_k^T b\right) + \mathbb{E}[y] - U_k U_k^T \mathbb{E}[y]

    The first linear layer has weight W1=UkTW∈Rk×nW_1 = U_k^T W \in \mathbb{R}^{k \times n} and bias b1=UkTb∈Rkb_1 = U_k^T b \in \mathbb{R}^k. The second linear layer has weight W2=Uk∈Rm×kW_2 = U_k \in \mathbb{R}^{m \times k} and bias b2=E[y]−UkUkTE[y]∈Rmb_2 = \mathbb{E}[y] - U_k U_k^T \mathbb{E}[y] \in \mathbb{R}^m. This factorization reduces the layer parameter count from O(mn)O(mn) to O((m+n)k)O((m+n)k).

  3. Knowl 3 — Atomic Feature Mimicking Algorithm

    algorithm

    Atomic Feature Mimicking (AFM) calculates the decomposed weights and biases for a given linear layer using an unlabeled proxy dataset D\mathcal{D} without backpropagation.

    Input: Model MM with weight W∈Rm×nW \in \mathbb{R}^{m \times n} and bias b∈Rmb \in \mathbb{R}^m at layer ii, proxy dataset D\mathcal{D}, rank kk.
    Output: Decomposed layer parameters (W1,b1)(W_1, b_1) and (W2,b2)(W_2, b_2).
    Initialize running moments for E[y]\mathbb{E}[y] and E[yyT]\mathbb{E}[yy^T].
    for each sample x∈Dx \in \mathcal{D} do
        Forward propagate M(x)M(x) to obtain layer ii output features y∈Rm×cy \in \mathbb{R}^{m \times c}.
        Update running estimates of E[yyT]\mathbb{E}[yy^T] and E[y]\mathbb{E}[y] in a streaming fashion.
    end for
    Compute covariance Cov(y)=E[yyT]−E[y]E[y]T\text{Cov}(y) = \mathbb{E}[yy^T] - \mathbb{E}[y]\mathbb{E}[y]^T.
    Compute eigendecomposition Cov(y)=USUT\text{Cov}(y) = U S U^T, where U∈Rm×mU \in \mathbb{R}^{m \times m} is orthonormal and eigenvalues in SS are in descending order.
    Extract top kk columns of UU into Uk∈Rm×kU_k \in \mathbb{R}^{m \times k}.
    W1←UkTWW_1 \leftarrow U_k^T W
    b1←UkTbb_1 \leftarrow U_k^T b
    W2←UkW_2 \leftarrow U_k
    b2←E[y]−UkUkTE[y]b_2 \leftarrow \mathbb{E}[y] - U_k U_k^T \mathbb{E}[y]
    return (W1,b1),(W2,b2)(W_1, b_1), (W_2, b_2)

    Moments E[y]\mathbb{E}[y] and E[yyT]\mathbb{E}[yy^T] are accumulated across all layers simultaneously in a single forward pass over D\mathcal{D}, avoiding the storage of all intermediate activations in GPU memory.

  4. Knowl 4 — Adaptive Atomic Feature Mimicking Layer Rank Search

    model/method

    Adaptive Atomic Feature Mimicking (AAFM) determines layer-specific ranks kik_i across ll linear layers under an overall model parameter budget PtarP_{\text{tar}}.

    The sensitivity Si(k)S_i(k) of the ii-th linear layer at rank kk is evaluated on an unlabeled proxy dataset D\mathcal{D} using the Kullback-Leibler (KL) divergence between original model output logits and logits obtained after applying AFM solely to layer ii:

    Si(k)=∑x∈DDKL(M(x,w)∥M(x,wi(k)))S_i(k) = \sum_{x \in \mathcal{D}} D_{\text{KL}}\left(M(x, w) \parallel M(x, w_i(k))\right)

    where M(x,w)M(x, w) represents the output class probability distribution of the original model with full weights ww, and M(x,wi(k))M(x, w_i(k)) represents the distribution when only layer ii is replaced by AFM at rank kk. To maximize GPU hardware efficiency and constrain search time, rank candidate values kk are constrained to multiples of 32.

    The global layer rank allocation is formulated as an integer programming problem:

    min⁡{ki}i=1l∑i=1lSi(ki)s.t.∑i=1lPi(ki)≤Ptar\min_{\{k_i\}_{i=1}^l} \sum_{i=1}^l S_i(k_i) \quad \text{s.t.} \quad \sum_{i=1}^l P_i(k_i) \le P_{\text{tar}}

    where Pi(ki)P_i(k_i) is the number of parameters in the ii-th layer configured at rank kik_i. Under the simplifying assumption of layer sensitivity independence, this optimization is solved with a greedy search algorithm that iteratively reduces ranks in the least sensitive layers.

  5. Knowl 5 — Global Feature Mimicking Fine-Tuning

    model/method

    Global Feature Mimicking (GFM) fine-tunes the entire compressed transformer model following Adaptive Atomic Feature Mimicking (AAFM) to correct accumulated inter-layer approximation errors using only the few-shot unlabeled proxy dataset D\mathcal{D}.

    The optimization minimizes the Mean Squared Error (MSE) between the penultimate-layer activation representations of the compressed and original models:

    LGFM=LMSE(fcL,foL)=1N∥fcL−foL∥F2\mathcal{L}_{\text{GFM}} = \mathcal{L}_{\text{MSE}}\left(f_c^L, f_o^L\right) = \frac{1}{N} \left\| f_c^L - f_o^L \right\|_F^2

    where fcLf_c^L and foLf_o^L are the activation feature maps at layer index LL from the compressed network and the original network, respectively, and NN is the total number of elements. For vision transformers, LL corresponds to the activations following the final LayerNorm before global average pooling; for transformer language models, LL is the feature representation after the final transformer block prior to the adaptive softmax layer. The final linear classification head is excluded from GFM fine-tuning.

    During few-shot GFM fine-tuning, heavy regularizers that impair late-stage convergence (random erasing, RandAugment, and layer dropout/stochastic depth) are removed, while label-free augmentations (Mixup, CutMix, horizontal flipping, color jittering) and cosine learning rate schedules are preserved.

  6. Knowl 6 — Low-Rank Compression Performance on Vision Transformers

    data/table

    DeiT-B, Swin-B, and Swin-L were compressed using AAFM and fine-tuned with GFM using an unlabeled proxy dataset of 2,000 randomly selected ImageNet-1K training images. Inference throughput was measured on a single NVIDIA GeForce RTX 3090 GPU with a fixed batch size of 512.

    Model Throughput (img/s) #Param. (M) Top-1 Acc. (%)
    DeiT-B 619.46 86.57 81.85
    +SVD 741.07 (+19.6%) 58.27 (-33%) 77.21
    +GFM – 58.27 (-33%) 80.36
    +AAFM 682.23 (+10.1%) 69.25 (-20%) 81.76
    +GFM – 69.25 (-20%) 81.83
    +AAFM 735.97 (+18.8%) 58.26 (-33%) 81.21
    +GFM – 58.26 (-33%) 81.62
    +AAFM 771.07 (+24.5%) 51.95 (-40%) 80.33
    +GFM – 51.95 (-40%) 81.28
    Swin-B 458.86 88.10 83.47
    +SVD 489.95 (+6.8%) 60.20 (-33%) 74.30
    +GFM – 60.20 (-33%) 81.13
    +AAFM 471.64 (+2.8%) 70.50 (-20%) 82.89
    +GFM – 70.50 (-20%) 83.19
    +AAFM 477.71 (+4.1%) 66.09 (-25%) 82.41
    +GFM – 66.09 (-25%) 83.00
    +AAFM 489.46 (+6.7%) 60.20 (-33%) 81.15
    +GFM – 60.20 (-33%) 82.68
    Swin-L 257.40 196.87 86.25
    +SVD 288.82 (+12.2%) 134.09 (-33%) 82.02
    +GFM – 134.09 (-33%) 84.52
    +AAFM 275.14 (+6.9%) 157.52 (-20%) 85.94
    +GFM – 157.52 (-20%) 86.01
    +AAFM 282.58 (+9.8%) 147.67 (-25%) 85.73
    +GFM – 147.67 (-25%) 85.83
    +AAFM 292.04 (+13.5%) 134.09 (-33%) 85.04
    +GFM – 134.09 (-33%) 85.44

    Prior to fine-tuning, AAFM outperforms traditional SVD low-rank decomposition by 4.0 percentage points on DeiT-B (81.21% vs. 77.21%) and by 6.85 percentage points on Swin-B (81.15% vs. 74.30%) at 33% parameter reduction. After GFM fine-tuning, DeiT-B with 33% fewer parameters retains 81.62% accuracy (only 0.23% drop from the 81.85% uncompressed baseline) while accelerating throughput by 18.8%.

  7. Knowl 7 — Transfer Performance of Compressed Vision Transformers on Downstream Tasks

    data/table

    Models compressed via AAFM and GFM retain feature generalization when used as backbones in downstream object detection/instance segmentation (Cascade Mask R-CNN trained on MS COCO 2017 with a 3×3\times 36-epoch schedule) and small-scale image classification (fine-tuned for 100 epochs).

    On MS COCO 2017 validation using Swin-B compressed by 33%:

    Backbone Tasks AP AP50_{50} AP75_{75} APS_S APM_M APL_L
    Swin-B Detection 52.0 70.8 56.4 35.0 55.6 67.4
    Ours (-33% Param.) Detection 51.9 70.5 56.4 35.5 55.8 67.0
    Swin-B Segmentation 45.0 68.3 48.8 28.5 48.6 60.6
    Ours (-33% Param.) Segmentation 44.7 67.9 48.5 28.8 48.4 59.7

    On small-scale classification datasets using DeiT-B and its compressed sub-models (reporting Top-1 accuracy in %):

    Dataset DeiT-B (86.57M) Ours (69.25M) Ours (58.26M) Ours (51.95M)
    CIFAR-100 90.99 90.67 90.37 90.17
    CUB-200 85.88 85.07 85.38 84.85
    Cars 90.45 91.18 90.66 90.72
    Aircraft 79.87 80.92 81.19 80.80
    Pets 94.74 94.22 93.98 93.95
    Flowers 97.77 97.45 97.30 97.02
    iNaturalist-2019 77.39 77.56 76.70 77.13

    Compressed models closely match original baseline detection/segmentation mAPs and fine-tuning classification accuracies across all datasets, even surpassing the uncompressed DeiT-B on Stanford Cars, FGVC Aircraft, and iNaturalist-2019.

  8. Knowl 8 — Compression Performance on Transformer Language Modeling

    data/table

    A 16-decoder-block transformer language model with adaptive input representations and 8-head multi-head self-attention on WikiText-103 (246.9M parameters, baseline perplexity 18.66 with context window 2560) was compressed via AAFM and fine-tuned via GFM using an unlabeled proxy dataset D\mathcal{D} of 4,000 sampled sentences.

    Model Throughput #Param. (M) Perplexity
    Baseline Transformer 3137.3 246.9 18.66
    +SVD 3303.3 (+5.3%) 196.6 (-20%) 29.76
    +GFM – 196.6 (-20%) 20.24
    +AAFM 3252.3 (+3.7%) 209.9 (-15%) 20.23
    +GFM – 209.9 (-15%) 19.07
    +AAFM 3293.2 (+5.0%) 196.6 (-20%) 22.34
    +GFM – 196.6 (-20%) 19.46
    +AAFM 3356.4 (+7.0%) 185.2 (-25%) 26.20
    +GFM – 185.2 (-25%) 20.05

    At 20% parameter reduction, AAFM achieves a perplexity of 22.34 prior to fine-tuning, outperforming standard SVD (29.76) by 7.42 perplexity points. After GFM fine-tuning, the compressed model reaches 19.46 perplexity with 5.0% higher throughput and 20% fewer parameters.

  9. Knowl 9 — Distillation Strategy Comparison in Few-Shot Compression

    data/table

    Fine-tuning strategies were compared on Swin-B compressed by 33% on ImageNet-1K (using 2,000 proxy images) and Transformer compressed by 25% on WikiText-103 (using 4,000 proxy sentences). For supervised variants, ground-truth cross-entropy LCE(p,y)\mathcal{L}_{\text{CE}}(p, y) was combined with the distillation objective with weight α=1.0\alpha = 1.0.

    Distillation Strategy Formulation Top-1 Acc. (%) Perplexity
    Soft Distillation w/ label: LCE(p,y)+αLKL(p,q)\mathcal{L}_{\text{CE}}(p, y) + \alpha \mathcal{L}_{\text{KL}}(p, q) 82.34 24.81
    Soft Distillation w/o label: LKL(p,q)\mathcal{L}_{\text{KL}}(p, q) 82.38 20.53
    Hard Distillation w/ label: LCE(p,y)+αLCE(p,yt)\mathcal{L}_{\text{CE}}(p, y) + \alpha \mathcal{L}_{\text{CE}}(p, y_t) 81.05 323.95
    Hard Distillation w/o label: LCE(p,yt)\mathcal{L}_{\text{CE}}(p, y_t) 82.08 1.2×1051.2 \times 10^5
    GFM w/ label: LCE(p,y)+αLMSE(fcL,foL)\mathcal{L}_{\text{CE}}(p, y) + \alpha \mathcal{L}_{\text{MSE}}(f_c^L, f_o^L) 78.81 20.10
    GFM w/o label: LMSE(fcL,foL)\mathcal{L}_{\text{MSE}}(f_c^L, f_o^L) 82.68 20.05

    where pp is the compressed model output distribution, qq is the teacher output distribution, yy is the ground-truth label, yt=arg⁡max⁡(q)y_t = \arg\max(q) is the teacher hard decision, and fcL,foLf_c^L, f_o^L are the penultimate-layer features of the compressed and original networks.

    Incorporating ground-truth label loss LCE(p,y)\mathcal{L}_{\text{CE}}(p, y) under the few-shot proxy regime induces overfitting, degrading Top-1 accuracy in GFM from 82.68% to 78.81%. Unsupervised penultimate-layer feature mimicking (GFM w/o label) achieves the best performance across both computer vision and NLP tasks.

  10. Knowl 10 — Sensitivity of Feature Mimicking to Proxy Dataset Size

    data/table

    The effect of the proxy dataset size ∣D∣|\mathcal{D}| on AAFM decomposition and subsequent GFM fine-tuning was evaluated on Swin-B with 33% parameter removal on ImageNet-1K.

    Proxy Dataset Size ∣D∣|\mathcal{D}| 1K 2K 5K 10K 100K 1.28M (Full)
    AAFM Top-1 Acc. (%) 81.21 81.15 81.20 81.31 81.30 81.22
    GFM Top-1 Acc. (%) 82.38 82.68 82.75 82.88 82.97 82.99

    AAFM layer sensitivity evaluation and decomposition quality are stable across all proxy sizes from 1,000 samples to the full 1.28M dataset (varying within 81.15% to 81.31%). GFM fine-tuning accuracy exhibits diminishing returns beyond 2,000 samples: increasing the dataset size from 2,000 samples to 1.28 million samples yields only a +0.31% increase in final Top-1 accuracy (82.68% vs. 82.99%), validating that 2,000 unlabeled images are sufficient for effective compression.

  11. Knowl 11 — Practical Limitations of Low-Rank Transformer Approximation

    limitation

    The low-rank feature mimicking compression approach has two primary limitations:

    1. Modest Throughput Gains: Decomposing a linear layer y=Wx+by = Wx + b into two sequential lower-rank linear layers y≈W2(W1x+b1)+b2y \approx W_2(W_1 x + b_1) + b_2 doubles the layer count and increases kernel invocation and memory access overhead. Consequently, relative throughput increases (e.g., +6.7% for Swin-B and +18.8% for DeiT-B at 33% parameter reduction) are substantially lower than the corresponding parameter reductions.
    2. Random Proxy Sampling: Samples in the proxy dataset D\mathcal{D} are selected uniformly at random without assessing the impact of data distribution, class balance, or targeted sample selection on decomposition quality.

Coverage note — Ablation results in Table 5 comparing uniform AFM against Adaptive SVD are omitted as their findings (AFM outperforms SVD and adaptive layer allocation outperforms uniform rank allocation) are fully subsumed by Knowls 4, 6, and 8.

References

  1. 1.Baevski, A.; and Auli, M. 2018. Adaptive Input Representations for Neural Language Modeling. In International Conference on Learning Representations (ICLR).
  2. 2.Bai, H.; Wu, J.; King, I.; and Lyu, M. 2020. Few shot network compression via cross distillation. In Proceedings of the AAAI Conference on Artificial Intelligence, 3203–3210.
  3. 3.Cai, Z.; and Vasconcelos, N. 2018. Cascade R-cnn: Delving into high quality object detection. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6154–6162.
  4. 4.Chen, P.; Yu, H.-F.; Dhillon, I.; and Hsieh, C.-J. 2021. Drone: Data-aware low-rank compression for large nlp models. In Advances in Neural Information Processing Systems, volume 34, 29321–29334.
  5. 5.Cubuk, E. D.; Zoph, B.; Shlens, J.; and Le, Q. V. 2020. RandAugment: Practical automated data augmentation with a reduced search space. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 702–703.
  6. 6.Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. ImageNet: A large-scale hierarchical image database. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 248–255.
  7. 7.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations (ICLR).
  8. 8.Golub, G. H.; and Van Loan, C. F. 2013. Matrix computations. Johns Hopkins University Press.
  9. 9.Hinton, G.; Vinyals, O.; and Dean, J. 2015. Distilling the Knowledge in a Neural Network. arXiv:1503.02531.
  10. 10.Hsu, Y.-C.; Hua, T.; Chang, S.; Lou, Q.; Shen, Y.; and Jin, H. 2022. Language model compression with weighted low-rank factorization. In International Conference on Learning Representations (ICLR).
  11. 11.Huang, G.; Sun, Y.; Liu, Z.; Sedra, D.; and Weinberger, K. Q. 2016. Deep Networks with Stochastic Depth. In The European Conference on Computer Vision (ECCV), volume 9908 of LNCS, 646–661. Springer.
  12. 12.Kenton, J. D. M.-W. C.; and Toutanova, L. K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT, 4171–4186.
  13. 13.Li, T.; Li, J.; Liu, Z.; and Zhang, C. 2020. Few sample knowledge distillation for efficient network compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14639–14647.
  14. 14.Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014. Microsoft COCO: Common Objects in Context. In The European Conference on Computer Vision (ECCV), volume 8693 of LNCS, 740–755. Springer.
  15. 15.Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10012–10022.
  16. 16.Loshchilov, I.; and Hutter, F. 2017. Sgdr: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations (ICLR).
  17. 17.Loshchilov, I.; and Hutter, F. 2018. Decoupled Weight Decay Regularization. In International Conference on Learning Representations (ICLR).
  18. 18.Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2017. Pointer Sentinel Mixture Models. In International Conference on Learning Representations (ICLR).
  19. 19.Noach, M. B.; and Goldberg, Y. 2020. Compressing pre-trained language models by matrix decomposition. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, 884–889.
  20. 20.Ott, M.; Auli, M.; Grangier, D.; and Ranzato, M. A. 2018. Analyzing uncertainty in neural machine translation. In International Conference on Machine Learning (ICML), 3956–3965.
  21. 21.Ott, M.; Edunov, S.; Baevski, A.; Fan, A.; Gross, S.; Ng, N.; Grangier, D.; and Auli, M. 2019. fairseq: A Fast, Extensible Toolkit for Sequence Modeling. In Proceedings of NAACL-HLT 2019: Demonstrations.
  22. 22.Romero, A.; Ballas, N.; Kahou, S. E.; Chassang, A.; Gatta, C.; and Bengio, Y. 2014. Fitnets: Hints for thin deep nets. In International Conference on Learning Representations (ICLR).
  23. 23.Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jegou, H. 2021. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning (ICML), 10347–10357.
  24. 24.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems, volume 30, 5998–6008.
  25. 25.Wang, G.-H.; Ge, Y.; and Wu, J. 2021. Distilling knowledge by mimicking features. IEEE Transactions on Pattern Analysis and Machine Intelligence.
  26. 26.Wang, H.; Liu, J.; Ma, X.; Yong, Y.; Chai, Z.; and Wu, J. 2022. Compressing Models With Few Samples: Mimicking Then Replacing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 701–710.
  27. 27.Wu, J. 2020. Essentials of Pattern Recognition: An Accessible Approach. Cambridge University Press.
  28. 28.Yu, H.; Wang, H.; and Wu, J. 2021. Mixup without hesitation. In International Conference on Image and Graphics, 143–154. Springer.
  29. 29.Yun, S.; Han, D.; Oh, S. J.; Chun, S.; Choe, J.; and Yoo, Y. 2019. CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 6023–6032.
  30. 30.Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2018. Mixup: Beyond Empirical Risk Minimization. In International Conference on Learning Representations (ICLR).
  31. 31.Zhong, Z.; Zheng, L.; Kang, G.; Li, S.; and Yang, Y. 2020. Random Erasing Data Augmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, 13001–13008.

Citation

MLA
Yu, H., and J. Wu. “Compressing Transformers: Features Are Low-Rank, but Weights Are Not!”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 9, 2023, pp. 11007–15, https://doi.org/10.1609/AAAI.V37I9.26304.
APA
Yu, H., & Wu, J. (2023). Compressing Transformers: Features Are Low-Rank, but Weights Are Not!. Proceedings of the AAAI Conference on Artificial Intelligence, 37(9), 11007–11015. https://doi.org/10.1609/AAAI.V37I9.26304
Chicago
Yu, H., and J. Wu. 2023. “Compressing Transformers: Features Are Low-Rank, but Weights Are Not!”. Proceedings of the AAAI Conference on Artificial Intelligence 37 (9): 11007–15. https://doi.org/10.1609/AAAI.V37I9.26304.
Harvard
Yu, H. and Wu, J. (2023) “Compressing Transformers: Features Are Low-Rank, but Weights Are Not!”, Proceedings of the AAAI Conference on Artificial Intelligence, 37(9), pp. 11007–11015. Available at: https://doi.org/10.1609/AAAI.V37I9.26304.
Vancouver
1. Yu H, Wu J (2023) Compressing Transformers: Features Are Low-Rank, but Weights Are Not!. Proceedings of the AAAI Conference on Artificial Intelligence 37:11007–11015

BibTeX

@article{Yu_2023, title={Compressing Transformers: Features Are Low-Rank, but Weights Are Not!}, volume={37}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V37I9.26304}, DOI={10.1609/aaai.v37i9.26304}, number={9}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Yu, Hao and Wu, Jianxin}, year={2023}, month=June, pages={11007–11015} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF