Compressing Transformers: Features Are Low-Rank, but Weights Are Not!
Hao YuJianxin Wu
Reveals that transformer activations are low-rank even when their weights are not, introducing an unsupervised, few-shot feature-mimicking framework that sharply reduces model parameters and increases throughput across vision and language tasks with minimal accuracy loss.
Modern transformer models drive leading achievements in computer vision and natural language processing, but their enormous parameter counts and high computational demands restrict practical deployment on resource-constrained platforms. Conventional low-rank compression methods focus on factorizing model weight matrices, but they typically suffer sharp accuracy drops and require slow, data-intensive retraining on entire datasets—an obstacle in scenarios demanding rapid deployment or strict data privacy.
The article demonstrates that while transformer weight matrices are nearly full-rank and difficult to compress directly, their internal output features (activations) exhibit strong low-rank characteristics. Based on this insight, the article evaluates a fast, unsupervised compression framework that factorizes layer activations rather than weights using only a tiny fraction of unlabeled data.
The approach introduces Atomic Feature Mimicking to approximate output activations layer-by-layer, an adaptive search mechanism (Adaptive Atomic Feature Mimicking) that allocates compression levels based on each layer's sensitivity, and Global Feature Mimicking to correct accumulated network errors by aligning penultimate representations. The authors evaluated this framework across computer vision tasks—including standard benchmarks (ImageNet-1K), downstream image classification across multiple datasets, and object detection and segmentation (MS COCO2017)—as well as language modeling (WikiText-103), using proxy datasets of only 2,000 to 4,000 unlabeled samples.
The findings confirm substantial performance advantages over traditional weight-factorization techniques. When applied to the DeiT-B vision model, the proposed method eliminated 33% of parameters and increased processing throughput by 18.8% with only a 0.23% loss in classification accuracy; removing 40% of parameters caused just a 0.57% accuracy reduction while improving throughput by 24.5%. Across Swin Transformer variants, the method outperformed standard Singular Value Decomposition by roughly 4 to nearly 7 percentage points in retained accuracy for the same 33% parameter reduction. Furthermore, the compressed models transferred effectively to downstream vision and language tasks without substantial degradation, occasionally outperforming original full-size baselines on small classification datasets.
These results demonstrate that engineering teams can rapidly compress state-of-the-art transformer architectures in roughly one GPU hour without requiring expensive labeled data or risking privacy exposure. Decision-makers can achieve notable reductions in memory and hardware costs while retaining baseline predictive accuracy, solving a persistent trade-off in edge and cloud deployments.
Organizations aiming to reduce deployment footprints should adopt feature-mimicking strategies in place of standard weight decomposition for transformer compression pipelines. When applying this method, teams should avoid using supervised labels during distillation, as the article finds label-based fine-tuning on few-shot data increases overfitting risk relative to unsupervised feature alignment.
While the framework consistently preserves model accuracy, stakeholders should note that decomposing single linear layers into two sequential layers limits inference throughput gains relative to raw parameter reduction. The evaluations rely on randomly sampled proxy data and a greedy layer-allocation heuristic, meaning further throughput optimization, broader model architecture validation (such as convolutional networks), and formal proxy sampling strategies represent necessary areas for ongoing development.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). Introduces Data-efficient Image Transformers (DeiT) and attention-based distillation, providing the primary vision transformer baseline and distillation foundation compressed and evaluated by the source.
- Paper: Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation, Emily L. Denton et al. (2014). Pioneers post-training low-rank matrix and tensor decompositions to compress neural network layers, establishing the traditional weight-factorization paradigm that the source contrasts against and improves upon.
- Paper: Predicting Parameters in Deep Learning, Misha Denil et al. (2013). Demonstrates structural parameter redundancy and low-rank factorizability in deep neural networks, laying the theoretical groundwork for parameter reduction via decomposition.
- Paper: Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer, Sergey Zagoruyko et al. (2017). Establishes intermediate feature and attention map alignment between teacher and student representations, inspiring the layer-wise feature mimicking strategies developed in the source.
- Paper: Linformer: Self-Attention with Linear Complexity, Sinong Wang et al. (2020). Demonstrates that self-attention mechanisms exhibit low-rank properties, motivating the source's exploration of low-rank structure within transformer internal features.
- Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). Provides a comprehensive taxonomy and evaluation of efficient transformer methods, including low-rank factorizations and sequence approximations that frame the source's design context.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). Extends the principle that transformer activations dictate compressibility to low-bit post-training quantization without backpropagation.
- Paper: SliceGPT: Compress Large Language Models by Deleting Rows and Columns, Saleh Ashkboos et al. (2024). Leverages internal representation variance and principal component analysis on calibration samples to perform structured pruning on large transformer models.
- Paper: DepGraph: Towards Any Structural Pruning, Gongfan Fang et al. (2023). Develops a generalized dependency graph framework to automate structural parameter reduction across interconnected transformer and neural layers.
- Paper: A Smaller Transformer in Your Transformer, Dhananjay Tomar et al. (2026). Addresses transformer redundancy by collapsing contiguous vision transformer layers through deep distillation, complementing layer-wise feature factorization.
