Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Adaptive Atomic Feature

An adaptive atomic feature is an intermediate activation representation within a neural network that is dynamically compressed or factorized at the individual layer level based on localized sensitivity metrics. Instead of uniformly modifying static model weights, this approach leverages the low-rank structure of activation outputs across fundamental network components, adaptively assigning rank capacities and compression ratios to preserve representation quality in sensitive layers while aggressively reducing redundancy in less critical ones. By focusing on localized feature factorization rather than direct weight manipulation, adaptive atomic features facilitate efficient neural network compression and computational acceleration while maintaining overall model performance.

1 item

Compressing Transformers: Features Are Low-Rank, but Weights Are Not!

Compressing Transformers: Features Are Low-Rank, but Weights Are Not!

Hao Yu, Jianxin Wu

OrganizationsNanjing University

Why you should read this

Reveals that transformer activations are low-rank even when their weights are not, introducing an unsupervised, few-shot feature-mimicking framework that sharply reduces model parameters and increases throughput across vision and language tasks with minimal accuracy loss.

Transformer and its variants achieve excellent results in various computer vision and natural language processing tasks, but high computational costs and reliance on large training datasets restrict their deployment in resource-constrained settings. Low-rank approximation of model weights has been effective in compressing CNN models, but its application to transformers has been less explored and is less effective. Existing methods require the complete dataset to fine-tune compressed models, which are both time-consuming and data-hungry. This paper reveals that the features (i.e., activations) are low-rank, but model weights are surprisingly not low-rank. Hence, AAFM is proposed, which adaptively determines the compressed model structure and locally compresses each linear layer’s output features rather than the model weights. A second stage, GFM, optimizes the entire compressed network holistically. Both AAFM and GFM only use few training samples without labels, that is, they are few-shot, unsupervised, fast and effective. For example, with only 2K images without labels, 33% of the parameters are removed in DeiT-B with 18.8% relative throughput increase, but only a 0.23% accuracy loss for ImageNet recognition. The proposed methods are successfully applied to the language modeling task in NLP, too. Besides, the few-shot compressed models generalize well in downstream tasks.

Added

2026-09-26