Frustratingly Easy Transferability Estimation
Long-Kai HuangJunzhou HuangYu RongQiang YangYing Wei
Proposes TransRate, an extremely lightweight, training-free metric based on coding rate that accurately estimates transferability and guides optimal layer selection for pre-trained models using only a single forward pass over target examples.
Transfer learning from existing models is essential for solving complex machine learning tasks with limited labeled data, but finding the right pre-trained architecture and identifying which specific internal layers to transfer remains a significant operational hurdle. Traditional techniques require expensive fine-tuning or full retraining to test compatibility, creating substantial computational costs and development delays. Meanwhile, earlier lightweight estimation methods fail to evaluate intermediate layers or rely on restricted assumptions about source data accessibility.
The article introduces and evaluates TransRate, an optimization-free transferability metric designed to rapidly predict target performance before training begins. TransRate measures the mutual information between extracted target features and their class labels using coding rate as an efficient proxy for entropy, assessing both feature completeness across classes and compactness within each class.
The approach was validated through extensive empirical testing across 32 pre-trained models—spanning supervised, self-supervised, convolutional, and graph neural networks—and 16 diverse downstream tasks covering image classification, molecular regression, and molecular classification. The authors benchmarked TransRate against several leading alternatives without requiring access to source training datasets or iterative target optimization.
The findings show that TransRate delivers superior rank and linear correlations with final downstream performance across source dataset selection, model architecture comparison, and individual layer selection. TransRate was the only evaluated method capable of consistently identifying the highest-performing internal layers to transfer, achieving perfect layer-ranking correlation in 9 out of 15 layer-selection experiments. Furthermore, it operates up to roughly 3,000 times faster than full fine-tuning grid searches, remains computationally stable across a wide range of distortion parameters, and maintains robust ranking accuracy even when target labeled sample sizes are sharply reduced.
These results indicate that organizations can drastically reduce compute costs, eliminate negative transfer risks, and shorten model selection cycles from days to minutes by adopting TransRate prior to downstream fine-tuning. Unlike prior lightweight metrics, TransRate accommodates unsupervised representations and custom layer selections without requiring access to proprietary source datasets.
Teams should deploy TransRate as a standardized, pre-training screening filter to rank candidate architectures and select optimal layer cutoffs before initiating resource-heavy tuning. While the framework demonstrates high empirical confidence across diverse modalities, practitioners should exercise care when applying it to continuous regression problems, as target values must be discretized into discrete bins, and in extreme few-shot scenarios where sample representations may degrade.
- Paper: How transferable are features in deep neural networks?, Jason Yosinski et al. (2014). Provides foundational empirical analyses on how feature specificity changes across network layers and how layer depth affects transferability, motivating TransRate's specific focus on layer-level transfer selection.
- Paper: Do Better ImageNet Models Transfer Better?, Simon Kornblith et al. (2018). Establishes systematic benchmarking for how pre-trained representation quality relates to downstream transfer performance, setting the baseline questions TransRate seeks to evaluate efficiently without expensive fine-tuning.
- Paper: Similarity of Neural Network Representations Revisited, Simon Kornblith et al. (2019). Introduces representation similarity analysis across neural network layers that directly informs how intermediate feature representations are evaluated prior to downstream transfer.
- Paper: Taskonomy: Disentangling Task Transfer Learning, Amir Zamir et al. (2018). Formalizes the problem of mapping transferability and task relationships across visual representations that lightweight transferability metrics aim to solve computationally.
- Paper: On the Role of Neural Collapse in Transfer Learning, Tomer Galanti et al. (2022). Provides theoretical insight into feature compactness and class separation in transfer learning that complements and explains TransRate's empirical coding rate measurements.
- Paper: Assaying Out-Of-Distribution Generalization in Transfer Learning, Florian Wenzel et al. (2022). Extends downstream transfer evaluation by assaying how architecture selection and fine-tuning transferability hold up under out-of-distribution shifts.
- Paper: Cross-Modal Fine-Tuning: Align then Refine, Junhong Shen et al. (2023). Applies cross-modal alignment and fine-tuning workflows across diverse non-vision target tasks where transferability estimation principles can be tested.
