Frustratingly Easy Transferability Estimation

Long-Kai HuangJunzhou HuangYu RongQiang YangYing Wei

article2022ICML78 citations

Proposes TransRate, an extremely lightweight, training-free metric based on coding rate that accurately estimates transferability and guides optimal layer selection for pre-trained models using only a single forward pass over target examples.

Listen

Transfer learning from existing models is essential for solving complex machine learning tasks with limited labeled data, but finding the right pre-trained architecture and identifying which specific internal layers to transfer remains a significant operational hurdle. Traditional techniques require expensive fine-tuning or full retraining to test compatibility, creating substantial computational costs and development delays. Meanwhile, earlier lightweight estimation methods fail to evaluate intermediate layers or rely on restricted assumptions about source data accessibility.

The article introduces and evaluates TransRate, an optimization-free transferability metric designed to rapidly predict target performance before training begins. TransRate measures the mutual information between extracted target features and their class labels using coding rate as an efficient proxy for entropy, assessing both feature completeness across classes and compactness within each class.

The approach was validated through extensive empirical testing across 32 pre-trained models—spanning supervised, self-supervised, convolutional, and graph neural networks—and 16 diverse downstream tasks covering image classification, molecular regression, and molecular classification. The authors benchmarked TransRate against several leading alternatives without requiring access to source training datasets or iterative target optimization.

The findings show that TransRate delivers superior rank and linear correlations with final downstream performance across source dataset selection, model architecture comparison, and individual layer selection. TransRate was the only evaluated method capable of consistently identifying the highest-performing internal layers to transfer, achieving perfect layer-ranking correlation in 9 out of 15 layer-selection experiments. Furthermore, it operates up to roughly 3,000 times faster than full fine-tuning grid searches, remains computationally stable across a wide range of distortion parameters, and maintains robust ranking accuracy even when target labeled sample sizes are sharply reduced.

These results indicate that organizations can drastically reduce compute costs, eliminate negative transfer risks, and shorten model selection cycles from days to minutes by adopting TransRate prior to downstream fine-tuning. Unlike prior lightweight metrics, TransRate accommodates unsupervised representations and custom layer selections without requiring access to proprietary source datasets.

Teams should deploy TransRate as a standardized, pre-training screening filter to rank candidate architectures and select optimal layer cutoffs before initiating resource-heavy tuning. While the framework demonstrates high empirical confidence across diverse modalities, practitioners should exercise care when applying it to continuous regression problems, as target values must be discretized into discrete bins, and in extreme few-shot scenarios where sample representations may degrade.

arXiv: 2106.09362
  • Paper: How transferable are features in deep neural networks?, Jason Yosinski et al. (2014). Provides foundational empirical analyses on how feature specificity changes across network layers and how layer depth affects transferability, motivating TransRate's specific focus on layer-level transfer selection.
  • Paper: Do Better ImageNet Models Transfer Better?, Simon Kornblith et al. (2018). Establishes systematic benchmarking for how pre-trained representation quality relates to downstream transfer performance, setting the baseline questions TransRate seeks to evaluate efficiently without expensive fine-tuning.
  • Paper: Similarity of Neural Network Representations Revisited, Simon Kornblith et al. (2019). Introduces representation similarity analysis across neural network layers that directly informs how intermediate feature representations are evaluated prior to downstream transfer.
  • Paper: Taskonomy: Disentangling Task Transfer Learning, Amir Zamir et al. (2018). Formalizes the problem of mapping transferability and task relationships across visual representations that lightweight transferability metrics aim to solve computationally.
Cover for Frustratingly Easy Transferability Estimation

Abstract

Transferability estimation has been an essential tool in selecting a pre-trained model and the layers in it for transfer learning, so as to maximize the performance on a target task and prevent negative transfer. Existing estimation algorithms either require intensive training on target tasks or have difficulties in evaluating the transferability between layers. To this end, we propose a simple, efficient, and effective transferability measure named TransRate. Through a single pass over examples of a target task, TransRate measures the transferability as the mutual information between features of target examples extracted by a pre-trained model and their labels. We overcome the challenge of efficient mutual information estimation by resorting to coding rate that serves as an effective alternative to entropy. From the perspective of feature representation, the resulting TransRate evaluates both completeness (whether features contain sufficient information of a target task) and compactness (whether features of each class are compact enough for good generalization) of pre-trained features. Theoretically, we have analyzed the close connection of TransRate to the performance after transfer learning. Despite its extraordinary simplicity in 10 lines of codes, TransRate performs remarkably well in extensive evaluations on 32 pre-trained models and 16 downstream tasks.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. TransRate
  • 3.1. Notations and Problem Settings
  • 3.2. Computation-Efficient Transferability Estimation
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Results
  • 4.3. Discussion on Sensitivity to ϵ and Sample Size
  • 5. Conclusion
  • Acknowledgements
  • References
  • A. Omitted Experiment Details in Section 4
  • A.1. Image Datasets Description
  • A.2. Molecule Datasets Description
  • A.3. Pre-trained Models
  • A.4. Performance measure
  • A.5. Details about the Layer Selection Experiments in Section 4.3
  • A.6. Details about Applying TransRate on Regression Tasks
  • A.7. Source Codes of TransRate
  • B. Extra Experiments
  • B.1. Extra Experiments on Source Selection
  • B.2. Extra Results of Layer Selection
  • B.3. Extra Results on Model Selection
  • B.4. Extra Results on Self-supervised Model Selection
  • B.5. Extra Experiments on Sample Size Sensitivity Study
  • B.6. Time Complexity
  • B.7. Sensitivity to Value of ϵ
  • B.8. Target Selection
  • C. Theoretical studies of TransRate
  • C.1. Coding Rate and Shannon Entropy of a Quantized Continuous Random Variable
  • C.2. TransRate Score and Transfer Performance
  • D. Theoretical Details Omitted in Section 3
  • D.1. Proof of Proposition 1.
  • D.2. Properties of Coding Rate and TransRate Score
  • D.3. Proof of the toy case in Section 3.2
  • D.4. The Influence of ϵ
  • D.5. Time Complexity

Knowls

  1. Knowl 1 — TransRate Transferability Metric

    equation

    TransRate measures the transferability of a pre-trained feature extractor gg to a target classification task TtT_t by estimating the mutual information between the extracted features and target task labels using coding rate distortion. Given nn target samples (xi,yi)i=1n{(x_i, y_i)}_{i=1}^n belonging to CC classes, let Z^=[z1,…,zn]∈Rd×n\hat{Z} = [z_1, \dots, z_n] \in \mathbb{R}^{d \times n} denote the zero-mean feature matrix where zi=g(xi)∈Rdz_i = g(x_i) \in \mathbb{R}^d, scaled such that tr(Z^Z^⊤)=1\text{tr}(\hat{Z}\hat{Z}^\top) = 1. Let Z^c∈Rd×nc\hat{Z}^c \in \mathbb{R}^{d \times n_c} denote the submatrix of ncn_c sample features belonging to class cc.

    The empirical coding rate R(Z^,ϵ)R(\hat{Z}, \epsilon), representing the average bit rate required to encode Z^\hat{Z} under distortion error ϵ>0\epsilon > 0, is:

    R(Z^,ϵ)=12log⁡det⁡(Id+1nϵZ^Z^⊤)R(\hat{Z}, \epsilon) = \frac{1}{2} \log\det\left(I_d + \frac{1}{n\epsilon} \hat{Z}\hat{Z}^\top\right)

    The class-conditional coding rate R(Z^,ϵ∣Y)R(\hat{Z}, \epsilon \mid Y) is:

    R(Z^,ϵ∣Y)=∑c=1CncnR(Z^c,ϵ)=∑c=1Cnc2nlog⁡det⁡(Id+1ncϵZ^c(Z^c)⊤)R(\hat{Z}, \epsilon \mid Y) = \sum_{c=1}^C \frac{n_c}{n} R(\hat{Z}^c, \epsilon) = \sum_{c=1}^C \frac{n_c}{2n} \log\det\left(I_d + \frac{1}{n_c \epsilon} \hat{Z}^c (\hat{Z}^c)^\top\right)

    The TransRate score is defined as the difference between the total feature coding rate and the class-conditional coding rate:

    TrRTs→Tt(g,ϵ)=R(Z^,ϵ)−R(Z^,ϵ∣Y)\text{TrR}_{T_s \to T_t}(g, \epsilon) = R(\hat{Z}, \epsilon) - R(\hat{Z}, \epsilon \mid Y)

    where ϵ\epsilon is a distortion parameter typically set to 10−410^{-4}.

  2. Knowl 2 — Expected Log-Likelihood Definition of Feature Transferability

    definition

    Let TsT_s and TtT_t be a source task and a target task, respectively. A pre-trained neural network consists of an LL-layer feature extractor and a classification head. When transferring the first KK layers (K≤LK \le L) as a pre-trained feature extractor g=fK∘⋯∘f1g = f_K \circ \dots \circ f_1, the downstream target model comprises gg and a task-specific head w∈Ww \in \mathcal{W} consisting of layers (K+1)(K+1) to LL along with a newly initialized classifier head fL+1tf^t_{L+1}. Both the feature extractor gg and the head ww are fine-tuned on the target dataset (xi,yi)i=1n{(x_i, y_i)}_{i=1}^n.

    The fine-tuning optimization problem over the target training dataset is:

    g∗,w∗=arg⁡max⁡g~∈G,w∈W1n∑i=1nlog⁡p(yi∣g~(xi);g~,w)subject to g~(0)=gg^*, w^* = \arg\max_{\tilde{g} \in \mathcal{G}, w \in \mathcal{W}} \frac{1}{n}\sum_{i=1}^n \log p(y_i \mid \tilde{g}(x_i); \tilde{g}, w) \quad \text{subject to } \tilde{g}^{(0)} = g

    The transferability of the pre-trained feature extractor gg from TsT_s to TtT_t, denoted by TrfTs→Tt(g)\text{Trf}_{T_s \to T_t}(g), is defined as the expected log-likelihood of the optimal fine-tuned model w∗∘g∗w^* \circ g^* evaluated on a random test sample (x,y)∼Tt(x, y) \sim T_t:

    TrfTs→Tt(g):=E(x,y)[log⁡p(y∣g∗(x);g∗,w∗)]\text{Trf}_{T_s \to T_t}(g) := \mathbb{E}_{(x,y)}\left[\log p(y \mid g^*(x); g^*, w^*)\right]

  3. Knowl 3 — Log-Likelihood Bounds via TransRate

    theoretical result

    Assuming the target task has a uniform class label distribution, p(Y=yc)=1Cp(Y = y_c) = \frac{1}{C} for all c∈{1,…,C}c \in \{1, \dots, C\}, the log-likelihood L(g,h∗)\mathcal{L}(g, h^*) of the pre-trained feature extractor gg combined with an optimal downstream classifier h∗h^* satisfies the two-sided bound:

    TrRTs→Tt(g)−H(Y)≥L(g,h∗)≥TrRTs→Tt(g)−H(Y)−H(ZΔ)\text{TrR}_{T_s \to T_t}(g) - H(Y) \ge \mathcal{L}(g, h^*) \ge \text{TrR}_{T_s \to T_t}(g) - H(Y) - H(Z^\Delta)

    where H(Y)=log⁡CH(Y) = \log C is the Shannon entropy of the uniform target labels, ZΔZ^\Delta is the continuous feature variable Z=g(X)Z = g(X) quantized with grid size Δ=2πeϵ\Delta = \sqrt{2\pi e}\epsilon as ϵ→0\epsilon \to 0, H(ZΔ)H(Z^\Delta) is its discrete Shannon entropy, and TrRTs→Tt(g)=H(ZΔ)−H(ZΔ∣Y)≈h(Z)−h(Z∣Y)=I(Y;Z)\text{TrR}_{T_s \to T_t}(g) = H(Z^\Delta) - H(Z^\Delta \mid Y) \approx h(Z) - h(Z \mid Y) = I(Y; Z) is the mutual information between features and target labels.

    The upper bound is tight because the maximal log-likelihood is a variational approximation of the mutual information I(Y;Z)I(Y; Z) under an optimal classifier.

  4. Knowl 4 — TransRate Computation Algorithm

    algorithm

    TransRate is an optimization-free algorithm that assesses transferability with a single forward pass over target examples. Features are first centered, the unconditioned coding rate is evaluated, the class-conditional coding rates are summed across each class subset, and the difference is returned.

    import numpy as np
    
    def coding_rate(Z, eps=1e-4):
        n, d = Z.shape
        (_, rate) = np.linalg.slogdet(np.eye(d) + 1.0 / (n * eps) * (Z.T @ Z))
        return 0.5 * rate
    
    def transrate(Z, y, eps=1e-4):
        # Z: numpy array of shape (n, d) representing extracted features
        # y: numpy array of shape (n,) containing integer labels in {0, ..., C-1}
        Z = Z - np.mean(Z, axis=0, keepdims=True)
        RZ = coding_rate(Z, eps)
        RZY = 0.0
        num_classes = int(y.max() + 1)
        for c in range(num_classes):
            class_features = Z[(y == c).flatten()]
            RZY += coding_rate(class_features, eps)
        return RZ - RZY / num_classes
    

    The function takes feature matrix Z∈Rn×dZ \in \mathbb{R}^{n \times d}, class label vector y∈{0,…,C−1}ny \in \{0, \dots, C-1\}^n, and distortion threshold ϵ=10−4\epsilon = 10^{-4} (default). It returns the scalar TransRate score.

  5. Knowl 5 — Dual Formulation and Computational Complexity of Coding Rate

    theoretical result

    Applying Sylvester's determinant identity to the empirical coding rate formulation yields a dual form that enables efficient computation in low sample-size or high feature-dimension regimes. For a zero-mean feature matrix Z^∈Rd×n\hat{Z} \in \mathbb{R}^{d \times n}:

    R(Z^,ϵ)=12log⁡det⁡(Id+1nϵZ^Z^⊤)=12log⁡det⁡(In+1nϵZ^⊤Z^)R(\hat{Z}, \epsilon) = \frac{1}{2} \log\det\left(I_d + \frac{1}{n\epsilon}\hat{Z}\hat{Z}^\top\right) = \frac{1}{2} \log\det\left(I_n + \frac{1}{n\epsilon}\hat{Z}^\top\hat{Z}\right)

    Equivalently, if Z^\hat{Z} possesses rr singular values σ1,…,σr\sigma_1, \dots, \sigma_r, the coding rate simplifies to:

    R(Z^,ϵ)=12∑i=1rlog⁡(1+1nϵσi2)R(\hat{Z}, \epsilon) = \frac{1}{2}\sum_{i=1}^r \log\left(1 + \frac{1}{n\epsilon}\sigma_i^2\right)

    When n<dn < d, switching the determinant computation from the d×dd \times d feature covariance matrix to the n×nn \times n Gram matrix reduces the computational complexity from O(d2.373+nd2)\mathcal{O}(d^{2.373} + nd^2) to O(n2.373+dn2)\mathcal{O}(n^{2.373} + dn^2). Across all CC classes, the total time complexity of TransRate is O(min⁡{(C+1)d2.373+2nd2,  n2.373+dn2})\mathcal{O}(\min\{(C+1)d^{2.373} + 2nd^2, \; n^{2.373} + dn^2\}), which reduces to O(min⁡{d2.373+nd2,  n2.373+dn2})\mathcal{O}(\min\{d^{2.373} + nd^2, \; n^{2.373} + dn^2\}) when class-conditional coding rates are computed in parallel.

  6. Knowl 6 — Completeness, Compactness, and Separability Bounds of TransRate

    theoretical result

    For a target dataset partitioned by class with feature matrices Z^c∈Rd×nc\hat{Z}^c \in \mathbb{R}^{d \times n_c} for c∈{1,…,C}c \in \{1, \dots, C\} and full feature matrix Z^=[Z^1,…,Z^C]∈Rd×n\hat{Z} = [\hat{Z}^1, \dots, \hat{Z}^C] \in \mathbb{R}^{d \times n}, the working TransRate TrRTs→Tt(g,ϵ)=R(Z^,ϵ)−∑c=1CncnR(Z^c,ϵ)\text{TrR}_{T_s \to T_t}(g, \epsilon) = R(\hat{Z}, \epsilon) - \sum_{c=1}^C \frac{n_c}{n} R(\hat{Z}^c, \epsilon) satisfies the theoretical bounds:

    0≤TrRTs→Tt(g,ϵ)≤12∑c=1C[log⁡det⁡(Id+1nϵZ^c(Z^c)⊤)−ncnlog⁡det⁡(Id+1ncϵZ^c(Z^c)⊤)]0 \le \text{TrR}_{T_s \to T_t}(g, \epsilon) \le \frac{1}{2}\sum_{c=1}^C \left[ \log\det\left(I_d + \frac{1}{n\epsilon}\hat{Z}^c(\hat{Z}^c)^\top\right) - \frac{n_c}{n}\log\det\left(I_d + \frac{1}{n_c\epsilon}\hat{Z}^c(\hat{Z}^c)^\top\right) \right]

    • The lower bound of 0 is achieved if and only if 1nZ^Z^⊤=1ncZ^c(Z^c)⊤\frac{1}{n}\hat{Z}\hat{Z}^\top = \frac{1}{n_c}\hat{Z}^c(\hat{Z}^c)^\top for all cc, indicating that class feature distributions completely overlap and are inseparable.
    • The upper bound is achieved if and only if features belonging to different classes are mutually orthogonal, (Z^c1)⊤Z^c2=0(\hat{Z}^{c_1})^\top \hat{Z}^{c_2} = 0 for all c1≠c2c_1 \ne c_2.

    In a two-class setting with fixed within-class covariances (Z^1)⊤Z^1(\hat{Z}^1)^\top\hat{Z}^1 and (Z^2)⊤Z^2(\hat{Z}^2)^\top\hat{Z}^2, TransRate monotonically decreases with increasing inter-class cross-covariance (Z^1)⊤Z^2(\hat{Z}^1)^\top\hat{Z}^2, penalizing inter-class overlap (favoring completeness) while rewarding small intra-class dispersion (favoring compactness).

  7. Knowl 7 — TransRate Adaptation to Continuous Regression Tasks

    model/method

    To estimate transferability on target regression tasks where label values yi∈Ry_i \in \mathbb{R} are continuous rather than discrete categories, TransRate quantizes the continuous target variable into discrete quantile intervals:

    1. Rank all nn continuous scalar labels {yi}i=1n\{y_i\}_{i=1}^n in ascending order.
    2. Divide the ordered dataset evenly into C=10C = 10 contiguous bins of equal sample size nc=n/Cn_c = n / C.
    3. Assign samples in the cc-th bin the discrete group label cc, and collect their extracted feature vectors into submatrix Z^c∈Rd×nc\hat{Z}^c \in \mathbb{R}^{d \times n_c}.
    4. Compute the conditional coding rate as R(Z^,ϵ∣Y)=∑c=1CncnR(Z^c,ϵ)R(\hat{Z}, \epsilon \mid Y) = \sum_{c=1}^C \frac{n_c}{n} R(\hat{Z}^c, \epsilon).
    5. Subtract the conditional coding rate from the total coding rate: TrR(g,ϵ)=R(Z^,ϵ)−R(Z^,ϵ∣Y)\text{TrR}(g, \epsilon) = R(\hat{Z}, \epsilon) - R(\hat{Z}, \epsilon \mid Y).

    This binning strategy evaluates whether pre-trained features separate examples having distinct continuous target values while maintaining tight clusters for examples with similar target values.

  8. Knowl 8 — Layer Selection Transferability Across Deep Architectures

    empirical result

    TransRate enables selecting the optimal cut-off layer KK to transfer from a pre-trained network (where layers 11 to KK are transferred and subsequent layers are initialized and trained from scratch). Across 15 layer selection transfer experiments (including ResNet-20 pre-trained on SVHN or CIFAR-10, ResNet-18 pre-trained on 11 distinct source domains, and ResNet-34 pre-trained on ImageNet, transferred to CIFAR-100 and Fashion-MNIST):

    • TransRate achieved perfect rank correlation (Kendall's τK=1.0\tau_K = 1.0 and weighted τω=1.0\tau_\omega = 1.0) in 9 out of 15 layer selection benchmarks, whereas the strongest baseline (Label-Feature Correlation, LFC) achieved perfect rank prediction in only 3 benchmarks.
    • For ResNet-20 pre-trained on SVHN transferred to CIFAR-100 across candidate cut-off layers K∈{9,11,13,15,17,19}K \in \{9, 11, 13, 15, 17, 19\}, TransRate achieved Pearson correlation Rp=0.9769R_p = 0.9769, τK=0.8667\tau_K = 0.8667, and τω=0.9265\tau_\omega = 0.9265. Baseline methods produced negative correlations: LFC (Rp=−0.1895R_p = -0.1895), H-Score (Rp=−0.5320R_p = -0.5320), and LogME (Rp=−0.3352R_p = -0.3352).
    • For ResNet-18 pre-trained on Birdsnap transferred to CIFAR-100 across layers K∈{11,13,15,17}K \in \{11, 13, 15, 17\}, TransRate achieved Rp=0.9871R_p = 0.9871, τK=0.6667\tau_K = 0.6667, and τω=0.8133\tau_\omega = 0.8133, compared to LogME (Rp=−0.5207R_p = -0.5207) and H-Score (Rp=0.3166,τK=0.0000R_p = 0.3166, \tau_K = 0.0000).
  9. Knowl 9 — Pre-trained Model Architecture and Source Dataset Selection

    empirical result

    TransRate was evaluated on selecting optimal pre-trained models across diverse source datasets and neural network architectures:

    • Source Selection (11 Source Datasets): Evaluating ResNet-18 models pre-trained on 11 source datasets (ImageNet, Caltech-101/256, DTD, Flowers, SUN397, Pets, Food, Aircraft, Birds, Cars) transferred to 12 downstream datasets, TransRate achieved 20 out of 36 best performances across Pearson RpR_p, Kendall's τK\tau_K, and weighted τω\tau_\omega. In transferring to CIFAR-100, TransRate obtained Rp=0.7262,τK=0.8182,τω=0.9055R_p = 0.7262, \tau_K = 0.8182, \tau_\omega = 0.9055, substantially outperforming LogME (Rp=0.4947,τK=0.7091R_p = 0.4947, \tau_K = 0.7091) and LEEP (Rp=0.2883,τK=0.0909R_p = 0.2883, \tau_K = 0.0909).
    • Architecture Selection (7 Architectures): For 7 ImageNet models (ResNet-18/34/50 and MobileNet-v2 width multipliers 0.25/0.5/0.75/1.0) transferred to CIFAR-100, TransRate achieved τK=0.9048\tau_K = 0.9048 and τω=0.9421\tau_\omega = 0.9421, outperforming NCE (τω=0.7322\tau_\omega = 0.7322), H-Score (τω=0.5041\tau_\omega = 0.5041), and LogME (τω=0.6186\tau_\omega = 0.6186).
    • Architecture Selection (12 Architectures): Extending candidate architectures to include DenseNet-121/169/201, Inception-v3, and NASNet across 8 target datasets, TransRate achieved 18 out of 24 best performance scores across evaluation metrics.
  10. Knowl 10 — Transferability Estimation for Self-Supervised Vision Models and Molecular Graph Neural Networks

    empirical result

    TransRate reliably predicts downstream transfer performance for self-supervised visual encoders and Graph Neural Networks (GNNs):

    • Self-Supervised Vision Encoders: Evaluating ResNet-50 models pre-trained on ImageNet using four self-supervised methods (SimCLR, BYOL, SwaV, MoCo) transferred to CIFAR-100, TransRate was the only method that correctly identified the highest-performing model, achieving Rp=0.8550R_p = 0.8550, τK=0.6667\tau_K = 0.6667, and τω=0.8133\tau_\omega = 0.8133. In comparison, H-Score and LogME exhibited negative correlation (Rp=−0.9006R_p = -0.9006 and Rp=−0.8595R_p = -0.8595). Across 9 target datasets, TransRate obtained perfect rank correlations (τK=1.0,τω=1.0\tau_K = 1.0, \tau_\omega = 1.0) on Fashion-MNIST, Caltech-256, Flowers, and SUN397.
    • Self-Supervised Molecular GNNs: Evaluating GROVER graph transformer models pre-trained on 2M (ChemBL) or 11M (ZINC) molecules with 12M or 48M parameters transferred to molecular property prediction tasks, TransRate achieved perfect ranking (τK=1.0,τω=1.0\tau_K = 1.0, \tau_\omega = 1.0) on BACE classification (Rp=0.9424R_p = 0.9424), FreeSolv regression (Rp=0.9582R_p = 0.9582), and ESOL regression (Rp=0.7422R_p = 0.7422), while baseline methods (H-Score, LogME, LFC) produced negative rank correlations.
  11. Knowl 11 — Wall-Clock Runtime and Speedup of TransRate Versus Fine-Tuning

    data/table

    The table below reports wall-clock execution times (in seconds) and speedup factors relative to full fine-tuning on an Intel Xeon Platinum 8255C CPU with a single NVIDIA P40 GPU. Three transfer scenarios are measured: (1) ResNet-18 transferred to CIFAR-100 with full data (n=50,000,d=512n = 50{,}000, d = 512), (2) ResNet-18 transferred to CIFAR-100 with 10% data (n=5,000,d=512n = 5{,}000, d = 512), and (3) ResNet-50 transferred to CIFAR-100 with full data (n=50,000,d=2048n = 50{,}000, d = 2048).

    ResNet-18, Full Data ResNet-18, Small Data ResNet-50, Full Data
    Measure Time (s) Speedup Time (s) Speedup Time (s) Speedup
    Fine-tune 8399.65 1×1\times 882.33 1×1\times 2.3×1042.3 \times 10^4 1×1\times
    Feature Extraction 30.1416 – 3.2986 – 72.7870 –
    NCE 0.9126 9,204×9{,}204\times 0.2119 4,164×4{,}164\times 2.1220 10,839×10{,}839\times
    LEEP 0.7771 10,808×10{,}808\times 0.1211 7,286×7{,}286\times 1.9152 12,009×12{,}009\times
    LFC 30.1416 279×279\times 0.7987 1,106×1{,}106\times 149.3040 154×154\times
    H-Score 1.6285 5,158×5{,}158\times 0.3998 2,207×2{,}207\times 13.0700 1,760×1{,}760\times
    LogME 9.2737 906×906\times 2.0224 436×436\times 50.1797 458×458\times
    TransRate 1.3410 6,264×6{,}264\times 0.2697 3,272×3{,}272\times 10.6498 2,160×2{,}160\times

    TransRate completes scoring in 1.34s for ResNet-18 and 10.65s for ResNet-50 on full CIFAR-100, providing 2,160×2{,}160\times to 6,264×6{,}264\times speedup over a single fine-tuning run. Compared to other feature-based methods that support layer and unsupervised transfer (H-Score and LogME), TransRate has the lowest computation time.

Coverage note — None was omitted; all contributed formulations, theoretical bounds, algorithmic procedures, and experimental benchmarks across source, layer, model, and self-supervised selection are captured.

References

  1. 1.Achille, A., Lam, M., Tewari, R., Ravichandran, A., Maji, S., Fowlkes, C. C., Soatto, S., and Perona, P. Task2vec: Task embedding for meta-learning. In ICCV, pp. 6430–6439, 2019.
  2. 2.Agakov, D. B. F. The im algorithm: a variational approach to information maximization. NeurIPS, 16:201, 2004.
  3. 3.Bao, Y., Li, Y., Huang, S.-L., Zhang, L., Zheng, L., Zamir, A., and Guibas, L. An information-theoretic approach to transferability in task transfer learning. In ICIP, pp. 2309–2313, 2019.
  4. 4.Beirlant, J., Dudewicz, E. J., Györfi, L., and Van der Meulen, E. C. Nonparametric entropy estimation: An overview. International Journal of Mathematical and Statistical Sciences, 6(1):17–39, 1997.
  5. 5.Belghazi, M. I., Baratin, A., Rajeshwar, S., Ozair, S., Bengio, Y., Courville, A., and Hjelm, D. Mutual information neural estimation. In ICML, pp. 531–540, 2018.
  6. 6.Berg, T., Liu, J., Woo Lee, S., Alexander, M. L., Jacobs, D. W., and Belhumeur, P. N. Birdsnap: Large-scale fine-grained visual categorization of birds. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2011–2018, 2014.
  7. 7.Binia, J., Zakai, M., and Ziv, J. On the epsilon-entropy and the rate-distortion function of certain non-gaussian processes. IEEE Transactions on Information Theory, 20 (4):517–524, 1974.
  8. 8.Bossard, L., Guillaumin, M., and Van Gool, L. Food-101–mining discriminative components with random forests. In European conference on computer vision, pp. 446–461. Springer, 2014.
  9. 9.Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. Unsupervised learning of visual features by contrasting cluster assignments. arXiv preprint arXiv:2006.09882, 2020.
  10. 10.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pp. 1597–1607. PMLR, 2020.
  11. 11.Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A. Describing textures in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3606–3613, 2014.
  12. 12.Cover, T. M. Elements of information theory. John Wiley & Sons, 1999.
  13. 13.Cui, Y., Song, Y., Sun, C., Howard, A., and Belongie, S. Large scale fine-grained categorization and domain-specific transfer learning. In CVPR, pp. 4109–4118, 2018.
  14. 14.Delaney, J. S. Esol: estimating aqueous solubility directly from molecular structure. Journal of chemical information and computer sciences, 44(3):1000–1005, 2004.
  15. 15.Deshpande, A., Achille, A., Ravichandran, A., Li, H., Zancato, L., Fowlkes, C., Bhotika, R., Soatto, S., and Perona, P. A linearized framework and a new benchmark for model selection for fine-tuning. arXiv preprint arXiv:2102.00084, 2021.
  16. 16.Dwivedi, K. and Roig, G. Representation similarity analysis for efficient task taxonomy & transfer learning. In CVPR, pp. 12387–12396, 2019.
  17. 17.Fei-Fei, L., Fergus, R., and Perona, P. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In 2004 conference on computer vision and pattern recognition workshop, pp. 178–178. IEEE, 2004.
  18. 18.Gaulton, A., Bellis, L. J., Bento, A. P., Chambers, J., Davies, M., Hersey, A., Light, Y., McGlinchey, S., Michalovich, D., Al-Lazikani, B., et al. Chembl: a large-scale bioactivity database for drug discovery. Nucleic acids research, 40(D1):D1100–D1107, 2012.
  19. 19.Griffin, G., Holub, A., and Perona, P. Caltech-256 object category dataset. 2007.
  20. 20.Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., et al. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733, 2020.
  21. 21.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, pp. 770–778, 2016.
  22. 22.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9729–9738, 2020.
  23. 23.Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y. Learning deep representations by mutual information estimation and maximization. In ICLR, 2019.
  24. 24.Kendall, M. G. A new measure of rank correlation. Biometrika, 30(1/2):81–93, 1938.
  25. 25.Kraskov, A., Stögbauer, H., and Grassberger, P. Estimating mutual information. Physical review E, 69(6):066138, 2004.
  26. 26.Krause, J., Deng, J., Stark, M., and Fei-Fei, L. Collecting a large-scale dataset of fine-grained cars. 2013.
  27. 27.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009.
  28. 28.Li, Y., Jia, X., Sang, R., Zhu, Y., Green, B., Wang, L., and Gong, B. Ranking neural checkpoints. CVPR, pp. 2663–2673, 2021.
  29. 29.Ma, Y., Derksen, H., Hong, W., and Wright, J. Segmentation of multivariate mixed data via lossy data coding and compression. IEEE transactions on pattern analysis and machine intelligence, 29(9):1546–1562, 2007.
  30. 30.Maji, S., Rahtu, E., Kannala, J., Blaschko, M., and Vedaldi, A. Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151, 2013.
  31. 31.Martins, I. F., Teixeira, A. L., Pinheiro, L., and Falcao, A. O. A bayesian approach to in silico blood-brain barrier penetration modeling. Journal of chemical information and modeling, 52(6):1686–1697, 2012.
  32. 32.Mobley, D. L. and Guthrie, J. P. Freesolv: a database of experimental and calculated hydration free energies, with input files. Journal of computer-aided molecular design, 28(7):711–720, 2014.
  33. 33.Moon, Y.-I., Rajagopalan, B., and Lall, U. Estimation of mutual information using kernel density estimators. Physical Review E, 52(3):2318, 1995.
  34. 34.Nguyen, C. V., Hassner, T., Archambeau, C., and Seeger, M. Leep: A new measure to evaluate transferability of learned representations. ICML, 2020.
  35. 35.Nilsback, M.-E. and Zisserman, A. Automated flower classification over a large number of classes. In 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, pp. 722–729. IEEE, 2008.
  36. 36.Pan, S. J. and Yang, Q. A survey on transfer learning. IEEE Transactions on knowledge and data engineering, 22(10): 1345–1359, 2009.
  37. 37.Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pp. 3498–3505. IEEE, 2012.
  38. 38.Qin, Z., Kim, D., and Gedeon, T. Rethinking softmax with cross-entropy: Neural network classifier as mutual information estimator. arXiv preprint arXiv:1911.10688, 2019.
  39. 39.Quinlan, J. R. Induction of decision trees. Machine learning, 1(1):81–106, 1986.
  40. 40.Rong, Y., Bian, Y., Xu, T., Xie, W., Wei, Y., Huang, W., and Huang, J. Self-supervised graph transformer on large-scale molecular data. Advances in Neural Information Processing Systems, 33, 2020.
  41. 41.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3): 211–252, 2015.
  42. 42.Salman, H., Ilyas, A., Engstrom, L., Kapoor, A., and Madry, A. Do adversarially robust imagenet models transfer better? arXiv preprint arXiv:2007.08489, 2020.
  43. 43.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, pp. 4510–4520, 2018.
  44. 44.Shalev, Y., Painsky, A., and Ben-Gal, I. Neural joint entropy estimation. arXiv preprint arXiv:2012.11197, 2020.
  45. 45.Song, J., Chen, Y., Ye, J., Wang, X., Shen, C., Mao, F., and Song, M. Depara: Deep attribution graph for deep knowledge transferability. In CVPR, pp. 3922–3930, 2020.
  46. 46.Sterling, T. and Irwin, J. J. Zinc 15–ligand discovery for everyone. Journal of chemical information and modeling, 55(11):2324–2337, 2015.
  47. 47.Subramanian, G., Ramsundar, B., Pande, V., and Denny, R. A. Computational modeling of β-secretase 1 (bace-1) inhibitors using ligand based approaches. Journal of chemical information and modeling, 56(10):1936–1949, 2016.
  48. 48.Tishby, N. and Zaslavsky, N. Deep learning and the information bottleneck principle. In 2015 IEEE Information Theory Workshop (ITW), pp. 1–5. IEEE, 2015.
  49. 49.Tong, X., Xu, X., Huang, S.-L., and Zheng, L. A mathematical framework for quantifying transferability in multi-source transfer learning. Advances in Neural Information Processing Systems, 34, 2021.
  50. 50.Tran, A. T., Nguyen, C. V., and Hassner, T. Transferability and hardness of supervised classification tasks. In ICCV, pp. 1395–1405, 2019.
  51. 51.Wang, Z., Dai, Z., Póczos, B., and Carbonell, J. Characterizing and avoiding negative transfer. In CVPR, pp. 11293–11302, 2019.
  52. 52.Wu, Z., Ramsundar, B., Feinberg, E. N., Gomes, J., Geniesse, C., Pappu, A. S., Leswing, K., and Pande, V. Moleculenet: a benchmark for molecular machine learning. Chemical science, 9(2):513–530, 2018.
  53. 53.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017.
  54. 54.Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., and Torralba, A. Sun database: Large-scale scene recognition from abbey to zoo. In 2010 IEEE computer society conference on computer vision and pattern recognition, pp. 3485–3492. IEEE, 2010.
  55. 55.Yosinski, J., Clune, J., Bengio, Y., and Lipson, H. How transferable are features in deep neural networks? In NeurIPS, pp. 3320–3328, 2014.
  56. 56.You, K., Liu, Y., Long, M., and Wang, J. Logme: Practical assessment of pre-trained models for transfer learning. arXiv preprint arXiv:2102.11005, 2021.
  57. 57.Yu, Y., Chan, K. H. R., You, C., Song, C., and Ma, Y. Learning diverse and discriminative representations via the principle of maximal coding rate reduction. NeurIPS, 33, 2020.
  58. 58.Zamir, A. R., Sax, A., Shen, W., Guibas, L. J., Malik, J., and Savarese, S. Taskonomy: Disentangling task transfer learning. In CVPR, pp. 3712–3722, 2018.
  59. 59.Zhang, G., Zhao, H., Yu, Y., and Poupart, P. Quantifying and improving transferability in domain generalization. NeurIPS, 2021.
  60. 60.Zhang, W., Deng, L., and Wu, D. Overcoming negative transfer: A survey. arXiv preprint arXiv:2009.00909, 2020.

Citation

MLA
Huang, L.-K., et al. “Frustratingly Easy Transferability Estimation”. International Conference on Machine Learning, vol. 162, 2022, pp. 9201–25, https://proceedings.mlr.press/v162/huang22d.html.
APA
Huang, L.-K., Huang, J., Rong, Y., Yang, Q., & Wei, Y. (2022). Frustratingly Easy Transferability Estimation. International Conference on Machine Learning, 162, 9201–9225. https://proceedings.mlr.press/v162/huang22d.html
Chicago
Huang, L.-K., J. Huang, Y. Rong, Q. Yang, and Y. Wei. 2022. “Frustratingly Easy Transferability Estimation”. International Conference on Machine Learning 162: 9201–25. https://proceedings.mlr.press/v162/huang22d.html.
Harvard
Huang, L.-K. et al. (2022) “Frustratingly Easy Transferability Estimation”, International Conference on Machine Learning. PMLR, pp. 9201–9225. Available at: https://proceedings.mlr.press/v162/huang22d.html.
Vancouver
1. Huang L-K, Huang J, Rong Y, Yang Q, Wei Y (2022) Frustratingly Easy Transferability Estimation. In: International Conference on Machine Learning. PMLR, pp 9201–9225

BibTeX

@InProceedings{pmlr-v162-huang22d,
  title = 	 {Frustratingly Easy Transferability Estimation},
  author =       {Huang, Long-Kai and Huang, Junzhou and Rong, Yu and Yang, Qiang and Wei, Ying},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {9201--9225},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/huang22d/huang22d.pdf},
  url = 	 {https://proceedings.mlr.press/v162/huang22d.html},
  abstract = 	 {Transferability estimation has been an essential tool in selecting a pre-trained model and the layers in it for transfer learning, to transfer, so as to maximize the performance on a target task and prevent negative transfer. Existing estimation algorithms either require intensive training on target tasks or have difficulties in evaluating the transferability between layers. To this end, we propose a simple, efficient, and effective transferability measure named TransRate. Through a single pass over examples of a target task, TransRate measures the transferability as the mutual information between features of target examples extracted by a pre-trained model and their labels. We overcome the challenge of efficient mutual information estimation by resorting to coding rate that serves as an effective alternative to entropy. From the perspective of feature representation, the resulting TransRate evaluates both completeness (whether features contain sufficient information of a target task) and compactness (whether features of each class are compact enough for good generalization) of pre-trained features. Theoretically, we have analyzed the close connection of TransRate to the performance after transfer learning. Despite its extraordinary simplicity in 10 lines of codes, TransRate performs remarkably well in extensive evaluations on 35 pre-trained models and 16 downstream tasks.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/