Boosting for transfer learning

Wenyuan DaiQiang YangGui-Rong XueYong Yu

article2007ICML1,884 citations

Proposes TrAdaBoost, an extension of AdaBoost that iteratively reweights source-domain instances to filter out conflicting data, enabling accurate classification in a new target domain using only a minimal set of labeled target samples.

Listen

Organizations often struggle when applying machine learning to new operational domains because newly collected data is scarce and expensive to label manually. While large amounts of historical data may exist from older or related domains, data patterns frequently shift over time, violating standard assumptions that training and testing conditions are identical. Standard models trained strictly on scarce new data perform poorly, while blindly adding outdated data can degrade model quality due to negative interference.

The article evaluates a novel transfer learning framework, called TrAdaBoost, designed to construct high-performance classification models for new target domains by intelligently reusing large volumes of older, differently distributed historical data alongside a minimal quantity of newly labeled data.

The authors conducted both theoretical mathematical proofs and empirical experiments across nine benchmark datasets, including text and tabular classifications. Their approach extends traditional iterative boosting methods by automatically adjusting data importance over 100 iterations. In each cycle, older instances that conflict with the newly labeled target data are down-weighted to minimize their influence, while older instances that align well with the target distribution retain higher influence to expand the training set.

The experimental and theoretical results demonstrate several key findings. First, TrAdaBoost consistently outperformed baseline support vector machine algorithms across supervised and semi-supervised evaluations, often cutting error rates by 20% to over 60% compared to models trained on combined data without transfer mechanisms. Second, the framework delivers its highest value when newly labeled data is very scarce—specifically when the ratio of new to old data is below 0.10. Third, the theoretical convergence proofs established that weighted error on the older data approaches zero over time while test accuracy on the target domain systematically improves.

These findings indicate that organizations can significantly lower data labeling costs, accelerate deployment timelines, and reduce model error by safely reusing legacy data assets rather than discarding them. Leaders must note, however, that when new domain data is already abundant (exceeding a 0.20 ratio to historical data), the benefits of transfer learning diminish, and training directly on new data becomes equally or slightly more effective.

Technical leaders looking to reduce annotation expenses should pilot TrAdaBoost on classification tasks with severe data scarcity, provided an older related dataset is available. Before broad deployment, engineering teams should evaluate the quality and similarity of the historical data, as poor-quality auxiliary data yields smaller accuracy gains. Next steps should focus on benchmarking algorithm runtimes, as the method requires approximately 50 to 100 iterations to achieve optimal convergence, and extending the framework to handle legacy data drawn from multiple distinct historical distributions simultaneously.

  • Paper: A Survey on Transfer Learning, Sinno Jialin Pan et al. (2010). This foundational survey provides a comprehensive taxonomy of transfer learning, contextualizing instance-reweighting methods like TrAdaBoost within the broader transfer landscape.
  • Paper: A theory of learning from different domains, Shai Ben-David et al. (2010). This paper expands learning theory for domain adaptation by establishing formal generalization bounds for hypotheses trained on mixtures of source and target data.
  • Paper: A Comprehensive Survey on Transfer Learning, Fuzhen Zhuang et al. (2019). This comprehensive survey categorizes modern transfer learning mechanisms, systematically reviewing and comparing instance-weighting paradigms against newer feature and deep transfer strategies.
  • Paper: Adapting Visual Category Models to New Domains, Kate Saenko et al. (2010). This work extends domain adaptation beyond sample reweighting by learning asymmetric metric transformations between source and target feature spaces.
  • Paper: Return of Frustratingly Easy Domain Adaptation, Baochen Sun et al. (2015). This article builds on domain adaptation principles by offering an efficient statistical alignment method that matches second-order feature correlations across domains.
  • Paper: A Survey on Deep Transfer Learning, Chuanqi Tan et al. (2018). This survey examines how classical transfer learning paradigms, including instance-based reweighting, have evolved into deep neural network architectures.
  • Paper: Unsupervised Domain Adaptation by Backpropagation, Yaroslav Ganin et al. (2015). This work advances domain adaptation from instance reweighting to deep adversarial feature learning via gradient reversal.
  • Paper: Moment Matching for Multi-Source Domain Adaptation, Xingchao Peng et al. (2018). This paper addresses the specific future direction noted in the source by transferring knowledge from multiple distinct source distributions simultaneously.
Cover for Boosting for transfer learning

Abstract

Traditional machine learning makes a basic assumption: the training and test data should be under the same distribution. However, in many cases, this identical-distribution assumption does not hold. The assumption might be violated when a task from one new domain comes, while there are only labeled data from a similar old domain. Labeling the new data can be costly and it would also be a waste to throw away all the old data. In this paper, we present a novel transfer learning framework called TrAdaBoost, which extends boosting-based learning algorithms (Freund & Schapire, 1997). TrAdaBoost allows users to utilize a small amount of newly labeled data to leverage the old data to construct a high-quality classification model for the new data. We show that this method can allow us to learn an accurate model using only a tiny amount of new data and a large amount of old data, even when the new data are not sufficient to train a model alone. We show that TrAdaBoost allows knowledge to be effectively transferred from the old data to the new. The effectiveness of our algorithm is analyzed theoretically and empirically to show that our iterative algorithm can converge well to an accurate model.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Transfer Learning through TrAdaBoost
  • 4. Theoretical Analysis of TrAdaBoost
  • 5. Experimental Evaluation
  • 5.1. Data Sets
  • 5.2. Comparison Methods
  • 5.3. Experimental Results
  • 6. Conclusions and Future Work
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — TrAdaBoost Algorithm

    algorithm

    TrAdaBoost is an iterative boosting framework for transfer learning in binary classification tasks. It leverages a small amount of labeled target-domain data (same-distribution data) and a large amount of labeled auxiliary source-domain data (different-distribution data) to learn a classifier for unlabeled target test instances. In each iteration, if a source-domain training instance is misclassified by the current base hypothesis, its weight is reduced by a factor β∣ht(xi)−c(xi)∣\beta^{|h_t(x_i) - c(x_i)|} where β∈(0,1]\beta \in (0, 1], diminishing the influence of source instances that conflict with the target distribution. Meanwhile, target-domain instances that are misclassified have their weights increased according to the standard AdaBoost mechanism by βt−∣ht(xi)−c(xi)∣\beta_t^{-|h_t(x_i) - c(x_i)|} where βt=ϵt/(1−ϵt)\beta_t = \epsilon_t / (1 - \epsilon_t). The final classification rule is formed by voting hypotheses produced during the second half of the boosting process (iterations ⌈N/2⌉\lceil N/2 \rceil to NN).

    Input: Diff-distribution labeled set Td={(xid,c(xid))}i=1nT_d = \{(x_i^d, c(x_i^d))\}_{i=1}^n, same-distribution labeled set Ts={(xjs,c(xjs))}j=1mT_s = \{(x_j^s, c(x_j^s))\}_{j=1}^m, unlabeled test set SS, base learning algorithm Learner\text{Learner}, maximum iterations NN.
    Output: Final hypothesis hf:X→{0,1}h_f: X \to \{0, 1\}.
    Initialize instance weight vector w1=(w11,…,wn+m1)w^1 = (w_1^1, \dots, w_{n+m}^1), where instances 1≤i≤n1 \le i \le n belong to TdT_d and instances n+1≤i≤n+mn+1 \le i \le n+m belong to TsT_s.
    Set β=1/(1+2ln⁡n/N)\beta = 1 / (1 + \sqrt{2 \ln n / N}).
    for t=1t = 1 to NN do
        Set probability distribution pt=wt/∑i=1n+mwitp^t = w^t / \sum_{i=1}^{n+m} w_i^t.
        Call Learner\text{Learner} with combined training set T=Td∪TsT = T_d \cup T_s under distribution ptp^t and unlabeled set SS to obtain hypothesis ht:X→{0,1}h_t: X \to \{0, 1\} (or confidence in [0,1][0, 1]).
        Calculate error of hth_t exclusively on same-distribution set TsT_s:
            ϵt=∑i=n+1n+mwit∣ht(xi)−c(xi)∣∑i=n+1n+mwit\epsilon_t = \frac{\sum_{i=n+1}^{n+m} w_i^t |h_t(x_i) - c(x_i)|}{\sum_{i=n+1}^{n+m} w_i^t}
        if ϵt≥1/2\epsilon_t \ge 1/2 then
            Set ϵt=1/2−10−5\epsilon_t = 1/2 - 10^{-5} or terminate iteration early.
        end if
        Set βt=ϵt/(1−ϵt)\beta_t = \epsilon_t / (1 - \epsilon_t).
        Update weight vector for next round t+1t+1:
            for i=1i = 1 to nn do
                wit+1=wit⋅β∣ht(xi)−c(xi)∣w_i^{t+1} = w_i^t \cdot \beta^{|h_t(x_i) - c(x_i)|}
            end for
            for i=n+1i = n+1 to n+mn+m do
                wit+1=wit⋅βt−∣ht(xi)−c(xi)∣w_i^{t+1} = w_i^t \cdot \beta_t^{-|h_t(x_i) - c(x_i)|}
            end for
    end for
    Return hypothesis:
        hf(x)=1h_f(x) = 1 if ∏t=⌈N/2⌉Nβt−ht(x)≥∏t=⌈N/2⌉Nβt−1/2\prod_{t=\lceil N/2 \rceil}^N \beta_t^{-h_t(x)} \ge \prod_{t=\lceil N/2 \rceil}^N \beta_t^{-1/2}, and hf(x)=0h_f(x) = 0 otherwise.
  2. Knowl 2 — Transfer Learning Problem Formulation with Auxiliary Different-Distribution Data

    definition

    Let XsX_s denote the same-distribution (target) instance space, XdX_d denote the different-distribution (source/auxiliary) instance space, and X=Xs∪XdX = X_s \cup X_d be the combined instance space. The label set is Y={0,1}Y = \{0, 1\}, and the target concept is a boolean mapping c:X→Yc: X \to Y. The learning problem consists of three data sets:

    1. An unlabeled test set S={xit}i=1kS = \{x_i^t\}_{i=1}^k where each xit∈Xsx_i^t \in X_s.
    2. A small labeled same-distribution training set Ts={(xjs,c(xjs))}j=1mT_s = \{(x_j^s, c(x_j^s))\}_{j=1}^m where each xjs∈Xsx_j^s \in X_s, with sample size mm being too small to train an accurate classifier independently.
    3. An abundant labeled diff-distribution training set Td={(xid,c(xid))}i=1nT_d = \{(x_i^d, c(x_i^d))\}_{i=1}^n where each xid∈Xdx_i^d \in X_d (n≫mn \gg m), which shares task relevance with XsX_s but follows a different data distribution.

    The combined training set is T={(xi,c(xi))}i=1n+mT = \{(x_i, c(x_i))\}_{i=1}^{n+m}, where xi=xidx_i = x_i^d for i∈{1,…,n}i \in \{1, \dots, n\} and xi=xi−nsx_i = x_{i-n}^s for i∈{n+1,…,n+m}i \in \{n+1, \dots, n+m\}. The objective is to train a classifier c^:X→Y\hat{c}: X \to Y that minimizes the classification error on the unlabeled target test set SS.

  3. Knowl 3 — Upper Bound on Cumulative Loss over Diff-Distribution Training Data

    theoretical result

    In the TrAdaBoost framework, let Td={(xi,c(xi))}i=1nT_d = \{(x_i, c(x_i))\}_{i=1}^n be the diff-distribution training set, and let NN be the total number of boosting iterations. For iteration t∈{1,…,N}t \in \{1, \dots, N\} and instance xi∈Tdx_i \in T_d, let lit=∣ht(xi)−c(xi)∣l_i^t = |h_t(x_i) - c(x_i)| be the loss suffered by hypothesis hth_t on instance xix_i. Let L(xi)=∑t=1NlitL(x_i) = \sum_{t=1}^N l_i^t be the total loss suffered on instance xix_i over NN rounds.

    Let dit=wit/∑j=1nwjtd_i^t = w_i^t / \sum_{j=1}^n w_j^t be the normalized distribution over TdT_d at iteration tt, and let Ld=∑t=1N∑i=1nditlitL_d = \sum_{t=1}^N \sum_{i=1}^n d_i^t l_i^t denote the cumulative weighted training loss on TdT_d.

    When the parameter β\beta is chosen as β=1/(1+2ln⁡n/N)\beta = 1 / (1 + \sqrt{2 \ln n / N}), the average weighted loss on TdT_d satisfies:

    LdN≤min⁡1≤i≤nL(xi)N+2ln⁡nN+ln⁡nN\frac{L_d}{N} \le \min_{1 \le i \le n} \frac{L(x_i)}{N} + \sqrt{\frac{2 \ln n}{N}} + \frac{\ln n}{N}

    This bound shows that the average training loss on TdT_d is at most 2ln⁡n/N+ln⁡n/N\sqrt{2 \ln n / N} + \ln n / N greater than the loss of the single best instance in TdT_d, establishing that the worst-case convergence rate of TrAdaBoost over the diff-distribution data is O(ln⁡nN)O\left(\sqrt{\frac{\ln n}{N}}\right).

  4. Knowl 4 — Asymptotic Convergence of Average Weighted Loss on Diff-Distribution Instances

    theoretical result

    In the TrAdaBoost framework, let pit=wit/∑j=1n+mwjtp_i^t = w_i^t / \sum_{j=1}^{n+m} w_j^t be the globally normalized instance weight of xix_i across the full combined dataset T=Td∪TsT = T_d \cup T_s at round tt, and let lit=∣ht(xi)−c(xi)∣l_i^t = |h_t(x_i) - c(x_i)| denote the loss of hypothesis hth_t on instance xix_i.

    As the number of boosting iterations N→∞N \to \infty, the average weighted training loss suffered on the diff-distribution data Td={x1,…,xn}T_d = \{x_1, \dots, x_n\} over the second half of the boosting process converges to zero:

    lim⁡N→∞∑t=⌈N/2⌉N∑i=1npitlitN−⌈N/2⌉=0\lim_{N \to \infty} \frac{\sum_{t=\lceil N/2 \rceil}^N \sum_{i=1}^n p_i^t l_i^t}{N - \lceil N/2 \rceil} = 0

    This result demonstrates that TrAdaBoost successfully reduces the effective training weight and impact of misaligned diff-distribution instances over successive iterations, justifying the inclusion of hypotheses from round ⌈N/2⌉\lceil N/2 \rceil to NN in the final voting ensemble.

  5. Knowl 5 — Upper Bound on Same-Distribution Training Error

    theoretical result

    Let Ts={(xi,c(xi))}i=n+1n+mT_s = \{(x_i, c(x_i))\}_{i=n+1}^{n+m} be the labeled same-distribution dataset of size mm. Let hfh_f be the final hypothesis produced by TrAdaBoost after NN rounds, and let I={i:hf(xi)≠c(xi) and n+1≤i≤n+m}I = \{i : h_f(x_i) \ne c(x_i) \text{ and } n+1 \le i \le n+m\} be the set of misclassified same-distribution instances.

    The training error of hfh_f on TsT_s, defined as ϵ=Pr⁡x∈Ts[hf(x)≠c(x)]=∣I∣/m\epsilon = \Pr_{x \in T_s}[h_f(x) \ne c(x)] = |I| / m, satisfies the upper bound:

    ϵ≤2N/2∏t=⌈N/2⌉Nϵt(1−ϵt)\epsilon \le 2^{N/2} \prod_{t=\lceil N/2 \rceil}^N \sqrt{\epsilon_t(1 - \epsilon_t)}

    where ϵt=∑i=n+1n+mwit∣ht(xi)−c(xi)∣∑i=n+1n+mwit\epsilon_t = \frac{\sum_{i=n+1}^{n+m} w_i^t |h_t(x_i) - c(x_i)|}{\sum_{i=n+1}^{n+m} w_i^t} is the weighted classification error of hypothesis hth_t on TsT_s at iteration tt. Whenever ϵt<0.5\epsilon_t < 0.5 for each iteration t≥⌈N/2⌉t \ge \lceil N/2 \rceil, the upper bound on the training error on TsT_s decreases monotonically as NN increases.

  6. Knowl 6 — VC Generalization Error Upper Bound on Same-Distribution Data

    theoretical result

    Let dVCd_{VC} be the Vapnik–Chervonenkis (VC) dimension of the hypothesis space from which each base learner hth_t is selected, mm be the number of labeled same-distribution training examples in TsT_s, NN be the number of boosting iterations, and ϵ\epsilon be the empirical error of the final ensemble hypothesis hfh_f on TsT_s.

    With high probability, the generalization error of the final TrAdaBoost hypothesis hfh_f on the target distribution is bounded by:

    ϵ+O(NdVCm)\epsilon + O\left(\sqrt{\frac{N d_{VC}}{m}}\right)

    This bound shows that while running more iterations NN drives down the same-distribution empirical training error ϵ\epsilon, it inflates the complexity penalty O(NdVC/m)O\left(\sqrt{N d_{VC} / m}\right), indicating a theoretical possibility of overfitting when the same-distribution sample size mm is small.

  7. Knowl 7 — Supervised Transfer Classification Performance Comparison

    data/table

    In a supervised transfer learning experiment, TrAdaBoost with a linear Support Vector Machine (SVM) base learner was compared against standard SVM trained only on target data TsT_s, SVMt trained on the naive combination Ts∪TdT_s \cup T_d, and the auxiliary data learning method (AUX) of Wu and Dietterich (2004) with parameter Cp/Ca=4C_p / C_a = 4. The ratio of same-distribution training instances to diff-distribution training instances was fixed to ∣Ts∣/∣Td∣=0.01|T_s| / |T_d| = 0.01. Linear SVMlight was used, with total boosting iterations N=100N = 100. Classification error rates are averaged over 10 random runs across 9 cross-domain benchmark datasets generated from 20 Newsgroups, SRAA, Reuters-21578, and UCI Mushroom:

    Data Set SVM SVMt AUX TrAdaBoost(SVM)
    rec vs talk 0.222 0.127 0.127 0.080
    rec vs sci 0.240 0.164 0.153 0.097
    sci vs talk 0.234 0.177 0.173 0.125
    auto vs aviation 0.131 0.192 0.188 0.096
    real vs simulated 0.140 0.219 0.210 0.119
    orgs vs people 0.494 0.285 0.287 0.280
    orgs vs places 0.423 0.440 0.433 0.315
    people vs places 0.412 0.255 0.257 0.216
    edible vs poisonous 0.127 0.135 0.082 0.071

    TrAdaBoost(SVM) achieved strictly lower classification error rates than SVM, SVMt, and AUX across all nine benchmark datasets under the target-scarce condition ∣Ts∣/∣Td∣=0.01|T_s|/|T_d| = 0.01.

  8. Knowl 8 — Semi-Supervised Transfer Classification with Minimal Target Labels

    data/table

    In a semi-supervised transfer learning setting with extreme label scarcity, TrAdaBoost was implemented using Transductive Support Vector Machines (TSVM) as the base learner and evaluated with only two labeled same-distribution examples (one positive, one negative) in TsT_s, along with the unlabeled target test set SS and diff-distribution labeled data TdT_d. The baseline TSVM was trained on TsT_s and SS, while TSVMt was trained on Ts∪TdT_s \cup T_d and SS. TrAdaBoost(TSVM) ran for N=100N = 100 iterations. Classification error rates were averaged over 10 random repetitions:

    Data Set TSVM TSVMt TrAdaBoost(TSVM)
    rec vs talk 0.059 0.040 0.021
    rec vs sci 0.067 0.062 0.013
    sci vs talk 0.173 0.106 0.075
    auto vs aviation 0.043 0.103 0.038
    real vs simulated 0.144 0.131 0.102
    orgs vs people 0.358 0.292 0.248
    orgs vs places 0.424 0.436 0.304
    people vs places 0.307 0.225 0.179
    edible vs poisonous 0.439 0.179 0.160

    Across all nine tasks, TrAdaBoost(TSVM) substantially outperformed both TSVM and TSVMt, demonstrating that TrAdaBoost successfully extracts transferable patterns from diff-distribution data even when target supervision is limited to a single labeled example per class.

  9. Knowl 9 — Influence of Target-to-Source Sample Ratio on Transfer Gain

    empirical result

    Varying the ratio of labeled same-distribution instances to diff-distribution instances ∣Ts∣/∣Td∣|T_s| / |T_d| from 0.010.01 to 0.50.5 reveals distinct regimes of transfer effectiveness:

    1. Target-scarce regime (∣Ts∣/∣Td∣<0.1|T_s| / |T_d| < 0.1): TrAdaBoost provides large reductions in classification error compared to standard target-only SVM. In this range, the target data alone is insufficient to build a high-performance classifier, and reweighting diff-distribution data provides necessary supervision without letting conflicting source instances dominate.
    2. Target-sufficient regime (∣Ts∣/∣Td∣>0.2|T_s| / |T_d| > 0.2): Target data becomes adequate to train a competitive supervised model independently. The error rate of target-only SVM becomes comparable to or slightly lower than TrAdaBoost(SVM), because the marginal benefit of auxiliary diff-distribution data diminishes while residual source noise slightly perturbs the boundary.
    3. Baseline comparison: TrAdaBoost consistently outperforms the unweighted union model SVMt across the entire ratio range. The relative reduction in error of TrAdaBoost over SVMt tends to increase as the Kullback–Leibler (KL) divergence between source and target feature distributions increases.
  10. Knowl 10 — Limitations of the TrAdaBoost Framework

    limitation

    The TrAdaBoost transfer learning framework possesses four key limitations acknowledged and analyzed by the authors:

    1. Sensitivity to source data quality: Performance gains are sensitive to the quality and relevance of the diff-distribution examples; when the auxiliary data contains high noise or poor transferability, TrAdaBoost may achieve only comparable performance to a target-only model and cannot guarantee performance improvements in every setting.
    2. Convergence speed: The theoretical convergence rate on diff-distribution training instances is O(ln⁡nN)O\left(\sqrt{\frac{\ln n}{N}}\right), which is relatively slow; empirical evaluation indicates that the algorithm requires 50 or more iterations before stabilizing.
    3. Single source domain constraint: The formulation transfers knowledge from only one different-distribution dataset at a time and cannot handle auxiliary data originating from multiple heterogeneous source distributions simultaneously.
    4. Theoretical risk of overfitting: The generalization error bound contains an O(NdVC/m)O\left(\sqrt{N d_{VC} / m}\right) term that grows with iteration count NN when the target sample size mm is small, though empirical results demonstrate more robustness against overfitting than the theoretical bound predicts.

Coverage note — No substantial contributed material was omitted from the extracted knowls.

References

  1. 1.Ben-David, S., & Schuller, R. (2003). Exploiting task relatedness for multiple task learning. Proceedings of the Sixteenth Annual Conference on Learning Theory.
  2. 2.Bickel, S., & Scheffer, T. (2007). Dirichlet-enhanced spam filtering based on biased samples. In Advances in neural information processing systems 19.
  3. 3.Boser, B. E., Guyon, I., & Vapnik, V. (1992). A training algorithm for optimal margin classifiers. Proceedings of the Fifth Annual Workshop on Computational Learning Theory.
  4. 4.Caruana, R. (1997). Multitask learning. Machine Learning, 28(1), 41–75.
  5. 5.DaumeIII, H., & Marcu, D. (2006). Domain adaptation for statistical classifiers. Journal of Artificial Intelligence Research, 26, 101–126.
  6. 6.Dudık, M., Schapire, R., & Phillips, S. (2006). Correcting sample selection bias in maximum entropy density estimation. In Advances in neural information processing systems 18.
  7. 7.Freund, Y., & Schapire, R. E. (1997). A decisiontheoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences, 55(1), 119–139.
  8. 8.Heckman, J. J. (1979). Sample selection bias as a specification error. Econometrica, 47, 153–161.
  9. 9.Huang, J., Smola, A., Gretton, A., Borgwardt, K. M., & Scholkopf, B. (2007). Correcting sample selection bias by unlabeled data. In Advances in neural information processing systems 19.
  10. 10.Joachims, T. (1999). Transductive inference for text classification using support vector machines. Proceedings of Sixteenth International Conference on Machine Learning.
  11. 11.Joachims, T. (2002). Learning to classify text using support vector machines. Dissertation, Kluwer.
  12. 12.Kullback, S., & Leibler, R. A. (1951). On information and sufficiency. Annals of Mathematical Statistics, 22(1), 79–86.
  13. 13.Liao, X., Xue, Y., & Carin, L. (2005). Logistic regression with an auxiliary data source. Proceedings of the Twenty-Second International Conference on Machine Learning.
  14. 14.Rosenstein, M. T., Marx, Z., Kaelbling, L. P., & Dietterich, T. G. (2005). To transfer or not to transfer. Proceedings of NIPS 2005 Workshop on Inductive Transfer: 10 Years Later.
  15. 15.Schapire, R. E. (1999). A brief introduction to boosting. Proceedings of the Sixteenth International Joint Conference on Artificial Intelligence.
  16. 16.Schapire, R. E., Freund, Y., Bartlett, P., & Lee, W. S. (1997). Boosting the margin: A new explanation for the effectiveness of voting methods. Proceedings of the Fourteenth International Conference on Machine Learning.
  17. 17.Schmidhuber, J. (1994). On learning how to learn learning strategies (Technical Report FKI-198-94). Fakultat fur Informatik.
  18. 18.Shimodaira, H. (2000). Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of Statistical Planning and Inference, 90, 227–244.
  19. 19.Thrun, S., & Mitchell, T. M. (1995). Learning one more thing. Proceedings of the Fourteenth International Joint Conference on Artificial Intelligence.
  20. 20.Wu, P., & Dietterich, T. G. (2004). Improving svm accuracy by training on auxiliary data sources. Proceedings of the Twenty-First International Conference on Machine Learning.
  21. 21.Zadrozny, B. (2004). Learning and evaluating classifiers under sample selection bias. Proceedings of the Twenty-First International Conference on Machine Learning.

Citation

MLA
Dai, W., et al. “Boosting for Transfer Learning”. Proceedings of the 24th International Conference on Machine Learning, 2007, pp. 193–200, https://doi.org/10.1145/1273496.1273521.
APA
Dai, W., Yang, Q., Xue, G.-R., & Yu, Y. (2007). Boosting for transfer learning. Proceedings of the 24th International Conference on Machine Learning, 193–200. https://doi.org/10.1145/1273496.1273521
Chicago
Dai, W., Q. Yang, G.-R. Xue, and Y. Yu. 2007. “Boosting for Transfer Learning”. Proceedings of the 24th International Conference on Machine Learning, 193–200. https://doi.org/10.1145/1273496.1273521.
Harvard
Dai, W. et al. (2007) “Boosting for transfer learning”, Proceedings of the 24th international conference on Machine learning. ACM, pp. 193–200. Available at: https://doi.org/10.1145/1273496.1273521.
Vancouver
1. Dai W, Yang Q, Xue G-R, Yu Y (2007) Boosting for transfer learning. In: Proceedings of the 24th international conference on Machine learning. ACM, pp 193–200

BibTeX

@inproceedings{Dai_2007, series={ICML ’07 & ILP ’07}, title={Boosting for transfer learning}, url={http://dx.doi.org/10.1145/1273496.1273521}, DOI={10.1145/1273496.1273521}, booktitle={Proceedings of the 24th international conference on Machine learning}, publisher={ACM}, author={Dai, Wenyuan and Yang, Qiang and Xue, Gui-Rong and Yu, Yong}, year={2007}, month=June, pages={193–200}, collection={ICML ’07 & ILP ’07} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors