Rotation Forest: A New Classifier Ensemble Method

Juan J. RodríguezLudmila I. KunchevaCarlos J. Alonso

article2006TPAMI1,974 citations

Introduces Rotation Forest, a classifier ensemble technique that applies Principal Component Analysis to random feature subsets to simultaneously boost individual decision tree accuracy and ensemble diversity, consistently outperforming Bagging, AdaBoost, and Random Forest across 33 benchmark datasets.

Listen

Combining multiple machine learning models into an ensemble is a proven way to improve prediction accuracy across various real-world tasks. However, building effective ensembles involves managing an accuracy-diversity trade-off: individual models must make accurate predictions while also making different types of errors from one another so their combined judgment works effectively. Standard approaches often struggle to balance this dynamic, particularly when deployed in compact ensembles that require fast operational execution.

The article sets out to develop and evaluate a new ensemble method called Rotation Forest, demonstrating that feature extraction via axis rotation can build base classifiers that are both highly accurate and diverse.

The approach divides the original feature set into smaller subsets and applies Principal Component Analysisa standard mathematical technique that rotates coordinate axesto subsets of training data. Rather than discarding components to reduce dimensions, the method retains all components to preserve total data variability. Transformed features are then used to train decision trees, which are naturally sensitive to axis rotations. The authors conducted rigorous empirical experiments across 33 standard benchmark datasets from the University of California, Irvine repository, comparing Rotation Forest against established ensemble techniques including Bagging, AdaBoost, and Random Forest using fixed ensemble sizes.

The findings show that Rotation Forest consistently outperformed the competing methods by a substantial margin. In head-to-head comparisons across the 33 datasets, Rotation Forest achieved the highest classification accuracy on the vast majority of tasks, winning nearly 70% of evaluations at a baseline ensemble size of 10 trees. Statistical significance tests confirmed its dominant performance over Bagging, AdaBoost, and Random Forest in both pruned and unpruned decision tree configurations. Detailed diversity-error visual analyses revealed why: Rotation Forest retains the high individual tree accuracy typical of Bagging while generating greater model diversity, effectively capturing the strengths of existing methods without incurring their typical error penalties.

These results demonstrate that organizations can achieve superior classification performance using smaller, computationally lean ensembles. This delivers immediate operational value by cutting the latency and memory footprint required for real-time predictions without sacrificing decision quality. Contrary to the prevailing belief that diversity must come at the expense of individual model accuracy, the findings prove that targeted axis rotations can simultaneously optimize both factors.

Decision-makers and practitioners deploying classification systems on small-to-moderate datasets should consider implementing Rotation Forest when prediction accuracy is critical. For subsequent development, technical teams should evaluate tuning the feature subset size parameter and explore applying this rotation framework to other base model architectures, such as neural networks. Further testing is also recommended before applying the method to massive, web-scale datasets containing millions of features or instances, as very large-scale data and automated feature importance ranking fall outside the current implementation's validated scope.

Cover for Rotation Forest: A New Classifier Ensemble Method

Abstract

We propose a method for generating classifier ensembles based on feature extraction. To create the training data for a base classifier, the feature set is randomly split into K subsets (K is a parameter of the algorithm) and Principal Component Analysis (PCA) is applied to each subset. All principal components are retained in order to preserve the variability information in the data. Thus, K axis rotations take place to form the new features for a base classifier. The idea of the rotation approach is to encourage simultaneously individual accuracy and diversity within the ensemble. Diversity is promoted through the feature extraction for each base classifier. Decision trees were chosen here because they are sensitive to rotation of the feature axes, hence the name "forest." Accuracy is sought by keeping all principal components and also using the whole data set to train each base classifier. Using WEKA, we examined the Rotation Forest ensemble on a random selection of 33 benchmark data sets from the UCI repository and compared it with Bagging, AdaBoost, and Random Forest. The results were favorable to Rotation Forest and prompted an investigation into the diversity-accuracy landscape of the ensemble models. Diversity-error diagrams revealed that Rotation Forest ensembles construct individual classifiers which are more accurate than these in AdaBoost and Random Forest, and more diverse than these in Bagging, sometimes more accurate as well.

Table of Contents

  • 1 INTRODUCTION
  • 2 ROTATION FORESTS
  • 3 EXPERIMENTAL VALIDATION
  • 4 DIVERSITY-ERROR DIAGRAMS
  • 5 CONCLUSIONS AND FUTURE WORK
  • ACKNOWLEDGMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — Rotation Forest Ensemble Algorithm

    algorithm

    Rotation Forest is a classifier ensemble method that builds diverse and accurate base classifiers by applying Principal Component Analysis (PCA) to randomly chosen subsets of features while retaining all principal components to preserve data variability.

    Input: Training dataset XX (N×nN \times n matrix), class labels YY (N×1N \times 1 vector where yj{ω1,,ωc}y_j \in \{\omega_1, \dots, \omega_c\}), ensemble size LL, feature subset size parameter MM (or number of subsets K=n/MK = \lceil n/M \rceil).
    Output: Ensemble of trained decision tree classifiers D1,,DLD_1, \dots, D_L and rotation matrices R1a,,RLaR_1^a, \dots, R_L^a.
    // Training Phase
    for i=1i = 1 to LL do
        Split the feature set FF randomly into KK disjoint subsets Fi,1,,Fi,KF_{i,1}, \dots, F_{i,K} (each containing MM features, with any remainder completed randomly)
        for j=1j = 1 to KK do
            Let Xi,jX_{i,j} be the training data restricted to the features in Fi,jF_{i,j}
            Select a random non-empty subset of classes from {ω1,,ωc}\{\omega_1, \dots, \omega_c\}
            Draw a bootstrap sample of objects belonging to the selected classes of size 75%75\% of the total object count in Xi,jX_{i,j}, yielding Xi,jX'_{i,j}
            Run Principal Component Analysis on Xi,jX'_{i,j}
            Store the principal component coefficient column vectors ai,j(1),,ai,j(Mj)a_{i,j}^{(1)}, \dots, a_{i,j}^{(M_j)} as matrix Ci,jC_{i,j} (MjMM_j \le M)
        end for
        Construct the sparse block diagonal rotation matrix RiR_i with diagonal blocks Ci,1,,Ci,KC_{i,1}, \dots, C_{i,K}
        Rearrange the columns of RiR_i to match the original feature order of FF, yielding matrix RiaRn×nR_i^a \in \mathbb{R}^{n \times n}
        Train decision tree base classifier DiD_i on the rotated training set (XRia,Y)(X R_i^a, Y)
    end for
    // Classification Phase for an unlabelled test point xR1×nx \in \mathbb{R}^{1 \times n}
    for j=1j = 1 to cc do
        Compute the average posterior probability confidence for class ωj\omega_j:
        μj(x)=1Li=1Ldi,j(xRia)\mu_j(x) = \frac{1}{L} \sum_{i=1}^L d_{i,j}(x R_i^a), where di,j(xRia)d_{i,j}(x R_i^a) is the probability assigned by DiD_i to class ωj\omega_j
    end for
    Assign xx to the class label with the largest confidence: argmaxj{1,,c}μj(x)\arg\max_{j \in \{1, \dots, c\}} \mu_j(x)

    Decision trees serve as the base classifiers because they are sensitive to rotations of the feature axes while retaining high individual accuracy. In standard evaluations, feature subset size is fixed to M=3M = 3 and probability aggregation across ensemble members uses mean aggregation.

  2. Knowl 2 — Rotation Matrix Formulation and Full-Component Retention

    model/method

    In Rotation Forest, each base classifier DiD_i (i=1,,Li = 1, \dots, L) is trained on a linearly transformed feature space defined by a rotation matrix RiR_i. The feature set FF is randomly partitioned into KK subsets Fi,1,,Fi,KF_{i,1}, \dots, F_{i,K}, each containing MM features (n=KMn = KM). For each subset Fi,jF_{i,j}, PCA produces MjMM_j \le M principal component loading vectors ai,j(1),,ai,j(Mj)a_{i,j}^{(1)}, \dots, a_{i,j}^{(M_j)}, each of dimension M×1M \times 1.

    These vectors are assembled into a block-diagonal rotation matrix RiR_i of dimension n×j=1KMjn \times \sum_{j=1}^K M_j:

    Ri=[ai,1(1),,ai,1(M1)[0][0][0]ai,2(1),,ai,2(M2)[0][0][0]ai,K(1),,ai,K(MK)]R_i = \begin{bmatrix} a_{i,1}^{(1)}, \dots, a_{i,1}^{(M_1)} & [0] & \dots & [0] \\ [0] & a_{i,2}^{(1)}, \dots, a_{i,2}^{(M_2)} & \dots & [0] \\ \vdots & \vdots & \ddots & \vdots \\ [0] & [0] & \dots & a_{i,K}^{(1)}, \dots, a_{i,K}^{(M_K)} \end{bmatrix}

    The columns of RiR_i are subsequently rearranged to align with the original ordering of the features in FF, producing the rearranged rotation matrix RiaR_i^a of size n×nn \times n (assuming all Mj=MM_j = M). The training data matrix XRN×nX \in \mathbb{R}^{N \times n} is transformed into XRiaX R_i^a prior to training DiD_i.

    Unlike traditional dimensionality reduction applications of PCA where low-variance components are discarded, Rotation Forest retains all principal component vectors. This guarantees that total dataset variance and discriminative information are completely preserved in the transformed feature space, even if discriminatory power lies along directions of minor variance. The rotation heuristic induces diversity across ensemble members by presenting axis-aligned decision trees with differently rotated coordinate systems.

  3. Knowl 3 — Combinatorial Feature Partitioning and Necessity of Secondary Randomization

    theoretical result

    For a dataset with nn features partitioned into KK disjoint subsets of size MM such that n=KMn = KM, the total number TT of possible distinct partitions is given by:

    T=n!K!(M!)KT = \frac{n!}{K! (M!)^K}

    Assuming each feature partition is chosen independently and with equal probability across an ensemble of LL classifiers, the probability that all LL classifiers receive strictly distinct feature partitions is:

    P(all distinct classifiers)=T!(TL)!TLP(\text{all distinct classifiers}) = \frac{T!}{(T-L)! \, T^L}

    When the feature dimension nn is small, the number of distinct partitions is limited. For example, with n=9n = 9, K=3K = 3, and M=3M = 3:

    T=9!3!(3!)3=3628806×216=280T = \frac{9!}{3! (3!)^3} = \frac{362880}{6 \times 216} = 280

    For an ensemble of size L=50L = 50, the probability that all 50 classifiers receive unique feature subsets is P(all distinct classifiers)<0.01P(\text{all distinct classifiers}) < 0.01.

    Because pure feature partitioning does not provide sufficient diversity for small or moderate feature dimensions, Rotation Forest incorporates a two-step secondary randomization heuristic prior to running PCA on each subset Fi,jF_{i,j}:

    1. A random non-empty subset of the class labels is selected.
    2. A bootstrap sample of instances containing 75%75\% of the dataset size restricted to the selected classes is drawn.

    This secondary randomization ensures that identical feature subsets generate different PCA transformation coefficients across classifiers.

  4. Knowl 4 — Pairwise Interrater Agreement (Kappa) and Kappa-Error Formulation

    equation

    To analyze ensemble diversity and individual classifier accuracy, pairwise relationships between base classifiers DiD_i and DjD_j are quantified using Cohen's kappa statistic κi,j\kappa_{i,j} alongside the average pair error Ei,jE_{i,j}.

    For a cc-class classification problem, let MM be the c×cc \times c coincidence matrix of the two classifiers evaluated on a test set, where entry mk,sm_{k,s} is the proportion of test instances classified as class ωk\omega_k by DiD_i and as class ωs\omega_s by DjD_j. The observed agreement is k=1cmkk\sum_{k=1}^c m_{kk}, and the agreement-by-chance ABCABC is defined as:

    ABC=k=1c(s=1cmk,s)(s=1cms,k)ABC = \sum_{k=1}^c \left( \sum_{s=1}^c m_{k,s} \right) \left( \sum_{s=1}^c m_{s,k} \right)

    The interrater agreement κi,j\kappa_{i,j} is given by:

    κi,j=k=1cmkkABC1ABC\kappa_{i,j} = \frac{\sum_{k=1}^c m_{kk} - ABC}{1 - ABC}

    For binary classification (c=2c = 2), κi,j\kappa_{i,j} simplifies to:

    κi,j=2(m1,1m2,2m1,2m2,1)(m1,1+m1,2)(m1,1+m2,1)+(m1,2+m2,2)(m2,1+m2,2)\kappa_{i,j} = \frac{2(m_{1,1}m_{2,2} - m_{1,2}m_{2,1})}{(m_{1,1}+m_{1,2})(m_{1,1}+m_{2,1}) + (m_{1,2}+m_{2,2})(m_{2,1}+m_{2,2})}

    The average individual classification error of the pair is:

    Ei,j=Ei+Ej2E_{i,j} = \frac{E_i + E_j}{2}

    where EiE_i and EjE_j are the test errors of classifiers DiD_i and DjD_j, respectively.

    In a κ\kappa-error diagram, each of the L(L1)/2L(L-1)/2 classifier pairs in an ensemble of size LL is plotted as a point (κi,j,Ei,j)(\kappa_{i,j}, E_{i,j}). Smaller values of κ\kappa denote greater diversity (with κ=1\kappa = 1 indicating identical predictions, κ=0\kappa = 0 indicating chance independence, and κ<0\kappa < 0 indicating negative dependence), while smaller values of Ei,jE_{i,j} denote higher individual accuracy. The most effective ensemble pairs occupy the lower-left region of the diagram.

  5. Knowl 5 — Experimental Protocol and Corrected Hypothesis Testing for Ensemble Comparison

    experimental setup

    The experimental validation evaluates Rotation Forest against Bagging, AdaBoost.M1, Random Forest, and single decision trees across 33 benchmark datasets from the UCI Machine Learning Repository with varied feature and class dimensions.

    Experimental configurations:

    • Base classifier: WEKA J48 decision tree (C4.5 implementation), tested under both error-based pruned (default 25% confidence threshold) and unpruned settings.
    • Rotation Forest parameter: feature subset size fixed to M=3M = 3 without hyperparameter tuning.
    • Random Forest: node split feature pool size set to log2(n)+1\lfloor \log_2(n) \rfloor + 1.
    • Bagging: uses probability average aggregation across ensemble members.
    • Ensemble size: fixed to L=10L = 10 base trees across all methods for controlled operating complexity.
    • Validation: 15 independent repetitions of 10-fold cross-validation, producing T=150T = 150 test folds per dataset and method.

    Statistical significance testing at the α=0.05\alpha = 0.05 significance level uses the conservative variance estimator proposed by Nadeau and Bengio:

    σ^=σ1T+NtestingNtraining\hat{\sigma} = \sigma \sqrt{\frac{1}{T} + \frac{N_{\text{testing}}}{N_{\text{training}}}}

    where σ\sigma is the standard deviation across all T=150T = 150 test runs, and NtrainingN_{\text{training}} and NtestingN_{\text{testing}} are the respective training and testing set sample sizes within each fold. This correction accounts for data overlap across folds, reducing Type I error rates.

  6. Knowl 6 — Comparative Performance and Statistical Dominance of Rotation Forest

    data/table

    Across 33 UCI datasets using an ensemble size of L=10L = 10, Rotation Forest demonstrates consistent statistical dominance over Bagging, AdaBoost.M1, Random Forest, and single J48 trees in both pruned and unpruned configurations.

    Pruned trees Unpruned trees
    Method J48 Bagging AdaBoost Rot. Forest J48 Bagging AdaBoost Rand. Forest Rot. Forest
    J48 (pruned) - 29 (9) 25 (12) 29 (18) 14 (2) 26 (9) 23 (12) 22 (8) 28 (18)
    Bagging (pruned) 4 (0) - 21 (7) 27 (10) 3 (0) 18 (2) 17 (6) 16 (4) 25 (9)
    AdaBoost (pruned) 8 (1) 12 (3) - 25 (9) 8 (0) 12 (2) 15 (0) 16 (1) 26 (7)
    Rot. Forest (pruned) 4 (0) 6 (0) 8 (1) - 2 (0) 7 (0) 7 (1) 5 (0) 17 (0)
    J48 (unpruned) 19 (5) 30 (14) 25 (14) 31 (19) - 31 (12) 26 (13) 28 (9) 31 (21)
    Bagging (unpruned) 7 (1) 15 (0) 21 (4) 26 (10) 2 (0) - 20 (5) 20 (4) 28 (10)
    AdaBoost (unpruned) 10 (1) 16 (3) 18 (0) 26 (8) 7 (1) 13 (1) - 15 (1) 26 (8)
    Rand. Forest 11 (2) 17 (2) 17 (5) 28 (10) 5 (2) 13 (2) 18 (4) - 28 (10)
    Rot. Forest (unpruned) 5 (0) 8 (0) 7 (1) 15 (0) 2 (0) 5 (0) 7 (1) 5 (0) -

    In the pairwise comparison matrix, table entry ai,ja_{i,j} indicates the number of datasets on which the column method jj achieved higher mean accuracy than the row method ii, with the number of statistically significant differences (p<0.05p < 0.05) shown in parentheses.

    Aggregated non-dominance ranking based on significant wins minus significant losses across all pairwise comparisons:

    1. Rotation Forest (pruned J48): 84 wins, 2 losses (Dominance score: +82+82)
    2. Rotation Forest (unpruned J48): 83 wins, 2 losses (Dominance score: +81+81)
    3. AdaBoost.M1 (pruned J48): 44 wins, 23 losses (Dominance score: +21+21)
    4. AdaBoost.M1 (unpruned J48): 42 wins, 23 losses (Dominance score: +19+19)
    5. Bagging (unpruned J48): 28 wins, 34 losses (Dominance score: 6-6)
    6. Bagging (pruned J48): 31 wins, 38 losses (Dominance score: 7-7)
    7. Random Forest: 27 wins, 37 losses (Dominance score: 10-10)
    8. Single J48 tree (pruned): 10 wins, 88 losses (Dominance score: 78-78)
    9. Single J48 tree (unpruned): 5 wins, 107 losses (Dominance score: 102-102)
  7. Knowl 7 — Diversity-Accuracy Dynamics in Rotation Forest versus Classic Ensembles

    empirical result

    Analysis of κ\kappa-error diagrams reveals the distinct operational mechanisms that distinguish Rotation Forest from Bagging, AdaBoost, and Random Forest:

    • AdaBoost generates high diversity (broad spread toward low κ\kappa values) by iteratively reweighting difficult instances, but this diversity comes at the expense of individual classifier accuracy, resulting in elevated pair errors Ei,jE_{i,j}.
    • Bagging maintains high individual classifier accuracy (low Ei,jE_{i,j}) because base trees are trained on full-feature bootstrap samples resembling the original data distribution, but it achieves comparatively low diversity (high κ\kappa).
    • Random Forest introduces diversity relative to Bagging by selecting random feature subsets at each tree split, but this randomization typically increases individual base tree error.
    • Rotation Forest achieves a favorable position in the diversity-accuracy landscape: it builds individual decision trees that match or exceed the high accuracy of Bagging while shifting the cloud of (κi,j,Ei,j)(\kappa_{i,j}, E_{i,j}) points leftward toward lower κ\kappa (higher diversity).

    Because PCA rotation preserves full dataset variance while altering coordinate orientations, base decision trees encounter different discriminative splits without losing relevant information. This combination of Bagging-level individual accuracy and heightened diversity accounts for Rotation Forest's superior performance, especially in small ensemble regimes (L20L \le 20).

  8. Knowl 8 — Nominal Attribute Preprocessing via Binary Indicator Encoding

    model/method

    To enable Principal Component Analysis on datasets containing categorical features, Rotation Forest converts nominal variables into numerical dummy variables using 1-of-ss binary encoding, where ss is the number of distinct categories in a given nominal attribute.

    This transformation replaces each categorical attribute with ss binary features taking values in {0,1}\{0, 1\}. The encoding is executed as an initial preprocessing step prior to feature subset partitioning. As a result, binary indicator columns originating from the same nominal variable can be assigned to different feature subsets Fi,jF_{i,j}.

    Standard PCA is applied directly to these binary variables without discarding components. While 1-of-ss encoding introduces linear dependencies among the indicators and non-Gaussian distributions (which violates assumptions required when PCA is used for econometric scoring indices or variance estimation), these properties do not degrade classification performance in Rotation Forest. In this ensemble framework, PCA is not employed for statistical inference or dimensionality reduction, but solely as an axis rotation heuristic to generate distinct tree geometries while retaining all information.

  9. Knowl 9 — Stated Limitations and Boundary Conditions of Rotation Forest

    limitation

    The Rotation Forest ensemble methodology possesses several documented limitations and operating constraints:

    1. Hyperparameter Sensitivity: The method introduces an additional parameter MM (the feature subset size, or equivalently K=n/MK = \lceil n/M \rceil subsets). Although setting M=3M = 3 performs well across diverse benchmark datasets, the optimal subset size is problem-dependent and requires tuning for specialized domains.
    2. Absence of Direct Feature Ranking: Unlike Random Forest, which provides intrinsic out-of-bag feature importance rankings, Rotation Forest linearly transforms and mixes original features across subsets, precluding straightforward ranking or direct attribution of individual feature importance.
    3. Scalability on Large-Scale Datasets: The algorithm relies on multiple PCA computations (one per feature subset per base tree) followed by dense linear transformations (XRiaX R_i^a). While computationally tractable on small to moderate datasets, it requires modifications for ultra-large datasets containing millions of instances or millions of features.
    4. Diminishing Performance Advantage for Large Ensemble Sizes: The accuracy advantage of Rotation Forest over Bagging, AdaBoost, and Random Forest is most pronounced at small ensemble sizes (L20L \le 20). As ensemble size LL approaches hundreds or thousands of trees, the performance differences between ensemble paradigms tend to diminish.

Coverage note — None was omitted; all key theoretical formulations, algorithmic specifications, experimental designs, empirical results, and stated limitations have been fully extracted.

References

  1. 1.E.L. Allwein, R.E. Schapire, and Y. Singer, “Reducing Multiclass to Binary: A Unifying Approach for Margin Classifiers,” J. Machine Learning Research, vol. 1, pp. 113-141, 2000.
  2. 2.R.E. Banfield, L.O. Hall, K.W. Bowyer, D. Bhadoria, W.P. Kegelmeyer, and S. Eschrich, “A Comparison of Ensemble Creation Techniques,” Proc Fifth Int’l Workshop Multiple Classifier Systems (MCS ’04), 2004.
  3. 3.E. Bauer and R. Kohavi, “An Empirical Comparison of Voting Classification Algorithms: Bagging, Boosting, and Variants,” Machine Learning, vol. 36, nos. 1-2, pp. 105-139, 1999.
  4. 4.C.L. Blake and C.J. Merz, “UCI Repository of Machine Learning Databases,” 1998, http://www.ics.uci.edu/mlearn/MLRepository.html.
  5. 5.L. Breiman, “Bagging Predictors,” Machine Learning, vol. 24, no. 2, pp. 123-140, 1996.
  6. 6.L. Breiman, “Arcing Classifiers,” Annals of Statistics, vol. 26, no. 3, pp. 801-849, 1998.
  7. 7.L. Breiman, “Random Forests,” Machine Learning, vol. 45, no. 1, pp. 5-32, 2001.
  8. 8.T.G. Dietterich, “Ensemble Methods in Machine Learning,” Proc. Conf. Multiple Classifier Systems, pp. 1-15, 2000.
  9. 9.X.Z. Fern and C.E. Brodley, “Random Projection for High Dimensional Data Clustering: A Cluster Ensemble Approach,” Proc. 20th Int’l Conf. Machine Learning (ICML), pp. 186-193, 2003.
  10. 10.J.L. Fleiss, Statistical Methods for Rates and Proportions. John Wiley and Sons, 1981.
  11. 11.F.H. Foley and J.W. Sammon, “An Optimal Set of Discriminant Vectors,” IEEE Trans. Computers, vol. 24, no. 3, pp. 281-289, Mar. 1975.
  12. 12.Y. Freund and R.E. Schapire, “A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting,” J. Computer and System Sciences, vol. 55, no. 1, pp. 119-139, 1997.
  13. 13.J. Friedman, T. Hastie, and R. Tibshirani, “Additive Logistic Regression: A Statistical View of Boosting,” Annals of Statistics, vol. 28, no. 2, pp. 337-374, 2000.
  14. 14.K. Fukunaga and W.L.G. Koontz, “Application of the Karhunen-Loeve Expansion to Feature Selection and Ordering,” IEEE Trans. Computers, vol. 19, no. 4, pp. 311-318, Apr. 1970.
  15. 15.J. Han and M. Kamber, Data Mining: Concepts and Techniques. Morgan Kaufmann, 2001.
  16. 16.L.K. Hansen and P. Salamon, “Neural Network Ensembles,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 12, no. 10, pp. 993-1001, Oct. 1990.
  17. 17.T.K. Ho, “The Random Subspace Method for Constructing Decision Forests,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 20, no. 8, pp. 832-844, Aug. 1998.
  18. 18.T.K. Ho, “A Data Complexity Analysis of Comparative Advantages of Decision Forest Constructors,” Pattern Analysis and Applications, vol. 5, pp. 102-112, 2002.
  19. 19.J. Kittler and P.C. Young, “A New Approach to Feature Selection Based on the Karhunen-Loeve Expansion,” Pattern Recognition, vol. 5, no. 4, pp. 335-352, Dec. 1973.
  20. 20.S. Kolenikov and G. Angeles, “The Use of Discrete Data in PCA: Theory, Simulations, and Applications to Socioeconomic Indices,” Proc. 2004 Joint Statistical Meeting, 2004.
  21. 21.L.I. Kuncheva, Combining Pattern Classifiers. Methods and Algorithms. John Wiley and Sons, 2004.
  22. 22.L.I. Kuncheva and C.J. Whitaker, “Measures of Diversity in Classifier Ensembles,” Machine Learning, vol. 51, pp. 181-207, 2003.
  23. 23.L.I. Kuncheva, “Diversity in Multiple Classifier Systems (editorial),” Information Fusion, vol. 6, no. 1, pp. 3-4, 2004.
  24. 24.L.I. Kuncheva, C.J. Whitaker, C.A. Shipp, and R.P.W. Duin, “Is Independence Good for Combining Classifiers?” Proc. 15th Int’l Conf. Pattern Recognition, vol. 2, pp. 169-171, 2000.
  25. 25.P.M. Long and V.B. Vega, “Boosting and Microarray Data,” Machine Learning, vol. 52, pp. 31-44, 2003.
  26. 26.D.D. Margineantu and T.G. Dietterich, “Pruning Adaptive Boosting,” Proc. 14th Int’l Conf. Machine Learning, pp. 211-218, 1997.
  27. 27.L. Mason, P.L. Bartlet, and J. Baxter, “Improved Generalization through Explicit Optimization of Margins,” Machine Learning, vol. 38, no. 3, pp. 243-255, 2000.
  28. 28.P. Melville, N. Shah, L. Mihalkova, and R.J. Mooney, “Experiments with Ensembles with Missing and Noisy Data,” Proc Fifth Int’l Workshop Multiple Classifier Systems, pp. 293-302, 2004.
  29. 29.C. Nadeau and Y. Bengio, “Inference for the Generalization Error,” Machine Learning, vol. 62, pp. 239-281, 2003.
  30. 30.N.C. Oza, “Boosting with Averaged Weight Vectors,” Proc Fourth Int’l Workshop Multiple Classifier Systems (MCS 2003), 2003.
  31. 31.Multiple Classifier Systems, Proc. Sixth Int’l Workshop, MCS 2005, N.C. Oza et al., eds., 2005.
  32. 32.J.R. Quinlan, C4.5: Programs for Machine Learning. Morgan Kaufmann, 1993.
  33. 33.Proc. First Int’l Workshop Multiple Classifier Systems (MCS 2000), F. Roli and J. Kittler, eds., 2001.
  34. 34.Proc. Second Int’l Workshop Multiple Classifier Systems (MCS 2001), F. Roli and J. Kittler, eds., 2001.
  35. 35.Proc. Third Int’l Workshop Multiple Classifier Systems (MCS 2002), F. Roli and J. Kittler, eds., 2002.
  36. 36.Proc. Fifth Int’l Workshop Multiple Classifier Systems (MCS 2004), F. Roli et al., eds., 2004.
  37. 37.R.E. Schapire, “Theoretical Views of Boosting,” Proc. Fourth European Conf. Computational Learning Theory, pp. 1-10, 1999.
  38. 38.R.E. Schapire, “The Boosting Approach to Machine Learning: An Overview,” Proc. MSRI Workshop Nonlinear Estimation and Classification, 2002.
  39. 39.R.E. Schapire, Y. Freund, P. Bartlett, and W.S. Lee, “Boosting the Margin: A New Explanation for the Effectiveness of Voting Methods,” Annals of Statistics, vol. 26, no. 5, pp. 1651-1686, 1998.
  40. 40.R.E. Schapire and Y. Singer, “Improved Boosting Algorithms Using Confidence-Rated Predictions,” Machine Learning, vol. 37, no. 3, pp. 397-336, 1999.
  41. 41.M. Skurichina and R.P.W. Duin, “Combining Feature Subsets in Feature Selection,” Proc. Sixth Int’l Workshop Multiple Classifier Systems, (MCS ’05), pp. 165-175, 2005.
  42. 42.K. Tumer and N.C. Oza, “Input Decimated Ensembles,” Pattern Analysis Applications, vol. 6, pp. 65-77, 2003.
  43. 43.F. van der Heijden, R.P.W. Duin, D. de Ridder, and D.M.J. Tax, Classification, Parameter Estimation and State Estimation. Wiley, 2004.
  44. 44.A. Webb, Statistical Pattern Recognition. London: Arnold, 1999.
  45. 45.G.I. Webb, “MultiBoosting: A Technique for Combining Boosting and Wagging,” Machine Learning, vol. 40, no. 2, pp. 159-196, 2000.
  46. 46.Proc. Fourth Int’l Workshop Multiple Classifier Systems (MCS 2003), T. Windeatt and F. Roli, eds., 2003.
  47. 47.I.H. Witten and E. Frank, Data Mining: Practical Machine Learning Tools and Techniques, second ed. Morgan Kaufmann, 2005.

Citation

MLA
Rodriguez, J. J., et al. “Rotation Forest: A New Classifier Ensemble Method”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 28, no. 10, 2006, pp. 1619–30, https://doi.org/10.1109/TPAMI.2006.211.
APA
Rodriguez, J. J., Kuncheva, L. I., & Alonso, C. J. (2006). Rotation Forest: A New Classifier Ensemble Method. IEEE Transactions on Pattern Analysis and Machine Intelligence, 28(10), 1619–1630. https://doi.org/10.1109/TPAMI.2006.211
Chicago
Rodriguez, J. J., L. I. Kuncheva, and C. J. Alonso. 2006. “Rotation Forest: A New Classifier Ensemble Method”. IEEE Transactions on Pattern Analysis and Machine Intelligence 28 (10): 1619–30. https://doi.org/10.1109/TPAMI.2006.211.
Harvard
Rodriguez, J.J., Kuncheva, L.I. and Alonso, C.J. (2006) “Rotation Forest: A New Classifier Ensemble Method”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 28(10), pp. 1619–1630. Available at: https://doi.org/10.1109/TPAMI.2006.211.
Vancouver
1. Rodriguez JJ, Kuncheva LI, Alonso CJ (2006) Rotation Forest: A New Classifier Ensemble Method. IEEE Transactions on Pattern Analysis and Machine Intelligence 28:1619–1630

BibTeX

@article{Rodriguez_2006, title={Rotation Forest: A New Classifier Ensemble Method}, volume={28}, ISSN={2160-9292}, url={http://dx.doi.org/10.1109/TPAMI.2006.211}, DOI={10.1109/tpami.2006.211}, number={10}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Rodriguez, J.J. and Kuncheva, L.I. and Alonso, C.J.}, year={2006}, month=Oct, pages={1619–1630} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF