Feature Selection for High-Dimensional Data: A Fast Correlation-Based Filter Solution

Lei YuHuan Liu

article2003ICML2,787 citations

Introduces the Fast Correlation-Based Filter algorithm, which efficiently removes both irrelevant and redundant features from high-dimensional data in sub-quadratic time using the concept of predominant correlation without requiring exhaustive pairwise comparisons.

Listen

High-dimensional datasets in applications such as genomics, text categorization, and image retrieval contain many irrelevant and redundant features that degrade the speed and accuracy of machine learning algorithms. Feature selection serves as an essential preprocessing step, yet existing filter and wrapper methods struggle with scalability when the number of features reaches hundreds or thousands.

The article set out to develop and test a fast correlation-based filter algorithm that removes both irrelevant and redundant features without exhaustive pairwise comparisons.

The authors introduced the concept of predominant correlation and built the FCBF algorithm around symmetrical uncertainty, an entropy-based measure. They evaluated it on ten UCI benchmark datasets ranging from 57 to 650 features, comparing runtime, number of selected features, and classification accuracy against ReliefF, CorrSF, and ConsSF using C4.5 and naïve Bayes classifiers.

FCBF ran orders of magnitude faster than the alternatives, reduced the feature set more aggressively than competing methods in nine of ten cases, and produced average accuracy gains for both classifiers (C4.5 rose from 88.6 % to 89.1 %; naïve Bayes rose from 82.2 % to 86.9 %). Accuracy either held steady or improved on most individual datasets.

These results indicate that predominant-correlation filtering can deliver compact, high-performing feature subsets at low computational cost, making reliable feature selection practical for current high-dimensional tasks and supporting downstream gains in speed, interpretability, and generalization.

The authors recommend extending FCBF to datasets with thousands of features, examining the role of redundant features in greater depth, and integrating discretization routines for mixed data types. Further empirical validation on larger-scale problems is needed before broad deployment.

The study is limited to ten moderate-sized UCI datasets and requires a user-specified relevance threshold; performance on streaming or extremely sparse data remains untested. Confidence is moderate to high for the tested regime but should be tempered when extrapolating to substantially larger or qualitatively different data.

Cover for Feature Selection for High-Dimensional Data: A Fast Correlation-Based Filter Solution

Abstract

Feature selection, as a preprocessing step to machine learning, is effective in reducing dimensionality, removing irrelevant data, increasing learning accuracy, and improving result comprehensibility. However, the recent increase of dimensionality of data poses a severe challenge to many existing feature selection methods with respect to efficiency and effectiveness. In this work, we introduce a novel concept, predominant correlation, and propose a fast filter method which can identify relevant features as well as redundancy among relevant features without pairwise correlation analysis. The efficiency and effectiveness of our method is demonstrated through extensive comparisons with other methods using real-world data of high dimensionality.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Correlation-Based Measures
  • 4. A Correlation-Based Filter Approach
  • 4.1. Methodology
  • 4.2. Algorithm and Analysis
  • 5. Empirical Study
  • 5.1. Experiment Setup
  • 5.2. Results and Discussions
  • 6. Conclusions
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Fast Correlation-Based Filter (FCBF) Algorithm

    algorithm

    The Fast Correlation-Based Filter (FCBF) algorithm selects a subset of predominant features SbestS_{best} from a dataset with NN features and a target class CC using symmetrical uncertainty (SUSU) as the correlation measure.

    FCBF operates in two stages: (1) Relevance filtering: It computes the correlation between each feature FiF_i and the class CC (SUi,cSU_{i,c}), filters out features whose SUi,cSU_{i,c} is below a user-defined threshold δ\delta, and sorts the remaining candidate features into an ordered list Slist′S'_{list} in descending order of SUi,cSU_{i,c}. (2) Redundancy filtering: Starting from the highest-ranked feature FpF_p (a predominant feature), it scans all lower-ranked features FqF_q. If the feature-feature correlation SUp,q≥SUq,cSU_{p,q} \ge SU_{q,c}, FqF_q is identified as a redundant peer of FpF_p and is removed from Slist′S'_{list}. Once a pass is complete, the next remaining feature in the list is selected as the new reference feature FpF_p, and the filtering process repeats until no further features can be removed.

    Input: S(F_1, F_2, ..., F_N, C) // training data set
           δ\delta // predefined threshold
    Output: S_best // optimal feature subset
    begin
      for i = 1 to N do begin
        calculate SU_{i,c} for F_i;
        if (SU_{i,c} >= δ\delta)
          append F_i to S'_list;
      end;
      order S'_list in descending SU_{i,c} value;
      F_p = getFirstElement(S'_list);
      do begin
        F_q = getNextElement(S'_list, F_p);
        if (F_q <> NULL)
          do begin
            F'_q = F_q;
            if (SU_{p,q} >= SU_{q,c})
              remove F_q from S'_list;
              F_q = getNextElement(S'_list, F'_q);
            else
              F_q = getNextElement(S'_list, F_q);
          end until (F_q == NULL);
        F_p = getNextElement(S'_list, F_p);
      end until (F_p == NULL);
      S_best = S'_list;
    end;
  2. Knowl 2 — Predominant Correlation and Predominant Features

    definition

    Let S={F1,F2,…,FN}S = \{F_1, F_2, \dots, F_N\} be the full set of features, CC be the target class concept, and SU(X,Y)SU(X, Y) denote the symmetrical uncertainty between two variables. Let S′⊆SS' \subseteq S be the set of relevant features whose correlation with the class satisfies SU(Fi,C)≥δSU(F_i, C) \ge \delta for a chosen threshold δ>0\delta > 0.

    • Predominant Correlation: The correlation between a feature Fi∈S′F_i \in S' and the class CC is defined as predominant if and only if SU(Fi,C)≥δSU(F_i, C) \ge \delta, and for all Fj∈S′F_j \in S' (j≠ij \neq i), there exists no feature FjF_j such that SU(Fj,Fi)≥SU(Fi,C)SU(F_j, F_i) \ge SU(F_i, C).

    • Redundant Peer: If there exists a feature Fj∈S′F_j \in S' (j≠ij \neq i) such that SU(Fj,Fi)≥SU(Fi,C)SU(F_j, F_i) \ge SU(F_i, C), then FjF_j is called a redundant peer to FiF_i. The set of all redundant peers for FiF_i is denoted by SPiS_{P_i}. This set is partitioned into two disjoint subsets: SPi+={Fj∣Fj∈SPi,  SU(Fj,C)>SU(Fi,C)}S_{P_i}^+ = \{F_j \mid F_j \in S_{P_i},\; SU(F_j, C) > SU(F_i, C)\} SPi−={Fj∣Fj∈SPi,  SU(Fj,C)≤SU(Fi,C)}S_{P_i}^- = \{F_j \mid F_j \in S_{P_i},\; SU(F_j, C) \le SU(F_i, C)\}

    • Predominant Feature: A feature is predominant to the class if and only if its correlation to the class is predominant or becomes predominant after removing its redundant peers.

  3. Knowl 3 — Heuristics for Redundancy Elimination in FCBF

    model/method

    To eliminate redundant features without computing full pairwise feature-feature correlation matrices (O(N2)O(N^2) pairs), the Fast Correlation-Based Filter (FCBF) relies on the principle that if two features are redundant to each other, removing the one with lower class correlation preserves more predictive information. This is operationalized via three heuristics:

    • Heuristic 1 (Empty SPi+S_{P_i}^+): If SPi+=∅S_{P_i}^+ = \emptyset, feature FiF_i has no redundant peers with higher correlation to the class CC. Therefore, FiF_i is treated as a predominant feature, all features in its lower-correlation redundant peer set SPi−S_{P_i}^- are removed from the candidate pool, and identifying redundant peers for those removed features is skipped.

    • Heuristic 2 (Non-empty SPi+S_{P_i}^+): If SPi+≠∅S_{P_i}^+ \neq \emptyset, all features in SPi+S_{P_i}^+ must be processed before deciding the status of FiF_i. If none of the features in SPi+S_{P_i}^+ becomes predominant, Heuristic 1 is applied to FiF_i; otherwise, FiF_i is removed, and whether to remove features in SPi−S_{P_i}^- is determined based on the remaining candidate features.

    • Heuristic 3 (Starting Point): The candidate feature with the highest correlation to the class SU(Fi,C)SU(F_i, C) in the relevant feature set S′S' is unconditionally a predominant feature (SPi+=∅S_{P_i}^+ = \emptyset) and serves as the deterministic starting point for filtering redundant features.

  4. Knowl 4 — Symmetrical Uncertainty Correlation Measure

    equation

    For two discrete random variables XX and YY, symmetrical uncertainty SU(X,Y)SU(X, Y) normalizes information gain to the range [0,1][0, 1] and compensates for information gain's inherent bias toward features with larger numbers of values:

    SU(X,Y)=2[IG(X∣Y)H(X)+H(Y)]SU(X, Y) = 2 \left[ \frac{IG(X \mid Y)}{H(X) + H(Y)} \right]

    where the Shannon entropy of XX is defined as:

    H(X)=−∑iP(xi)log⁡2P(xi)H(X) = -\sum_{i} P(x_i) \log_2 P(x_i)

    the conditional entropy of XX given YY is defined as:

    H(X∣Y)=−∑jP(yj)∑iP(xi∣yj)log⁡2P(xi∣yj)H(X \mid Y) = -\sum_{j} P(y_j) \sum_{i} P(x_i \mid y_j) \log_2 P(x_i \mid y_j)

    and the information gain IG(X∣Y)IG(X \mid Y) is:

    IG(X∣Y)=H(X)−H(X∣Y)IG(X \mid Y) = H(X) - H(X \mid Y)

    A value of SU(X,Y)=1SU(X, Y) = 1 indicates that knowing either variable completely predicts the other, whereas SU(X,Y)=0SU(X, Y) = 0 indicates that XX and YY are entirely independent. Symmetrical uncertainty is symmetric, satisfying SU(X,Y)=SU(Y,X)SU(X, Y) = SU(Y, X). Continuous variables must be discretized prior to computing SUSU.

  5. Knowl 5 — Computational Time Complexity of the FCBF Algorithm

    theoretical result

    For a dataset with MM instances and NN features, the overall average computational time complexity of the Fast Correlation-Based Filter (FCBF) algorithm is O(MNlog⁡N)O(M N \log N).

    • Relevance Step: Computing the symmetrical uncertainty between every feature FiF_i and the class label CC requires O(MN)O(M N) operations. Sorting the filtered subset of relevant features takes O(Nlog⁡N)O(N \log N) operations.
    • Redundancy Step: In each filtering iteration, a predominant feature FpF_p removes a fraction of the remaining candidate features that are redundant peers to FpF_p. Under the average-case assumption that a constant fraction (e.g., half) of the remaining features are pruned in each round, the number of pairwise feature comparisons is O(Nlog⁡N)O(N \log N). Because computing each pairwise symmetrical uncertainty takes O(M)O(M) time, the redundancy filtering stage runs in O(MNlog⁡N)O(M N \log N) time on average.

    In the worst case (where no redundant features are removed), the complexity is O(MN2)O(M N^2), while in the best case (where all subsequent features are redundant to the first), it is O(MN)O(M N).

  6. Knowl 6 — Symmetry of Information Gain

    theoretical result

    For any two discrete random variables XX and YY, the information gain measure is symmetric:

    IG(X∣Y)=IG(Y∣X)IG(X \mid Y) = IG(Y \mid X)

    This holds because the joint entropy decomposes as H(X,Y)=H(X)+H(Y∣X)=H(Y)+H(X∣Y)H(X, Y) = H(X) + H(Y \mid X) = H(Y) + H(X \mid Y), which implies H(X)−H(X∣Y)=H(Y)−H(Y∣X)H(X) - H(X \mid Y) = H(Y) - H(Y \mid X).

  7. Knowl 7 — Experimental Setup for Filter Feature Selection Benchmark

    experimental setup

    The Fast Correlation-Based Filter (FCBF) was evaluated against three filter-based feature selection baselines using 10 benchmark datasets from the UCI Machine Learning Repository and UCI KDD Archive: Lung-cancer (57 features, 32 instances, 3 classes), Promoters (59 features, 106 instances, 2 classes), Splice (62 features, 3,190 instances, 3 classes), USCensus90 (68 features, 9,338 instances, 3 classes), CoIL2000 (86 features, 5,822 instances, 2 classes), Chemical (151 features, 936 instances, 3 classes), Musk2 (169 features, 6,598 instances, 2 classes), Arrhythmia (280 features, 452 instances, 16 classes), Isolet (618 features, 1,560 instances, 26 classes), and Multi-features (650 features, 2,000 instances, 10 classes).

    The compared feature selection algorithms are:

    1. ReliefF: A feature weighting algorithm configured with k=5k = 5 nearest neighbors and m=30m = 30 sampled instances.
    2. CorrSF: A correlation-based subset search algorithm using sequential forward search guided by CFS correlation merit.
    3. ConsSF: A consistency-based subset search algorithm using sequential forward search guided by inconsistency rate.

    Performance was assessed on two classification algorithms: the C4.5 decision tree and the Naive Bayes Classifier (NBC). Classification accuracy was evaluated using 10-fold cross-validation on both the full feature sets and the reduced subsets obtained by each feature selection method.

  8. Knowl 8 — Execution Time and Dimensionality Reduction Comparison Across Feature Selection Algorithms

    data/table

    Execution time (in milliseconds) and the number of selected features were measured for FCBF, CorrSF, ReliefF, and ConsSF across 10 benchmark datasets. FCBF demonstrated superior computational speed (averaging 995 ms compared to 32,028 ms for CorrSF, 7,911 ms for ReliefF, and 109,617 ms for ConsSF) and achieved the highest degree of dimensionality reduction (retaining an average of 7 features, compared to 30 for CorrSF, 11 for ReliefF, and 12 for ConsSF).

    Dataset Running Time (ms) # Selected Features
    FCBF CorrSF ReliefF ConsSF FCBF CorrSF ReliefF ConsSF
    Lung-cancer 20 50 50 110 5 8 5 4
    Promoters 20 50 100 190 4 4 4 4
    Splice 200 961 2343 34920 6 6 11 10
    USCensus90 541 932 7601 161121 2 1 2 13
    CoIL2000 470 3756 7751 341231 3 10 12 29
    Chemical 121 450 2234 14000 4 7 7 11
    Musk2 971 8903 18066 175453 2 10 2 11
    Arrhythmia 151 2002 2233 31235 6 25 25 24
    Isolet 3174 177986 17025 203973 23 137 23 11
    Multi-Features 4286 125190 21711 133932 14 87 14 7
    Average 995 32028 7911 109617 7 30 11 12
  9. Knowl 9 — Classification Accuracy of C4.5 and Naive Bayes Classifiers After Feature Selection

    data/table

    Classification accuracy (percentage ±\pm standard deviation from 10-fold cross-validation) was evaluated for C4.5 and Naive Bayes Classifier (NBC) on the full feature set and on subsets selected by FCBF, CorrSF, ReliefF, and ConsSF. Feature selection with FCBF improved the average accuracy of both C4.5 (from 88.62% on the full feature set to 89.13%) and NBC (from 82.20% on the full feature set to 86.92%), matching or exceeding the baseline feature selectors while using fewer features.

    Dataset Full Set FCBF CorrSF ReliefF ConsSF
    C4.5 Accuracy (%)
    Lung-cancer 80.83 ±22.92\pm 22.92 87.50 ±16.32\pm 16.32 84.17 ±16.87\pm 16.87 80.83 ±22.92\pm 22.92 84.17 ±16.87\pm 16.87
    Promoters 86.91 ±6.45\pm 6.45 87.73 ±6.55\pm 6.55 87.73 ±6.55\pm 6.55 89.64 ±5.47\pm 5.47 84.00 ±6.15\pm 6.15
    Splice 94.14 ±1.57\pm 1.57 93.48 ±2.20\pm 2.20 93.48 ±2.20\pm 2.20 89.25 ±1.94\pm 1.94 93.92 ±1.53\pm 1.53
    USCensus90 98.27 ±0.19\pm 0.19 98.08 ±0.22\pm 0.22 97.95 ±0.15\pm 0.15 98.08 ±0.22\pm 0.22 98.22 ±0.30\pm 0.30
    CoIL2000 93.97 ±0.21\pm 0.21 94.02 ±0.07\pm 0.07 94.02 ±0.07\pm 0.07 94.02 ±0.07\pm 0.07 93.99 ±0.20\pm 0.20
    Chemical 94.65 ±2.03\pm 2.03 95.51 ±2.31\pm 2.31 96.47 ±2.15\pm 2.15 93.48 ±1.79\pm 1.79 95.72 ±2.09\pm 2.09
    Musk2 96.79 ±0.81\pm 0.81 91.33 ±0.51\pm 0.51 95.56 ±0.73\pm 0.73 94.62 ±0.92\pm 0.92 95.38 ±0.75\pm 0.75
    Arrhythmia 67.25 ±3.68\pm 3.68 72.79 ±6.30\pm 6.30 68.58 ±7.41\pm 7.41 65.90 ±8.23\pm 8.23 67.48 ±4.49\pm 4.49
    Isolet 79.10 ±2.79\pm 2.79 75.77 ±4.07\pm 4.07 80.70 ±4.94\pm 4.94 52.44 ±3.61\pm 3.61 69.23 ±4.53\pm 4.53
    Multi-Features 94.30 ±1.49\pm 1.49 95.06 ±0.86\pm 0.86 94.95 ±0.96\pm 0.96 80.45 ±2.41\pm 2.41 90.80 ±1.75\pm 1.75
    Average 88.62 ±9.99\pm 9.99 89.13 ±8.52\pm 8.52 89.36 ±9.24\pm 9.24 83.87 ±14.56\pm 14.56 87.29 ±11.04\pm 11.04
    NBC Accuracy (%)
    Lung-cancer 80.00 ±23.31\pm 23.31 90.00 ±16.10\pm 16.10 90.00 ±16.10\pm 16.10 80.83 ±22.92\pm 22.92 86.67 ±17.21\pm 17.21
    Promoters 90.45 ±7.94\pm 7.94 94.45 ±8.83\pm 8.83 94.45 ±8.83\pm 8.83 87.82 ±10.99\pm 10.99 92.64 ±7.20\pm 7.20
    Splice 95.33 ±0.88\pm 0.88 93.60 ±1.74\pm 1.74 93.60 ±1.74\pm 1.74 88.40 ±1.97\pm 1.97 94.48 ±1.39\pm 1.39
    USCensus90 93.38 ±0.90\pm 0.90 97.93 ±0.16\pm 0.16 97.95 ±0.15\pm 0.15 97.93 ±0.16\pm 0.16 97.87 ±0.26\pm 0.26
    CoIL2000 79.03 ±2.08\pm 2.08 93.94 ±0.21\pm 0.21 92.94 ±0.80\pm 0.80 93.58 ±0.43\pm 0.43 83.18 ±1.94\pm 1.94
    Chemical 60.79 ±5.98\pm 5.98 72.11 ±2.51\pm 2.51 70.72 ±4.20\pm 4.20 78.20 ±3.58\pm 3.58 67.20 ±2.51\pm 2.51
    Musk2 84.69 ±2.01\pm 2.01 84.59 ±0.07\pm 0.07 64.85 ±2.09\pm 2.09 84.59 ±0.07\pm 0.07 83.56 ±1.05\pm 1.05
    Arrhythmia 60.61 ±3.32\pm 3.32 66.61 ±5.89\pm 5.89 68.80 ±4.22\pm 4.22 66.81 ±3.62\pm 3.62 68.60 ±7.64\pm 7.64
    Isolet 83.72 ±2.38\pm 2.38 80.06 ±2.52\pm 2.52 86.28 ±2.14\pm 2.14 52.37 ±3.17\pm 3.17 71.67 ±3.08\pm 3.08
    Multi-Features 93.95 ±1.50\pm 1.50 95.95 ±1.06\pm 1.06 96.15 ±0.94\pm 0.94 76.05 ±3.26\pm 3.26 93.75 ±1.95\pm 1.95
    Average 82.20 ±12.70\pm 12.70 86.92 ±10.79\pm 10.79 85.57 ±12.53\pm 12.53 80.66 ±13.38\pm 13.38 83.96 ±11.31\pm 11.31

Coverage note — None was omitted; all key theoretical definitions, the FCBF algorithm, time complexity derivations, and experimental results have been captured as self-contained knowls.

References

  1. 1.Bay, S. D. (1999). The UCI KDD Archive. http://kdd.ics.uci.edu.
  2. 2.Blake, C., & Merz, C. (1998). UCI repository of machine learning databases. http://www.ics.uci.edu/~mlearn/MLRepository.html.
  3. 3.Blum, A., & Langley, P. (1997). Selection of relevant features and examples in machine learning. Artificial Intelligence, 97, 245–271.
  4. 4.Das, S. (2001). Filters, wrappers and a boosting-based hybrid for feature selection. Proceedings of the Eighteenth International Conference on Machine Learning (pp. 74–81).
  5. 5.Das, S. K. (1971). Feature selection with a linear dependence measure. IEEE Transactions on Computers.
  6. 6.Dash, M., & Liu, H. (1997). Feature selection for classifications. Intelligent Data Analysis: An International Journal, 1, 131–156.
  7. 7.Dash, M., Liu, H., & Motoda, H. (2000). Consistency based feature selection. Proceedings of the Fourth Pacific Asia Conference on Knowledge Discovery and Data Mining (pp. 98–109). Springer-Verlag.
  8. 8.Fayyad, U., & Irani, K. (1993). Multi-interval discretization of continuous-valued attributes for classification learning. Proceedings of the Thirteenth International Joint Conference on Artificial Intelligence (pp. 1022–1027). Morgan Kaufmann.
  9. 9.Hall, M. (1999). Correlation based feature selection for machine learning. Doctoral dissertation, University of Waikato, Dept. of Computer Science.
  10. 10.Hall, M. (2000). Correlation-based feature selection for discrete and numeric class machine learning. Proceedings of the Seventeenth International Conference on Machine Learning (pp. 359–366).
  11. 11.Kira, K., & Rendell, L. (1992). The feature selection problem: Traditional methods and a new algorithm. Proceedings of the Tenth National Conference on Artificial Intelligence (pp. 129–134). Menlo Park: AAAI Press/The MIT Press.
  12. 12.Kohavi, R., & John, G. (1997). Wrappers for feature subset selection. Artificial Intelligence, 97, 273–324.
  13. 13.Kononenko, I. (1994). Estimating attributes : Analysis and extension of RELIEF. Proceedings of the European Conference on Machine Learning (pp. 171–182). Catania, Italy: Berlin: Springer-Verlag.
  14. 14.Langley, P. (1994). Selection of relevant features in machine learning. Proceedings of the AAAI Fall Symposium on Relevance. AAAI Press.
  15. 15.Liu, H., Hussain, F., Tan, C., & Dash, M. (2002a). Discretization: An enabling technique. Data Mining and Knowledge Discovery, 6, 393–423.
  16. 16.Liu, H., & Motoda, H. (1998). Feature selection for knowledge discovery and data mining. Boston: Kluwer Academic Publishers.
  17. 17.Liu, H., Motoda, H., & Yu, L. (2002b). Feature selection with selective sampling. Proceedings of the Nineteenth International Conference on Machine Learning (pp. 395 – 402).
  18. 18.Mitra, P., Murthy, C. A., & Pal, S. K. (2002). Unsupervised feature selection using feature similarity. IEEE Transactions on Pattern Analysis and Machine Intelligence, 24, 301–312.
  19. 19.Ng, A. Y. (1998). On feature selection: learning with exponentially many irrelevant features as training examples. Proceedings of the Fifteenth International Conference on Machine Learning (pp. 404–412).
  20. 20.Ng, K., & Liu, H. (2000). Customer retention via data mining. AI Review, 14, 569 – 590.
  21. 21.Press, W. H., Flannery, B. P., Teukolsky, S. A., & Vetterling, W. T. (1988). Numerical recipes in C. Cambridge University Press, Cambridge.
  22. 22.Quinlan, J. (1993). C4.5: Programs for machine learning. Morgan Kaufmann.
  23. 23.Rui, Y., Huang, T. S., & Chang, S. (1999). Image retrieval: Current techniques, promising directions and open issues. Journal of Visual Communication and Image Representation, 10, 39–62.
  24. 24.Witten, I., & Frank, E. (2000). Data mining - pracitcal machine learning tools and techniques with JAVA implementations. Morgan Kaufmann Publishers.
  25. 25.Xing, E., Jordan, M., & Karp, R. (2001). Feature selection for high-dimensional genomic microarray data. Proceedings of the Eighteenth International Conference on Machine Learning (pp. 601–608).
  26. 26.Yang, Y., & Pederson, J. O. (1997). A comparative study on feature selection in text categorization. Proceedings of the Fourteenth International Conference on Machine Learning (pp. 412–420).

Citation

MLA
Yu, L., and H. Liu. “Feature Selection for High-dimensional Data: A Fast Correlation-based Filter Solution”. International Conference on Machine Learning, 2003, pp. 856–63, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.566.3227.
APA
Yu, L., & Liu, H. (2003). Feature selection for high-dimensional data: a fast correlation-based filter solution. International Conference on Machine Learning, 856–863. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.566.3227
Chicago
Yu, L., and H. Liu. 2003. “Feature Selection for High-dimensional Data: A Fast Correlation-based Filter Solution”. International Conference on Machine Learning, 856–63. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.566.3227.
Harvard
Yu, L. and Liu, H. (2003) “Feature selection for high-dimensional data: a fast correlation-based filter solution”, International Conference on Machine Learning, pp. 856–863. Available at: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.566.3227.
Vancouver
1. Yu L, Liu H (2003) Feature selection for high-dimensional data: a fast correlation-based filter solution. International Conference on Machine Learning 856–863

BibTeX

@article{yu2003feature,
  title = {Feature selection for high-dimensional data: a fast correlation-based filter solution},
  author = {Yu, Lei and Liu, Huan},
  year = {2003},
  journal = {International Conference on Machine Learning},
  pages = {856-863},
  url = {http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.566.3227}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF