Feature Selection for High-Dimensional Data: A Fast Correlation-Based Filter Solution
Lei YuHuan Liu
Introduces the Fast Correlation-Based Filter algorithm, which efficiently removes both irrelevant and redundant features from high-dimensional data in sub-quadratic time using the concept of predominant correlation without requiring exhaustive pairwise comparisons.
High-dimensional datasets in applications such as genomics, text categorization, and image retrieval contain many irrelevant and redundant features that degrade the speed and accuracy of machine learning algorithms. Feature selection serves as an essential preprocessing step, yet existing filter and wrapper methods struggle with scalability when the number of features reaches hundreds or thousands.
The article set out to develop and test a fast correlation-based filter algorithm that removes both irrelevant and redundant features without exhaustive pairwise comparisons.
The authors introduced the concept of predominant correlation and built the FCBF algorithm around symmetrical uncertainty, an entropy-based measure. They evaluated it on ten UCI benchmark datasets ranging from 57 to 650 features, comparing runtime, number of selected features, and classification accuracy against ReliefF, CorrSF, and ConsSF using C4.5 and naïve Bayes classifiers.
FCBF ran orders of magnitude faster than the alternatives, reduced the feature set more aggressively than competing methods in nine of ten cases, and produced average accuracy gains for both classifiers (C4.5 rose from 88.6 % to 89.1 %; naïve Bayes rose from 82.2 % to 86.9 %). Accuracy either held steady or improved on most individual datasets.
These results indicate that predominant-correlation filtering can deliver compact, high-performing feature subsets at low computational cost, making reliable feature selection practical for current high-dimensional tasks and supporting downstream gains in speed, interpretability, and generalization.
The authors recommend extending FCBF to datasets with thousands of features, examining the role of redundant features in greater depth, and integrating discretization routines for mixed data types. Further empirical validation on larger-scale problems is needed before broad deployment.
The study is limited to ten moderate-sized UCI datasets and requires a user-specified relevance threshold; performance on streaming or extremely sparse data remains untested. Confidence is moderate to high for the tested regime but should be tempered when extrapolating to substantially larger or qualitatively different data.
- Paper: Irrelevant Features and the Subset Selection Problem, George H. John et al. (1994). This paper establishes the foundational formal definitions of feature relevance and redundancy that underpin the filter selection criteria formalized by FCBF.
- Paper: Supervised and Unsupervised Discretization of Continuous Features, James Dougherty et al. (1995). Understanding entropy-based discretization provides essential background for how FCBF preprocesses continuous attributes to compute symmetrical uncertainty and evaluate classification models like Naive Bayes.
- Paper: Gene Selection for Cancer Classification using Support Vector Machines, ISABELLE GUYON et al. (2002). This work introduces seminal embedded feature elimination techniques for high-dimensional microarray data, representing the primary alternative paradigm that fast filter methods like FCBF aim to outperform in efficiency.
- Paper: Feature Selection: Evaluation, Application, and Small Sample Performance, Anil K. Jain et al. (1997). This comprehensive evaluation of classical subset search strategies highlights the computational bottlenecks that motivate FCBF's fast heuristic filter design.
- Paper: Theoretical and Empirical Analysis of ReliefF and RReliefF, M. Robnik-Sikonja et al. (2003). This study analyzes ReliefF, which serves as one of the primary benchmark algorithms evaluated against FCBF for handling high-dimensional attribute quality estimation.
- Paper: Toward integrating feature selection algorithms for classification and clustering, Huan Liu et al. (2005). Co-authored by the creators of FCBF, this survey synthesizes the filter framework into an integrated taxonomy covering search strategies and data mining tasks.
- Paper: An Introduction to Variable and Feature Selection, Isabelle M Guyon et al. (2003). This comprehensive tutorial expands on correlation, ranking, and subset selection methods across modern high-dimensional bioinformatics and text classification benchmarks.
- Paper: Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy, Hanchuan Peng et al. (2003). This paper advances the theoretical and empirical treatment of relevance versus redundancy using mutual information criteria in the minimal-redundancy-maximal-relevance (mRMR) algorithm.
- Paper: On Model Selection Consistency of Lasso, P. Zhao et al. (2006). This work formalizes the mathematical conditions required for sparse model consistency under regularized regression, extending the study of correlated predictor redundancy.
