Correlation-based Feature Selection for Discrete and Numeric Class Machine Learning
Mark A. Hall
Introduces Correlation-based Feature Selection (CFS), a fast filter algorithm for discrete and continuous learning tasks that evaluates feature subsets by balancing class correlation against inter-feature redundancy to cut data dimensionality in half and improve model compactness.
High-dimensional datasets containing irrelevant or redundant information present a major operational challenge in machine learning, slowing down model training and increasing the risk of overfitting. While wrapper-based selection methods are computationally expensive and traditional filter algorithms largely focus on discrete classification problems, there is a strong need for efficient, universal preprocessing techniques that scale to large datasets across both discrete and continuous prediction tasks.
The article sets out to develop and evaluate Correlation-based Feature Selection (CFS), a fast filter method designed to assess feature subsets rather than individual features, and to demonstrate its effectiveness across both discrete classification and continuous regression problems.
The analysis evaluated CFS across 16 discrete and 19 continuous benchmark datasets from the UCI repository, comparing its performance against the established ReliefF algorithm. Evaluated over multiple ten-fold cross-validation runs, the filtered datasets were processed using diverse learning algorithms, including probabilistic methods, tree-based models, and nearest-neighbor approaches.
The primary findings show that CFS is an aggressive and effective feature selector. First, CFS reduced data dimensionality by approximately 47% on discrete datasets and 54% on continuous datasets, systematically outperforming ReliefF in pruning unneeded variables. Second, CFS maintained or significantly improved the predictive accuracy of downstream learning algorithms in the majority of cases, enhancing classification accuracy on up to six datasets and lowering regression error across eight to nine datasets. Third, preprocessing with CFS produced substantially smaller decision trees in 56% of discrete cases and reduced the number of linear models in 42% of regression tree models, without ever increasing tree size in classification tasks.
These results demonstrate that organizations can lower computational costs and streamline machine learning pipelines by stripping away roughly half of their data features without sacrificing predictive power. By generating smaller, less complex decision trees, CFS improves model interpretability, which reduces deployment risk and simplifies compliance in regulated environments.
Technical leaders and data science teams should consider adopting CFS as a standard preprocessing filter for large-scale classification and regression workflows, particularly when wrapper methods are too slow. Practitioners should, however, evaluate whether their specific domain relies heavily on complex multi-feature interactions, as CFS evaluates features based on individual and pairwise correlations and cannot detect strongly interacting attributes such as parity problems. Given the strong evidence base across diverse standard benchmarks, stakeholders can have high confidence in using CFS for general tabular data tasks.
- Paper: Irrelevant Features and the Subset Selection Problem, George H. John et al. (1994). It establishes the foundational distinction between filter and wrapper models of feature selection and formalizes concepts of attribute relevance that motivate fast correlation-based filters.
- Paper: Multi-Interval Discretization of Continuous-Valued Attributes for Classification Learning, Usama M. Fayyad et al. (1993). It develops the standard minimum description length discretization method for continuous attributes that enables discrete-oriented feature selection metrics to evaluate continuous data.
- Paper: Supervised and Unsupervised Discretization of Continuous Features, James Dougherty et al. (1995). It provides the foundational study on supervised entropy-based discretization heuristics for continuous features in classifiers like Naive Bayes and decision trees.
- Paper: Feature Selection: Evaluation, Application, and Small Sample Performance, Anil K. Jain et al. (1997). It systematically analyzes feature subset search algorithms and evaluations across varying sample sizes, framing the computational trade-offs addressed by heuristic filters.
- Paper: Instance-based learning algorithms, D. Aha et al. (1991). It introduces standard instance-based learning algorithms (IB1–IB3), which serve as key evaluation benchmarks in the source paper's experiments.
- Paper: On the Optimality of the Simple Bayesian Classifier under Zero-One Loss, Pedro M. Domingos et al. (1997). It analyzes the behavior and optimality of simple Bayesian classifiers under feature dependencies, directly motivating the removal of redundant attributes.
- Paper: Feature Selection for High-Dimensional Data: A Fast Correlation-Based Filter Solution, Lei Yu et al. (2003). It advances correlation-based filtering to very high-dimensional settings by introducing predominant correlation and the Fast Correlation-Based Filter (FCBF) algorithm based on symmetrical uncertainty.
- Paper: Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy, Hanchuan Peng et al. (2003). It formulates the minimal-redundancy-maximal-relevance (mRMR) framework using mutual information, extending the concept of balancing feature-target relevance with inter-feature redundancy.
- Paper: An Introduction to Variable and Feature Selection, Isabelle M Guyon et al. (2003). It provides a comprehensive post-2000 taxonomy and practical methodology synthesizing correlation ranking, filter approaches, wrappers, and embedded selection techniques.
- Paper: Toward integrating feature selection algorithms for classification and clustering, Huan Liu et al. (2005). It synthesizes filter and wrapper algorithms, including correlation-based criteria, into a unified evaluation framework spanning classification and clustering tasks.
- Paper: Theoretical and Empirical Analysis of ReliefF and RReliefF, M. Robnik-Sikonja et al. (2003). It provides theoretical and empirical analysis of the ReliefF and RReliefF attribute estimation family, the primary baseline against which the source algorithm is compared.
- Paper: Gene Selection for Cancer Classification using Support Vector Machines, ISABELLE GUYON et al. (2002). It introduces SVM Recursive Feature Elimination, establishing an embedded alternative to correlation-based filtering for high-dimensional feature selection.
