Benchmarking Attribute Selection Techniques for Discrete Class Data Mining
Mark A. HallGeoffrey Holmes
Evaluates six major attribute selection methods across multiple benchmark and high-dimensional datasets to determine how ranking-based feature selection impacts the classification performance and efficiency of C4.5 and naive Bayes.
Modern data mining applications frequently suffer from poor predictive accuracy and heavy computational overhead due to datasets contaminated with irrelevant, redundant, and noisy variables. Identifying a compact subset of highly predictive attributes is critical for building efficient, interpretable models, yet practitioners lack comprehensive benchmarks to guide the selection among numerous competing techniques.
The article systematically evaluates and compares six major attribute selection methods for supervised classification on discrete class data. The primary objective is to demonstrate how these ranking-based dimensionality reduction techniques affect predictive accuracy, model complexity, and computational runtime across different types of learning algorithms.
The authors conducted rigorous experimental evaluations using sixteen standard benchmark classification datasets alongside three large-scale datasets from the UCI repository containing up to 1,557 features. Six attribute selection techniques were implemented: Information Gain, ReliefF, Principal Component Analysis, Correlation-based Feature Selection (CFS), a Consistency-based evaluator, and a Wrapper method. These techniques produced ranked lists of variables that were cross-validated to identify optimal subsets for two fundamentally distinct learning schemes: a probabilistic model (naive Bayes) and a decision tree inducer (C4.5).
The benchmark revealed four key findings. First, no single attribute selection method was best across all situations, as performance strongly depended on the underlying inductive bias of the learning algorithm. Second, the Wrapper approach achieved the highest classification accuracy for naive Bayes (scoring 34 wins and 4 losses), but it was computationally prohibitive on large datasets, taking an estimated 140 days on the largest benchmark. Third, for decision trees (C4.5), ReliefF and the Consistency method delivered the top predictive performance because they effectively capture attribute interactions that decision trees exploit during early node splitting. Fourth, CFS consistently eliminated the most unnecessary data, retaining only 42% to 48% of the original attributes on average, while producing the smallest, most interpretable decision trees and executing significantly faster than wrapper or instance-based techniques. Conversely, unsupervised Principal Component Analysis consistently degraded predictive accuracy across models.
These findings indicate that reducing data dimensionality by roughly 50% generally maintains or improves classification performance while producing simpler, more explainable models. For decision-makers, choosing the right attribute selection technique directly affects operational efficiency, project timelines, and model interpretability. Implementing computationally heavy wrapper models introduces severe timeline risks on large datasets without guaranteeing better results than faster heuristic filters.
Organizations should align their feature selection method with their modeling approach and resource constraints. When computational budget allows, Wrapper methods remain the top choice for naive Bayes, whereas ReliefF or Consistency methods are recommended for tree-based models where feature interactions matter. When processing speed, scalability, and compact model size are essential, teams should deploy CFS as a fast, balanced default. Highly complex pipelines should avoid unsupervised Principal Component Analysis for supervised discrete classification.
Confidence in these findings is high for standard tabular classification tasks given the extensive cross-validation methodology. However, caution is warranted when extrapolating these results to non-discrete regression problems, extremely high-dimensional sparse datasets, or architectures beyond naive Bayes and decision trees, which were outside the scope of this evaluation.
- Paper: Correlation-based Feature Selection for Discrete and Numeric Class Machine Learning, Mark A. Hall (1999). Introduces Correlation-based Feature Selection (CFS) and ranking-based attribute evaluation metrics that serve as direct baselines for the attribute selection benchmark.
- Paper: Irrelevant Features and the Subset Selection Problem, George H. John et al. (1994). Formalizes the distinction between filter and wrapper approaches for subset selection and evaluates their performance on standard decision tree and naive Bayes classifiers.
- Paper: Toward Optimal Feature Selection, Daphne Koller et al. (1996). Provides the foundational information-theoretic framework for feature utility estimation and redundant attribute elimination evaluated in classification benchmarks.
- Paper: Feature Selection: Evaluation, Application, and Small Sample Performance, Anil K. Jain et al. (1997). Establishes a systematic benchmark methodology for comparing variable selection techniques across synthetic and empirical datasets.
- Paper: A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection, Ron Kohavi (1995). Defines the cross-validation and accuracy estimation protocols used to evaluate the utility of attribute selection with C4.5 and Naive Bayes.
- Paper: Supervised and Unsupervised Discretization of Continuous Features, James Dougherty et al. (1995). Explores preprocessing and discretization strategies for continuous features in discrete-class learning environments using C4.5 and Naive Bayes.
- Paper: On the Optimality of the Simple Bayesian Classifier under Zero-One Loss, Pedro Domingos et al. (1997). Analyzes the behavior and optimality conditions of the Naive Bayes classifier under varying degrees of attribute dependence.
- Paper: Toward integrating feature selection algorithms for classification and clustering, Huan Liu et al. (2005). Synthesizes and extends benchmark evaluations of feature selection by proposing an integrated categorization framework across classification and clustering tasks.
- Paper: Efficient Feature Selection via Analysis of Relevance and Redundancy, Lei Yu et al. (2004). Extends attribute evaluation techniques by explicitly decoupling relevance ranking from redundancy analysis to scale feature selection for ultra-high-dimensional datasets.
- Paper: Laplacian Score for Feature Selection, Xiaofei He et al. (2005). Generalizes attribute ranking and filter methods from discrete-class supervised mining to graph-based unsupervised feature selection.
- Paper: Auto-WEKA: combined selection and hyperparameter optimization of classification algorithms, Chris J. Thornton et al. (2012). Builds on discrete-class benchmark evaluations by automating the combined selection of feature selection methods and classification algorithm hyperparameters.
