Feature Selection: Evaluation, Application, and Small Sample Performance
Anil K. JainDouglas E. Zongker
Compares prominent feature subset selection methods on synthetic benchmarks and SAR satellite imagery, establishing the superior performance of sequential forward floating selection while identifying critical pitfalls of selection algorithms in small-sample scenarios.
Modern pattern recognition and data mining systems often gather hundreds of measurements across multiple sources or mathematical representations. While capturing more dimensions can theoretically improve system performance, including too many irrelevant or redundant inputs increases processing costs and can paradoxically degrade classification accuracy. Consequently, selecting an optimal subset of informative variables is critical for building efficient, high-performing automated decision systems.
The article evaluates the practical performance of fifteen major variable selection techniques across synthetic and real-world datasets. It specifically assesses how these algorithms balance computational effort and accuracy, how multi-model information can be combined effectively, and how limited training sample sizes impact the reliability of subset selection.
To conduct this evaluation, the researchers tested a range of statistical and neural-network-based selection techniques using controlled synthetic distributions as well as a real-world satellite terrain classification dataset comprising eighteen texture features from four different mathematical models. Controlled simulations were also run with varying training sample sizes, ranging from very small to large datasets, to measure how accurately selection routines recover known optimal subsets.
The primary finding is that the sequential forward floating selection method consistently delivers near-optimal accuracy while remaining computationally efficient, clearly outperforming standard forward, backward, and genetic search algorithms on complex tasks. Second, combining measurements from diverse mathematical models and applying subset selection substantially boosted satellite image classification accuracy from a baseline of under seventy percent up to approximately eighty-nine percent. Third, the results demonstrate that small training sample sizes severely degrade selection quality, causing algorithms to pick suboptimal variable subsets due to estimation errors in high-dimensional spaces.
These findings indicate that integrating multiple complementary data representations combined with robust subset selection provides a practical pathway to superior operational accuracy. However, practitioners must recognize the high risk of over-fitting and poor generalization when variable selection is attempted on small training sets, as the apparent performance gains during training may fail to hold up on new data.
Organizations developing automated classification systems should adopt floating selection algorithms as a reliable standard for moderate-to-high dimensional problems. Teams should also ensure that sample sizes are sufficiently large relative to the number of measured variables before performing selection, or otherwise acquire additional ground-truth data prior to deploying critical classification systems.
The primary limitations of this study include its reliance on synthetic Gaussian distributions for mathematical validation and an empirical focus on a single satellite imagery application. While the comparative rankings among search methods are highly robust, performance trade-offs may vary depending on the specific classifier architecture and underlying data distributions encountered in other operational contexts.
- Paper: Irrelevant Features and the Subset Selection Problem, George H. John et al. (1994). Reading this foundational study on wrapper-based subset selection provides essential context for the feature evaluation strategies tested in the source paper.
- Paper: Feature Selection for High-Dimensional Data: A Fast Correlation-Based Filter Solution, Lei Yu et al. (2003). This paper extends the source's exploration of feature selection by introducing a fast correlation-based filter solution designed to scale efficiently to high-dimensional datasets.
