Built independently by an author, for readers. Read the story and support ChapterPal

keyword

feature selection

Feature selection is the process in machine learning, statistics, and pattern recognition of identifying and selecting a subset of relevant, informative variables from a larger pool of data for use in model construction. Unlike feature extraction techniques that transform the original data into a new feature space, feature selection retains the original variables while systematically removing irrelevant, redundant, and noisy attributes. This form of dimensionality reduction aims to improve model generalization and predictive accuracy, reduce the risk of overfitting, lower computational and memory requirements, and enhance the interpretability of the learned models. The primary methodological strategies for feature selection include filter methods, which evaluate and score features using statistical or information-theoretic properties independently of a learning algorithm; wrapper methods, which evaluate candidate subsets using the predictive performance of a specific machine learning model; and embedded methods, which perform selection directly as an intrinsic part of model training.

26 items

A Rigorous Information-Theoretic Definition of Redundancy and Relevancy in Feature Selection Based on (Partial) Information Decomposition

A Rigorous Information-Theoretic Definition of Redundancy and Relevancy in Feature Selection Based on (Partial) Information Decomposition

Patricia Wollstadt, Sebastian Schmitt, Michael Wibral

OrganizationsCampus Institute for Dynamics of Biological NetworksHonda Research InstituteUniversity of Göttingen

Why you should read this

Establishes a rigorous definition of feature relevancy and redundancy using partial information decomposition and introduces an iterative conditional mutual information algorithm that isolates unique, redundant, and synergistic feature contributions for machine learning.

Selecting a minimal feature set that is maximally informative about a target variable is a central task in machine learning and statistics. Information theory provides a powerful framework for formulating feature selection algorithms—yet, a rigorous, information-theoretic definition of feature relevancy, which accounts for feature interactions such as redundant and synergistic contributions, is still missing. We argue that this lack is inherent to classical information theory which does not provide measures to decompose the information a set of variables provides about a target into unique, redundant, and synergistic contributions. Such a decomposition has been introduced only recently by the partial information decomposition (PID) framework. Using PID, we clarify why feature selection is a conceptually difficult problem when approached using information theory and provide a novel definition of feature relevancy and redundancy in PID terms. From this definition, we show that the conditional mutual information (CMI) maximizes relevancy while minimizing redundancy and propose an iterative, CMI-based algorithm for practical feature selection. We demonstrate the power of our CMI-based algorithm in comparison to the unconditional mutual information on benchmark examples and provide corresponding PID estimates to highlight how PID allows to quantify information contribution of features and their interactions in feature-selection problems.

Added

2026-10-05

Sparse Invariant Risk Minimization

Sparse Invariant Risk Minimization

Xiao Zhou, Yong Lin, Weizhong Zhang, Tong Zhang

OrganizationsGoogleThe Hong Kong University of Science and Technology

Why you should read this

Proposes Sparse Invariant Risk Minimization, a method that enforces continuous global sparsity constraints during training to prevent overparameterized networks from learning spurious features under distribution shifts, outperforming existing techniques by up to 29% in accuracy.

Invariant Risk Minimization (IRM) is an emerging invariant feature extracting technique to help generalization with distributional shift. However, we find that there exists a basic and intractable contradiction between the model trainability and generalization ability in IRM. On one hand, recent studies on deep learning theory indicate the importance of large-sized or even overparameterized neural networks to make the model easy to train. On the other hand, unlike empirical risk minimization that can be benefited from overparameterization, our empirical and theoretical analyses show that the generalization ability of IRM is much easier to be demolished by overfitting caused by overparameterization. In this paper, we propose a simple yet effective paradigm named Sparse Invariant Risk Minimization (SparseIRM) to address this contradiction. Our key idea is to employ a global sparsity constraint as a defense to prevent spurious features from leaking in during the whole IRM process. Compared with sparsify-after-training prototype by prior work which can discard invariant features, the global sparsity constraint limits the budget for feature selection and enforces SparseIRM to select the invariant features. We illustrate the benefit of SparseIRM through a theoretical analysis on a simple linear case. Empirically we demonstrate the power of SparseIRM through various datasets and models and surpass state-of-the-art methods with a gap up to 29%.

Added

2026-10-02

Benchmarking Attribute Selection Techniques for Discrete Class Data Mining

Benchmarking Attribute Selection Techniques for Discrete Class Data Mining

Mark A. Hall, Geoffrey Holmes

Why you should read this

Evaluates six major attribute selection methods across multiple benchmark and high-dimensional datasets to determine how ranking-based feature selection impacts the classification performance and efficiency of C4.5 and naive Bayes.

Data engineering is generally considered to be a central issue in the development of data mining applications. The success of many learning schemes, in their attempts to construct models of data, hinges on the reliable identification of a small set of highly predictive attributes. The inclusion of irrelevant, redundant and noisy attributes in the model building process phase can result in poor predictive performance and increased computation. Attribute selection generally involves a combination of search and attribute utility estimation plus evaluation with respect to specific learning schemes. This leads to a large number of possible permutations and has led to a situation where very few benchmark studies have been conducted. This paper presents a benchmark comparison of several attribute selection methods for supervised classification. All the methods produce an attribute ranking, a useful devise for isolating the individual merit of an attribute. Attribute selection is achieved by cross-validating the attribute rankings with respect to a classification learner to find the best attributes. Results are reported for a selection of standard data sets and two diverse learning schemes C4.5 and naive Bayes.

Added

2026-09-25

Inducing Features of Random Fields

Inducing Features of Random Fields

Stephen Della Pietra, Vincent J. Della Pietra, J. Lafferty

OrganizationsCarnegie Mellon UniversityRenaissance Technologies

Why you should read this

Introduces a principled framework for incrementally inducing complex features in random fields via maximum entropy and introduces the Improved Iterative Scaling algorithm to train non-Markovian exponential models with thousands of parameters.

We present a technique for constructing random fields from a set of training samples. The learning paradigm builds increasingly complex fields by allowing potential functions, or features, that are supported by increasingly large subgraphs. Each feature has a weight that is trained by minimizing the Kullback-Leibler divergence between the model and the empirical distribution of the training data. A greedy algorithm determines how features are incrementally added to the field and an iterative scaling algorithm is used to estimate the optimal values of the weights. The random field models and techniques introduced in this paper differ from those common to much of the computer vision literature in that the underlying random fields are non-Markovian and have a large number of parameters that must be estimated. Relations to other learning approaches, including decision trees, are given. As a demonstration of the method, we describe its application to the problem of automatic word classification in natural language processing.

Added

2026-09-25

Small Sample Size Effects in Statistical Pattern Recognition: Recommendations for Practitioners

Small Sample Size Effects in Statistical Pattern Recognition: Recommendations for Practitioners

S. Raudys, Anil K. Jain

OrganizationsInstitute of Mathematics and Cybernetics, Lithuanian Academy of SciencesMichigan State University

Why you should read this

Presents practical guidelines and quantitative analyses to help practitioners choose appropriate training and test sample sizes, avoid small-sample bias in classifier design and feature selection, and accurately estimate classification error rates.

During the last two decades a considerable amount of effort has been devoted to the analysis of the influence of both training and testing sample size on the design and performance of pattern recognition systems. These questions are interesting to practitioners as well as theoreticians, because the small-sample effects can easily contaminate the design and evaluation of a proposed system. For applications with a large number of features and a complex classification rule, the training sample size must be quite large. A large test sample is required to accurately evaluate a classifier with a low error rate. The design of a pattern recognition system consists of several stages: data collection, formation of the pattern classes, feature selection, specification of the classification algorithm, and estimation of the classification error. In this paper, we will discuss the effects of sample size on feature selection and error estimation for several types of classifier. In addition to surveying prior work in this area, our emphasis is on giving practical advice to today's designers and users of statistical pattern recognition systems.

Added

2026-09-25

Auto-WEKA: combined selection and hyperparameter optimization of classification algorithms

Auto-WEKA: combined selection and hyperparameter optimization of classification algorithms

Chris J. Thornton, F. Hutter, H. Hoos, Kevin Leyton-Brown

OrganizationsUniversity of British Columbia

Why you should read this

Introduces a Bayesian optimization framework that automatically selects the best classification algorithm, feature selection method, and hyperparameter settings from the WEKA library to maximize predictive performance on a given dataset without manual intervention.

Many different machine learning algorithms exist; taking into account each algorithm's hyperparameters, there is a staggeringly large number of possible alternatives overall. We consider the problem of simultaneously selecting a learning algorithm and setting its hyperparameters, going beyond previous work that addresses these issues in isolation. We show that this problem can be addressed by a fully automated approach, leveraging recent innovations in Bayesian optimization. Specifically, we consider a wide range of feature selection techniques (combining 3 search and 8 evaluator methods) and all classification approaches implemented in WEKA, spanning 2 ensemble methods, 10 meta-methods, 27 base classifiers, and hyperparameter settings for each classifier. On each of 21 popular datasets from the UCI repository, the KDD Cup 09, variants of the MNIST dataset and CIFAR-10, we show classification performance often much better than using standard selection/hyperparameter optimization methods. We hope that our approach will help non-expert users to more effectively identify machine learning algorithms and hyperparameter settings appropriate to their applications, and hence to achieve improved performance.

Added

2026-09-24

Rotation Forest: A New Classifier Ensemble Method

Rotation Forest: A New Classifier Ensemble Method

Juan J. Rodríguez, Ludmila I. Kuncheva, Carlos J. Alonso

OrganizationsBangor UniversityUniversidad de BurgosUniversidad de Valladolid

Why you should read this

Introduces Rotation Forest, a classifier ensemble technique that applies Principal Component Analysis to random feature subsets to simultaneously boost individual decision tree accuracy and ensemble diversity, consistently outperforming Bagging, AdaBoost, and Random Forest across 33 benchmark datasets.

We propose a method for generating classifier ensembles based on feature extraction. To create the training data for a base classifier, the feature set is randomly split into K subsets (K is a parameter of the algorithm) and Principal Component Analysis (PCA) is applied to each subset. All principal components are retained in order to preserve the variability information in the data. Thus, K axis rotations take place to form the new features for a base classifier. The idea of the rotation approach is to encourage simultaneously individual accuracy and diversity within the ensemble. Diversity is promoted through the feature extraction for each base classifier. Decision trees were chosen here because they are sensitive to rotation of the feature axes, hence the name "forest." Accuracy is sought by keeping all principal components and also using the whole data set to train each base classifier. Using WEKA, we examined the Rotation Forest ensemble on a random selection of 33 benchmark data sets from the UCI repository and compared it with Bagging, AdaBoost, and Random Forest. The results were favorable to Rotation Forest and prompted an investigation into the diversity-accuracy landscape of the ensemble models. Diversity-error diagrams revealed that Rotation Forest ensembles construct individual classifiers which are more accurate than these in AdaBoost and Random Forest, and more diverse than these in Bagging, sometimes more accurate as well.

Added

2026-09-18

Machine learning for neuroimaging with scikit-learn

Machine learning for neuroimaging with scikit-learn

Alexandre Abraham, Fabian Pedregosa, Michael Eickenberg, Philippe Gervais, Andreas Muller, Jean Kossaifi, Alexandre Gramfort, Bertrand Thirion, Gäel Varoquaux

OrganizationsCEAImperial College LondonINRIATélécom ParisUniversity of Bonn

Why you should read this

Demonstrates how to apply scikit-learn to functional neuroimaging datasets, providing practical implementations for high-dimensional brain decoding, encoding, and resting-state fMRI analysis.

Statistical machine learning methods are increasingly used for neuroimaging data analysis. Their main virtue is their ability to model high-dimensional datasets, e.g. multivariate analysis of activation images or resting-state time series. Supervised learning is typically used in decoding or encoding settings to relate brain images to behavioral or clinical observations, while unsupervised learning can uncover hidden structures in sets of images (e.g. resting state functional MRI) or find sub-populations in large cohorts. By considering different functional neuroimaging applications, we illustrate how scikit-learn, a Python machine learning library, can be used to perform some key analysis steps. Scikit-learn contains a very large set of statistical learning algorithms, both supervised and unsupervised, and its application to neuroimaging data provides a versatile tool to study the brain.

Added

2026-09-16

Laplacian Score for Feature Selection

Laplacian Score for Feature Selection

Xiaofei He, Deng Cai, P. Niyogi

OrganizationsUniversity of ChicagoUniversity of Illinois Urbana-Champaign

Why you should read this

Introduces Laplacian Score, an efficient filter-based feature selection algorithm that evaluates features based on their ability to preserve local geometric structures modeled via nearest-neighbor graphs in both supervised and unsupervised settings.

In supervised learning scenarios, feature selection has been studied widely in the literature. Selecting features in unsupervised learning scenarios is a much harder problem, due to the absence of class labels that would guide the search for relevant information. And, almost all of previous unsupervised feature selection methods are “wrapper” techniques that require a learning algorithm to evaluate the candidate feature subsets. In this paper, we propose a “filter” method for feature selection which is independent of any learning algorithm. Our method can be performed in either supervised or unsupervised fashion. The proposed method is based on the observation that, in many real world classification problems, data from the same class are often close to each other. The importance of a feature is evaluated by its power of locality preserving, or, Laplacian Score. We compare our method with data variance (unsupervised) and Fisher score (supervised) on two data sets. Experimental results demonstrate the effectiveness and efficiency of our algorithm.

Added

2026-09-16