keyword
feature selection
Feature selection is the process in machine learning, statistics, and pattern recognition of identifying and selecting a subset of relevant, informative variables from a larger pool of data for use in model construction. Unlike feature extraction techniques that transform the original data into a new feature space, feature selection retains the original variables while systematically removing irrelevant, redundant, and noisy attributes. This form of dimensionality reduction aims to improve model generalization and predictive accuracy, reduce the risk of overfitting, lower computational and memory requirements, and enhance the interpretability of the learned models. The primary methodological strategies for feature selection include filter methods, which evaluate and score features using statistical or information-theoretic properties independently of a learning algorithm; wrapper methods, which evaluate candidate subsets using the predictive performance of a specific machine learning model; and embedded methods, which perform selection directly as an intrinsic part of model training.
26 items

A Rigorous Information-Theoretic Definition of Redundancy and Relevancy in Feature Selection Based on (Partial) Information Decomposition
Patricia Wollstadt, Sebastian Schmitt, Michael Wibral
Why you should read this
Establishes a rigorous definition of feature relevancy and redundancy using partial information decomposition and introduces an iterative conditional mutual information algorithm that isolates unique, redundant, and synergistic feature contributions for machine learning.
Selecting a minimal feature set that is maximally informative about a target variable is a central task in machine learning and statistics. Information theory provides a powerful framework for formulating feature selection algorithms—yet, a rigorous, information-theoretic definition of feature relevancy, which accounts for feature interactions such as redundant and synergistic contributions, is still missing. We argue that this lack is inherent to classical information theory which does not provide measures to decompose the information a set of variables provides about a target into unique, redundant, and synergistic contributions. Such a decomposition has been introduced only recently by the partial information decomposition (PID) framework. Using PID, we clarify why feature selection is a conceptually difficult problem when approached using information theory and provide a novel definition of feature relevancy and redundancy in PID terms. From this definition, we show that the conditional mutual information (CMI) maximizes relevancy while minimizing redundancy and propose an iterative, CMI-based algorithm for practical feature selection. We demonstrate the power of our CMI-based algorithm in comparison to the unconditional mutual information on benchmark examples and provide corresponding PID estimates to highlight how PID allows to quantify information contribution of features and their interactions in feature-selection problems.
Added
2026-10-05

Sparse Invariant Risk Minimization
Xiao Zhou, Yong Lin, Weizhong Zhang, Tong Zhang
Why you should read this
Proposes Sparse Invariant Risk Minimization, a method that enforces continuous global sparsity constraints during training to prevent overparameterized networks from learning spurious features under distribution shifts, outperforming existing techniques by up to 29% in accuracy.
Invariant Risk Minimization (IRM) is an emerging invariant feature extracting technique to help generalization with distributional shift. However, we find that there exists a basic and intractable contradiction between the model trainability and generalization ability in IRM. On one hand, recent studies on deep learning theory indicate the importance of large-sized or even overparameterized neural networks to make the model easy to train. On the other hand, unlike empirical risk minimization that can be benefited from overparameterization, our empirical and theoretical analyses show that the generalization ability of IRM is much easier to be demolished by overfitting caused by overparameterization. In this paper, we propose a simple yet effective paradigm named Sparse Invariant Risk Minimization (SparseIRM) to address this contradiction. Our key idea is to employ a global sparsity constraint as a defense to prevent spurious features from leaking in during the whole IRM process. Compared with sparsify-after-training prototype by prior work which can discard invariant features, the global sparsity constraint limits the budget for feature selection and enforces SparseIRM to select the invariant features. We illustrate the benefit of SparseIRM through a theoretical analysis on a simple linear case. Empirically we demonstrate the power of SparseIRM through various datasets and models and surpass state-of-the-art methods with a gap up to 29%.
Added
2026-10-02

Benchmarking Attribute Selection Techniques for Discrete Class Data Mining
Mark A. Hall, Geoffrey Holmes
Why you should read this
Evaluates six major attribute selection methods across multiple benchmark and high-dimensional datasets to determine how ranking-based feature selection impacts the classification performance and efficiency of C4.5 and naive Bayes.
Data engineering is generally considered to be a central issue in the development of data mining applications. The success of many learning schemes, in their attempts to construct models of data, hinges on the reliable identification of a small set of highly predictive attributes. The inclusion of irrelevant, redundant and noisy attributes in the model building process phase can result in poor predictive performance and increased computation. Attribute selection generally involves a combination of search and attribute utility estimation plus evaluation with respect to specific learning schemes. This leads to a large number of possible permutations and has led to a situation where very few benchmark studies have been conducted. This paper presents a benchmark comparison of several attribute selection methods for supervised classification. All the methods produce an attribute ranking, a useful devise for isolating the individual merit of an attribute. Attribute selection is achieved by cross-validating the attribute rankings with respect to a classification learner to find the best attributes. Results are reported for a selection of standard data sets and two diverse learning schemes C4.5 and naive Bayes.
Source
https://researchcommons.waikato.ac.nz/bitstreams/180e332c-f21c-4c5d-9a47-e13d48f71e70/downloadAdded
2026-09-25

Learning and Revising User Profiles: The Identification of Interesting Web Sites
MICHAEL PAZZANI, DANIEL BILLSUS
Why you should read this
Presents Syskill & Webert, an intelligent agent that uses naive Bayesian classification combined with user background knowledge and lexical feature selection to accurately predict and identify web pages matching a user's long-term interests.
We discuss algorithms for learning and revising user profiles that can determine which World Wide Web sites on a given topic would be interesting to a user. We describe the use of a naive Bayesian classifier for this task, and demonstrate that it can incrementally learn profiles from user feedback on the interestingness of Web sites. Furthermore, the Bayesian classifier may easily be extended to revise user provided profiles. In an experimental evaluation we compare the Bayesian classifier to computationally more intensive alternatives, and show that it performs at least as well as these approaches throughout a range of different domains. In addition, we empirically analyze the effects of providing the classifier with background knowledge in form of user defined profiles and examine the use of lexical knowledge for feature selection. We find that both approaches can substantially increase the prediction accuracy.
Added
2026-09-25

Inducing Features of Random Fields
Stephen Della Pietra, Vincent J. Della Pietra, J. Lafferty
Why you should read this
Introduces a principled framework for incrementally inducing complex features in random fields via maximum entropy and introduces the Improved Iterative Scaling algorithm to train non-Markovian exponential models with thousands of parameters.
We present a technique for constructing random fields from a set of training samples. The learning paradigm builds increasingly complex fields by allowing potential functions, or features, that are supported by increasingly large subgraphs. Each feature has a weight that is trained by minimizing the Kullback-Leibler divergence between the model and the empirical distribution of the training data. A greedy algorithm determines how features are incrementally added to the field and an iterative scaling algorithm is used to estimate the optimal values of the weights. The random field models and techniques introduced in this paper differ from those common to much of the computer vision literature in that the underlying random fields are non-Markovian and have a large number of parameters that must be estimated. Relations to other learning approaches, including decision trees, are given. As a demonstration of the method, we describe its application to the problem of automatic word classification in natural language processing.
Added
2026-09-25

A streaming ensemble algorithm (SEA) for large-scale classification
W. Street, YongSeog Kim
Why you should read this
Proposes a fast, constant-memory streaming ensemble algorithm that processes continuous data chunks and uses a targeted replacement heuristic to match batch classifier accuracy while rapidly adapting to concept drift.
Ensemble methods have recently garnered a great deal of attention in the machine learning community. Techniques such as Boosting and Bagging have proven to be highly effective but require repeated resampling of the training data, making them inappropriate in a data mining context. The methods presented in this paper take advantage of plentiful data, building separate classifiers on sequential chunks of training points. These classifiers are combined into a fixed-size ensemble using a heuristic replacement strategy. The result is a fast algorithm for large-scale or streaming data that classifies as well as a single decision tree built on all the data, requires approximately constant memory, and adjusts quickly to concept drift.
Added
2026-09-25

API design for machine learning software: experiences from the scikit-learn project
Lars Buitinck, Gilles Louppe, Mathieu Blondel, Fabian Pedregosa, Andreas Mueller, Olivier Grisel, Vlad Niculae, Peter Prettenhofer, Alexandre Gramfort, Jaques Grobler, Robert Layton, Jake Vanderplas, Arnaud Joly, Brian Holt, Gaël Varoquaux
Why you should read this
Presents the core API design principles behind scikit-learn, providing practical lessons on building consistent, composable, and user-friendly machine learning software.
Scikit-learn is an increasingly popular machine learning li- brary. Written in Python, it is designed to be simple and efficient, accessible to non-experts, and reusable in various contexts. In this paper, we present and discuss our design choices for the application programming interface (API) of the project. In particular, we describe the simple and elegant interface shared by all learning and processing units in the library and then discuss its advantages in terms of composition and reusability. The paper also comments on implementation details specific to the Python ecosystem and analyzes obstacles faced by users and developers of the library.
Added
2026-09-25

Heterogeneous Uncertainty Sampling for Supervised Learning
David D. Lewis, Jason Catlett
Why you should read this
Demonstrates that using a fast probabilistic classifier to actively select training instances for a more complex C4.5 rule induction model achieves lower error rates on text categorization tasks than random sampling sets ten times larger.
Uncertainty sampling methods iteratively request class labels for training instances whose classes are uncertain despite the previous labeled instances. These methods can greatly reduce the number of instances that an expert need label. One problem with this approach is that the classifier best suited for an application may be too expensive to train or use during the selection of instances. We test the use of one classifier (a highly efficient probabilistic one) to select examples for training another (the C4.5 rule induction program). Despite being chosen by this heterogeneous approach, the uncertainty samples yielded classifiers with lower error rates than random samples ten times larger.
Added
2026-09-25

Small Sample Size Effects in Statistical Pattern Recognition: Recommendations for Practitioners
S. Raudys, Anil K. Jain
Why you should read this
Presents practical guidelines and quantitative analyses to help practitioners choose appropriate training and test sample sizes, avoid small-sample bias in classifier design and feature selection, and accurately estimate classification error rates.
During the last two decades a considerable amount of effort has been devoted to the analysis of the influence of both training and testing sample size on the design and performance of pattern recognition systems. These questions are interesting to practitioners as well as theoreticians, because the small-sample effects can easily contaminate the design and evaluation of a proposed system. For applications with a large number of features and a complex classification rule, the training sample size must be quite large. A large test sample is required to accurately evaluate a classifier with a low error rate. The design of a pattern recognition system consists of several stages: data collection, formation of the pattern classes, feature selection, specification of the classification algorithm, and estimation of the classification error. In this paper, we will discuss the effects of sample size on feature selection and error estimation for several types of classifier. In addition to surveying prior work in this area, our emphasis is on giving practical advice to today's designers and users of statistical pattern recognition systems.
Added
2026-09-25

A Bayesian Approach to Filtering Junk E-Mail
M. Sahami, S. Dumais, D. Heckerman, E. Horvitz
Why you should read this
Presents a decision-theoretic Bayesian approach that combines text classification with domain-specific email features and asymmetric misclassification costs to filter spam accurately in real-world scenarios.
In addressing the growing problem of junk E-mail on the Internet, we examine methods for the automated construction of filters to eliminate such unwanted messages from a user's mail stream. By casting this problem in a decision theoretic framework, we are able to make use of probabilistic learning methods in conjunction with a notion of differential misclassification cost to produce filters which are especially appropriate for the nuances of this task. While this may appear, at first, to be a straight-forward text classification problem, we show that by considering domain-specific features of this problem in addition to the raw text of E-mail messages, we can produce much more accurate filters. Finally, we show the efficacy of such filters in a real world usage scenario, arguing that this technology is mature enough for deployment.
Added
2026-09-24

Auto-WEKA: combined selection and hyperparameter optimization of classification algorithms
Chris J. Thornton, F. Hutter, H. Hoos, Kevin Leyton-Brown
Why you should read this
Introduces a Bayesian optimization framework that automatically selects the best classification algorithm, feature selection method, and hyperparameter settings from the WEKA library to maximize predictive performance on a given dataset without manual intervention.
Many different machine learning algorithms exist; taking into account each algorithm's hyperparameters, there is a staggeringly large number of possible alternatives overall. We consider the problem of simultaneously selecting a learning algorithm and setting its hyperparameters, going beyond previous work that addresses these issues in isolation. We show that this problem can be addressed by a fully automated approach, leveraging recent innovations in Bayesian optimization. Specifically, we consider a wide range of feature selection techniques (combining 3 search and 8 evaluator methods) and all classification approaches implemented in WEKA, spanning 2 ensemble methods, 10 meta-methods, 27 base classifiers, and hyperparameter settings for each classifier. On each of 21 popular datasets from the UCI repository, the KDD Cup 09, variants of the MNIST dataset and CIFAR-10, we show classification performance often much better than using standard selection/hyperparameter optimization methods. We hope that our approach will help non-expert users to more effectively identify machine learning algorithms and hyperparameter settings appropriate to their applications, and hence to achieve improved performance.
Added
2026-09-24

Using Discriminant Eigenfeatures for Image Retrieval
D. Swets, J. Weng
Why you should read this
Proposes a Discriminant Karhunen-Loève projection framework that combines principal component analysis with linear discriminant analysis to eliminate non-informative variations like lighting and substantially improve content-based image retrieval across diverse object classes.
This paper describes the automatic selection of features from an image training set using the theories of multidimensional discriminant analysis and the associated optimal linear projection. We demonstrate the effectiveness of these Most Discriminating Features for view-based class retrieval from a large database of widely varying real-world objects presented as “well-framed” views, and compare it with that of the principal component analysis.
Added
2026-09-18

Rotation Forest: A New Classifier Ensemble Method
Juan J. Rodríguez, Ludmila I. Kuncheva, Carlos J. Alonso
Why you should read this
Introduces Rotation Forest, a classifier ensemble technique that applies Principal Component Analysis to random feature subsets to simultaneously boost individual decision tree accuracy and ensemble diversity, consistently outperforming Bagging, AdaBoost, and Random Forest across 33 benchmark datasets.
We propose a method for generating classifier ensembles based on feature extraction. To create the training data for a base classifier, the feature set is randomly split into K subsets (K is a parameter of the algorithm) and Principal Component Analysis (PCA) is applied to each subset. All principal components are retained in order to preserve the variability information in the data. Thus, K axis rotations take place to form the new features for a base classifier. The idea of the rotation approach is to encourage simultaneously individual accuracy and diversity within the ensemble. Diversity is promoted through the feature extraction for each base classifier. Decision trees were chosen here because they are sensitive to rotation of the feature axes, hence the name "forest." Accuracy is sought by keeping all principal components and also using the whole data set to train each base classifier. Using WEKA, we examined the Rotation Forest ensemble on a random selection of 33 benchmark data sets from the UCI repository and compared it with Bagging, AdaBoost, and Random Forest. The results were favorable to Rotation Forest and prompted an investigation into the diversity-accuracy landscape of the ensemble models. Diversity-error diagrams revealed that Rotation Forest ensembles construct individual classifiers which are more accurate than these in AdaBoost and Random Forest, and more diverse than these in Bagging, sometimes more accurate as well.
Added
2026-09-18

Feature selection, L1 vs. L2 regularization, and rotational invariance
Andrew Y. Ng
Why you should read this
Proves that L1-regularized logistic regression requires only logarithmically many training examples in the number of irrelevant features, whereas rotationally invariant methods like L2 regularization and SVMs suffer from sample complexity that scales at least linearly.
This document is a presentation slide deck and does not contain an abstract.
Added
2026-09-16

Orange: data mining toolbox in python
Janez Demšar, Tomaž Curk, Aleš Erjavec, Črt Gorup, Tomaž Hočevar, Mitar Milutinovič, Martin Možina, Matija Polajnar, Marko Toplak, Anže Starič, Miha Štajdohar, Lan Umek, Lan Žagar, Jure Žbontar, Marinka Žitnik, Blaž Zupan
Why you should read this
Presents Orange, an open-source Python data mining library that enables rapid prototyping and interactive data analysis by combining high-level scriptable components with fast C++ implementations for core machine learning tasks.
Orange is a machine learning and data mining suite for data analysis through Python scripting and visual programming. Here we report on the scripting part, which features interactive data analysis and component-based assembly of data mining procedures. In the selection and design of components, we focus on the flexibility of their reuse: our principal intention is to let the user write simple and clear scripts in Python, which build upon C++ implementations of computationally-intensive tasks. Orange is intended both for experienced users and programmers, as well as for students of data mining.
Added
2026-09-16

Machine learning for neuroimaging with scikit-learn
Alexandre Abraham, Fabian Pedregosa, Michael Eickenberg, Philippe Gervais, Andreas Muller, Jean Kossaifi, Alexandre Gramfort, Bertrand Thirion, Gäel Varoquaux
Why you should read this
Demonstrates how to apply scikit-learn to functional neuroimaging datasets, providing practical implementations for high-dimensional brain decoding, encoding, and resting-state fMRI analysis.
Statistical machine learning methods are increasingly used for neuroimaging data analysis. Their main virtue is their ability to model high-dimensional datasets, e.g. multivariate analysis of activation images or resting-state time series. Supervised learning is typically used in decoding or encoding settings to relate brain images to behavioral or clinical observations, while unsupervised learning can uncover hidden structures in sets of images (e.g. resting state functional MRI) or find sub-populations in large cohorts. By considering different functional neuroimaging applications, we illustrate how scikit-learn, a Python machine learning library, can be used to perform some key analysis steps. Scikit-learn contains a very large set of statistical learning algorithms, both supervised and unsupervised, and its application to neuroimaging data provides a versatile tool to study the brain.
Added
2026-09-16

The Feature Selection Problem: Traditional Methods and a New Algorithm
Kenji Kira, Larry Rendell
Why you should read this
Introduces the Relief algorithm, a noise-tolerant feature selection method that identifies relevant attributes in linear time without relying on heuristics, even when strong feature interactions are present.
For real-world concept learning problems, feature selection is important to speed up learning and to improve concept quality. We review and analyze past approaches to feature selection and note their strengths and weaknesses. We then introduce and theoretically examine a new algorithm Relief which selects relevant features using a statistical method. Relief does not depend on heuristics, is accurate even if features interact, and is noise-tolerant. It requires only linear time in the number of given features and the number of training instances, regardless of the target concept complexity. The algorithm also has certain limitations such as non-optimal feature set size. Ways to overcome the limitations are suggested. We also report the test results of comparison between Relief and other feature selection algorithms. The empirical results support the theoretical analysis, suggesting a practical approach to feature selection for real-world problems.
Added
2026-09-16

Laplacian Score for Feature Selection
Xiaofei He, Deng Cai, P. Niyogi
Why you should read this
Introduces Laplacian Score, an efficient filter-based feature selection algorithm that evaluates features based on their ability to preserve local geometric structures modeled via nearest-neighbor graphs in both supervised and unsupervised settings.
In supervised learning scenarios, feature selection has been studied widely in the literature. Selecting features in unsupervised learning scenarios is a much harder problem, due to the absence of class labels that would guide the search for relevant information. And, almost all of previous unsupervised feature selection methods are “wrapper” techniques that require a learning algorithm to evaluate the candidate feature subsets. In this paper, we propose a “filter” method for feature selection which is independent of any learning algorithm. Our method can be performed in either supervised or unsupervised fashion. The proposed method is based on the observation that, in many real world classification problems, data from the same class are often close to each other. The importance of a feature is evaluated by its power of locality preserving, or, Laplacian Score. We compare our method with data variance (unsupervised) and Fisher score (supervised) on two data sets. Experimental results demonstrate the effectiveness and efficiency of our algorithm.
Added
2026-09-16

Feature Selection: Evaluation, Application, and Small Sample Performance
Anil K. Jain, Douglas E. Zongker
Why you should read this
Compares prominent feature subset selection methods on synthetic benchmarks and SAR satellite imagery, establishing the superior performance of sequential forward floating selection while identifying critical pitfalls of selection algorithms in small-sample scenarios.
A large number of algorithms have been proposed for feature subset selection. Our experimental results show that the sequential forward floating selection (SFFS) algorithm, proposed by Pudil et al., dominates the other algorithms tested. We study the problem of choosing an optimal feature set for land use classification based on SAR satellite images using four different texture models. Pooling features derived from different texture models, followed by a feature selection results in a substantial improvement in the classification accuracy. We also illustrate the dangers of using feature selection in small sample size situations.
Added
2026-09-14

Choosing Multiple Parameters for Support Vector Machines
OLIVIER CHAPELLE, VLADIMIR VAPNIK, OLIVIER BOUSQUET, SAYAN MUKHERJEE
Why you should read this
Proposes a gradient descent framework to automatically tune large numbers of support vector machine hyperparameters and kernel scaling factors by minimizing generalization error bounds, enabling simultaneous kernel optimization and feature selection.
The problem of automatically tuning multiple parameters for pattern recognition Support Vector Machines (SVMs) is considered. This is done by minimizing some estimates of the generalization error of SVMs using a gradient descent algorithm over the set of parameters. Usual methods for choosing parameters, based on exhaustive search become intractable as soon as the number of parameters exceeds two. Some experimental results assess the feasibility of our approach for a large number of parameters (more than 100) and demonstrate an improvement of generalization performance.
Added
2026-09-14
