keyword
induction algorithms
An induction algorithm is a computational procedure in machine learning that generalizes from a set of specific training examples to produce a predictive model or decision rule capable of classifying unseen data. Working primarily within supervised learning frameworks, these algorithms analyze labeled instances characterized by descriptive features and infer an underlying mapping function, hypothesis, or concept description. Common examples include decision tree generators, rule learners, and probabilistic classifiers like naive Bayes. The resulting model captures statistical regularities and patterns present in the training data, allowing the system to perform automated predictions, concept learning, and feature evaluation across new observations.
5 items

An Analysis of Bayesian Classifiers
Pat Langley, Wayne Iba, Kevin Thompson
Why you should read this
Presents an average-case theoretical analysis explaining why simple Bayesian classifiers achieve high accuracy across diverse learning domains despite their strong attribute independence assumptions.
In this paper we present an average-case analysis of the Bayesian classifier, a simple induction algorithm that fares remarkably well on many learning tasks. Our analysis assumes a monotone conjunctive target concept, and independent, noise-free Boolean attributes. We calculate the probability that the algorithm will induce an arbitrary pair of concept descriptions and then use this to compute the probability of correct classification over the instance space. The analysis takes into account the number of training instances, the number of attributes, the distribution of these attributes, and the level of class noise. We also explore the behavioral implications of the analysis by presenting predicted learning curves for artificial domains, and give experimental results on these domains as a check on our reasoning.
Added
2026-09-25

Toward Optimal Feature Selection
Daphne Koller, Mehran Sahami
Why you should read this
Develops an information-theoretic filter algorithm that efficiently eliminates both irrelevant and redundant features, providing a theoretically grounded solution for high-dimensional classification tasks without the computational burden of wrapper methods.
In this paper, we examine a method for feature subset selection based on Information Theory. Initially, a framework for defining the theoretically optimal, but computationally intractable, method for feature subset selection is presented. We show that our goal should be to eliminate a feature if it gives us little or no additional information beyond that subsumed by the remaining features. In particular, this will be the case for both irrelevant and redundant features. We then give an efficient algorithm for feature selection which computes an approximation to the optimal feature selection criterion. The conditions under which the approximate algorithm is successful are examined. Empirical results are given on a number of data sets, showing that the algorithm effectively handles datasets with a very large number of features.
Added
2026-09-18

Scaling Up the Accuracy of Naive-Bayes Classifiers: A Decision-Tree Hybrid
Ron Kohavi
Why you should read this
Introduces NBTree, a hybrid algorithm that embeds Naive-Bayes classifiers at the leaves of decision trees to significantly boost classification accuracy on large datasets while retaining model interpretability.
Naive-Bayes induction algorithms were previously shown to be surprisingly accurate on many classification tasks even when the conditional independence assumption on which they are based is violated. However, most studies were done on small databases. We show that in some larger databases, the accuracy of Naive-Bayes does not scale up as well as decision trees. We then propose a new algorithm, NBTree, which induces a hybrid of decision-tree classifiers and Naive-Bayes classifiers: the decision-tree nodes contain univariate splits as regular decision-trees, but the leaves contain Naive-Bayesian classifiers. The approach retains the interpretability of Naive-Bayes and decision trees, while resulting in classifiers that frequently outperform both constituents, especially in the larger databases tested.
Added
2026-09-17

Irrelevant Features and the Subset Selection Problem
George H. John, Ron Kohavi, Karl Pfleger
Why you should read this
Formalizes definitions of strong and weak feature relevance and introduces the wrapper framework for feature subset selection, demonstrating how evaluating candidate subsets directly with cross-validated induction algorithms eliminates deceptive attributes and produces more compact, accurate models.
We address the problem of finding a subset of features that allows a supervised induction algorithm to induce small high-accuracy concepts. We examine notions of relevance and irrelevance, and show that the definitions used in the machine learning literature do not adequately partition the features into useful categories of relevance. We present definitions for irrelevance and for two degrees of relevance. These definitions improve our understanding of the behavior of previous subset selection algorithms, and help define the subset of features that should be sought. The features selected should depend not only on the features and the target concept, but also on the induction algorithm. We describe a method for feature subset selection using cross-validation that is applicable to any induction algorithm, and discuss experiments conducted with ID3 and C4.5 on artificial and real datasets.
Added
2026-09-14

A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection
Ron Kohavi
Why you should read this
Demonstrates through over half a million experimental runs that ten-fold stratified cross-validation is the most effective approach for model selection and accuracy estimation on real-world datasets, outperforming both bootstrap and leave-one-out methods.
We review accuracy estimation methods and compare the two most common methods: cross-validation and bootstrap. Recent experimental results on artificial data and theoretical results in restricted settings have shown that for selecting a good classifier from a set of classifiers (model selection), ten-fold cross-validation may be better than the more expensive leave-one-out cross-validation. We report on a large-scale experiment—over half a million runs of C4.5 and a Naive-Bayes algorithm—to estimate the effects of different parameters on these algorithms on real-world datasets. For cross-validation, we vary the number of folds and whether the folds are stratified or not; for bootstrap, we vary the number of bootstrap samples. Our results indicate that for real-word datasets similar to ours, the best method to use for model selection is ten-fold stratified cross validation, even if computation power allows using more folds.
Added
2026-09-06
