Built independently by an author, for readers. Read the story and support ChapterPal

keyword

induction algorithms

An induction algorithm is a computational procedure in machine learning that generalizes from a set of specific training examples to produce a predictive model or decision rule capable of classifying unseen data. Working primarily within supervised learning frameworks, these algorithms analyze labeled instances characterized by descriptive features and infer an underlying mapping function, hypothesis, or concept description. Common examples include decision tree generators, rule learners, and probabilistic classifiers like naive Bayes. The resulting model captures statistical regularities and patterns present in the training data, allowing the system to perform automated predictions, concept learning, and feature evaluation across new observations.

5 items

A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection

A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection

Ron Kohavi

OrganizationsStanford University

Why you should read this

Demonstrates through over half a million experimental runs that ten-fold stratified cross-validation is the most effective approach for model selection and accuracy estimation on real-world datasets, outperforming both bootstrap and leave-one-out methods.

We review accuracy estimation methods and compare the two most common methods: cross-validation and bootstrap. Recent experimental results on artificial data and theoretical results in restricted settings have shown that for selecting a good classifier from a set of classifiers (model selection), ten-fold cross-validation may be better than the more expensive leave-one-out cross-validation. We report on a large-scale experiment—over half a million runs of C4.5 and a Naive-Bayes algorithm—to estimate the effects of different parameters on these algorithms on real-world datasets. For cross-validation, we vary the number of folds and whether the folds are stratified or not; for bootstrap, we vary the number of bootstrap samples. Our results indicate that for real-word datasets similar to ours, the best method to use for model selection is ten-fold stratified cross validation, even if computation power allows using more folds.

Added

2026-09-06