Built independently by an author, for readers. Read the story and support ChapterPal

keyword

information gain

Information gain is a metric in information theory and machine learning that measures the reduction in uncertainty or entropy achieved by learning the value of an additional feature or piece of data. Equivalent to mutual information, it quantifies how much knowledge about a target variable increases after observing another variable. In supervised learning algorithms, particularly during decision tree induction, information gain serves as a standard criterion to evaluate and select the best attribute for splitting a dataset. It is calculated by subtracting the weighted average entropy of the resulting subsets from the total entropy of the original data, where higher values indicate features that provide stronger predictive power and better separate distinct classes.

7 items

Understanding In-Context Learning via Supportive Pretraining Data

Understanding In-Context Learning via Supportive Pretraining Data

Xiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov, Asli Celikyilmaz, Tianlu Wang

OrganizationsMetaUniversity of Washington

Why you should read this

Reveals that in-context learning in large language models is driven by specific, challenging pretraining instances rich in long-tail tokens and difficult long-range contexts rather than domain-relevant text, providing actionable criteria to guide future pretraining data selection.

In-context learning (ICL) improves language models’ performance on a variety of NLP tasks by simply demonstrating a handful of examples at inference time. It is not well understood why ICL ability emerges, as the model has never been specifically trained on such demonstrations. Unlike prior work that explores implicit mechanisms behind ICL, we study ICL via investigating the pretraining data. Specifically, we first adapt an iterative, gradient-based approach to find a small subset of pretraining data that supports ICL. We observe that a continued pretraining on this small subset significantly improves the model’s ICL ability, by up to 18%. We then compare the supportive subset contrastively with random subsets of pretraining data and discover: (1) The supportive pretraining data to ICL do not have a higher domain relevance to downstream tasks. (2) The supportive pretraining data have a higher mass of rarely occurring, long-tail tokens. (3) The supportive pretraining data are challenging examples where the information gain from long-range context is below average, indicating learning to incorporate difficult long-range context encourages ICL. Our work takes a first step towards understanding ICL via analyzing instance-level pretraining data. Our insights have a potential to enhance the ICL ability of language models by actively guiding the construction of pretraining data in the future.

Added

2026-10-03

An Experimental Comparison of Three Methods for Constructing Ensembles of Decision Trees: Bagging, Boosting, and Randomization

An Experimental Comparison of Three Methods for Constructing Ensembles of Decision Trees: Bagging, Boosting, and Randomization

Thomas G. Dietterich

OrganizationsOregon State University

Why you should read this

Demonstrates across 33 benchmark datasets that while boosting generates the most accurate decision tree ensembles on clean data, bagging remains superior under classification noise because boosting assigns excessive weight to mislabeled training instances.

Bagging and boosting are methods that generate a diverse ensemble of classifiers by manipulating the training data given to a "base" learning algorithm. Breiman has pointed out that they rely for their effectiveness on the instability of the base learning algorithm. An alternative approach to generating an ensemble is to randomize the internal decisions made by the base algorithm. This general approach has been studied previously by Ali and Pazzani and by Dietterich and Kong. This paper compares the effectiveness of randomization, bagging, and boosting for improving the performance of the decision-tree algorithm C4.5. The experiments show that in situations with little or no classification noise, randomization is competitive with (and perhaps slightly superior to) bagging but not as accurate as boosting. In situations with substantial classification noise, bagging is much better than boosting, and sometimes better than randomization.

Added

2026-09-12

Induction of Decision Trees

Induction of Decision Trees

J. R. Quinlan

OrganizationsUniversity of Technology Sydney

Why you should read this

Introduces the ID3 algorithm for efficiently constructing interpretable decision trees from data using information gain principles.

The technology for building knowledge-based systems by inductive inference from examples has been demonstrated successfully in several practical applications. This paper summarizes an approach to synthesizing decision trees that has been used in a variety of systems, and it describes one such system, ID3, in detail. Results from recent studies show ways in which the methodology can be modified to deal with information that is noisy and/or incomplete. A reported shortcoming of the basic algorithm is discussed and two means of overcoming it are compared. The paper concludes with illustrations of current research directions.

Added

2026-05-14