keyword
information gain
Information gain is a metric in information theory and machine learning that measures the reduction in uncertainty or entropy achieved by learning the value of an additional feature or piece of data. Equivalent to mutual information, it quantifies how much knowledge about a target variable increases after observing another variable. In supervised learning algorithms, particularly during decision tree induction, information gain serves as a standard criterion to evaluate and select the best attribute for splitting a dataset. It is calculated by subtracting the weighted average entropy of the resulting subsets from the total entropy of the original data, where higher values indicate features that provide stronger predictive power and better separate distinct classes.
7 items

Understanding In-Context Learning via Supportive Pretraining Data
Xiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov, Asli Celikyilmaz, Tianlu Wang
Why you should read this
Reveals that in-context learning in large language models is driven by specific, challenging pretraining instances rich in long-tail tokens and difficult long-range contexts rather than domain-relevant text, providing actionable criteria to guide future pretraining data selection.
In-context learning (ICL) improves language models’ performance on a variety of NLP tasks by simply demonstrating a handful of examples at inference time. It is not well understood why ICL ability emerges, as the model has never been specifically trained on such demonstrations. Unlike prior work that explores implicit mechanisms behind ICL, we study ICL via investigating the pretraining data. Specifically, we first adapt an iterative, gradient-based approach to find a small subset of pretraining data that supports ICL. We observe that a continued pretraining on this small subset significantly improves the model’s ICL ability, by up to 18%. We then compare the supportive subset contrastively with random subsets of pretraining data and discover: (1) The supportive pretraining data to ICL do not have a higher domain relevance to downstream tasks. (2) The supportive pretraining data have a higher mass of rarely occurring, long-tail tokens. (3) The supportive pretraining data are challenging examples where the information gain from long-range context is below average, indicating learning to incorporate difficult long-range context encourages ICL. Our work takes a first step towards understanding ICL via analyzing instance-level pretraining data. Our insights have a potential to enhance the ICL ability of language models by actively guiding the construction of pretraining data in the future.
Added
2026-10-03

Learning and Revising User Profiles: The Identification of Interesting Web Sites
MICHAEL PAZZANI, DANIEL BILLSUS
Why you should read this
Presents Syskill & Webert, an intelligent agent that uses naive Bayesian classification combined with user background knowledge and lexical feature selection to accurately predict and identify web pages matching a user's long-term interests.
We discuss algorithms for learning and revising user profiles that can determine which World Wide Web sites on a given topic would be interesting to a user. We describe the use of a naive Bayesian classifier for this task, and demonstrate that it can incrementally learn profiles from user feedback on the interestingness of Web sites. Furthermore, the Bayesian classifier may easily be extended to revise user provided profiles. In an experimental evaluation we compare the Bayesian classifier to computationally more intensive alternatives, and show that it performs at least as well as these approaches throughout a range of different domains. In addition, we empirically analyze the effects of providing the classifier with background knowledge in form of user defined profiles and examine the use of lexical knowledge for feature selection. We find that both approaches can substantially increase the prediction accuracy.
Added
2026-09-25

Unifying Count-Based Exploration and Intrinsic Motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, Remi Munos
Why you should read this
Develops a pseudo-count mechanism derived from density models that generalizes count-based exploration to high-dimensional reinforcement learning, enabling agents to effectively tackle sparse-reward benchmarks like Montezuma's Revenge directly from raw pixels.
We consider an agent's uncertainty about its environment and the problem of generalizing this uncertainty across observations. Specifically, we focus on the problem of exploration in non-tabular reinforcement learning. Drawing inspiration from the intrinsic motivation literature, we use density models to measure uncertainty, and propose a novel algorithm for deriving a pseudo-count from an arbitrary density model. This technique enables us to generalize count-based exploration algorithms to the non-tabular case. We apply our ideas to Atari 2600 games, providing sensible pseudo-counts from raw pixels. We transform these pseudo-counts into intrinsic rewards and obtain significantly improved exploration in a number of hard games, including the infamously difficult Montezuma's Revenge.
Added
2026-09-24

Improved Use of Continuous Attributes in C4.5
J. R. Quinlan
A reported weakness of C4.5 in domains with continuous attributes is addressed by modifying the formation and evaluation of tests on continuous attributes. An MDL-inspired penalty is applied to such tests, eliminating some of them from consideration and altering the relative desirability of all tests. Empirical trials show that the modifications lead to smaller decision trees with higher predictive accuracy. Results also confirm that a new version of C4.5 incorporating these changes is superior to recent approaches that use global discretization and that construct small trees with multi-interval splits.
Added
2026-09-16

Query by committee
H. Seung, Manfred Opper, Haim Sompolinsky
Why you should read this
Introduces the Query by Committee active learning framework, demonstrating that selecting unlabeled examples based on maximal disagreement among an ensemble of models achieves an exponential reduction in generalization error compared to passive random sampling.
We propose an algorithm called query by committee, in which a committee of students is trained on the same data set. The next query is chosen according to the principle of maximal disagreement. The algorithm is studied for two toy models: the high-low game and perceptron learning of another perceptron. As the number of queries goes to infinity, the committee algorithm yields asymptotically finite information gain. This leads to generalization error that decreases exponentially with the number of examples. This in marked contrast to learning from randomly chosen inputs, for which the information gain approaches zero and the generalization error decreases with a relatively slow inverse power law. We suggest that asymptotically finite information gain may be an important characteristic of good query algorithms.
Added
2026-09-16

An Experimental Comparison of Three Methods for Constructing Ensembles of Decision Trees: Bagging, Boosting, and Randomization
Thomas G. Dietterich
Why you should read this
Demonstrates across 33 benchmark datasets that while boosting generates the most accurate decision tree ensembles on clean data, bagging remains superior under classification noise because boosting assigns excessive weight to mislabeled training instances.
Bagging and boosting are methods that generate a diverse ensemble of classifiers by manipulating the training data given to a "base" learning algorithm. Breiman has pointed out that they rely for their effectiveness on the instability of the base learning algorithm. An alternative approach to generating an ensemble is to randomize the internal decisions made by the base algorithm. This general approach has been studied previously by Ali and Pazzani and by Dietterich and Kong. This paper compares the effectiveness of randomization, bagging, and boosting for improving the performance of the decision-tree algorithm C4.5. The experiments show that in situations with little or no classification noise, randomization is competitive with (and perhaps slightly superior to) bagging but not as accurate as boosting. In situations with substantial classification noise, bagging is much better than boosting, and sometimes better than randomization.
Added
2026-09-12

Induction of Decision Trees
J. R. Quinlan
Why you should read this
Introduces the ID3 algorithm for efficiently constructing interpretable decision trees from data using information gain principles.
The technology for building knowledge-based systems by inductive inference from examples has been demonstrated successfully in several practical applications. This paper summarizes an approach to synthesizing decision trees that has been used in a variety of systems, and it describes one such system, ID3, in detail. Results from recent studies show ways in which the methodology can be modified to deal with information that is noisy and/or incomplete. A reported shortcoming of the basic algorithm is discussed and two means of overcoming it are compared. The paper concludes with illustrations of current research directions.
Added
2026-05-14
