keyword
class noise
Class noise refers to the presence of incorrect, corrupted, or contradictory category labels assigned to instances within a dataset used for supervised machine learning. Unlike attribute noise, which involves errors in the input feature values, class noise occurs when an instance is mistakenly assigned to the wrong target class due to factors such as human annotation mistakes, data entry errors, or ambiguous classification criteria. In machine learning workflows, class noise can mislead induction algorithms, distort decision boundaries, reduce overall classification accuracy, and unnecessarily increase model complexity as algorithms attempt to accommodate mislabeled examples.
2 items

An Analysis of Bayesian Classifiers
Pat Langley, Wayne Iba, Kevin Thompson
Why you should read this
Presents an average-case theoretical analysis explaining why simple Bayesian classifiers achieve high accuracy across diverse learning domains despite their strong attribute independence assumptions.
In this paper we present an average-case analysis of the Bayesian classifier, a simple induction algorithm that fares remarkably well on many learning tasks. Our analysis assumes a monotone conjunctive target concept, and independent, noise-free Boolean attributes. We calculate the probability that the algorithm will induce an arbitrary pair of concept descriptions and then use this to compute the probability of correct classification over the instance space. The analysis takes into account the number of training instances, the number of attributes, the distribution of these attributes, and the level of class noise. We also explore the behavioral implications of the analysis by presenting predicted learning curves for artificial domains, and give experimental results on these domains as a check on our reasoning.
Added
2026-09-25

Generating Accurate Rule Sets Without Global Optimization
Eibe Frank, Ian H. Witten
Why you should read this
Presents PART, a fast rule-learning algorithm that avoids complex global optimization by deriving rules from partial decision trees within a separate-and-conquer framework to achieve accuracy and compact model sizes matching or exceeding C4.5 and RIPPER.
The two dominant schemes for rule-learning, C4.5 and RIPPER, both operate in two stages. First they induce an initial rule set and then they refine it using a rather complex optimization stage that discards (C4.5) or adjusts (RIPPER) individual rules to make them work better together. In contrast, this paper shows how good rule sets can be learned one rule at a time, without any need for global optimization. We present an algorithm for inferring rules by repeatedly generating partial decision trees, thus combining the two major paradigms for rule generation—creating rules from decision trees and the separate-and-conquer rule-learning technique. The algorithm is straightforward and elegant: despite this, experiments on standard datasets show that it produces rule sets that are as accurate as and of similar size to those generated by C4.5, and more accurate than RIPPER's. Moreover, it operates efficiently, and because it avoids postprocessing, does not suffer the extremely slow performance on pathological example sets for which the C4.5 method has been criticized.
Source
https://researchcommons.waikato.ac.nz/bitstreams/2e1b230f-cab4-471b-8076-915fd9a2d79c/downloadAdded
2026-09-25
