Built independently by an author, for readers. Read the story and support ChapterPal

keyword

class noise

Class noise refers to the presence of incorrect, corrupted, or contradictory category labels assigned to instances within a dataset used for supervised machine learning. Unlike attribute noise, which involves errors in the input feature values, class noise occurs when an instance is mistakenly assigned to the wrong target class due to factors such as human annotation mistakes, data entry errors, or ambiguous classification criteria. In machine learning workflows, class noise can mislead induction algorithms, distort decision boundaries, reduce overall classification accuracy, and unnecessarily increase model complexity as algorithms attempt to accommodate mislabeled examples.

2 items

Generating Accurate Rule Sets Without Global Optimization

Generating Accurate Rule Sets Without Global Optimization

Eibe Frank, Ian H. Witten

OrganizationsUniversity of Waikato

Why you should read this

Presents PART, a fast rule-learning algorithm that avoids complex global optimization by deriving rules from partial decision trees within a separate-and-conquer framework to achieve accuracy and compact model sizes matching or exceeding C4.5 and RIPPER.

The two dominant schemes for rule-learning, C4.5 and RIPPER, both operate in two stages. First they induce an initial rule set and then they refine it using a rather complex optimization stage that discards (C4.5) or adjusts (RIPPER) individual rules to make them work better together. In contrast, this paper shows how good rule sets can be learned one rule at a time, without any need for global optimization. We present an algorithm for inferring rules by repeatedly generating partial decision trees, thus combining the two major paradigms for rule generation—creating rules from decision trees and the separate-and-conquer rule-learning technique. The algorithm is straightforward and elegant: despite this, experiments on standard datasets show that it produces rule sets that are as accurate as and of similar size to those generated by C4.5, and more accurate than RIPPER's. Moreover, it operates efficiently, and because it avoids postprocessing, does not suffer the extremely slow performance on pathological example sets for which the C4.5 method has been criticized.

Added

2026-09-25