keyword
minimum description length principle
The minimum description length principle is an information-theoretic framework for statistical inference and model selection that asserts the best model for a given set of data is the one that minimizes the total combined length needed to encode both the model and the data described by that model. Serving as a formal, mathematical formulation of Occams razor, it views learning as data compression, positing that identifying regular patterns in data allows for more compact representations. By penalizing excessively complex models that require longer descriptions, the principle balances goodness of fit against complexity to prevent overfitting and ensure robust generalization, making it a foundational criterion in machine learning for tasks such as feature discretization, rule induction, and structure optimization.
3 items

Generating Accurate Rule Sets Without Global Optimization
Eibe Frank, Ian H. Witten
Why you should read this
Presents PART, a fast rule-learning algorithm that avoids complex global optimization by deriving rules from partial decision trees within a separate-and-conquer framework to achieve accuracy and compact model sizes matching or exceeding C4.5 and RIPPER.
The two dominant schemes for rule-learning, C4.5 and RIPPER, both operate in two stages. First they induce an initial rule set and then they refine it using a rather complex optimization stage that discards (C4.5) or adjusts (RIPPER) individual rules to make them work better together. In contrast, this paper shows how good rule sets can be learned one rule at a time, without any need for global optimization. We present an algorithm for inferring rules by repeatedly generating partial decision trees, thus combining the two major paradigms for rule generation—creating rules from decision trees and the separate-and-conquer rule-learning technique. The algorithm is straightforward and elegant: despite this, experiments on standard datasets show that it produces rule sets that are as accurate as and of similar size to those generated by C4.5, and more accurate than RIPPER's. Moreover, it operates efficiently, and because it avoids postprocessing, does not suffer the extremely slow performance on pathological example sets for which the C4.5 method has been criticized.
Source
https://researchcommons.waikato.ac.nz/bitstreams/2e1b230f-cab4-471b-8076-915fd9a2d79c/downloadAdded
2026-09-25

Multi-Interval Discretization of Continuous-Valued Attributes for Classification Learning
Usama M. Fayyad, Keki B. Irani
Why you should read this
Develops a minimum description length criterion for multi-interval discretization of continuous attributes in decision tree learning and proves that entropy-minimizing cut points always lie on class boundaries, producing smaller and more accurate classification models.
Since most real-world applications of classification learning involve continuous-valued attributes, properly addressing the discretization process is an important problem. This paper addresses the use of the entropy minimization heuristic for discretizing the range of a continuous-valued attribute into multiple intervals. We briefly present theoretical evidence for the appropriateness of this heuristic for use in the binary discretization algorithm used in ID3, C4, CART, and other learning algorithms. The results serve to justify extending the algorithm to derive multiple intervals. We formally derive a criterion based on the minimum description length principle for deciding the partitioning of intervals. We demonstrate via empirical evaluation on several real-world data sets that better decision trees are obtained using the new multi-interval algorithm.
Added
2026-09-11

A Tutorial Introduction to the Minimum Description Length Principle
Peter D. Grünwald
Why you should read this
This tutorial teaches you how to choose the right level of complexity when building models from data—neither oversimplifying nor overfitting—by connecting the intuitive idea that better explanations compress information more efficiently to rigorous mathematical tools you can actually use.
This tutorial provides an overview of and introduction to Rissanen's Minimum Description Length (MDL) Principle. The first chapter provides a conceptual, entirely non-technical introduction to the subject. It serves as a basis for the technical introduction given in the second chapter, in which all the ideas of the first chapter are made mathematically precise. The main ideas are discussed in great conceptual and technical detail. This tutorial is an extended version of the first two chapters of the collection "Advances in Minimum Description Length: Theory and Application" (edited by this http URL, I.J. Myung and M. Pitt, to be published by the MIT Press, Spring 2005).
Added
2026-02-21
