Logistic Model Trees
Niels LandwehrMark HallEibe Frank
Presents Logistic Model Trees (LMT), an algorithm that blends decision tree structures with incrementally refined leaf-level logistic regression models to achieve classification accuracy competitive with boosted trees while maintaining model interpretability and compactness.
Modern data-driven decision-making frequently requires classification models that deliver both high predictive accuracy and clear interpretability. Standard decision trees provide easily understood rules but suffer from instability and high variance, whereas linear logistic regression models offer stable predictions but cannot capture nonlinear patterns in complex data. The article addresses this trade-off by introducing and evaluating Logistic Model Trees (LMT), an automated machine learning algorithm that integrates standard tree structures with linear logistic regression models situated directly at the leaf nodes.
The primary objective of the article is to demonstrate that LMT produces compact, highly accurate classifiers that automatically scale model complexity to match domain characteristics without requiring manual parameter tuning. To evaluate this approach, the authors tested LMT across 36 diverse benchmark datasets spanning small to large sample sizes. They benchmarked LMT against standard decision trees (C4.5 and CART), linear logistic regression variants, other hybrid tree learners (such as Functional Trees and Naive Bayes Trees), and ensemble methods including boosted trees (AdaBoost) and multi-tree regression systems (M5'). Credibility was reinforced using ten runs of ten-fold cross-validation combined with corrected statistical significance testing.
The findings show that LMT consistently equals or outperforms standard decision trees and standalone logistic regression, never losing a statistically significant comparison against them across all 36 datasets. Specifically, LMT achieved significant accuracy wins over C4.5 in 16 datasets and CART in 17 datasets, while producing drastically smaller trees—often reducing leaf counts from hundreds or thousands down to fewer than a dozen. Furthermore, LMT significantly outperformed other enhanced tree learners and proved highly competitive with boosted decision trees (AdaBoost with 100 iterations), matching their accuracy across most datasets while delivering a single, interpretable tree rather than an opaque voting ensemble of 100 separate trees. On 18 of the 36 datasets, LMT automatically pruned the tree entirely back to the root, selecting a simple linear model when additional tree structure was unjustified.
These results demonstrate that organizations do not necessarily have to sacrifice model interpretability to achieve state-of-the-art predictive performance. By incorporating stepwise attribute selection and incrementally refining logistic models down the tree hierarchy, LMT avoids overfitting and isolates the most critical predictive factors. This reduces the risk of relying on misleading variables, lowers operational complexity, and facilitates regulatory compliance or stakeholder auditing through transparent decision paths.
For practical implementation, teams seeking robust, off-the-shelf classification should adopt LMT as an alternative to both basic decision trees and black-box ensemble methods. However, decision-makers should account for training time constraints, as LMT is several orders of magnitude slower to train than standard C4.5 due to nested cross-validation procedures. Future work recommended by the source includes developing faster fitting procedures to bypass repeated cross-validation and implementing more sophisticated imputation methods for missing data. Confidence in the empirical results is high given the rigorous cross-validation and statistical controls across varied real-world benchmarks.
- Paper: Scaling Up the Accuracy of Naive-Bayes Classifiers: A Decision-Tree Hybrid, Ron Kohavi (1996). It introduces the foundational concept of hybrid decision trees with probabilistic models at leaf nodes, which Logistic Model Trees directly builds upon.
- Paper: Improved Use of Continuous Attributes in C4.5, J. R. Quinlan (1996). It details continuous attribute splitting and pruning mechanisms in C4.5 that serve as core architectural baselines for Logistic Model Trees.
- Paper: A Brief Introduction to Boosting, R. Schapire (1999). It provides the foundational principles of boosting algorithms like AdaBoost and LogitBoost used to fit logistic models stagewise within LMT nodes.
- Paper: Induction of Decision Trees, J. R. Quinlan (1986). It establishes the foundational top-down induction of decision trees framework upon which tree-structured classifiers rely.
- Paper: Greedy function approximation: A gradient boosting machine, Jerome H. Friedman (2001). It develops gradient boosting and stagewise additive expansion techniques that underpin the fitting of logistic models at tree nodes.
- Paper: On Discriminative vs. Generative Classifiers: A comparison of logistic regression and naive Bayes, Andrew Ng et al. (2001). It provides fundamental theoretical and empirical comparisons between discriminative logistic regression and generative classifiers.
- Paper: Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms, Thomas G. Dietterich (1998). It outlines the corrected cross-validation and statistical significance testing methodology used directly to benchmark Logistic Model Trees.
- Paper: An empirical comparison of supervised learning algorithms, Rich Caruana et al. (2006). It provides an extensive multi-metric empirical benchmark of supervised classification algorithms, contextualizing tree and logistic regression performance.
- Paper: Intelligible Models for HealthCare: Predicting Pneumonia Risk and Hospital 30-day Readmission, Rich Caruana et al. (2015). It advances the pursuit of combining high predictive accuracy with interpretability by introducing generalized additive models with pairwise interactions.
- Paper: Auto-WEKA: combined selection and hyperparameter optimization of classification algorithms, Chris J. Thornton et al. (2012). It automates the selection and hyperparameter optimization of classification algorithms across the WEKA framework, which houses LMT.
- Paper: Do we need hundreds of classifiers to solve real world classification problems?, Manuel Fernández Delgado et al. (2014). It evaluates real-world classification performance across hundreds of classifiers to assess the competitive standing of tree-based and linear models.
- Paper: Predicting good probabilities with supervised learning, Alexandru Niculescu-Mizil et al. (2005). It investigates probability calibration across supervised learning models, evaluating how tree-based and logistic models predict true probabilities.
