Auto-WEKA: combined selection and hyperparameter optimization of classification algorithms
Chris J. ThorntonF. HutterH. HoosKevin Leyton-Brown
Introduces a Bayesian optimization framework that automatically selects the best classification algorithm, feature selection method, and hyperparameter settings from the WEKA library to maximize predictive performance on a given dataset without manual intervention.
Organizations increasingly rely on machine learning for critical predictive tasks, yet non-expert practitioners often choose algorithms based on intuition and leave settings at arbitrary default values. This practice risks severe misclassification errors, as algorithm performance varies widely depending on data characteristics. While previous research has optimized model selection and algorithm hyperparameters separately, real-world deployment requires solving both problems simultaneously to achieve reliable accuracy.
The article evaluates an automated approach to the combined algorithm selection and hyperparameter optimization problem by framing it as a single hierarchical optimization task. Specifically, it demonstrates how Bayesian optimization techniques can automatically select the best learning pipeline from dozens of algorithms and tune over 700 associated hyperparameters.
To demonstrate this capability, the authors developed Auto-WEKA, an automated tool incorporating 39 classification models and 11 feature selection methods from the WEKA machine learning library. The authors benchmarked the system across 21 diverse datasets—ranging from small tabular sets to large image datasets—comparing its performance against standard default models and exhaustive random grid search baselines over fixed compute budgets.
The evaluation produced four primary findings. First, selecting algorithms with default parameters is highly unreliable: across the evaluated datasets, misclassification rates between the best and worst default algorithms differed by over 20% on 14 datasets. Second, the automated tool significantly outperformed traditional random grid search, finding lower cross-validation errors in 20 out of 21 datasets while using less than one-third of the computation time (120 CPU hours versus an average of 400 CPU hours). Third, generalization on held-out test data was superior in 15 of 21 datasets, achieving the highest accuracy on all 11 large datasets and reducing test error rates by over 15% in five cases. Fourth, the optimizer utilizing random-forest-based Sequential Model-based Algorithm Configuration (SMAC) consistently outperformed the alternative Tree-structured Parzen Estimator (TPE) approach, winning on 12 datasets compared to 6 for TPE.
These results demonstrate that automating algorithm selection and tuning dramatically reduces human effort and engineering risk while improving model performance, especially as dataset sizes grow. Organizations can systematically avoid suboptimal model deployments without requiring dedicated machine learning specialists to manually test combinations of algorithms and tuning parameters.
Decision-makers adopting automated machine learning workflows should implement SMAC-based optimization pipelines and allow parallelized multi-core exploration to maximize solution quality. For future enhancements, organizations and researchers should investigate advanced methods to mitigate overfitting on smaller datasets and incorporate ensemble learning strategies that construct blended multi-model portfolios.
While confidence is high regarding performance gains on medium-to-large datasets, users should exercise caution with small sample sizes. On smaller datasets, the enormous search space increases the risk of overfitting, meaning that cross-validation improvements may not translate into equivalent test-set performance.
- Paper: Algorithms for Hyper-Parameter Optimization, James Bergstra et al. (2011). This foundational work introduces the Tree-structured Parzen Estimator (TPE) algorithm for Sequential Model-based Algorithm Configuration, one of the primary Bayesian optimization search mechanisms benchmarked in Auto-WEKA.
- Paper: Random Search for Hyper-Parameter Optimization, James Bergstra et al. (2012). This paper establishes random search as a robust, scalable baseline for hyperparameter optimization across high-dimensional spaces, against which Auto-WEKA evaluates its Bayesian optimization approach.
- Paper: A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning, Eric Brochu et al. (2010). This tutorial covers the core principles of Bayesian optimization and acquisition functions for expensive black-box objective functions that underpin automated machine learning pipelines.
- Paper: On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation, Gavin C. Cawley et al. (2010). This study analyzes how tuning hyperparameters on finite validation sets introduces overfitting and selection bias, directly informing Auto-WEKA's caveats regarding small dataset evaluation.
- Paper: A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection, Ron Kohavi (1995). This work establishes the methodological basis of stratified k-fold cross-validation for model selection and performance evaluation used as Auto-WEKA's optimization objective.
- Paper: An empirical comparison of supervised learning algorithms, Rich Caruana et al. (2006). This large-scale empirical study provides critical precedent for systematically benchmarking diverse supervised learning algorithms across standard datasets and metrics.
- Paper: Efficient and Robust Automated Machine Learning, Matthias Feurer et al. (2015). This work directly extends Auto-WEKA's combined algorithm selection and hyperparameter optimization framework to the scikit-learn ecosystem while introducing meta-learning warmstarts and post-hoc ensembling.
- Paper: Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization, Lisha Li et al. (2016). This paper proposes a bandit-based multi-fidelity optimization framework (Hyperband) that serves as an alternative and extension to the standard Bayesian optimization methods evaluated in Auto-WEKA.
- Paper: AutoML: A Survey of the State-of-the-Art, Xin He et al. (2019). This comprehensive survey contextualizes the evolution of automated machine learning from pipeline optimization systems like Auto-WEKA to modern neural architecture search.
- Paper: On Hyperparameter Optimization of Machine Learning Algorithms: Theory and Practice, Li Yang et al. (2020). This survey reviews the broader landscape of automatic hyperparameter optimization techniques and software tools that developed following early AutoML systems.
- Paper: Do we need hundreds of classifiers to solve real world classification problems?, Manuel Fernández Delgado et al. (2014). This extensive benchmark analyzes hundreds of classifiers across the UCI repository to determine whether complex model spaces like those searched by Auto-WEKA are necessary in practice.
- Paper: Deep feature synthesis: Towards automating data science endeavors, James Max Kanter et al. (2015). This paper expands end-to-end automation beyond model selection and parameter tuning to include automated relational feature engineering.
