Auto-WEKA: combined selection and hyperparameter optimization of classification algorithms

Chris J. ThorntonF. HutterH. HoosKevin Leyton-Brown

article2012KDD1,728 citationsTest of Time Award for Research

Introduces a Bayesian optimization framework that automatically selects the best classification algorithm, feature selection method, and hyperparameter settings from the WEKA library to maximize predictive performance on a given dataset without manual intervention.

Listen

Organizations increasingly rely on machine learning for critical predictive tasks, yet non-expert practitioners often choose algorithms based on intuition and leave settings at arbitrary default values. This practice risks severe misclassification errors, as algorithm performance varies widely depending on data characteristics. While previous research has optimized model selection and algorithm hyperparameters separately, real-world deployment requires solving both problems simultaneously to achieve reliable accuracy.

The article evaluates an automated approach to the combined algorithm selection and hyperparameter optimization problem by framing it as a single hierarchical optimization task. Specifically, it demonstrates how Bayesian optimization techniques can automatically select the best learning pipeline from dozens of algorithms and tune over 700 associated hyperparameters.

To demonstrate this capability, the authors developed Auto-WEKA, an automated tool incorporating 39 classification models and 11 feature selection methods from the WEKA machine learning library. The authors benchmarked the system across 21 diverse datasets—ranging from small tabular sets to large image datasets—comparing its performance against standard default models and exhaustive random grid search baselines over fixed compute budgets.

The evaluation produced four primary findings. First, selecting algorithms with default parameters is highly unreliable: across the evaluated datasets, misclassification rates between the best and worst default algorithms differed by over 20% on 14 datasets. Second, the automated tool significantly outperformed traditional random grid search, finding lower cross-validation errors in 20 out of 21 datasets while using less than one-third of the computation time (120 CPU hours versus an average of 400 CPU hours). Third, generalization on held-out test data was superior in 15 of 21 datasets, achieving the highest accuracy on all 11 large datasets and reducing test error rates by over 15% in five cases. Fourth, the optimizer utilizing random-forest-based Sequential Model-based Algorithm Configuration (SMAC) consistently outperformed the alternative Tree-structured Parzen Estimator (TPE) approach, winning on 12 datasets compared to 6 for TPE.

These results demonstrate that automating algorithm selection and tuning dramatically reduces human effort and engineering risk while improving model performance, especially as dataset sizes grow. Organizations can systematically avoid suboptimal model deployments without requiring dedicated machine learning specialists to manually test combinations of algorithms and tuning parameters.

Decision-makers adopting automated machine learning workflows should implement SMAC-based optimization pipelines and allow parallelized multi-core exploration to maximize solution quality. For future enhancements, organizations and researchers should investigate advanced methods to mitigate overfitting on smaller datasets and incorporate ensemble learning strategies that construct blended multi-model portfolios.

While confidence is high regarding performance gains on medium-to-large datasets, users should exercise caution with small sample sizes. On smaller datasets, the enormous search space increases the risk of overfitting, meaning that cross-validation improvements may not translate into equivalent test-set performance.

Cover for Auto-WEKA: combined selection and hyperparameter optimization of classification algorithms

Abstract

Many different machine learning algorithms exist; taking into account each algorithm's hyperparameters, there is a staggeringly large number of possible alternatives overall. We consider the problem of simultaneously selecting a learning algorithm and setting its hyperparameters, going beyond previous work that addresses these issues in isolation. We show that this problem can be addressed by a fully automated approach, leveraging recent innovations in Bayesian optimization. Specifically, we consider a wide range of feature selection techniques (combining 3 search and 8 evaluator methods) and all classification approaches implemented in WEKA, spanning 2 ensemble methods, 10 meta-methods, 27 base classifiers, and hyperparameter settings for each classifier. On each of 21 popular datasets from the UCI repository, the KDD Cup 09, variants of the MNIST dataset and CIFAR-10, we show classification performance often much better than using standard selection/hyperparameter optimization methods. We hope that our approach will help non-expert users to more effectively identify machine learning algorithms and hyperparameter settings appropriate to their applications, and hence to achieve improved performance.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Model Selection
  • 2.2 Hyperparameter Optimization
  • 3 Combined Algorithm Selection and Hyperparameter Optimization (CASH)
  • 3.1 Sequential Model-based Algorithm Configuration (SMAC)
  • 3.2 Tree-structured Parzen Estimator (TPE)
  • 4 Auto-WEKA
  • 5 Evaluating Auto-WEKA
  • 5.1 Experimental setup
  • 5.2 Algorithm Selection and CASH: Baseline Methods
  • 5.3 Results for Cross-Validation Performance
  • 5.4 Results for Test Performance
  • 5.5 Classifiers Selected by Auto-WEKA
  • 6 Conclusion and future work
  • References

Knowls

  1. Knowl 1 — Combined Algorithm Selection and Hyperparameter Optimization Problem Formulation

    definition

    The Combined Algorithm Selection and Hyperparameter Optimization (CASH) problem addresses the simultaneous selection of a learning algorithm and its hyperparameter settings to optimize empirical generalization performance.

    Let A={A(1),…,A(k)}\mathcal{A} = \{A^{(1)}, \dots, A^{(k)}\} be a set of learning algorithms, where each algorithm A(j)A^{(j)} has an associated hyperparameter space Λ(j)\Lambda^{(j)}. Let a training dataset D={(x1,y1),…,(xn,yn)}D = \{(x_1, y_1), \dots, (x_n, y_n)\} with input features xi∈Xx_i \in \mathcal{X} and discrete class labels yi∈Yy_i \in \mathcal{Y} be partitioned into kk disjoint validation sets Dvalid(1),…,Dvalid(k)D_{\text{valid}}^{(1)}, \dots, D_{\text{valid}}^{(k)} and corresponding training sets Dtrain(i)=D∖Dvalid(i)D_{\text{train}}^{(i)} = D \setminus D_{\text{valid}}^{(i)} for i=1,…,ki = 1, \dots, k. Given a loss function L(Aλ,Dtrain,Dvalid)\mathcal{L}(A_\lambda, D_{\text{train}}, D_{\text{valid}}) denoting the validation misclassification rate of algorithm AA instantiated with hyperparameters λ\lambda after training on DtrainD_{\text{train}}, the CASH problem is defined as:

    Aλ∗∗∈arg⁡min⁡A(j)∈A, λ∈Λ(j)1k∑i=1kL(Aλ(j),Dtrain(i),Dvalid(i))A^*_{\lambda^*} \in \arg\min_{A^{(j)} \in \mathcal{A}, \, \lambda \in \Lambda^{(j)}} \frac{1}{k} \sum_{i=1}^k \mathcal{L}\left(A^{(j)}_\lambda, D_{\text{train}}^{(i)}, D_{\text{valid}}^{(i)}\right)

    This formulation is equivalent to a single combined hierarchical hyperparameter optimization problem over the parameter space Λ=Λ(1)∪⋯∪Λ(k)∪{λr}\Lambda = \Lambda^{(1)} \cup \dots \cup \Lambda^{(k)} \cup \{\lambda_r\}, where λr\lambda_r is a root-level categorical hyperparameter that selects among the algorithms A(1),…,A(k)A^{(1)}, \dots, A^{(k)}, and the root-level parameters of each subspace Λ(j)\Lambda^{(j)} are conditional upon λr\lambda_r selecting algorithm A(j)A^{(j)}.

  2. Knowl 2 — Auto-WEKA Hierarchical Parameter Space Architecture

    model/method

    Auto-WEKA structures the CASH problem for the entire WEKA machine learning library as a four-layer hierarchical parameter tree containing 786 hyperparameters.

    The top-level structure consists of two independent branches:

    1. Classification Pipeline Branch:

      • Controlled by a root Boolean parameter is_base.
      • If is_base is true, a categorical parameter base chooses one of 27 base classification algorithms (e.g., Support Vector Machines, Random Forest, Multi-Layer Perceptron, Gaussian Processes, Decision Trees).
      • If is_base is false, a categorical parameter class chooses between 10 meta-methods (e.g., AdaBoostM1, Bagging, LogitBoost) or 2 ensemble methods (Voting and Stacking).
      • For meta-methods, a child parameter meta_base selects one of the 27 base classifiers to serve as the base learner.
      • For ensemble methods, an integer parameter num_classes ∈{1,…,5}\in \{1, \dots, 5\} determines the ensemble size, activating parameters base_1 through base_5 which each independently select one of the 27 base classifiers.
      • Each selected base classifier activates its own conditional categorical and numerical hyperparameters.
    2. Feature Selection Pipeline Branch:

      • Controlled by a top-level Boolean parameter feat_sel.
      • If feat_sel is false, the raw dataset is passed directly to the classifier.
      • If feat_sel is true, feature preprocessing is executed prior to classifier training. A categorical parameter feat_search selects 1 of 3 search strategies (Best First, Greedy Stepwise, Ranker; with up to 5 subparameters), and feat_eval selects 1 of 8 feature evaluators (CFS Subset Eval, Pearson Correlation Eval, Gain Ratio Eval, Info Gain Eval, 1-R Eval, Principal Components Eval, RELIEF Eval, Symmetrical Uncertainty Eval; with up to 4 subparameters).

    Numerical parameters are assigned uniform or log-uniform priors depending on their semantics (e.g., log-uniform for regularization penalties, uniform for maximum tree depths). Discretizing each hyperparameter domain to 10 values results in a space of over 104710^{47} configurations.

  3. Knowl 3 — Sequential Model-Based Optimization Framework for CASH

    algorithm

    Sequential Model-Based Optimization (SMBO) solves hierarchical hyperparameter optimization by iteratively fitting a surrogate response surface model MLM_\mathcal{L} to observed loss values and optimizing an acquisition function to select subsequent candidate configurations.

    Input: Dataset DD, loss function L\mathcal{L}, hyperparameter space Λ\Lambda, total optimization time budget TT
    Output: Configuration λ∗∈Λ\lambda^* \in \Lambda with lowest observed loss
    Initialize surrogate model MLM_{\mathcal{L}}
    Initialize observation history H←∅\mathcal{H} \leftarrow \emptyset
    while time budget TT has not been exhausted do
        Select candidate λ←arg⁡max⁡λ′∈ΛaML(λ′)\lambda \leftarrow \arg\max_{\lambda' \in \Lambda} a_{M_{\mathcal{L}}}(\lambda') via acquisition function aMLa_{M_{\mathcal{L}}}
        Evaluate loss c=1k∑i=1kL(Aλ,Dtrain(i),Dvalid(i))c = \frac{1}{k} \sum_{i=1}^k \mathcal{L}(A_\lambda, D_{\text{train}}^{(i)}, D_{\text{valid}}^{(i)})
        Update history H←H∪{(λ,c)}\mathcal{H} \leftarrow \mathcal{H} \cup \{(\lambda, c)\}
        Refit surrogate model MLM_{\mathcal{L}} using H\mathcal{H}
    end while
    return λ∗=arg⁡min⁡(λ,c)∈Hc\lambda^* = \arg\min_{(\lambda, c) \in \mathcal{H}} c

    The acquisition function aML(λ)a_{M_\mathcal{L}}(\lambda) computes the positive Expected Improvement (EI) over the incumbent minimum loss cmin⁡c_{\min}:

    Icmin⁡(λ):=max⁡{cmin⁡−c(λ),0}I_{c_{\min}}(\lambda) := \max\{c_{\min} - c(\lambda), 0\}

    EML[Icmin⁡(λ)]=∫−∞cmin⁡max⁡{cmin⁡−c,0}⋅pML(c∣λ) dc\mathbb{E}_{M_\mathcal{L}}[I_{c_{\min}}(\lambda)] = \int_{-\infty}^{c_{\min}} \max\{c_{\min} - c, 0\} \cdot p_{M_\mathcal{L}}(c \mid \lambda) \, dc

    In Sequential Model-based Algorithm Configuration (SMAC), MLM_\mathcal{L} is a random forest that outputs frequentist mean μλ\mu_\lambda and variance σλ2\sigma_\lambda^2 across individual trees, modeling pML(c∣λ)=N(μλ,σλ2)p_{M_\mathcal{L}}(c \mid \lambda) = \mathcal{N}(\mu_\lambda, \sigma_\lambda^2). The expected improvement evaluates in closed form as:

    EML[Icmin⁡(λ)]=σλ⋅[u⋅Φ(u)+ϕ(u)],u=cmin⁡−μλσλ\mathbb{E}_{M_\mathcal{L}}[I_{c_{\min}}(\lambda)] = \sigma_\lambda \cdot [u \cdot \Phi(u) + \phi(u)], \quad u = \frac{c_{\min} - \mu_\lambda}{\sigma_\lambda}

    where ϕ\phi and Φ\Phi denote the standard normal PDF and CDF.

    In the Tree-structured Parzen Estimator (TPE), the density is modeled as p(λ∣c)=ℓ(λ)p(\lambda \mid c) = \ell(\lambda) if c<c∗c < c^* and p(λ∣c)=g(λ)p(\lambda \mid c) = g(\lambda) if c≥c∗c \ge c^* (where c∗c^* is the γ\gamma-quantile of observed losses, γ=0.15\gamma = 0.15), and EI is maximized by finding arg⁡min⁡λg(λ)/ℓ(λ)\arg\min_\lambda g(\lambda) / \ell(\lambda).

  4. Knowl 4 — Progressive Multi-Fold Evaluation and Diversification in SMAC

    model/method

    To minimize computational cost during the optimization of cross-validation objectives, SMAC implements a progressive evaluation mechanism over cross-validation folds rather than evaluating all folds for every candidate configuration.

    When evaluating a candidate hyperparameter configuration λ\lambda:

    1. The loss terms L(Aλ,Dtrain(i),Dvalid(i))\mathcal{L}(A_\lambda, D_{\text{train}}^{(i)}, D_{\text{valid}}^{(i)}) are computed one fold at a time.
    2. In order for λ\lambda to replace the current incumbent configuration λinc\lambda_{\text{inc}}, λ\lambda must achieve a lower cumulative average error than λinc\lambda_{\text{inc}} across each subset of folds 1,2,…,m1, 2, \dots, m, where mm is the number of folds on which λinc\lambda_{\text{inc}} has been evaluated.
    3. If λ\lambda performs worse than λinc\lambda_{\text{inc}} on any intermediate comparison, evaluation terminates immediately, allowing poor configurations to be discarded after evaluating as few as a single fold.
    4. Whenever an incumbent survives a comparison, it is evaluated on an additional fold until all kk folds have been evaluated.

    To prevent the search from getting trapped in local optima due to surrogate model inaccuracies, SMAC incorporates a diversification mechanism that selects every second candidate hyperparameter configuration uniformly at random from Λ\Lambda rather than maximizing the acquisition function.

  5. Knowl 5 — Baseline Methods for Algorithm Selection and CASH

    model/method

    To evaluate combined algorithm selection and hyperparameter optimization against standard practices, two baseline methods are defined:

    1. Exhaustive Default Selection (Ex-Def):

      • Evaluates all 39 WEKA classification algorithms using their default hyperparameter settings.
      • Measures performance using exhaustive 10-fold cross-validation on the training set.
      • Selects the classifier with the lowest cross-validation error rate.
    2. Random Grid Search:

      • Constructs a discrete grid of hyperparameter values for each of WEKA's 27 base classifiers.
      • Samples uniformly at random from the union of these 27 hyperparameter grids together with the 39 default classifier configurations.
      • Serves as a standard, unguided CASH baseline and is allocated an average of 400 CPU hours per dataset.
  6. Knowl 6 — Incumbent Trajectory Spearman Correlation for Overfitting Diagnostics

    model/method

    To diagnose the risk of overfitting during the CASH search without querying the final test set, the available training data is partitioned into two disjoint subsets:

    • 70% optimization training data, used internally by the SMBO optimizer for 10-fold cross-validation.
    • 30% validation data, held out during the optimization process.

    During optimization, SMBO maintains the sequence of incumbent configurations λ1,λ2,…,λm\lambda_1, \lambda_2, \dots, \lambda_m (configurations that achieved a new best 10-fold cross-validation error rate ecv(λj)e_{\text{cv}}(\lambda_j) on the 70% split). After the optimization budget is exhausted, each incumbent λj\lambda_j is trained on the entire 70% optimization split and evaluated on the held-out 30% validation split to compute validation error eval(λj)e_{\text{val}}(\lambda_j).

    The diagnostic score is the Spearman rank correlation coefficient rsr_s between the sequence of cross-validation error values (ecv(λj))j=1m(e_{\text{cv}}(\lambda_j))_{j=1}^m and validation error values (eval(λj))j=1m(e_{\text{val}}(\lambda_j))_{j=1}^m. A decaying or negative correlation coefficient indicates that improvements in cross-validation loss no longer correspond to improved generalization on unseen data, providing an automated warning of overfitting.

  7. Knowl 7 — Cross-Validation and Test Generalization Performance Across 21 Benchmarks

    data/table

    Auto-WEKA (leveraging SMAC and TPE) was evaluated on 21 classification datasets against Ex-Def (exhaustive default selection across 39 classifiers) and Random Grid Search (400 CPU hours average). Auto-WEKA was executed with 4 parallel runs for 30 hours per run (median percent error across 100,000 bootstrap samples).

    Dataset 10-Fold C.V. Performance (%) Test Performance (%) SC
    Ex-Def Rand. Grid TPE SMAC Ex-Def Rand. Grid TPE SMAC TPE SMAC
    Dexter 10.20 7.48 9.90 5.48 8.89 5.00 9.44 7.22 0.82 0.25
    GermanCredit 22.45 22.45 21.43 19.59 27.33 27.33 27.67 28.33 0.31 0.20
    Dorothea 6.03 6.03 6.93 5.52 6.96 6.96 6.96 6.38 0.95 0.40
    Yeast 39.43 38.87 35.03 36.27 40.45 40.90 41.12 40.45 0.36 0.49
    Amazon 43.94 43.94 48.43 48.30 28.44 28.44 37.56 37.56 0.92 0.97
    Secom 6.25 6.12 6.25 5.34 8.09 8.30 7.87 7.87 -0.10 -0.56
    Semeion 6.52 6.52 6.91 4.86 8.18 8.18 8.18 5.03 0.84 0.73
    Car 2.71 1.54 0.94 0.71 0.77 0.19 0.00 0.58 0.12 0.75
    Madelon 25.98 24.26 24.26 20.87 21.38 20.77 20.77 21.15 0.44 0.43
    KR-vs-KP 0.89 0.70 0.45 0.32 0.31 0.52 0.52 0.31 0.22 0.32
    Abalone 73.33 72.45 72.20 71.76 73.18 72.79 72.71 73.02 0.15 0.10
    Wine Quality 38.94 37.28 35.94 34.74 37.51 36.08 33.56 33.70 0.73 0.85
    Waveform 12.73 12.73 12.57 11.71 14.40 14.40 14.20 14.40 0.36 0.26
    Gisette 3.62 3.27 3.70 2.42 2.81 2.38 2.57 2.24 0.69 0.79
    Convex 28.68 28.50 29.04 24.70 25.96 26.76 25.45 22.05 0.98 0.84
    CIFAR-10-Small 66.59 65.11 57.97 57.76 65.91 64.54 56.65 55.93 0.93 0.80
    MNIST Basic 5.12 4.00 13.64 3.64 5.19 3.79 18.03 3.56 1.00 0.87
    Rot. MNIST + BI 66.15 59.75 73.04 59.61 63.14 58.16 69.86 55.84 0.50 0.95
    Shuttle 0.0328 0.0263 0.0230 0.0230 0.0138 0.0276 0.0069 0.0069 0.60 0.73
    KDD09-Appentency 1.88 1.88 1.88 1.75 1.75 1.77 1.74 1.74 0.89 1.00
    CIFAR-10 65.54 65.54 66.68 63.21 64.27 64.27 64.80 62.39 0.33 0.69

    Key observations:

    • Cross-Validation Optimization: Auto-WEKA with SMAC achieved the lowest CV error on 19 of 21 datasets (tying on 1), outperforming Random Grid Search on 20/21 datasets despite using significantly fewer CPU hours (120 vs. 400 hours).
    • Test Generalization: Auto-WEKA (SMAC or TPE) outperformed Ex-Def and Random Grid Search on 15 of 21 datasets (3 ties, 3 losses). On all 11 largest datasets (N≥10,000N \ge 10{,}000 training instances: Convex, CIFAR-10-Small, MNIST Basic, Rot. MNIST + BI, Shuttle, KDD09-Appentency, CIFAR-10, Gisette, Waveform, Wine Quality, Abalone), Auto-WEKA achieved the best test performance.
    • Optimizer Comparison: SMAC outperformed TPE on 19/21 datasets in CV error and 12/21 datasets on test error (with 3 ties).
  8. Knowl 8 — Empirical Classifier and Feature Selector Selection Distributions

    empirical result

    Analysis of the configurations selected by Auto-WEKA (SMAC and TPE) across runs reveals distinct patterns based on dataset size:

    • Classifier Diversity: No single classification algorithm dominates the search space. The three most frequently selected standalone classifiers—Random Forest, Single-Layer Perceptron (SLP), and Support Vector Machines (SVM)—were each selected in only approximately 12% of runs. Most other classifiers were chosen in at least a small fraction of cases.
    • Scale-Dependent Algorithm Preference: Meta-methods and ensemble methods were selected substantially more often on large datasets (N≥10,000N \ge 10{,}000) than on small datasets.
    • Base Classifiers within Meta-Methods: Inside AdaBoostM1, Single-Layer Perceptrons were frequently chosen for small datasets but never for large datasets, where Reduced Error Pruning Trees (REP Tree) dominated. In Random Subspace ensembles, Naive Bayes and Decision Table were the most frequent base learners, despite rarely being chosen on their own as standalone classifiers.
    • Feature Selection Usage: Feature selection was selected much more frequently on small datasets than on large datasets, functioning as a regularizer to prevent overfitting. When active, the Ranker search method was chosen predominantly. For feature evaluators, small datasets used all evaluators with roughly equal frequency, whereas large datasets almost exclusively selected Information Gain (Info Gain Eval).

Coverage note — No substantial contributed material was omitted. All theoretical formulations, model architectures, algorithms, experimental baselines, quantitative benchmark tables, and empirical analyses have been captured.

References

  1. 1.M. Adankon and M. Cheriet. Model selection for the LS-SVM. application to handwriting recognition. Pattern Recognition, 42(12):3264–3270, 2009.
  2. 2.Y. Bengio. Gradient-based optimization of hyperparameters. Neural Computation, 12(8):1889–1900, 2000.
  3. 3.J. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl. Algorithms for Hyper-Parameter Optimization. In Proc. of NIPS-11, 2011.
  4. 4.J. Bergstra and Y. Bengio. Random search for hyperparameter optimization. JMLR, 13:281–305, 2012.
  5. 5.A. Biem. A model selection criterion for classification: Application to HMM topology optimization. In Proc. of ICDAR-03, pages 104–108. IEEE, 2003.
  6. 6.H. Bozdogan. Model selection and Akaike’s information criterion (AIC): The general theory and its analytical extensions. Psychometrika, 52(3):345–370, 1987.
  7. 7.P. Brazdil, C. Soares, and J. Da Costa. Ranking learning algorithms: Using IBL and meta-learning on accuracy and time results. Machine Learning, 50(3):251–277, 2003.
  8. 8.E. Brochu, V. M. Cora, and N. de Freitas. A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning. Technical Report UBC TR-2009-23 and arXiv:1012.2599v1, Department of Computer Science, University of British Columbia, 2009.
  9. 9.O. Chapelle, V. Vapnik, and Y. Bengio. Model selection for small sample regression. Machine Learning, 2001.
  10. 10.T. Desautels, A. Krause, and J. Burdick. Parallelizing exploration-exploitation tradeoffs with gaussian process bandit optimization. In Proc. of ICML-12, 2012.
  11. 11.A. Frank and A. Asuncion. UCI machine learning repository, 2010. URL: http://archive.ics.uci.edu/ml. University of California, Irvine, School of Information and Computer Sciences.
  12. 12.X. Guo, J. Yang, C. Wu, C. Wang, and Y. Liang. A novel LS-SVMs hyper-parameter selection based on particle swarm optimization. Neurocomputing, 71(16):3211–3215, 2008.
  13. 13.M. Hall, E. Frank, G. Holmes, B. Pfahringer, P. Reutemann, and I. Witten. The WEKA data mining software: an update. ACM SIGKDD Explorations Newsletter, 11(1):10–18, 2009.
  14. 14.F. Hutter, H. Hoos, and K. Leyton-Brown. Sequential model-based optimization for general algorithm configuration. Proc. of LION-5, pages 507–523, 2011.
  15. 15.F. Hutter, H. Hoos, K. Leyton-Brown, and T. Stützle. ParamILS: an automatic algorithm configuration framework. JAIR, 36(1):267–306, 2009.
  16. 16.F. Hutter, H. H. Hoos, and K. Leyton-Brown. Parallel algorithm configuration. In Proc. of LION-6, 2012.
  17. 17.D. R. Jones, M. Schonlau, and W. J. Welch. Efficient global optimization of expensive black box functions. Journal of Global Optimization, 13:455–492, 1998.
  18. 18.R. Kohavi. A study of cross-validation and bootstrap for accuracy estimation and model selection. In Proc. of IJCAI-95, pages 1137–1145, 1995.
  19. 19.A. Krizhevsky and G. Hinton. Learning multiple layers of features from tiny images. Master’s thesis, Department of Computer Science, University of Toronto, 2009.
  20. 20.O. Maron and A. Moore. Hoeffding races: Accelerating model selection search for classification and function approximation. In Proc. of NIPS-94, pages 59–66, 1994.
  21. 21.A. McQuarrie and C. Tsai. Regression and time series model selection. World Scientific, 1998.
  22. 22.T. Schaul, J. Bayer, D. Wierstra, Y. Sun, M. Felder, F. Sehnke, T. Rückstieß, and J. Schmidhuber. PyBrain. JMLR, 2010.
  23. 23.M. Schonlau, W. J. Welch, and D. R. Jones. Global versus local search in constrained optimization of computer models. In N. Flournoy, W. Rosenberger, and W. Wong, editors, New Developments and Applications in Experimental Design, volume 34, pages 11–25. Institute of Mathematical Statistics, Hayward, California, 1998.
  24. 24.J. Snoek, H. Larochelle, and R. Adams. Opportunity cost in Bayesian optimization. In NIPS Workshop on Bayesian Optimization, Sequential Experimental Design, and Bandits, 2011. Published online.
  25. 25.J. Snoek, H. Larochelle, and R. P. Adams. Practical bayesian optimization of machine learning algorithms. In Proc. of NIPS-12, 2012.
  26. 26.N. Srinivas, A. Krause, S. Kakade, and M. Seeger. Gaussian process optimization in the bandit setting: No regret and experimental design. In Proc. of ICML-10, pages 1015–1022, 2010.
  27. 27.V. Strijov and G. Weber. Nonlinear regression model generation using hyperparameter optimization. Computers & Mathematics with Applications, 60(4):981–988, 2010.
  28. 28.L. Xu, H. H. Hoos, and K. Leyton-Brown. Hydra: Automatically configuring algorithms for portfolio-based selection. In Proc. of AAAI-10, pages 210–216, 2010.
  29. 29.P. Zhao and B. Yu. On model selection consistency of lasso. JMLR, 7:2541–2563, Dec. 2006.

Citation

MLA
Thornton, C., et al. “Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms”. arXiv, 2012, http://arxiv.org/abs/1208.3719v2.
APA
Thornton, C., Hutter, F., Hoos, H. H., & Leyton-Brown, K. (2012). Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms. arXiv. http://arxiv.org/abs/1208.3719v2
Chicago
Thornton, C., F. Hutter, H. H. Hoos, and K. Leyton-Brown. 2012. “Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms”. arXiv. http://arxiv.org/abs/1208.3719v2.
Harvard
Thornton, C. et al. (2012) “Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1208.3719v2.
Vancouver
1. Thornton C, Hutter F, Hoos HH, Leyton-Brown K (2012) Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms. arXiv

BibTeX

@article{thornton2012auto,
  title = {Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms},
  author = {Thornton, Chris and Hutter, Frank and Hoos, Holger H. and Leyton-Brown, Kevin},
  year = {2012},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1208.3719v2},
  eprint = {1208.3719}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF