AutoML: A Survey of the State-of-the-Art
Xin HeKaiyong ZhaoXiaowen Chu
Surveys state-of-the-art automated machine learning across the entire deep learning pipeline, offering direct performance comparisons of leading neural architecture search methods on CIFAR-10 and ImageNet alongside key open challenges.
Deep learning has driven breakthroughs across visual recognition and natural language processing, but designing high-performing models remains heavily reliant on costly, trial-and-error human expertise. Automated Machine Learning (AutoML) addresses this bottleneck by automating the construction and optimization of machine learning pipelines under constrained computational budgets. The article provides a comprehensive evaluation of the state of the art in AutoML, detailing advancements across data preparation, feature engineering, hyperparameter optimization, and neural architecture search (NAS).
The authors conducted a structured literature review and comparative analysis of representative AutoML techniques and search algorithms. Their evaluation synthesized methodological frameworks across each pipeline stage and benchmarked the empirical accuracy, search time, and hardware costs of leading architecture search algorithms on standard reference datasets, including CIFAR-10 and ImageNet.
The analysis reveals several core findings. First, efficiency-focused search paradigms have radically reduced computational overhead: early reinforcement learning and evolutionary techniques required upwards of 2,000 to 22,000 GPU days, whereas recent weight-sharing, continuous gradient-based, and one-stage methods complete searches in less than one GPU day (often under 0.5 GPU days) while matching or exceeding human-level accuracy. Second, randomized search baselines paired with weight-sharing perform competitively against complex search controllers, highlighting that search space design often dictates success more than the optimization heuristic itself. Third, while automated architectures achieve parity with or outperform human-designed systems in computer vision, a substantial performance gap remains in natural language processing. Fourth, common cell-based two-stage search routines suffer from a structural gap between shallow search architectures and deeper evaluation networks, causing performance rank inconsistencies across stages.
These findings indicate that AutoML, particularly efficient neural architecture search, has become commercially viable by dramatically lowering development timelines, hardware expenses, and specialized talent requirements. However, deploying automated methods introduces strategic trade-offs, as coupled optimization schemes and supernet weight-sharing can suffer from evaluation bias and catastrophic forgetting among candidate models. Furthermore, the lack of theoretical interpretability and high sensitivity to random seeds can create reproducibility and compliance risks in critical operating environments.
To move forward, organizations should adopt modern weight-sharing and gradient-based one-stage frameworks to optimize computing budgets while benchmarking new algorithms against standardized resources such as NAS-Bench. Future research and development must focus on designing flexible, bias-free search spaces, closing the performance gap in language modeling, jointly optimizing hyperparameters with model architectures, and integrating lifelong learning mechanisms to prevent catastrophic forgetting.
The findings are constrained by variations in experimental protocols, training hyperparameters, and random seed reporting across original studies, as well as the literature's prevailing reliance on clean, benchmark vision datasets. Confidence is high regarding the substantial efficiency gains and vision-task capabilities of modern AutoML methods, but stakeholders should exercise caution when extrapolating these results to noisy production data or non-vision domains without initial pilot evaluations.
- Paper: Neural Architecture Search with Reinforcement Learning, Barret Zoph et al. (2016). Introduces the foundational reinforcement learning framework for neural architecture search that underpins the NAS taxonomy and methods reviewed in the survey.
- Paper: Efficient and Robust Automated Machine Learning, Matthias Feurer et al. (2015). Pioneers automated end-to-end machine learning pipelines combining Bayesian optimization, meta-learning, and ensembling, establishing the classic AutoML system structure described in the survey.
- Paper: Learning Transferable Architectures for Scalable Image Recognition, Barret Zoph et al. (2018). Establishes the cell-based transferable search space on CIFAR-10 and ImageNet that serves as a standard benchmark throughout the survey's NAS performance analysis.
- Paper: ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware, Han Cai et al. (2018). Demonstrates path-binarized super-network optimization directly on target hardware, exemplifying the efficient one-shot and hardware-aware NAS strategies analyzed in the survey.
- Paper: MnasNet: Platform-Aware Neural Architecture Search for Mobile, Mingxing Tan et al. (2018). Integrates real-world device latency directly into multi-objective reinforcement learning search, forming a primary baseline for platform-aware NAS evaluated in the survey.
- Paper: Progressive Neural Architecture Search, Chenxi Liu et al. (2017). Introduces surrogate-guided progressive search to dramatically reduce NAS compute costs, directly informing the survey's coverage of efficient search strategies.
- Paper: A Tutorial on Bayesian Optimization, Peter I. Frazier (2018). Provides a comprehensive tutorial on Bayesian optimization and acquisition functions, which constitute the core mathematical foundations of hyperparameter optimization in AutoML.
- Paper: Algorithms for Hyper-Parameter Optimization, James Bergstra et al. (2011). Introduces the Tree-structured Parzen Estimator and sequential model-based optimization algorithms essential to modern automated hyperparameter tuning pipelines.
- Paper: Practical Bayesian Optimization of Machine Learning Algorithms, Jasper Snoek et al. (2012). Formalizes practical Bayesian optimization methodologies with Gaussian processes and acquisition functions for tuning complex machine learning algorithms.
- Paper: On Hyperparameter Optimization of Machine Learning Algorithms: Theory and Practice, Li Yang et al. (2020). Expands on the hyperparameter optimization stage of AutoML with an extensive theoretical and empirical taxonomy of optimization algorithms and tools.
- Paper: Knowledge Distillation: A Survey, Jianping Gou et al. (2020). Surveys knowledge distillation as a complementary technique for compressing and optimizing the complex architectures discovered by AutoML and NAS systems.
