AutoML: A Survey of the State-of-the-Art

Xin HeKaiyong ZhaoXiaowen Chu

article2019Knowledge-Based Systems1,881 citations

Surveys state-of-the-art automated machine learning across the entire deep learning pipeline, offering direct performance comparisons of leading neural architecture search methods on CIFAR-10 and ImageNet alongside key open challenges.

Listen

Deep learning has driven breakthroughs across visual recognition and natural language processing, but designing high-performing models remains heavily reliant on costly, trial-and-error human expertise. Automated Machine Learning (AutoML) addresses this bottleneck by automating the construction and optimization of machine learning pipelines under constrained computational budgets. The article provides a comprehensive evaluation of the state of the art in AutoML, detailing advancements across data preparation, feature engineering, hyperparameter optimization, and neural architecture search (NAS).

The authors conducted a structured literature review and comparative analysis of representative AutoML techniques and search algorithms. Their evaluation synthesized methodological frameworks across each pipeline stage and benchmarked the empirical accuracy, search time, and hardware costs of leading architecture search algorithms on standard reference datasets, including CIFAR-10 and ImageNet.

The analysis reveals several core findings. First, efficiency-focused search paradigms have radically reduced computational overhead: early reinforcement learning and evolutionary techniques required upwards of 2,000 to 22,000 GPU days, whereas recent weight-sharing, continuous gradient-based, and one-stage methods complete searches in less than one GPU day (often under 0.5 GPU days) while matching or exceeding human-level accuracy. Second, randomized search baselines paired with weight-sharing perform competitively against complex search controllers, highlighting that search space design often dictates success more than the optimization heuristic itself. Third, while automated architectures achieve parity with or outperform human-designed systems in computer vision, a substantial performance gap remains in natural language processing. Fourth, common cell-based two-stage search routines suffer from a structural gap between shallow search architectures and deeper evaluation networks, causing performance rank inconsistencies across stages.

These findings indicate that AutoML, particularly efficient neural architecture search, has become commercially viable by dramatically lowering development timelines, hardware expenses, and specialized talent requirements. However, deploying automated methods introduces strategic trade-offs, as coupled optimization schemes and supernet weight-sharing can suffer from evaluation bias and catastrophic forgetting among candidate models. Furthermore, the lack of theoretical interpretability and high sensitivity to random seeds can create reproducibility and compliance risks in critical operating environments.

To move forward, organizations should adopt modern weight-sharing and gradient-based one-stage frameworks to optimize computing budgets while benchmarking new algorithms against standardized resources such as NAS-Bench. Future research and development must focus on designing flexible, bias-free search spaces, closing the performance gap in language modeling, jointly optimizing hyperparameters with model architectures, and integrating lifelong learning mechanisms to prevent catastrophic forgetting.

The findings are constrained by variations in experimental protocols, training hyperparameters, and random seed reporting across original studies, as well as the literature's prevailing reliance on clean, benchmark vision datasets. Confidence is high regarding the substantial efficiency gains and vision-task capabilities of modern AutoML methods, but stakeholders should exercise caution when extrapolating these results to noisy production data or non-vision domains without initial pilot evaluations.

Cover for AutoML: A Survey of the State-of-the-Art

Abstract

Deep learning (DL) techniques have penetrated all aspects of our lives and brought us great convenience. However, building a high-quality DL system for a specific task highly relies on human expertise, hindering the applications of DL to more areas. Automated machine learning (AutoML) becomes a promising solution to build a DL system without human assistance, and a growing number of researchers focus on AutoML. In this paper, we provide a comprehensive and up-to-date review of the state-of-the-art (SOTA) in AutoML. First, we introduce AutoML methods according to the pipeline, covering data preparation, feature engineering, hyperparameter optimization, and neural architecture search (NAS). We focus more on NAS, as it is currently very hot sub-topic of AutoML. We summarize the performance of the representative NAS algorithms on the CIFAR-10 and ImageNet datasets and further discuss several worthy studying directions of NAS methods: one/two-stage NAS, one-shot NAS, and joint hyperparameter and architecture optimization. Finally, we discuss some open problems of the existing AutoML methods for future research.

Table of Contents

  • 1 Introduction
  • 2 Data Preparation
  • 2.1 Data Collection
  • 2.1.1 Data Searching
  • 2.1.2 Data Synthesis
  • 2.2 Data Cleaning
  • 2.3 Data Augmentation
  • 3 Feature Engineering
  • 3.1 Feature Selection
  • 3.2 Feature Construction
  • 3.3 Feature Extraction
  • 4 Model Generation
  • 4.1 Search Space
  • 4.1.1 Entire-structured Search Space
  • 4.1.2 Cell-based Search Space
  • 4.1.3 Hierarchical Search Space
  • 4.1.4 Morphism-based Search Space
  • 4.2 Architecture Optimization
  • 4.2.1 Evolutionary Algorithm
  • 4.2.2 Reinforcement Learning
  • 4.2.3 Gradient Descent
  • 4.2.4 Surrogate Model-based Optimization
  • 4.2.5 Grid and Random Search
  • 4.2.6 Hybrid Optimization Method
  • 4.3 Hyperparameter Optimization
  • 4.3.1 Grid and Random Search
  • 4.3.2 Bayesian Optimization
  • 4.3.3 Gradient-based Optimization
  • 5 Model Evaluation
  • 5.1 Low fidelity
  • 5.2 Weight sharing
  • 5.3 Surrogate
  • 5.4 Early stopping
  • 6 NAS Discussion
  • 6.1 NAS Performance Comparison
  • 6.1.1 Kendall Tau Metric
  • 6.1.2 NAS-Bench Dataset
  • 6.2 One-stage vs. Two-stage
  • 6.3 One-shot/Weight-sharing
  • 6.4 Joint Hyperparameter and Architecture Optimization
  • 6.5 Resource-aware NAS
  • 7 Open Problems and Future Directions
  • 7.1 Flexible Search Space
  • 7.2 Exploring More Areas
  • 7.3 Interpretability
  • 7.4 Reproducibility
  • 7.5 Robustness
  • 7.6 Joint Hyperparameter and Architecture Optimization
  • 7.7 Complete AutoML Pipeline
  • 7.8 Lifelong Learning
  • 7.8.1 Learn New Data
  • 7.8.2 Remember Old Knowledge
  • 8 Conclusions
  • References

Knowls

  1. Knowl 1 — Differentiable Architecture Search and Multi-Level Optimization Formulations

    model/method

    Differentiable Architecture Search (DARTS) relaxes the discrete candidate operation selection over a directed acyclic graph (DAG) into a continuous mixture of operations. For a pair of nodes (i,j)(i, j) with input feature tensor xx, the mixed operation oˉi,j(x)\bar{o}_{i,j}(x) over a set of KK predefined candidate operations O={o1,o2,…,oK}\mathcal{O} = \{o^1, o^2, \dots, o^K\} is defined as:

    oˉi,j(x)=∑k=1Kexp⁡(αi,jk)∑l=1Kexp⁡(αi,jl)ok(x)\bar{o}_{i,j}(x) = \sum_{k=1}^K \frac{\exp(\alpha_{i,j}^k)}{\sum_{l=1}^K \exp(\alpha_{i,j}^l)} o^k(x)

    where αi,jk∈R\alpha_{i,j}^k \in \mathbb{R} is the continuous architectural parameter assigned to operation oko^k on edge (i,j)(i, j). The search is framed as a bilevel optimization problem where the architecture parameters α\alpha and supernet weights θ\theta are alternately updated on the validation loss Lval\mathcal{L}_{\text{val}} and training loss Ltrain\mathcal{L}_{\text{train}}, respectively:

    min⁡αLval(θ∗(α),α)subject toθ∗(α)=arg⁡min⁡θLtrain(θ,α)\min_{\alpha} \mathcal{L}_{\text{val}}(\theta^*(\alpha), \alpha) \quad \text{subject to} \quad \theta^*(\alpha) = \arg\min_{\theta} \mathcal{L}_{\text{train}}(\theta, \alpha)

    To address the computational complexity and gradient estimation errors of bilevel optimization, two alternative formulations exist:

    1. Single-level optimization, which optimizes both parameters jointly on the training set:

    min⁡θ,αLtrain(θ,α)\min_{\theta, \alpha} \mathcal{L}_{\text{train}}(\theta, \alpha)

    1. Mixed-level optimization (as in MiLeNAS), which balances training and validation objectives using a regularization factor λ≥0\lambda \ge 0:

    min⁡α,θ[Ltrain(θ∗,α)+λLval(θ∗,α)]\min_{\alpha, \theta} \left[ \mathcal{L}_{\text{train}}(\theta^*, \alpha) + \lambda \mathcal{L}_{\text{val}}(\theta^*, \alpha) \right]

    To eliminate the linear GPU memory scaling of computing all candidate operations simultaneously, the Gumbel-Softmax reparameterization trick enables single-path sampling during forward propagation while remaining differentiable:

    oi,j(x)=∑k=1Kexp⁡((log⁡αi,jk+Gi,jk)/τ)∑l=1Kexp⁡((log⁡αi,jl+Gi,jl)/τ)ok(x)o_{i,j}(x) = \sum_{k=1}^K \frac{\exp\left( (\log \alpha_{i,j}^k + G_{i,j}^k) / \tau \right)}{\sum_{l=1}^K \exp\left( (\log \alpha_{i,j}^l + G_{i,j}^l) / \tau \right)} o^k(x)

    where Gi,jk=−log⁡(−log⁡(ui,jk))G_{i,j}^k = -\log(-\log(u_{i,j}^k)) is the kk-th Gumbel sample generated from an i.i.d. uniform random variable ui,jk∼Uniform(0,1)u_{i,j}^k \sim \text{Uniform}(0, 1), and τ>0\tau > 0 is the Softmax temperature parameter.

  2. Knowl 2 — Combinatorial Complexity Bounds for Entire-Structured versus Cell-Based Search Spaces

    equation

    In neural architecture search (NAS), a candidate network is modeled as a directed acyclic graph (DAG) of BB ordered nodes. The intermediate representation at node ZkZ_k (k∈{1,2,…,B}k \in \{1, 2, \dots, B\}) is computed by aggregating incoming edges:

    Zk=∑i=1Nkoi(Ii),oi∈OZ_k = \sum_{i=1}^{N_k} o_i(I_i), \quad o_i \in \mathcal{O}

    where NkN_k is the in-degree of node ZkZ_k, IiI_i is the ii-th input tensor, and O\mathcal{O} is the set of M=∣O∣M = |\mathcal{O}| candidate operations.

    For an entire-structured search space across an LL-layer deep network where each layer can connect to any preceding layer, the total number of candidate network architectures NentireN_{\text{entire}} is:

    Nentire=ML×2L(L−1)2N_{\text{entire}} = M^L \times 2^{\frac{L(L-1)}{2}}

    For a cell-based search space containing two repeating cell motifs (a normal cell and a reduction cell), each composed of BB ordered blocks where each block receives two inputs and applies two candidate operations combined via addition or concatenation, the total search space size NcellN_{\text{cell}} is:

    Ncell=(MB×(B+2)!)4N_{\text{cell}} = \left( M^B \times (B+2)! \right)^4

    Under typical hyperparameters (M=5M = 5 candidate operations, L=10L = 10 layers, and B=3B = 3 blocks per cell), the entire-structured space yields Nentire=510×245≈3.44×1020N_{\text{entire}} = 5^{10} \times 2^{45} \approx 3.44 \times 10^{20} possible configurations, whereas the cell-based space yields Ncell=(53×5!)4=(15000)4≈5.06×1016N_{\text{cell}} = (5^3 \times 5!)^4 = (15000)^4 \approx 5.06 \times 10^{16} configurations, demonstrating the combinatorial efficiency of repeating motif spaces.

  3. Knowl 3 — Sequential Model-Based Optimization for Black-Box AutoML

    algorithm

    Sequential Model-Based Optimization (SMBO) is an iterative paradigm used in hyperparameter optimization (HPO) and surrogate model-based neural architecture search (AO). It maintains a statistical surrogate model of an expensive objective function (such as network validation accuracy) to guide candidate selection.

    Input: Black-box objective evaluation function ff, candidate search space Θ\Theta, acquisition function SS, probabilistic surrogate model MM, total evaluation budget TT
    Output: Optimal configuration θ∗∈Θ\theta^* \in \Theta with minimum loss / maximum performance
    D←InitSamples(f,Θ)D \leftarrow \text{InitSamples}(f, \Theta)
    for i=1i = 1 to TT do
        p(y∣θ,D)←FitModel(M,D)p(y|\theta, D) \leftarrow \text{FitModel}(M, D)
        θi←arg⁡max⁡θ∈ΘS(θ,p(y∣θ,D))\theta_i \leftarrow \arg\max_{\theta \in \Theta} S(\theta, p(y|\theta, D))
        yi←f(θi)y_i \leftarrow f(\theta_i)
        D←D∪{(θi,yi)}D \leftarrow D \cup \{(\theta_i, y_i)\}
    end for
    θ∗←arg⁡max⁡(θ,y)∈Dy\theta^* \leftarrow \arg\max_{(\theta, y) \in D} y
    return θ∗\theta^*

    The algorithm operates as follows:

    1. Initialize sample dataset D={(θk,yk)}D = \{(\theta_k, y_k)\} with a set of evaluated configurations where θk∈Θ\theta_k \in \Theta and yk=f(θk)y_k = f(\theta_k).
    2. In each iteration i∈{1,…,T}i \in \{1, \dots, T\}, train or update the probabilistic surrogate model MM (e.g., Gaussian Process, Random Forest, Tree-structured Parzen Estimator, or an LSTM/MLP predictor) on DD to estimate the conditional distribution p(y∣θ,D)p(y|\theta, D).
    3. Optimize the acquisition function SS (e.g., Expected Improvement) over the search space Θ\Theta to select the next configuration θi\theta_i.
    4. Evaluate θi\theta_i with the expensive black-box function ff (which involves training and validating the network).
    5. Append the resulting record (θi,yi)(\theta_i, y_i) to DD.
  4. Knowl 4 — Resource-Constrained and Differentiable Hardware-Aware NAS Objectives

    model/method

    To trade off validation accuracy against device latency, memory footprint, or floating-point operations (FLOPs), hardware-aware NAS integrates resource constraints directly into the search objective.

    In reinforcement learning and multi-objective evolutionary search (e.g., MnasNet), a customized weighted product objective approximates Pareto optimality:

    max⁡mACC(m)×[LAT(m)T]w\max_{m} ACC(m) \times \left[ \frac{LAT(m)}{T} \right]^w

    where ACC(m)ACC(m) is the validation accuracy of candidate model mm, LAT(m)LAT(m) is the measured or estimated inference latency on the target hardware platform, TT is the user-specified target latency constraint, and ww is a penalty scaling factor defined piecewise as:

    w={α,if LAT(m)≤Tβ,otherwisew = \begin{cases} \alpha, & \text{if } LAT(m) \le T \\ \beta, & \text{otherwise} \end{cases}

    with recommended parameter values α=β=−0.07\alpha = \beta = -0.07.

    In differentiable neural architecture search (e.g., FBNet), where hardware constraints must be differentiable with respect to architecture parameters, an operator latency lookup table is combined with cross-entropy loss CE(a,θa)CE(a, \theta_a):

    L(a,θa)=CE(a,θa)⋅αlog⁡(LAT(a))β\mathcal{L}(a, \theta_a) = CE(a, \theta_a) \cdot \alpha \log(LAT(a))^\beta

    where aa represents the architecture parameters, θa\theta_a are the model weights, LAT(a)LAT(a) is the total latency computed by summing operator-level lookup latencies, and α>0\alpha > 0 and β>0\beta > 0 are hyperparameter coefficients controlling total loss magnitude and latency sensitivity.

  5. Knowl 5 — Coupled versus Decoupled Supernet Optimization in One-Shot NAS

    model/method

    One-shot neural architecture search embeds the entire search space into a single overparameterized supernet where all candidate sub-architectures share a single pool of weights. One-shot methods are categorized by how they optimize the architecture parameters α\alpha and supernet weights θ\theta:

    1. Coupled Optimization (e.g., ENAS, DARTS, GDAS, ProxylessNAS): Architecture selection and weight updates are performed simultaneously or in an alternating bilevel loop during a single training run.

      • Limitations: Creates an optimization bias where sub-networks that converge rapidly receive more gradient updates, starving slower-converging but potentially superior architectures. Furthermore, continuously training newly sampled sub-networks leads to 'multi-model forgetting', where previously learned weights degrade across steps.
    2. Decoupled Optimization (e.g., SPOS, FairNAS, SETN, Single-Path NAS): Separates the process into two sequential phases:

      • Phase 1 (Supernet Training): The supernet weights are trained using uniform random sampling, fair path sampling (activating each candidate operator an equal number of times per batch), or path dropout (randomly dropping active paths with increasing probability). This reduces inter-architecture co-adaptation and ensures fair weight estimation without optimizing architecture parameters.
      • Phase 2 (Architecture Search): The trained supernet acts purely as a frozen, zero-training performance estimator. Search algorithms (such as Evolutionary Algorithms, Random Search, or surrogate predictors) evaluate and rank sub-networks via inference only on validation data.
  6. Knowl 6 — Kendall Rank Correlation Metric for Evaluating NAS Search Fidelity

    equation

    In neural architecture search, proxy metrics (such as early stopping, reduced channel size, reduced dataset resolution, or shared-weight performance) are used to rank candidate networks during search. The fidelity of these proxy rankings relative to true standalone training performance is quantified using the Kendall rank correlation coefficient τ∈[−1,1]\tau \in [-1, 1]:

    τ=NC−NDNC+ND\tau = \frac{N_C - N_D}{N_C + N_D}

    where NCN_C is the number of concordant pairs (pairs of architectures whose relative performance order is preserved between proxy search evaluation and standalone retraining) and NDN_D is the number of discordant pairs (pairs whose relative performance order is inverted).

    The values of τ\tau correspond to the following search conditions:

    • τ=1\tau = 1: The proxy evaluation perfectly predicts true standalone rankings.
    • τ=−1\tau = -1: The proxy ranking is completely inverted relative to true standalone performance.
    • τ=0\tau = 0: The proxy ranking exhibits no statistical relationship with true standalone performance.
  7. Knowl 7 — Symmetrized Kullback-Leibler Divergence for Architecture Ranking in One-Shot NAS

    equation

    In decoupled one-shot architecture search, candidate subnetworks can be ranked for validation quality without full evaluation by measuring the divergence between the output probability distribution of the sampled candidate subnetwork p=(p1,p2,…,pn)p = (p_1, p_2, \dots, p_n) and that of the full one-shot supernet q=(q1,q2,…,qn)q = (q_1, q_2, \dots, q_n) across nn classes. This is calculated using the symmetrized Kullback-Leibler (KL) divergence DSKLD_{\text{SKL}}:

    DSKL=DKL(p∥q)+DKL(q∥p)D_{\text{SKL}} = D_{\text{KL}}(p \parallel q) + D_{\text{KL}}(q \parallel p)

    where the standard discrete Kullback-Leibler divergence DKL(p∥q)D_{\text{KL}}(p \parallel q) is defined as:

    DKL(p∥q)=∑i=1npilog⁡piqiD_{\text{KL}}(p \parallel q) = \sum_{i=1}^n p_i \log \frac{p_i}{q_i}

    Architectures exhibiting smaller DSKLD_{\text{SKL}} values relative to the parent supernet preserve functional agreement and achieve higher standalone validation accuracy. Evaluating DSKLD_{\text{SKL}} requires only a small batch of training data (e.g., 64 samples), providing an extremely fast proxy for architecture selection.

  8. Knowl 8 — Benchmark Performance and Search Cost of Representative NAS Methods on CIFAR-10

    data/table

    The table below compares representative Neural Architecture Search algorithms across architecture optimization (AO) paradigms on the CIFAR-10 image classification benchmark against manually designed convolutional networks.

    Method Published #Params (M) Top-1 Acc (%) GPU Days AO Category
    ResNet-110 ECCV16 1.7 93.57 - Manual
    DenseNet CVPR17 25.6 96.54 - Manual
    AmoebaNet-B + c/o AAAI19 34.9 97.87 3,150 Evolutionary (EA)
    EENA ICCV19 8.47 97.44 0.65 Evolutionary (EA)
    NASNet-A + c/o CVPR18 3.3 97.35 2,000 Reinforcement Learning (RL)
    ENAS + micro + c/o ICML18 4.6 97.11 0.45 Reinforcement Learning (RL)
    DARTS (second order) + c/o ICLR19 3.3 97.23 4.0 Gradient Descent (GD)
    P-DARTS + c/o ICCV19 3.4 97.50 0.3 Gradient Descent (GD)
    PC-DARTS + c/o CVPR20 3.6 97.43 0.1 Gradient Descent (GD)
    PNAS ECCV18 3.2 96.59 225 Surrogate (SMBO)
    RandomNAS UAI19 4.3 97.15 2.7 Random Search (RS)
    CARS CVPR20 3.6 97.38 0.4 Hybrid (EA+GD)

    Notes:

    • + c/o denotes models trained with Cutout data augmentation.
    • GPU Days measures search efficiency computed as N×DN \times D, where NN is the number of GPUs used and DD is elapsed search days.
    • The empirical progression demonstrates that weight-sharing and differentiable search (e.g., ENAS, P-DARTS, PC-DARTS, CARS) reduced search costs from thousands of GPU days (AmoebaNet, NASNet) to under 1 GPU day while achieving comparable or superior top-1 accuracy (>97.4%>97.4\%).
  9. Knowl 9 — Benchmark Performance and Efficiency of Representative NAS Methods on ImageNet

    data/table

    The table below compares the performance, parameter complexity, and search overhead of representative automated architecture search methods transferred to or searched directly on the ImageNet dataset.

    Method Published #Params (M) Top-1 / Top-5 Acc (%) GPU Days AO Category
    ResNet-152 CVPR16 230 70.62 / 95.51 - Manual
    MobileNetV2 CVPR18 6.9 74.70 / - - Manual
    AmoebaNet-C AAAI19 6.4 75.70 / 92.40 3,150 Evolutionary (EA)
    GreedyNAS CVPR20 6.5 77.10 / 93.30 1.0 Evolutionary (EA)
    NASNet-A (6@4032) ICLR17 88.9 82.70 / 96.20 2,000 Reinforcement Learning (RL)
    MnasNet CVPR19 5.2 76.70 / 93.30 1,666 Reinforcement Learning (RL)
    DARTS (searched on CIFAR-10) ICLR19 4.7 73.30 / 81.30 4.0 Gradient Descent (GD)
    P-DARTS ICCV19 4.9 75.60 / 92.60 0.3 Gradient Descent (GD)
    PC-DARTS CVPR20 5.3 75.80 / 92.70 3.8 Gradient Descent (GD)
    PNAS-5 ECCV18 5.1 74.20 / 91.90 225 Surrogate (SMBO)
    SemiNAS CVPR20 6.32 76.50 / 93.20 4.0 Surrogate (SMBO)
    CARS CVPR20 5.1 75.20 / 92.50 0.4 Hybrid (EA+GD)

    Notes:

    • The results highlight two main paradigms: (1) proxy-transferred architectures (e.g., DARTS, PNAS), which search on CIFAR-10 and transfer the discovered cell to ImageNet; and (2) direct target search methods (e.g., MnasNet, PC-DARTS, GreedyNAS), which optimize directly for target data resolution and hardware latency constraints.
  10. Knowl 10 — Two-Stage Search-Evaluation Gap versus One-Stage Direct Deployment NAS

    model/method

    NAS pipelines are distinguished by their stage workflow:

    1. Two-Stage NAS: Comprises a searching stage followed by an independent evaluation/retraining stage. During the search stage, candidate cell structures are searched on shallow networks (e.g., 8 cells in DARTS) to minimize GPU memory consumption. In the evaluation stage, the discovered optimal cell is stacked into a substantially deeper network (e.g., 20 cells in DARTS) and retrained from scratch.

      • Depth Gap Limitation: The shallow architecture optimal for the proxy search space is often suboptimal when expanded to deeper networks. Progressive DARTS (P-DARTS) bridges this gap by dividing the search phase into consecutive multi-stage steps, progressively increasing network depth (e.g., 5 to 11 to 17 cells) while shrinking the candidate operation set (e.g., from 5 to 3 to 2 operations) via search space approximation.
    2. One-Stage NAS (e.g., Once-for-All, BigNAS, AtomNAS, DSNAS): Unifies architecture optimization and parameter training into a single end-to-end process. The supernet is trained such that all constituent sub-networks achieve fully converged, high-performing weights simultaneously (often via progressive shrinking or in-place distillation). Once training concludes, specialized sub-networks tailored to specific latency or memory constraints can be extracted and deployed immediately without additional retraining or fine-tuning.

  11. Knowl 11 — Joint Hyperparameter and Architecture Optimization via Basis Discretization

    model/method

    Traditional NAS methods evaluate candidate architectures using a fixed, manually chosen set of training hyperparameters (e.g., learning rate, weight decay, optimizer type). However, different neural architectures achieve optimal performance under distinct hyperparameter configurations, introducing bias into architecture ranking. Joint Hyperparameter and Architecture Optimization (HAO) formulates the search space as the Cartesian product of architecture choices and hyperparameter choices.

    To apply continuous differentiable search methods (such as AutoHAS) across mixed search spaces containing discrete architectural choices (e.g., layer operations) and continuous hyperparameters (e.g., learning rate lrlr), continuous hyperparameters are discretized into a linear combination of categorical base values:

    lr=∑k=1Kwk⋅bklr = \sum_{k=1}^K w_k \cdot b_k

    where {b1,b2,…,bK}\{b_1, b_2, \dots, b_K\} is a predefined discrete set of hyperparameter base values (such as {0.1,0.2,0.3}\{0.1, 0.2, 0.3\}) and {w1,w2,…,wK}\{w_1, w_2, \dots, w_K\} are continuous architectural weights normalized via Softmax (∑k=1Kwk=1\sum_{k=1}^K w_k = 1, wk≥0w_k \ge 0). This parameterization enables joint gradient-based optimization of continuous training hyperparameters and discrete structural choices.

  12. Knowl 12 — Structural Taxonomy of the End-to-End Automated Machine Learning Pipeline

    definition

    An end-to-end Automated Machine Learning (AutoML) pipeline consists of four modular processes structured to construct high-performing machine learning systems under a fixed computational budget without human manual tuning:

    1. Data Preparation: Automates the ingestion, curation, and expansion of raw data via:
      • Data Collection/Synthesis: Web scraping, self-labeling semi-supervised techniques, and simulator/GAN-based synthetic generation.
      • Data Cleaning: Automated missing value imputation, noise detection, and pipeline generation (e.g., BoostClean, AlphaClean).
      • Data Augmentation: Policy search for automated affine, elastic, and neural transformations (e.g., AutoAugment, Fast AutoAugment).
    2. Feature Engineering: Maximizes signal extraction from raw tabular, image, or text representations through:
      • Feature Selection: Filter, wrapper, and embedded subset evaluations.
      • Feature Construction: Non-linear attribute generation (Cartesian products, decision trees, genetic feature programming).
      • Feature Extraction: Unsupervised and dimensionality-reducing projections (PCA, autoencoders).
    3. Model Generation: Defines and explores the model configuration space via:
      • Search Space Design: Traditional models (SVM, KNN) vs. Deep Neural Network spaces (entire-structured, cell-based, hierarchical, morphism-based).
      • Optimization Methods: Hyperparameter Optimization (HPO) and Architecture Optimization (AO) using Grid/Random Search, Evolutionary Algorithms, Reinforcement Learning, Bayesian Optimization, and Gradient Descent.
    4. Model Evaluation: Accelerates the validation of candidate models to reduce computational overhead using low-fidelity approximations (sub-sampling, downscaled resolution), weight-sharing supernets, surrogate predictors (PNAS, SemiNAS), and early-stopping learning curve extrapolators.

Coverage note — Deliberately omitted individual specialized manual data cleaning heuristic systems (e.g., Katara, SampleClean) and generic open-source library listings (e.g., lists of tool URLs), focusing instead on the conceptual methodologies, formal mathematical models, optimization algorithms, and comparative empirical benchmarks that constitute the paper's primary analytical survey contribution.

References

  1. 1.A. Krizhevsky, I. Sutskever, G. E. Hinton, Imagenet classification with deep convolutional neural networks, in: P. L. Bartlett, F. C. N. Pereira, C. J. C. Burges, L. Bottou, K. Q. Weinberger (Eds.), Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting held December 3-6, 2012, Lake Tahoe, Nevada, United States, 2012, pp. 1106–1114. URL https://proceedings.neurips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html
  2. 2.K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, IEEE Computer Society, 2016, pp. 770–778. doi:10.1109/CVPR.2016.90. URL https://doi.org/10.1109/CVPR.2016.90
  3. 3.J. Redmon, S. K. Divvala, R. B. Girshick, A. Farhadi, You only look once: Unified, real-time object detection, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, IEEE Computer Society, 2016, pp. 779–788. doi:10.1109/CVPR.2016.91. URL https://doi.org/10.1109/CVPR.2016.91
  4. 4.C. Gong, D. He, X. Tan, T. Qin, L. Wang, T. Liu, FRAGE: frequency-agnostic word representation, in: S. Bengio, H. M. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, R. Garnett (Eds.), Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montr´eal, Canada, 2018, pp. 1341–1352. URL https://proceedings.neurips.cc/paper/2018/hash/e555ebe0ce426f7f9b2bef0706315e0c-Abstract.html
  5. 5.Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le, R. Salakhutdinov, Transformer-XL: Attentive language models beyond a fixed-length context, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, Florence, Italy, 2019, pp. 2978–2988. doi:10.18653/v1/P19-1285. URL https://www.aclweb.org/anthology/P19-1285
  6. 6.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, L. Fei-Fei, ImageNet Large Scale Visual Recognition Challenge, International Journal of Computer Vision (IJCV) 115 (3) (2015) 211–252. doi:10.1007/s11263-015-0816-y.
  7. 7.K. Simonyan, A. Zisserman, Very deep convolutional networks for large-scale image recognition, in: Y. Bengio, Y. LeCun (Eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015. URL http://arxiv.org/abs/1409.1556
  8. 8.M. Zoller, M. F. Huber, Benchmark and survey of automated machine learning frameworks, arXiv preprint arXiv:1904.12054.
  9. 9.Q. Yao, M. Wang, Y. Chen, W. Dai, H. Yi-Qi, L. Yu-Feng, T. Wei-Wei, Y. Qiang, Y. Yang, Taking human out of learning applications: A survey on automated machine learning, arXiv preprint arXiv:1810.13306.
  10. 10.T. Elsken, J. H. Metzen, F. Hutter, Neural architecture search: A survey, arXiv preprint arXiv:1808.05377.
  11. 11.K. Yu, C. Sciuto, M. Jaggi, C. Musat, M. Salzmann, Evaluating the search phase of neural architecture search, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=H1loF2NFwr
  12. 12.B. Zoph, Q. V. Le, Neural architecture search with reinforcement learning, in: 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings, OpenReview.net, 2017. URL https://openreview.net/forum?id=r1Ue8Hcxg
  13. 13.H. Pham, M. Y. Guan, B. Zoph, Q. V. Le, J. Dean, Efficient neural architecture search via parameter sharing, in: J. G. Dy, A. Krause (Eds.), Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm¨assan, Stockholm, Sweden, July 10-15, 2018, Vol. 80 of Proceedings of Machine Learning Research, PMLR, 2018, pp. 4092–4101. URL http://proceedings.mlr.press/v80/pham18a.html
  14. 14.A. Brock, T. Lim, J. M. Ritchie, N. Weston, SMASH: one-shot model architecture search through hypernetworks, in: 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings, OpenReview.net, 2018. URL https://openreview.net/forum?id=rydeCEhs-
  15. 15.B. Zoph, V. Vasudevan, J. Shlens, Q. V. Le, Learning transferable architectures for scalable image recognition, in: 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, IEEE Computer Society, 2018, pp. 8697–8710. doi:10.1109/CVPR.2018.00907. URL http://openaccess.thecvf.com/content_cvpr_2018/html/Zoph_Learning_Transferable_Architectures_CVPR_2018_paper.html
  16. 16.Z. Zhong, J. Yan, W. Wu, J. Shao, C. Liu, Practical block-wise neural network architecture generation, in: 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, IEEE Computer Society, 2018, pp. 2423–2432. doi:10.1109/CVPR.2018.00257. URL http://openaccess.thecvf.com/content_cvpr_2018/html/Zhong_Practical_Block-Wise_Neural_CVPR_2018_paper.html
  17. 17.H. Liu, K. Simonyan, Y. Yang, DARTS: differentiable architecture search, in: 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, OpenReview.net, 2019. URL https://openreview.net/forum?id=S1eYHoC5FX
  18. 18.C. Liu, B. Zoph, M. Neumann, J. Shlens, W. Hua, L.-J. Li, L. Fei-Fei, A. Yuille, J. Huang, K. Murphy, Progressive neural architecture search (2018) 19–34.
  19. 19.H. Liu, K. Simonyan, O. Vinyals, C. Fernando, K. Kavukcuoglu, Hierarchical representations for efficient architecture search, in: 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings, OpenReview.net, 2018. URL https://openreview.net/forum?id=BJQRKzbA-
  20. 20.T. Chen, I. J. Goodfellow, J. Shlens, Net2net: Accelerating learning via knowledge transfer, in: Y. Bengio, Y. LeCun (Eds.), 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, 2016. URL http://arxiv.org/abs/1511.05641
  21. 21.T. Wei, C. Wang, Y. Rui, C. W. Chen, Network morphism, in: M. Balcan, K. Q. Weinberger (Eds.), Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, Vol. 48 of JMLR Workshop and Conference Proceedings, JMLR.org, 2016, pp. 564–572. URL http://proceedings.mlr.press/v48/wei16.html
  22. 22.H. Jin, Q. Song, X. Hu, Auto-keras: An efficient neural architecture search system, in: A. Teredesai, V. Kumar, Y. Li, R. Rosales, E. Terzi, G. Karypis (Eds.), Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK, USA, August 4-8, 2019, ACM, 2019, pp. 1946–1956. doi:10.1145/3292500.3330648. URL https://doi.org/10.1145/3292500.3330648
  23. 23.B. Baker, O. Gupta, N. Naik, R. Raskar, Designing neural network architectures using reinforcement learning, in: 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings, OpenReview.net, 2017. URL https://openreview.net/forum?id=S1c2cvqee
  24. 24.K. O. Stanley, R. Miikkulainen, Evolving neural networks through augmenting topologies, Evolutionary computation 10 (2) (2002) 99–127.
  25. 25.E. Real, S. Moore, A. Selle, S. Saxena, Y. L. Suematsu, J. Tan, Q. V. Le, A. Kurakin, Large-scale evolution of image classifiers, in: D. Precup, Y. W. Teh (Eds.), Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, Vol. 70 of Proceedings of Machine Learning Research, PMLR, 2017, pp. 2902–2911. URL http://proceedings.mlr.press/v70/real17a.html
  26. 26.E. Real, A. Aggarwal, Y. Huang, Q. V. Le, Regularized evolution for image classifier architecture search, in: The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, AAAI Press, 2019, pp. 4780–4789. doi:10.1609/aaai.v33i01.33014780. URL https://doi.org/10.1609/aaai.v33i01.33014780
  27. 27.T. Elsken, J. H. Metzen, F. Hutter, Efficient multi-objective neural architecture search via lamarckian evolution, in: 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, OpenReview.net, 2019. URL https://openreview.net/forum?id=ByME42AqK7
  28. 28.M. Suganuma, S. Shirakawa, T. Nagao, A genetic programming approach to designing convolutional neural network architectures, in: J. Lang (Ed.), Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden, ijcai.org, 2018, pp. 5369–5373. doi:10.24963/ijcai.2018/755. URL https://doi.org/10.24963/ijcai.2018/755
  29. 29.R. Miikkulainen, J. Liang, E. Meyerson, A. Rawal, D. Fink, O. Francon, B. Raju, H. Shahrzad, A. Navruzyan, N. Duffy, et al., Evolving deep neural networks (2019) 293–312.
  30. 30.L. Xie, A. L. Yuille, Genetic CNN, in: IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, IEEE Computer Society, 2017, pp. 1388–1397. doi:10.1109/ICCV.2017.154. URL https://doi.org/10.1109/ICCV.2017.154
  31. 31.K. Ahmed, L. Torresani, Maskconnect: Connectivity learning by gradient descent (2018) 349–365.
  32. 32.R. Shin, C. Packer, D. Song, Differentiable neural network architecture search.
  33. 33.H. Mendoza, A. Klein, M. Feurer, J. T. Springenberg, F. Hutter, Towards automatically-tuned neural networks (2016) 58–65.
  34. 34.A. Zela, A. Klein, S. Falkner, F. Hutter, Towards automated deep learning: Efficient joint neural architecture and hyperparameter search, arXiv preprint arXiv:1807.06906.
  35. 35.A. Klein, S. Falkner, S. Bartels, P. Hennig, F. Hutter, Fast bayesian optimization of machine learning hyperparameters on large datasets, in: A. Singh, X. J. Zhu (Eds.), Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, AISTATS 2017, 20-22 April 2017, Fort Lauderdale, FL, USA, Vol. 54 of Proceedings of Machine Learning Research, PMLR, 2017, pp. 528–536. URL http://proceedings.mlr.press/v54/klein17a.html
  36. 36.S. Falkner, A. Klein, F. Hutter, Practical hyperparameter optimization for deep learning.
  37. 37.F. Hutter, H. H. Hoos, K. Leyton-Brown, Sequential model-based optimization for general algorithm configuration, in: International conference on learning and intelligent optimization, 2011, pp. 507–523.
  38. 38.S. Falkner, A. Klein, F. Hutter, BOHB: robust and efficient hyperparameter optimization at scale, in: J. G. Dy, A. Krause (Eds.), Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm¨assan, Stockholm, Sweden, July 10-15, 2018, Vol. 80 of Proceedings of Machine Learning Research, PMLR, 2018, pp. 1436–1445. URL http://proceedings.mlr.press/v80/falkner18a.html
  39. 39.J. Bergstra, D. Yamins, D. D. Cox, Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures, in: Proceedings of the 30th International Conference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16-21 June 2013, Vol. 28 of JMLR Workshop and Conference Proceedings, JMLR.org, 2013, pp. 115–123. URL http://proceedings.mlr.press/v28/bergstra13.html
  40. 40.Z. Yang, Y. Wang, X. Chen, B. Shi, C. Xu, C. Xu, Q. Tian, C. Xu, CARS: continuous evolution for efficient neural architecture search, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 1826–1835. doi:10.1109/CVPR42600.2020.00190. URL https://doi.org/10.1109/CVPR42600.2020.00190
  41. 41.K. Maziarz, M. Tan, A. Khorlin, M. Georgiev, A. Gesmundo, Evolutionary-neural hybrid agents for architecture searcharXiv:1811.09828.
  42. 42.Y. Chen, G. Meng, Q. Zhang, S. Xiang, C. Huang, L. Mu, X. Wang, Reinforced evolutionary neural architecture search, arXiv preprint arXiv:1808.00193.
  43. 43.Y. Sun, H. Wang, B. Xue, Y. Jin, G. G. Yen, M. Zhang, Surrogate-assisted evolutionary deep learning using an end-to-end random forest-based performance predictor, IEEE Transactions on Evolutionary Computation.
  44. 44.B. Wang, Y. Sun, B. Xue, M. Zhang, A hybrid differential evolution approach to designing deep convolutional neural networks for image classification, in: Australasian Joint Conference on Artificial Intelligence, Springer, 2018, pp. 237–250.
  45. 45.M. Wistuba, A. Rawat, T. Pedapati, A survey on neural architecture search, arXiv preprint arXiv:1905.01392.
  46. 46.P. Ren, Y. Xiao, X. Chang, P.-Y. Huang, Z. Li, X. Chen, X. Wang, A comprehensive survey of neural architecture search: Challenges and solutions (2020). arXiv:2006.02903.
  47. 47.R. Elshawi, M. Maher, S. Sakr, Automated machine learning: State-of-the-art and open challenges, arXiv preprint arXiv:1906.02287.
  48. 48.Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, Gradient-based learning applied to document recognition, Proceedings of the IEEE 86 (11) (1998) 2278–2324.
  49. 49.A. Krizhevsky, V. Nair, G. Hinton, The cifar-10 dataset, online: http://www. cs. toronto. edu/kriz/cifar. html.
  50. 50.J. Deng, W. Dong, R. Socher, L. Li, K. Li, F. Li, Imagenet: A large-scale hierarchical image database, in: 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA, IEEE Computer Society, 2009, pp. 248–255. doi:10.1109/CVPR.2009.5206848. URL https://doi.org/10.1109/CVPR.2009.5206848
  51. 51.J. Yang, X. Sun, Y.-K. Lai, L. Zheng, M.-M. Cheng, Recognition from web data: a progressive filtering approach, IEEE Transactions on Image Processing 27 (11) (2018) 5303–5315.
  52. 52.X. Chen, A. Shrivastava, A. Gupta, NEIL: extracting visual knowledge from web data, in: IEEE International Conference on Computer Vision, ICCV 2013, Sydney, Australia, December 1-8, 2013, IEEE Computer Society, 2013, pp. 1409–1416. doi:10.1109/ICCV.2013.178. URL https://doi.org/10.1109/ICCV.2013.178
  53. 53.Y. Xia, X. Cao, F. Wen, J. Sun, Well begun is half done: Generating high-quality seeds for automatic image dataset construction from web, in: European Conference on Computer Vision, Springer, 2014, pp. 387–400.
  54. 54.N. H. Do, K. Yanai, Automatic construction of action datasets using web videos with density-based cluster analysis and outlier detection, in: Pacific-Rim Symposium on Image and Video Technology, Springer, 2015, pp. 160–172.
  55. 55.J. Krause, B. Sapp, A. Howard, H. Zhou, A. Toshev, T. Duerig, J. Philbin, L. Fei-Fei, The unreasonable effectiveness of noisy data for fine-grained recognition, in: European Conference on Computer Vision, Springer, 2016, pp. 301–320.
  56. 56.P. D. Vo, A. Ginsca, H. Le Borgne, A. Popescu, Harnessing noisy web images for deep representation, Computer Vision and Image Understanding 164 (2017) 68–81.
  57. 57.B. Collins, J. Deng, K. Li, L. Fei-Fei, Towards scalable dataset construction: An active learning approach, in: European conference on computer vision, Springer, 2008, pp. 86–98.
  58. 58.Y. Roh, G. Heo, S. E. Whang, A survey on data collection for machine learning: a big data-ai integration perspective, IEEE Transactions on Knowledge and Data Engineering.
  59. 59.D. Yarowsky, Unsupervised word sense disambiguation rivaling supervised methods, in: 33rd Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, Cambridge, Massachusetts, USA, 1995, pp. 189–196. doi:10.3115/981658.981684. URL https://www.aclweb.org/anthology/P95-1026
  60. 60.I. Triguero, J. A. S´aez, J. Luengo, S. Garc´ıa, F. Herrera, On the characterization of noise filters for self-training semi-supervised in nearest neighbor classification, Neurocomputing 132 (2014) 30–41.
  61. 61.M. F. A. Hady, F. Schwenker, Combining committee-based semi-supervised learning and active learning, Journal of Computer Science and Technology 25 (4) (2010) 681–698.
  62. 62.A. Blum, T. Mitchell, Combining labeled and unlabeled data with co-training, in: Proceedings of the eleventh annual conference on Computational learning theory, ACM, 1998, pp. 92–100.
  63. 63.Y. Zhou, S. Goldman, Democratic co-learning, in: Tools with Artificial Intelligence, 2004. ICTAI 2004. 16th IEEE International Conference on, IEEE, 2004, pp. 594–602.
  64. 64.X. Chen, A. Gupta, Webly supervised learning of convolutional networks, in: 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, IEEE Computer Society, 2015, pp. 1431–1439. doi:10.1109/ICCV.2015.168. URL https://doi.org/10.1109/ICCV.2015.168
  65. 65.Z. Xu, S. Huang, Y. Zhang, D. Tao, Augmenting strong supervision using web data for fine-grained categorization, in: 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, IEEE Computer Society, 2015, pp. 2524–2532. doi:10.1109/ICCV.2015.290. URL https://doi.org/10.1109/ICCV.2015.290
  66. 66.N. V. Chawla, K. W. Bowyer, L. O. Hall, W. P. Kegelmeyer, Smote: synthetic minority over-sampling technique, Journal of artificial intelligence research 16 (2002) 321–357.
  67. 67.H. Guo, H. L. Viktor, Learning from imbalanced data sets with boosting and data generation: the databoost-im approach, ACM Sigkdd Explorations Newsletter 6 (1) (2004) 30–39.
  68. 68.G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, W. Zaremba, Openai gym, arXiv preprint arXiv:1606.01540.
  69. 69.Q. Wang, S. Zheng, Q. Yan, F. Deng, K. Zhao, X. Chu, Irs: A large synthetic indoor robotics stereo dataset for disparity and surface normal estimation, arXiv preprint arXiv:1912.09678.
  70. 70.N. Ruiz, S. Schulter, M. Chandraker, Learning to simulate, in: 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, OpenReview.net, 2019. URL https://openreview.net/forum?id=HJgkx2Aqt7
  71. 71.I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, Y. Bengio, Generative adversarial nets, in: Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, K. Q. Weinberger (Eds.), Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada, 2014, pp. 2672–2680. URL https://proceedings.neurips.cc/paper/2014/hash/5ca3e9b122f61f8f06494c97b1afccf3-Abstract.html
  72. 72.T.-H. Oh, R. Jaroensri, C. Kim, M. Elgharib, F. Durand, W. T. Freeman, W. Matusik, Learning-based video motion magnification, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 633–648.
  73. 73.L. Sixt, Rendergan: Generating realistic labeled data–with an application on decoding bee tags, unpublished Bachelor Thesis, Freie Universit¨at, Berlin.
  74. 74.C. Bowles, L. Chen, R. Guerrero, P. Bentley, R. Gunn, A. Hammers, D. A. Dickie, M. V. Hern´andez, J. Wardlaw, D. Rueckert, Gan augmentation: Augmenting training data using generative adversarial networks, arXiv preprint arXiv:1810.10863.
  75. 75.N. Park, M. Mohammadi, K. Gorde, S. Jajodia, H. Park, Y. Kim, Data synthesis based on generative adversarial networks, Proceedings of the VLDB Endowment 11 (10) (2018) 1071–1083.
  76. 76.L. Xu, K. Veeramachaneni, Synthesizing tabular data using generative adversarial networks, arXiv preprint arXiv:1811.11264.
  77. 77.D. Donahue, A. Rumshisky, Adversarial text generation without reinforcement learning, arXiv preprint arXiv:1810.06640.
  78. 78.T. Karras, S. Laine, T. Aila, A style-based generator architecture for generative adversarial networks, in: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, Computer Vision Foundation / IEEE, 2019, pp. 4401–4410. doi:10.1109/CVPR.2019.00453. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Karras_A_Style-Based_Generator_Architecture_for_Generative_Adversarial_Networks_CVPR_2019_paper.html
  79. 79.X. Chu, I. F. Ilyas, S. Krishnan, J. Wang, Data cleaning: Overview and emerging challenges, in: F. Ozcan, G. Koutrika, ¨ S. Madden (Eds.), Proceedings of the 2016 International Conference on Management of Data, SIGMOD Conference 2016, San Francisco, CA, USA, June 26 - July 01, 2016, ACM, 2016, pp. 2201–2206. doi:10.1145/2882903.2912574. URL https://doi.org/10.1145/2882903.2912574
  80. 80.M. Jesmeen, J. Hossen, S. Sayeed, C. Ho, K. Tawsif, A. Rahman, E. Arif, A survey on cleaning dirty data using machine learning paradigm for big data analytics, Indonesian Journal of Electrical Engineering and Computer Science 10 (3) (2018) 1234–1243.
  81. 81.X. Chu, J. Morcos, I. F. Ilyas, M. Ouzzani, P. Papotti, N. Tang, Y. Ye, KATARA: A data cleaning system powered by knowledge bases and crowdsourcing, in: T. K. Sellis, S. B. Davidson, Z. G. Ives (Eds.), Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data, Melbourne, Victoria, Australia, May 31 - June 4, 2015, ACM, 2015, pp. 1247–1261. doi:10.1145/2723372.2749431. URL https://doi.org/10.1145/2723372.2749431
  82. 82.S. Krishnan, J. Wang, M. J. Franklin, K. Goldberg, T. Kraska, T. Milo, E. Wu, Sampleclean: Fast and reliable analytics on dirty data., IEEE Data Eng. Bull. 38 (3) (2015) 59–75.
  83. 83.S. Krishnan, M. J. Franklin, K. Goldberg, J. Wang, E. Wu, Activeclean: An interactive data cleaning framework for modern machine learning, in: F. Ozcan, G. Koutrika, S. Madden (Eds.), ¨ Proceedings of the 2016 International Conference on Management of Data, SIGMOD Conference 2016, San Francisco, CA, USA, June 26 - July 01, 2016, ACM, 2016, pp. 2117–2120. doi:10.1145/2882903.2899409. URL https://doi.org/10.1145/2882903.2899409
  84. 84.S. Krishnan, M. J. Franklin, K. Goldberg, E. Wu, Boostclean: Automated error detection and repair for machine learning, arXiv preprint arXiv:1711.01299.
  85. 85.S. Krishnan, E. Wu, Alphaclean: Automatic generation of data cleaning pipelines, arXiv preprint arXiv:1904.11827.
  86. 86.I. Gemp, G. Theocharous, M. Ghavamzadeh, Automated data cleansing through meta-learning, in: S. P. Singh, S. Markovitch (Eds.), Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9, 2017, San Francisco, California, USA, AAAI Press, 2017, pp. 4760–4761. URL http://aaai.org/ocs/index.php/IAAI/IAAI17/paper/view/14236
  87. 87.I. F. Ilyas, Effective data cleaning with continuous evaluation., IEEE Data Eng. Bull. 39 (2) (2016) 38–46.
  88. 88.M. Mahdavi, F. Neutatz, L. Visengeriyeva, Z. Abedjan, Towards automated data cleaning workflows, Machine Learning 15 (2019) 16.
  89. 89.T. DeVries, G. W. Taylor, Improved regularization of convolutional neural networks with cutout, arXiv preprint arXiv:1708.04552.
  90. 90.H. Zhang, M. Ciss´e, Y. N. Dauphin, D. Lopez-Paz, mixup: Beyond empirical risk minimization, in: 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings, OpenReview.net, 2018. URL https://openreview.net/forum?id=r1Ddp1-Rb
  91. 91.A. B. Jung, K. Wada, J. Crall, S. Tanaka, J. Graving, C. Reinders, S. Yadav, J. Banerjee, G. Vecsei, A. Kraft, Z. Rui, J. Borovec, C. Vallentin, S. Zhydenko, K. Pfeiffer, B. Cook, I. Fern´andez, F.-M. De Rainville, C.-H. Weng, A. Ayala-Acevedo, R. Meudec, M. Laporte, et al., imgaug, https://github.com/aleju/imgaug, online; accessed 01-Feb-2020 (2020).
  92. 92.A. Buslaev, A. Parinov, E. Khvedchenya, V. I. Iglovikov, A. A. Kalinin, Albumentations: fast and flexible image augmentations, ArXiv e-printsarXiv:1809.06839.
  93. 93.A. Miko lajczyk, M. Grochowski, Data augmentation for improving deep learning in image classification problem, in: 2018 international interdisciplinary PhD workshop (IIPhDW), IEEE, 2018, pp. 117–122.
  94. 94.A. Miko lajczyk, M. Grochowski, Style transfer-based image synthesis as an efficient regularization technique in deep learning, in: 2019 24th International Conference on Methods and Models in Automation and Robotics (MMAR), IEEE, 2019, pp. 42–47.
  95. 95.A. Antoniou, A. Storkey, H. Edwards, Data augmentation generative adversarial networks, arXiv preprint arXiv:1711.04340.
  96. 96.S. C. Wong, A. Gatt, V. Stamatescu, M. D. McDonnell, Understanding data augmentation for classification: when to warp?, arXiv preprint arXiv:1609.08764.
  97. 97.Z. Xie, S. I. Wang, J. Li, D. L´evy, A. Nie, D. Jurafsky, A. Y. Ng, Data noising as smoothing in neural network language models, in: 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings, OpenReview.net, 2017. URL https://openreview.net/forum?id=H1VyHY9gg
  98. 98.A. W. Yu, D. Dohan, M. Luong, R. Zhao, K. Chen, M. Norouzi, Q. V. Le, Qanet: Combining local convolution with global self-attention for reading comprehension, in: 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings, OpenReview.net, 2018. URL https://openreview.net/forum?id=B14TlG-RW
  99. 99.E. Ma, Nlp augmentation, https://github.com/makcedward/nlpaug (2019).
  100. 100.E. D. Cubuk, B. Zoph, D. Man´e, V. Vasudevan, Q. V. Le, Autoaugment: Learning augmentation strategies from data, in: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, Computer Vision Foundation / IEEE, 2019, pp. 113–123. doi:10.1109/CVPR.2019.00020. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Cubuk_AutoAugment_Learning_Augmentation_Strategies_From_Data_CVPR_2019_paper.html
  101. 101.Y. Li, G. Hu, Y. Wang, T. Hospedales, N. M. Robertson, Y. Yang, Dada: Differentiable automatic data augmentation, arXiv preprint arXiv:2003.03780.
  102. 102.R. Hataya, J. Zdenek, K. Yoshizoe, H. Nakayama, Faster autoaugment: Learning augmentation strategies using backpropagation, arXiv preprint arXiv:1911.06987.
  103. 103.S. Lim, I. Kim, T. Kim, C. Kim, S. Kim, Fast autoaugment, in: H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alch´e-Buc, E. B. Fox, R. Garnett (Eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, 2019, pp. 6662–6672. URL https://proceedings.neurips.cc/paper/2019/hash/6add07cf50424b14fdf649da87843d01-Abstract.html
  104. 104.A. Naghizadeh, M. Abavisani, D. N. Metaxas, Greedy autoaugment, arXiv preprint arXiv:1908.00704.
  105. 105.D. Ho, E. Liang, X. Chen, I. Stoica, P. Abbeel, Population based augmentation: Efficient learning of augmentation policy schedules, in: K. Chaudhuri, R. Salakhutdinov (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, PMLR, 2019, pp. 2731–2741. URL http://proceedings.mlr.press/v97/ho19b.html
  106. 106.T. Niu, M. Bansal, Automatically learning data augmentation policies for dialogue tasks, in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Association for Computational Linguistics, Hong Kong, China, 2019, pp. 1317–1323. doi:10.18653/v1/D19-1132. URL https://www.aclweb.org/anthology/D19-1132
  107. 107.M. Geng, K. Xu, B. Ding, H. Wang, L. Zhang, Learning data augmentation policies using augmented random search, arXiv preprint arXiv:1811.04768.
  108. 108.X. Zhang, Q. Wang, J. Zhang, Z. Zhong, Adversarial autoaugment, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=ByxdUySKvS
  109. 109.C. Lin, M. Guo, C. Li, X. Yuan, W. Wu, J. Yan, D. Lin, W. Ouyang, Online hyper-parameter learning for auto-augmentation strategy, in: 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, IEEE, 2019, pp. 6578–6587. doi:10.1109/ICCV.2019.00668. URL https://doi.org/10.1109/ICCV.2019.00668
  110. 110.T. C. LingChen, A. Khonsari, A. Lashkari, M. R. Nazari, J. S. Sambee, M. A. Nascimento, Uniformaugment: A search-free probabilistic data augmentation approach, arXiv preprint arXiv:2003.14348.
  111. 111.H. Motoda, H. Liu, Feature selection, extraction and construction, Communication of IICM (Institute of Information and Computing Machinery, Taiwan) Vol 5 (67-72) (2002) 2.
  112. 112.M. Dash, H. Liu, Feature selection for classification, Intelligent data analysis 1 (1-4) (1997) 131–156.
  113. 113.M. J. Pazzani, Constructive induction of cartesian product attributes, in: Feature Extraction, Construction and Selection, Springer, 1998, pp. 341–354.
  114. 114.Z. Zheng, A comparison of constructing different types of new feature for decision tree learning, in: Feature Extraction, Construction and Selection, Springer, 1998, pp. 239–255.
  115. 115.J. Gama, Functional trees, Machine Learning 55 (3) (2004) 219–250.
  116. 116.H. Vafaie, K. De Jong, Evolutionary feature space transformation, in: Feature Extraction, Construction and Selection, Springer, 1998, pp. 307–323.
  117. 117.P. Sondhi, Feature construction methods: a survey, sifaka. cs. uiuc. edu 69 (2009) 70–71.
  118. 118.D. Roth, K. Small, Interactive feature space construction using semantic information, in: Proceedings of the Thirteenth Conference on Computational Natural Language Learning (CoNLL-2009), Association for Computational Linguistics, Boulder, Colorado, 2009, pp. 66–74. URL https://www.aclweb.org/anthology/W09-1110
  119. 119.Q. Meng, D. Catchpoole, D. Skillicom, P. J. Kennedy, Relational autoencoder for feature extraction, in: 2017 International Joint Conference on Neural Networks (IJCNN), IEEE, 2017, pp. 364–371.
  120. 120.O. Irsoy, E. Alpaydın, Unsupervised feature extraction with autoencoder trees, Neurocomputing 258 (2017) 63–73.
  121. 121.C. Cortes, V. Vapnik, Support-vector networks, Machine learning 20 (3) (1995) 273–297.
  122. 122.N. S. Altman, An introduction to kernel and nearest-neighbor nonparametric regression, The American Statistician 46 (3) (1992) 175–185.
  123. 123.A. Yang, P. M. Esperan¸ca, F. M. Carlucci, NAS evaluation is frustratingly hard, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=HygrdpVKvr
  124. 124.F. Chollet, Xception: Deep learning with depthwise separable convolutions, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, IEEE Computer Society, 2017, pp. 1800–1807. doi:10.1109/CVPR.2017.195. URL https://doi.org/10.1109/CVPR.2017.195
  125. 125.F. Yu, V. Koltun, Multi-scale context aggregation by dilated convolutions, in: Y. Bengio, Y. LeCun (Eds.), 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, 2016. URL http://arxiv.org/abs/1511.07122
  126. 126.J. Hu, L. Shen, G. Sun, Squeeze-and-excitation networks, in: 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, IEEE Computer Society, 2018, pp. 7132–7141. doi:10.1109/CVPR.2018.00745. URL http://openaccess.thecvf.com/content_cvpr_2018/html/Hu_Squeeze-and-Excitation_Networks_CVPR_2018_paper.html
  127. 127.G. Huang, Z. Liu, L. van der Maaten, K. Q. Weinberger, Densely connected convolutional networks, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, IEEE Computer Society, 2017, pp. 2261–2269. doi:10.1109/CVPR.2017.243. URL https://doi.org/10.1109/CVPR.2017.243
  128. 128.X. Chen, L. Xie, J. Wu, Q. Tian, Progressive differentiable architecture search: Bridging the depth gap between search and evaluation, in: 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, IEEE, 2019, pp. 1294–1303. doi:10.1109/ICCV.2019.00138. URL https://doi.org/10.1109/ICCV.2019.00138
  129. 129.C. Liu, L. Chen, F. Schroff, H. Adam, W. Hua, A. L. Yuille, F. Li, Auto-deeplab: Hierarchical neural architecture search for semantic image segmentation, in: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, Computer Vision Foundation / IEEE, 2019, pp. 82–92. doi:10.1109/CVPR.2019.00017. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Liu_Auto-DeepLab_Hierarchical_Neural_Architecture_Search_for_Semantic_Image_Segmentation_CVPR_2019_paper.html
  130. 130.M. Tan, B. Chen, R. Pang, V. Vasudevan, M. Sandler, A. Howard, Q. V. Le, Mnasnet: Platform-aware neural architecture search for mobile, in: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, Computer Vision Foundation / IEEE, 2019, pp. 2820–2828. doi:10.1109/CVPR.2019.00293. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Tan_MnasNet_Platform-Aware_Neural_Architecture_Search_for_Mobile_CVPR_2019_paper.html
  131. 131.B. Wu, X. Dai, P. Zhang, Y. Wang, F. Sun, Y. Wu, Y. Tian, P. Vajda, Y. Jia, K. Keutzer, Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search, in: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, Computer Vision Foundation / IEEE, 2019, pp. 10734–10742. doi:10.1109/CVPR.2019.01099. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Wu_FBNet_Hardware-Aware_Efficient_ConvNet_Design_via_Differentiable_Neural_Architecture_Search_CVPR_2019_paper.html
  132. 132.H. Cai, L. Zhu, S. Han, Proxylessnas: Direct neural architecture search on target task and hardware, in: 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, OpenReview.net, 2019. URL https://openreview.net/forum?id=HylVB3AqYm
  133. 133.M. Courbariaux, Y. Bengio, J. David, Binaryconnect: Training deep neural networks with binary weights during propagations, in: C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, R. Garnett (Eds.), Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, 2015, pp. 3123–3131. URL https://proceedings.neurips.cc/paper/2015/hash/3e15cc11f979ed25912dff5b0669f2cd-Abstract.html
  134. 134.G. Hinton, O. Vinyals, J. Dean, Distilling the knowledge in a neural network, arXiv preprint arXiv:1503.02531.
  135. 135.J. Yosinski, J. Clune, Y. Bengio, H. Lipson, How transferable are features in deep neural networks?, in: Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, K. Q. Weinberger (Eds.), Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada, 2014, pp. 3320–3328. URL https://proceedings.neurips.cc/paper/2014/hash/375c71349b295fbe2dcdca9206f20a06-Abstract.html
  136. 136.T. Wei, C. Wang, C. W. Chen, Modularized morphing of neural networks, arXiv preprint arXiv:1701.03281.
  137. 137.H. Cai, T. Chen, W. Zhang, Y. Yu, J. Wang, Efficient architecture search by network transformation, in: S. A. McIlraith, K. Q. Weinberger (Eds.), Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, AAAI Press, 2018, pp. 2787–2794. URL https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/16755
  138. 138.A. Kwasigroch, M. Grochowski, M. Mikolajczyk, Deep neural network architecture search using network morphism, in: 2019 24th International Conference on Methods and Models in Automation and Robotics (MMAR), IEEE, 2019, pp. 30–35.
  139. 139.H. Cai, J. Yang, W. Zhang, S. Han, Y. Yu, Path-level network transformation for efficient architecture search, in: J. G. Dy, A. Krause (Eds.), Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm¨assan, Stockholm, Sweden, July 10-15, 2018, Vol. 80 of Proceedings of Machine Learning Research, PMLR, 2018, pp. 677–686. URL http://proceedings.mlr.press/v80/cai18a.html
  140. 140.J. Fang, Y. Sun, K. Peng, Q. Zhang, Y. Li, W. Liu, X. Wang, Fast neural network adaptation via parameter remapping and architecture search, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=rklTmyBKPH
  141. 141.A. Gordon, E. Eban, O. Nachum, B. Chen, H. Wu, T. Yang, E. Choi, Morphnet: Fast & simple resource-constrained structure learning of deep networks, in: 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, IEEE Computer Society, 2018, pp. 1586–1595. doi:10.1109/CVPR.2018.00171. URL http://openaccess.thecvf.com/content_cvpr_2018/html/Gordon_MorphNet_Fast__CVPR_2018_paper.html
  142. 142.M. Tan, Q. V. Le, Efficientnet: Rethinking model scaling for convolutional neural networks, in: K. Chaudhuri, R. Salakhutdinov (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, PMLR, 2019, pp. 6105–6114. URL http://proceedings.mlr.press/v97/tan19a.html
  143. 143.J. F. Miller, S. L. Harding, Cartesian genetic programming, in: Proceedings of the 10th annual conference companion on Genetic and evolutionary computation, ACM, 2008, pp. 2701–2726.
  144. 144.J. F. Miller, S. L. Smith, Redundancy and computational efficiency in cartesian genetic programming, IEEE Transactions on Evolutionary Computation 10 (2) (2006) 167–174.
  145. 145.F. Gruau, Cellular encoding as a graph grammar, in: IEEE Colloquium on Grammatical Inference: Theory, Applications & Alternatives, 1993.
  146. 146.C. Fernando, D. Banarse, M. Reynolds, F. Besse, D. Pfau, M. Jaderberg, M. Lanctot, D. Wierstra, Convolution by evolution: Differentiable pattern producing networks, in: Proceedings of the Genetic and Evolutionary Computation Conference 2016, ACM, 2016, pp. 109–116.
  147. 147.M. Kim, L. Rigazio, Deep clustered convolutional kernels, in: Feature Extraction: Modern Questions and Challenges, 2015, pp. 160–172.
  148. 148.J. K. Pugh, K. O. Stanley, Evolving multimodal controllers with hyperneat, in: Proceedings of the 15th annual conference on Genetic and evolutionary computation, ACM, 2013, pp. 735–742.
  149. 149.H. Zhu, Z. An, C. Yang, K. Xu, E. Zhao, Y. Xu, Eena: Efficient evolution of neural architecture (2019). arXiv:1905.07320.
  150. 150.R. J. Williams, Simple statistical gradient-following algorithms for connectionist reinforcement learning, Machine learning 8 (3-4) (1992) 229–256.
  151. 151.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal policy optimization algorithms, arXiv preprint arXiv:1707.06347.
  152. 152.M. Marcus, G. Kim, M. A. Marcinkiewicz, R. MacIntyre, A. Bies, M. Ferguson, K. Katz, B. Schasberger, The Penn Treebank: Annotating predicate argument structure, in: Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8-11, 1994, 1994. URL https://www.aclweb.org/anthology/H94-1020
  153. 153.C. He, H. Ye, L. Shen, T. Zhang, Milenas: Efficient neural architecture search via mixed-level reformulation, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 11990–11999. doi:10.1109/CVPR42600.2020.01201. URL https://doi.org/10.1109/CVPR42600.2020.01201
  154. 154.X. Dong, Y. Yang, Searching for a robust neural architecture in four GPU hours, in: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, Computer Vision Foundation / IEEE, 2019, pp. 1761–1770. doi:10.1109/CVPR.2019.00186. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Dong_Searching_for_a_Robust_Neural_Architecture_in_Four_GPU_Hours_CVPR_2019_paper.html
  155. 155.S. Xie, H. Zheng, C. Liu, L. Lin, SNAS: stochastic neural architecture search, in: 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, OpenReview.net, 2019. URL https://openreview.net/forum?id=rylqooRqK7
  156. 156.B. Wu, Y. Wang, P. Zhang, Y. Tian, P. Vajda, K. Keutzer, Mixed precision quantization of convnets via differentiable neural architecture search (2018). arXiv:1812.00090.
  157. 157.E. Jang, S. Gu, B. Poole, Categorical reparameterization with gumbel-softmax, in: 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings, OpenReview.net, 2017. URL https://openreview.net/forum?id=rkE3y85ee
  158. 158.C. J. Maddison, A. Mnih, Y. W. Teh, The concrete distribution: A continuous relaxation of discrete random variables, in: 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings, OpenReview.net, 2017. URL https://openreview.net/forum?id=S1jE5L5gl
  159. 159.H. Liang, S. Zhang, J. Sun, X. He, W. Huang, K. Zhuang, Z. Li, Darts+: Improved differentiable architecture search with early stopping, arXiv preprint arXiv:1909.06035.
  160. 160.K. Kandasamy, W. Neiswanger, J. Schneider, B. P´oczos, E. P. Xing, Neural architecture search with bayesian optimisation and optimal transport, in: S. Bengio, H. M. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, R. Garnett (Eds.), Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montr´eal, Canada, 2018, pp. 2020–2029. URL https://proceedings.neurips.cc/paper/2018/hash/f33ba15effa5c10e873bf3842afb46a6-Abstract.html
  161. 161.R. Negrinho, G. Gordon, Deeparchitect: Automatically designing and training deep architectures (2017). arXiv:1704.08792.
  162. 162.R. Negrinho, M. R. Gormley, G. J. Gordon, D. Patil, N. Le, D. Ferreira, Towards modular and programmable architecture search, in: H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alch´e-Buc, E. B. Fox, R. Garnett (Eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, 2019, pp. 13715–13725. URL https://proceedings.neurips.cc/paper/2019/hash/4ab50afd6dcc95fcba76d0fe04295632-Abstract.html
  163. 163.G. Dikov, J. Bayer, Bayesian learning of neural network architectures, in: K. Chaudhuri, M. Sugiyama (Eds.), The 22nd International Conference on Artificial Intelligence and Statistics, AISTATS 2019, 16-18 April 2019, Naha, Okinawa, Japan, Vol. 89 of Proceedings of Machine Learning Research, PMLR, 2019, pp. 730–738. URL http://proceedings.mlr.press/v89/dikov19a.html
  164. 164.C. White, W. Neiswanger, Y. Savani, Bananas: Bayesian optimization with neural architectures for neural architecture search (2019). arXiv:1910.11858.
  165. 165.M. Wistuba, Bayesian optimization combined with incremental evaluation for neural network architecture optimization, in: Proceedings of the International Workshop on Automatic Selection, Configuration and Composition of Machine Learning Algorithms, 2017.
  166. 166.J. Perez-Rua, M. Baccouche, S. Pateux, Efficient progressive neural architecture search, in: British Machine Vision Conference 2018, BMVC 2018, Newcastle, UK, September 3-6, 2018, BMVA Press, 2018, p. 150. URL http://bmvc2018.org/contents/papers/0291.pdf
  167. 167.C. E. Rasmussen, Gaussian processes in machine learning, Lecture Notes in Computer Science (2003) 63–71.
  168. 168.J. Bergstra, R. Bardenet, Y. Bengio, B. K´egl, Algorithms for hyper-parameter optimization, in: J. Shawe-Taylor, R. S. Zemel, P. L. Bartlett, F. C. N. Pereira, K. Q. Weinberger (Eds.), Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011. Proceedings of a meeting held 12-14 December 2011, Granada, Spain, 2011, pp. 2546–2554. URL https://proceedings.neurips.cc/paper/2011/hash/86e8f7ab32cfd12577bc2619bc635690-Abstract.html
  169. 169.R. Luo, F. Tian, T. Qin, E. Chen, T. Liu, Neural architecture optimization, in: S. Bengio, H. M. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, R. Garnett (Eds.), Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montr´eal, Canada, 2018, pp. 7827–7838. URL https://proceedings.neurips.cc/paper/2018/hash/933670f1ac8ba969f32989c312faba75-Abstract.html
  170. 170.M. M. Ian Dewancker, S. Clark, Bayesian optimization primer. URL https://app.sigopt.com/static/pdf/SigOpt_Bayesian_Optimization_Primer.pdf
  171. 171.B. Shahriari, K. Swersky, Z. Wang, R. P. Adams, N. De Freitas, Taking the human out of the loop: A review of bayesian optimization, Proceedings of the IEEE 104 (1) (2016) 148–175.
  172. 172.J. Snoek, O. Rippel, K. Swersky, R. Kiros, N. Satish, N. Sundaram, M. M. A. Patwary, Prabhat, R. P. Adams, Scalable bayesian optimization using deep neural networks, in: F. R. Bach, D. M. Blei (Eds.), Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, Vol. 37 of JMLR Workshop and Conference Proceedings, JMLR.org, 2015, pp. 2171–2180. URL http://proceedings.mlr.press/v37/snoek15.html
  173. 173.J. Snoek, H. Larochelle, R. P. Adams, Practical bayesian optimization of machine learning algorithms, in: P. L. Bartlett, F. C. N. Pereira, C. J. C. Burges, L. Bottou, K. Q. Weinberger (Eds.), Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting held December 3-6, 2012, Lake Tahoe, Nevada, United States, 2012, pp. 2960–2968. URL https://proceedings.neurips.cc/paper/2012/hash/05311655a15b75fab86956663e1819cd-Abstract.html
  174. 174.J. Stork, M. Zaefferer, T. Bartz-Beielstein, Improving neuroevolution efficiency by surrogate model-based optimization with phenotypic distance kernels (2019). arXiv:1902.03419.
  175. 175.K. Swersky, D. Duvenaud, J. Snoek, F. Hutter, M. A. Osborne, Raiders of the lost architecture: Kernels for bayesian optimization in conditional parameter spaces (2014). arXiv:1409.4011.
  176. 176.A. Camero, H. Wang, E. Alba, T. B¨ack, Bayesian neural architecture search using a training-free performance metric (2020). arXiv:2001.10726.
  177. 177.C. Thornton, F. Hutter, H. H. Hoos, K. Leyton-Brown, Auto-weka: combined selection and hyperparameter optimization of classification algorithms, in: I. S. Dhillon, Y. Koren, R. Ghani, T. E. Senator, P. Bradley, R. Parekh, J. He, R. L. Grossman, R. Uthurusamy (Eds.), The 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD 2013, Chicago, IL, USA, August 11-14, 2013, ACM, 2013, pp. 847–855. doi:10.1145/2487575.2487629. URL https://doi.org/10.1145/2487575.2487629
  178. 178.A. sharpdarts, V. Jain, G. D. Hager, sharpdarts: Faster and more accurate differentiable architecture search, Tech. rep. (2019).
  179. 179.Y. Geifman, R. El-Yaniv, Deep active learning with a neural architecture search, in: H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alch´e-Buc, E. B. Fox, R. Garnett (Eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, 2019, pp. 5974–5984. URL https://proceedings.neurips.cc/paper/2019/hash/b59307fdacf7b2db12ec4bd5ca1caba8-Abstract.html
  180. 180.L. Li, A. Talwalkar, Random search and reproducibility for neural architecture search, in: A. Globerson, R. Silva (Eds.), Proceedings of the Thirty-Fifth Conference on Uncertainty in Artificial Intelligence, UAI 2019, Tel Aviv, Israel, July 22-25, 2019, Vol. 115 of Proceedings of Machine Learning Research, AUAI Press, 2019, pp. 367–377. URL http://proceedings.mlr.press/v115/li20c.html
  181. 181.J. Bergstra, Y. Bengio, Random search for hyper-parameter optimization, Journal of machine learning research 13 (Feb) (2012) 281–305.
  182. 182.C.-W. Hsu, C.-C. Chang, C.-J. Lin, et al., A practical guide to support vector classification.
  183. 183.J. Y. Hesterman, L. Caucci, M. A. Kupinski, H. H. Barrett, L. R. Furenlid, Maximum-likelihood estimation with a contracting-grid search algorithm, IEEE transactions on nuclear science 57 (3) (2010) 1077–1084.
  184. 184.L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, A. Talwalkar, Hyperband: A novel bandit-based approach to hyperparameter optimization, The Journal of Machine Learning Research 18 (1) (2017) 6765–6816.
  185. 185.M. Feurer, F. Hutter, Hyperparameter Optimization, Springer International Publishing, Cham, 2019, pp. 3–33. URL https://doi.org/10.1007/978-3-030-05318-5_1
  186. 186.T. Yu, H. Zhu, Hyper-parameter optimization: A review of algorithms and applications, arXiv preprint arXiv:2003.05689.
  187. 187.Y. Bengio, Gradient-based optimization of hyperparameters, Neural computation 12 (8) (2000) 1889–1900.
  188. 188.J. Domke, Generic methods for optimization-based modeling, in: Artificial Intelligence and Statistics, 2012, pp. 318–326.
  189. 189.D. Maclaurin, D. Duvenaud, R. P. Adams, Gradient-based hyperparameter optimization through reversible learning, in: F. R. Bach, D. M. Blei (Eds.), Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, Vol. 37 of JMLR Workshop and Conference Proceedings, JMLR.org, 2015, pp. 2113–2122. URL http://proceedings.mlr.press/v37/maclaurin15.html
  190. 190.F. Pedregosa, Hyperparameter optimization with approximate gradient, in: M. Balcan, K. Q. Weinberger (Eds.), Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, Vol. 48 of JMLR Workshop and Conference Proceedings, JMLR.org, 2016, pp. 737–746. URL http://proceedings.mlr.press/v48/pedregosa16.html
  191. 191.L. Franceschi, M. Donini, P. Frasconi, M. Pontil, Forward and reverse gradient-based hyperparameter optimization, in: D. Precup, Y. W. Teh (Eds.), Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, Vol. 70 of Proceedings of Machine Learning Research, PMLR, 2017, pp. 1165–1173. URL http://proceedings.mlr.press/v70/franceschi17a.html
  192. 192.K. Chandra, E. Meijer, S. Andow, E. Arroyo-Fang, I. Dea, J. George, M. Grueter, B. Hosmer, S. Stumpos, A. Tempest, et al., Gradient descent: The ultimate optimizer, arXiv preprint arXiv:1909.13371.
  193. 193.D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, in: Y. Bengio, Y. LeCun (Eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015. URL http://arxiv.org/abs/1412.6980
  194. 194.P. Chrabaszcz, I. Loshchilov, F. Hutter, A downsampled variant of imagenet as an alternative to the CIFAR datasets, CoRR abs/1707.08819. arXiv:1707.08819. URL http://arxiv.org/abs/1707.08819
  195. 195.Y. Hu, Y. Yu, W. Tu, Q. Yang, Y. Chen, W. Dai, Multi-fidelity automatic hyper-parameter tuning via transfer series expansion, in: The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, AAAI Press, 2019, pp. 3846–3853. doi:10.1609/aaai.v33i01.33013846. URL https://doi.org/10.1609/aaai.v33i01.33013846
  196. 196.C. Wong, N. Houlsby, Y. Lu, A. Gesmundo, Transfer learning with neural automl, in: S. Bengio, H. M. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, R. Garnett (Eds.), Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montr´eal, Canada, 2018, pp. 8366–8375. URL https://proceedings.neurips.cc/paper/2018/hash/bdb3c278f45e6734c35733d24299d3f4-Abstract.html
  197. 197.D. Stamoulis, R. Ding, D. Wang, D. Lymberopoulos, B. Priyantha, J. Liu, D. Marculescu, Single-path nas: Designing hardware-efficient convnets in less than 4 hours, arXiv preprint arXiv:1904.02877.
  198. 198.K. Eggensperger, F. Hutter, H. H. Hoos, K. Leyton-Brown, Surrogate benchmarks for hyperparameter optimization., in: MetaSel@ ECAI, 2014, pp. 24–31.
  199. 199.C. Wang, Q. Duan, W. Gong, A. Ye, Z. Di, C. Miao, An evaluation of adaptive surrogate modeling based optimization with two benchmark problems, Environmental Modelling & Software 60 (2014) 167–179.
  200. 200.K. Eggensperger, F. Hutter, H. H. Hoos, K. Leyton-Brown, Efficient benchmarking of hyperparameter optimizers via surrogates, in: B. Bonet, S. Koenig (Eds.), Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, January 25-30, 2015, Austin, Texas, USA, AAAI Press, 2015, pp. 1114–1120. URL http://www.aaai.org/ocs/index.php/AAAI/AAAI15/paper/view/9993
  201. 201.K. K. Vu, C. D’Ambrosio, Y. Hamadi, L. Liberti, Surrogate-based methods for black-box optimization, International Transactions in Operational Research 24 (3) (2017) 393–424.
  202. 202.R. Luo, X. Tan, R. Wang, T. Qin, E. Chen, T.-Y. Liu, Semi-supervised neural architecture search (2020). arXiv:2002.10389.
  203. 203.A. Klein, S. Falkner, J. T. Springenberg, F. Hutter, Learning curve prediction with bayesian neural networks, in: 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings, OpenReview.net, 2017. URL https://openreview.net/forum?id=S11KBYclx
  204. 204.B. Deng, J. Yan, D. Lin, Peephole: Predicting network performance before training, arXiv preprint arXiv:1712.03351.
  205. 205.T. Domhan, J. T. Springenberg, F. Hutter, Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves, in: Q. Yang, M. J. Wooldridge (Eds.), Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence, IJCAI 2015, Buenos Aires, Argentina, July 25-31, 2015, AAAI Press, 2015, pp. 3460–3468. URL http://ijcai.org/Abstract/15/487
  206. 206.M. Mahsereci, L. Balles, C. Lassner, P. Hennig, Early stopping without a validation set, arXiv preprint arXiv:1703.09580.
  207. 207.D. Han, J. Kim, J. Kim, Deep pyramidal residual networks, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, IEEE Computer Society, 2017, pp. 6307–6315. doi:10.1109/CVPR.2017.668. URL https://doi.org/10.1109/CVPR.2017.668
  208. 208.J. Cui, P. Chen, R. Li, S. Liu, X. Shen, J. Jia, Fast and practical neural architecture search, in: 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, IEEE, 2019, pp. 6508–6517. doi:10.1109/ICCV.2019.00661. URL https://doi.org/10.1109/ICCV.2019.00661
  209. 209.X. Dong, Y. Yang, One-shot neural architecture search via self-evaluated template network, in: 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, IEEE, 2019, pp. 3680–3689. doi:10.1109/ICCV.2019.00378. URL https://doi.org/10.1109/ICCV.2019.00378
  210. 210.H. Zhou, M. Yang, J. Wang, W. Pan, Bayesnas: A bayesian approach for neural architecture search, in: K. Chaudhuri, R. Salakhutdinov (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, PMLR, 2019, pp. 7603–7613. URL http://proceedings.mlr.press/v97/zhou19e.html
  211. 211.Y. Xu, L. Xie, X. Zhang, X. Chen, G. Qi, Q. Tian, H. Xiong, PC-DARTS: partial channel connections for memory-efficient architecture search, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=BJlS634tPr
  212. 212.G. Li, G. Qian, I. C. Delgadillo, M. M¨uller, A. K. Thabet, B. Ghanem, SGAS: sequential greedy architecture search, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 1617–1627. doi:10.1109/CVPR42600.2020.00169. URL https://doi.org/10.1109/CVPR42600.2020.00169
  213. 213.M. Zhang, H. Li, S. Pan, X. Chang, S. W. Su, Overcoming multi-model forgetting in one-shot NAS with diversity maximization, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 7806–7815. doi:10.1109/CVPR42600.2020.00783. URL https://doi.org/10.1109/CVPR42600.2020.00783
  214. 214.C. Zhang, M. Ren, R. Urtasun, Graph hypernetworks for neural architecture search, in: 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, OpenReview.net, 2019. URL https://openreview.net/forum?id=rkgW0oA9FX
  215. 215.M. Sandler, A. G. Howard, M. Zhu, A. Zhmoginov, L. Chen, Mobilenetv2: Inverted residuals and linear bottlenecks, in: 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, IEEE Computer Society, 2018, pp. 4510–4520. doi:10.1109/CVPR.2018.00474. URL http://openaccess.thecvf.com/content_cvpr_2018/html/Sandler_MobileNetV2_Inverted_Residuals_CVPR_2018_paper.html
  216. 216.S. You, T. Huang, M. Yang, F. Wang, C. Qian, C. Zhang, Greedynas: Towards fast one-shot NAS with greedy supernet, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 1996–2005. doi:10.1109/CVPR42600.2020.00207. URL https://doi.org/10.1109/CVPR42600.2020.00207
  217. 217.H. Cai, C. Gan, T. Wang, Z. Zhang, S. Han, Once-for-all: Train one network and specialize it for efficient deployment, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=HylxE1HKwS
  218. 218.J. Mei, Y. Li, X. Lian, X. Jin, L. Yang, A. L. Yuille, J. Yang, Atomnas: Fine-grained end-to-end neural architecture search, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=BylQSxHFwr
  219. 219.S. Hu, S. Xie, H. Zheng, C. Liu, J. Shi, X. Liu, D. Lin, DSNAS: direct neural architecture search without parameter retraining, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 12081–12089. doi:10.1109/CVPR42600.2020.01210. URL https://doi.org/10.1109/CVPR42600.2020.01210
  220. 220.J. Fang, Y. Sun, Q. Zhang, Y. Li, W. Liu, X. Wang, Densely connected search space for more flexible neural architecture search, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 10625–10634. doi:10.1109/CVPR42600.2020.01064. URL https://doi.org/10.1109/CVPR42600.2020.01064
  221. 221.A. Wan, X. Dai, P. Zhang, Z. He, Y. Tian, S. Xie, B. Wu, M. Yu, T. Xu, K. Chen, P. Vajda, J. E. Gonzalez, Fbnetv2: Differentiable neural architecture search for spatial and channel dimensions, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 12962–12971. doi:10.1109/CVPR42600.2020.01298. URL https://doi.org/10.1109/CVPR42600.2020.01298
  222. 222.R. Istrate, F. Scheidegger, G. Mariani, D. S. Nikolopoulos, C. Bekas, A. C. I. Malossi, TAPAS: train-less accuracy predictor for architecture search, in: The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, AAAI Press, 2019, pp. 3927–3934. doi:10.1609/aaai.v33i01.33013927. URL https://doi.org/10.1609/aaai.v33i01.33013927
  223. 223.M. G. Kendall, A new measure of rank correlation, Biometrika 30 (1/2) (1938) 81–93. URL http://www.jstor.org/stable/2332226
  224. 224.C. Ying, A. Klein, E. Christiansen, E. Real, K. Murphy, F. Hutter, Nas-bench-101: Towards reproducible neural architecture search, in: K. Chaudhuri, R. Salakhutdinov (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, PMLR, 2019, pp. 7105–7114. URL http://proceedings.mlr.press/v97/ying19a.html
  225. 225.X. Dong, Y. Yang, Nas-bench-201: Extending the scope of reproducible neural architecture search, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=HJxyZkBKDr
  226. 226.N. Klyuchnikov, I. Trofimov, E. Artemova, M. Salnikov, M. Fedorov, E. Burnaev, Nas-bench-nlp: Neural architecture search benchmark for natural language processing (2020). arXiv:2006.07116.
  227. 227.X. Zhang, Z. Huang, N. Wang, You only search once: Single shot neural architecture search via direct sparse optimization, arXiv preprint arXiv:1811.01567.
  228. 228.J. Yu, P. Jin, H. Liu, G. Bender, P.-J. Kindermans, M. Tan, T. Huang, X. Song, R. Pang, Q. Le, Bignas: Scaling up neural architecture search with big single-stage models, arXiv preprint arXiv:2003.11142.
  229. 229.X. Chu, B. Zhang, R. Xu, J. Li, Fairnas: Rethinking evaluation fairness of weight sharing neural architecture search, arXiv preprint arXiv:1907.01845.
  230. 230.Y. Benyahia, K. Yu, K. Bennani-Smires, M. Jaggi, A. C. Davison, M. Salzmann, C. Musat, Overcoming multi-model forgetting, in: K. Chaudhuri, R. Salakhutdinov (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, PMLR, 2019, pp. 594–603. URL http://proceedings.mlr.press/v97/benyahia19a.html
  231. 231.M. Zhang, H. Li, S. Pan, X. Chang, S. W. Su, Overcoming multi-model forgetting in one-shot NAS with diversity maximization, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 7806–7815. doi:10.1109/CVPR42600.2020.00783. URL https://doi.org/10.1109/CVPR42600.2020.00783
  232. 232.G. Bender, P. Kindermans, B. Zoph, V. Vasudevan, Q. V. Le, Understanding and simplifying one-shot architecture search, in: J. G. Dy, A. Krause (Eds.), Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm¨assan, Stockholm, Sweden, July 10-15, 2018, Vol. 80 of Proceedings of Machine Learning Research, PMLR, 2018, pp. 549–558. URL http://proceedings.mlr.press/v80/bender18a.html
  233. 233.X. Dong, M. Tan, A. W. Yu, D. Peng, B. Gabrys, Q. V. Le, Autohas: Differentiable hyper-parameter and architecture search (2020). arXiv:2006.03656.
  234. 234.A. Klein, F. Hutter, Tabular benchmarks for joint architecture and hyperparameter optimization, arXiv preprint arXiv:1905.04970.
  235. 235.X. Dai, A. Wan, P. Zhang, B. Wu, Z. He, Z. Wei, K. Chen, Y. Tian, M. Yu, P. Vajda, et al., Fbnetv3: Joint architecture-recipe search using neural acquisition function, arXiv preprint arXiv:2006.02049.
  236. 236.C.-H. Hsu, S.-H. Chang, J.-H. Liang, H.-P. Chou, C.-H. Liu, S.-C. Chang, J.-Y. Pan, Y.-T. Chen, W. Wei, D.-C. Juan, Monas: Multi-objective neural architecture search using reinforcement learning, arXiv preprint arXiv:1806.10332.
  237. 237.X. He, S. Wang, S. Shi, X. Chu, J. Tang, X. Liu, C. Yan, J. Zhang, G. Ding, Benchmarking deep learning models and automated model design for covid-19 detection with chest ct scans, medRxiv.
  238. 238.L. Faes, S. K. Wagner, D. J. Fu, X. Liu, E. Korot, J. R. Ledsam, T. Back, R. Chopra, N. Pontikos, C. Kern, et al., Automated deep learning design for medical image classification by health-care professionals with no coding experience: a feasibility study, The Lancet Digital Health 1 (5) (2019) e232–e242.
  239. 239.X. He, S. Wang, X. Chu, S. Shi, J. Tang, X. Liu, C. Yan, J. Zhang, G. Ding, Automated model design and benchmarking of 3d deep learning models for covid-19 detection with chest ct scans (2021). arXiv:2101.05442.
  240. 240.G. Ghiasi, T. Lin, Q. V. Le, NAS-FPN: learning scalable feature pyramid architecture for object detection, in: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, Computer Vision Foundation / IEEE, 2019, pp. 7036–7045. doi:10.1109/CVPR.2019.00720. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Ghiasi_NAS-FPN_Learning_Scalable_Feature_Pyramid_Architecture_for_Object_Detection_CVPR_2019_paper.html
  241. 241.H. Xu, L. Yao, Z. Li, X. Liang, W. Zhang, Auto-fpn: Automatic network architecture adaptation for object detection beyond classification, in: 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, IEEE, 2019, pp. 6648–6657. doi:10.1109/ICCV.2019.00675. URL https://doi.org/10.1109/ICCV.2019.00675
  242. 242.M. Tan, R. Pang, Q. V. Le, Efficientdet: Scalable and efficient object detection, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 10778–10787. doi:10.1109/CVPR42600.2020.01079. URL https://doi.org/10.1109/CVPR42600.2020.01079
  243. 243.Y. Chen, T. Yang, X. Zhang, G. Meng, C. Pan, J. Sun, Detnas: Neural architecture search on object detection, arXiv preprint arXiv:1903.10979 1 (2) (2019) 4–1.
  244. 244.J. Guo, K. Han, Y. Wang, C. Zhang, Z. Yang, H. Wu, X. Chen, C. Xu, Hit-detector: Hierarchical trinity architecture search for object detection, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 11402–11411. doi:10.1109/CVPR42600.2020.01142. URL https://doi.org/10.1109/CVPR42600.2020.01142
  245. 245.C. Jiang, H. Xu, W. Zhang, X. Liang, Z. Li, SP-NAS: serial-to-parallel backbone search for object detection, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 11860–11869. doi:10.1109/CVPR42600.2020.01188. URL https://doi.org/10.1109/CVPR42600.2020.01188
  246. 246.Y. Weng, T. Zhou, Y. Li, X. Qiu, Nas-unet: Neural architecture search for medical image segmentation, IEEE Access 7 (2019) 44247–44257.
  247. 247.V. Nekrasov, H. Chen, C. Shen, I. D. Reid, Fast neural architecture search of compact semantic segmentation models via auxiliary cells, in: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, Computer Vision Foundation / IEEE, 2019, pp. 9126–9135. doi:10.1109/CVPR.2019.00934. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Nekrasov_Fast_Neural_Architecture_Search_of_Compact_Semantic_Segmentation_Models_via_CVPR_2019_paper.html
  248. 248.W. Bae, S. Lee, Y. Lee, B. Park, M. Chung, K.-H. Jung, Resource optimized neural architecture search for 3d medical image segmentation, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2019, pp. 228–236.
  249. 249.D. Yang, H. Roth, Z. Xu, F. Milletari, L. Zhang, D. Xu, Searching learning strategy with reinforcement learning for 3d medical image segmentation, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2019, pp. 3–11.
  250. 250.N. Dong, M. Xu, X. Liang, Y. Jiang, W. Dai, E. Xing, Neural architecture search for adversarial medical image segmentation, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2019, pp. 828–836.
  251. 251.S. Kim, I. Kim, S. Lim, W. Baek, C. Kim, H. Cho, B. Yoon, T. Kim, Scalable neural architecture search for 3d medical image segmentation, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2019, pp. 220–228.
  252. 252.R. Quan, X. Dong, Y. Wu, L. Zhu, Y. Yang, Auto-reid: Searching for a part-aware convnet for person re-identification, in: 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, IEEE, 2019, pp. 3749–3758. doi:10.1109/ICCV.2019.00385. URL https://doi.org/10.1109/ICCV.2019.00385
  253. 253.D. Song, C. Xu, X. Jia, Y. Chen, C. Xu, Y. Wang, Efficient residual dense block search for image super-resolution., in: AAAI, 2020, pp. 12007–12014.
  254. 254.X. Chu, B. Zhang, H. Ma, R. Xu, J. Li, Q. Li, Fast, accurate and lightweight super-resolution with neural architecture search, arXiv preprint arXiv:1901.07261.
  255. 255.Y. Guo, Y. Luo, Z. He, J. Huang, J. Chen, Hierarchical neural architecture search for single image super-resolution, arXiv preprint arXiv:2003.04619.
  256. 256.H. Zhang, Y. Li, H. Chen, C. Shen, Ir-nas: Neural architecture search for image restoration, arXiv preprint arXiv:1909.08228.
  257. 257.X. Gong, S. Chang, Y. Jiang, Z. Wang, Autogan: Neural architecture search for generative adversarial networks, in: 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, IEEE, 2019, pp. 3223–3233. doi:10.1109/ICCV.2019.00332. URL https://doi.org/10.1109/ICCV.2019.00332
  258. 258.Y. Fu, W. Chen, H. Wang, H. Li, Y. Lin, Z. Wang, Autogandistiller: Searching to compress generative adversarial networks, in: Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Vol. 119 of Proceedings of Machine Learning Research, PMLR, 2020, pp. 3292–3303. URL http://proceedings.mlr.press/v119/fu20b.html
  259. 259.M. Li, J. Lin, Y. Ding, Z. Liu, J. Zhu, S. Han, GAN compression: Efficient architectures for interactive conditional gans, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 5283–5293. doi:10.1109/CVPR42600.2020.00533. URL https://doi.org/10.1109/CVPR42600.2020.00533
  260. 260.C. Gao, Y. Chen, S. Liu, Z. Tan, S. Yan, Adversarialnas: Adversarial neural architecture search for gans, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 5679–5688. doi:10.1109/CVPR42600.2020.00572. URL https://doi.org/10.1109/CVPR42600.2020.00572
  261. 261.T. Saikia, Y. Marrakchi, A. Zela, F. Hutter, T. Brox, Autodispnet: Improving disparity estimation with automl, in: 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, IEEE, 2019, pp. 1812–1823. doi:10.1109/ICCV.2019.00190. URL https://doi.org/10.1109/ICCV.2019.00190
  262. 262.W. Peng, X. Hong, G. Zhao, Video action recognition via neural architecture searching, in: 2019 IEEE International Conference on Image Processing (ICIP), IEEE, 2019, pp. 11–15.
  263. 263.M. S. Ryoo, A. J. Piergiovanni, M. Tan, A. Angelova, Assemblenet: Searching for multi-stream neural connectivity in video architectures, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=SJgMK64Ywr
  264. 264.V. Nekrasov, H. Chen, C. Shen, I. Reid, Architecture search of dynamic cells for semantic video segmentation, in: The IEEE Winter Conference on Applications of Computer Vision, 2020, pp. 1970–1979.
  265. 265.A. J. Piergiovanni, A. Angelova, A. Toshev, M. S. Ryoo, Evolving space-time neural architectures for videos, in: 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, IEEE, 2019, pp. 1793–1802. doi:10.1109/ICCV.2019.00188. URL https://doi.org/10.1109/ICCV.2019.00188
  266. 266.Y. Fan, F. Tian, Y. Xia, T. Qin, X.-Y. Li, T.-Y. Liu, Searching better architectures for neural machine translation, IEEE/ACM Transactions on Audio, Speech, and Language Processing.
  267. 267.Y. Jiang, C. Hu, T. Xiao, C. Zhang, J. Zhu, Improved differentiable architecture search for language modeling and named entity recognition, in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Association for Computational Linguistics, Hong Kong, China, 2019, pp. 3585–3590. doi:10.18653/v1/D19-1367. URL https://www.aclweb.org/anthology/D19-1367
  268. 268.J. Chen, K. Chen, X. Chen, X. Qiu, X. Huang, Exploring shared structures and hierarchies for multiple nlp tasks, arXiv preprint arXiv:1808.07658.
  269. 269.H. Mazzawi, X. Gonzalvo, A. Kracun, P. Sridhar, N. Subrahmanya, I. Lopez-Moreno, H.-J. Park, P. Violette, Improving keyword spotting and language identification via neural architecture search at scale., in: INTERSPEECH, 2019, pp. 1278–1282.
  270. 270.Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, S. Han, Amc: Automl for model compression and acceleration on mobile devices, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 784–800.
  271. 271.X. Xiao, Z. Wang, S. Rajasekaran, Autoprune: Automatic network pruning by regularizing auxiliary parameters, in: H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alch´e-Buc, E. B. Fox, R. Garnett (Eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, 2019, pp. 13681–13691. URL https://proceedings.neurips.cc/paper/2019/hash/4efc9e02abdab6b6166251918570a307-Abstract.html
  272. 272.R. Zhao, W. Luk, Efficient structured pruning and architecture searching for group convolution, in: Proceedings of the IEEE International Conference on Computer Vision Workshops, 2019, pp. 0–0.
  273. 273.T. Wang, K. Wang, H. Cai, J. Lin, Z. Liu, H. Wang, Y. Lin, S. Han, APQ: joint search for network architecture, pruning and quantization policy, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 2075–2084. doi:10.1109/CVPR42600.2020.00215. URL https://doi.org/10.1109/CVPR42600.2020.00215
  274. 274.X. Dong, Y. Yang, Network pruning via transformable architecture search, in: H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alch´e-Buc, E. B. Fox, R. Garnett (Eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, 2019, pp. 759–770. URL https://proceedings.neurips.cc/paper/2019/hash/a01a0380ca3c61428c26a231f0e49a09-Abstract.html
  275. 275.Q. Huang, K. Zhou, S. You, U. Neumann, Learning to prune filters in convolutional neural networks (2018). arXiv:1801.07365.
  276. 276.Y. He, P. Liu, L. Zhu, Y. Yang, Meta filter pruning to accelerate deep convolutional neural networks (2019). arXiv:1904.03961.
  277. 277.T.-W. Chin, C. Zhang, D. Marculescu, Layer-compensated pruning for resource-constrained convolutional neural networks (2018). arXiv:1810.00518.
  278. 278.K. Zhou, Q. Song, X. Huang, X. Hu, Auto-gnn: Neural architecture search of graph neural networks, arXiv preprint arXiv:1909.03184.
  279. 279.C. He, M. Annavaram, S. Avestimehr, Fednas: Federated deep learning via neural architecture search (2020). arXiv:2004.08546.
  280. 280.H. Zhu, Y. Jin, Real-time federated evolutionary neural architecture search, arXiv preprint arXiv:2003.02793.
  281. 281.C. Li, X. Yuan, C. Lin, M. Guo, W. Wu, J. Yan, W. Ouyang, AM-LFS: automl for loss function search, in: 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, IEEE, 2019, pp. 8409–8418. doi:10.1109/ICCV.2019.00850. URL https://doi.org/10.1109/ICCV.2019.00850
  282. 282.B. Ru, C. Lyle, L. Schut, M. van der Wilk, Y. Gal, Revisiting the train loss: an efficient performance estimator for neural architecture search, arXiv preprint arXiv:2006.04492.
  283. 283.P. Ramachandran, B. Zoph, Q. V. Le, Searching for activation functions (2017). arXiv:1710.05941.
  284. 284.H. Wang, H. Wang, K. Xu, Evolutionary recurrent neural network for image captioning, Neurocomputing.
  285. 285.L. Wang, Y. Zhao, Y. Jinnai, Y. Tian, R. Fonseca, Neural architecture search using deep neural networks and monte carlo tree search, arXiv preprint arXiv:1805.07440.
  286. 286.P. Zhao, K. Xiao, Y. Zhang, K. Bian, W. Yan, Amer: Automatic behavior modeling and interaction exploration in recommender system, arXiv preprint arXiv:2006.05933.
  287. 287.X. Zhao, C. Wang, M. Chen, X. Zheng, X. Liu, J. Tang, Autoemb: Automated embedding dimensionality search in streaming recommendations, arXiv preprint arXiv:2002.11252.
  288. 288.W. Cheng, Y. Shen, L. Huang, Differentiable neural input search for recommender systems, arXiv preprint arXiv:2006.04466.
  289. 289.E. Real, C. Liang, D. R. So, Q. V. Le, Automl-zero: Evolving machine learning algorithms from scratch, in: Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Vol. 119 of Proceedings of Machine Learning Research, PMLR, 2020, pp. 8007–8019. URL http://proceedings.mlr.press/v119/real20a.html
  290. 290.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, Language models are unsupervised multitask learners, OpenAI Blog 1 (2019) 8.
  291. 291.D. Wang, C. Gong, Q. Liu, Improving neural language modeling via adversarial training, in: K. Chaudhuri, R. Salakhutdinov (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, PMLR, 2019, pp. 6555–6565. URL http://proceedings.mlr.press/v97/wang19f.html
  292. 292.A. Zela, T. Elsken, T. Saikia, Y. Marrakchi, T. Brox, F. Hutter, Understanding and robustifying differentiable architecture search, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=H1gDNyrKDS
  293. 293.S. KOTYAN, D. V. VARGAS, Is neural architecture search a way forward to develop robust neural networks?, Proceedings of the Annual Conference of JSAI JSAI2020 (2020) 2K1ES203–2K1ES203.
  294. 294.M. Guo, Y. Yang, R. Xu, Z. Liu, D. Lin, When NAS meets robustness: In search of robust architectures against adversarial attacks, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 628–637. doi:10.1109/CVPR42600.2020.00071. URL https://doi.org/10.1109/CVPR42600.2020.00071
  295. 295.Y. Chen, Q. Song, X. Liu, P. S. Sastry, X. Hu, On robustness of neural architecture search under label noise, in: Frontiers in Big Data, 2020.
  296. 296.D. V. Vargas, S. Kotyan, Evolving robust neural architectures to defend from adversarial attacks, arXiv preprint arXiv:1906.11667.
  297. 297.J. Yim, D. Joo, J. Bae, J. Kim, A gift from knowledge distillation: Fast optimization, network minimization and transfer learning, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, IEEE Computer Society, 2017, pp. 7130–7138. doi:10.1109/CVPR.2017.754. URL https://doi.org/10.1109/CVPR.2017.754
  298. 298.G. Squillero, P. Burelli, Applications of Evolutionary Computation: 19th European Conference, EvoApplications 2016, Porto, Portugal, March 30–April 1, 2016, Proceedings, Vol. 9597, Springer, 2016.
  299. 299.M. Feurer, A. Klein, K. Eggensperger, J. T. Springenberg, M. Blum, F. Hutter, Efficient and robust automated machine learning, in: C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, R. Garnett (Eds.), Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, 2015, pp. 2962–2970. URL https://proceedings.neurips.cc/paper/2015/hash/11d0e6287202fced83f79975ec59a3a6-Abstract.html
  300. 300.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, E. Duchesnay, Scikit-learn: Machine learning in Python, Journal of Machine Learning Research 12 (2011) 2825–2830.
  301. 301.A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, S. Chintala, Pytorch: An imperative style, high-performance deep learning library, in: H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alch´e-Buc, E. B. Fox, R. Garnett (Eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, 2019, pp. 8024–8035. URL https://proceedings.neurips.cc/paper/2019/hash/bdbca288fee7f92f2bfa9f7012727740-Abstract.html
  302. 302.F. Chollet, et al., Keras, https://github.com/fchollet/keras (2015).
  303. 303.NNI (Neural Network Intelligence), 2020. URL https://github.com/microsoft/nni
  304. 304.M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, X. Zheng, Tensorflow: A system for large-scale machine learning (2016). arXiv:1605.08695.
  305. 305.Vega, 2020. URL https://github.com/huawei-noah/vega
  306. 306.R. Pasunuru, M. Bansal, Continual and multi-task architecture search, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, Florence, Italy, 2019, pp. 1911–1922. doi:10.18653/v1/P19-1185. URL https://www.aclweb.org/anthology/P19-1185
  307. 307.J. Kim, S. Lee, S. Kim, M. Cha, J. K. Lee, Y. Choi, Y. Choi, D.-Y. Cho, J. Kim, Auto-meta: Automated gradient based meta learner search, arXiv preprint arXiv:1806.06927.
  308. 308.D. Lian, Y. Zheng, Y. Xu, Y. Lu, L. Lin, P. Zhao, J. Huang, S. Gao, Towards fast adaptation of neural architectures with meta learning, in: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, OpenReview.net, 2020. URL https://openreview.net/forum?id=r1eowANFvr
  309. 309.T. Elsken, B. Staffler, J. H. Metzen, F. Hutter, Meta-learning of neural architectures for few-shot learning, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, IEEE, 2020, pp. 12362–12372. doi:10.1109/CVPR42600.2020.01238. URL https://doi.org/10.1109/CVPR42600.2020.01238
  310. 310.C. Liu, P. Doll´ar, K. He, R. Girshick, A. Yuille, S. Xie, Are labels necessary for neural architecture search? (2020). arXiv:2003.12056.
  311. 311.Z. Li, D. Hoiem, Learning without forgetting, IEEE transactions on pattern analysis and machine intelligence 40 (12) (2018) 2935–2947.
  312. 312.S. Rebuffi, A. Kolesnikov, G. Sperl, C. H. Lampert, icarl: Incremental classifier and representation learning, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, IEEE Computer Society, 2017, pp. 5533–5542. doi:10.1109/CVPR.2017.587. URL https://doi.org/10.1109/CVPR.2017.587

Citation

MLA
He, X., et al. “AutoML: A Survey of the State-of-the-art”. Knowledge-Based Systems, vol. 212, 2021, p. 106622, https://doi.org/10.1016/j.knosys.2020.106622.
APA
He, X., Zhao, K., & Chu, X. (2021). AutoML: A survey of the state-of-the-art. Knowledge-Based Systems, 212, 106622. https://doi.org/10.1016/j.knosys.2020.106622
Chicago
He, X., K. Zhao, and X. Chu. 2021. “AutoML: A Survey of the State-of-the-art”. Knowledge-Based Systems 212: 106622. https://doi.org/10.1016/j.knosys.2020.106622.
Harvard
He, X., Zhao, K. and Chu, X. (2021) “AutoML: A survey of the state-of-the-art”, Knowledge-Based Systems, 212, p. 106622. Available at: https://doi.org/10.1016/j.knosys.2020.106622.
Vancouver
1. He X, Zhao K, Chu X (2021) AutoML: A survey of the state-of-the-art. Knowledge-Based Systems 212:106622

BibTeX

@article{He_2021, title={AutoML: A survey of the state-of-the-art}, volume={212}, ISSN={0950-7051}, url={http://dx.doi.org/10.1016/j.knosys.2020.106622}, DOI={10.1016/j.knosys.2020.106622}, journal={Knowledge-Based Systems}, publisher={Elsevier BV}, author={He, Xin and Zhao, Kaiyong and Chu, Xiaowen}, year={2021}, month=Jan, pages={106622} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/