A Framework and Benchmark for Deep Batch Active Learning for Regression

David HolzmüllerViktor ZaverkinJohannes KästnerIngo Steinwart

article2023JMLR73 citations

Presents a modular kernel-based framework, an efficient neural tangent kernel sketching technique paired with a novel clustering selection method, and an open-source 15-dataset benchmark to advance deep batch active learning for regression.

Listen

Supervised deep learning models have achieved remarkable success across diverse predictive applications, but their performance typically depends on access to large volumes of labeled data. In many practical engineering, physical, and scientific regression problems, obtaining accurate labels requires costly physical experiments, complex numerical simulations, or intensive manual measurements. While active learning addresses this by selectively querying informative data points, evaluating neural networks sequentially after each single label is computationally prohibitive and blocks parallel workflows. Batch-mode deep active learning overcomes this bottleneck by selecting batches of unlabeled data simultaneously, yet systematic frameworks and comprehensive benchmarks for continuous regression tasks remain underdeveloped.

To address this gap, the article establishes a modular, unified framework for constructing batch active learning algorithms and introduces a standardized tabular regression benchmark to evaluate their effectiveness. The core objective is to systematically analyze and improve sample efficiency in deep regression models without requiring changes to underlying neural network architectures or training routines.

Methodologically, the framework decomposes active learning workflows into three modular components: a base kernel capturing model representations, kernel transformations for statistical modeling and efficiency, and an iterative selection method that queries candidate points. The authors evaluate this approach using an open-source benchmark of 15 large tabular regression datasets spanning up to 379 features and hundreds of thousands of candidate points. Experiments evaluate fully connected neural networks across 20 randomized splits and 16 sequential batch acquisitions, tracking metrics including root mean squared error, mean absolute error, tail quantiles, and maximum error.

The findings establish that replacing standard last-layer feature representations with a sketched, finite-width neural tangent kernel consistently improves predictive accuracy across selection strategies. Furthermore, random feature sketching drastically reduces computational overhead with negligible loss in accuracy. Crucially, the authors' newly developed clustering selection strategy—Largest Cluster Maximum Distance—consistently outperforms existing state-of-the-art methods in both root mean squared error and mean absolute error. The top-performing active learning configurations achieve the predictive accuracy of random sampling using roughly half as much labeled data, while executing batch selections in just seconds on modern hardware.

These performance gains demonstrate substantial real-world value by significantly cutting experimental labeling costs, reducing turnaround times, and boosting model reliability. Tail error improvements also suggest stronger robustness against extreme modeling failures. Additionally, the article identifies that datasets displaying higher initial prediction error variance derive the greatest relative benefit from active batch selection, offering practitioners a practical indicator for when to implement these methods.

Organizations training deep regression models on expensive data should adopt batch active learning using sketched neural tangent representations combined with balanced clustering-based selection methods. For standard accuracy targets, Largest Cluster Maximum Distance offers the best performance profile, whereas geometric distance maximization remains a strong alternative if worst-case errors or covariate shifts are primary concerns.

Confidence in these findings is bolstered by rigorous statistical testing across multiple random seeds, datasets, and activation functions. However, practitioners should note that the evaluation is confined to fully connected networks on tabular data without explicit distribution shifts between candidate and operational test sets. Further validation is warranted before generalizing these conclusions to non-tabular modalities, custom domain architectures, or streaming data environments.

  • Paper: Active Learning with Statistical Models, David Cohn et al. (1996). Its variance-based query criteria for locally weighted regression provide a foundation for understanding how active selection can target informative observations in continuous prediction tasks.
  • Paper: Active Learning for Convolutional Neural Networks: A Core-Set Approach, Ozan Sener et al. (2018). Its core-set formulation makes the batch-selection challenge concrete: selecting a diverse subset helps avoid redundant queries, a central concern in the source’s clustering methods.
  • Paper: Neural Network Ensembles, Cross Validation, and Active Learning, Anders Krogh et al. (1994). Its use of neural-network disagreement to guide active selection of continuous-function examples introduces a model-based selection idea that helps contextualize the source’s regression strategies.
  • Paper: A Survey of Deep Active Learning, Pengzhen Ren et al. (2020). Its survey organizes deep active learning’s batch, uncertainty, and diversity approaches, giving readers the methodological map needed to place the source’s framework and benchmark.

No sufficiently relevant recommendations were found.

Cover for A Framework and Benchmark for Deep Batch Active Learning for Regression

Abstract

The acquisition of labels for supervised learning can be expensive. To improve the sample efficiency of neural network regression, we study active learning methods that adaptively select batches of unlabeled data for labeling. We present a framework for constructing such methods out of (network-dependent) base kernels, kernel transformations, and selection methods. Our framework encompasses many existing Bayesian methods based on Gaussian process approximations of neural networks as well as non-Bayesian methods. Additionally, we propose to replace the commonly used last-layer features with sketched finite-width neural tangent kernels and to combine them with a novel clustering method. To evaluate different methods, we introduce an open-source benchmark consisting of 15 large tabular regression data sets. Our proposed method outperforms the state-of-the-art on our benchmark, scales to large data sets, and works out-of-the-box without adjusting the network architecture or training code. We provide open-source code that includes efficient implementations of all kernels, kernel transformations, and selection methods, and can be used for reproducing our results.

Table of Contents

  • 1. Introduction
  • 1.1 Contributions
  • 2. Problem Setting
  • 2.1 Regression with Fully-Connected Neural Networks
  • 2.2 Batch Mode Active Learning
  • 3. Related Work
  • 3.1 Uncertainty Measures and Kernel Approximations
  • 3.2 Selection Methods
  • 3.3 Data Sets
  • 4. Kernels
  • 4.1 Base Kernels
  • 4.1.1 Linear Kernel
  • 4.1.3 Last-layer Kernel
  • 4.1.4 Infinite-width NNGP
  • 4.2 Kernel Transformations
  • 4.2.1 Scaling
  • 4.2.2 Gaussian Process Posterior Transformation
  • 4.2.3 Sketching
  • 4.2.4 Ensembling
  • 4.2.5 ACS Random Features Transformation
  • 4.2.6 ACS Gradient Transformation
  • 4.3 Discussion
  • 5. Selection Methods
  • 5.1 Iterative Selection Methods
  • 5.2 Specific Methods
  • 5.2.1 Random Selection
  • 5.2.2 Naive Active Learning
  • 5.2.3 Greedy Determinant Maximization
  • 5.2.4 Greedy total uncertainty minimization
  • 5.2.5 Frank-Wolfe optimization
  • 5.2.6 Greedy distance maximization
  • 5.2.7 k-means++ seeding
  • 5.2.8 Largest cluster maximum distance
  • 5.2.9 Other options
  • 5.3 Discussion
  • 6. Experiments
  • 6.1 Comparison to Existing Methods
  • 6.2 Evaluated Combinations
  • 6.3 Best Kernels and Modes for each Selection Method
  • 6.4 Comparison of Selection Methods
  • 6.5 When should BMDAL be Applied?
  • 7. Conclusion
  • 7.1 Limitations
  • 7.2 Remaining Questions
  • Acknowledgments
  • Appendix A. Overview
  • Appendix B. Details on Base Kernels
  • B.1 NNGP Kernel
  • Appendix C. Details on Kernel Transformations
  • C.1 Gaussian Process Posterior Transformation
  • C.2 Sketching
  • C.3 ACS Random Features Transformation
  • Appendix D. Details on Selection Methods
  • D.1 Iterative Selection Scheme
  • D.2 Random
  • D.3 MaxDiag
  • D.4 MaxDet
  • D.4.1 Equivalence of MaxDet to Non-batch Mode Active Learning With Fixed Kernel
  • D.4.2 Equivalence of MaxDet to BatchBALD on a GP
  • D.4.3 Equivalence of MaxDet to the P-greedy Algorithm
  • D.4.4 Relation to the Greedy Algorithm for D-optimal Design
  • D.4.5 Kernel-space Implementation of MaxDet
  • D.4.6 Feature-space Implementation of MaxDet
  • D.5 Bait
  • D.5.2 Forward Version
  • D.5.3 Forward-Backward Version
  • D.6 FrankWolfe
  • D.7 MaxDist
  • D.8 KMeansPP
  • D.9 LCMD
  • Appendix E. Details on Experiments
  • E.1 Data Sets
  • E.2 Preprocessing
  • E.3 Neural Network Configuration
  • E.4 Results
  • References

Knowls

  1. Knowl 1 — Composable kernel framework for batch active learning

    model/method

    The paper represents a pool-based batch active-learning method as three components: a base kernel, a sequence of kernel transformations, and a selection rule. The base kernel measures similarity between inputs and may depend on a trained neural network; transformations can, for example, encode posterior uncertainty or reduce feature dimension; the selection rule uses the resulting kernel to choose a batch from the unlabeled pool. This composition includes Bayesian approaches based on Gaussian-process or Laplace approximations as well as geometric selection methods, and allows the same kernel to be paired with different selection rules.

    Two modes distinguish how a method accounts for information already in the labeled data. In P mode, informativeness is incorporated through a kernel transformed using the current training inputs. In TP mode, selection promotes diversity from the union of the training inputs and the batch selected so far. The selection rule receives the kernel, training inputs, pool inputs, and requested batch size; it does not directly receive training labels.

  2. Knowl 2 — LCMD: deterministic selection from the largest distance-weighted cluster

    algorithm

    LCMD (largest cluster maximum distance) is a deterministic, iterative selection method intended to promote both diversity and representativity of the pool. Let kk be a positive semidefinite kernel and define the induced distance dk(x,x′)=k(x,x)+k(x′,x′)−2k(x,x′)d_k(x,x')=\sqrt{k(x,x)+k(x',x')-2k(x,x')}. Let XremX_{\mathrm{rem}} be the pool points not yet selected, and let XselX_{\mathrm{sel}} contain already selected points. In TP mode, XselX_{\mathrm{sel}} initially contains the labeled training inputs; in P mode it is initially empty.

    Input: Kernel kk, initial selected set XselX_{\mathrm{sel}}, pool XpoolX_{\mathrm{pool}}, batch size NbatchN_{\mathrm{batch}}
    Output: A batch XbatchX_{\mathrm{batch}} of NbatchN_{\mathrm{batch}} pool points
    Set XbatchX_{\mathrm{batch}} to the empty set
    If XselX_{\mathrm{sel}} is empty, add to XbatchX_{\mathrm{batch}} the remaining point maximizing k(x,x)k(x,x)
    While ∣Xbatch∣<Nbatch|X_{\mathrm{batch}}| < N_{\mathrm{batch}}:
        Set Xrem=Xpool∖XbatchX_{\mathrm{rem}} = X_{\mathrm{pool}} \setminus X_{\mathrm{batch}}
        For each xx in XremX_{\mathrm{rem}}, assign xx to a nearest center c(x)c(x) in Xsel∪XbatchX_{\mathrm{sel}} \cup X_{\mathrm{batch}}
        For each center zz in Xsel∪XbatchX_{\mathrm{sel}} \cup X_{\mathrm{batch}}, compute
            s(z)=∑x∈Xrem:c(x)=zdk(x,z)2s(z) = \sum_{x \in X_{\mathrm{rem}}: c(x)=z} d_k(x,z)^2
        Find a center z∗z^* with maximum s(z)s(z)
        Add to XbatchX_{\mathrm{batch}} a point in the cluster of z∗z^* maximizing dk(x,z∗)d_k(x,z^*)
    Return XbatchX_{\mathrm{batch}}

    Ties may be resolved arbitrarily. The cluster score is the sum of squared distances of its remaining points to its center, so the method focuses on a cluster with substantial distance-weighted mass, then chooses its most distant point. With NpoolN_{\mathrm{pool}} pool points, NselN_{\mathrm{sel}} selected or initial-center points, and kernel-evaluation time TkT_k, the stated runtime is O(NpoolNsel(Tk+1))O(N_{\mathrm{pool}}N_{\mathrm{sel}}(T_k+1)) and memory use is O(Ncand)O(N_{\mathrm{cand}}), where Ncand=Npool+∣Xsel∣N_{\mathrm{cand}}=N_{\mathrm{pool}}+|X_{\mathrm{sel}}|.

  3. Knowl 3 — Finite-width neural tangent kernel as a neural-network-dependent similarity

    model/method

    For a trained neural network fθT:Rd→Rf_{\theta_T}:\mathbb{R}^d\to\mathbb{R} with parameter vector θT\theta_T, the full-gradient feature map is ϕgrad(x)=∇θfθT(x)\phi_{\mathrm{grad}}(x)=\nabla_\theta f_{\theta_T}(x), and its kernel is the finite-width neural tangent kernel

    kgrad(x,x′)=⟨∇θfθT(x),∇θfθT(x′)⟩.k_{\mathrm{grad}}(x,x')=\langle \nabla_\theta f_{\theta_T}(x),\nabla_\theta f_{\theta_T}(x')\rangle.

    This kernel measures similarity in the network's parameter-gradient features at the trained parameters; it is motivated by the first-order linearization of the network around θT\theta_T. In a fully connected network, the full parameter-gradient feature need not be explicitly formed to evaluate the kernel: each layer contributes a product of an inner product between its augmented input activations and an inner product between output derivatives with respect to that layer's preactivations. Summing these products over layers yields the full-gradient kernel. For a network with LL layers of hidden width mm, cached activations and derivatives permit reducing the kernel-evaluation cost from Θ(m2L)\Theta(m^2L) to Θ(mL)\Theta(mL), with a corresponding reduction for stored precomputed features. The paper's experiments use a sketched version of this trained, finite-width kernel rather than the commonly used last-layer-only feature kernel.

  4. Knowl 4 — Gaussian sketching preserves kernel distances with a dimension-independent feature bound

    theoretical result

    For a finite-dimensional kernel with feature map ϕ:Rd→Rdfeat\phi:\mathbb{R}^d\to\mathbb{R}^{d_{\mathrm{feat}}}, a Gaussian sketch with pp features is ϕsketch(x)=p−1/2Uϕ(x)\phi_{\mathrm{sketch}}(x)=p^{-1/2}U\phi(x), where U∈Rp×dfeatU\in\mathbb{R}^{p\times d_{\mathrm{feat}}} has independent standard-normal entries. The induced kernel is an unbiased estimate of the original kernel. Define dk(x,x′)=∥ϕ(x)−ϕ(x′)∥2d_k(x,x')=\|\phi(x)-\phi(x')\|_2 and let X⊂RdX\subset\mathbb{R}^d be finite. For ε,δ∈(0,1)\varepsilon,\delta\in(0,1), if

    p≥8log⁡(∣X∣2/δ)ε2,p\geq \frac{8\log(|X|^2/\delta)}{\varepsilon^2},

    then, with probability at least 1−δ1-\delta, every pair in XX satisfies

    (1−ε)dk(x,x′)≤dksketch(x,x′)≤(1+ε)dk(x,x′).(1-\varepsilon)d_k(x,x')\leq d_{k_{\mathrm{sketch}}}(x,x')\leq(1+\varepsilon)d_k(x,x').

    The feature-count bound does not depend on the original feature dimension dfeatd_{\mathrm{feat}}. For a product of kernels, the paper also gives an efficient sketch: independently sketch the two factor feature maps to pp dimensions, take their element-wise product, and multiply by p\sqrt{p}. This avoids materializing the tensor-product feature space. In the benchmark, sketching the finite-width neural tangent kernel to 512 features retains its selection accuracy while substantially reducing runtime.

  5. Knowl 5 — Posterior kernel transformation and automatic scale normalization

    equation

    Given a kernel kk with feature map ϕ\phi, training inputs XtrainX_{\mathrm{train}} of size NtrainN_{\mathrm{train}}, and assumed Gaussian observation-noise variance σ2>0\sigma^2>0, the paper uses a Gaussian-process posterior covariance transformation. Before conditioning, the kernel can be rescaled to give it mean unit diagonal on the training inputs:

    kscale(x,x′)=λ2k(x,x′),λ=(1Ntrain∑x∈Xtraink(x,x))−1/2.k_{\mathrm{scale}}(x,x')=\lambda^2 k(x,x'),\qquad \lambda=\left(\frac{1}{N_{\mathrm{train}}}\sum_{x\in X_{\mathrm{train}}}k(x,x)\right)^{-1/2}.

    For a feature map ϕ\phi of the kernel being transformed, the posterior covariance of the latent function is

    kpost(x,x′)=ϕ(x)⊤(σ−2ϕ(Xtrain)⊤ϕ(Xtrain)+I)−1ϕ(x′),k_{\mathrm{post}}(x,x')=\phi(x)^\top\left(\sigma^{-2}\phi(X_{\mathrm{train}})^\top\phi(X_{\mathrm{train}})+I\right)^{-1}\phi(x'),

    where ϕ(Xtrain)\phi(X_{\mathrm{train}}) is the matrix whose rows are the training-input feature vectors and II is the identity matrix of the matching feature dimension. The transformed covariance depends on training inputs but not on their observed labels. For the last-layer kernel this is Bayesian linear regression on the network's last-layer features; applying the same transformation to the full-gradient kernel gives a generalized Gauss–Newton/Laplace-style approximation under the paper's stated linearization assumptions.

  6. Knowl 6 — Benchmark design and neural-network training protocol

    experimental setup

    The benchmark contains 15 large tabular regression datasets: sgemm, wec_sydney, ct_slices, kegg_undir, online_video, query, poker, road, mlr_knn_rng, fried, diamonds, methane, stock, protein, and sarcos. Their input dimensions range from 2 to 379, and initial pool sizes range from 31,335 to 198,720. The evaluated regressor is a fully connected neural network with two hidden layers of 512 units each. Networks are trained with Adam for 256 epochs using mini-batches of 256 and early stopping on a 1,024-example validation set. The main results use ReLU; most comparisons were also repeated with SiLU.

    Each active-learning run starts with 256 labeled training examples and acquires 16 batches of 256 examples. The authors repeat runs 20 times with different initializations and data splits. They measure test MAE, RMSE, the 95% and 99% error quantiles, and maximum error, averaging logarithms of each metric across repetitions and, where specified, datasets and acquisition steps. Posterior-based methods use σ2=10−6\sigma^2=10^{-6}; sketched kernels and random-feature transformations target 512 features. The benchmark was designed for scalable pool-based comparisons without changing the network architecture or training procedure.

  7. Knowl 7 — LCMD with the sketched full-gradient kernel beats literature baselines on mean log RMSE

    empirical result

    On the 15-dataset benchmark, the proposed combination LCMD-TP with the 512-feature sketched finite-width neural tangent kernel achieved the lowest averaged logarithmic RMSE among the compared methods, for both ReLU and SiLU networks. The table reports the mean log RMSE averaged across datasets, repetitions, and active-learning steps; lower values indicate lower RMSE. The compared methods were adapted to regression within the paper's kernel framework.

    Method ReLU SiLU
    Random (no active learning) -1.401 -1.406
    BALD with last-layer GP -1.285 -1.300
    BatchBALD with last-layer GP -1.463 -1.467
    Bait -1.541 -1.522
    ACS-FW -1.439 -1.437
    Core-Set / FF-Active -1.491 -1.515
    BADGE with last-layer GP uncertainty -1.530 -1.484
    LCMD-TP with kgrad→sketch(512)k_{\mathrm{grad}\to\mathrm{sketch}(512)} -1.590 -1.597

    The corresponding LCMD configuration was not just better than random selection: it also improved on the best reported literature baseline in this comparison, Bait for ReLU and Bait for SiLU.

  8. Knowl 8 — Best-performing kernel and mode vary by selection rule

    data/table

    The authors selected a kernel and P/TP mode for each selection method using average log RMSE on the ReLU benchmark. These configurations show that selection-rule comparisons depend on kernel choice: LCMD-TP paired with the sketched full-gradient kernel was best overall, while posterior-based transformations or ACS transformations were selected for several other rules. Average selection time is seconds per batch across datasets and acquisition steps on an NVIDIA RTX 3090; mean log RMSE is averaged across datasets, repetitions, and steps.

    Selection rule Selected kernel Mean log RMSE Time (s)
    Random none -1.401 0.001
    MaxDiag kgrad→sketch(512)→acs−rf(512)k_{\mathrm{grad}\to\mathrm{sketch}(512)\to\mathrm{acs-rf}(512)} -1.370 0.650
    MaxDet-P kgrad→sketch(512)→Xtraink_{\mathrm{grad}\to\mathrm{sketch}(512)\to X_{\mathrm{train}}} -1.512 0.770
    Bait-F-P kgrad→sketch(512)→Xtraink_{\mathrm{grad}\to\mathrm{sketch}(512)\to X_{\mathrm{train}}} -1.585 1.508
    FrankWolfe-P kgrad→sketch(512)→acs−rf−hyper(512)k_{\mathrm{grad}\to\mathrm{sketch}(512)\to\mathrm{acs-rf-hyper}(512)} -1.542 0.823
    MaxDist-P kgrad→sketch(512)→Xtraink_{\mathrm{grad}\to\mathrm{sketch}(512)\to X_{\mathrm{train}}} -1.514 0.713
    KMeansPP-P kgrad→sketch(512)→acs−rf(512)k_{\mathrm{grad}\to\mathrm{sketch}(512)\to\mathrm{acs-rf}(512)} -1.569 0.836
    LCMD-TP kgrad→sketch(512)k_{\mathrm{grad}\to\mathrm{sketch}(512)} -1.590 0.981

    Here, kgradk_{\mathrm{grad}} is the finite-width full-gradient kernel, “sketch(512)” reduces its feature map to 512 dimensions, and XtrainX_{\mathrm{train}} denotes the rescale-then-posterior transformation. P and TP denote pool-only and training-plus-pool selection modes, respectively.

  9. Knowl 9 — Performance patterns across metrics, datasets, and batch sizes

    empirical result

    Across the benchmark, replacing the last-layer feature kernel with the finite-width full-gradient kernel generally improved results across selection methods, and sketching that kernel to 512 features preserved accuracy while making selection faster. LCMD-TP with this sketched kernel reduced RMSE relative to random selection on 13 of 15 datasets and matched or outperformed the other selected active-learning methods on 8 of 15 datasets. In aggregate learning curves, the strongest active-learning methods reached random selection's average RMSE with about half as many labeled examples; for maximum error they needed even fewer. LCMD-TP performed best overall for MAE, RMSE, and the 95% quantile, while MaxDist, MaxDet, and Bait-F were among the strongest for maximum error.

    With a fixed final training-set size of 4,352, MaxDiag was particularly sensitive to acquisition batch size. The other compared methods showed almost no degradation through batch sizes of 1,024; the paper notes that this threshold may depend on training-set size and feature dimension. On these datasets, the initial-training-set ratio RMSE/MAE was strongly associated with the benefit of LCMD-TP over random selection: the reported Pearson correlation between error-distribution variation and relative sample-efficiency gain was approximately 0.880.88. This association is an empirical observation on the benchmark, not a guarantee for other datasets.

  10. Knowl 10 — Limits on generalizing the benchmark findings

    limitation

    The empirical conclusions come from large tabular regression datasets and fully connected neural networks, primarily with ReLU or SiLU activations. The paper states that it is unclear whether the findings transfer to substantially different applications such as drug discovery or atomistic machine learning, to smaller tabular datasets, or to newer tabular-network architectures. The benchmark also contains no distribution shift between the pool and test data, so it does not establish how the methods compare when those distributions differ.

Coverage note — The paper's detailed implementations and comparisons of established selection methods and the full per-dataset appendix results are omitted because they support the framework and benchmark rather than constitute separate new contributions.

References

  1. 1.Thomas D. Ahle, Michael Kapralov, Jakob BT Knudsen, Rasmus Pagh, Ameya Velingker, David P. Woodruff, and Amir Zandieh. Oblivious sketching of high-degree polynomial kernels. In ACM-SIAM Symposium on Discrete Algorithms, 2020.
  2. 2.Rahaf Aljundi, Nikolay Chumerin, and Daniel Olmeda Reino. Identifying wrongly predicted samples: A method for active learning. In Winter Conference on Applications of Computer Vision, 2022.
  3. 3.Christos Anagnostopoulos, Fotis Savva, and Peter Triantafillou. Scalable aggregation predictive analytics. Applied Intelligence, 48(9):2546–2567, 2018.
  4. 4.Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, Russ R. Salakhutdinov, and Ruosong Wang. On exact computation with an infinitely wide neural net. In Neural Information Processing Systems, 2019.
  5. 5.Rosa I. Arriaga and Santosh Vempala. Algorithmic theories of learning. In Foundations of Computer Science, 1999.
  6. 6.David Arthur and Sergei Vassilvitskii. k-means++: The advantages of careful seeding. In ACM-SIAM Symposium on Discrete Algorithms, 2007.
  7. 7.Jordan Ash, Surbhi Goel, Akshay Krishnamurthy, and Sham Kakade. Gone fishing: Neural active learning with Fisher embeddings. In Neural Information Processing Systems, 2021.
  8. 8.Jordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal. Deep batch active learning by diverse, uncertain gradient lower bounds. In International Conference on Learning Representations, 2019.
  9. 9.Alexander Atanasov, Blake Bordelon, and Cengiz Pehlevan. Neural networks as kernel learners: The silent alignment effect. In International Conference on Learning Representations, 2021.
  10. 10.Rafael Ballester-Ripoll, Enrique G. Paredes, and Renato Pajarola. Sobol tensor trains for global sensitivity analysis. Reliability Engineering & System Safety, 183:311–322, 2019.
  11. 11.Jörg Behler. Perspective: Machine learning potentials for atomistic simulations. Journal of Chemical Physics, 145(17):170901, 2016.
  12. 12.William H. Beluch, Tim Genewein, Andreas Nürnberger, and Jan M. Köhler. The power of ensembles for active learning in image classification. In Conference on Computer Vision and Pattern Recognition, 2018.
  13. 13.Marshall Bern and David Eppstein. Approximation algorithms for geometric problems. In Approximation Algorithms for NP-hard Problems, pages 296–345. PWS Publishing Company, 1996.
  14. 14.Christopher M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006.
  15. 15.Zalán Borsos, Mojmir Mutny, and Andreas Krause. Coresets via bilevel optimization for continual learning and streaming. In Neural Information Processing Systems, 2020.
  16. 16.Zalán Borsos, Marco Tagliasacchi, and Andreas Krause. Semi-supervised batch active learning via bilevel optimization. In Conference on Acoustics, Speech and Signal Processing, 2021.
  17. 17.Erdem Bıyık, Kenneth Wang, Nima Anari, and Dorsa Sadigh. Batch active learning using determinantal point processes. arXiv:1906.07975, 2019.
  18. 18.William F. Caselton and James V. Zidek. Optimal monitoring network designs. Statistics & Probability Letters, 2(4):223–227, 1984.
  19. 19.M. Emre Celebi, Hassan A. Kingravi, and Patricio A. Vela. A comparative study of efficient initialization methods for the k-means clustering algorithm. Expert Systems with Applications, 40(1):200–210, 2013.
  20. 20.Kathryn Chaloner and Isabella Verdinelli. Bayesian experimental design: A review. Statistical Science, pages 273–304, 1995.
  21. 21.Laming Chen, Guoxin Zhang, and Eric Zhou. Fast greedy map inference for determinantal point process to improve recommendation diversity. In Neural Information Processing Systems, 2018.
  22. 22.Lenaic Chizat, Edouard Oyallon, and Francis Bach. On lazy training in differentiable programming. In Neural Information Processing Systems, 2019.
  23. 23.Ali Civril and Malik Magdon-Ismail. Exponential inapproximability of selecting a maximum volume sub-matrix. Algorithmica, 65(1):159–176, 2013.
  24. 24.David A. Cohn. Neural network exploration using optimal experiment design. Neural Networks, 9(6):1071–1083, 1996.
  25. 25.Cody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman, Peter Bailis, Percy Liang, Jure Leskovec, and Matei Zaharia. Selection via proxy: Efficient data selection for deep learning. In International Conference on Learning Representations, 2019.
  26. 26.Erik Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen, Matthias Bauer, and Philipp Hennig. Laplace Redux-Effortless Bayesian Deep Learning. In Neural Information Processing Systems, 2021.
  27. 27.Stefano De Marchi, Robert Schaback, and Holger Wendland. Near-optimal data-independent point locations for radial basis function interpolation. Advances in Computational Mathematics, 23(3):317–330, 2005.
  28. 28.Tewodors Deneke, Habtegebreil Haile, Sébastien Lafond, and Johan Lilius. Video transcoding time prediction for proactive load balancing. In International Conference on Multimedia and Expo, 2014.
  29. 29.Dheeru Dua and Casey Graff. UCI Machine Learning Repository, 2017. URL http://archive.ics.uci.edu/ml.
  30. 30.Stefan Elfwing, Eiji Uchibe, and Kenji Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural Networks, 107:3–11, 2018.
  31. 31.David Eppstein, Sariel Har-Peled, and Anastasios Sidiropoulos. Approximate greedy clustering and distance selection for graph metrics. Journal of Computational Geometry, 11(1):629–652, 2020.
  32. 32.Runa Eschenhagen, Erik Daxberger, Philipp Hennig, and Agustinus Kristiadi. Mixtures of Laplace approximations for improved post-hoc uncertainty in deep learning. In NeurIPS 2021 Workshop on Bayesian Deep Learning, 2021.
  33. 33.Sebastian Farquhar, Yarin Gal, and Tom Rainforth. On statistical bias in active learning: How and when to fix it. In International Conference on Learning Representations, 2021.
  34. 34.Tomás Feder and Daniel Greene. Optimal algorithms for approximate clustering. In ACM Symposium on Theory of Computing, 1988.
  35. 35.Valerii V. Fedorov. Theory of Optimal Experiments. Academic Press, New York, 1972.
  36. 36.Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M. Roy, and Surya Ganguli. Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel. In Neural Information Processing Systems, 2020.
  37. 37.Marguerite Frank and Philip Wolfe. An algorithm for quadratic programming. Naval Research Logistics Quarterly, 3(1-2):95–110, 1956.
  38. 38.Jerome H. Friedman. Multivariate adaptive regression splines. The Annals of Statistics, pages 1–67, 1991.
  39. 39.Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In International Conference on Machine Learning, 2016.
  40. 40.Yarin Gal, Riashat Islam, and Zoubin Ghahramani. Deep bayesian active learning with image data. In International Conference on Machine Learning, 2017.
  41. 41.Yonatan Geifman and Ran El-Yaniv. Deep active learning over the long tail. arXiv:1711.00941, 2017.
  42. 42.Amirata Ghorbani, James Zou, and Andre Esteva. Data shapley valuation for efficient batch active learning. In Asilomar Conference on Signals, Systems, and Computers. IEEE, 2022.
  43. 43.Teofilo F. Gonzalez. Clustering to minimize the maximum intercluster distance. Theoretical Computer Science, 38:293–306, 1985.
  44. 44.Yury Gorishniy, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. Revisiting deep learning models for tabular data. In Neural Information Processing Systems, 2021.
  45. 45.Franz Graf, Hans-Peter Kriegel, Matthias Schubert, Sebastian Pölsterl, and Alexander Cavallaro. 2D image registration in ct images using radial image descriptors. In International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2011.
  46. 46.Insu Han, Haim Avron, Neta Shoham, Chaewon Kim, and Jinwoo Shin. Random features for the neural tangent kernel. arXiv:2104.01351, 2021.
  47. 47.Lars Kai Hansen and Peter Salamon. Neural network ensembles. IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(10):993–1001, 1990.
  48. 48.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In IEEE Conference on Computer Vision, 2015.
  49. 49.Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani, and Máté Lengyel. Bayesian active learning for classification and preference learning. arXiv:1112.5745, 2011.
  50. 50.Alexander Immer, Maciej Korzepa, and Matthias Bauer. Improving predictions of Bayesian neural nets via local linearization. In International Conference on Artificial Intelligence and Statistics, 2021.
  51. 51.Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural Tangent Kernel: Convergence and generalization in neural networks. In Neural Information Processing Systems, 2018.
  52. 52.William B. Johnson and Joram Lindenstrauss. Extensions of Lipschitz mappings into a Hilbert space. Contemporary Mathematics, 26, 1984.
  53. 53.Arlind Kadra, Marius Lindauer, Frank Hutter, and Josif Grabocka. Well-tuned simple nets excel on tabular datasets. In Neural Information Processing Systems, 2021.
  54. 54.Purushottam Kar and Harish Karnick. Random feature maps for dot product kernels. In Artificial Intelligence and Statistics, 2012.
  55. 55.Ioannis Katsavounidis, C.-C. Jay Kuo, and Zhen Zhang. A new initialization technique for generalized Lloyd iteration. IEEE Signal Processing Letters, 1(10):144–146, 1994.
  56. 56.Leonard Kaufman and Peter J. Rousseeuw. Finding groups in data: an introduction to cluster analysis. Wiley Series in Probability and Mathematical Statistics. Applied Probability and Statistics, 1990.
  57. 57.Manohar Kaul, Bin Yang, and Christian S. Jensen. Building accurate 3d spatial networks to enable next generation intelligent transportation systems. In International Conference on Mobile Data Management, 2013.
  58. 58.Ronald W. Kennard and Larry A. Stone. Computer aided design of experiments. Technometrics, 11(1):137–148, 1969.
  59. 59.Mohammad Emtiyaz E. Khan, Alexander Immer, Ehsan Abedi, and Maciej Korzepa. Approximate inference turns deep networks into gaussian processes. In Neural Information Processing Systems, 2019.
  60. 60.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015.
  61. 61.Andreas Kirsch, Joost Van Amersfoort, and Yarin Gal. BatchBALD: Efficient and diverse batch acquisition for deep Bayesian active learning. In Neural Information Processing Systems, 2019.
  62. 62.Andreas Krause, Ajit Singh, and Carlos Guestrin. Near-optimal sensor placements in Gaussian processes: Theory, efficient algorithms and empirical studies. Journal of Machine Learning Research, 9(2), 2008.
  63. 63.Agustinus Kristiadi, Matthias Hein, and Philipp Hennig. Being Bayesian, even just a bit, fixes overconfidence in relu networks. In International Conference on Machine Learning, 2020.
  64. 64.Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009.
  65. 65.Anders Krogh and Jesper Vedelsby. Neural network ensembles, cross validation, and active learning. In Neural Information Processing Systems, 1994.
  66. 66.Punit Kumar and Atul Gupta. Active learning query strategies for classification, regression, and clustering: a survey. Journal of Computer Science and Technology, 35(4):913–945, 2020.
  67. 67.J. Nathan Kutz. Deep learning in fluid dynamics. Journal of Fluid Mechanics, 814:1–4, 2017.
  68. 68.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In Neural Information Processing Systems, volume 30, 2017.
  69. 69.Pierre Simon Laplace. Mémoire sur la probabilité de causes par les évènements. Mémoires de Mathématique et de Physique, Presentés à l’Académie Royale des Sciences, par divers Savants & lus dans ses Assemblées. Tome Sixième, pages 621–656, 1774.
  70. 70.Alexander Lavin, Hector Zenil, Brooks Paige, David Krakauer, Justin Gottschlich, Tim Mattson, Anima Anandkumar, Sanjay Choudry, Kamil Rocki, Atılım Güneş Baydin, and others. Simulation intelligence: Towards a new generation of scientific methods. arXiv:2112.03235, 2021.
  71. 71.Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein. Deep neural networks as Gaussian processes. In International Conference on Learning Representations, 2018.
  72. 72.Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington. Wide neural networks of any depth evolve as linear models under gradient descent. In Neural Information Processing Systems, 2019.
  73. 73.Stuart Lloyd. Least squares quantization in PCM. IEEE Transactions on Information Theory, 28(2):129–137, 1982.
  74. 74.Philip M. Long. Properties of the after kernel. arXiv:2105.10585, 2021.
  75. 75.Miguel Lázaro-Gredilla and Aníbal R. Figueiras-Vidal. Marginalized neural network mixtures for large-scale regression. IEEE Transactions on Neural Networks, 21(8):1345–1351, 2010.
  76. 76.David JC MacKay. Bayesian interpolation. Neural Computation, 4(3):415–447, 1992a.
  77. 77.David JC MacKay. Information-based objective functions for active data selection. Neural Computation, 4(4):590–604, 1992b.
  78. 78.Vivek Madan, Mohit Singh, Uthaipon Tantipongpipat, and Weijun Xie. Combinatorial algorithms for optimal design. In Conference on Learning Theory, 2019.
  79. 79.Wesley J. Maddox, Pavel Izmailov, Timur Garipov, Dmitry P. Vetrov, and Andrew Gordon Wilson. A simple baseline for bayesian uncertainty in deep learning. In Neural Information Processing Systems, 2019.
  80. 80.Alexander G. de G. Matthews, Jiri Hron, Mark Rowland, Richard E. Turner, and Zoubin Ghahramani. Gaussian process behaviour in wide deep neural networks. In International Conference on Learning Representations, 2018.
  81. 81.Arash Mehrjou, Ashkan Soleymani, Andrew Jesson, Pascal Notin, Yarin Gal, Stefan Bauer, and Patrick Schwab. GeneDisco: A benchmark for experimental design in drug discovery. In International Conference on Learning Representations, 2021.
  82. 82.Mohamad Amin Mohamadi, Wonho Bae, and Danica J. Sutherland. Making look-ahead active learning strategies feasible with Neural Tangent Kernels. In Advances in Neural Information Processing Systems, 2022.
  83. 83.Douglas C. Montgomery. Design and Analysis of Experiments. John Wiley & Sons, 2017.
  84. 84.Radford M. Neal. Priors for Infinite Networks. Technical Report CRG-TR-94-1, Dept. of Computer Science, University of Toronto, 1994.
  85. 85.Mehdi Neshat, Bradley Alexander, Markus Wagner, and Yuanzhong Xia. A detailed comparison of meta-heuristic methods for optimising wave energy converter placements. In Proceedings of the Genetic and Evolutionary Computation Conference, 2018.
  86. 86.Manuel Nonnenmacher, David Reeb, and Ingo Steinwart. Which minimizer does my neural network converge to? In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, 2021.
  87. 87.Roman Novak, Jascha Sohl-Dickstein, and Samuel S. Schoenholz. Fast finite width neural tangent kernel. In International Conference on Machine Learning, 2022.
  88. 88.Sebastian W. Ober and Carl Edward Rasmussen. Benchmarking the neural linear model for regression. In Symposium on Advances in Approximate Bayesian Inference, 2019.
  89. 89.Rafail Ostrovsky, Yuval Rabani, Leonard J. Schulman, and Chaitanya Swamy. The effectiveness of Lloyd-type methods for the k-means problem. In IEEE Symposium on Foundations of Computer Science, 2006.
  90. 90.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, and Luca Antiga. Pytorch: An imperative style, high-performance deep learning library. In Neural Information Processing Systems, 2019.
  91. 91.Maryam Pazouki and Robert Schaback. Bases for kernel-based spaces. Journal of Computational and Applied Mathematics, 236(4):575–588, 2011.
  92. 92.Robert Pinsler, Jonathan Gordon, Eric Nalisnick, and José Miguel Hernández-Lobato. Bayesian batch active learning as sparse subset approximation. In Neural Information Processing Systems, 2019.
  93. 93.Remus Pop and Patric Fulop. Deep ensemble bayesian active learning: Addressing the mode collapse issue in monte carlo dropout via ensembles. arXiv:1811.03897, 2018.
  94. 94.Maziar Raissi, Paris Perdikaris, and George E. Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686–707, 2019.
  95. 95.Hiranmayi Ranganathan, Hemanth Venkateswara, Shayok Chakraborty, and Sethuraman Panchanathan. Deep active learning for image regression. In Deep Learning Applications, pages 113–135. Springer, Singapore, 2020.
  96. 96.Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Brij B. Gupta, Xiaojiang Chen, and Xin Wang. A survey of deep active learning. ACM Computing Surveys, 54(9):1–40, 2021.
  97. 97.Daniel J. Rosenkrantz, Richard E. Stearns, and Philip M. Lewis. An analysis of several heuristics for the traveling salesman problem. SIAM Journal on Computing, 6(3):563–581, 1977.
  98. 98.Gabriele Santin, Toni Karvonen, and Bernard Haasdonk. Sampling based approximation of linear functionals in reproducing kernel Hilbert spaces. BIT Numerical Mathematics, pages 1–32, 2021. Publisher: Springer.
  99. 99.Fotis Savva, Christos Anagnostopoulos, and Peter Triantafillou. Explaining aggregates for exploratory analytics. In International Conference on Big Data, 2018.
  100. 100.Nicol N. Schraudolph. Fast curvature matrix-vector products for second-order gradient descent. Neural Computation, 14(7):1723–1738, 2002.
  101. 101.Ozan Sener and Silvio Savarese. Active learning for convolutional neural networks: A core-set approach. In International Conference on Learning Representations, 2018.
  102. 102.Sambu Seo, Marko Wallat, Thore Graepel, and Klaus Obermayer. Gaussian process regression: Active data selection and test point rejection. In Mustererkennung 2000, pages 27–34. Springer, 2000.
  103. 103.Burr Settles. Active learning literature survey. Computer Sciences Technical Report 1648, University of Wisconsin–Madison, 2009.
  104. 104.H. Sebastian Seung, Manfred Opper, and Haim Sompolinsky. Query by committee. In Workshop on Computational Learning Theory, 1992.
  105. 105.Haozhe Shan and Blake Bordelon. A theory of neural tangent kernel alignment and its influence on training. arXiv:2105.14301, 2021.
  106. 106.Claude Elwood Shannon. A mathematical theory of communication. The Bell System Technical Journal, 27(3):379–423, 1948.
  107. 107.Paul Shannon, Andrew Markiel, Owen Ozier, Nitin S. Baliga, Jonathan T. Wang, Daniel Ramage, Nada Amin, Benno Schwikowski, and Trey Ideker. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome research, 13(11):2498–2504, 2003.
  108. 108.Apoorva Sharma, Navid Azizan, and Marco Pavone. Sketching curvature for efficient out-of-distribution detection for deep neural networks. In Uncertainty in Artificial Intelligence, 2021.
  109. 109.Neta Shoham and Haim Avron. Experimental design for overparameterized learning with application to single shot deep active learning. arXiv:2009.12820, 2020.
  110. 110.Kirstine Smith. On the standard deviations of adjusted and interpolated values of an observed polynomial function and its constants and the guidance they give towards a proper choice of the distribution of observations. Biometrika, 12(1/2):1–85, 1918.
  111. 111.Jasper Snoek, Oren Rippel, Kevin Swersky, Ryan Kiros, Nadathur Satish, Narayanan Sundaram, Mostofa Patwary, Mr Prabhat, and Ryan Adams. Scalable bayesian optimization using deep neural networks. In International Conference on Machine Learning, 2015.
  112. 112.Gowthami Somepalli, Micah Goldblum, Avi Schwarzschild, C. Bayan Bruss, and Tom Goldstein. SAINT: Improved neural networks for tabular data via row attention and contrastive pre-training. arXiv:2106.01342, 2021.
  113. 113.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research, 15(1):1929–1958, 2014.
  114. 114.Ingo Steinwart and Andreas Christmann. Support vector machines. Springer Science & Business Media, 2008.
  115. 115.Evgenii Tsymbalov, Maxim Panov, and Alexander Shapeev. Dropout-based active learning for regression. In International Conference on Analysis of Images, Social Networks and Texts, 2018.
  116. 116.Joaquin Vanschoren, Jan N. van Rijn, Bernd Bischl, and Luis Torgo. OpenML: Networked science in machine learning. SIGKDD Explorations, 15(2):49–60, 2013.
  117. 117.Sethu Vijayakumar and Stefan Schaal. Locally weighted projection regression: An o(n) algorithm for incremental real time learning in high dimensional space. In International Conference on Machine Learning, 2000.
  118. 118.Abraham Wald. On the efficient design of statistical investigations. The Annals of Mathematical Statistics, 14(2):134–140, 1943.
  119. 119.Haonan Wang, Wei Huang, Ziwei Wu, Hanghang Tong, Andrew J. Margenot, and Jingrui He. Deep active learning by leveraging training dynamics. In Advances in Neural Information Processing Systems, 2022.
  120. 120.Zhilei Wang, Pranjal Awasthi, Christoph Dann, Ayush Sekhari, and Claudio Gentile. Neural active learning with performance guarantees. In Neural Information Processing Systems, 2021.
  121. 121.Lilian Weng. Learning with not enough data part 2: Active learning. lilianweng.github.io/lil-log, 2022. URL https://lilianweng.github.io/lil-log/2022/02/20/active-learning.html.
  122. 122.Tizian Wenzel, Gabriele Santin, and Bernard Haasdonk. A novel class of stabilized greedy kernel approximation algorithms: Convergence, stability and uniform point distribution. Journal of Approximation Theory, 262:105508, 2021.
  123. 123.Andrew G. Wilson and Pavel Izmailov. Bayesian deep learning and a probabilistic perspective of generalization. In Neural Information Processing Systems, 2020.
  124. 124.David P. Woodruff. Sketching as a tool for numerical linear algebra. Foundations and Trends in Theoretical Computer Science, 10(1–2):1–157, 2014.
  125. 125.Dongrui Wu. Pool-based sequential active learning for regression. IEEE Transactions on Neural Networks and Learning Systems, 30(5):1348–1359, 2018.
  126. 126.Henry P. Wynn. The sequential generation of D-optimum experimental designs. The Annals of Mathematical Statistics, 41(5):1655–1664, 1970.
  127. 127.Hwanjo Yu and Sungchul Kim. Passive sampling for regression. In International Conference on Data Mining, 2010.
  128. 128.Amir Zandieh, Insu Han, Haim Avron, Neta Shoham, Chaewon Kim, and Jinwoo Shin. Scaling neural tangent kernels via sketching and random features. In Neural Information Processing Systems, 2021.
  129. 129.Viktor Zaverkin and Johannes Kästner. Exploration of transferable and uniformly accurate neural network interatomic potentials using optimal experimental design. Machine Learning: Science and Technology, 2(3):035009, 2021.
  130. 130.Fedor Zhdanov. Diverse mini-batch active learning. arXiv:1901.05954, 2019.
  131. 131.Yilun Zhou, Adithya Renduchintala, Xian Li, Sida Wang, Yashar Mehdad, and Asish Ghoshal. Towards Understanding the Behaviors of Optimal Deep Active Learning Algorithms. In International Conference on Artificial Intelligence and Statistics, 2021.
  132. 132.Dominik Ślęzak, Marek Grzegorowski, Andrzej Janusz, Michał Kozielski, Sinh Hoa Nguyen, Marek Sikora, Sebastian Stawicki, and Łukasz Wróbel. A framework for learning and embedding multi-sensor forecasting models into a decision support system: A case study of methane concentration in coal mines. Information Sciences, 451:112–133, 2018.

Citation

MLA
Holzmüller, D., et al. “A Framework and Benchmark for Deep Batch Active Learning for Regression”. Journal of Machine Learning Research, vol. 24, no. 164, 2023, pp. 1–1, https://www.jmlr.org/papers/v24/22-0937.html.
APA
Holzmüller, D., Zaverkin, V., Kästner, J., & Steinwart, I. (2023). A Framework and Benchmark for Deep Batch Active Learning for Regression. Journal of Machine Learning Research, 24(164), 1–81. https://www.jmlr.org/papers/v24/22-0937.html
Chicago
Holzmüller, D., V. Zaverkin, J. Kästner, and I. Steinwart. 2023. “A Framework and Benchmark for Deep Batch Active Learning for Regression”. Journal of Machine Learning Research 24 (164): 1–81. https://www.jmlr.org/papers/v24/22-0937.html.
Harvard
Holzmüller, D. et al. (2023) “A Framework and Benchmark for Deep Batch Active Learning for Regression”, Journal of Machine Learning Research, 24(164), pp. 1–81. Available at: https://www.jmlr.org/papers/v24/22-0937.html.
Vancouver
1. Holzmüller D, Zaverkin V, Kästner J, Steinwart I (2023) A Framework and Benchmark for Deep Batch Active Learning for Regression. Journal of Machine Learning Research 24:1–81

BibTeX

@article{JMLR:v24:22-0937,
  author  = {David Holzmüller and Viktor Zaverkin and Johannes Kästner and Ingo Steinwart},
  title   = {A Framework and Benchmark for Deep Batch Active Learning for Regression},
  journal = {Journal of Machine Learning Research},
  year    = {2023},
  volume  = {24},
  number  = {164},
  pages   = {1--81},
  url     = {http://jmlr.org/papers/v24/22-0937.html}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/