A Framework and Benchmark for Deep Batch Active Learning for Regression
David HolzmüllerViktor ZaverkinJohannes KästnerIngo Steinwart
Presents a modular kernel-based framework, an efficient neural tangent kernel sketching technique paired with a novel clustering selection method, and an open-source 15-dataset benchmark to advance deep batch active learning for regression.
Supervised deep learning models have achieved remarkable success across diverse predictive applications, but their performance typically depends on access to large volumes of labeled data. In many practical engineering, physical, and scientific regression problems, obtaining accurate labels requires costly physical experiments, complex numerical simulations, or intensive manual measurements. While active learning addresses this by selectively querying informative data points, evaluating neural networks sequentially after each single label is computationally prohibitive and blocks parallel workflows. Batch-mode deep active learning overcomes this bottleneck by selecting batches of unlabeled data simultaneously, yet systematic frameworks and comprehensive benchmarks for continuous regression tasks remain underdeveloped.
To address this gap, the article establishes a modular, unified framework for constructing batch active learning algorithms and introduces a standardized tabular regression benchmark to evaluate their effectiveness. The core objective is to systematically analyze and improve sample efficiency in deep regression models without requiring changes to underlying neural network architectures or training routines.
Methodologically, the framework decomposes active learning workflows into three modular components: a base kernel capturing model representations, kernel transformations for statistical modeling and efficiency, and an iterative selection method that queries candidate points. The authors evaluate this approach using an open-source benchmark of 15 large tabular regression datasets spanning up to 379 features and hundreds of thousands of candidate points. Experiments evaluate fully connected neural networks across 20 randomized splits and 16 sequential batch acquisitions, tracking metrics including root mean squared error, mean absolute error, tail quantiles, and maximum error.
The findings establish that replacing standard last-layer feature representations with a sketched, finite-width neural tangent kernel consistently improves predictive accuracy across selection strategies. Furthermore, random feature sketching drastically reduces computational overhead with negligible loss in accuracy. Crucially, the authors' newly developed clustering selection strategy—Largest Cluster Maximum Distance—consistently outperforms existing state-of-the-art methods in both root mean squared error and mean absolute error. The top-performing active learning configurations achieve the predictive accuracy of random sampling using roughly half as much labeled data, while executing batch selections in just seconds on modern hardware.
These performance gains demonstrate substantial real-world value by significantly cutting experimental labeling costs, reducing turnaround times, and boosting model reliability. Tail error improvements also suggest stronger robustness against extreme modeling failures. Additionally, the article identifies that datasets displaying higher initial prediction error variance derive the greatest relative benefit from active batch selection, offering practitioners a practical indicator for when to implement these methods.
Organizations training deep regression models on expensive data should adopt batch active learning using sketched neural tangent representations combined with balanced clustering-based selection methods. For standard accuracy targets, Largest Cluster Maximum Distance offers the best performance profile, whereas geometric distance maximization remains a strong alternative if worst-case errors or covariate shifts are primary concerns.
Confidence in these findings is bolstered by rigorous statistical testing across multiple random seeds, datasets, and activation functions. However, practitioners should note that the evaluation is confined to fully connected networks on tabular data without explicit distribution shifts between candidate and operational test sets. Further validation is warranted before generalizing these conclusions to non-tabular modalities, custom domain architectures, or streaming data environments.
- Paper: Active Learning with Statistical Models, David Cohn et al. (1996). Its variance-based query criteria for locally weighted regression provide a foundation for understanding how active selection can target informative observations in continuous prediction tasks.
- Paper: Active Learning for Convolutional Neural Networks: A Core-Set Approach, Ozan Sener et al. (2018). Its core-set formulation makes the batch-selection challenge concrete: selecting a diverse subset helps avoid redundant queries, a central concern in the source’s clustering methods.
- Paper: Neural Network Ensembles, Cross Validation, and Active Learning, Anders Krogh et al. (1994). Its use of neural-network disagreement to guide active selection of continuous-function examples introduces a model-based selection idea that helps contextualize the source’s regression strategies.
- Paper: A Survey of Deep Active Learning, Pengzhen Ren et al. (2020). Its survey organizes deep active learning’s batch, uncertainty, and diversity approaches, giving readers the methodological map needed to place the source’s framework and benchmark.
No sufficiently relevant recommendations were found.
