DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery - a Focus on Affinity Prediction Problems with Noise Annotations
Yuanfeng JiLu ZhangJiaxiang WuBingzhe WuLanqing LiLong-Kai HuangTingyang XuYu RongJie RenDing Xue
Presents DrugOOD, an automated dataset generation pipeline and evaluation suite that enables researchers to build customized drug-target binding affinity benchmarks with realistic label noise and domain shifts to test the out-of-distribution generalization of graph neural networks.
Artificial intelligence holds significant promise for accelerating drug discovery and cutting research costs, especially in predicting the binding affinity between drug compounds and target proteins. However, real-world deployment frequently encounters two major bottlenecks: distribution shifts, where models must evaluate unfamiliar molecular structures or entirely new disease targets, and label noise within public experimental data repositories. Standard AI benchmarks typically evaluate performance using random data splits, which artificially inflate accuracy estimates and obscure severe performance drops encountered in practice.
The article introduces DrugOOD, an automated dataset curation engine and benchmark suite designed to evaluate model robustness against realistic domain shifts and experimental noise in AI-aided drug discovery. The primary objective is to systematically measure and address the performance gap that occurs when machine learning models are deployed on unseen biological and chemical domains.
The authors constructed an open-source curation pipeline using bioactivity data from the ChEMBL database, generating 45 distinct benchmark datasets. These datasets span three affinity measurement types and three noise severity tiers based on assay reliability and experimental thresholds. To simulate realistic distribution shifts, the data were partitioned into separate training, validation, and testing sets across five biochemically meaningful domains: molecular scaffold, molecular size, assay environment, protein target, and protein family. Using standard graph and sequence backbones, the study rigorously evaluated standard training alongside five state-of-the-art domain generalization algorithms across multiple random trials.
The evaluation revealed several critical findings. First, models suffered severe performance drops when tested on out-of-distribution data. For example, standard training accuracy on the binding classification task dropped by 15.0 to 26.4 percentage points when evaluated on novel target proteins, scaffolds, assays, or molecular sizes, with size variations causing the steepest declines. Second, specialized domain generalization algorithms designed to handle distribution shifts failed to outperform standard empirical risk minimization baselines, with some methods struggling to fit training data effectively. Third, introducing broader datasets with higher noise levels provided additional information that improved generalization up to a point, but further volume expansion hit a performance plateau where data corruption counteracted the benefits of scale.
These findings indicate that current commercial and academic drug discovery models likely operate with substantial unmeasured risk when applied to novel chemical space or emerging biological targets. Relying on conventional out-of-distribution algorithms developed for vision or text is insufficient for molecular graph applications. Furthermore, simply discarding imperfect data or relying exclusively on pristine subsets restricts model scale, while uncurated data scaling leads to diminishing returns.
To address these challenges, development teams and research organizations should adopt systematic out-of-distribution evaluation pipelines like DrugOOD to establish realistic performance baselines prior to deploying AI models in wet-lab pipelines. Future research should prioritize building domain-aware generalization methods tailored to molecular graphs and developing specialized denoising techniques that account for the unique generation mechanisms of experimental bioassays.
The article's conclusions are supported by structured empirical evaluations across diverse splits and baselines. However, current results rely on 2D graph representations of molecules and sequence data for proteins, leaving the integration of detailed 3D structural targets and continuous binding regression models as important areas for future investigation.
- Paper: In Search of Lost Domain Generalization, Ishaan Gulrajani et al. (2020). DomainBed establishes the rigorous domain-generalization comparison that DrugOOD adapts to molecular distribution shifts, including its caution that empirical risk minimization can match specialized methods.
- Paper: WILDS: A Benchmark of in-the-Wild Distribution Shifts, Pang Wei Koh et al. (2020). WILDS provides the broader benchmark framework for evaluating realistic, naturally occurring distribution shifts that motivates DrugOOD’s domain-based splits.
- Paper: MoleculeNet: a benchmark for molecular machine learning, Zhenqin Wu et al. (2017). MoleculeNet supplies the standardized molecular datasets, task definitions, metrics, and splitting conventions that DrugOOD extends with chemically and biologically meaningful out-of-distribution partitions.
- Paper: Open Graph Benchmark: Datasets for Machine Learning on Graphs, Weihua Hu et al. (2020). The Open Graph Benchmark demonstrates why random graph splits can overstate generalization and why application-driven scaffold, species, and time splits are needed.
- Paper: DeepDTA: deep drug–target binding affinity prediction, Hakime Öztürk et al. (2018). DeepDTA establishes the sequence-based drug–target affinity prediction setting and benchmarks that DrugOOD broadens to systematic target, assay, scaffold, and noise shifts.
- Paper: Analyzing Learned Molecular Representations for Property Prediction, Kevin Yang et al. (2019). This molecular-representation study shows that scaffold and chronological splits materially change apparent generalization, preparing the reader for DrugOOD’s stricter evaluation design.
- Paper: Strategies for Pre-training Graph Neural Networks, Weihua Hu et al. (2020). Strategies for Pre-training Graph Neural Networks frames molecular scaffold shifts as a central out-of-distribution challenge and supplies the graph-learning context for DrugOOD’s robustness tests.
- Paper: Learning Causally Invariant Representations for Out-of-Distribution Generalization on Graphs, Yongqiang Chen et al. (2022). CIGA introduces causal invariant graph representations and directly uses DrugOOD as an evaluation setting, making it an important conceptual precursor to the benchmark’s domain-shift problem.
- Paper: E3Bind: An End-to-End Equivariant Network for Protein-Ligand Docking, Yangtian Zhang et al. (2023). E3Bind advances beyond DrugOOD’s 2D molecular and sequence representations by applying end-to-end equivariant modeling to flexible 3D protein–ligand docking.
- Paper: Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs, Rui Jiao et al. (2023). 3D-EMGP follows DrugOOD’s call for richer structural modeling by pretraining equivariant molecular representations on 3D conformations for downstream property and force prediction.
- Paper: ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design, Pascal Notin et al. (2023). ProteinGym extends the benchmark philosophy to large-scale protein fitness and design, testing models across diverse assays, protein families, mutation types, and experimental conditions.
- Paper: Change is Hard: A Closer Look at Subpopulation Shift, Yuzhe Yang et al. (2023). Change is Hard generalizes DrugOOD’s warning about distribution shift by dissecting which subpopulation-shift interventions actually improve robustness across unseen deployment attributes.
