DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery - a Focus on Affinity Prediction Problems with Noise Annotations

Yuanfeng JiLu ZhangJiaxiang WuBingzhe WuLanqing LiLong-Kai HuangTingyang XuYu RongJie RenDing Xue

article2023AAAI77 citations

Presents DrugOOD, an automated dataset generation pipeline and evaluation suite that enables researchers to build customized drug-target binding affinity benchmarks with realistic label noise and domain shifts to test the out-of-distribution generalization of graph neural networks.

Listen

Artificial intelligence holds significant promise for accelerating drug discovery and cutting research costs, especially in predicting the binding affinity between drug compounds and target proteins. However, real-world deployment frequently encounters two major bottlenecks: distribution shifts, where models must evaluate unfamiliar molecular structures or entirely new disease targets, and label noise within public experimental data repositories. Standard AI benchmarks typically evaluate performance using random data splits, which artificially inflate accuracy estimates and obscure severe performance drops encountered in practice.

The article introduces DrugOOD, an automated dataset curation engine and benchmark suite designed to evaluate model robustness against realistic domain shifts and experimental noise in AI-aided drug discovery. The primary objective is to systematically measure and address the performance gap that occurs when machine learning models are deployed on unseen biological and chemical domains.

The authors constructed an open-source curation pipeline using bioactivity data from the ChEMBL database, generating 45 distinct benchmark datasets. These datasets span three affinity measurement types and three noise severity tiers based on assay reliability and experimental thresholds. To simulate realistic distribution shifts, the data were partitioned into separate training, validation, and testing sets across five biochemically meaningful domains: molecular scaffold, molecular size, assay environment, protein target, and protein family. Using standard graph and sequence backbones, the study rigorously evaluated standard training alongside five state-of-the-art domain generalization algorithms across multiple random trials.

The evaluation revealed several critical findings. First, models suffered severe performance drops when tested on out-of-distribution data. For example, standard training accuracy on the binding classification task dropped by 15.0 to 26.4 percentage points when evaluated on novel target proteins, scaffolds, assays, or molecular sizes, with size variations causing the steepest declines. Second, specialized domain generalization algorithms designed to handle distribution shifts failed to outperform standard empirical risk minimization baselines, with some methods struggling to fit training data effectively. Third, introducing broader datasets with higher noise levels provided additional information that improved generalization up to a point, but further volume expansion hit a performance plateau where data corruption counteracted the benefits of scale.

These findings indicate that current commercial and academic drug discovery models likely operate with substantial unmeasured risk when applied to novel chemical space or emerging biological targets. Relying on conventional out-of-distribution algorithms developed for vision or text is insufficient for molecular graph applications. Furthermore, simply discarding imperfect data or relying exclusively on pristine subsets restricts model scale, while uncurated data scaling leads to diminishing returns.

To address these challenges, development teams and research organizations should adopt systematic out-of-distribution evaluation pipelines like DrugOOD to establish realistic performance baselines prior to deploying AI models in wet-lab pipelines. Future research should prioritize building domain-aware generalization methods tailored to molecular graphs and developing specialized denoising techniques that account for the unique generation mechanisms of experimental bioassays.

The article's conclusions are supported by structured empirical evaluations across diverse splits and baselines. However, current results rely on 2D graph representations of molecules and sequence data for proteins, leaving the integration of detailed 3D structural targets and continuous binding regression models as important areas for future investigation.

Cover for DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery - a Focus on Affinity Prediction Problems with Noise Annotations

Abstract

AI-aided drug discovery (AIDD) is gaining popularity due to its potential to make the search for new pharmaceuticals faster, less expensive, and more effective. Despite its extensive use in numerous fields (e.g., ADMET prediction, virtual screening), little research has been conducted on the out-of-distribution (OOD) learning problem with noise. We present DrugOOD, a systematic OOD dataset curator and benchmark for AIDD. Particularly, we focus on the drug-target binding affinity prediction problem, which involves both macromolecule (protein target) and small-molecule (drug compound). DrugOOD offers an automated dataset curator with user-friendly customization scripts, rich domain annotations aligned with biochemistry knowledge, realistic noise level annotations, and rigorous benchmarking of SOTA OOD algorithms, as opposed to only providing fixed datasets. Since the molecular data is often modeled as irregular graphs using graph neural network (GNN) backbones, DrugOOD also serves as a valuable testbed for graph OOD learning problems. Extensive empirical studies have revealed a significant performance gap between in-distribution and out-of-distribution experiments, emphasizing the need for the development of more effective schemes that permit OOD generalization under noise for AIDD.

Table of Contents

  • Introduction
  • Related Work
  • DrugOOD
  • Automated Data Curator
  • Benchmarking State-of-the-art OOD Algorithms
  • Empirical Studies
  • Implementations
  • Experimental Results
  • Discussions and Future Work
  • References

Knowls

  1. Knowl 1 — Structure-Based Affinity Prediction Curation Pipeline in DrugOOD

    model/method

    DrugOOD provides an automated curation pipeline for generating out-of-distribution (OOD) benchmark datasets from the ChEMBL bioactivity repository for structure-based affinity prediction (SBAP). The curation process converts raw bioassay measurements into standardized OOD domain generalization datasets through three sequential steps:

    1. Noise Level Filtering: Bioactivity records are filtered by measurement types (e.g., IC50\text{IC}_{50}, EC50\text{EC}_{50}, KiK_i), target types, assay confidence scores, value relation boundaries (such as exact vs. inequality measurements), and chemical validity via SMILES to partition data into three noise categories: core, refined, and general.
    2. Uncertainty and Multiple Measurement Processing: Activity values recorded at testing boundary limits (with relations <<, ≤\le, ≈\approx, >>, ≥\ge) are offset by 10-fold to address assay concentration bounds. Multiple reported measurement values for the identical protein--ligand pair across different assays are averaged to eliminate redundant recording bias. Continuous values are converted into binary classification labels (active vs. inactive) using adaptive thresholding tailored to individual targets.
    3. Domain Partitioning and Splitting: Samples are categorized into domain clusters using biochemistry-driven descriptors (molecular scaffold, molecular size, protein target, protein family, or bioassay identifier). Domains are sorted by their descriptors and sequentially partitioned into training, OOD validation, and OOD testing sets using a fixed 60%:20%:20%60\% : 20\% : 20\% sample proportion.
  2. Knowl 2 — Biochemical Domain Partitioning Strategies in DrugOOD

    definition

    DrugOOD establishes five biochemically grounded domain definitions to evaluate model robustness against distribution shifts in drug-target interaction prediction:

    • Molecular Scaffold Domain: Molecules sharing the same core Murcko scaffold are grouped into the same domain. The model is trained on a set of scaffolds and evaluated on disjoint, unseen scaffolds.
    • Molecular Size Domain: Molecules are grouped by heavy atom count. The model is trained on specific molecular size ranges and evaluated on unseen molecular sizes.
    • Protein Target Domain: Paired compound--protein samples with the identical target protein are clustered into one domain, testing generalization to completely unseen target proteins during inference.
    • Protein Family Domain: Data samples with targets belonging to the same evolutionary/structural protein family form a single domain, evaluating broad cross-family transfer.
    • Bioassay Domain: Samples generated from the same binding assay environment/protocol are grouped into one domain, evaluating generalization across disparate experimental protocols and laboratory conditions.
  3. Knowl 3 — Multi-Tier Noise Aggregation Framework

    definition

    DrugOOD aggregates heterogeneous bioassay data from ChEMBL into three distinct noise tiers reflecting varying levels of experimental fidelity and data volume:

    • Core Level: The cleanest and most stringent subset. It retains only high-confidence experimental assays with exact measurement values, well-defined protein target types, and standard units, yielding the highest label fidelity at the cost of lower sample count.
    • Refined Level: An intermediate tier that incorporates additional assay records with slightly relaxed confidence score criteria and broader measurement classifications, expanding sample coverage while introducing moderate experimental variation.
    • General Level: The most inclusive and noisy tier. It encompasses all bioassay entries meeting minimal formatting requirements, including records with approximate relations (<,≤,≈,>,≥<, \le, \approx, >, \ge) and varying measurement confidence, reflecting realistic, noisy real-world repositories where the majority of publicly available bioactivity data resides.

    For instance, for IC50\text{IC}_{50} bioassay domains, moving from Core to Refined to General scales the dataset from 123,028 samples (1,503 domains) to 348,248 samples (6,635 domains) to 552,347 samples (22,376 domains).

  4. Knowl 4 — Two-Tower Architecture for Structure-Based Affinity Prediction

    experimental setup

    The benchmark uses a two-tower deep neural network for structure-based affinity prediction (SBAP) formulated as binary classification:

    • Compound Graph Tower: Small molecules are represented as 2D graphs generated via RDKit. Atom nodes contain 39-dimensional feature vectors (atomic symbol, hybridization, hydrogen count, etc.), and bond edges contain 10-dimensional feature vectors (bond type, conjugation, ring membership, stereochemistry). A Graph Isomorphism Network (GIN) backbone extracts a 256-dimensional molecular representation.
    • Protein Sequence Tower: Protein targets are represented as amino acid sequences. A pre-trained BERT encoder (bert-base-uncased) extracts a 768-dimensional protein representation.
    • Readout and Classification Head: The 256-dimensional molecular embedding and 768-dimensional protein embedding are concatenated into a 1024-dimensional joint representation and fed into a multi-layer perceptron (MLP) with a readout function to output the probability of binding activity.
    • Optimization: Models are trained from scratch using Adam with batch sizes in {64,128,256,512,1024}\{64, 128, 256, 512, 1024\} (default 256) and learning rates in {3×10−5,1×10−4,5×10−4,1×10−3,1×10−2}\{3\times 10^{-5}, 1\times 10^{-4}, 5\times 10^{-4}, 1\times 10^{-3}, 1\times 10^{-2}\} (default 1×10−41\times 10^{-4}), without L2L_2 regularization, evaluated by the Area Under the Receiver Operating Characteristic curve (AUROC) and accuracy (ACC).
  5. Knowl 5 — In-Distribution versus Out-of-Distribution Performance Gap of ERM on SBAP

    empirical result

    Empirical Risk Minimization (ERM) evaluated with a GIN-BERT two-tower architecture on DrugOOD core-noise IC50\text{IC}_{50} datasets shows severe performance degradation between in-distribution (ID) and out-of-distribution (OOD) test splits across all five domain types:

    Domain Shift Val (ID) ACC (%) Test (ID) AUC (%) Test (OOD) AUC (%) AUC Drop (%)
    Assay 87.90 ±\pm 0.48 89.73 ±\pm 2.52 70.86 ±\pm 0.38 18.87
    Molecular Scaffold 94.56 ±\pm 0.26 84.24 ±\pm 1.83 68.74 ±\pm 1.07 15.50
    Molecular Size 93.10 ±\pm 0.05 92.75 ±\pm 0.24 66.37 ±\pm 0.35 26.38
    Protein 88.92 ±\pm 0.24 90.32 ±\pm 1.49 68.62 ±\pm 0.45 21.70
    Protein Family 88.30 ±\pm 0.38 86.79 ±\pm 2.85 71.84 ±\pm 1.01 14.95

    The drops in AUROC (ranging from 14.95 to 26.38 percentage points) demonstrate that standard random ID splits strongly overestimate binding affinity prediction accuracy. Molecular size shifts cause the largest performance degradation among all tested domain shifts.

  6. Knowl 6 — Benchmarking Domain Generalization Algorithms on DrugOOD SBAP Datasets

    empirical result

    Six domain generalization (DG) algorithms—Empirical Risk Minimization (ERM), Invariant Risk Minimization (IRM), DeepCORAL, Domain-Adversarial Neural Networks (DANN), Mixup, and Group Distributionally Robust Optimization (GroupDRO)—were benchmarked on DrugOOD core-noise IC50\text{IC}_{50} datasets with identical GIN-BERT backbones:

    Algorithm Assay Scaffold Size Protein Protein Family
    Test (ID) Test (OOD) Test (ID) Test (OOD) Test (ID) Test (OOD) Test (ID) Test (OOD) Test (ID) Test (OOD)
    ERM 89.73 70.86 84.24 68.74 92.75 66.37 90.32 68.62 86.79 71.84
    IRM 83.55 68.72 83.87 67.74 74.92 56.62 91.29 67.66 85.63 70.44
    DeepCORAL 84.91 68.68 80.21 67.83 79.34 59.41 90.33 67.26 – –
    DANN 75.98 65.16 75.86 64.18 90.60 66.05 78.12 62.58 86.00 70.28
    Mixup 88.07 70.85 86.05 68.61 91.98 66.21 91.11 68.25 86.14 73.10
    GroupDRO 84.31 68.49 81.11 67.79 82.71 59.92 89.36 67.62 87.72 72.76

    Note: Values represent AUROC (%). Missing entries (--) denote settings infeasible due to limited domain counts.

    Standard ERM and Mixup consistently equal or outperform specialized OOD algorithms (IRM, DeepCORAL, DANN, GroupDRO) on OOD test performance across all domain partitions. Several DG algorithms impair the model's capacity to fit the training data (e.g., DeepCORAL and IRM achieve only 79.34% and 74.92% ID test AUC on the Size split compared to ERM's 92.75%).

  7. Knowl 7 — Impact of Bioassay Noise Levels on Model Generalization

    empirical result

    Empirical evaluation of ERM across the three noise tiers on assay-shifted IC50\text{IC}_{50} datasets reveals the following trade-offs:

    Noise Level Val (ID) AUC (%) Val (OOD) AUC (%) Test (ID) AUC (%) Test (OOD) AUC (%)
    Core 89.62 72.26 90.32 68.62
    Refined 82.87 73.53 82.92 68.00
    General 78.97 71.63 78.94 68.06
    1. In-Distribution Degradation: Label noise progressively degrades ID test performance from 90.32% AUC (Core) to 78.94% AUC (General).
    2. ID--OOD Gap Reduction: Expanding sample size by including moderately noisy data (Refined level) narrows the relative gap between ID and OOD metrics.
    3. Plateau at High Noise: Moving from Refined to General noise level yields no additional OOD test AUC improvement (68.06% vs. 68.00%), demonstrating that massive label noise offsets the statistical benefit of larger training volume.

Coverage note — Ligand-based affinity prediction (LBAP) and specific baseline loss formulas (e.g., standard IRM, DANN, DeepCORAL objectives) are omitted because LBAP details are deferred to external project pages and the DG objectives are standard prior literature.

References

  1. 1.Angluin, D.; and Laird, P. 1988. Learning from noisy examples. Machine Learning, 2(4): 343–370.
  2. 2.Arjovsky, M.; Bottou, L.; Gulrajani, I.; and Lopez-Paz, D. 2019. Invariant risk minimization. arXiv preprint arXiv:1907.02893.
  3. 3.Berman, H. M.; Westbrook, J.; Feng, Z.; Gilliland, G.; Bhat, T. N.; Weissig, H.; Shindyalov, I. N.; and Bourne, P. E. 2000. The protein data bank. Nucleic acids research, 28(1): 235–242.
  4. 4.Bevilacqua, B.; Zhou, Y.; and Ribeiro, B. 2021. Size-invariant graph representations for graph classification extrapolations. In International Conference on Machine Learning, 837–851. PMLR.
  5. 5.Burley, S. K.; Bhikadiya, C.; Bi, C.; Bittrich, S.; Chen, L.; Crichlow, G. V.; Christie, C. H.; Dalenberg, K.; Di Costanzo, L.; Duarte, J. M.; Dutta, S.; Feng, Z.; Ganesan, S.; Goodsell, D. S.; Ghosh, S.; Green, R. K.; Guranovic, V.; Guzenko, D.; Hudson, B. P.; Lawson, ´ C.; Liang, Y.; Lowe, R.; Namkoong, H.; Peisach, E.; Persikova, I.; Randle, C.; Rose, A.; Rose, Y.; Sali, A.; Segura, J.; Sekharan, M.; Shao, C.; Tao, Y.-P.; Voigt, M.; Westbrook, J.; Young, J. Y.; Zardecki, C.; and Zhuravleva, M. 2020. RCSB Protein Data Bank: powerful new tools for exploring 3D structures of biological macromolecules for basic and applied research and education in fundamental biology, biomedicine, biotechnology, bioengineering and energy sciences. Nucleic Acids Research, 49(D1): D437–D451.
  6. 6.Chuang, C.-Y.; and Mroueh, Y. 2021. Fair Mixup: Fairness via Interpolation. In International Conference on Learning Representations.
  7. 7.Consortium, T. U. 2014. UniProt: a hub for protein information. Nucleic Acids Research, 43(D1): D204–D212.
  8. 8.Cortés-Ciriano, I.; and Bender, A. 2016. How consistent are publicly reported cytotoxicity data? Large-scale statistical analysis of the concordance of public independent cytotoxicity measurements. ChemMedChem, 11(1): 57–71.
  9. 9.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  10. 10.Duvenaud, D.; Maclaurin, D.; Aguilera-Iparraguirre, J.; Gómez-Bombarelli, R.; Hirzel, T.; Aspuru-Guzik, A.; and Adams, R. P. 2015. Convolutional Networks on Graphs for Learning Molecular Fingerprints. In Cortes, C.; Lawrence, N. D.; Lee, D. D.; Sugiyama, M.; and Garnett, R., eds., Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, 2224–2232.
  11. 11.Ganin, Y.; Ustinova, E.; Ajakan, H.; Germain, P.; Larochelle, H.; Laviolette, F.; Marchand, M.; and Lempitsky, V. 2016a. Domain-adversarial training of neural networks. The journal of machine learning research, 17(1): 2096–2030.
  12. 12.Ganin, Y.; Ustinova, E.; Ajakan, H.; Germain, P.; Larochelle, H.; Laviolette, F.; Marchand, M.; and Lempitsky, V. 2016b. Domain-adversarial training of neural networks. The journal of machine learning research, 17(1): 2096–2030.
  13. 13.Gilson, M. K.; Liu, T.; Baitaluk, M.; Nicola, G.; Hwang, L.; and Chong, J. 2016. BindingDB in 2015: a public database for medicinal chemistry, computational chemistry and systems pharmacology. Nucleic acids research, 44(D1): D1045–D1053.
  14. 14.Han, B.; Yao, Q.; Liu, T.; Niu, G.; Tsang, I. W.; Kwok, J. T.; and Sugiyama, M. 2020. A survey of label-noise representation learning: Past, present and future. arXiv preprint arXiv:2011.04406.
  15. 15.Hu, P.-W.; Chan, K. C.; and You, Z.-H. 2016. Large-scale prediction of drug-target interactions from deep representations. In 2016 international joint conference on neural networks (IJCNN), 1236–1243. IEEE.
  16. 16.Hu, R.; Xu, H.; Jia, P.; and Zhao, Z. 2020. KinaseMD: kinase mutations and drug response database. Nucleic Acids Research, 49(D1): D552–D561.
  17. 17.Hu, W.; Fey, M.; Zitnik, M.; Dong, Y.; Ren, H.; Liu, B.; Catasta, M.; and Leskovec, J. 2020. Open Graph Benchmark: Datasets for Machine Learning on Graphs. arXiv e-prints, arXiv:2005.00687.
  18. 18.Huang, K.; Fu, T.; Gao, W.; Zhao, Y.; Roohani, Y.; Leskovec, J.; Coley, C. W.; Xiao, C.; Sun, J.; and Zitnik, M. 2021. Therapeutics data Commons: machine learning datasets and tasks for therapeutics. arXiv preprint arXiv:2102.09548.
  19. 19.Kalliokoski, T.; Kramer, C.; Vulpetti, A.; and Gedeck, P. 2013. Comparability of mixed IC50 data–a statistical analysis. PloS one, 8(4): e61007.
  20. 20.Karimi, M.; Wu, D.; Wang, Z.; and Shen, Y. 2019. DeepAffinity: interpretable deep learning of compound–protein affinity through unified recurrent and convolutional neural networks. Bioinformatics, 35(18): 3329–3338.
  21. 21.Kearnes, S.; McCloskey, K.; Berndl, M.; Pande, V.; and Riley, P. 2016. Molecular graph convolutions: moving beyond fingerprints. Journal of computer-aided molecular design, 30(8): 595–608.
  22. 22.Kipf, T. N.; and Welling, M. 2016. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907.
  23. 23.Koh, P. W.; Sagawa, S.; Marklund, H.; Xie, S. M.; Zhang, M.; Balsubramani, A.; Hu, W.; Yasunaga, M.; Phillips, R. L.; Gao, I.; Lee, T.; David, E.; Stavness, I.; Guo, W.; Earnshaw, B.; Haque, I.; Beery, S. M.; Leskovec, J.; Kundaje, A.; Pierson, E.; Levine, S.; Finn, C.; and Liang, P. 2021. WILDS: A Benchmark of in-the-Wild Distribution Shifts. In Meila, M.; and Zhang, T., eds., Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, 5637–5664. PMLR.
  24. 24.Koyama, M.; and Yamaguchi, S. 2021. When is invariance useful in an Out-of-Distribution Generalization problem ? arXiv:2008.01883.
  25. 25.Kramer, C.; Kalliokoski, T.; Gedeck, P.; and Vulpetti, A. 2012. The experimental uncertainty of heterogeneous public K i data. Journal of medicinal chemistry, 55(11): 5165–5173.
  26. 26.Krueger, D.; Caballero, E.; Jacobsen, J.-H.; Zhang, A.; Binas, J.; Zhang, D.; Priol, R. L.; and Courville, A. 2021. Out-of-Distribution Generalization via Risk Extrapolation (REx). In Proceedings of the 38th International Conference on Machine Learning, volume 139, 5815–5826. PMLR.
  27. 27.Li, M.; Zhou, J.; Hu, J.; Fan, W.; Zhang, Y.; Gu, Y.; and Karypis, G. 2021. Dgl-lifesci: An open-source toolkit for deep learning on graphs in life science. ACS omega, 6(41): 27233–27238.
  28. 28.Lim, J.; Ryu, S.; Park, K.; Choe, Y. J.; Ham, J.; and Kim, W. Y. 2019. Predicting drug–target interaction using a novel graph neural network with 3D structure-embedded graph representation. Journal of chemical information and modeling, 59(9): 3981–3988.
  29. 29.Liu, Z.; Li, Y.; Han, L.; Li, J.; Liu, J.; Zhao, Z.; Nie, W.; Liu, Y.; and Wang, R. 2014. PDB-wide collection of binding data: current status of the PDBbind database. Bioinformatics, 31(3): 405–412.
  30. 30.Lu, C.; Liu, Q.; Wang, C.; Huang, Z.; Lin, P.; and He, L. 2019. Molecular property prediction: A multilevel quantum interactions modeling perspective. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, 1052–1060.
  31. 31.Martin, E. J.; Polyakov, V. R.; Tian, L.; and Perez, R. C. 2017. Profile-QSAR 2.0: Kinase Virtual Screening Accuracy Comparable to Four-Concentration IC50s for Realistically Novel Compounds. Journal of Chemical Information and Modeling, 57(8): 2077–2088. PMID: 28651433.
  32. 32.Mendez, D.; Gaulton, A.; Bento, A. P.; Chambers, J.; De Veij, M.; Félix, E.; Magariños, M. P.; Mosquera, J. F.; Mutowo, P.; Nowotka, M.; et al. 2019. ChEMBL: towards direct deposition of bioassay data. Nucleic acids research, 47(D1): D930–D940.
  33. 33.Muratov, E. N.; Bajorath, J.; Sheridan, R. P.; Tetko, I. V.; Filimonov, D.; Poroikov, V.; Oprea, T. I.; Baskin, I. I.; Varnek, A.; Roitberg, A.; et al. 2020. QSAR without borders. Chemical Society Reviews, 49(11): 3525–3564.
  34. 34.Qiao, F.; Zhao, L.; and Peng, X. 2020. Learning to Learn Single Domain Generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  35. 35.Rong, Y.; Bian, Y.; Xu, T.; Xie, W.; Wei, Y.; Huang, W.; and Huang, J. 2020. Self-Supervised Graph Transformer on Large-Scale Molecular Data. In NeurIPS.
  36. 36.Sagawa, S.; Koh, P. W.; Hashimoto, T. B.; and Liang, P. 2019. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. arXiv preprint arXiv:1911.08731.
  37. 37.Schneider, G. 2018. Automating drug discovery. Nature reviews drug discovery, 17(2): 97–113.
  38. 38.Schütt, K. T.; Kindermans, P.-J.; Sauceda, H. E.; Chmiela, S.; Tkatchenko, A.; and Müller, K.-R. 2017. Schnet: A continuous-filter convolutional neural network for modeling quantum interactions. arXiv preprint arXiv:1706.08566.
  39. 39.Shen, T.; Wu, J.; Lan, H.; Zheng, L.; Pei, J.; Wang, S.; Liu, W.; and Huang, J. 2021. When homologous sequences meet structural decoys: Accurate contact prediction by tFold in CASP14—(tFold for CASP14 contact prediction). Proteins: Structure, Function, and Bioinformatics, 89(12): 1901–1910.
  40. 40.Sliwoski, G.; Kothiwale, S.; Meiler, J.; and Lowe Jr., E. W. 2013. Computational methods in drug discovery. Pharmacological reviews, 66(1): 334–395. 24381236[pmid].
  41. 41.Stanley, M.; Bronskill, J. F.; Maziarz, K.; Misztela, H.; Lanini, J.; Segler, M.; Schneider, N.; and Brockschmidt, M. 2021. FS-Mol: A Few-Shot Learning Dataset of Molecules. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  42. 42.Sun, B.; and Saenko, K. 2016a. Deep coral: Correlation alignment for deep domain adaptation. In European conference on computer vision, 443–450. Springer.
  43. 43.Sun, B.; and Saenko, K. 2016b. Deep coral: Correlation alignment for deep domain adaptation. In European conference on computer vision, 443–450. Springer.
  44. 44.Tzeng, E.; Hoffman, J.; Zhang, N.; Saenko, K.; and Darrell, T. 2014. Deep Domain Confusion: Maximizing for Domain Invariance. arXiv:1412.3474.
  45. 45.Wu, Z.; Ramsundar, B.; Feinberg, E. N.; Gomes, J.; Geniesse, C.; Pappu, A. S.; Leswing, K.; and Pande, V. 2018. MoleculeNet: a benchmark for molecular machine learning. Chemical science, 9(2): 513–530.
  46. 46.Xiong, Z.; Wang, D.; Liu, X.; Zhong, F.; Wan, X.; Li, X.; Li, Z.; Luo, X.; Chen, K.; Jiang, H.; et al. 2019. Pushing the boundaries of molecular representation for drug discovery with the graph attention mechanism. Journal of medicinal chemistry, 63(16): 8749–8760.
  47. 47.Xu, K.; Hu, W.; Leskovec, J.; and Jegelka, S. 2018. How Powerful are Graph Neural Networks? In International Conference on Learning Representations.
  48. 48.Xu, M.; Zhang, J.; Ni, B.; Li, T.; Wang, C.; Tian, Q.; and Zhang, W. 2020. Adversarial Domain Adaptation with Domain Mixup. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, 6502–6509. AAAI Press.
  49. 49.Yan, S.; Song, H.; Li, N.; Zou, L.; and Ren, L. 2020. Improve Unsupervised Domain Adaptation with Mixup Training. arXiv:2001.00677.
  50. 50.Yehudai, G.; Fetaya, E.; Meirom, E.; Chechik, G.; and Maron, H. 2021. From local structures to size generalization in graph neural networks. In International Conference on Machine Learning, 11975–11986. PMLR.
  51. 51.Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2017. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412.
  52. 52.Zhao, L.; Liu, T.; Peng, X.; and Metaxas, D. 2020. Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M. F.; and Lin, H., eds., Advances in Neural Information Processing Systems, volume 33, 14435–14447. Curran Associates, Inc.
  53. 53.Zhou, F.; Jiang, Z.; Shui, C.; Wang, B.; and Chaib-draa, B. 2021. Domain Generalization via Optimal Transport with Metric Similarity Learning. Neurocomputing, 456(C): 469–480.
  54. 54.Zhuang, F.; Qi, Z.; Duan, K.; Xi, D.; Zhu, Y.; Zhu, H.; Xiong, H.; and He, Q. 2020. A comprehensive survey on transfer learning. Proceedings of the IEEE, 109(1): 43–76.

Citation

MLA
Ji, Y., et al. “DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery – a Focus on Affinity Prediction Problems with Noise Annotations”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 7, 2023, pp. 8023–31, https://doi.org/10.1609/AAAI.V37I7.25970.
APA
Ji, Y., Zhang, L., Wu, J., Wu, B., Li, L., Huang, L.-K., Xu, T., Rong, Y., Ren, J., Xue, D., Lai, H., Liu, W., Huang, J., Zhou, S., Luo, P., Zhao, P., & Bian, Y. (2023). DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery – a Focus on Affinity Prediction Problems with Noise Annotations. Proceedings of the AAAI Conference on Artificial Intelligence, 37(7), 8023–8031. https://doi.org/10.1609/AAAI.V37I7.25970
Chicago
Ji, Y., L. Zhang, J. Wu, et al. 2023. “DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery – a Focus on Affinity Prediction Problems with Noise Annotations”. Proceedings of the AAAI Conference on Artificial Intelligence 37 (7): 8023–31. https://doi.org/10.1609/AAAI.V37I7.25970.
Harvard
Ji, Y. et al. (2023) “DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery – a Focus on Affinity Prediction Problems with Noise Annotations”, Proceedings of the AAAI Conference on Artificial Intelligence, 37(7), pp. 8023–8031. Available at: https://doi.org/10.1609/AAAI.V37I7.25970.
Vancouver
1. Ji Y, Zhang L, Wu J, et al (2023) DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery – a Focus on Affinity Prediction Problems with Noise Annotations. Proceedings of the AAAI Conference on Artificial Intelligence 37:8023–8031

BibTeX

@article{Ji_2023, title={DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery – a Focus on Affinity Prediction Problems with Noise Annotations}, volume={37}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V37I7.25970}, DOI={10.1609/aaai.v37i7.25970}, number={7}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Ji, Yuanfeng and Zhang, Lu and Wu, Jiaxiang and Wu, Bingzhe and Li, Lanqing and Huang, Long-Kai and Xu, Tingyang and Rong, Yu and Ren, Jie and Xue, Ding and Lai, Houtim and Liu, Wei and Huang, Junzhou and Zhou, Shuigeng and Luo, Ping and Zhao, Peilin and Bian, Yatao}, year={2023}, month=June, pages={8023–8031} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF