TANKBind: Trigonometry-Aware Neural NetworKs for Drug-Protein Binding Structure Prediction

Wei LuQifeng WuJixian ZhangJiahua RaoChengtao LiShuangjia Zheng

article2022NeurIPS253 citations

Proposes a deep learning framework that incorporates physical geometric constraints and functional protein block segmentation to accurately predict drug-protein binding structures and affinities without extensive conformational sampling.

Listen

Understanding how small-molecule drugs interact with target proteins is central to modern drug discovery and assessing off-target risks. Conventional computational molecular docking methods face severe limitations due to expensive conformational sampling and simplified scoring functions, while recent one-shot deep learning alternatives often overlook physical geometric constraints and produce unrealistic molecular structures. The article introduces TANKBind (Trigonometry-Aware Neural NetworKs), an artificial intelligence framework designed to accurately predict both the three-dimensional binding structures and the corresponding binding affinities of drug-protein pairs with high computational speed.

To overcome previous shortcomings, the approach combines geometric constraints with a biological divide-and-conquer strategy. The model segments proteins into functional structural blocks and uses a specialized trigonometry module to enforce triangle inequalities and prevent atomic overlap. Protein and small-molecule structures are embedded as graphs, and the system is trained via contrastive loss functions that treat non-binding protein regions as negative decoys. The framework was evaluated on the standard PDBbind database against state-of-the-art physics-based docking packages and modern deep learning models, assessing performance on both familiar proteins and entirely unseen targets.

The evaluation demonstrated substantial improvements across key performance benchmarks. For blind flexible self-docking, TANKBind achieved a 22% improvement over existing deep learning baselines in generating high-quality ligand poses with error below 5 Ångströms. When evaluated on novel, previously unobserved proteins, this improvement grew to 42%, while simultaneously outperforming traditional docking suites by large margins. Additionally, the model surpassed existing sequence-, structure-, and complex-based methods in predicting binding affinities, achieving a Pearson correlation of 0.726 and an absolute error of 1.070.

These results demonstrate that incorporating physical geometric constraints directly into neural networks substantially enhances their ability to generalize to novel biological targets. For drug development pipelines, this approach offers an avenue to accelerate virtual screening timelines and reduce computational overhead while minimizing false-positive pose predictions. By successfully identifying unobserved binding sites on known proteins and resolving novel complexes, the framework provides a viable computational tool for discovering new mechanisms of action.

To build upon these results, development teams should explore integrating ligand conformation generation modules, training on broader datasets expanded with predicted protein structures and structure-activity relationship data, and incorporating protein backbone dynamics. Decision-makers should note that the model currently relies on residue-level protein approximations and pre-segmented functional sites, requiring standard experimental validation for critical discovery candidates.

Lu et al (2022).pdf
  • Paper: Neural Message Passing for Quantum Chemistry, Justin Gilmer et al. (2017). This foundational paper establishes message passing neural networks on molecular graphs, providing the baseline graph-learning formulation adapted by TANKBind for small-molecule and protein representations.
  • Paper: DeepDTA: deep drug–target binding affinity prediction, Hakime Öztürk et al. (2018). It introduces continuous drug-target binding affinity prediction with deep learning, establishing the affinity prediction problem formulation that TANKBind advances using 3D geometric constraints.
  • Paper: Do Transformers Really Perform Badly for Graph Representation?, Chengxuan Ying et al. (2021). It formalizes incorporating structural spatial distances and graph inductive biases into attention mechanisms, a core conceptual prerequisite for TANKBind's trigonometry-aware modules.
  • Paper: Strategies for Pre-training Graph Neural Networks, Weihua Hu et al. (2020). This study develops pretraining and out-of-distribution evaluation paradigms for molecular graph networks, informing TANKBind's contrastive block training and unseen-target validation methodology.
  • Paper: MoleculeNet: a benchmark for molecular machine learning, Zhenqin Wu et al. (2017). It establishes standardized biophysical and physiological benchmarks for molecular machine learning, framing the evaluation protocols and data splits used in drug-target interaction research.
Cover for TANKBind: Trigonometry-Aware Neural NetworKs for Drug-Protein Binding Structure Prediction

Abstract

Illuminating interactions between proteins and small drug molecules is a long-standing challenge in the field of drug discovery. Despite the importance of understanding these interactions, most previous works are limited by hand-designed scoring functions and insufficient conformation sampling. The recently-proposed graph neural network-based methods provides alternatives to predict protein-ligand complex conformation in a one-shot manner. However, these methods neglect the geometric constraints of the complex structure and weaken the role of local functional regions. As a result, they might produce unreasonable conformations for challenging targets and generalize poorly to novel proteins. In this paper, we propose Trigonometry-Aware Neural networKs for binding structure prediction, TANKBind, that builds trigonometry constraint as a vigorous inductive bias into the model and explicitly attends to all possible binding sites for each protein by segmenting the whole protein into functional blocks. We construct novel contrastive losses with local region negative sampling to jointly optimize the binding interaction and affinity. Extensive experiments show substantial performance gains in comparison to state-of-the-art physics-based and deep learning-based methods on commonly-used benchmark datasets for both binding structure and affinity predictions with variant settings.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 TankBind Model
  • 3.1 Overview of TankBind model
  • 3.2 Structural encoders of protein and drug
  • 3.3 Details of trigonometry module
  • 3.4 Design of binding interaction and affinity loss functions
  • 3.5 Generation of drug coordinates based on predicted inter-molecular distance map.
  • 4 Evaluation
  • 4.1 Protein-ligand binding structure prediction
  • 4.2 Protein-ligand binding affinity prediction
  • 4.3 Ablation study
  • 4.4 Case studies
  • 5 Conclusion
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — Trigonometry-Aware Interaction Embedding Update in TANKBind

    equation

    TANKBind incorporates geometric triangle inequality constraints into the protein-ligand interaction embedding by updating the joint residue-atom pair representation using intra-protein and intra-ligand spatial distance matrices. For layer ℓ\ell, given the protein block node count nn, ligand heavy-atom count mm, and embedding dimension ss, the intermediate interaction representation z~ij(ℓ)∈Rs\tilde{z}_{ij}^{(\ell)} \in \mathbb{R}^s between protein residue ii and ligand atom jj is updated as:

    z~ij(ℓ)=zij(ℓ)+Φ(∑k=1npiktkj(ℓ)+∑k′=1mtik′′(ℓ)ck′j)⊙g(zij(ℓ))\tilde{z}_{ij}^{(\ell)} = z_{ij}^{(\ell)} + \Phi\left(\sum_{k=1}^n p_{ik} t_{kj}^{(\ell)} + \sum_{k'=1}^m t_{ik'}'^{(\ell)} c_{k'j}\right) \odot g(z_{ij}^{(\ell)})

    where:

    • pik=ϕ(Dikp)∈Rsp_{ik} = \phi(D_{ik}^p) \in \mathbb{R}^s is a linear embedding of the encoded intra-protein residue-pair distance Dikp=∥xip−xkp∥∈RD_{ik}^p = \|x_i^p - x_k^p\| \in \mathbb{R} between CαC_\alpha coordinates xip,xkp∈R3x_i^p, x_k^p \in \mathbb{R}^3.
    • ck′j=ϕ(Djk′c)∈Rsc_{k'j} = \phi(D_{jk'}^c) \in \mathbb{R}^s is a linear embedding of the encoded intra-ligand atom-pair distance Djk′c=∥xjc−xk′c∥∈RD_{jk'}^c = \|x_j^c - x_{k'}^c\| \in \mathbb{R} between ligand heavy-atom coordinates xjc,xk′c∈R3x_j^c, x_{k'}^c \in \mathbb{R}^3.
    • tij(ℓ)=Linear(zij(ℓ))⊙g(zij(ℓ))∈Rst_{ij}^{(\ell)} = \text{Linear}(z_{ij}^{(\ell)}) \odot g(z_{ij}^{(\ell)}) \in \mathbb{R}^s and tij′(ℓ)t_{ij}'^{(\ell)} are gated linear transformations with separate parameters.
    • g(zij(ℓ))=σ(Linear(zij(ℓ)))∈Rsg(z_{ij}^{(\ell)}) = \sigma(\text{Linear}(z_{ij}^{(\ell)})) \in \mathbb{R}^s, where σ\sigma is the element-wise sigmoid function.
    • Φ\Phi denotes layer normalization followed by a linear transformation.
    • ⊙\odot represents the Hadamard (element-wise) product.
  2. Knowl 2 — Max-Margin Contrastive Affinity and Distance Map Loss Formulation

    equation

    TANKBind is trained jointly to predict block-level binding affinities and native inter-molecular distance maps across segmented protein functional blocks. Let ζ\zeta index a protein functional block, and let 1(ζ)=1\mathbf{1}(\zeta) = 1 if block ζ\zeta encloses the native crystallized ligand and 1(ζ)=0\mathbf{1}(\zeta) = 0 otherwise (acting as a decoy). For an experimental affinity aa and predicted affinity a^ζ\hat{a}_\zeta, the max-margin contrastive affinity loss is defined as:

    Laffinity(a^ζ,a)=1(ζ)(a^ζ−a)2+(1−1(ζ))max⁡(0,a^ζ−(a−ϵ))2\mathcal{L}_{\text{affinity}}(\hat{a}_\zeta, a) = \mathbf{1}(\zeta)(\hat{a}_\zeta - a)^2 + (1 - \mathbf{1}(\zeta)) \max(0, \hat{a}_\zeta - (a - \epsilon))^2

    where ϵ>0\epsilon > 0 denotes a predefined margin enforcing non-binding decoy regions to exhibit weaker affinity than the native binding pocket.

    The inter-molecular distance map loss is evaluated exclusively on the native binding block ζ\zeta with nn residues and mm ligand heavy atoms:

    Ldistance=1(ζ)1nm∑i=1n∑j=1m(Dijpred−Dij)2\mathcal{L}_{\text{distance}} = \mathbf{1}(\zeta) \frac{1}{nm} \sum_{i=1}^n \sum_{j=1}^m (D_{ij}^{\text{pred}} - D_{ij})^2

    where DijD_{ij} is the true inter-molecular distance between residue CαC_\alpha coordinate ii and ligand heavy atom jj, and Dijpred=g(zij(L))Linear(zij(L))D_{ij}^{\text{pred}} = g(z_{ij}^{(L)})\text{Linear}(z_{ij}^{(L)}) is the predicted distance from the final layer LL representation zij(L)z_{ij}^{(L)}.

    The composite training objective is:

    L=Laffinity+Ldistance\mathcal{L} = \mathcal{L}_{\text{affinity}} + \mathcal{L}_{\text{distance}}

  3. Knowl 3 — TANKBind Architecture and Multi-Scale Graph Representations

    model/method

    TANKBind is a two-stage deep learning framework for drug-protein binding conformation and affinity prediction that operates on segmented protein graphs and molecular ligand graphs:

    1. Protein Representation and Pocket Segmentation: The protein structure is modeled as a 3D proximity graph Gp=(Vp,Ep)G^p = (V^p, E^p), where each node vip∈Vpv_i^p \in V^p represents an amino acid residue located at its CαC_\alpha coordinate xip∈R3x_i^p \in \mathbb{R}^3. Edges eikp∈Epe_{ik}^p \in E^p connect each node to its 30 nearest spatial neighbors with scalar and vector directional features. Node embeddings hp∈Rn×sh^p \in \mathbb{R}^{n \times s} are extracted via Geometric Vector Perceptrons (GVP). To scale to entire proteins and leverage conserved functional zones, the protein is partitioned into overlapping subgraphs Gp′=({vip,eikp}∣∥xip−x0∥≤20A˚,∥xkp−x0∥≤20A˚)G^{p\prime} = (\{v_i^p, e_{ik}^p\} \mid \|x_i^p - x_0\| \le 20\text{\AA}, \|x_k^p - x_0\| \le 20\text{\AA}), where x0x_0 is a pocket center predicted by P2rank.

    2. Ligand Representation: The small-molecule drug is represented as an all-heavy-atom 2D graph Gc=(Vc,Ec)G^c = (V^c, E^c) with atom nodes vjcv_j^c and bond edges ejk′ce_{jk'}^c. A Graph Isomorphism Network (GIN) processes GcG^c into node embeddings hc∈Rm×sh^c \in \mathbb{R}^{m \times s}.

    3. Joint Embedding Initialization: The initial interaction tensor z(0)∈Rn×m×sz^{(0)} \in \mathbb{R}^{n \times m \times s} between a protein block of size nn and a ligand of size mm is formed by the element-wise outer product zij(0)=hip⊙hjcz_{ij}^{(0)} = h_i^p \odot h_j^c.

    4. Affinity and Distance Prediction: After LL stacked blocks of trigonometry updates, self-attention modulation, and transition MLPs, the model outputs the overall binding affinity a^=∑i=1n∑j=1mLinear(zij(L))\hat{a} = \sum_{i=1}^n \sum_{j=1}^m \text{Linear}(z_{ij}^{(L)}) and predicted pairwise inter-atomic distances Dijpred=g(zij(L))Linear(zij(L))D_{ij}^{\text{pred}} = g(z_{ij}^{(L)})\text{Linear}(z_{ij}^{(L)}), selecting the block with the highest predicted affinity to generate binding poses.

  4. Knowl 4 — Excluded-Volume Self-Attention Modulation in TANKBind

    model/method

    To model physical steric exclusion (Van der Waals interactions) and chemical valency/saturation effects (such as the finite number of hydrogen bond donors/acceptors per residue), TANKBind modulates the trigonometry-updated interaction matrix z~ij(ℓ)\tilde{z}_{ij}^{(\ell)} across all ligand atoms k′∈{1,…,m}k' \in \{1, \dots, m\} interacting with residue ii.

    For attention head h∈{1,…,H}h \in \{1, \dots, H\}:

    wijk′(ℓ)h=softmaxk′(qij(ℓ)h⊤kik′(ℓ)h)w_{ijk'}^{(\ell)h} = \text{softmax}_{k'}\left({q_{ij}^{(\ell)h}}^\top k_{ik'}^{(\ell)h}\right)

    z˙ij(ℓ)=z~ij(ℓ)+Φ(concath(∑k′=1mwijk′(ℓ)hvik′(ℓ)h)⊙gh(z~ij(ℓ)))\dot{z}_{ij}^{(\ell)} = \tilde{z}_{ij}^{(\ell)} + \Phi\left(\text{concat}_h\left(\sum_{k'=1}^m w_{ijk'}^{(\ell)h} v_{ik'}^{(\ell)h}\right) \odot g^h(\tilde{z}_{ij}^{(\ell)})\right)

    where qij(ℓ)h,kik′(ℓ)h,vik′(ℓ)hq_{ij}^{(\ell)h}, k_{ik'}^{(\ell)h}, v_{ik'}^{(\ell)h} are linear projections of z~ij(ℓ)\tilde{z}_{ij}^{(\ell)} and z~ik′(ℓ)\tilde{z}_{ik'}^{(\ell)}, gh(⋅)g^h(\cdot) is a gated sigmoid transformation reshaped over attention heads, and Φ\Phi applies layer normalization followed by a linear projection.

    Following self-attention modulation, the representation transitions to the next layer via a multilayer perceptron:

    zij(ℓ+1)=MLP(z˙ij(ℓ))z_{ij}^{(\ell+1)} = \text{MLP}(\dot{z}_{ij}^{(\ell)})

  5. Knowl 5 — Numerical Ligand Coordinate Generation from Distance Maps

    algorithm

    To convert noisy predicted inter-molecular distance matrices Dpred∈Rn×mD^{\text{pred}} \in \mathbb{R}^{n \times m} into 3D Cartesian coordinates {x^jc}j=1m⊂R3\{\hat{x}_j^c\}_{j=1}^m \subset \mathbb{R}^3 for ligand heavy atoms relative to protein CαC_\alpha coordinates {xip}i=1n\{x_i^p\}_{i=1}^n, TANKBind minimizes a composite loss function Lgeneration\mathcal{L}_{\text{generation}}:

    Input: Protein coordinates {xip}i=1n\{x_i^p\}_{i=1}^n, predicted inter-molecular distances Dpred∈Rn×mD^{\text{pred}} \in \mathbb{R}^{n \times m}, intra-ligand reference distances Dc∈Rm×mD^c \in \mathbb{R}^{m \times m}, local atomic structures (LAS) mask 1(j,k′)\mathbf{1}(j, k')
    Output: Docked ligand coordinates {x^jc}j=1m⊂R3\{\hat{x}_j^c\}_{j=1}^m \subset \mathbb{R}^3
    Initialize {x^jc}j=1m\{\hat{x}_j^c\}_{j=1}^m from initial ligand conformation
    Define loss Lgeneration({x^jc})=Linteraction+Lconfiguration\mathcal{L}_{\text{generation}}(\{\hat{x}_j^c\}) = \mathcal{L}_{\text{interaction}} + \mathcal{L}_{\text{configuration}} where:
      D^ij=∥xip−x^jc∥\hat{D}_{ij} = \|x_i^p - \hat{x}_j^c\|
      D^jk′c=∥x^jc−x^k′c∥\hat{D}_{jk'}^c = \|\hat{x}_j^c - \hat{x}_{k'}^c\|
      Linteraction=∑i=1n∑j=1m∣min⁡(D^ij,10A˚)−min⁡(Dijpred,10A˚)∣\mathcal{L}_{\text{interaction}} = \sum_{i=1}^n \sum_{j=1}^m |\min(\hat{D}_{ij}, 10\text{\AA}) - \min(D_{ij}^{\text{pred}}, 10\text{\AA})|
      Lconfiguration=∑j=1m∑k′=1m1(j,k′)∣D^jk′c−Djk′c∣\mathcal{L}_{\text{configuration}} = \sum_{j=1}^m \sum_{k'=1}^m \mathbf{1}(j, k') |\hat{D}_{jk'}^c - D_{jk'}^c|
    Optimize {x^jc}j=1m\{\hat{x}_j^c\}_{j=1}^m by gradient descent to minimize Lgeneration\mathcal{L}_{\text{generation}}
    return {x^jc}j=1m\{\hat{x}_j^c\}_{j=1}^m

    The local atomic structure mask 1(j,k′)\mathbf{1}(j, k') equals 11 if atoms jj and k′k' share a covalent bond, are separated by two hops, or belong to the same ring structure, and 00 otherwise. Clamping inter-molecular distances to 10A˚10\text{\AA} ensures the optimization concentrates on contact-region interactions.

  6. Knowl 6 — Blind Flexible Self-Docking Performance on PDBbind

    data/table

    TANKBind was evaluated on blind flexible self-docking using the PDBbind v2020 test set of 363 protein-ligand complexes deposited after 2019. In self-docking, binding sites are unknown and ligands are initialized from RDKit-generated 3D conformations.

    Ligand RMSD Centroid Distance
    Percentiles ()Ä ↓\downarrow % Below ↑\uparrow Percentiles ()Ä ↓\downarrow
    Methods 25% 50% 75% Mean 2Ä 5Ä 25% 50% 75% Mean
    QVINA-W 2.5 7.7 23.7 13.6 20.9 40.2 0.9 3.7 22.9 11.9
    GNINA 2.8 8.7 22.1 13.3 21.2 37.1 1.0 4.5 21.2 11.5
    SMINA 3.8 8.1 17.9 12.1 13.5 33.9 1.3 3.7 16.2 9.8
    GLIDE(c.) 2.6 9.3 28.1 16.2 21.8 33.6 0.8 5.6 26.9 14.4
    VINA 5.7 10.7 21.4 14.7 5.5 21.2 1.9 6.2 20.1 12.1
    EQUIBIND-U 3.3 5.7 9.7 7.8 7.2 42.4 1.3 2.6 7.4 5.6
    EQUIBIND 3.8 6.2 10.3 8.2 5.5 39.1 1.3 2.6 7.4 5.6
    TANKBind-R 2.8 5.2 11.2 9.4 16.0 47.9 1.0 2.3 7.7 7.3
    TANKBind-C 2.4 4.5 8.4 8.2 19.6 54.8 0.9 1.9 5.4 6.3
    TANKBind-P 2.6 4.5 8.1 8.5 16.3 54.0 0.9 1.9 5.2 6.4
    TANKBind 2.4 4.0 7.7 7.4 19.3 61.7 0.9 1.7 4.2 5.5

    TANKBind produces ligand conformations with RMSD <5A˚< 5\text{\AA} in 61.7% of test cases, representing an absolute increase of 22.6% over EquiBind (39.1%) and 28.1% over GLIDE (33.6%). Centroid distance below 2A˚2\text{\AA} reached 56.5% for TANKBind compared to 40.0% for EquiBind and 26.5% for Vina.

  7. Knowl 7 — Blind Self-Docking Generalization on Unseen Protein Targets

    data/table

    To assess generalization to novel proteins, methods were evaluated on a subset of 142 PDBbind v2020 test complexes whose protein targets were not present in the training set.

    Ligand RMSD Centroid Distance
    Percentiles ()Ä ↓\downarrow % Below ↑\uparrow Percentiles ()Ä ↓\downarrow
    Methods 25% 50% 75% Mean 2Ä 5Ä 25% 50% 75% Mean
    QVINA-W 3.4 10.3 28.1 16.9 15.3 31.9 1.3 6.5 26.8 15.2
    GNINA 4.5 13.4 27.8 16.7 13.9 27.8 2.0 10.1 27.0 15.1
    SMINA 4.8 10.9 26.0 15.7 9.0 25.7 1.6 6.5 25.7 13.6
    GLIDE 3.4 18.0 31.4 19.6 19.6 28.7 1.1 17.6 29.1 18.1
    VINA 7.9 16.6 27.1 18.7 1.4 12.0 2.4 15.7 26.2 16.1
    EQUIBIND-U 5.7 8.8 14.1 11.0 1.4 21.5 2.6 6.3 12.9 8.9
    EQUIBIND 5.9 9.1 14.3 11.3 0.7 18.8 2.6 6.3 12.9 8.9
    TANKBind-R 3.6 6.9 17.0 12.6 5.6 35.2 1.3 3.6 15.7 10.3
    TANKBind-C 3.4 5.5 9.8 9.9 9.2 43.0 1.1 2.6 8.1 7.9
    TANKBind-P 3.3 5.5 10.9 11.2 5.6 45.1 1.3 2.3 7.9 9.1
    TANKBind 2.9 4.7 8.8 9.1 4.9 55.6 1.3 2.3 4.8 7.0

    On unseen receptors, TANKBind achieves a 55.6% fraction of ligand RMSD <5A˚< 5\text{\AA}, outperforming EquiBind (18.8%), GLIDE (28.7%), and Vina (12.0%). The median ligand RMSD is 4.7\AA for TANKBind versus 9.1\AA for EquiBind and 16.6\AA for Vina.

  8. Knowl 8 — Drug-Protein Binding Affinity Prediction Benchmark

    data/table

    Binding affinity predictions on the PDBbind v2020 test set (evaluating negative log-transformed affinity metrics pKdpK_d, pKipK_i, or pIC50pIC_{50}) were compared across sequence-based, structure-based, complex-based, and TANKBind models.

    Methods RMSE ↓\downarrow Pearson ↑\uparrow Spearman ↑\uparrow MAE ↓\downarrow
    TransCPI 1.741 0.576 0.540 1.404
    MONN 1.438 0.624 0.589 1.143
    PIGNet 2.640 0.511 0.489 2.110
    IGN 1.433 0.698 0.641 1.169
    HOLOPROT 1.546 0.602 0.571 1.208
    STAMPDPI 1.658 0.545 0.411 1.325
    TANKBind 1.346 0.726 0.703 1.070

    TANKBind achieves the lowest RMSE (1.346) and MAE (1.070), and highest Pearson (r=0.726r = 0.726) and Spearman (ρ=0.703\rho = 0.703) correlation coefficients, surpassing complex-based models (IGN, PIGNet) that require true experimental holo-structures as inputs.

  9. Knowl 9 — Ablation Analysis of TANKBind Structural Components

    data/table

    Ablation studies evaluate the contribution of functional pocket segmentation, the bidirectional trigonometry module, and graph encoder architectures on PDBbind binding structure prediction:

    Methods Ligand RMSD ()Ä ↓\downarrow Centroid ()Ä ↓\downarrow Below 2Ä (%) ↑\uparrow Below 5Ä (%) ↑\uparrow
    w/o P2Rank 9.37 7.30 44.90 69.42
    w/o Trig 8.73 6.44 44.08 74.93
    TAPE 8.81 6.89 50.69 73.00
    GAT 8.27 6.23 56.47 78.51
    TankBind-P 8.47 6.44 53.17 74.38
    TankBind-C 8.20 6.27 53.17 73.28
    Origin 7.43 5.51 56.47 77.41

    Key observations from the ablation experiments:

    • Random vs. P2Rank functional segmentation: Replacing P2rank functional blocks with random block segmentation (w/o P2Rank) causes the largest performance drop (mean ligand RMSD rises from 7.43\AA to 9.37\AA).
    • Trigonometry module: Removing trigonometry updates (w/o Trig) increases mean ligand RMSD to 8.73\AA.
    • Bidirectionality: Uni-modal variants summing only over protein nodes (TankBind-P, RMSD 8.47\AA) or compound nodes (TankBind-C, RMSD 8.20\AA) demonstrate that bidirectional geometric message passing is essential.
  10. Knowl 10 — Limitations of the TANKBind Architecture

    limitation

    The authors identify several methodological limitations in TANKBind:

    1. Decoupled Pocket Segmentation: Functional block division relies on a separate upstream heuristic tool (P2rank) rather than learning end-to-end differentiable pocket segmentation.
    2. Rigid Backbone Assumption: Protein structures are modeled at the residue level using static CαC_\alpha coordinates, which does not explicitly account for large-scale backbone conformational adjustments (induced fit) upon ligand binding.
    3. Decoupled Conformation Generation: TANKBind relies on an external conformer generator (such as RDKit) to initialize compound configurations rather than incorporating an internal end-to-end ligand conformer generation module.

Coverage note — Case study details on PDB complexes (e.g., 6K1S, 6QRG, 6KQI) were omitted as specific qualitative illustrations of general empirical results.

References

  1. 1.Michael S Kinch, Zachary Kraft, and Tyler Schwartz. 2021 in review: Fda approvals of new medicines. Drug discovery today, 2022.
  2. 2.Subramanian Boopathi, Adolfo B Poma, and Ponmalai Kolandaivel. Novel 2019 coronavirus structure, mechanism of action, antiviral drug promises and rule out against its treatment. Journal of Biomolecular Structure and Dynamics, 39(9):3409–3418, 2021.
  3. 3.Lei Xie, Li Xie, and Philip E Bourne. Structure-based systems biology for analyzing off-target binding. Current opinion in structural biology, 21(2):189–199, 2011.
  4. 4.Zhihai Liu, Yan Li, Li Han, Jie Li, Jie Liu, Zhixiong Zhao, Wei Nie, Yuchen Liu, and Renxiao Wang. Pdb-wide collection of binding data: current status of the pdbbind database. Bioinformatics, 31(3): 405–412, 2015.
  5. 5.Jean-Louis Reymond, Ruud Van Deursen, Lorenz C Blum, and Lars Ruddigkeit. Chemical space as a source for new drugs. MedChemComm, 1(1):30–38, 2010.
  6. 6.Elena A Ponomarenko, Ekaterina V Poverennaya, Ekaterina V Ilgisonis, Mikhail A Pyatnitskiy, Arthur T Kopylov, Victor G Zgoda, Andrey V Lisitsa, and Alexander I Archakov. The size of the human proteome: the width and depth. International journal of analytical chemistry, 2016, 2016.
  7. 7.Oleg Trott and Arthur J Olson. Autodock vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. Journal of computational chemistry, 31(2):455–461, 2010.
  8. 8.Richard A Friesner, Jay L Banks, Robert B Murphy, Thomas A Halgren, Jasna J Klicic, Daniel T Mainz, Matthew P Repasky, Eric H Knoll, Mee Shelley, Jason K Perry, et al. Glide: a new approach for rapid, accurate docking and scoring. 1. method and assessment of docking accuracy. Journal of medicinal chemistry, 47(7):1739–1749, 2004.
  9. 9.Suzanne Ackloo, Rima Al-awar, Rommie E Amaro, Cheryl H Arrowsmith, Hatylas Azevedo, Robert A Batey, Yoshua Bengio, Ulrich AK Betz, Cristian G Bologa, John D Chodera, et al. Cache (critical assessment of computational hit-finding experiments): A public–private partner-ship benchmarking initiative to enable the development of computational methods for hit-finding. Nature Reviews Chemistry, pages 1–9, 2022.
  10. 10.Francesco Gentile, Jean Charle Yaacoub, James Gleave, Michael Fernandez, Anh-Tien Ton, Fuqiang Ban, Abraham Stern, and Artem Cherkasov. Artificial intelligence–enabled virtual screening of ultra-large chemical libraries with deep docking. Nature Protocols, pages 1–26, 2022.
  11. 11.Amy C Anderson. The process of structure-based drug design. Chemistry & biology, 10(9):787–797, 2003.
  12. 12.Ajay N Jain. Scoring functions for protein-ligand docking. Current Protein and Peptide Science, 7 (5):407–420, 2006.
  13. 13.John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  14. 14.Mingchen Chen, Xun Chen, Nicholas P Schafer, Cecilia Clementi, Elizabeth A Komives, Diego U Ferreiro, and Peter G Wolynes. Surveying biomolecular frustration at atomic resolution. Nature communications, 11(1):1–9, 2020a.
  15. 15.José Nelson Onuchic, Zaida Luthey-Schulten, and Peter G Wolynes. Theory of protein folding: the energy landscape perspective. Annual review of physical chemistry, 48(1):545–600, 1997.
  16. 16.David De Juan, Florencio Pazos, and Alfonso Valencia. Emerging methods in protein co-evolution. Nature Reviews Genetics, 14(4):249–261, 2013.
  17. 17.Fabian Glaser, Tal Pupko, Inbal Paz, Rachel E Bell, Dalit Bechor-Shental, Eric Martz, and Nir Ben-Tal. Consurf: identification of functional regions in proteins by surface-mapping of phylogenetic information. Bioinformatics, 19(1):163–164, 2003.
  18. 18.Minkyung Baek, Frank DiMaio, Ivan Anishchenko, Justas Dauparas, Sergey Ovchinnikov, Gyu Rie Lee, Jue Wang, Qian Cong, Lisa N Kinch, R Dustin Schaeffer, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science, 373(6557):871–876, 2021.
  19. 19.Bowen Jing, Stephan Eismann, Pratham N Soni, and Ron O Dror. Equivariant graph neural networks for 3d macromolecular structure. arXiv preprint arXiv:2106.03843, 2021.
  20. 20.Octavian-Eugen Ganea, Xinyuan Huang, Charlotte Bunne, Yatao Bian, Regina Barzilay, Tommi Jaakkola, and Andreas Krause. Independent se (3)-equivariant models for end-to-end rigid protein docking. arXiv preprint arXiv:2111.07786, 2021.
  21. 21.Wengong Jin, Jeremy Wohlwend, Regina Barzilay, and Tommi Jaakkola. Iterative refinement graph neural network for antibody sequence-structure co-design. arXiv preprint arXiv:2110.04624, 2021.
  22. 22.John Ingraham, Vikas Garg, Regina Barzilay, and Tommi Jaakkola. Generative models for graph-based protein design. Advances in Neural Information Processing Systems, 32, 2019.
  23. 23.Mohammed AlQuraishi. End-to-end differentiable learning of protein structure. Cell systems, 8(4): 292–301, 2019.
  24. 24.Kristof Schütt, Pieter-Jan Kindermans, Huziel Enoc Sauceda Felix, Stefan Chmiela, Alexandre Tkatchenko, and Klaus-Robert Müller. Schnet: A continuous-filter convolutional neural network for modeling quantum interactions. Advances in neural information processing systems, 30, 2017.
  25. 25.Vignesh Ram Somnath, Charlotte Bunne, and Andreas Krause. Multi-scale representation learning on proteins. Advances in Neural Information Processing Systems, 34, 2021.
  26. 26.Chence Shi, Shitong Luo, Minkai Xu, and Jian Tang. Learning gradient fields for molecular conformation generation. In International Conference on Machine Learning, pages 9558–9568. PMLR, 2021.
  27. 27.Minkai Xu, Lantao Yu, Yang Song, Chence Shi, Stefano Ermon, and Jian Tang. Geodiff: A geometric diffusion model for molecular conformation generation. arXiv preprint arXiv:2203.02923, 2022.
  28. 28.Oscar Méndez-Lucio, Mazen Ahmad, Ehecatl Antonio del Rio-Chanona, and Jörg Kurt Wegner. A geometric deep learning approach to predict binding conformations of bioactive molecules. Nature Machine Intelligence, 3(12):1033–1039, 2021.
  29. 29.Jian Wang and Nikolay V Dokholyan. Yuel: Compound-protein interaction prediction with high generalizability. bioRxiv, 2021.
  30. 30.Shuya Li, Fangping Wan, Hantao Shu, Tao Jiang, Dan Zhao, and Jianyang Zeng. Monn: a multi-objective neural network for predicting compound-protein interactions and affinities. Cell Systems, 10(4):308–322, 2020.
  31. 31.Masashi Tsubaki, Kentaro Tomii, and Jun Sese. Compound–protein interaction prediction with end-to-end learning of neural networks for graphs and sequences. Bioinformatics, 35(2):309–318, 2019.
  32. 32.Kyle Yingkai Gao, Achille Fokoue, Heng Luo, Arun Iyengar, Sanjoy Dey, and Ping Zhang. Interpretable drug target prediction using deep neural representation. In IJCAI, volume 2018, pages 3371–3377, 2018.
  33. 33.Mostafa Karimi, Di Wu, Zhangyang Wang, and Yang Shen. Deepaffinity: interpretable deep learning of compound–protein affinity through unified recurrent and convolutional neural networks. Bioinformatics, 35(18):3329–3338, 2019.
  34. 34.Shuangjia Zheng, Yongjian Li, Sheng Chen, Jun Xu, and Yuedong Yang. Predicting drug–protein interaction using quasi-visual question answering system. Nature Machine Intelligence, 2(2): 134–140, 2020.
  35. 35.José Jiménez, Miha Skalic, Gerard Martinez-Rosell, and Gianni De Fabritiis. K deep: protein–ligand absolute binding affinity prediction via 3d-convolutional neural networks. Journal of chemical information and modeling, 58(2):287–296, 2018.
  36. 36.Jaechang Lim, Seongok Ryu, Kyubyong Park, Yo Joong Choe, Jiyeon Ham, and Woo Youn Kim. Predicting drug–target interaction using a novel graph neural network with 3d structure-embedded graph representation. Journal of chemical information and modeling, 59(9):3981–3988, 2019.
  37. 37.Joseph A Morrone, Jeffrey K Weber, Tien Huynh, Heng Luo, and Wendy D Cornell. Combining docking pose rank and structure with deep learning improves protein–ligand binding mode prediction over a baseline docking approach. Journal of chemical information and modeling, 60(9): 4170–4179, 2020.
  38. 38.Hannes Stärk, Octavian-Eugen Ganea, Lagnajit Pattanaik, Regina Barzilay, and Tommi Jaakkola. Equibind: Geometric deep learning for drug binding structure prediction. arXiv preprint arXiv:2202.05146, 2022.
  39. 39.Bowen Jing, Stephan Eismann, Patricia Suriana, Raphael JL Townshend, and Ron Dror. Learning from protein structure with geometric vector perceptrons. arXiv preprint arXiv:2009.01411, 2020.
  40. 40.Dejun Jiang, Chang-Yu Hsieh, Zhenxing Wu, Yu Kang, Jike Wang, Ercheng Wang, Ben Liao, Chao Shen, Lei Xu, Jian Wu, et al. Interactiongraphnet: A novel and efficient deep graph representation learning framework for accurate protein–ligand interaction predictions. Journal of medicinal chemistry, 64(24):18209–18232, 2021.
  41. 41.Pablo Gainza, Freyr Sverrisson, Frederico Monti, Emanuele Rodola, D Boscaini, MM Bronstein, and BE Correia. Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning. Nature Methods, 17(2):184–192, 2020.
  42. 42.John M Jumper, Nabil F Faruk, Karl F Freed, and Tobin R Sosnick. Accurate calculation of side chain packing and free energy with applications to protein molecular dynamics. PLoS computational biology, 14(12):e1006342, 2018.
  43. 43.Radoslav Krivák and David Hoksza. P2rank: machine learning based tool for rapid and accurate prediction of ligand binding sites from protein structure. Journal of cheminformatics, 10(1):1–12, 2018.
  44. 44.Zhaocheng Zhu, Chence Shi, Zuobai Zhang, Shengchao Liu, Minghao Xu, Xinyu Yuan, Yangtian Zhang, Junkun Chen, Huiyu Cai, Jiarui Lu, et al. Torchdrug: A powerful and flexible machine learning platform for drug discovery. arXiv preprint arXiv:2202.08320, 2022.
  45. 45.Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks? arXiv preprint arXiv:1810.00826, 2018.
  46. 46.Raphael JL Townshend, Martin Vögele, Patricia Suriana, Alexander Derry, Alexander Powers, Yianni Laloudakis, Sidhika Balachandar, Bowen Jing, Brandon Anderson, Stephan Eismann, et al. Atom3d: Tasks on molecules in three dimensions. arXiv preprint arXiv:2012.04035, 2020.
  47. 47.Raia Hadsell, Sumit Chopra, and Yann LeCun. Dimensionality reduction by learning an invariant mapping. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), volume 2, pages 1735–1742. IEEE, 2006.
  48. 48.Matthew Masters, Amr H Mahmoud, Yao Wei, and Markus Lill. Deep learning model for flexible and efficient protein-ligand docking. In ICLR2022 Machine Learning for Drug Discovery, 2022.
  49. 49.Moritz Hoffmann and Frank Noé. Generating valid euclidean distance matrices. arXiv preprint arXiv:1910.03131, 2019.
  50. 50.Zsolt Zsoldos, Darryl Reid, Aniko Simon, Sayyed Bashir Sadjad, and A Peter Johnson. ehits: a new fast, exhaustive flexible ligand docking system. Journal of Molecular Graphics and Modelling, 26 (1):198–212, 2007.
  51. 51.Stephen K Burley, Charmi Bhikadiya, Chunxiao Bi, Sebastian Bittrich, Li Chen, Gregg V Crichlow, Cole H Christie, Kenneth Dalenberg, Luigi Di Costanzo, Jose M Duarte, et al. Rcsb protein data bank: powerful new tools for exploring 3d structures of biological macromolecules for basic and applied research and education in fundamental biology, biomedicine, biotechnology, bioengineering and energy sciences. Nucleic acids research, 49(D1):D437–D451, 2021.
  52. 52.Greg Landrum et al. Rdkit: A software suite for cheminformatics, computational chemistry, and predictive modeling, 2013.
  53. 53.Andrew T McNutt, Paul Francoeur, Rishal Aggarwal, Tomohide Masuda, Rocco Meli, Matthew Ragoza, Jocelyn Sunseri, and David Ryan Koes. Gnina 1.0: molecular docking with deep learning. Journal of cheminformatics, 13(1):1–20, 2021.
  54. 54.David Ryan Koes, Matthew P Baumgartner, and Carlos J Camacho. Lessons learned in empirical scoring with smina from the csar 2011 benchmarking exercise. Journal of chemical information and modeling, 53(8):1893–1904, 2013.
  55. 55.Lifan Chen, Xiaoqin Tan, Dingyan Wang, Feisheng Zhong, Xiaohong Liu, Tianbiao Yang, Xiaomin Luo, Kaixian Chen, Hualiang Jiang, and Mingyue Zheng. Transformercpi: improving compound–protein interaction prediction by sequence-based deep learning with self-attention mechanism and label reversal experiments. Bioinformatics, 36(16):4406–4414, 2020b.
  56. 56.Seokhyun Moon, Wonho Zhung, Soojung Yang, Jaechang Lim, and Woo Youn Kim. Pignet: a physics-informed deep learning model toward generalized drug–target interaction predictions. Chemical Science, 13(13):3661–3673, 2022.
  57. 57.Penglei Wang, Shuangjia Zheng, Yize Jiang, Chengtao Li, Junhong Liu, Chang Wen, Atanas Patronov, Dahong Qian, Hongming Chen, and Yuedong Yang. Structure-aware multimodal deep learning for drug–protein interaction prediction. Journal of Chemical Information and Modeling, 62(5): 1308–1317, 2022.
  58. 58.Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Peter Chen, John Canny, Pieter Abbeel, and Yun Song. Evaluating protein transfer learning with tape. Advances in neural information processing systems, 32, 2019.
  59. 59.Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017.
  60. 60.Timo Lassmann. Kalign 3: multiple sequence alignment of large datasets, 2020.
  61. 61.Peter JA Cock, Tiago Antao, Jeffrey T Chang, Brad A Chapman, Cymon J Cox, Andrew Dalke, Iddo Friedberg, Thomas Hamelryck, Frank Kauff, Bartek Wilczynski, et al. Biopython: freely available python tools for computational molecular biology and bioinformatics. Bioinformatics, 25(11): 1422–1423, 2009.
  62. 62.Mengyao Zhao, Wan-Ping Lee, Erik P Garrison, and Gabor T Marth. Ssw library: an simd smith-waterman c/c++ library for use in genomic applications. PloS one, 8(12):e82138, 2013.

Citation

MLA
Lu, W., et al. “TANKBind: Trigonometry-Aware Neural NetworKs for Drug-Protein Binding Structure Prediction”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 7236–49, https://proceedings.neurips.cc/paper_files/paper/2022/file/2f89a23a19d1617e7fb16d4f7a049ce2-Paper-Conference.pdf.
APA
Lu, W., Wu, Q., Zhang, J., Rao, J., Li, C., & Zheng, S. (2022). TANKBind: Trigonometry-Aware Neural NetworKs for Drug-Protein Binding Structure Prediction. Advances in Neural Information Processing Systems, 35, 7236–7249. https://proceedings.neurips.cc/paper_files/paper/2022/file/2f89a23a19d1617e7fb16d4f7a049ce2-Paper-Conference.pdf
Chicago
Lu, W., Q. Wu, J. Zhang, J. Rao, C. Li, and S. Zheng. 2022. “TANKBind: Trigonometry-Aware Neural NetworKs for Drug-Protein Binding Structure Prediction”. Advances in Neural Information Processing Systems 35: 7236–49. https://proceedings.neurips.cc/paper_files/paper/2022/file/2f89a23a19d1617e7fb16d4f7a049ce2-Paper-Conference.pdf.
Harvard
Lu, W. et al. (2022) “TANKBind: Trigonometry-Aware Neural NetworKs for Drug-Protein Binding Structure Prediction”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 7236–7249. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/2f89a23a19d1617e7fb16d4f7a049ce2-Paper-Conference.pdf.
Vancouver
1. Lu W, Wu Q, Zhang J, Rao J, Li C, Zheng S (2022) TANKBind: Trigonometry-Aware Neural NetworKs for Drug-Protein Binding Structure Prediction. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 7236–7249

BibTeX

@inproceedings{lu2022tankbind,
  title = {TANKBind: Trigonometry-Aware Neural NetworKs for Drug-Protein Binding Structure Prediction},
  author = {Lu, Wei and Wu, Qifeng and Zhang, Jixian and Rao, Jiahua and Li, Chengtao and Zheng, Shuangjia},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {7236-7249},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/2f89a23a19d1617e7fb16d4f7a049ce2-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors