Protein Representation Learning by Geometric Structure Pretraining

Zuobai ZhangMinghao XuArian Rokkum JamasbVijil ChenthamarakshanAurélie C. LozanoPayel DasJian Tang

article2023ICLR338 citations

Develops a geometric pretraining framework for protein graphs using multiview contrastive learning and self-prediction tasks, matching or exceeding sequence-based language models on function and fold classification while requiring far less training data.

Listen

Understanding protein function and structure is central to modern biotechnology, therapeutics, and drug discovery. While recent computational advances have trained powerful machine learning models on massive databases of linear protein sequences, a protein's biological role is fundamentally determined by its three-dimensional folded shape. Until recently, the scarcity of experimentally verified structures prevented researchers from training models directly on geometric data at scale. The article addresses this limitation by establishing a self-supervised pretraining framework that learns protein representations directly from 3D structures. The main objective is to demonstrate that pretraining on 3D geometric configurations allows models to achieve competitive or superior predictive performance on functional and structural tasks while using significantly less training data than sequence-based alternatives.

To accomplish this, the authors designed a specialized graph neural network called the Geometry-Aware Relational Graph Neural Network, along with an enhanced variant that performs sparse message passing between edges to explicitly model spatial angles and distances between protein residues. The model was pretrained without human labels using two main strategies: a multiview contrastive learning approach that aligns representations of biologically related structural subregions, and four self-prediction tasks that reconstruct masked biochemical attributes, distances, and angles. Pretraining was conducted on approximately 805,000 structures from the AlphaFold database, and models were evaluated across standard benchmarks encompassing enzyme function prediction, Gene Ontology term classification, fold classification, and reaction categorization.

The findings show that geometric pretraining provides substantial performance advantages across all evaluated domains. First, the multiview contrastive pretraining method achieved state-of-the-art results on seven out of eight downstream benchmark datasets, substantially outperforming models trained from scratch. Second, despite being pretrained on fewer than one million structural samples, the model performed on par with or better than leading protein language models trained on tens of millions to billions of sequences; for example, achieving an F-score of 0.874 on enzyme classification compared to 0.864 for a benchmark model trained on 24 million sequences. Third, while sequence-based models struggle to categorize protein folds, the structure-based encoder achieved up to 78.1% average accuracy on fold classification. Finally, ensembling the neural network with existing structure retrieval tools produced further additive gains across multiple functional tasks.

These results imply that capturing 3D geometry directly provides a far more data-efficient pathway for biological modeling, significantly reducing computational requirements while improving functional annotation accuracy. For organizations developing computational pipelines for protein engineering and discovery, adopting geometric pretraining represents a viable strategy to enhance prediction reliability and streamline candidate screening. Practitioners are encouraged to explore hybrid configurations combining structural encoders, sequence features, and retrieval tools to maximize performance. However, decision-makers should recognize that the study focused primarily on single-chain functional and fold classification using under one million structures. Expanding validation to multi-protein interactions, ligand binding, and larger structural repositories will be an essential next step before full deployment in critical therapeutic development.

Cover for Protein Representation Learning by Geometric Structure Pretraining

Abstract

Learning effective protein representations is critical in a variety of tasks in biology such as predicting protein function or structure. Existing approaches usually pretrain protein language models on a large number of unlabeled amino acid sequences and then finetune the models with some labeled data in downstream tasks. Despite the effectiveness of sequence-based approaches, the power of pretraining on known protein structures, which are available in smaller numbers only, has not been explored for protein property prediction, though protein structures are known to be determinants of protein function. In this paper, we propose to pretrain protein representations according to their 3D structures. We first present a simple yet effective encoder to learn the geometric features of a protein. We pretrain the protein graph encoder by leveraging multiview contrastive learning and different self-prediction tasks. Experimental results on both function prediction and fold classification tasks show that our proposed pretraining methods outperform or are on par with the state-of-the-art sequence-based methods, while using much less pretraining data. Our implementation is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Structure-Based Protein Encoder
  • 3.1 Geometry-Aware Relational Graph Neural Network
  • 3.2 Edge Message Passing Layer
  • 4 Geometric Pretraining Methods
  • 4.1 Multiview Contrastive Learning
  • 4.2 Straightforward Baselines: Self-Prediction Methods
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Results
  • 5.3 Ablation Studies
  • 6 Conclusions and Future Work
  • References
  • A More Related Work
  • A.1 Sequence-Based Methods for Protein Representation Learning
  • A.2 Structure-based Methods for Biological Molecules
  • A.3 Pretraining Graph Neural Networks
  • B Potential Negative Impact
  • C More Details of GearNet
  • C.1 Protein Graph Construction
  • C.2 Enhance GearNet with IEConv Layers
  • D Self-Prediction Methods
  • E Experimental Details
  • E.1 Dataset Statistics
  • E.2 Evaluation Metrics
  • E.3 Implementation Details
  • F Additional Experimental Results on EC and GO Prediction
  • G Structure Pretraining on EGNN
  • H Combine Sequence- and Structure-Based Encoders
  • I Combine Neural and Retrieval-Based Methods
  • J Latent Space Visualization
  • K Residue-Level Explanation

Knowls

  1. Knowl 1 — GearNet: Geometry-Aware Relational Graph Neural Network for Protein Structure Encoding

    model/method

    GearNet (GeomEtry-Aware Relational Graph Neural Network) models a protein as a residue-level directed relational graph G=(V,E,R)G = (V, \mathcal{E}, \mathcal{R}), where each node i∈Vi \in V represents the Cα\text{C}_\alpha atom of a residue with 3D coordinate vector xi∈R3x_i \in \mathbb{R}^3, and R\mathcal{R} denotes the set of distinct relational edge types.

    Graph construction includes three classes of directed edges:

    1. Sequential edges: Directed edges connecting residue ii to residue jj when ∣j−i∣<dseq|j - i| < d_{\text{seq}} (with dseq=3d_{\text{seq}} = 3). The relation type corresponds to the signed relative sequence distance d=j−i∈{−2,−1,0,1,2}d = j - i \in \{-2, -1, 0, 1, 2\}, yielding 2dseq−1=52d_{\text{seq}} - 1 = 5 sequential relation types.
    2. Radius edges: Directed edges between residues ii and jj if their Euclidean distance satisfies ∥xi−xj∥2<dradius\|x_i - x_j\|_2 < d_{\text{radius}} (with dradius=10.0 A˚d_{\text{radius}} = 10.0\,\text{Å}) and ∣i−j∣≥dlong=5|i - j| \ge d_{\text{long}} = 5.
    3. KK-nearest neighbor (KNN) edges: Directed edges connecting each residue ii to its k=10k = 10 nearest spatial neighbors satisfying ∣i−j∣≥dlong=5|i - j| \ge d_{\text{long}} = 5.

    In total, there are ∣R∣=2dseq+1=7|\mathcal{R}| = 2d_{\text{seq}} + 1 = 7 edge types. Node features fi∈{0,1}21f_i \in \{0, 1\}^{21} are 21-dimensional one-hot encodings of residue types (20 standard amino acids plus one unknown indicator). Edge features are computed as f(i,j,r)=Cat(fi,fj,onehot(r),∣i−j∣,∥xi−xj∥2)f_{(i,j,r)} = \text{Cat}(f_i, f_j, \text{onehot}(r), |i - j|, \|x_i - x_j\|_2).

    The relational graph convolutional layer updates node representations via:

    hi(0)=fih_i^{(0)} = f_i

    ui(l)=σ(BN(∑r∈RWr∑j∈Nr(i)hj(l−1)))u_i^{(l)} = \sigma\left(\text{BN}\left(\sum_{r \in \mathcal{R}} W_r \sum_{j \in \mathcal{N}_r(i)} h_j^{(l-1)}\right)\right)

    hi(l)=hi(l−1)+ui(l)h_i^{(l)} = h_i^{(l-1)} + u_i^{(l)}

    where hi(l)h_i^{(l)} is the hidden representation of residue ii at layer ll, Nr(i)={j∈V∣(j,i,r)∈E}\mathcal{N}_r(i) = \{j \in V \mid (j, i, r) \in \mathcal{E}\} is the relational neighborhood of node ii under relation rr, WrW_r is a learnable relation-specific weight matrix, BN(⋅)\text{BN}(\cdot) denotes batch normalization, and σ(⋅)\sigma(\cdot) is the ReLU activation function. Because node and edge inputs use relative coordinates and invariant features, GearNet is E(3)\text{E}(3)-invariant.

  2. Knowl 2 — GearNet-Edge Message Passing Mechanism

    model/method

    GearNet-Edge extends GearNet with an edge-level message passing mechanism defined over the line graph G′=(V′,E′,R′)G' = (V', \mathcal{E}', \mathcal{R}') of the residue graph G=(V,E,R)G = (V, \mathcal{E}, \mathcal{R}), modeling spatial dependencies and angular interactions between connected edges.

    Each node in V′V' corresponds to an original directed edge (i,j,r1)∈E(i, j, r_1) \in \mathcal{E}. A directed edge connects node (w,k,r2)(w, k, r_2) to node (i,j,r1)(i, j, r_1) in G′G' if and only if j=wj = w and i≠ki \ne k. The relation type r′∈R′r' \in \mathcal{R}' in graph G′G' is determined by discretizing the geometric angle θ∈[0,π]\theta \in [0, \pi] between edges (w,k,r2)(w, k, r_2) and (i,j,r1)(i, j, r_1) into 8 equal bins.

    The edge message passing layer updates edge representations via:

    m(i,j,r1)(0)=f(i,j,r1)m_{(i,j,r_1)}^{(0)} = f_{(i,j,r_1)}

    m(i,j,r1)(l)=σ(BN(∑r′∈R′Wr′′∑(w,k,r2)∈Nr′′((i,j,r1))m(w,k,r2)(l−1)))m_{(i,j,r_1)}^{(l)} = \sigma\left(\text{BN}\left(\sum_{r' \in \mathcal{R}'} W'_{r'} \sum_{(w,k,r_2) \in \mathcal{N}'_{r'}((i,j,r_1))} m_{(w,k,r_2)}^{(l-1)}\right)\right)

    where m(i,j,r1)(l)m_{(i,j,r_1)}^{(l)} is the message vector for edge (i,j,r1)(i, j, r_1) at layer ll, Nr′′((i,j,r1))={(w,k,r2)∈V′∣((w,k,r2),(i,j,r1),r′)∈E′}\mathcal{N}'_{r'}((i,j,r_1)) = \{(w, k, r_2) \in V' \mid ((w, k, r_2), (i, j, r_1), r') \in \mathcal{E}'\} is the incoming neighborhood of (i,j,r1)(i, j, r_1) under relation r′r', Wr′′W'_{r'} is a relation-specific kernel matrix on G′G', BN(⋅)\text{BN}(\cdot) denotes batch normalization, and σ(⋅)\sigma(\cdot) is the ReLU activation.

    The edge messages are then integrated into the residue node aggregation step in the primary graph:

    ui(l)=σ(BN(∑r∈RWr∑j∈Nr(i)(hj(l−1)+FC(m(j,i,r)(l)))))u_i^{(l)} = \sigma\left(\text{BN}\left(\sum_{r \in \mathcal{R}} W_r \sum_{j \in \mathcal{N}_r(i)} \left(h_j^{(l-1)} + \text{FC}\left(m_{(j,i,r)}^{(l)}\right)\right)\right)\right)

    where FC(⋅)\text{FC}(\cdot) denotes a linear projection layer mapping edge message embeddings to the node hidden dimension.

  3. Knowl 3 — Multiview Contrastive Learning for Protein Structure Pretraining

    model/method

    Multiview Contrastive Learning pretrains protein structure encoders by maximizing mutual information between biologically correlated substructures extracted from the same protein while separating substructures from different proteins.

    Given a protein residue graph G=(V,E,R)G = (V, \mathcal{E}, \mathcal{R}), two views GxG_x and GyG_y are constructed using two substructure sampling schemes and two noise functions:

    1. Subsequence Cropping: Randomly selects start and end residues ll and rr with length r−l=50r - l = 50, extracting the subgraph induced by Vl,r(seq)={i∈V∣l≤i≤r}V_{l,r}^{(\text{seq})} = \{i \in V \mid l \le i \le r\} to model sequential domains.
    2. Subspace Cropping: Randomly selects center residue pp and collects all residues within Euclidean distance d=15 A˚d = 15\,\text{Å}, extracting the subgraph induced by Vp,d(space)={i∈V∣∥xi−xp∥2≤d}V_{p,d}^{(\text{space})} = \{i \in V \mid \|x_i - x_p\|_2 \le d\} to capture 3D structural motifs.

    Each sampled subgraph is perturbed by one of two noise functions chosen uniformly at random: the identity transformation, or random edge masking with edge dropout probability 0.150.15.

    Graph representations hx,hyh_x, h_y are obtained from the encoder and projected to latent vectors zx,zyz_x, z_y via a two-layer MLP projection head. For a mini-batch of BB proteins yielding 2B2B views, the InfoNCE loss for positive pair (x,y)(x, y) is:

    Lx,y=−log⁡exp⁡(sim(zx,zy)/τ)∑k=12B1[k≠x]exp⁡(sim(zx,zk)/τ)\mathcal{L}_{x,y} = -\log \frac{\exp(\text{sim}(z_x, z_y)/\tau)}{\sum_{k=1}^{2B} \mathbf{1}_{[k \ne x]} \exp(\text{sim}(z_x, z_k)/\tau)}

    where sim(u,v)=u⊤v∥u∥2∥v∥2\text{sim}(u, v) = \frac{u^\top v}{\|u\|_2 \|v\|_2} is cosine similarity, τ=0.07\tau = 0.07 is the temperature hyperparameter, and views from other proteins in the mini-batch serve as negative samples.

  4. Knowl 4 — Self-Prediction Pretraining Objectives for Protein Structures

    model/method

    Four self-supervised pretraining objectives reconstruct masked geometric and physicochemical attributes of protein residue graphs:

    1. Residue Type Prediction: Evaluates masked amino acid identity prediction (masked inverse folding). A random subset of residue node features are masked, and the model minimizes cross-entropy:

    Li=CE(fresidue(hi′),fi)\mathcal{L}_i = \text{CE}\left(f_{\text{residue}}(h'_i), f_i\right)

    where hi′h'_i is the representation of node ii obtained from the masked graph, fif_i is the true residue one-hot label, fresidue(⋅)f_{\text{residue}}(\cdot) is an MLP prediction head, and CE(⋅)\text{CE}(\cdot) is cross-entropy loss.

    1. Distance Prediction: A subset of 256 edges are masked from the graph, and the Euclidean distance between connected residue pairs is predicted via mean squared error:

    L(i,j,r)=(fdist(hi′,hj′)−∥xi−xj∥2)2\mathcal{L}_{(i,j,r)} = \left(f_{\text{dist}}(h'_i, h'_j) - \|x_i - x_j\|_2\right)^2

    where fdist(⋅)f_{\text{dist}}(\cdot) is an MLP head taking the concatenated representations (hi′,hj′)(h'_i, h'_j).

    1. Angle Prediction: For 512 sampled adjacent edge pairs ((i,j,r1),(j,k,r2))((i, j, r_1), (j, k, r_2)), the spatial angle ∠ijk∈[0,π]\angle ijk \in [0, \pi] is discretized into 8 bins via function bin(⋅)\text{bin}(\cdot) and predicted via cross-entropy:

    L(i,j,r1),(j,k,r2)=CE(fangle(hi′,hj′,hk′),bin(∠ijk))\mathcal{L}_{(i,j,r_1),(j,k,r_2)} = \text{CE}\left(f_{\text{angle}}(h'_i, h'_j, h'_k), \text{bin}(\angle ijk)\right)

    1. Dihedral Prediction: For 512 sampled adjacent edge triplets ((i,j,r1),(j,k,r2),(k,t,r3))((i, j, r_1), (j, k, r_2), (k, t, r_3)), the dihedral angle ∠ijkt∈[0,π]\angle ijkt \in [0, \pi] is discretized into 8 bins and predicted via cross-entropy:

    L(i,j,r1),(j,k,r2),(k,t,r3)=CE(fdih(hi′,hj′,hk′,ht′),bin(∠ijkt))\mathcal{L}_{(i,j,r_1),(j,k,r_2),(k,t,r_3)} = \text{CE}\left(f_{\text{dih}}(h'_i, h'_j, h'_k, h'_t), \text{bin}(\angle ijkt)\right)

  5. Knowl 5 — GearNet-Edge-IEConv Hybrid Architecture for Fold Classification

    model/method

    To improve structural fold classification, GearNet-Edge is combined with a simplified Intrinsic-Extrinsic Convolution (IEConv) layer that dynamically computes convolution kernel matrices from relative spatial and sequential edge features.

    For node ii at layer ll, the simplified IEConv message aggregation is defined as:

    h~i(l)=∑j∈N(i)ko(f(G,i,j))⋅hj(l−1)\tilde{h}_i^{(l)} = \sum_{j \in \mathcal{N}(i)} k_o\left(f(G, i, j)\right) \cdot h_j^{(l-1)}

    where N(i)\mathcal{N}(i) is the neighborhood of residue node ii, f(G,i,j)f(G, i, j) is the feature vector of relative positions between residues ii and jj, and ko(⋅)k_o(\cdot) is an MLP mapping geometric features directly to a convolution kernel matrix. The dynamic receptive fields, continuous pooling layers, and smoothing heuristics of original IEConv are omitted so that convolution operates over the static residue relational graph.

    The combined GearNet-Edge-IEConv node update incorporates residual connections, relational convolutions, edge message passing, and dynamic IEConv representations:

    hi(l)=hi(l−1)+ui(l)+h~i(l)h_i^{(l)} = h_i^{(l-1)} + u_i^{(l)} + \tilde{h}_i^{(l)}

    where ui(l)u_i^{(l)} is the relational edge-enhanced update from GearNet-Edge and h~i(l)\tilde{h}_i^{(l)} is the output of the simplified IEConv layer.

  6. Knowl 6 — Downstream Performance of GearNet and Structure-Pretrained Models

    data/table

    GearNet, GearNet-Edge, and their pretrained variants were evaluated against sequence-based, structure-based, and pretrained baselines on Enzyme Commission (EC) number prediction, Gene Ontology (GO) term prediction across Biological Process (BP), Molecular Function (MF), and Cellular Component (CC), SCOPe 1.75 Fold Classification (evaluated across Fold, Superfamily, and Family test splits), and Reaction classification. EC and GO tasks are evaluated by protein-centric maximum F-score (Fmax⁡F_{\max}), while Fold and Reaction tasks are evaluated by classification accuracy (%). Models pretrained on the AlphaFold Database (805K structures: 365K proteome-wide predictions + 440K Swiss-Prot predictions) outperform structure baselines trained from scratch and match or exceed sequence-based models trained on orders of magnitude larger datasets (e.g., ESM-1b on 24M sequences and ProtBERT-BFD on 2.1B sequences).

    Method Dataset (Size) EC GO Fold Classification Reaction
    BP MF CC Fold Super. Fam. Avg.
    w/o pretraining
    CNN - 0.545 0.244 0.354 0.287 11.3 13.4 53.4 26.0 51.7
    ResNet - 0.605 0.280 0.405 0.304 10.1 7.21 23.5 13.6 24.1
    LSTM - 0.425 0.225 0.321 0.283 6.41 4.33 18.1 9.61 11.0
    Transformer - 0.238 0.264 0.211 0.405 9.22 8.81 40.4 19.4 26.6
    GCN - 0.320 0.252 0.195 0.329 16.8 21.3 82.8 40.3 67.3
    GAT - 0.368 0.284 0.317 0.385 12.4 16.5 72.7 33.8 55.6
    GVP - 0.489 0.326 0.426 0.420 16.0 22.5 83.8 40.7 65.5
    3DCNN_MQA - 0.077 0.240 0.147 0.305 31.6 45.4 92.5 56.5 72.2
    GraphQA - 0.509 0.308 0.329 0.413 23.7 32.5 84.4 46.9 60.8
    New IEConv - 0.735 0.374 0.544 0.444 47.6 70.2 99.2 72.3 87.2
    GearNet - 0.730 0.356 0.503 0.414 28.4 42.6 95.3 55.4 79.4
    GearNet-IEConv - 0.800 0.381 0.563 0.422 42.3 64.1 99.1 68.5 83.7
    GearNet-Edge - 0.810 0.403 0.580 0.450 44.0 66.7 99.1 69.9 86.6
    GearNet-Edge-IEConv - 0.810 0.400 0.581 0.430 48.3 70.3 99.5 72.7 85.3
    w/ pretraining
    DeepFRI Pfam (10M) 0.631 0.399 0.465 0.460 15.3 20.6 73.2 36.4 63.3
    ESM-1b UniRef50 (24M) 0.864 0.452 0.657 0.477 26.8 60.1 97.8 61.5 83.1
    ProtBERT-BFD BFD (2.1B) 0.838 0.279 0.456 0.408 26.6 55.8 97.6 60.0 72.2
    LM-GVP UniRef100 (216M) 0.664 0.417 0.545 0.527 - - - - -
    New IEConv PDB (476K) - - - - 50.3 80.6 99.7 76.9 87.6
    Residue Type Pred. AlphaFoldDB (805K) 0.843 0.430 0.604 0.465 48.8 71.0 99.4 73.0 86.6
    Distance Pred. AlphaFoldDB (805K) 0.839 0.448 0.616 0.464 50.9 73.5 99.4 74.6 87.5
    Angle Pred. AlphaFoldDB (805K) 0.853 0.458 0.625 0.473 56.5 76.3 99.6 77.4 86.8
    Dihedral Pred. AlphaFoldDB (805K) 0.859 0.458 0.626 0.465 51.8 77.8 99.6 75.9 87.0
    Multiview Contrast AlphaFoldDB (805K) 0.874 0.490 0.654 0.488 54.1 80.5 99.9 78.1 87.5
  7. Knowl 7 — Combining Pretrained Protein Sequence Encoders with GearNet Structure Modeling

    empirical result

    Integrating pretrained sequence representations from protein language models directly as input node features into GearNet architectures improves property prediction over sequence or structure models used in isolation.

    Replacing the default one-hot residue node features in GearNet with frozen or fine-tuned sequence embeddings from ESM-1b (pretrained on 24 million sequences) yields the hybrid ESM-1b+GearNet model. Evaluated on Enzyme Commission (EC) and Gene Ontology (GO) benchmarks using maximum F-score (Fmax⁡F_{\max}):

    • GearNet (trained from scratch): EC Fmax⁡=0.730F_{\max} = 0.730, GO-BP Fmax⁡=0.356F_{\max} = 0.356, GO-MF Fmax⁡=0.503F_{\max} = 0.503, GO-CC Fmax⁡=0.414F_{\max} = 0.414.
    • GearNet-Edge (trained from scratch): EC Fmax⁡=0.810F_{\max} = 0.810, GO-BP Fmax⁡=0.403F_{\max} = 0.403, GO-MF Fmax⁡=0.580F_{\max} = 0.580, GO-CC Fmax⁡=0.450F_{\max} = 0.450.
    • GearNet-Edge with Multiview Contrast (structure pretraining): EC Fmax⁡=0.874F_{\max} = 0.874, GO-BP Fmax⁡=0.490F_{\max} = 0.490, GO-MF Fmax⁡=0.654F_{\max} = 0.654, GO-CC Fmax⁡=0.488F_{\max} = 0.488.
    • ESM-1b alone (sequence pretraining): EC Fmax⁡=0.864F_{\max} = 0.864, GO-BP Fmax⁡=0.452F_{\max} = 0.452, GO-MF Fmax⁡=0.657F_{\max} = 0.657, GO-CC Fmax⁡=0.477F_{\max} = 0.477.
    • ESM-1b + GearNet (hybrid sequence + structure): EC Fmax⁡=0.883F_{\max} = 0.883, GO-BP Fmax⁡=0.491F_{\max} = 0.491, GO-MF Fmax⁡=0.677F_{\max} = 0.677, GO-CC Fmax⁡=0.501F_{\max} = 0.501.

    Even without structure pretraining on the hybrid backbone, ESM-1b+GearNet achieves superior performance across all four functional benchmark metrics compared to both standalone ESM-1b and GearNet-Edge with Multiview Contrast.

  8. Knowl 8 — Sensitivity Comparison of Neural Embeddings and Alignment Tools on SCOPe40 Structure Search

    data/table

    Structure search sensitivity was evaluated on the SCOPe40 benchmark (comprising 11,211 non-redundant protein chains clustered at 40%40\% sequence identity). An all-versus-all search tests the ability to retrieve proteins belonging to the same SCOPe Family, Superfamily, and Fold. Sensitivity is quantified as the area under the cumulative ROC curve up to the first false positive match.

    GearNet-Edge-IEConv representations evaluated using cosine similarity retrieval outperform standard sequence and structure alignment algorithms on average, showing a substantial advantage at the Fold level due to learned structural generalization. Ensembling neural embeddings with DALI yields state-of-the-art retrieval sensitivity across all structural hierarchy tiers.

    Method Fold Superfamily Family Average
    MMseqs2 0.001 0.082 0.542 0.208
    3D-BLAST 0.009 0.126 0.572 0.235
    CLE-SW 0.021 0.293 0.763 0.359
    CE 0.131 0.529 0.885 0.515
    TMalign-fast 0.188 0.618 0.906 0.571
    TMalign 0.188 0.610 0.901 0.566
    Foldseek 0.155 0.593 0.914 0.554
    DALI 0.310 0.751 0.942 0.667
    GearNet-Edge-IEConv 0.474 0.722 0.936 0.710
    GearNet-Edge-IEConv + DALI 0.481 0.779 0.963 0.741
  9. Knowl 9 — Ablation Analysis of Relational Message Passing, Augmentations, and Pretraining Corpora

    empirical result

    Ablation experiments isolate the contributions of specific architecture components, contrastive data augmentations, and structure pretraining corpora:

    1. Relational Graph Convolutional Layers: On GearNet-Edge (6 layers, 42M parameters, EC Fmax⁡=0.810F_{\max} = 0.810), replacing type-specific relational kernel matrices with a single shared kernel matrix degrades EC Fmax⁡F_{\max} across all layer depths tested: 6 layers (23M parameters) achieves 0.7520.752; 8 layers (39M parameters) achieves 0.7540.754; and 10 layers (60M parameters) achieves 0.7440.744.

    2. Contrastive Augmentation Schemes in Multiview Contrast: Evaluating fixed combinations of cropping and noise operations on downstream benchmarks demonstrates:

    • Subsequence Cropping + Identity: EC Fmax⁡=0.866F_{\max} = 0.866, GO-BP Fmax⁡=0.477F_{\max} = 0.477, GO-MF Fmax⁡=0.627F_{\max} = 0.627, GO-CC Fmax⁡=0.473F_{\max} = 0.473.
    • Subspace Cropping + Identity: EC Fmax⁡=0.872F_{\max} = 0.872, GO-BP Fmax⁡=0.480F_{\max} = 0.480, GO-MF Fmax⁡=0.640F_{\max} = 0.640, GO-CC Fmax⁡=0.468F_{\max} = 0.468.
    • Subsequence Cropping + Random Edge Masking: EC Fmax⁡=0.869F_{\max} = 0.869, GO-BP Fmax⁡=0.484F_{\max} = 0.484, GO-MF Fmax⁡=0.641F_{\max} = 0.641, GO-CC Fmax⁡=0.471F_{\max} = 0.471.
    • Subspace Cropping + Random Edge Masking: EC Fmax⁡=0.876F_{\max} = 0.876, GO-BP Fmax⁡=0.481F_{\max} = 0.481, GO-MF Fmax⁡=0.645F_{\max} = 0.645, GO-CC Fmax⁡=0.470F_{\max} = 0.470.
    • Random mixture over all cropping and noise choices: EC Fmax⁡=0.874F_{\max} = 0.874, GO-BP Fmax⁡=0.490F_{\max} = 0.490, GO-MF Fmax⁡=0.654F_{\max} = 0.654, GO-CC Fmax⁡=0.488F_{\max} = 0.488.
    1. Pretraining Dataset Choice: Pretraining GearNet-Edge via Multiview Contrast on experimental versus predicted structure sets yields robust downstream EC performance: AlphaFoldDB v1 (365,198 proteins) achieves AUPRpair=0.890,Fmax⁡=0.874\text{AUPR}_{\text{pair}} = 0.890, F_{\max} = 0.874; AlphaFoldDB v2 (439,674 proteins) achieves AUPRpair=0.890,Fmax⁡=0.874\text{AUPR}_{\text{pair}} = 0.890, F_{\max} = 0.874; combined AlphaFoldDB v1+v2 (804,872 proteins) achieves AUPRpair=0.892,Fmax⁡=0.874\text{AUPR}_{\text{pair}} = 0.892, F_{\max} = 0.874; and experimental PDB structures (305,265 chains, resolution ≤2.5 A˚\le 2.5\,\text{Å}) achieves AUPRpair=0.881,Fmax⁡=0.859\text{AUPR}_{\text{pair}} = 0.881, F_{\max} = 0.859.
  10. Knowl 10 — Limitations of Geometric Structure Pretraining

    limitation

    The methodology and experimental evaluation exhibit two primary limitations:

    1. Pretraining Data Scale: The pretraining experiments utilize up to 805,000 protein structures from AlphaFoldDB and PDB. While this is sufficient to achieve competitive performance against sequence models, it does not leverage the complete AlphaFold database (which contains over 100 million predicted structures), leaving the scaling behavior of geometric structure pretraining to massive datasets uncharacterized.

    2. Scope of Downstream Tasks: Evaluation is restricted to single-protein functional annotation (Enzyme Commission classification, Gene Ontology term prediction, reaction classification) and fold recognition. Downstream biological applications involving multi-entity complexes—such as protein-protein interaction prediction, interface modeling, and structure-based ligand docking/design—are not evaluated.

Coverage note — None was omitted; all key architectural components (GearNet, GearNet-Edge, GearNet-Edge-IEConv), pretraining methods (Multiview Contrast, self-prediction objectives), downstream benchmarks, hybrid sequence-structure combinations, retrieval comparisons on SCOPe40, ablation analyses, and stated limitations are covered.

References

  1. 1.Mehmet Akdel, Douglas EV Pires, Eduard Porta Pardo, Jürgen Jänes, Arthur O Zalevsky, Bálint Mészáros, Patrick Bryant, Lydia L Good, Roman A Laskowski, Gabriele Pozzati, et al. A structural biology community assessment of alphafold 2 applications. bioRxiv, 2021.
  2. 2.Ethan C Alley, Grigory Khimulya, Surojit Biswas, Mohammed AlQuraishi, and George M Church. Unified rational protein engineering with sequence-based deep representation learning. Nature methods, 16(12):1315–1322, 2019.
  3. 3.Stephen F. Altschul, Thomas L. Madden, Alejandro A. Schäffer, J Zhang, Z Zhang, Webb Miller, and David J. Lipman. Gapped blast and psi-blast: a new generation of protein database search programs. Nucleic acids research, 25 17:3389–402, 1997.
  4. 4.Sarp Aykent and Tian Xia. Gbpnet: Universal geometric representation learning on protein structures. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022.
  5. 5.Minkyung Baek, Frank DiMaio, Ivan Anishchenko, Justas Dauparas, Sergey Ovchinnikov, Gyu Rie Lee, Jue Wang, Qian Cong, Lisa N Kinch, R Dustin Schaeffer, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science, 373(6557):871–876, 2021.
  6. 6.Federico Baldassarre, David Menéndez Hurtado, Arne Elofsson, and Hossein Azizpour. Graphqa: protein model quality assessment using graph convolutional networks. Bioinformatics, 37(3): 360–366, 2021.
  7. 7.Albert P Bartók, Mike C Payne, Risi Kondor, and Gábor Csányi. Gaussian approximation potentials: The accuracy of quantum mechanics, without the electrons. Physical review letters, 104(13): 136403, 2010.
  8. 8.Albert P Bartók, Risi Kondor, and Gábor Csányi. On representing chemical environments. Physical Review B, 87(18):184115, 2013.
  9. 9.Jörg Behler and Michele Parrinello. Generalized neural-network representation of high-dimensional potential-energy surfaces. Physical review letters, 98(14):146401, 2007.
  10. 10.Tristan Bepler and Bonnie Berger. Learning protein sequence embeddings using information from structure. arXiv preprint arXiv:1902.08661, 2019.
  11. 11.Tristan Bepler and Bonnie Berger. Learning the protein language: Evolution, structure, and function. Cell Systems, 12(6):654–669, 2021.
  12. 12.Helen M Berman, John Westbrook, Zukang Feng, Gary Gilliland, Talapady N Bhat, Helge Weissig, Ilya N Shindyalov, and Philip E Bourne. The protein data bank. Nucleic acids research, 28(1): 235–242, 2000.
  13. 13.Surojit Biswas, Grigory Khimulya, Ethan C Alley, Kevin M Esvelt, and George M Church. Low-n protein engineering with data-efficient deep learning. Nature Methods, 18(4):389–396, 2021.
  14. 14.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  15. 15.Yue Cao, Payel Das, Vijil Chenthamarakshan, Pin-Yu Chen, Igor Melnyk, and Yang Shen. Fold2seq: A joint sequence(1d)-fold(3d) embedding-based generative model for protein design. In Marina Meila and Tong Zhang (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 1261–1271. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/cao21a.html.
  16. 16.Can Chen, Jingbo Zhou, Fan Wang, Xue Liu, and Dejing Dou. Structure-aware protein self-supervised learning. arXiv preprint arXiv:2204.04213, 2022.
  17. 17.Chi Chen, Weike Ye, Yunxing Zuo, Chen Zheng, and Shyue Ping Ong. Graph networks as a universal machine learning framework for molecules and crystals. Chemistry of Materials, 31(9):3564–3572, 2019.
  18. 18.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pp. 1597–1607. PMLR, 2020.
  19. 19.Stefan Chmiela, Alexandre Tkatchenko, Huziel E Sauceda, Igor Poltavsky, Kristof T Schütt, and Klaus-Robert Müller. Machine learning of accurate energy-conserving molecular force fields. Science advances, 3(5):e1603015, 2017.
  20. 20.The UniProt Consortium. Uniprot: the universal protein knowledgebase in 2021. Nucleic Acids Research, 49(D1):D480–D489, 2021.
  21. 21.Bowen Dai and Chris Bailey-Kellogg. Protein interaction interface region prediction by geometric deep learning. Bioinformatics, 2021.
  22. 22.Georgy Derevyanko, Sergei Grudinin, Yoshua Bengio, and Guillaume Lamoureux. Deep convolutional networks for quality assessment of protein folds. Bioinformatics, 34(23):4046–4053, 2018.
  23. 23.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  24. 24.Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Wang Yu, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, Debsindhu Bhowmik, and Burkhard Rost. Prottrans: Towards cracking the language of lifes code through self-supervised deep learning and high performance computing. IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–1, 2021. doi: 10.1109/TPAMI.2021.3095381.
  25. 25.Pablo Gainza, Freyr Sverrisson, Frederico Monti, Emanuele Rodola, D Boscaini, MM Bronstein, and BE Correia. Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning. Nature Methods, 17(2):184–192, 2020.
  26. 26.Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. Neural message passing for quantum chemistry. In International conference on machine learning, pp. 1263–1272. PMLR, 2017.
  27. 27.Vladimir Gligorijević, P Douglas Renfrew, Tomasz Kosciolek, Julia Koehler Leman, Daniel Berenberg, Tommi Vatanen, Chris Chandler, Bryn C Taylor, Ian M Fisk, Hera Vlamakis, et al. Structure-based protein function prediction using graph convolutional networks. Nature communications, 12 (1):1–14, 2021.
  28. 28.Yuzhi Guo, Jiaxiang Wu, Hehuan Ma, and Junzhou Huang. Self-supervised pre-training for protein embeddings using tertiary structures. In AAAI, 2022.
  29. 29.William L Hamilton, Rex Ying, and Jure Leskovec. Inductive representation learning on large graphs. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 1025–1035, 2017.
  30. 30.Frank Harary and Robert Z Norman. Some properties of line digraphs. Rendiconti del circolo matematico di palermo, 9(2):161–168, 1960.
  31. 31.Kaveh Hassani and Amir Hosein Khasahmadi. Contrastive multi-view representation learning on graphs. In International Conference on Machine Learning, pp. 4116–4126. PMLR, 2020.
  32. 32.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9729–9738, 2020.
  33. 33.Liang He, Shizhuo Zhang, Lijun Wu, Huanhuan Xia, Fusong Ju, He Zhang, Siyuan Liu, Yingce Xia, Jianwei Zhu, Pan Deng, et al. Pre-training co-evolutionary protein representation via a pairwise masked language model. arXiv preprint arXiv:2110.15527, 2021.
  34. 34.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  35. 35.Pedro Hermosilla and Timo Ropinski. Contrastive representation learning for 3d protein structures. In Submitted to The Tenth International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=VINWzIM6_6.
  36. 36.Pedro Hermosilla, Marco Schäfer, Matej Lang, Gloria Fackelmann, Pere Pau Vázquez, Barbora Kozlíková, Michael Krone, Tobias Ritschel, and Timo Ropinski. Intrinsic-extrinsic convolution and pooling for learning on 3d protein structures. International Conference on Learning Representations, 2021.
  37. 37.Liisa Holm. Benchmarking fold detection by dalilite v.5. Bioinformatics, 2019.
  38. 38.Jie Hou, Badri Adhikari, and Jianlin Cheng. Deepsf: deep convolutional neural network for mapping protein sequences to folds. Bioinformatics, 34(8):1295–1303, 2018.
  39. 39.Min Hu, Fajie Yuan, Kevin Kaichuang Yang, Fusong Ju, Jingyu Su, Hongya Wang, Fei Yang, and Qiuyang Ding. Exploring evolution-based & -free protein language models as protein function predictors. ArXiv, abs/2206.06583, 2022.
  40. 40.Weihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay Pande, and Jure Leskovec. Strategies for pre-training graph neural networks. arXiv preprint arXiv:1905.12265, 2019.
  41. 41.John Ingraham, Vikas Garg, Regina Barzilay, and Tommi Jaakkola. Generative models for graph-based protein design. Advances in Neural Information Processing Systems, 32:15820–15831, 2019.
  42. 42.Bowen Jing, Stephan Eismann, Pratham N. Soni, and Ron O. Dror. Learning from protein structure with geometric vector perceptrons. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=1YLJDvSx6J4.
  43. 43.Peter Bjørn Jørgensen, Karsten Wedel Jacobsen, and Mikkel N Schmidt. Neural message passing with edge updates for predicting properties of molecules and materials. arXiv preprint arXiv:1806.03146, 2018.
  44. 44.John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  45. 45.Thomas N Kipf and Max Welling. Variational graph auto-encoders. arXiv preprint arXiv:1611.07308, 2016.
  46. 46.Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations, 2017.
  47. 47.Johannes Klicpera, Janek Groß, and Stephan Günnemann. Directional message passing for molecular graphs. In International Conference on Learning Representations (ICLR), 2020.
  48. 48.Johannes Klicpera, Florian Becker, and Stephan Günnemann. Gemnet: Universal directional graph neural networks for molecules. arXiv preprint arXiv:2106.08903, 2021.
  49. 49.Yi Liu, Limei Wang, Meng Liu, Xuan Zhang, Bora Oztekin, and Shuiwang Ji. Spherical message passing for 3d graph networks. arXiv preprint arXiv:2102.05013, 2021.
  50. 50.Amy X Lu, Haoran Zhang, Marzyeh Ghassemi, and Alan M Moses. Self-supervised contrastive learning of protein representations by mutual information maximization. BioRxiv, 2020.
  51. 51.Bin Ma. Novor: real-time peptide de novo sequencing software. Journal of the American Society for Mass Spectrometry, 26(11):1885–1894, 2015.
  52. 52.Bin Ma and Richard Johnson. De novo sequencing and homology searching. Molecular & cellular proteomics, 11(2), 2012.
  53. 53.Craig O. Mackenzie and Gevorg Grigoryan. Protein structural motifs in prediction and design. Current opinion in structural biology, 44:161–167, 2017.
  54. 54.Leland McInnes, John Healy, and James Melville. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018.
  55. 55.Joshua Meier, Roshan Rao, Robert Verkuil, Jason Liu, Tom Sercu, and Alexander Rives. Language models enable zero-shot prediction of the effects of mutations on protein function. bioRxiv, 2021.
  56. 56.Jaina Mistry, Sara Chuguransky, Lowri Williams, Matloob Qureshi, Gustavo A Salazar, Erik LL Sonnhammer, Silvio CE Tosatto, Lisanna Paladin, Shriya Raj, Lorna J Richardson, et al. Pfam: The protein families database in 2021. Nucleic Acids Research, 49(D1):D412–D419, 2021.
  57. 57.Bhaskar Mitra and Nick Craswell. Neural models for information retrieval. ArXiv, abs/1705.01509, 2017.
  58. 58.Alexey G Murzin, Steven E Brenner, Tim Hubbard, and Cyrus Chothia. Scop: a structural classification of proteins database for the investigation of sequences and structures. Journal of molecular biology, 247(4):536–540, 1995.
  59. 59.Pascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado, Aidan N. Gomez, Debora S. Marks, and Yarin Gal. Tranception: protein fitness prediction with autoregressive transformers and inference-time retrieval. In ICML, 2022.
  60. 60.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  61. 61.Chris Paul Ponting and Robert R Russell. The natural history of protein domains. Annual review of biophysics and biomolecular structure, 31:45–71, 2002.
  62. 62.Jiezhong Qiu, Qibin Chen, Yuxiao Dong, Jing Zhang, Hongxia Yang, Ming Ding, Kuansan Wang, and Jie Tang. Gcc: Graph contrastive coding for graph neural network pre-training. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 1150–1160, 2020.
  63. 63.Predrag Radivojac, Wyatt T Clark, Tal Ronnen Oron, Alexandra M Schnoes, Tobias Wittkop, Artem Sokolov, Kiley Graim, Christopher Funk, Karin Verspoor, Asa Ben-Hur, et al. A large-scale evaluation of computational protein function prediction. Nature methods, 10(3):221–227, 2013.
  64. 64.Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Xi Chen, John Canny, Pieter Abbeel, and Yun S Song. Evaluating protein transfer learning with tape. In Advances in Neural Information Processing Systems, 2019.
  65. 65.Roshan Rao, Jason Liu, Robert Verkuil, Joshua Meier, John F Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives. Msa transformer. bioRxiv, 2021.
  66. 66.Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences, 118(15), 2021.
  67. 67.Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying Wei, Wenbing Huang, and Junzhou Huang. Self-supervised graph transformer on large-scale molecular data. arXiv preprint arXiv:2007.02835, 2020.
  68. 68.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015.
  69. 69.Victor Garcia Satorras, Emiel Hoogeboom, and Max Welling. E (n) equivariant graph neural networks. arXiv preprint arXiv:2102.09844, 2021.
  70. 70.Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne Van Den Berg, Ivan Titov, and Max Welling. Modeling relational data with graph convolutional networks. In European semantic web conference, pp. 593–607. Springer, 2018.
  71. 71.Kristof T Schütt, Farhad Arbabzadah, Stefan Chmiela, Klaus R Müller, and Alexandre Tkatchenko. Quantum-chemical insights from deep tensor neural networks. Nature communications, 8(1):1–8, 2017a.
  72. 72.Kristof T Schütt, Pieter-Jan Kindermans, Huziel E Sauceda, Stefan Chmiela, Alexandre Tkatchenko, and Klaus-Robert Müller. Schnet: A continuous-filter convolutional neural network for modeling quantum interactions. arXiv preprint arXiv:1706.08566, 2017b.
  73. 73.Amir Shanehsazzadeh, David Belanger, and David Dohan. Is transfer learning necessary for protein landscape prediction? arXiv preprint arXiv:2011.03443, 2020.
  74. 74.Ilya N. Shindyalov and Philip E. Bourne. Protein structure alignment by incremental combinatorial extension (ce) of the optimal path. Protein engineering, 11 9:739–47, 1998.
  75. 75.Vignesh Ram Somnath, Charlotte Bunne, and Andreas Krause. Multi-scale representation learning on proteins. Advances in Neural Information Processing Systems, 34, 2021.
  76. 76.Martin Steinegger and Johannes Söding. Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature Biotechnology, 35:1026–1028, 2017.
  77. 77.Martin Steinegger and Johannes Söding. Clustering huge protein sequence sets in linear time. Nature communications, 9(1):1–8, 2018.
  78. 78.Martin Steinegger, Markus Meier, Milot Mirdita, Harald Voehringer, Stephan J. Haunsberger, and Johannes Soeding. Hh-suite3 for fast remote homology detection and deep protein annotation. BMC Bioinformatics, 20, 2019.
  79. 79.Alexey Strokach, David Becerra, Carles Corbi-Verge, Albert Perez-Riba, and Philip M Kim. Fast and flexible protein design using deep graph neural networks. Cell Systems, 11(4):402–411, 2020.
  80. 80.Haitian Sun, Tania Bedrax-Weiss, and William W. Cohen. Pullnet: Open domain question answering with iterative retrieval on knowledge bases and text. EMNLP, abs/1904.09537, 2019.
  81. 81.Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In International Conference on Machine Learning, pp. 3319–3328. PMLR, 2017.
  82. 82.Baris E Suzek, Hongzhan Huang, Peter McGarvey, Raja Mazumder, and Cathy H Wu. Uniref: comprehensive and non-redundant uniprot reference clusters. Bioinformatics, 23(10):1282–1288, 2007.
  83. 83.Freyr Sverrisson, Jean Feydy, Bruno E Correia, and Michael M Bronstein. Fast end-to-end learning on protein surfaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15272–15281, 2021.
  84. 84.Y Tateno, K Ikeo, T Imanishi, H Watanabe, T Endo, Y Yamaguchi, Y Suzuki, K Takahashi, K Tsunoyama, M Kawai, et al. Evolutionary motif and its biological and structural significance. Journal of molecular evolution, 44(1):S38–S43, 1997.
  85. 85.Michel van Kempen, Stephanie Kim, Charlotte Tumescheit, Milot Mirdita, Johannes Soeding, and Martin Steinegger. Foldseek: fast and accurate protein structure search. bioRxiv, 2022.
  86. 86.Mihaly Varadi, Stephen Anyango, Mandar Deshpande, Sreenath Nair, Cindy Natassia, Galabina Yordanova, David Yuan, Oana Stroe, Gemma Wood, Agata Laydon, et al. Alphafold protein structure database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic acids research, 2021.
  87. 87.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.
  88. 88.Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph Attention Networks. International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=rJXMpikCZ. accepted as poster.
  89. 89.Limei Wang, Haoran Liu, Yi Liu, Jerry Kurtin, and Shuiwang Ji. Learning protein representations via complete 3d graph networks. ArXiv, abs/2207.12600, 2022a.
  90. 90.Yanbin Wang, Zhu-Hong You, Shan Yang, Xiao Li, Tong-Hai Jiang, and Xi Zhou. A high efficient biological language model for predicting protein–protein interactions. Cells, 8(2):122, 2019.
  91. 91.Zichen Wang, Steven A Combs, Ryan Brand, Miguel Romero Calvo, Panpan Xu, George Price, Nataliya Golovach, Emmanuel O Salawu, Colby J Wise, Sri Priya Ponnapalli, et al. Lm-gvp: an extensible sequence and structure informed deep learning framework for protein property prediction. Scientific reports, 12(1):1–12, 2022b.
  92. 92.Oren F Webb, Tommy J Phelps, Paul R Bienkowski, Philip M Digrazia, David C White, and Gary S Sayler. Enzyme nomenclature. 1992.
  93. 93.Minghao Xu, Hang Wang, Bingbing Ni, Hongyu Guo, and Jian Tang. Self-supervised graph-level representation learning with local and global structure. arXiv preprint arXiv:2106.04113, 2021.
  94. 94.Jinn-Moon Yang and Chi-Hua Tung. Protein structure database search and evolutionary classification. Nucleic Acids Research, 34:3646 – 3659, 2006.
  95. 95.Kevin Kaichuang Yang, Niccol’o Zanichelli, and Hugh Yeh. Masked inverse folding with sequence transfer for protein representation learning. bioRxiv, 2022.
  96. 96.Yuning You, Tianlong Chen, Yongduo Sui, Ting Chen, Zhangyang Wang, and Yang Shen. Graph contrastive learning with augmentations. Advances in Neural Information Processing Systems, 33: 5812–5823, 2020.
  97. 97.Yang Zhang and Jeffrey Skolnick. Tm-align: a protein structure alignment algorithm based on the tm-score. Nucleic Acids Research, 33:2302 – 2309, 2005.
  98. 98.Mengyao Zhao, Wan-Ping Lee, Erik P Garrison, and Gabor T. Marth. Ssw library: An simd smith-waterman c/c++ library for use in genomic applications. PLoS ONE, 8, 2013.
  99. 99.Zhaocheng Zhu, Chence Shi, Zuobai Zhang, Shengchao Liu, Minghao Xu, Xinyu Yuan, Yangtian Zhang, Junkun Chen, Huiyu Cai, Jiarui Lu, et al. Torchdrug: A powerful and flexible machine learning platform for drug discovery. arXiv preprint arXiv:2202.08320, 2022.

Citation

MLA
Zhang, Z., et al. “Protein Representation Learning by Geometric Structure Pretraining”. arXiv, 2022, http://arxiv.org/abs/2203.06125v5.
APA
Zhang, Z., Xu, M., Jamasb, A., Chenthamarakshan, V., Lozano, A., Das, P., & Tang, J. (2022). Protein Representation Learning by Geometric Structure Pretraining. arXiv. http://arxiv.org/abs/2203.06125v5
Chicago
Zhang, Z., M. Xu, A. Jamasb, et al. 2022. “Protein Representation Learning by Geometric Structure Pretraining”. arXiv. http://arxiv.org/abs/2203.06125v5.
Harvard
Zhang, Z. et al. (2022) “Protein Representation Learning by Geometric Structure Pretraining”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.06125v5.
Vancouver
1. Zhang Z, Xu M, Jamasb A, Chenthamarakshan V, Lozano A, Das P, Tang J (2022) Protein Representation Learning by Geometric Structure Pretraining. arXiv

BibTeX

@article{zhang2022protein,
  title = {Protein Representation Learning by Geometric Structure Pretraining},
  author = {Zhang, Zuobai and Xu, Minghao and Jamasb, Arian and Chenthamarakshan, Vijil and Lozano, Aurelie and Das, Payel and Tang, Jian},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.06125v5},
  eprint = {2203.06125}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors