Linkless Link Prediction via Relational Distillation

Zhichun GuoWilliam ShiaoShichang ZhangYozen LiuNitesh V. ChawlaNeil ShahTong Zhao

article2023ICML58 citations

Proposes a relational knowledge distillation framework that transfers topological graph knowledge from GNNs to lightweight MLPs using rank- and distribution-based matching, achieving competitive link prediction accuracy alongside a 70x inference speedup.

Listen

Modern online services rely heavily on predicting connections across vast networks, such as recommending friends on social platforms or items in e-commerce. Graph Neural Networks deliver high accuracy for these link prediction tasks by analyzing surrounding network neighborhoods. However, this neighborhood data dependency causes severe latency and computational bottlenecks, making real-time deployment difficult. In contrast, simpler multi-layer perceptron models offer negligible latency because they process data points independently, but they typically suffer from poor accuracy due to the complete lack of relational graph context.

The article evaluates whether relational knowledge distillation can effectively transfer complex structural information from a high-capacity Graph Neural Network teacher into a lightweight Multi-Layer Perceptron student. To achieve this, the article introduces Linkless Link Prediction, a training framework that distills relational topology by anchoring comparisons around individual nodes rather than matching isolated link scores or node representations. The framework combines two complementary objectives: a margin-based ranking loss to teach the student relative candidate ordering, and a distribution loss based on relative probability values to capture magnitude differences across local and global node samples. The authors evaluated the approach across eight standard benchmark datasets under standard transductive conditions, a realistic production setting with emerging nodes, and a cold-start scenario.

The results demonstrate substantial gains across accuracy and operational speed. Linkless Link Prediction achieved up to a 70.68-fold speedup in inference time over standard Graph Neural Networks on the large-scale Collab dataset, operating at under two milliseconds. In predictive accuracy, the framework outperformed standalone Multi-Layer Perceptrons by an average of 18.18 points in standard evaluations and 12.01 points in production settings. Furthermore, it matched or outperformed the teacher Graph Neural Network on seven out of eight benchmarks in standard settings. In cold-start scenarios with newly isolated nodes, the distilled student surpassed teacher Graph Neural Networks by an average of 25.29 points in Hits@20 and standalone MLPs by 9.42 points.

These findings indicate that organizations can eliminate the substantial infrastructure costs and latency bottlenecks of real-time graph traversal without sacrificing recommendation accuracy. In high-throughput industrial environments where rapid response times are mandatory, deploying distilled Multi-Layer Perceptrons removes the need for complex neighborhood aggregation during live queries. In addition, the method substantially improves cold-start handling, where traditional graph models struggle due to a lack of immediate structural connections.

Engineering and data science teams running real-time link prediction or recommendation systems should consider piloting Linkless Link Prediction as a low-latency alternative to live Graph Neural Network pipelines. Decision-makers should weigh the clear operational trade-offs: while this framework dramatically accelerates online inference and excels on existing nodes, its performance on newly appearing nodes depends directly on the richness of raw node features. Further validation should be conducted on company-specific datasets to confirm feature quality before full production rollout.

  • Paper: Link Prediction Based on Graph Neural Networks, Muhan Zhang et al. (2018). Zhang and Chen establish the standard paradigm of using graph neural networks to predict links from local enclosing subgraphs, which the source builds upon to formulate its teacher models.
  • Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). Kipf and Welling introduce foundational graph convolutional networks whose recursive neighborhood aggregation latency the source seeks to circumvent via distillation into multi-layer perceptrons.
  • Paper: Inductive Representation Learning on Large Graphs, William L. Hamilton et al. (2017). Hamilton et al. define inductive neighborhood aggregation and sampling frameworks on large graphs, establishing the classical GNN inference mechanics replaced by the source's linkless architecture.
  • Paper: Simplifying Graph Convolutional Networks, Felix Wu et al. (2019). Wu et al. explore the computational simplification of graph convolutions by decoupling relational propagation and neural transformations, motivating low-latency graph representation learning.
  • Paper: Link Prediction in Complex Networks: A Survey, Linyuan Lu et al. (2010). Lu and Zhou provide a comprehensive survey of topological heuristics and similarity indices for link prediction that underpin the evaluation benchmarks and structural baselines analyzed in the source.
  • Paper: Beyond Homophily in Graph Neural Networks: Current Limitations and Effective Designs, Jiong Zhu et al. (2020). Zhu et al. analyze the trade-offs between neighborhood aggregation and standalone multi-layer perceptrons across diverse graph environments, directly informing the source's investigation into feature-rich MLP capabilities.
Cover for Linkless Link Prediction via Relational Distillation

Abstract

Graph Neural Networks (GNNs) have shown exceptional performance in the task of link prediction. Despite their effectiveness, the high latency brought by non-trivial neighborhood dependency limits GNNs in practical deployments. Conversely, the known efficient MLPs are much less effective than GNNs due to the lack of relational knowledge. In this work, to combine the advantages of GNNs and MLPs, we start with exploring direct knowledge distillation (KD) methods for link prediction, i.e., predicted logit-based matching and node representation-based matching. Upon observing direct KD analogs do not perform well for link prediction, we propose a relational KD framework, Linkless Link Prediction (LLP), to distill knowledge for link prediction with MLPs. Unlike simple KD methods that match independent link logits or node representations, LLP distills relational knowledge that is centered around each (anchor) node to the student MLP. Specifically, we propose rank-based matching and distribution-based matching strategies that complement each other. Extensive experiments demonstrate that LLP boosts the link prediction performance of MLPs with significant margins, and even outperforms the teacher GNNs on 7 out of 8 benchmarks. LLP also achieves a 70.68× speedup in link prediction inference compared to GNNs on the large-scale OGB dataset.

Table of Contents

  • 1. Introduction
  • 2. Related Work and Preliminaries
  • 3. Cross-model Knowledge Distillation for Link Prediction
  • 3.1. Direct Methods
  • 3.2. Link Prediction with Relational Distillation
  • 3.3. Proposed Framework: Linkless Link Prediction
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Link Prediction Results
  • 4.3. Inference Acceleration Comparison
  • 4.4. Link Prediction Results on Cold Start Nodes
  • 4.5. Link Prediction Results on Large Scale Datasets
  • 4.6. Ablation Study
  • 5. Conclusion
  • Limitations
  • Ethical Impact
  • Acknowledgement
  • References
  • A. Further Related Work
  • B. Additional Datasets Details
  • C. Additional Evaluation Setting Details
  • C.1. Transductive Setting
  • C.2. Production Setting
  • C.3. Cold Start Setting
  • D. Additional Experimental Results
  • D.1. Detailed Hits@20 Results under Production Setting
  • D.2. Link Prediction Results Measured by AUC under Transductive and Production Settings
  • D.3. Performance with Different Teachers
  • D.4. Comparison among SAGE, PLNLP, and LLP
  • D.5. Analysis of Context Nodes Selection Strategy.
  • D.7. Analysis of δ in L LLP R
  • D.8. Further Comparison Results on Cold Start Nodes
  • D.9. Comparison between LLP and Other Distillation/Rank-based Matching Methods
  • E. Implementation Details

Knowls

  1. Knowl 1 — Linkless Link Prediction (LLP) Relational Distillation Framework

    model/method

    Linkless Link Prediction (LLP) is a cross-model knowledge distillation framework designed to transfer graph topological and relational knowledge from a graph neural network (GNN) teacher to a multi-layer perceptron (MLP) student for link prediction tasks.

    Given an attributed graph G=(V,E)G = (\mathcal{V}, \mathcal{E}) with node feature matrix X∈RN×FX \in \mathbb{R}^{N \times F}, a teacher GNN encoder generates node representations hi∈RDh_i \in \mathbb{R}^D and predicts link existence probabilities for node pairs (i,j)(i, j) as yi,j=σ(DECODER(hi,hj))y_{i,j} = \sigma(\text{DECODER}(h_i, h_j)), where σ\sigma is the Sigmoid function and DECODER\text{DECODER} is an MLP applied to the Hadamard product of the node representations. The student MLP encoder computes node representations h^i∈RD\hat{h}_i \in \mathbb{R}^D solely from node features xix_i (without graph message passing) and predicts link existence probabilities y^i,j=σ(DECODER(h^i,h^j))\hat{y}_{i,j} = \sigma(\text{DECODER}(\hat{h}_i, \hat{h}_j)).

    Instead of matching isolated pair-level logits or individual node embeddings, LLP distills relational knowledge centered around each anchor node v∈Vv \in \mathcal{V} across a context set of nodes Cv\mathcal{C}_v. The student MLP is trained using the joint objective:

    L=α⋅Lsup+β⋅LLLP_R+γ⋅LLLP_D\mathcal{L} = \alpha \cdot \mathcal{L}_{sup} + \beta \cdot \mathcal{L}_{LLP\_R} + \gamma \cdot \mathcal{L}_{LLP\_D}

    where Lsup\mathcal{L}_{sup} is the standard binary cross-entropy supervised link prediction loss on observed and sampled negative node pairs, LLLP_R\mathcal{L}_{LLP\_R} is an anchor-centered rank-based matching loss, LLLP_D\mathcal{L}_{LLP\_D} is an anchor-centered distribution-based matching loss, and α,β,γ≥0\alpha, \beta, \gamma \ge 0 are weighting hyperparameters.

  2. Knowl 2 — Anchor-Centered Rank-Based Distillation Loss

    equation

    The rank-based distillation objective LLLP_R\mathcal{L}_{LLP\_R} aligns the student MLP's predicted relative rankings of context nodes with respect to each anchor node against the rankings predicted by the teacher GNN:

    LLLP_R=∑v∈V∑{y^v,i,y^v,j}∈Y^vmax⁡(0,−r⋅(y^v,i−y^v,j)+δ)\mathcal{L}_{LLP\_R} = \sum_{v \in \mathcal{V}} \sum_{\{\hat{y}_{v,i}, \hat{y}_{v,j}\} \in \hat{\mathcal{Y}}_v} \max\left(0, -r \cdot (\hat{y}_{v,i} - \hat{y}_{v,j}) + \delta\right)

    where the ranking indicator rr is defined as:

    r={1,if yv,i−yv,j>δ−1,if yv,i−yv,j<−δ0,otherwiser = \begin{cases} 1, & \text{if } y_{v,i} - y_{v,j} > \delta \\ -1, & \text{if } y_{v,i} - y_{v,j} < -\delta \\ 0, & \text{otherwise} \end{cases}

    Here, v∈Vv \in \mathcal{V} denotes an anchor node, Cv\mathcal{C}_v is the set of sampled context nodes for vv, Yv={yv,i∣i∈Cv}\mathcal{Y}_v = \{y_{v,i} \mid i \in \mathcal{C}_v\} are the teacher GNN's predicted link probabilities between anchor vv and each context node ii, Y^v={y^v,i∣i∈Cv}\hat{\mathcal{Y}}_v = \{\hat{y}_{v,i} \mid i \in \mathcal{C}_v\} are the student MLP's predicted link probabilities, and δ>0\delta > 0 is a margin hyperparameter (e.g., δ=0.05\delta = 0.05). Setting r=0r = 0 when ∣yv,i−yv,j∣≤δ|y_{v,i} - y_{v,j}| \le \delta creates a tolerance band that prevents the student from fitting noise caused by minuscule probability differences in the teacher predictions.

  3. Knowl 3 — Anchor-Centered Distribution-Based Distillation Loss

    equation

    The distribution-based distillation objective LLLP_D\mathcal{L}_{LLP\_D} transfers value-level relational information by matching the soft link probability distribution over a context set Cv\mathcal{C}_v conditioned on an anchor node v∈Vv \in \mathcal{V}:

    LLLP_D=∑v∈V∑i∈Cvexp⁡(yv,i/τ)∑j∈Cvexp⁡(yv,j/τ)log⁡(exp⁡(y^v,i/τ)∑j∈Cvexp⁡(y^v,j/τ))\mathcal{L}_{LLP\_D} = \sum_{v \in \mathcal{V}} \sum_{i \in \mathcal{C}_v} \frac{\exp(y_{v,i}/\tau)}{\sum_{j \in \mathcal{C}_v} \exp(y_{v,j}/\tau)} \log \left( \frac{\exp(\hat{y}_{v,i}/\tau)}{\sum_{j \in \mathcal{C}_v} \exp(\hat{y}_{v,j}/\tau)} \right)

    where yv,i∈[0,1]y_{v,i} \in [0, 1] is the teacher GNN's predicted link probability between anchor vv and context node ii, y^v,i∈[0,1]\hat{y}_{v,i} \in [0, 1] is the student MLP's predicted link probability, and τ>0\tau > 0 is a temperature hyperparameter controlling the smoothness of the probability distribution. This objective complements rank-based matching by preserving the relative scale and difference magnitudes among context node link probabilities.

  4. Knowl 4 — Hybrid Context Node Sampling Strategy for Relational Distillation

    model/method

    To compute relational distillation losses without enumerating all NN nodes in the graph as context for each anchor node v∈Vv \in \mathcal{V}, LLP constructs a hybrid context node set Cv=CvN∪CvR\mathcal{C}_v = \mathcal{C}_v^N \cup \mathcal{C}_v^R:

    1. Local Context Set (CvN\mathcal{C}_v^N): Consists of pp nearby nodes sampled by executing repeated fixed-length random walks starting from anchor node vv. This preserves local graph topology and neighborhood structure.
    2. Global Context Set (CvR\mathcal{C}_v^R): Consists of qq nodes sampled uniformly at random from the entire node set V\mathcal{V}. This captures global graph context and relative distances to distant nodes.

    The sizes pp and qq are hyperparameters, typically selected with qq set as a multiple of pp (e.g., q/p∈{1,3,5,10,15,20}q/p \in \{1, 3, 5, 10, 15, 20\}).

  5. Knowl 5 — Direct Logit and Representation Distillation Objectives for Link Prediction

    model/method

    Direct knowledge distillation methods adapt standard instance-level distillation objectives to the link prediction setting:

    • Logit-Matching Distillation (LLM\mathcal{L}_{LM}): Trains the student MLP to match teacher link probability scores independently for each sampled node pair (i,j)(i, j):

    LLM=∑(i,j)∈E∪E−[λLsup(y^i,j,ai,j)+(1−λ)Lmatch(y^i,j,yi,j)]\mathcal{L}_{LM} = \sum_{(i,j) \in \mathcal{E} \cup \mathcal{E}^-} \left[ \lambda \mathcal{L}_{sup}(\hat{y}_{i,j}, a_{i,j}) + (1 - \lambda) \mathcal{L}_{match}(\hat{y}_{i,j}, y_{i,j}) \right]

    • Representation-Matching Distillation (LRM\mathcal{L}_{RM}): Aligns the student MLP's latent node representations h^i\hat{h}_i with the teacher GNN's latent node representations hih_i while optimizing the student decoder with supervised loss:

    LRM=∑(i,j)∈E∪E−λLsup(y^i,j,ai,j)+(1−λ)∑i∈VLmatch(h^i,hi)\mathcal{L}_{RM} = \sum_{(i,j) \in \mathcal{E} \cup \mathcal{E}^-} \lambda \mathcal{L}_{sup}(\hat{y}_{i,j}, a_{i,j}) + (1 - \lambda) \sum_{i \in \mathcal{V}} \mathcal{L}_{match}(\hat{h}_i, h_i)

    where E\mathcal{E} is the set of observed positive edges, E−\mathcal{E}^- is the set of sampled negative non-edges, ai,j∈{0,1}a_{i,j} \in \{0, 1\} is the ground-truth link label, λ∈[0,1]\lambda \in [0, 1] is a balancing weight, and Lmatch\mathcal{L}_{match} is a distance or divergence loss (e.g., MSE, KL divergence, or cosine distance).

  6. Knowl 6 — Production Evaluation Protocol for Graph Link Prediction

    experimental setup

    The production evaluation protocol benchmarks link prediction in dynamic network environments where new nodes and edges appear at inference time:

    1. Node Partition: 10%10\% of nodes in V\mathcal{V} (or 30%30\% for small datasets such as Cora and Citeseer) are randomly held out as new nodes VN\mathcal{V}^N; the remaining nodes form the existing node set VE\mathcal{V}^E.
    2. Edge Partition: Observed edges E\mathcal{E} are partitioned into three subsets based on endpoint node sets: existing–existing edges EE−E\mathcal{E}^{E-E}, existing–new edges EE−N\mathcal{E}^{E-N}, and new–new edges EN−N\mathcal{E}^{N-N}.
    3. Split Proportions:
      • EE−E\mathcal{E}^{E-E} is split 80/10/1080/10/10 into training message-passing edges, newly visible inference message-passing edges, and test evaluation edges.
      • EE−N\mathcal{E}^{E-N} and EN−N\mathcal{E}^{N-N} are each split 90/1090/10 into newly visible inference message-passing edges and test evaluation edges.
    4. Inference Visibility: During training, GNNs access only the 80%80\% training portion of EE−E\mathcal{E}^{E-E}. During inference, GNNs perform message passing over all non-test edges (90%90\% of all edges across all three partitions). Evaluation is reported as Hits@20 overall and stratified across EE−E\mathcal{E}^{E-E}, EE−N\mathcal{E}^{E-N}, and EN−N\mathcal{E}^{N-N}.
  7. Knowl 7 — Link Prediction Performance Under Transductive Evaluation

    data/table

    Under the transductive evaluation setting, LLP outperforms stand-alone MLPs and direct distillation baselines across all 8 benchmark datasets, and matches or outperforms the teacher GNN (GraphSAGE) on 7 out of 8 datasets. Metrics reported are Hits@50 for Collab and Hits@20 for all other benchmarks (mean ±\pm standard deviation over 10 runs):

    Dataset GNN (Teacher) MLP LLM\mathcal{L}_{LM} LRM\mathcal{L}_{RM} LLP (Ours)
    Cora 74.38±\pm1.54 78.06±\pm1.50 74.72±\pm4.27 75.75±\pm1.51 78.82±\pm1.74
    Citeseer 73.89±\pm0.95 71.21±\pm3.22 72.44±\pm1.52 65.19±\pm5.54 77.32±\pm2.42
    Pubmed 51.98±\pm5.25 42.89±\pm1.67 42.78±\pm3.15 44.44±\pm2.40 57.33±\pm2.42
    CS 59.51±\pm7.34 34.01±\pm9.37 40.69±\pm5.12 61.10±\pm2.83 68.62±\pm1.46
    Physics 66.74±\pm1.53 31.26±\pm9.12 52.11±\pm2.44 52.34±\pm3.78 72.01±\pm1.89
    Computers 31.66±\pm3.08 20.19±\pm1.58 12.81±\pm1.80 21.75±\pm1.96 35.32±\pm2.28
    Photos 51.50±\pm4.48 27.83±\pm4.90 24.24±\pm2.79 38.47±\pm2.76 49.32±\pm2.64
    Collab 48.69±\pm0.87 36.95±\pm1.37 35.97±\pm0.96 36.86±\pm0.45 49.10±\pm0.57

    On average, LLP improves performance over stand-alone MLP by 18.18 points and over direct KD methods by 10.59 points.

  8. Knowl 8 — Link Prediction Performance in Production and Cold-Start Settings

    empirical result

    In the production setting evaluated with Hits@20 across seven benchmarks, LLP achieves an average gain of 12.01 points over stand-alone MLPs and 6.67 points over direct distillation baselines:

    • Cora: GNN 27.80, MLP 22.90, LLP 27.87
    • Citeseer: GNN 38.78, MLP 31.21, LLP 34.75
    • Pubmed: GNN 52.71, MLP 38.01, LLP 53.48
    • CS: GNN 60.69, MLP 38.15, LLP 60.74
    • Physics: GNN 55.82, MLP 29.99, LLP 52.83
    • Computers: GNN 34.38, MLP 19.43, LLP 24.58
    • Photos: GNN 51.03, MLP 34.29, LLP 43.79

    In the strict cold-start setting (where new nodes have zero observable message-passing edges at test time), teacher GNNs experience catastrophic performance degradation due to inability to aggregate neighbors (e.g., GNN scores 6.39 on Cora, 4.63 on Pubmed, 0.87 on Photos). In contrast, LLP retains high accuracy without graph dependencies, scoring 22.01 on Cora, 37.68 on Pubmed, and 23.79 on Photos, outperforming teacher GNNs by an average of 25.29 Hits@20 and stand-alone MLPs by 9.42 Hits@20.

  9. Knowl 9 — Inference Latency and Speedup of LLP Versus GNN Acceleration Baselines

    empirical result

    On the large-scale OGB Collab dataset, LLP eliminates neighborhood dependency during inference, achieving major latency reductions over GNNs and common GNN acceleration methods while preserving predictive performance:

    • Vanilla GraphSAGE: Inference time 134.3 ms, Hits@50 = 48.69
    • Quantized SAGE (QSAGE, float32 to int8): Inference time 128.3 ms, Hits@50 = 45.36
    • Pruned SAGE (PSAGE, 50% weights pruned): Inference time 128.7 ms, Hits@50 = 48.34
    • SAGE with Neighbor Sampling (fan-out 15): Inference time 28.6 ms, Hits@50 = 31.50
    • Stand-alone MLP: Inference time 1.9 ms, Hits@50 = 36.95
    • LLP (Student MLP): Inference time 1.9 ms, Hits@50 = 49.10

    LLP achieves a 70.68×70.68\times inference speedup compared to GraphSAGE and a 15.05×15.05\times speedup compared to Neighbor Sampling. When using an enlarged student MLP (hidden dimension 4 times larger than the teacher), inference latency is 7.1 ms, maintaining an 18.9×18.9\times speedup over GraphSAGE while achieving 51.14 Hits@50 when trained with a PLNLP teacher.

  10. Knowl 10 — Criticality of Ranking Margin Delta and Complementarity of Distillation Losses

    empirical result

    Empirical ablations demonstrate the necessity of the margin parameter δ\delta in the rank-based distillation loss and the mutual complementarity of the ranking and distribution objectives:

    • Necessity of Margin δ\delta: Setting δ=0\delta = 0 in LLLP_R\mathcal{L}_{LLP\_R} results in complete model collapse (Hits@20 drops to 0.00 across all benchmarks: Cora, Citeseer, Pubmed, CS, Physics, Computers, and Photos). Setting δ>0\delta > 0 filters out noisy rank comparisons on teacher probability pairs that are indistinguishable (∣yv,i−yv,j∣≤δ|y_{v,i} - y_{v,j}| \le \delta).
    • Comparison with ListNet: LLLP_R\mathcal{L}_{LLP\_R} consistently outperforms standard listwise ranking loss (ListNet) across all datasets (e.g., CS: 68.30 vs 43.06 Hits@20; Physics: 60.28 vs 28.87 Hits@20; Pubmed: 53.30 vs 18.79 Hits@20), because ListNet enforces strict total orders that overfit arbitrary score variations among same-hop neighbors.
    • Loss Complementarity: On Pubmed (transductive), full LLP achieves 57.33 Hits@20. Removing LLLP_R\mathcal{L}_{LLP\_R} drops performance to 55.35; removing LLLP_D\mathcal{L}_{LLP\_D} drops performance to 54.97. Both components independently outperform stand-alone MLP (42.89) and teacher GNN (51.98).
  11. Knowl 11 — Sensitivity to Node Feature Quality in MLP-Based Distillation

    limitation

    Because the student model in LLP is an MLP that does not ingest graph structure during inference, its link prediction ability relies heavily on the quality and expressiveness of raw input node features XX. When node features are insufficiently informative, unaligned with graph connectivity, or missing, LLP cannot recover the structural inductive bias of GNNs. For example, on the Citation2 benchmark, teacher GNN achieves 82.56 Hits@200 whereas LLP achieves 53.20 Hits@200, reflecting a performance gap on feature-constrained distributions.

Coverage note — No substantial contributed material was omitted. Detailed parameter search grids across every individual dataset and secondary hardware-level configuration details have been condensed into the relevant methodological and empirical knowls.

References

  1. 1.Adamic, L. A. and Adar, E. Friends and neighbors on the web. Social networks, 2003.
  2. 2.Allen-Zhu, Z. and Li, Y. Towards understanding ensemble, knowledge distillation and self-distillation in deep learning. arXiv preprint arXiv:2012.09816, 2020.
  3. 3.Alon, U. and Yahav, E. On the bottleneck of graph neural networks and its practical implications. arXiv preprint arXiv:2006.05205, 2020.
  4. 4.Barabási, A.-L. and Albert, R. Emergence of scaling in random networks. science, 1999.
  5. 5.Berg, R. v. d., Kipf, T. N., and Welling, M. Graph convolutional matrix completion. arXiv preprint arXiv:1706.02263, 2017.
  6. 6.Bevilacqua, B., Frasca, F., Lim, D., Srinivasan, B., Cai, C., Balamurugan, G., Bronstein, M. M., and Maron, H. Equivariant subgraph aggregation networks. arXiv preprint arXiv:2110.02910, 2021.
  7. 7.Bojchevski, A. and Günnemann, S. Deep gaussian embedding of graphs: Unsupervised inductive learning via ranking. arXiv preprint arXiv:1707.03815, 2017.
  8. 8.Brin, S. and Page, L. Reprint of: The anatomy of a largescale hypertextual web search engine. Computer networks, 2012.
  9. 9.Cai, L. and Ji, S. A multi-scale approach for graph link prediction. In Proceedings of the AAAI conference on artificial intelligence, 2020.
  10. 10.Cai, L., Li, J., Wang, J., and Ji, S. Line graph neural networks for link prediction. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  11. 11.Cao, Z., Qin, T., Liu, T.-Y., Tsai, M.-F., and Li, H. Learning to rank: from pairwise approach to listwise approach. In Proceedings of the 24th international conference on Machine learning, pp. 129–136, 2007.
  12. 12.Chami, I., Ying, Z., Ré, C., and Leskovec, J. Hyperbolic graph convolutional neural networks. Advances in neural information processing systems, 2019.
  13. 13.Chen, D., Mei, J.-P., Zhang, Y., Wang, C., Wang, Z., Feng, Y., and Chen, C. Cross-layer distillation with semantic calibration. In Proceedings of the AAAI Conference on Artificial Intelligence, 2021a.
  14. 14.Chen, J., He, H., Wu, F., and Wang, J. Topology-aware correlations between relations for inductive link prediction in knowledge graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, 2021b.
  15. 15.Chen, M., Wei, Z., Huang, Z., Ding, B., and Li, Y. Simple and deep graph convolutional networks. In International Conference on Machine Learning, 2020.
  16. 16.Chen, T., Sui, Y., Chen, X., Zhang, A., and Wang, Z. A unified lottery ticket hypothesis for graph neural networks. In International Conference on Machine Learning, pp. 1695–1706. PMLR, 2021c.
  17. 17.Covington, P., Adams, J., and Sargin, E. Deep neural networks for youtube recommendations. In Proceedings of the 10th ACM conference on recommender systems, pp. 191–198, 2016.
  18. 18.Davidson, T. R., Falorsi, L., De Cao, N., Kipf, T., and Tomczak, J. M. Hyperspherical variational auto-encoders. arXiv preprint arXiv:1804.00891, 2018.
  19. 19.Deng, X. and Zhang, Z. Graph-free knowledge distillation for graph neural networks. arXiv preprint arXiv:2105.07519, 2021.
  20. 20.Ding, H., Ma, Y., Deoras, A., Wang, Y., and Wang, H. Zero-shot recommender systems. arXiv preprint arXiv:2105.08318, 2021.
  21. 21.Fan, W., Liu, X., Jin, W., Zhao, X., Tang, J., and Li, Q. Graph trend filtering networks for recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 112–121, 2022.
  22. 22.Fey, M. and Lenssen, J. E. Fast graph representation learning with pytorch geometric. arXiv preprint arXiv:1903.02428, 2019.
  23. 23.Fey, M., Lenssen, J. E., Weichert, F., and Leskovec, J. Gnnautoscale: Scalable and expressive graph neural networks via historical embeddings. In International Conference on Machine Learning, pp. 3294–3304. PMLR, 2021.
  24. 24.Furlanello, T., Lipton, Z., Tschannen, M., Itti, L., and Anandkumar, A. Born again neural networks. In International Conference on Machine Learning, pp. 1607–1616. PMLR, 2018.
  25. 25.Geerts, F. and Reutter, J. L. Expressiveness and approximation properties of graph neural networks. arXiv preprint arXiv:2204.04661, 2022.
  26. 26.Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K. A survey of quantization methods for efficient neural network inference. arXiv preprint arXiv:2103.13630, 2021.
  27. 27.Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., and Dahl, G. E. Neural message passing for quantum chemistry. In International conference on machine learning, pp. 1263–1272. PMLR, 2017.
  28. 28.Gou, J., Yu, B., Maybank, S. J., and Tao, D. Knowledge distillation: A survey. International Journal of Computer Vision, 2021.
  29. 29.Guo, Z., Zhang, C., Yu, W., Herr, J., Wiest, O., Jiang, M., and Chawla, N. V. Few-shot graph learning for molecular property prediction. In WWW, 2021.
  30. 30.Guo, Z., Zhang, C., Fan, Y., Tian, Y., Zhang, C., and Chawla, N. Boosting graph neural networks via adaptive knowledge distillation. arXiv preprint arXiv:2210.05920, 2022.
  31. 31.Hamilton, W., Ying, Z., and Leskovec, J. Inductive representation learning on large graphs. Advances in neural information processing systems, 2017.
  32. 32.Hao, Y., Cao, X., Fang, Y., Xie, X., and Wang, S. Inductive link prediction for nodes having only attribute information. arXiv preprint arXiv:2007.08053, 2020.
  33. 33.He, X., Deng, K., Wang, X., Li, Y., Zhang, Y., and Wang, M. Lightgcn: Simplifying and powering graph convolution network for recommendation. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, pp. 639–648, 2020.
  34. 34.Hinton, G., Vinyals, O., Dean, J., et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.
  35. 35.Hu, Y., You, H., Wang, Z., Wang, Z., Zhou, E., and Gao, Y. Graph-mlp: node classification without message passing in graph. arXiv preprint arXiv:2106.04051, 2021.
  36. 36.Huang, Z., Li, X., and Chen, H. Link prediction approach to collaborative filtering. In Proceedings of the 5th ACM/IEEE-CS joint conference on Digital libraries, pp. 141–142, 2005.
  37. 37.Jeh, G. and Widom, J. Simrank: a measure of structuralcontext similarity. In Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining, 2002.
  38. 38.Ji, G. and Zhu, Z. Knowledge distillation in wide neural networks: Risk bound, data efficiency and imperfect teacher. Advances in Neural Information Processing Systems, 2020.
  39. 39.Jia, Z., Lin, S., Ying, R., You, J., Leskovec, J., and Aiken, A. Redundancy-free computation for graph neural networks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2020.
  40. 40.Joshi, C. K., Liu, F., Xun, X., Lin, J., and Foo, C.-S. On representation knowledge distillation for graph neural networks. arXiv preprint arXiv:2111.04964, 2021.
  41. 41.Ju, M., Zhao, T., Wen, Q., Yu, W., Shah, N., Ye, Y., and Zhang, C. Multi-task self-supervised graph neural networks enable stronger task generalization. International Conference on Learning Representations, 2023.
  42. 42.Kang, S., Hwang, J., Kweon, W., and Yu, H. Topology distillation for recommender system. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp. 829–839, 2021.
  43. 43.Khatua, A., Mailthody, V. S., Taleka, B., Ma, T., Song, X., and Hwu, W.-m. Igb: Addressing the gaps in labeling, features, heterogeneity, and size of public graph datasets for deep learning research. In In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2023.
  44. 44.Kim, J., Park, S., and Kwak, N. Paraphrasing complex network: Network compression via factor transfer. Advances in neural information processing systems, 2018.
  45. 45.Kipf, T. N. and Welling, M. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016a.
  46. 46.Kipf, T. N. and Welling, M. Variational graph auto-encoders. arXiv preprint arXiv:1611.07308, 2016b.
  47. 47.Koren, Y., Bell, R., and Volinsky, C. Matrix factorization techniques for recommender systems. Computer, 2009.
  48. 48.Leskovec, J. and Faloutsos, C. Sampling from large graphs. In Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining, 2006.
  49. 49.Li, J., Jing, M., Lu, K., Zhu, L., Yang, Y., and Huang, Z. From zero-shot learning to cold-start recommendation. In Proceedings of the AAAI conference on artificial intelligence, 2019.
  50. 50.Li, P., Wang, Y., Wang, H., and Leskovec, J. Distance encoding: Design provably more powerful neural networks for graph representation learning. Advances in Neural Information Processing Systems, 33:4465–4478, 2020.
  51. 51.Liu, G., Zhao, T., Xu, J., Luo, T., and Jiang, M. Graph rationalization with environment-based augmentations. In Proceedings of the 28th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2022.
  52. 52.Liu, X., Jin, W., Ma, Y., Li, Y., Liu, H., Wang, Y., Yan, M., and Tang, J. Elastic graph neural networks. In International Conference on Machine Learning, pp. 6837–6849. PMLR, 2021.
  53. 53.Ma, J. and Mei, Q. Graph representation learning via multi-task knowledge distillation. arXiv preprint arXiv:1911.05700, 2019.
  54. 54.Ma, Y., Liu, X., Zhao, T., Liu, Y., Tang, J., and Shah, N. A unified view on graph neural networks as graph signal denoising. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, 2021.
  55. 55.Maron, H., Ben-Hamu, H., Serviansky, H., and Lipman, Y. Provably powerful graph networks. Advances in neural information processing systems, 32, 2019.
  56. 56.Martínez, V., Berzal, F., and Cubero, J.-C. A survey of link prediction in complex networks. ACM computing surveys (CSUR), 49(4):1–33, 2016.
  57. 57.McAuley, J., Targett, C., Shi, Q., and Van Den Hengel, A. Image-based recommendations on styles and substitutes. In Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval, 2015.
  58. 58.Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. Distributed representations of words and phrases and their compositionality. Advances in neural information processing systems, 2013.
  59. 59.Nathani, D., Chauhan, J., Sharma, C., and Kaul, M. Learning attention-based embeddings for relation prediction in knowledge graphs. arXiv preprint arXiv:1906.01195, 2019.
  60. 60.Park, W., Kim, D., Lu, Y., and Cho, M. Relational knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019.
  61. 61.Perozzi, B., Al-Rfou, R., and Skiena, S. Deepwalk: Online learning of social representations. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, 2014.
  62. 62.Philip, S. Y., Han, J., and Faloutsos, C. Link mining: Models, algorithms, and applications. Springer, 2010.
  63. 63.Phuong, M. and Lampert, C. Towards understanding knowledge distillation. In International Conference on Machine Learning. PMLR, 2019.
  64. 64.Reddi, S., Pasumarthi, R. K., Menon, A., Rawat, A. S., Yu, F., Kim, S., Veit, A., and Kumar, S. Rankdistil: Knowledge distillation for ranking. In International Conference on Artificial Intelligence and Statistics. PMLR, 2021.
  65. 65.Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y. Fitnets: Hints for thin deep nets. arXiv preprint arXiv:1412.6550, 2014.
  66. 66.Sankar, A., Liu, Y., Yu, J., and Shah, N. Graph neural networks for friend ranking in large-scale social platforms. In Proceedings of the Web Conference, pp. 2535–2546, 2021.
  67. 67.Schlichtkrull, M., Kipf, T. N., Bloem, P., Berg, R. v. d., Titov, I., and Welling, M. Modeling relational data with graph convolutional networks. In European semantic web conference, pp. 593–607. Springer, 2018.
  68. 68.Shchur, O., Mumme, M., Bojchevski, A., and Günnemann, S. Pitfalls of graph neural network evaluation. arXiv preprint arXiv:1811.05868, 2018.
  69. 69.Shiao, W. and Papalexakis, E. E. Adversarially generating rank-constrained graphs. In 2021 IEEE 8th International Conference on Data Science and Advanced Analytics (DSAA). IEEE, 2021.
  70. 70.Shiao, W., Guo, Z., Zhao, T., Papalexakis, E. E., Liu, Y., and Shah, N. Link prediction with non-contrastive learning. In International Conference on Learning Representations, 2023.
  71. 71.Tailor, S. A., Fernandez-Marques, J., and Lane, N. D. Degree-quant: Quantization-aware training for graph neural networks. arXiv preprint arXiv:2008.05000, 2020.
  72. 72.Tang, X., Liu, Y., He, X., Wang, S., and Shah, N. Friend story ranking with edge-contextual local graph convolutions. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, pp. 1007–1015, 2022.
  73. 73.Tian, Y., Zhang, C., Guo, Z., Zhang, X., and Chawla, N. V. Nosmog: Learning noise-robust and structure-aware mlps on graphs. arXiv preprint arXiv:2208.10010, 2022.
  74. 74.Trouillon, T., Welbl, J., Riedel, S., Gaussier, É., and Bouchard, G. Complex embeddings for simple link prediction. In International conference on machine learning, pp. 2071–2080. PMLR, 2016.
  75. 75.Tsitsulin, A., Mottin, D., Karras, P., and Müller, E. Verse: Versatile graph embeddings from similarity measures. In Proceedings of the 2018 world wide web conference, pp. 539–548, 2018.
  76. 76.Tung, F. and Mori, G. Similarity-preserving knowledge distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019.
  77. 77.Vashishth, S., Sanyal, S., Nitin, V., and Talukdar, P. Composition-based multi-relational graph convolutional networks. arXiv preprint arXiv:1911.03082, 2020.
  78. 78.Veličković, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017.
  79. 79.Wang, K., Shen, Z., Huang, C., Wu, C.-H., Dong, Y., and Kanakia, A. Microsoft academic graph: When experts are not enough. Quantitative Science Studies, 2020a.
  80. 80.Wang, X., Fu, T., Liao, S., Wang, S., Lei, Z., and Mei, T. Exclusivity-consistency regularized knowledge distillation for face recognition. In European Conference on Computer Vision, 2020b.
  81. 81.Wang, Z., Zhou, Y., Hong, L., Zou, Y., and Su, H. Pairwise learning for neural link prediction. arXiv preprint arXiv:2112.02936, 2021.
  82. 82.Xu, K., Li, C., Tian, Y., Sonobe, T., Kawarabayashi, K.-i., and Jegelka, S. Representation learning on graphs with jumping knowledge networks. In International conference on machine learning, pp. 5453–5462. PMLR, 2018.
  83. 83.Yan, B., Wang, C., Guo, G., and Lou, Y. Tinygnn: Learning efficient graph neural networks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2020.
  84. 84.Yang, C., Liu, J., and Shi, C. Extract the knowledge of graph neural networks and go beyond it: An effective knowledge distillation framework. In Proceedings of the Web Conference 2021, 2021.
  85. 85.Yang, Y., Qiu, J., Song, M., Tao, D., and Wang, X. Distilling knowledge from graph convolutional networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020a.
  86. 86.Yang, Z., Cohen, W., and Salakhudinov, R. Revisiting semi-supervised learning with graph embeddings. In International conference on machine learning, 2016.
  87. 87.Yang, Z., Shou, L., Gong, M., Lin, W., and Jiang, D. Model compression with two-stage multi-teacher knowledge distillation for web question answering system. In Proceedings of the 13th International Conference on Web Search and Data Mining, 2020b.
  88. 88.Yin, H., Zhang, M., Wang, Y., Wang, J., and Li, P. Algorithm and system co-design for efficient subgraph-based graph representation learning. arXiv preprint arXiv:2202.13538, 2022.
  89. 89.Ying, R., He, R., Chen, K., Eksombatchai, P., Hamilton, W. L., and Leskovec, J. Graph convolutional neural networks for web-scale recommender systems. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pp. 974–983, 2018a.
  90. 90.Ying, Z., You, J., Morris, C., Ren, X., Hamilton, W., and Leskovec, J. Hierarchical graph representation learning with differentiable pooling. Advances in neural information processing systems, 2018b.
  91. 91.You, J., Ying, R., Ren, X., Hamilton, W., and Leskovec, J. Graphrnn: Generating realistic graphs with deep autoregressive models. In International conference on machine learning. PMLR, 2018.
  92. 92.Yun, S., Kim, S., Lee, J., Kang, J., and Kim, H. J. Neo-gnns: Neighborhood overlap-aware graph neural networks for link prediction. Advances in Neural Information Processing Systems, 2021.
  93. 93.Zeng, H., Zhou, H., Srivastava, A., Kannan, R., and Prasanna, V. Graphsaint: Graph sampling based inductive learning method. arXiv preprint arXiv:1907.04931, 2019.
  94. 94.Zhang, C., Yao, H., Huang, C., Jiang, M., Li, Z., and Chawla, N. V. Few-shot knowledge graph completion. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020a.
  95. 95.Zhang, D., Huang, X., Liu, Z., Hu, Z., Song, X., Ge, Z., Zhang, Z., Wang, L., Zhou, J., Shuang, Y., et al. Agl: a scalable system for industrial-purpose graph machine learning. arXiv preprint arXiv:2003.02454, 2020b.
  96. 96.Zhang, H., Lin, S., Liu, W., Zhou, P., Tang, J., Liang, X., and Xing, E. P. Iterative graph self-distillation. arXiv preprint arXiv:2010.12609, 2020c.
  97. 97.Zhang, M. and Chen, Y. Weisfeiler-lehman neural machine for link prediction. In Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining, pp. 575–583, 2017.
  98. 98.Zhang, M. and Chen, Y. Link prediction based on graph neural networks. Advances in neural information processing systems, 2018.
  99. 99.Zhang, M., Cui, Z., Neumann, M., and Chen, Y. An end-to-end deep learning architecture for graph classification. In Proceedings of the AAAI conference on artificial intelligence, 2018.
  100. 100.Zhang, M., Li, P., Xia, Y., Wang, K., and Jin, L. Labeling trick: A theory of using graph neural networks for multi-node representation learning. Advances in Neural Information Processing Systems, 34:9061–9073, 2021a.
  101. 101.Zhang, S., Liu, Y., Sun, Y., and Shah, N. Graph-less neural networks: Teaching old mlps new tricks via distillation. arXiv preprint arXiv:2110.08727, 2021b.
  102. 102.Zhang, W., Miao, X., Shao, Y., Jiang, J., Chen, L., Ruas, O., and Cui, B. Reliable data distillation on graph convolutional network. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, 2020d.
  103. 103.Zhang, Z., Cui, P., and Zhu, W. Deep learning on graphs: A survey. IEEE Transactions on Knowledge and Data Engineering, 2020e.
  104. 104.Zhao, L., Jin, W., Akoglu, L., and Shah, N. From stars to subgraphs: Uplifting any gnn with local structure awareness. arXiv preprint arXiv:2110.03753, 2021a.
  105. 105.Zhao, T., Liu, Y., Neves, L., Woodford, O., Jiang, M., and Shah, N. Data augmentation for graph neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, 2021b.
  106. 106.Zhao, T., Jin, W., Liu, Y., Wang, Y., Liu, G., Günnemann, S., Shah, N., and Jiang, M. Graph data augmentation for graph machine learning: A survey. arXiv preprint arXiv:2202.08871, 2022a.
  107. 107.Zhao, T., Liu, G., Wang, D., Yu, W., and Jiang, M. Learning from counterfactual links for link prediction. In International Conference on Machine Learning, pp. 26911–26926. PMLR, 2022b.
  108. 108.Zhao, T., Tang, X., Zhang, D., Jiang, H., Rao, N., Song, Y., Agrawal, P., Subbian, K., Yin, B., and Jiang, M. Autogda: Automated graph data augmentation for node classification. In The First Learning on Graphs Conference, 2022c.
  109. 109.Zhao, Y., Wang, D., Bates, D., Mullins, R., Jamnik, M., and Lio, P. Learned low precision graph neural networks. arXiv preprint arXiv:2009.09232, 2020.
  110. 110.Zheng, W., Huang, E. W., Rao, N., Katariya, S., Wang, Z., and Subbian, K. Cold brew: Distilling graph node representations with incomplete or missing neighborhoods. arXiv preprint arXiv:2111.04840, 2021.
  111. 111.Zhou, H., Srivastava, A., Zeng, H., Kannan, R., and Prasanna, V. Accelerating large scale real-time gnn inference using channel pruning. arXiv preprint arXiv:2105.04528, 2021.
  112. 112.Zhu, Z., Zhang, Z., Xhonneux, L.-P., and Tang, J. Neural bellman-ford networks: A general graph neural network framework for link prediction. Advances in Neural Information Processing Systems, 2021.

Citation

MLA
Guo, Z., et al. “Linkless Link Prediction via Relational Distillation”. International Conference on Machine Learning, vol. 202, 2023, pp. 12012–33, https://proceedings.mlr.press/v202/guo23f.html.
APA
Guo, Z., Shiao, W., Zhang, S., Liu, Y., Chawla, N. V., Shah, N., & Zhao, T. (2023). Linkless Link Prediction via Relational Distillation. International Conference on Machine Learning, 202, 12012–12033. https://proceedings.mlr.press/v202/guo23f.html
Chicago
Guo, Z., W. Shiao, S. Zhang, et al. 2023. “Linkless Link Prediction via Relational Distillation”. International Conference on Machine Learning 202: 12012–33. https://proceedings.mlr.press/v202/guo23f.html.
Harvard
Guo, Z. et al. (2023) “Linkless Link Prediction via Relational Distillation”, International Conference on Machine Learning. PMLR, pp. 12012–12033. Available at: https://proceedings.mlr.press/v202/guo23f.html.
Vancouver
1. Guo Z, Shiao W, Zhang S, Liu Y, Chawla NV, Shah N, Zhao T (2023) Linkless Link Prediction via Relational Distillation. In: International Conference on Machine Learning. PMLR, pp 12012–12033

BibTeX

@InProceedings{pmlr-v202-guo23f,
  title = 	 {Linkless Link Prediction via Relational Distillation},
  author =       {Guo, Zhichun and Shiao, William and Zhang, Shichang and Liu, Yozen and Chawla, Nitesh V and Shah, Neil and Zhao, Tong},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {12012--12033},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/guo23f/guo23f.pdf},
  url = 	 {https://proceedings.mlr.press/v202/guo23f.html},
  abstract = 	 {Graph Neural Networks (GNNs) have shown exceptional performance in the task of link prediction. Despite their effectiveness, the high latency brought by non-trivial neighborhood data dependency limits GNNs in practical deployments. Conversely, the known efficient MLPs are much less effective than GNNs due to the lack of relational knowledge. In this work, to combine the advantages of GNNs and MLPs, we start with exploring direct knowledge distillation (KD) methods for link prediction, i.e., predicted logit-based matching and node representation-based matching. Upon observing direct KD analogs do not perform well for link prediction, we propose a relational KD framework, Linkless Link Prediction (LLP), to distill knowledge for link prediction with MLPs. Unlike simple KD methods that match independent link logits or node representations, LLP distills relational knowledge that is centered around each (anchor) node to the student MLP. Specifically, we propose rank-based matching and distribution-based matching strategies that complement each other. Extensive experiments demonstrate that LLP boosts the link prediction performance of MLPs with significant margins and even outperforms the teacher GNNs on 7 out of 8 benchmarks. LLP also achieves a 70.68x speedup in link prediction inference compared to GNNs on the large-scale OGB dataset.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/