Linkless Link Prediction via Relational Distillation
Zhichun GuoWilliam ShiaoShichang ZhangYozen LiuNitesh V. ChawlaNeil ShahTong Zhao
Proposes a relational knowledge distillation framework that transfers topological graph knowledge from GNNs to lightweight MLPs using rank- and distribution-based matching, achieving competitive link prediction accuracy alongside a 70x inference speedup.
Modern online services rely heavily on predicting connections across vast networks, such as recommending friends on social platforms or items in e-commerce. Graph Neural Networks deliver high accuracy for these link prediction tasks by analyzing surrounding network neighborhoods. However, this neighborhood data dependency causes severe latency and computational bottlenecks, making real-time deployment difficult. In contrast, simpler multi-layer perceptron models offer negligible latency because they process data points independently, but they typically suffer from poor accuracy due to the complete lack of relational graph context.
The article evaluates whether relational knowledge distillation can effectively transfer complex structural information from a high-capacity Graph Neural Network teacher into a lightweight Multi-Layer Perceptron student. To achieve this, the article introduces Linkless Link Prediction, a training framework that distills relational topology by anchoring comparisons around individual nodes rather than matching isolated link scores or node representations. The framework combines two complementary objectives: a margin-based ranking loss to teach the student relative candidate ordering, and a distribution loss based on relative probability values to capture magnitude differences across local and global node samples. The authors evaluated the approach across eight standard benchmark datasets under standard transductive conditions, a realistic production setting with emerging nodes, and a cold-start scenario.
The results demonstrate substantial gains across accuracy and operational speed. Linkless Link Prediction achieved up to a 70.68-fold speedup in inference time over standard Graph Neural Networks on the large-scale Collab dataset, operating at under two milliseconds. In predictive accuracy, the framework outperformed standalone Multi-Layer Perceptrons by an average of 18.18 points in standard evaluations and 12.01 points in production settings. Furthermore, it matched or outperformed the teacher Graph Neural Network on seven out of eight benchmarks in standard settings. In cold-start scenarios with newly isolated nodes, the distilled student surpassed teacher Graph Neural Networks by an average of 25.29 points in Hits@20 and standalone MLPs by 9.42 points.
These findings indicate that organizations can eliminate the substantial infrastructure costs and latency bottlenecks of real-time graph traversal without sacrificing recommendation accuracy. In high-throughput industrial environments where rapid response times are mandatory, deploying distilled Multi-Layer Perceptrons removes the need for complex neighborhood aggregation during live queries. In addition, the method substantially improves cold-start handling, where traditional graph models struggle due to a lack of immediate structural connections.
Engineering and data science teams running real-time link prediction or recommendation systems should consider piloting Linkless Link Prediction as a low-latency alternative to live Graph Neural Network pipelines. Decision-makers should weigh the clear operational trade-offs: while this framework dramatically accelerates online inference and excels on existing nodes, its performance on newly appearing nodes depends directly on the richness of raw node features. Further validation should be conducted on company-specific datasets to confirm feature quality before full production rollout.
- Paper: Link Prediction Based on Graph Neural Networks, Muhan Zhang et al. (2018). Zhang and Chen establish the standard paradigm of using graph neural networks to predict links from local enclosing subgraphs, which the source builds upon to formulate its teacher models.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). Kipf and Welling introduce foundational graph convolutional networks whose recursive neighborhood aggregation latency the source seeks to circumvent via distillation into multi-layer perceptrons.
- Paper: Inductive Representation Learning on Large Graphs, William L. Hamilton et al. (2017). Hamilton et al. define inductive neighborhood aggregation and sampling frameworks on large graphs, establishing the classical GNN inference mechanics replaced by the source's linkless architecture.
- Paper: Simplifying Graph Convolutional Networks, Felix Wu et al. (2019). Wu et al. explore the computational simplification of graph convolutions by decoupling relational propagation and neural transformations, motivating low-latency graph representation learning.
- Paper: Link Prediction in Complex Networks: A Survey, Linyuan Lu et al. (2010). Lu and Zhou provide a comprehensive survey of topological heuristics and similarity indices for link prediction that underpin the evaluation benchmarks and structural baselines analyzed in the source.
- Paper: Beyond Homophily in Graph Neural Networks: Current Limitations and Effective Designs, Jiong Zhu et al. (2020). Zhu et al. analyze the trade-offs between neighborhood aggregation and standalone multi-layer perceptrons across diverse graph environments, directly informing the source's investigation into feature-rich MLP capabilities.
- Paper: Label-free Node Classification on Graphs with Large Language Models (LLMs), Zhikai Chen et al. (2024). Chen et al. extend label-efficient graph learning by combining language model zero-shot distillation with structural graph learning for label-free node classification.
