Towards Better Evaluation for Dynamic Link Prediction

Farimah PoursafaeiShenyang HuangKellin PelrineReihaneh Rabbany

article2022NeurIPS170 citations

Exposes critical flaws in dynamic link prediction evaluations by introducing the surprisingly competitive memorization baseline EdgeBank, two harder negative sampling strategies, and six diverse dynamic graph benchmarks to enable meaningful model comparison.

Listen

Many critical real-world systems—such as financial transactions, transportation networks, and social platforms—are modeled as dynamic graphs where connections evolve over time. While recent machine learning methods report near-perfect accuracy in predicting future connections (dynamic link prediction), these results are misleading. Current evaluation benchmarks rely on simplistic tests and narrow datasets, masking model deficiencies and preventing practitioners from determining which models actually generalize to real-world conditions.

The main objective of the article is to expose the critical weaknesses in existing dynamic link prediction evaluations and establish a more rigorous, realistic benchmarking framework. To achieve this, the article introduces six new datasets from diverse domains, proposes two novel negative sampling strategies, and develops a pure memorization baseline to measure the true predictive value of complex neural architectures.

The investigation evaluated five state-of-the-art dynamic graph neural network models alongside the new memorization baseline across thirteen datasets (seven existing and six newly introduced from politics, economics, and transportation). Performance was measured using standard accuracy metrics under three distinct negative sampling settings: the conventional random approach, a historical setting (testing whether models know when previously seen connections temporarily disappear), and an inductive setting (testing unseen connections).

The analysis produced several vital findings. First, existing evaluation setups are largely testing basic memorization; a simple, parameter-free baseline named EdgeBank—which merely predicts that previously seen connections will reoccur—matched or exceeded state-of-the-art neural models on multiple benchmarks. Second, when evaluated under realistic negative sampling conditions with reoccurring or inductive connections, the performance of complex models degraded sharply, dropping substantially on datasets such as flight networks. Third, the relative performance ranking of state-of-the-art models shifted dramatically under different sampling strategies, demonstrating that models excelling under current standard benchmarks often fail to generalize. Finally, a strong correlation emerged between a model's reliance on memorization and its performance collapse under challenging evaluation settings.

These findings indicate that organizations deploying dynamic graph models based on standard benchmark scores risk severe operational underperformance and wasted investment in overly complex architectures. Highly parameterized deep learning models may offer minimal value over simple rule-based memorization when temporal dynamics are repetitive, while simultaneously failing when required to predict non-trivial connection changes.

Decision-makers and engineering teams should immediately adopt more stringent evaluation protocols, specifically testing models against both historical and inductive negative sampling before production deployment. In addition, practitioners should implement simple memorization baselines as mandatory performance benchmarks to verify whether complex deep learning approaches provide genuine return on investment.

The evaluation carries high confidence across transductive network environments, where all interacting entities are observed during training. However, key limitations remain: the study evaluated dynamic link prediction under a single-point train-test time split rather than continuous multi-point splits, and it restricted its core baseline evaluations to transductive settings. Further testing across broader continuous-time forecasting setups and related tasks, such as node classification, is recommended.

arXiv: 2207.10128fpour/DGB.git
  • Paper: Benchmarking Graph Neural Networks, Vijay Prakash Dwivedi et al. (2023). Expands comprehensive and controlled graph neural network benchmarking across diverse medium-scale tasks under standardized parameter budgets and evaluation protocols.
Cover for Towards Better Evaluation for Dynamic Link Prediction

Abstract

Despite the prevalence of recent success in learning from static graphs, learning from time-evolving graphs remains an open challenge. In this work, we design new, more stringent evaluation procedures for link prediction specific to dynamic graphs, which reflect real-world considerations, to better compare the strengths and weaknesses of methods. First, we create two visualization techniques to understand the reoccurring patterns of edges over time and show that many edges reoccur at later time steps. Based on this observation, we propose a pure memorization-based baseline called EdgeBank. EdgeBank achieves surprisingly strong performance across multiple settings which highlights that the negative edges used in the current evaluation are easy. To sample more challenging negative edges, we introduce two novel negative sampling strategies that improve robustness and better match real-world applications. Lastly, we introduce six new dynamic graph datasets from a diverse set of domains missing from current benchmarks, providing new challenges and opportunities for future research. Our code repository is accessible at https://github.com/fpour/DGB.git.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Understanding Dynamic Graph Datasets
  • 3.1 Temporal Graph Datasets
  • 3.2 Temporal Edge Appearance (TEA) Plot
  • 3.3 Temporal Edge Traffic (TET) Plot
  • 4 EdgeBank: A Baseline for Dynamic Link Prediction
  • 5 Revisiting Negative Sampling in Dynamic Graphs
  • 6 Experiments
  • 7 Conclusion
  • References
  • Checklist

Knowls

  1. Knowl 1 — EdgeBank Memorization Baseline for Dynamic Link Prediction

    model/method

    EdgeBank is a non-parametric, memorization-based baseline for dynamic link prediction in continuous-time dynamic graphs. In a dynamic graph represented as a stream of timestamped edges G={(s0,d0,t0),(s1,d1,t1),…,(sT,dT,T)}G = \{(s_0, d_0, t_0), (s_1, d_1, t_1), \dots, (s_T, d_T, T)\}, EdgeBank maintains an edge memory (a dictionary tracking previously observed node pairs). At test time tt, EdgeBank predicts that an edge (u,v)(u, v) is positive if (u,v)(u, v) is present in its memory dictionary, and negative otherwise.

    EdgeBank operates under two memory management strategies:

    1. EdgeBank∞\text{EdgeBank}_\infty (Infinite Memory): Stores all unique edges observed from timestamp t=0t=0 up to the current timestamp. It is capable of recalling long-past connections but is prone to false positives on transient edges that appeared once and never reoccur.
    2. EdgeBanktw\text{EdgeBank}_{tw} (Time-Window Memory): Retains only the edges observed within a sliding window of fixed temporal length WW immediately preceding the current evaluation timestamp (i.e., edges observed in [t−W,t)[t - W, t)). The window length WW is set to match the time span of the validation split, capturing recent edge activity.

    EdgeBank requires zero learned parameters, zero hyperparameter optimization (other than setting the window size from validation split duration in EdgeBanktw\text{EdgeBank}_{tw}), and its memory storage is bounded by the total number of unique observed edges.

  2. Knowl 2 — Historical Negative Sampling for Dynamic Graphs

    model/method

    Standard evaluation for dynamic link prediction samples negative edges uniformly at random from all non-existent node pairs, which tends to generate trivial negatives (pairs of nodes that never interact in reality). Historical Negative Sampling addresses this by evaluating whether a model can determine when an active edge reoccurs versus assuming that any previously seen edge is continuously active.

    Let EtrainE_{train} be the set of all training edges and EtE^t be the set of positive edges occurring at evaluation timestamp tt. At timestamp tt, historical negative edges are sampled from previously observed training edges that are not currently active at time tt: eneg∈Etrain∖Ete_{neg} \in E_{train} \setminus E^t If the number of available candidate historical negative edges is smaller than the number of positive edges at timestamp tt, the remaining required negative samples are filled using random negative sampling (sampling from U∖EallU \setminus E_{all}, where UU is the set of all node pairs and Eall=Etrain∪EtestE_{all} = E_{train} \cup E_{test}).

  3. Knowl 3 — Inductive Negative Sampling for Dynamic Graphs

    model/method

    Inductive Negative Sampling evaluates whether a dynamic graph model can learn the reoccurrence dynamics of newly introduced edges that were unseen during training.

    Let EtrainE_{train} denote all edges observed during training, EtestE_{test} denote all edges occurring during testing, and EtE^t denote the set of positive edges occurring at test timestamp tt. At test timestamp tt, inductive negative edges are sampled from edges that first appeared during the test period prior to timestamp tt but are absent at timestamp tt: eneg∈(Etest∖Etrain)∖Ete_{neg} \in (E_{test} \setminus E_{train}) \setminus E^t If the number of available candidate inductive negative edges is less than the required number of negative samples (equal to the number of positive edges at timestamp tt), the remaining negative edges are drawn using standard random negative sampling from non-dataset node pairs.

  4. Knowl 4 — Novelty Index for Temporal Edge Dynamics

    equation

    The novelty index quantifies the proportion of edges appearing at each timestamp in a continuous-time dynamic graph that have never been observed in prior timestamps. For a dynamic graph with discrete timestamp indices t∈{1,…,T}t \in \{1, \dots, T\}, the novelty index is defined as: novelty=1T∑t=1T∣Et∖Eseent∣∣Et∣\text{novelty} = \frac{1}{T} \sum_{t=1}^T \frac{|E^t \setminus E^t_{seen}|}{|E^t|} where Et={(s,d,te)∣te=t}E^t = \{(s, d, t_e) \mid t_e = t\} is the set of edges occurring at timestamp tt, and Eseent={(s,d,te)∣te<t}E^t_{seen} = \{(s, d, t_e) \mid t_e < t\} is the set of all edges observed strictly before timestamp tt.

    A higher novelty index signifies that a large portion of positive test edges are inductive (new), placing an upper bound on the maximum link prediction accuracy achievable by pure edge memorization.

  5. Knowl 5 — Reoccurrence and Surprise Indices for Evaluation Splits

    equation

    For a dynamic graph partitioned chronologically at timestamp tsplitt_{split} into a training edge set EtrainE_{train} (edges occurring at t≤tsplitt \le t_{split}) and a test edge set EtestE_{test} (edges occurring at t>tsplitt > t_{split}), the reoccurrence index and surprise index characterize the difficulty of transductive memorization and inductive generalization:

    1. Reoccurrence Index: The fraction of edges observed during training that reappear during the test period: reoccurrence=∣Etrain∩Etest∣∣Etrain∣\text{reoccurrence} = \frac{|E_{train} \cap E_{test}|}{|E_{train}|} A higher reoccurrence index indicates that a large fraction of the training graph structure remains active and relevant in the future.

    2. Surprise Index: The fraction of edges occurring in the test period that were never seen during the training period: surprise=∣Etest∖Etrain∣∣Etest∣\text{surprise} = \frac{|E_{test} \setminus E_{train}|}{|E_{test}|} A higher surprise index indicates a substantial presence of inductive links in the test period that cannot be predicted via pure historical edge memorization.

  6. Knowl 6 — Dynamic Link Prediction Benchmark Datasets

    data/table

    A benchmark suite of 13 continuous-time dynamic graph datasets across diverse domains (social, interaction, proximity, transport, politics, and economics), including seven established benchmarks and six newly introduced datasets (Flights, Can. Parl., US Legis., UN Trade, UN Vote, Contact):

    Dataset Domain # Nodes Total Edges Unique Edges Time Granularity Duration
    Wikipedia Social 9,227 157,474 18,257 Unix timestamp 1 month
    Reddit Social 10,984 672,447 78,516 Unix timestamp 1 month
    MOOC Interaction 7,144 411,749 178,443 Unix timestamp 17 months
    LastFM Interaction 1,980 1,293,103 154,993 Unix timestamp 1 month
    Enron Social 184 125,235 3,125 Unix timestamp 3 years
    Social Evo. Proximity 74 2,099,519 4,486 Unix timestamp 8 months
    UCI Social 1,899 59,835 20,296 Unix timestamp 196 days
    Flights Transport 13,169 1,927,145 395,072 days 4 months
    Can. Parl. Politics 734 74,478 51,331 years 14 years
    US Legis. Politics 225 60,396 26,423 congresses 12 congresses
    UN Trade Economics 255 507,497 36,182 years 32 years
    UN Vote Politics 201 1,035,742 31,516 years 72 years
    Contact Proximity 694 2,426,280 79,531 5 minutes 1 month

    The new datasets expand evaluation beyond online social/interaction networks:

    • Flights: airport flight tracking graph during COVID-19.
    • Can. Parl.: co-voting network among Canadian Members of Parliament.
    • US Legis.: bill co-sponsorship graph among US Senators.
    • UN Trade: international food and agriculture bilateral trade volume.
    • UN Vote: United Nations General Assembly roll-call voting alignment.
    • Contact: physical proximity sensor network between university students.
  7. Knowl 7 — Model Ranking Instability Under Historical and Inductive Negative Sampling

    empirical result

    When evaluating state-of-the-art dynamic graph neural networks (JODIE, DyRep, TGAT, TGN, CAWN) across 13 datasets under a 70%-15%-15% chronological train-val-test split:

    1. Standard random negative sampling produces artificially elevated AU-ROC scores (often >0.95>0.95) and suggests that CAWN is consistently the top-performing method across almost all benchmarks.
    2. Under Historical Negative Sampling and Inductive Negative Sampling, the performance of all deep dynamic graph models drops substantially (e.g., drops of 0.20 to over 0.50 AU-ROC points on datasets like Flights, Enron, and LastFM).
    3. Relative model rankings change significantly across sampling strategies. CAWN, which performs best under random NS, experiences severe performance collapse on datasets such as LastFM and Enron under historical and inductive NS, where models like TGAT, TGN, or DyRep outperform it.
  8. Knowl 8 — Competitiveness of EdgeBank Memorization Baselines Against Deep Dynamic Graph Models

    empirical result

    Under standard random negative sampling, the parameter-free memorization baseline EdgeBank achieves link prediction performance comparable to or exceeding complex dynamic graph neural networks (DGNNs like TGN, CAWN, TGAT, DyRep, and JODIE). On datasets with high edge repetition (such as LastFM, Enron, and UN Trade), EdgeBank outperforms highly parameterized models without requiring any training or representation learning.

    In the Historical Negative Sampling setting, the time-window variant EdgeBanktw\text{EdgeBank}_{tw} (which memorizes only edges seen within a window matching the validation split duration) achieves the second-highest average AU-ROC across all 13 benchmarks and establishes state-of-the-art performance on UN Trade, UN Vote, Flights, Enron, and Contact.

  9. Knowl 9 — Correlation Between Model Memorization Reliance and Performance Degradation

    empirical result

    The degree of performance degradation experienced by dynamic graph neural networks under hard negative sampling strategies (historical and inductive negative sampling) is directly correlated with how closely a model's predictions align with pure memorization (EdgeBank∞\text{EdgeBank}_\infty).

    Models whose test predictions exhibit the highest correlation with EdgeBank∞\text{EdgeBank}_\infty (such as CAWN, followed by JODIE) suffer the largest drops in AU-ROC when evaluated on historical and inductive negatives. Conversely, models with the lowest correlation to EdgeBank∞\text{EdgeBank}_\infty (such as DyRep) demonstrate the smallest performance degradation when moving from standard random negative sampling to hard negative sampling.

  10. Knowl 10 — Limitations of Dynamic Link Prediction Evaluation Methodology

    limitation

    The proposed dynamic graph evaluation framework has several documented limitations:

    1. Single Split Point: Link prediction evaluation relies on a single chronological split timestamp (tsplitt_{split}, 70%-15%-15%) to divide past and future edges, rather than evaluating continuous rolling time windows or exact time-of-occurrence prediction.
    2. Transductive Node Setup: Analysis of edge memorization and historical negative sampling assumes a transductive node setting where all nodes evaluated at test time are present in the training graph. Pure inductive node link prediction (predicting edges between previously unseen nodes) is not evaluated with EdgeBank or historical negative sampling.
    3. Task Scope: The investigation is confined strictly to continuous-time dynamic link prediction and does not extend to dynamic node classification.

Coverage note — TEA and TET graphical plotting formats were omitted in favor of the formal equations for their underlying metrics (novelty, reoccurrence, and surprise indices).

References

  1. 1.L. Backstrom and J. Leskovec. Supervised random walks: predicting and recommending links in social networks. In Proceedings of the fourth ACM international conference on Web search and data mining, pages 635–644, 2011.
  2. 2.M. A. Bailey, A. Strezhnev, and E. Voeten. Estimating dynamic state preferences from united nations voting data. Journal of Conflict Resolution, 61(2):430–456, 2017.
  3. 3.P. Bielak, K. Tagowski, M. Falkiewicz, T. Kajdanowicz, and N. V. Chawla. FILDNE: A framework for incremental learning of dynamic networks embeddings. Knowledge-Based Systems, 236:107453, 2022.
  4. 4.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  5. 5.V. P. Dwivedi, C. K. Joshi, T. Laurent, Y. Bengio, and X. Bresson. Benchmarking graph neural networks. arXiv preprint arXiv:2003.00982, 2020.
  6. 6.F. Errica, M. Podda, D. Bacciu, and A. Micheli. A fair comparison of graph neural networks for graph classification. In ICLR, 2020.
  7. 7.J. H. Fowler. Legislative cosponsorship networks in the us house and senate. Social networks, 28(4):454–465, 2006.
  8. 8.A. Grover and J. Leskovec. node2vec: Scalable feature learning for networks. In Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining, pages 855–864, 2016.
  9. 9.T. Gutiérrez-Bunster, U. Stege, A. Thomo, and J. Taylor. How do biological networks differ from social networks?(an experimental study). In 2014 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM 2014), pages 744–751. IEEE, 2014.
  10. 10.S. Haghani and M. R. Keyvanpour. A systemic analysis of link prediction in social network. Artificial Intelligence Review, 52(3):1961–1995, 2019.
  11. 11.W. Hu, M. Fey, H. Ren, M. Nakata, Y. Dong, and J. Leskovec. Ogb-lsc: A large-scale challenge for machine learning on graphs. arXiv preprint arXiv:2103.09430, 2021.
  12. 12.W. Hu, M. Fey, M. Zitnik, Y. Dong, H. Ren, B. Liu, M. Catasta, and J. Leskovec. Open graph benchmark: Datasets for machine learning on graphs. arXiv preprint arXiv:2005.00687, 2020.
  13. 13.S. Huang, Y. Hitti, G. Rabusseau, and R. Rabbany. Laplacian change point detection for dynamic graphs. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 349–358, 2020.
  14. 14.R. R. Junuthula, K. S. Xu, and V. K. Devabhaktuni. Evaluating link prediction accuracy in dynamic networks with added and removed edges. In 2016 IEEE international conferences on big data and cloud computing (BDCloud), social computing and networking (SocialCom), sustainable computing and communications (SustainCom)(BDCloud-SocialCom-SustainCom), pages 377–384. IEEE, 2016.
  15. 15.R. R. Junuthula, K. S. Xu, and V. K. Devabhaktuni. Leveraging friendship networks for dynamic link prediction in social interaction networks. In Twelfth International AAAI Conference on Web and Social Media, 2018.
  16. 16.S. M. Kazemi, R. Goel, K. Jain, I. Kobyzev, A. Sethi, P. Forsyth, and P. Poupart. Representation learning for dynamic graphs: A survey. J. Mach. Learn. Res., 21(70):1–73, 2020.
  17. 17.B. Kotnis and V. Nastase. Analysis of the impact of negative sampling on link prediction in knowledge graphs. arXiv:1708.06816, 2017.
  18. 18.S. Kumar, X. Zhang, and J. Leskovec. Predicting dynamic embedding trajectory in temporal interaction networks. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019.
  19. 19.D. Liben-Nowell and J. Kleinberg. The link-prediction problem for social networks. Journal of the American society for information science and technology, 58(7):1019–1031, 2007.
  20. 20.R. N. Lichtenwalter, J. T. Lussier, and N. V. Chawla. New perspectives and methods in link prediction. In Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 243–252, 2010.
  21. 21.L. Lü and T. Zhou. Link prediction in complex networks: A survey. Physica A: statistical mechanics and its applications, 390(6):1150–1170, 2011.
  22. 22.Q. Lv, M. Ding, Q. Liu, Y. Chen, W. Feng, S. He, C. Zhou, J. Jiang, Y. Dong, and J. Tang. Are we really making much progress? revisiting, benchmarking, and refining heterogeneous graph neural networks. 2021.
  23. 23.G. K. MacDonald, K. A. Brauman, S. Sun, K. M. Carlson, E. S. Cassidy, J. S. Gerber, and P. C. West. Rethinking agricultural trade relationships in an era of globalization. BioScience, 65(3):275–289, 2015.
  24. 24.A. Madan, M. Cebrian, S. Moturu, K. Farrahi, et al. Sensing the" health state" of a community. IEEE Pervasive Computing, 11(4), 2011.
  25. 25.X. Olive. Traffic, a toolbox for processing and analysing air traffic data. Journal of Open Source Software, 4(39):1518–1, 2019.
  26. 26.P. Panzarasa, T. Opsahl, and K. M. Carley. Patterns and dynamics of users’ behavior and interaction: Network analysis of an online community. Journal of the American Society for Information Science and Technology, 60(5):911–932, 2009.
  27. 27.J. W. Pennebaker, M. E. Francis, and R. J. Booth. Linguistic inquiry and word count: Liwc 2001. Mahway: Lawrence Erlbaum Associates, 71(2001):2001, 2001.
  28. 28.E. Rossi, B. Chamberlain, F. Frasca, D. Eynard, F. Monti, and M. Bronstein. Temporal graph networks for deep learning on dynamic graphs. arXiv preprint arXiv:2006.10637, 2020.
  29. 29.A. Sankar, Y. Wu, L. Gou, W. Zhang, and H. Yang. DySAT: Deep neural representation learning on dynamic graphs via self-attention networks. In Proceedings of the 13th international conference on web search and data mining, pages 519–527, 2020.
  30. 30.P. Sapiezynski, A. Stopczynski, D. D. Lassen, and S. Lehmann. Interaction data from the copenhagen networks study. Scientific Data, 6(1):1–10, 2019.
  31. 31.M. Schäfer, M. Strohmeier, V. Lenders, I. Martinovic, and M. Wilhelm. Bringing up opensky: A large-scale ads-b sensor network for research. In IPSN-14 Proceedings of the 13th International Symposium on Information Processing in Sensor Networks, pages 83–94. IEEE, 2014.
  32. 32.J. Scripps, P.-N. Tan, F. Chen, and A.-H. Esfahanian. A matrix alignment approach for link prediction. In 2008 19th International Conference on Pattern Recognition, pages 1–4. IEEE, 2008.
  33. 33.O. Shchur, M. Mumme, A. Bojchevski, and S. Günnemann. Pitfalls of graph neural network evaluation. arXiv preprint arXiv:1811.05868, 2018.
  34. 34.J. Shetty and J. Adibi. The enron email dataset database schema and brief statistical report. Information sciences institute technical report, University of Southern California, 4(1):120–128, 2004.
  35. 35.J. Skardinga, B. Gabrys, and K. Musial. Foundations and modelling of dynamic networks using dynamic graph neural networks: A survey. IEEE Access, 2021.
  36. 36.M. Strohmeier, X. Olive, J. Lübbe, M. Schäfer, and V. Lenders. Crowdsourced air traffic data from the opensky network 2019–2020. Earth System Science Data, 13(2):357–366, 2021.
  37. 37.S. Tian, T. Xiong, and L. Shi. Streaming dynamic graph neural networks for continuous-time temporal graph modeling. In 2021 IEEE International Conference on Data Mining (ICDM), pages 1361–1366. IEEE, 2021.
  38. 38.R. Trivedi, M. Farajtabar, P. Biswal, and H. Zha. Dyrep: Learning representations over dynamic graphs. In International conference on learning representations, 2019.
  39. 39.E. Voeten, A. Strezhnev, and M. Bailey. United Nations General Assembly Voting Data, 2009.
  40. 40.Y. Wang, Y. Cai, Y. Liang, H. Ding, C. Wang, S. Bhatia, and B. Hooi. Adaptive data augmentation on temporal graphs. Advances in Neural Information Processing Systems, 34, 2021.
  41. 41.Y. Wang, Y.-Y. Chang, Y. Liu, J. Leskovec, and P. Li. Inductive representation learning in temporal networks via causal anonymous walks. In International Conference on Learning Representations, 2020.
  42. 42.D. Xu, C. Ruan, E. Korpeoglu, S. Kumar, and K. Achan. Inductive representation learning on temporal graphs. arXiv preprint arXiv:2002.07962, 2020.
  43. 43.M. Yang, M. Zhou, M. Kalander, Z. Huang, and I. King. Discrete-time temporal network embedding via implicit hierarchical learning in hyperbolic space. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pages 1975–1985, 2021.
  44. 44.Y. Yang, R. N. Lichtenwalter, and N. V. Chawla. Evaluating link prediction methods. Knowledge and Information Systems, 45(3):751–782, 2015.
  45. 45.Z. Yang, M. Ding, C. Zhou, H. Yang, J. Zhou, and J. Tang. Understanding negative sampling in graph representation learning. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1666–1676, 2020.
  46. 46.Y. Zhang, Y. Xiong, D. Li, C. Shan, K. Ren, and Y. Zhu. Cope: Modeling continuous propagation and evolution on interaction graph. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, pages 2627–2636, 2021.

Citation

MLA
Poursafaei, F., et al. “Towards Better Evaluation for Dynamic Link Prediction”. arXiv, 2022, http://arxiv.org/abs/2207.10128v2.
APA
Poursafaei, F., Huang, S., Pelrine, K., & Rabbany, R. (2022). Towards Better Evaluation for Dynamic Link Prediction. arXiv. http://arxiv.org/abs/2207.10128v2
Chicago
Poursafaei, F., S. Huang, K. Pelrine, and R. Rabbany. 2022. “Towards Better Evaluation for Dynamic Link Prediction”. arXiv. http://arxiv.org/abs/2207.10128v2.
Harvard
Poursafaei, F. et al. (2022) “Towards Better Evaluation for Dynamic Link Prediction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2207.10128v2.
Vancouver
1. Poursafaei F, Huang S, Pelrine K, Rabbany R (2022) Towards Better Evaluation for Dynamic Link Prediction. arXiv

BibTeX

@article{poursafaei2022towards,
  title = {Towards Better Evaluation for Dynamic Link Prediction},
  author = {Poursafaei, Farimah and Huang, Shenyang and Pelrine, Kellin and Rabbany, Reihaneh},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2207.10128v2},
  eprint = {2207.10128}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors