Pitfalls of Graph Neural Network Evaluation
Oleksandr ShchurMaximilian MummeAleksandar BojchevskiStephan Günnemann
Demonstrates how standard graph neural network benchmarks produce misleading model comparisons through rigid data splits and inconsistent training setups, revealing that simpler architectures often outperform complex ones under fair evaluation.
Graph Neural Networks (GNNs)—machine learning models designed to classify interconnected data such as social networks, scientific literature, and e-commerce platforms—have seen rapid development and widespread adoption. However, assessing actual progress in the field has become unreliable. Many newly proposed architectures rely on flawed benchmarking practices, primarily testing on identical, single data splits and using inconsistent training procedures across models. This practice risks mistaking overfitted setups and tuning advantages for genuine architectural breakthroughs.
To establish a rigorous baseline, the article evaluated four leading GNN architectures (GCN, MoNet, GAT, and GraphSAGE) alongside four traditional baselines to assess how standardized training conditions and varied data splits impact model performance and rankings. The authors developed a unified benchmarking framework that enforces identical optimization rules, early-stopping criteria, and hyperparameter tuning protocols across all models. Testing was conducted across eight distinct graph datasets—comprising four established citation networks and four newly introduced co-authorship and co-purchase networks—evaluated over 100 random data splits with 20 random weight initializations per split.
The findings reveal that model rankings are highly fragile when evaluated on single data splits. Testing on the standard single benchmark split ranked GAT highest on two datasets, but evaluating the exact same models on a different random split completely inverted the rankings, placing GCN first. When aggregated across all splits and datasets, all GNN architectures significantly outperformed non-graph baselines, but simpler architectures proved superior to more complex ones: the simpler GCN model achieved the highest overall relative accuracy (99.4%) and best average rank (2.3), outperforming more sophisticated models like GAT (95.9% relative accuracy, rank 3.6). Additionally, the evaluation uncovered stability risks; for instance, GAT suffered severe performance drops (falling below 40% accuracy) on specific random weight initializations in certain datasets, showing high variance despite comparable median performance.
These results demonstrate that the perceived performance gains of increasingly complex graph architectures are often artifacts of specific data splits or hyperparameter tuning rather than architectural superiority. For organizations deploying machine learning on relational data, adopting complex models introduces unnecessary engineering overhead, compute costs, and operational risk when simpler, well-tuned models achieve equivalent or better generalization.
Organizations and practitioners evaluating or implementing graph-based models should immediately abandon single-split evaluations in favor of robust, multi-split benchmarks with unified hyperparameter tuning. When selecting models for production, decision-makers should default to simpler architectures like GCN as primary baselines before adopting complex mechanisms like graph attention. Future research and development should focus on investigating specific graph structural properties that drive performance differences across diverse domain applications.
The conclusions provide high confidence regarding the necessity of multi-split benchmarking and standardized training within the evaluated conditions. However, users should note that the evaluation was restricted to transductive semi-supervised node classification using two-layer architectures and fixed training sample sizes (20 labeled nodes per class), meaning findings should be validated separately if applying models to inductive tasks or distinct operational scales.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). This seminal paper introduces the Graph Convolutional Network (GCN) architecture and the standard citation network evaluation protocol that the source critically assesses and re-evaluates.
- Paper: Inductive Representation Learning on Large Graphs, William L. Hamilton et al. (2017). GraphSAGE establishes one of the foundational spatial graph neural network architectures and benchmarking practices evaluated and compared in the source study.
- Paper: Graph Attention Networks, Petar Veličković et al. (2018). This work introduces Graph Attention Networks (GAT), one of the primary prominent architectures whose empirical performance and evaluation fairness are analyzed by the source.
- Paper: Revisiting Semi-Supervised Learning with Graph Embeddings, Zhilin Yang et al. (2016). This paper established the widely adopted semi-supervised citation network benchmark splits (e.g., Planetoid splits for Cora, CiteSeer, and PubMed) whose methodological limitations are directly challenged in the source.
- Paper: Geometric Deep Learning on Graphs and Manifolds Using Mixture Model CNNs, Federico Monti et al. (2017). MoNet provides a key spatial generalization framework for graph convolutional architectures that serves as part of the core comparison suite in early GNN benchmarking.
- Paper: Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering, Michaël Defferrard et al. (2016). This foundational work introduces ChebNet and localized spectral filtering, establishing the computational formulation upon which first-order GCNs and subsequent evaluation baselines were built.
- Paper: Deeper Insights into Graph Convolutional Networks for Semi-Supervised Learning, Qimai Li et al. (2018). This paper analyzes the smoothing behavior and low-label-rate performance of GCNs, providing important context for why standard train/test splits can lead to misleading conclusions.
- Paper: Statistical Comparisons of Classifiers over Multiple Data Sets, Janez Demšar (2006). This methodological paper outlines the formal statistical testing procedures and experimental design principles necessary for conducting valid classifier comparisons across datasets.
- Paper: Open Graph Benchmark: Datasets for Machine Learning on Graphs, Weihua Hu et al. (2020). Open Graph Benchmark addresses the benchmarking and data-split pitfalls identified in the source by establishing realistic, large-scale standardized datasets and split strategies.
- Paper: Predict then Propagate: Graph Neural Networks meet Personalized PageRank, Johannes Gasteiger et al. (2019). This work directly responds to GNN evaluation pitfalls by adopting rigorous evaluation protocols (100 random splits, bootstrap confidence intervals) and showing that decoupled propagation outperforms complex architectures.
- Paper: Simplifying Graph Convolutional Networks, Felix Wu et al. (2019). This paper validates the source's finding that simpler models can outperform sophisticated ones by demonstrating that removing non-linearities from GCNs yields competitive performance at lower computational cost.
- Paper: How Powerful are Graph Neural Networks?, Keyulu Xu et al. (2019). This work establishes a formal theoretical framework to analyze the expressive capacity of GNN architectures beyond empirical benchmark rankings.
- Paper: Fast Graph Representation Learning with PyTorch Geometric, Matthias Fey et al. (2019). PyTorch Geometric implements standardized implementations and benchmarking splits across dozens of graph architectures, directly addressing the reproducibility challenges highlighted in the source.
- Paper: Simple and Deep Graph Convolutional Networks, Ming Chen et al. (2020). This paper explores deep GCN architectures using rigorous experimental protocols that guard against the hyperparameter and split-dependent pitfalls identified in the source.
- Paper: Analyzing Learned Molecular Representations for Property Prediction, Kevin Yang et al. (2019). This work examines the impact of scaffold and chronological data splitting methodologies on molecular GNN evaluation, confirming the critical role of split design on model rankings.
