Benchmarking Graph Neural Networks
Vijay Prakash DwivediChaitanya K. JoshiAnh Tuan LuuThomas LaurentYoshua BengioXavier Bresson
Establishes a standardized benchmarking suite with diverse datasets and strictly controlled parameter budgets to enable fair, reproducible evaluations across graph neural network architectures and positional encoding strategies.
Graph neural networks have become a vital tool for machine learning across chemistry, physics, social networks, and optimization. However, tracking progress has been historically difficult due to reliance on small legacy datasets, inconsistent evaluation settings, and arbitrary model parameter sizes. These shortcomings made it hard to determine whether performance gains came from true architectural advances or extra parameter capacity. To establish rigorous standards, the article introduces a reproducible, open-source benchmarking framework designed to evaluate graph architectures fairly under fixed parameter budgets across diverse, medium-scale tasks.
To conduct these evaluations, the article curates twelve medium-scale datasets spanning molecular chemistry, social networks, computer vision, and mathematical modeling, covering node, edge, and whole-graph prediction tasks. The framework enforces strict fairness by constraining model complexity to fixed budgets of approximately 100,000 and 500,000 parameters. It compares standard message-passing models, such as vanilla graph convolutional networks and attention-based networks, against theoretically expressive higher-order Weisfeiler-Lehman networks across identical training and hardware configurations.
The experimental findings show that message-passing architectures consistently outperform higher-order networks on medium-scale benchmarks while scaling far more effectively. Anisotropic models that leverage directional attention and edge gating, such as Gated Graph ConvNets and Graph Attention Networks, achieved top performance across multiple domains, outperforming isotropic models on tasks like the Travelling Salesman Problem and collaboration link prediction. In contrast, higher-order networks suffered from high computational and memory complexity, frequently running out of memory or failing to converge when scaled to deeper layers. In addition, the article demonstrated that adding Laplacian positional encodings—derived from graph eigenvectors—dramatically improves message-passing models, raising cycle detection accuracy from baseline levels to over 99% and enabling standard architectures to overcome fundamental symmetry limitations.
These results demonstrate that theoretical expressiveness does not automatically translate to practical success if models cannot scale or train stably using standard deep learning practices like batch normalization. For organizations investing in graph learning, prioritizing computationally efficient message-passing architectures augmented with attention mechanisms and positional encodings offers superior accuracy, lower computational cost, and faster turnaround times. Standardizing parameter budgets also provides an objective method to assess true algorithmic quality rather than superficial gains from model size.
Practitioners and researchers should adopt standardized parameter budgets and integrate Laplacian positional encodings or attention-based mechanisms into their graph pipelines. Before deploying higher-order architectures in production, teams must address their memory and training stability limitations. Future work should focus on principled graph normalization methods, specialized hardware handling for large dense graphs, and refining positional encoding techniques to handle arbitrary sign ambiguities.
- Paper: Pitfalls of Graph Neural Network Evaluation, Oleksandr Shchur et al. (2018). This foundational critique highlights the widespread flaws and ranking instability of early GNN evaluation practices, directly establishing the methodological need for standardized benchmarks.
- Paper: Open Graph Benchmark: Datasets for Machine Learning on Graphs, Weihua Hu et al. (2020). This seminal benchmark suite provides critical context for modern large-scale, split-driven graph neural network evaluation standards that parallel and contextualize the benchmark design.
- Paper: How Powerful are Graph Neural Networks?, Keyulu Xu et al. (2019). It formalizes the Weisfeiler-Lehman theoretical limits and expressive capabilities of standard message-passing architectures evaluated systematically within the benchmark.
- Paper: Neural Message Passing for Quantum Chemistry, Justin Gilmer et al. (2017). It unifies spatial graph architectures into the Message Passing Neural Network framework, which serves as the core paradigm tested across the benchmark's mathematical and chemical datasets.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). It introduces Graph Convolutional Networks, serving as a primary baseline architecture evaluated under controlled parameter budgets.
- Paper: Graph Attention Networks, Petar Veličković et al. (2018). It defines Graph Attention Networks, another essential baseline architecture rigorously compared across the benchmark tasks.
- Paper: Inductive Representation Learning on Large Graphs, William L. Hamilton et al. (2017). It introduces GraphSAGE and inductive neighborhood aggregation, establishing an important reference model assessed in the comparative experiments.
- Paper: Fast Graph Representation Learning with PyTorch Geometric, Matthias Fey et al. (2019). It details PyTorch Geometric, the primary open-source deep learning framework on top of which standardized GNN benchmarking suites are implemented.
- Paper: MoleculeNet: a benchmark for molecular machine learning, Zhenqin Wu et al. (2017). It establishes standardized molecular machine learning evaluation protocols that motivated the inclusion and format of chemical datasets like ZINC and AQSOL.
- Paper: Do Transformers Really Perform Badly for Graph Representation?, Chengxuan Ying et al. (2021). It demonstrates how graph structural encodings empower Transformer architectures on graph tasks, building upon the positional encoding paradigms explored in the benchmark.
- Paper: Learn from Global Correlations: Enhancing Evolutionary Algorithm via Spectral GNN, Kaichen Ouyang et al. (2026). This work applies spectral graph neural network formulations and frequency-based filtering beyond standard benchmark domains to enhance population-level evolutionary optimization algorithms.
