Do Transformers Really Perform Badly for Graph Representation?
Chengxuan YingTianle CaiShengjie LuoShuxin ZhengGuolin KeDi HeYanming ShenTie-Yan Liu
Introduces Graphormer, a standard Transformer-based architecture that integrates centrality, spatial, and edge encodings to outperform conventional graph neural networks across major graph representation benchmarks.
Standard deep learning architectures based on attention mechanisms, known as Transformers, have achieved state-of-the-art results across natural language processing and computer vision. However, adapting them directly to graph-structured data—such as molecular structures and social networks—has historically failed to outperform mainstream Graph Neural Networks (GNNs). This performance gap exists because standard self-attention calculates semantic similarity between nodes but fails to capture the essential topological structure and non-Euclidean connectivity inherent to graphs.
The article aims to evaluate whether a standard Transformer architecture can achieve superior performance on graph-level representation tasks when provided with structural inductive biases. To demonstrate this, the authors introduce Graphormer, an architecture that encodes structural information directly into the self-attention mechanism, and evaluate its mathematical expressiveness alongside its empirical performance against leading GNN benchmarks.
The approach integrates three core structural components into the standard Transformer framework: Centrality Encoding, which adds learnable degree-based vectors to node inputs to reflect node importance; Spatial Encoding, which inserts a learnable bias into the attention matrix based on the shortest path distance between node pairs; and Edge Encoding, which incorporates edge feature averages along those shortest paths into the attention calculation. The authors evaluated Graphormer across several premier public benchmarks, including the PCQM4M-LSC quantum chemistry regression dataset (exceeding 3.8 million graphs), as well as MolHIV, MolPCBA, and ZINC, comparing results against standard message-passing networks and earlier Transformer adaptations.
The findings show that structural encodings allow the Transformer to significantly outperform existing methods. First, on the large-scale PCQM4M-LSC challenge, Graphormer achieved a validation Mean Absolute Error of 0.1234, representing an approximate 11.5% relative error reduction over the previous best-performing model (GIN-VN at 0.1395) and ultimately winning first place in the competition. Second, Graphormer established new state-of-the-art results across all evaluated benchmarks, achieving an Average Precision of 31.39% on MolPCBA, an Area Under the Curve of 80.51% on MolHIV, and an error of 0.122 on ZINC. Third, ablation analyses confirmed that Spatial Encoding and degree-based Centrality Encoding are essential drivers of these performance gains, outperforming traditional Laplacian positional encodings. Finally, mathematical analysis proved that popular GNN variants are special cases of Graphormer, confirming it possesses strictly greater expressive capacity beyond standard message-passing limits without suffering from common degradation issues like over-smoothing.
These results indicate that specialized message-passing architectures are not fundamentally required for graph learning; rather, standard global attention architectures can effectively model complex graph topologies when supplied with appropriate relational biases. For organizations working on molecular property prediction, drug discovery, or materials science, Graphormer offers a unified, highly scalable modeling alternative with higher predictive accuracy. The primary actionable recommendation is to explore Graphormer as a baseline for large-scale graph-level regression and classification tasks, particularly when large pre-training datasets are available.
However, decision-makers should note key operational limitations before large-scale deployment. Because standard self-attention exhibits quadratic computational and memory complexity relative to the number of nodes, Graphormer is currently resource-intensive and restricted when applied to massive, individual graphs. Future initiatives must develop efficient attention approximations, investigate specialized graph sampling methods for node-level tasks, and explore domain-specific encodings before the architecture can be universally deployed across all graph scales.
- Paper: How Powerful are Graph Neural Networks?, Keyulu Xu et al. (2019). This paper establishes the theoretical expressive limits of standard message-passing GNNs via the Weisfeiler-Lehman test, providing the baseline criteria against which Graphormer's expressive superiority is proved.
- Paper: Neural Message Passing for Quantum Chemistry, Justin Gilmer et al. (2017). This work formalizes Neural Message Passing networks on molecular graphs, creating the primary framework that Graphormer re-evaluates and surpasses.
- Paper: Graph Attention Networks, Petar Veličković et al. (2018). This foundational paper introduces masked self-attention over graph neighborhoods, serving as an important precursor to applying global attention mechanisms to graph structures.
- Paper: Weisfeiler and Leman Go Neural: Higher-Order Graph Neural Networks, Christopher Morris et al. (2019). This study analyzes the equivalence of standard GNNs to the 1-WL test and motivates higher-order formulations, framing the expressiveness problems that Graphormer addresses.
- Paper: Relational inductive biases, deep learning, and graph networks, Peter W. Battaglia et al. (2018). This survey conceptualizes relational inductive biases in neural networks, which directly inspires Graphormer’s strategy of incorporating structural spatial and centrality encodings into standard Transformers.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). This core work introduces scalable graph convolutional networks, establishing the standard local neighborhood aggregation baseline compared against Graphormer.
- Paper: Strategies for Pre-training Graph Neural Networks, Weihua Hu et al. (2020). This research details strategies for pre-training graph neural networks on large-scale molecular datasets, contextualizing the pre-training paradigms utilized by Graphormer.
- Paper: Heterogeneous Graph Transformer, Ziniu Hu et al. (2020). This paper demonstrates adapting Transformer self-attention to heterogeneous graphs, serving as a key precedent for designing attention mechanisms tailored to graph topologies.
- Paper: How Attentive are Graph Attention Networks?, Shaked Brody et al. (2021). This work uncovers fundamental expressiveness constraints in standard graph attention mechanisms and introduces dynamic weighting, providing deeper insights into attention formulations on graphs.
- Paper: Rethinking Attention with Performers, Krzysztof Choromanski et al. (2021). This paper introduces linear-complexity kernel attention approximations, directly addressing the quadratic computational bottleneck highlighted as a primary limitation of Graphormer.
