Do Transformers Really Perform Badly for Graph Representation?

Chengxuan YingTianle CaiShengjie LuoShuxin ZhengGuolin KeDi HeYanming ShenTie-Yan Liu

article2021NeurIPS1,561 citations

Introduces Graphormer, a standard Transformer-based architecture that integrates centrality, spatial, and edge encodings to outperform conventional graph neural networks across major graph representation benchmarks.

Listen

Standard deep learning architectures based on attention mechanisms, known as Transformers, have achieved state-of-the-art results across natural language processing and computer vision. However, adapting them directly to graph-structured data—such as molecular structures and social networks—has historically failed to outperform mainstream Graph Neural Networks (GNNs). This performance gap exists because standard self-attention calculates semantic similarity between nodes but fails to capture the essential topological structure and non-Euclidean connectivity inherent to graphs.

The article aims to evaluate whether a standard Transformer architecture can achieve superior performance on graph-level representation tasks when provided with structural inductive biases. To demonstrate this, the authors introduce Graphormer, an architecture that encodes structural information directly into the self-attention mechanism, and evaluate its mathematical expressiveness alongside its empirical performance against leading GNN benchmarks.

The approach integrates three core structural components into the standard Transformer framework: Centrality Encoding, which adds learnable degree-based vectors to node inputs to reflect node importance; Spatial Encoding, which inserts a learnable bias into the attention matrix based on the shortest path distance between node pairs; and Edge Encoding, which incorporates edge feature averages along those shortest paths into the attention calculation. The authors evaluated Graphormer across several premier public benchmarks, including the PCQM4M-LSC quantum chemistry regression dataset (exceeding 3.8 million graphs), as well as MolHIV, MolPCBA, and ZINC, comparing results against standard message-passing networks and earlier Transformer adaptations.

The findings show that structural encodings allow the Transformer to significantly outperform existing methods. First, on the large-scale PCQM4M-LSC challenge, Graphormer achieved a validation Mean Absolute Error of 0.1234, representing an approximate 11.5% relative error reduction over the previous best-performing model (GIN-VN at 0.1395) and ultimately winning first place in the competition. Second, Graphormer established new state-of-the-art results across all evaluated benchmarks, achieving an Average Precision of 31.39% on MolPCBA, an Area Under the Curve of 80.51% on MolHIV, and an error of 0.122 on ZINC. Third, ablation analyses confirmed that Spatial Encoding and degree-based Centrality Encoding are essential drivers of these performance gains, outperforming traditional Laplacian positional encodings. Finally, mathematical analysis proved that popular GNN variants are special cases of Graphormer, confirming it possesses strictly greater expressive capacity beyond standard message-passing limits without suffering from common degradation issues like over-smoothing.

These results indicate that specialized message-passing architectures are not fundamentally required for graph learning; rather, standard global attention architectures can effectively model complex graph topologies when supplied with appropriate relational biases. For organizations working on molecular property prediction, drug discovery, or materials science, Graphormer offers a unified, highly scalable modeling alternative with higher predictive accuracy. The primary actionable recommendation is to explore Graphormer as a baseline for large-scale graph-level regression and classification tasks, particularly when large pre-training datasets are available.

However, decision-makers should note key operational limitations before large-scale deployment. Because standard self-attention exhibits quadratic computational and memory complexity relative to the number of nodes, Graphormer is currently resource-intensive and restricted when applied to massive, individual graphs. Future initiatives must develop efficient attention approximations, investigate specialized graph sampling methods for node-level tasks, and explore domain-specific encodings before the architecture can be universally deployed across all graph scales.

  • Paper: How Powerful are Graph Neural Networks?, Keyulu Xu et al. (2019). This paper establishes the theoretical expressive limits of standard message-passing GNNs via the Weisfeiler-Lehman test, providing the baseline criteria against which Graphormer's expressive superiority is proved.
  • Paper: Neural Message Passing for Quantum Chemistry, Justin Gilmer et al. (2017). This work formalizes Neural Message Passing networks on molecular graphs, creating the primary framework that Graphormer re-evaluates and surpasses.
  • Paper: Graph Attention Networks, Petar Veličković et al. (2018). This foundational paper introduces masked self-attention over graph neighborhoods, serving as an important precursor to applying global attention mechanisms to graph structures.
  • Paper: Weisfeiler and Leman Go Neural: Higher-Order Graph Neural Networks, Christopher Morris et al. (2019). This study analyzes the equivalence of standard GNNs to the 1-WL test and motivates higher-order formulations, framing the expressiveness problems that Graphormer addresses.
  • Paper: Relational inductive biases, deep learning, and graph networks, Peter W. Battaglia et al. (2018). This survey conceptualizes relational inductive biases in neural networks, which directly inspires Graphormer’s strategy of incorporating structural spatial and centrality encodings into standard Transformers.
  • Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). This core work introduces scalable graph convolutional networks, establishing the standard local neighborhood aggregation baseline compared against Graphormer.
  • Paper: Strategies for Pre-training Graph Neural Networks, Weihua Hu et al. (2020). This research details strategies for pre-training graph neural networks on large-scale molecular datasets, contextualizing the pre-training paradigms utilized by Graphormer.
  • Paper: Heterogeneous Graph Transformer, Ziniu Hu et al. (2020). This paper demonstrates adapting Transformer self-attention to heterogeneous graphs, serving as a key precedent for designing attention mechanisms tailored to graph topologies.
  • Paper: How Attentive are Graph Attention Networks?, Shaked Brody et al. (2021). This work uncovers fundamental expressiveness constraints in standard graph attention mechanisms and introduces dynamic weighting, providing deeper insights into attention formulations on graphs.
  • Paper: Rethinking Attention with Performers, Krzysztof Choromanski et al. (2021). This paper introduces linear-complexity kernel attention approximations, directly addressing the quadratic computational bottleneck highlighted as a primary limitation of Graphormer.
Cover for Do Transformers Really Perform Badly for Graph Representation?

Abstract

The Transformer architecture has become a dominant choice in many domains, such as natural language processing and computer vision. Yet, it has not achieved competitive performance on popular leaderboards of graph-level prediction compared to mainstream GNN variants. Therefore, it remains a mystery how Transformers could perform well for graph representation learning. In this paper, we solve this mystery by presenting Graphormer, which is built upon the standard Transformer architecture, and could attain excellent results on a broad range of graph representation learning tasks, especially on the recent OGB Large-Scale Challenge. Our key insight to utilizing Transformer in the graph is the necessity of effectively encoding the structural information of a graph into the model. To this end, we propose several simple yet effective structural encoding methods to help Graphormer better model graph-structured data. Besides, we mathematically characterize the expressive power of Graphormer and exhibit that with our ways of encoding the structural information of graphs, many popular GNN variants could be covered as the special cases of Graphormer. The code and models of Graphormer will be made publicly available at https://github.com/microsoft/Graphormer.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 3 Graphormer
  • 3.1 Structural Encodings in Graphormer
  • 3.1.1 Centrality Encoding
  • 3.1.2 Spatial Encoding
  • 3.1.3 Edge Encoding in the Attention
  • 3.2 Implementation Details of Graphormer
  • 3.3 How Powerful is Graphormer?
  • 4 Experiments
  • 4.1 OGB Large-Scale Challenge
  • 4.2 Graph Representation
  • 4.3 Ablation Studies
  • 5 Related Work
  • 5.1 Graph Transformer
  • 5.2 Structural Encodings in GNNs
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Graphormer Layer Architecture

    model/method

    Graphormer adapts the standard Transformer encoder for graph representation learning by employing a pre-Layer Normalization (Pre-LN) Transformer block equipped with graph-specific attention mechanisms.

    Let H(l−1)=[h1(l−1)op,…,hn(l−1)op]⊤∈Rn×dH^{(l-1)} = [h_1^{(l-1) op}, \dots, h_n^{(l-1) op}]^\top \in \mathbb{R}^{n \times d} denote the node representations input to the ll-th Graphormer layer, where n=∣V∣n = |V| is the number of nodes and dd is the hidden representation dimension. The layer computation is given by:

    h′(l)=MHA(LN(h(l−1)))+h(l−1)h'^{(l)} = \text{MHA}(\text{LN}(h^{(l-1)})) + h^{(l-1)}

    h(l)=FFN(LN(h′(l)))+h′(l)h^{(l)} = \text{FFN}(\text{LN}(h'^{(l)})) + h'^{(l)}

    where LN(⋅)\text{LN}(\cdot) denotes layer normalization applied prior to the multi-head self-attention module MHA(⋅)\text{MHA}(\cdot) and the position-wise feed-forward network FFN(⋅)\text{FFN}(\cdot). In the feed-forward network sub-layer, the input, output, and inner hidden representations all share the same dimensionality dd.

    In the self-attention calculation for a single attention head with projection matrices WQ,WK,WV∈Rd×dW_Q, W_K, W_V \in \mathbb{R}^{d \times d}, the attention score AijA_{ij} between node viv_i and node vjv_j is computed as:

    Aij=(hiWQ)(hjWK)⊤d+bϕ(vi,vj)+cijA_{ij} = \frac{(h_i W_Q)(h_j W_K)^\top}{\sqrt{d}} + b_{\phi(v_i, v_j)} + c_{ij}

    Attn(H)=softmax(A)HWV\text{Attn}(H) = \text{softmax}(A) H W_V

    where bϕ(vi,vj)b_{\phi(v_i, v_j)} is a learnable spatial relation bias and cijc_{ij} is an edge encoding bias.

  2. Knowl 2 — Centrality Encoding

    model/method

    Centrality Encoding incorporates node importance into the initial node representation of Graphormer. While standard self-attention calculates similarity based purely on semantic feature correlation, degree centrality provides structural significance signals (e.g., highly connected hubs in molecular or social graphs).

    For an input node vi∈Vv_i \in V with initial feature vector xi∈Rdx_i \in \mathbb{R}^d, the input representation hi(0)h_i^{(0)} to the first Transformer layer is defined as:

    hi(0)=xi+zdeg−(vi)−+zdeg+(vi)+h_i^{(0)} = x_i + z^{-}_{\text{deg}^-(v_i)} + z^{+}_{\text{deg}^+(v_i)}

    where deg−(vi)\text{deg}^-(v_i) and deg+(vi)\text{deg}^+(v_i) denote the indegree and outdegree of node viv_i, and zdeg−(vi)−,zdeg+(vi)+∈Rdz^{-}_{\text{deg}^-(v_i)}, z^{+}_{\text{deg}^+(v_i)} \in \mathbb{R}^d are learnable embedding vectors indexed by the degree values. For undirected graphs, indegree and outdegree coincide and are unified as a single degree embedding zdeg(vi)∈Rdz_{\text{deg}(v_i)} \in \mathbb{R}^d added to xix_i.

  3. Knowl 3 — Spatial Encoding via Shortest Path Distance

    model/method

    Spatial Encoding enables Graphormer to incorporate graph structural distances and topological inductive bias directly into the self-attention module.

    For any pair of nodes (vi,vj)(v_i, v_j) in a graph G=(V,E)G = (V, E), let ϕ(vi,vj):V×V→R\phi(v_i, v_j): V \times V \to \mathbb{R} denote the spatial relation function defined as the shortest path distance (SPD) between viv_i and vjv_j if a path exists, and a reserved value (e.g., −1-1) if the nodes are disconnected. The attention score matrix element AijA_{ij} between queries and keys is modulated by an additive bias:

    Aij=(hiWQ)(hjWK)⊤d+bϕ(vi,vj)A_{ij} = \frac{(h_i W_Q)(h_j W_K)^\top}{\sqrt{d}} + b_{\phi(v_i, v_j)}

    where bϕ(vi,vj)∈Rb_{\phi(v_i, v_j)} \in \mathbb{R} is a learnable scalar indexed by the shortest path distance ϕ(vi,vj)\phi(v_i, v_j) and shared across all Transformer layers. This allows all node pairs to attend to each other globally while adaptively adjusting attention weights based on topological graph distance.

  4. Knowl 4 — Edge Encoding in Self-Attention

    model/method

    Edge Encoding integrates edge structural features along paths connecting node pairs directly into the attention mechanism.

    For an ordered node pair (vi,vj)(v_i, v_j), let SPij=(e1,e2,…,eN)\text{SP}_{ij} = (e_1, e_2, \dots, e_N) denote a shortest path of length NN connecting viv_i to vjv_j, where ene_n is the nn-th edge on the path. The edge bias term cij∈Rc_{ij} \in \mathbb{R} added to the attention logit AijA_{ij} is defined as:

    cij=1N∑n=1Nxen(wnE)⊤c_{ij} = \frac{1}{N} \sum_{n=1}^{N} x_{e_n} (w_n^E)^\top

    where xen∈R1×dEx_{e_n} \in \mathbb{R}^{1 \times d_E} is the feature vector of edge ene_n, wnE∈R1×dEw_n^E \in \mathbb{R}^{1 \times d_E} is a learnable weight embedding corresponding to the nn-th hop position along the shortest path, and dEd_E is the edge feature dimensionality.

  5. Knowl 5 — Graph Readout via Virtual Special Node ([VNode])

    model/method

    Graphormer produces a whole-graph representation hGh_G by augmenting the graph with a dedicated virtual token, denoted [VNode][\text{VNode}], analogous to the [CLS][\text{CLS}] token in sequence Transformers.

    The [VNode][\text{VNode}] is connected to every node vj∈Vv_j \in V in the graph. To distinguish physical graph edges from virtual connections without distorting topological shortest path metrics, the spatial encoding for all pairs involving the virtual node is assigned a distinct learnable scalar parameter:

    bϕ([VNode],vj)=bϕ(vi,[VNode])=bvirtualb_{\phi([\text{VNode}], v_j)} = b_{\phi(v_i, [\text{VNode}])} = b_{\text{virtual}}

    During message aggregation and attention across LL layers, [VNode][\text{VNode}] is updated alongside all regular nodes. The final representation of the entire graph hGh_G is given directly by the representation of [VNode][\text{VNode}] at layer LL:

    hG=h[VNode](L)h_G = h_{[\text{VNode}]}^{(L)}

  6. Knowl 6 — Expressive Power and Subsumption of Classical GNNs

    theoretical result

    Graphormer with structural encodings strictly generalizes standard message passing graph neural networks:

    1. GNN Subsumption (Fact 1): By selecting appropriate weight matrices and spatial distance function ϕ\phi, a Graphormer layer can represent the AGGREGATE\text{AGGREGATE} and COMBINE\text{COMBINE} steps of popular GNN architectures including GIN, GCN, and GraphSAGE. Specifically, spatial encoding allows the self-attention mechanism to isolate 1-hop neighbor sets N(vi)\mathcal{N}(v_i) to compute mean statistics; degree centrality enables exact scaling from neighborhood mean to sum; and multi-head attention coupled with the feed-forward network allows processing node self-features and neighbor aggregates separately before combining them.

    2. Expressive Power Beyond 1-WL: Equipped with spatial encoding based on shortest path distances, Graphormer can distinguish non-isomorphic graphs that are indistinguishable by the 1-Weisfeiler-Lehman (1-WL) graph isomorphism test.

    3. Readout Simulation (Fact 2): Every node representation at the output of a vanilla Graphormer self-attention layer (even without additional encodings) can represent graph-level MEAN READOUT\text{MEAN}\ \text{READOUT} functions, enabling global information aggregation and broadcast across the graph without suffering from over-smoothing.

  7. Knowl 7 — Quantum Chemistry Regression Performance on PCQM4M-LSC

    empirical result

    On the PCQM4M-LSC dataset from the Open Graph Benchmark Large-Scale Challenge (comprising over 3.8 million molecular graphs for HOMO-LUMO gap prediction), Graphormer achieves state-of-the-art validation Mean Absolute Error (MAE), significantly outperforming standard message-passing GNNs and prior Graph Transformers.

    Method #Param. Train MAE Validate MAE
    GCN 2.0M 0.1318 0.1691
    GIN 3.8M 0.1203 0.1537
    GCN-VN 4.9M 0.1225 0.1485
    GIN-VN 6.7M 0.1150 0.1395
    GINE-VN 13.2M 0.1248 0.1430
    DeeperGCN-VN 25.5M 0.1059 0.1398
    GT 0.6M 0.0944 0.1400
    GT-Wide 83.2M 0.0955 0.1408
    GraphormerSMALL (L=6,d=512L=6, d=512) 12.5M 0.0778 0.1264
    Graphormer (L=12,d=768L=12, d=768) 47.1M 0.0582 0.1234

    Graphormer reduces validation MAE relative to the previous best GNN baseline (GIN-VN) by 11.5%. An ensemble of Graphormer with ExpC achieved 0.1200 test MAE on the complete PCQM4M-LSC test set, securing 1st place in the KDD Cup 2021 OGB-LSC graph track.

  8. Knowl 8 — Graph Classification and Regression Benchmarks on MolPCBA, MolHIV, and ZINC

    empirical result

    Graphormer demonstrates state-of-the-art transferability and training performance across OGB benchmarks (ogbg-molpcba, ogbg-molhiv) and the Benchmarking-GNN ZINC dataset.

    Dataset Method #Param. Metric
    MolPCBA DeeperGCN-VN+FLAG 5.6M 28.42 ±\pm 0.43 AP (%)
    DGN 6.7M 28.85 ±\pm 0.30 AP (%)
    GINE-APPNP 6.1M 29.79 ±\pm 0.30 AP (%)
    GIN-VN (pre-trained fine-tune) 3.4M 29.02 ±\pm 0.17 AP (%)
    Graphormer-FLAG 119.5M 31.39 ±\pm 0.32 AP (%)
    MolHIV GCN-GraphNorm 526K 78.83 ±\pm 1.00 AUC (%)
    DeeperGCN-FLAG 532K 79.42 ±\pm 1.20 AUC (%)
    DGN 114K 79.70 ±\pm 0.97 AUC (%)
    GIN-VN (pre-trained fine-tune) 3.3M 77.80 ±\pm 1.82 AUC (%)
    Graphormer-FLAG 47.0M 80.51 ±\pm 0.53 AUC (%)
    ZINC GIN 509K 0.526 ±\pm 0.051 MAE
    GatedGCN-PE 505K 0.214 ±\pm 0.006 MAE
    PNA 387K 0.142 ±\pm 0.010 MAE
    GT 589K 0.226 ±\pm 0.014 MAE
    SAN 508K 0.139 ±\pm 0.006 MAE
    GraphormerSLIM (L=12,d=80L=12, d=80) 489K 0.122 ±\pm 0.006 MAE

    On MolPCBA and MolHIV, Graphormer is pre-trained on PCQM4M-LSC and fine-tuned using FLAG data augmentation. On ZINC, GraphormerSLIM is trained from scratch under the 500k parameter budget.

  9. Knowl 9 — Ablation of Graphormer Structural Encodings

    empirical result

    Ablation experiments on the PCQM4M-LSC validation set (using 12-layer models trained for 100K iterations) verify the cumulative contributions of Spatial Encoding, Centrality Encoding, and Edge Encoding via attention bias.

    Node Relation Encoding Centrality Edge Encoding Validate MAE
    Laplacian PE Spatial via node via Aggr via attn bias
    - - - - - - 0.2276
    ✓ - - - - - 0.1483
    - ✓ - - - - 0.1427
    - ✓ ✓ - - - 0.1396
    - ✓ ✓ ✓ - - 0.1328
    - ✓ ✓ - ✓ - 0.1327
    - ✓ ✓ - - ✓ 0.1304

    Key observations:

    1. Shortest path distance-based Spatial Encoding (0.1427 MAE) outperforms Laplacian positional encoding (0.1483 MAE).
    2. Centrality encoding yields a notable boost from 0.1427 to 0.1396 MAE.
    3. Incorporating edge features as an attention bias along shortest paths achieves 0.1304 MAE, outperforming conventional node addition (0.1328 MAE) and neighborhood aggregation (0.1327 MAE) methods.
  10. Knowl 10 — Scalability and Complexity Limitations of Graphormer

    limitation

    Graphormer possesses key operational limitations:

    1. Quadratic Complexity: The standard full self-attention mechanism computes pairwise attention across all node pairs, resulting in O(∣V∣2)O(|V|^2) computational and memory complexity per layer. This quadratic scaling limits the direct application of Graphormer to very large graphs.
    2. General vs. Domain-Specific Encodings: While degree centrality and shortest path distance encodings provide general inductive biases, domain-specific geometric or biochemical encodings may be required for further performance gains on specialized datasets.
    3. Node Representation Extraction: Graphormer is primarily designed and evaluated for whole-graph prediction tasks; applying it efficiently to large-scale node classification requires dedicated graph sampling strategies.

Coverage note — None was omitted; all key architectural components, theoretical results, empirical benchmarks, ablations, and limitations are covered.

References

  1. 1.Jinheon Baek, Minki Kang, and Sung Ju Hwang. Accurate learning of graph representations with graph multiset pooling. ICLR, 2021.
  2. 2.Dominique Beaini, Saro Passaro, Vincent Létourneau, William L Hamilton, Gabriele Corso, and Pietro Liò. Directional graph networks. In International Conference on Machine Learning, 2021.
  3. 3.Mikhail Belkin and Partha Niyogi. Laplacian eigenmaps for dimensionality reduction and data representation. Neural computation, 15(6):1373–1396, 2003.
  4. 4.Xavier Bresson and Thomas Laurent. Residual gated graph convnets. arXiv preprint arXiv:1711.07553, 2017.
  5. 5.Rémy Brossard, Oriel Frigo, and David Dehaene. Graph convolutions that can finally model local structure. arXiv preprint arXiv:2011.15069, 2020.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc., 2020.
  7. 7.Deng Cai and Wai Lam. Graph transformer for graph-to-sequence learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7464–7471, 2020.
  8. 8.Tianle Cai, Shengjie Luo, Keyulu Xu, Di He, Tie-yan Liu, and Liwei Wang. Graphnorm: A principled approach to accelerating graph neural network training. In International Conference on Machine Learning, 2021.
  9. 9.Benson Chen, Regina Barzilay, and Tommi Jaakkola. Path-augmented graph transformer network. arXiv preprint arXiv:1905.12712, 2019.
  10. 10.Gabriele Corso, Luca Cavalleri, Dominique Beaini, Pietro Liò, and Petar Veličković. Principal neighbourhood aggregation for graph nets. Advances in Neural Information Processing Systems, 33, 2020.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, 2019.
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  13. 13.Vijay Prakash Dwivedi and Xavier Bresson. A generalization of transformer networks to graphs. AAAI Workshop on Deep Learning on Graphs: Methods and Applications, 2021.
  14. 14.Vijay Prakash Dwivedi, Chaitanya K Joshi, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. Benchmarking graph neural networks. arXiv preprint arXiv:2003.00982, 2020.
  15. 15.Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. Neural message passing for quantum chemistry. In International Conference on Machine Learning, pages 1263–1272. PMLR, 2017.
  16. 16.Liyu Gong and Qiang Cheng. Exploiting edge features for graph neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9211–9219, 2019.
  17. 17.Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al. Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100, 2020.
  18. 18.William L Hamilton, Zhitao Ying, and Jure Leskovec. Inductive representation learning on large graphs. In NIPS, 2017.
  19. 19.Vincent J Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis, and David Bieber. Global relational models of source code. In International conference on learning representations, 2019.
  20. 20.W Hu, B Liu, J Gomes, M Zitnik, P Liang, V Pande, and J Leskovec. Strategies for pre-training graph neural networks. In International Conference on Learning Representations (ICLR), 2020.
  21. 21.Weihua Hu, Matthias Fey, Hongyu Ren, Maho Nakata, Yuxiao Dong, and Jure Leskovec. Ogb-lsc: A large-scale challenge for machine learning on graphs. arXiv preprint arXiv:2103.09430, 2021.
  22. 22.Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. Open graph benchmark: Datasets for machine learning on graphs. arXiv preprint arXiv:2005.00687, 2020.
  23. 23.Ziniu Hu, Yuxiao Dong, Kuansan Wang, and Yizhou Sun. Heterogeneous graph transformer. In Proceedings of The Web Conference 2020, pages 2704–2710, 2020.
  24. 24.Katsuhiko Ishiguro, Shin-ichi Maeda, and Masanori Koyama. Graph warp module: an auxiliary module for boosting the power of graph neural networks in molecular graph analysis. arXiv preprint arXiv:1902.01020, 2019.
  25. 25.Guolin Ke, Di He, and Tie-Yan Liu. Rethinking the positional encoding in language pre-training. ICLR, 2020.
  26. 26.Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016.
  27. 27.Kezhi Kong, Guohao Li, Mucong Ding, Zuxuan Wu, Chen Zhu, Bernard Ghanem, Gavin Taylor, and Tom Goldstein. Flag: Adversarial data augmentation for graph neural networks. arXiv preprint arXiv:2010.09891, 2020.
  28. 28.Devin Kreuzer, Dominique Beaini, William Hamilton, Vincent Létourneau, and Prudencio Tossou. Rethinking graph transformers with spectral attention. arXiv preprint arXiv:2106.03893, 2021.
  29. 29.Tuan Le, Marco Bertolini, Frank Noé, and Djork-Arné Clevert. Parameterized hypercomplex graph neural networks for graph classification. arXiv preprint arXiv:2103.16584, 2021.
  30. 30.Guohao Li, Chenxin Xiong, Ali Thabet, and Bernard Ghanem. Deepergcn: All you need to train deeper gcns. arXiv preprint arXiv:2006.07739, 2020.
  31. 31.Junying Li, Deng Cai, and Xiaofei He. Learning graph-level representation for drug discovery. arXiv preprint arXiv:1709.03741, 2017.
  32. 32.Pan Li, Yanbang Wang, Hongwei Wang, and Jure Leskovec. Distance encoding: Design provably more powerful neural networks for graph representation learning. Advances in Neural Information Processing Systems, 33, 2020.
  33. 33.Yuan Li, Xiaodan Liang, Zhiting Hu, Yinbo Chen, and Eric P. Xing. Graph transformer, 2019.
  34. 34.Xi Victoria Lin, Richard Socher, and Caiming Xiong. Multi-hop knowledge graph reasoning with reward shaping. arXiv preprint arXiv:1808.10568, 2018.
  35. 35.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  36. 36.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  37. 37.P David Marshall. The promotion and presentation of the self: celebrity as marker of presentational media. Celebrity studies, 1(1):35–48, 2010.
  38. 38.Alice Marwick and Danah Boyd. To see and be seen: Celebrity practice on twitter. Convergence, 17(2):139–158, 2011.
  39. 39.Łukasz Maziarka, Tomasz Danel, Sławomir Mucha, Krzysztof Rataj, Jacek Tabor, and Stanisław Jastrzębski. Molecule attention transformer. arXiv preprint arXiv:2002.08264, 2020.
  40. 40.Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, et al. Do transformer modifications transfer across implementations and applications? arXiv preprint arXiv:2102.11972, 2021.
  41. 41.Dinglan Peng, Shuxin Zheng, Yatao Li, Guolin Ke, Di He, and Tie-Yan Liu. How could neural networks understand programs? In International Conference on Machine Learning. PMLR, 2021.
  42. 42.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020.
  43. 43.Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying Wei, Wenbing Huang, and Junzhou Huang. Self-supervised graph transformer on large-scale molecular data. Advances in Neural Information Processing Systems, 33, 2020.
  44. 44.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 464–468, 2018.
  45. 45.Yunsheng Shi, Zhengjie Huang, Wenjin Wang, Hui Zhong, Shikun Feng, and Yu Sun. Masked label prediction: Unified message passing model for semi-supervised classification. arXiv preprint arXiv:2009.03509, 2020.
  46. 46.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, 2017.
  47. 47.Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. Graph attention networks. ICLR, 2018.
  48. 48.Guangtao Wang, Rex Ying, Jing Huang, and Jure Leskovec. Direct multi-hop attention based graph neural network. arXiv preprint arXiv:2009.14332, 2020.
  49. 49.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. On layer normalization in the transformer architecture. In International Conference on Machine Learning, pages 10524–10533. PMLR, 2020.
  50. 50.Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks? In International Conference on Learning Representations, 2019.
  51. 51.Mingqi Yang, Yanming Shen, Heng Qi, and Baocai Yin. Breaking the expressive bottlenecks of graph neural networks. arXiv preprint arXiv:2012.07219, 2020.
  52. 52.Yiding Yang, Xinchao Wang, Mingli Song, Junsong Yuan, and Dacheng Tao. Spagan: Shortest path graph attention network. Advances in IJCAI, 2019.
  53. 53.Chengxuan Ying, Mingqi Yang, Shuxin Zheng, Guolin Ke, Shengjie Luo, Tianle Cai, Chenglin Wu, Yuxin Wang, Yanming Shen, and Di He. First place solution of kdd cup 2021 & ogb large-scale challenge graph-level track. arXiv preprint arXiv:2106.08279, 2021.
  54. 54.Jiaxuan You, Rex Ying, and Jure Leskovec. Position-aware graph neural networks. In International Conference on Machine Learning, pages 7134–7143. PMLR, 2019.
  55. 55.Seongjun Yun, Minbyul Jeong, Raehyun Kim, Jaewoo Kang, and Hyunwoo J Kim. Graph transformer networks. Advances in Neural Information Processing Systems, 32, 2019.
  56. 56.Jiawei Zhang, Haopeng Zhang, Congying Xia, and Li Sun. Graph-bert: Only attention is needed for learning graph representations. arXiv preprint arXiv:2001.05140, 2020.
  57. 57.Daniel Zügner, Tobias Kirschstein, Michele Catasta, Jure Leskovec, and Stephan Günnemann. Language-agnostic representation learning of source code from structure and context. In International Conference on Learning Representations, 2020.

Citation

MLA
Ying, C., et al. “Do Transformers Really Perform Badly for Graph Representation?”. Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 28877–88, https://proceedings.neurips.cc/paper_files/paper/2021/file/f1c1592588411002af340cbaedd6fc33-Paper.pdf.
APA
Ying, C., Cai, T., Luo, S., Zheng, S., Ke, G., He, D., Shen, Y., & Liu, T.-Y. (2021). Do Transformers Really Perform Badly for Graph Representation?. Advances in Neural Information Processing Systems, 34, 28877–28888. https://proceedings.neurips.cc/paper_files/paper/2021/file/f1c1592588411002af340cbaedd6fc33-Paper.pdf
Chicago
Ying, C., T. Cai, S. Luo, et al. 2021. “Do Transformers Really Perform Badly for Graph Representation?”. Advances in Neural Information Processing Systems 34: 28877–88. https://proceedings.neurips.cc/paper_files/paper/2021/file/f1c1592588411002af340cbaedd6fc33-Paper.pdf.
Harvard
Ying, C. et al. (2021) “Do Transformers Really Perform Badly for Graph Representation?”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 28877–28888. Available at: https://proceedings.neurips.cc/paper_files/paper/2021/file/f1c1592588411002af340cbaedd6fc33-Paper.pdf.
Vancouver
1. Ying C, Cai T, Luo S, Zheng S, Ke G, He D, Shen Y, Liu T-Y (2021) Do Transformers Really Perform Badly for Graph Representation?. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 28877–28888

BibTeX

@inproceedings{ying2021transformers,
  title = {Do Transformers Really Perform Badly for Graph Representation?},
  author = {Ying, Chengxuan and Cai, Tianle and Luo, Shengjie and Zheng, Shuxin and Ke, Guolin and He, Di and Shen, Yanming and Liu, Tie-Yan},
  year = {2021},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {34},
  pages = {28877-28888},
  url = {https://proceedings.neurips.cc/paper_files/paper/2021/file/f1c1592588411002af340cbaedd6fc33-Paper.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission