Graph Transformer Networks

Seongjun YunMinbyul JeongRaehyun KimJaewoo KangHyunwoo J. Kim

article2019NeurIPS1,445 citations

Introduces Graph Transformer Networks that dynamically learn task-specific meta-paths and multi-hop connections on heterogeneous graphs, eliminating the need for manual graph preprocessing while setting new performance standards in node classification.

Listen

Real-world data networks, such as citation indices, social networks, and recommendation systems, are typically heterogeneous—meaning they contain multiple types of entities (nodes) and diverse relationships (edges) between them. Standard machine learning models designed for graphs usually operate under the assumption that networks are uniform (homogeneous) and fixed. Existing techniques that handle heterogeneous structures rely heavily on human domain experts to manually define composite relationship paths (meta-paths) in a labor-intensive, multi-stage preprocessing workflow. This manual engineering limits model flexibility, risks missing important hidden patterns, and can degrade overall prediction accuracy.

The article demonstrates and evaluates Graph Transformer Networks, a novel machine learning framework that automatically learns optimal composite relationship paths and generates new network structures directly from data without requiring human preprocessing or predefined domain paths.

The researchers conducted an experimental evaluation across three standard benchmark datasets: two academic citation networks (DBLP and ACM) and one entertainment dataset (IMDB). The model was tested on node classification tasks, such as predicting research topics or movie genres. The evaluation benchmarked the proposed method against conventional random-walk network embedding techniques as well as leading modern graph neural network models, including those that rely on manually designed meta-paths.

The experimental findings show clear performance advantages. First, the proposed framework achieved the highest classification accuracy across all three benchmark datasets, scoring F1 scores of 94.18 on DBLP, 92.68 on ACM, and 60.92 on IMDB, outperforming all baseline models. Second, the model successfully outperformed specialized models that relied on manual domain-expert paths, proving that end-to-end learned structures are more effective than hand-crafted rules. Third, the framework demonstrated adaptive path discovery by effectively identifying both short and long relationship chains across different datasets, as evidenced by an ablation study showing noticeable performance drops when path-shortening mechanisms were removed. Finally, the framework maintained strong interpretability by assigning quantifiable attention scores to composite relationships, successfully discovering meaningful multi-step connections that domain experts had overlooked.

These findings indicate that organizations can eliminate the engineering bottleneck, labor costs, and operational risks associated with manually constructing graph schemas and relational paths. Because the framework learns useful connection patterns directly from heterogeneous data, it automates a critical step in relational data processing while achieving state-of-the-art predictive performance.

Organizations analyzing complex relational data should consider adopting end-to-end learned graph transformation layers over manual path-engineering pipelines. Next steps include validating the framework on operational tasks beyond node classification—specifically link prediction and whole-graph classification—and evaluating integration with other graph neural network architectures.

The primary boundary condition is that the evaluation was conducted on three academic benchmark datasets of moderate scale (up to approximately 18,400 nodes and 68,000 edges). While confidence in the benchmark results is high, stakeholders should conduct pilot validations when scaling to very large, enterprise-grade graphs or when applying the architecture to dynamic, continuously updating networks.

  • Paper: Graph Attention Networks, Petar Veličković et al. (2018). Graph Attention Networks established the foundational masked self-attention mechanism over graph neighbors that Graph Transformer Networks adapt to learn multi-hop composite relations.
  • Paper: Modeling Relational Data with Graph Convolutional Networks, Michael Schlichtkrull et al. (2018). Relational Graph Convolutional Networks introduced message passing across heterogeneous edge relations, motivating the need in GTNs to automatically discover composite meta-paths.
  • Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). This seminal work defined standard graph convolution operations on homogeneous graphs, forming the baseline message-passing scheme that GTN enhances via dynamic graph generation.
  • Paper: Heterogeneous Graph Neural Network, Chuxu Zhang et al. (2019). HetGNN addresses heterogeneous graph representation via type-specific aggregation, providing direct context for the multi-relational challenges GTNs resolve without predefined meta-paths.
  • Paper: Graph Neural Networks: A Review of Methods and Applications, Jie Zhou et al. (2018). This survey provides a comprehensive review of early graph neural network formulations and message-passing paradigms prerequisite to understanding advanced graph transformation models.
Cover for Graph Transformer Networks

Abstract

Graph neural networks (GNNs) have been widely used in representation learning on graphs and achieved state-of-the-art performance in tasks such as node classification and link prediction. However, most existing GNNs are designed to learn node representations on the fixed and homogeneous graphs. The limitations especially become problematic when learning representations on a misspecified graph or a heterogeneous graph that consists of various types of nodes and edges. In this paper, we propose Graph Transformer Networks (GTNs) that are capable of generating new graph structures, which involve identifying useful connections between unconnected nodes on the original graph, while learning effective node representation on the new graphs in an end-to-end fashion. Graph Transformer layer, a core layer of GTNs, learns a soft selection of edge types and composite relations for generating useful multi-hop connections so-called meta-paths. Our experiments show that GTNs learn new graph structures, based on data and tasks without domain knowledge, and yield powerful node representation via convolution on the new graphs. Without domain-specific graph preprocessing, GTNs achieved the best performance in all three benchmark node classification tasks against the state-of-the-art methods that require pre-defined meta-paths from domain knowledge.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Method
  • 3.1 Preliminaries
  • 3.2 Meta-Path Generation
  • 3.3 Graph Transformer Networks
  • 4 Experiments
  • 4.1 Baselines
  • 4.2 Results on Node Classification
  • 4.3 Interpretation of Graph Transformer Networks
  • 5 Conclusion
  • 6 Acknowledgement
  • References

Knowls

  1. Knowl 1 — Graph Transformer Layer for Meta-Path Generation

    model/method

    Given a heterogeneous graph G=(V,E)G = (V, E) with N=∣V∣N = |V| nodes and edge types Te\mathcal{T}^e, the graph structure is represented as a tensor of adjacency matrices A∈RN×N×K\mathbb{A} \in \mathbb{R}^{N \times N \times K}, where K=∣Te∣K = |\mathcal{T}^e| is the number of edge types and Ak∈RN×NA_k \in \mathbb{R}^{N \times N} is the adjacency matrix for the kk-th edge type. A Graph Transformer (GT) layer learns composite relations (meta-paths) by softly selecting adjacency matrices and composing them via matrix multiplication.

    First, soft selection is computed as a convex combination of adjacency matrices using a 1×11 \times 1 convolution ϕ\phi with weight vector Wϕ∈R1×1×KW_\phi \in \mathbb{R}^{1 \times 1 \times K}: Q=F(A;Wϕ)=ϕ(A;softmax(Wϕ))=∑t∈TeαtAtQ = F(\mathbb{A}; W_\phi) = \phi(\mathbb{A}; \text{softmax}(W_\phi)) = \sum_{t \in \mathcal{T}^e} \alpha_t A_t where αt=[softmax(Wϕ)]t\alpha_t = [\text{softmax}(W_\phi)]_t represents the attention weight assigned to edge type tt.

    Second, the GT layer selects two such intermediate adjacency matrices, Q1Q_1 and Q2Q_2, and computes the composite meta-path adjacency matrix via matrix multiplication followed by degree normalization: A(1)=D−1Q1Q2A^{(1)} = D^{-1} Q_1 Q_2 where D∈RN×ND \in \mathbb{R}^{N \times N} is the diagonal degree matrix of Q1Q2Q_1 Q_2 (Dii=∑j(Q1Q2)ijD_{ii} = \sum_j (Q_1 Q_2)_{ij}) used to ensure numerical stability. The resulting matrix A(1)A^{(1)} represents a soft meta-path graph connecting multi-hop neighbors.

  2. Knowl 2 — Variable-Length Meta-Path Learning via Identity Matrix Inclusion

    model/method

    Stacking ll successive Graph Transformer (GT) layers via matrix multiplication inherently produces meta-paths of length l+1l + 1. However, downstream tasks on heterogeneous graphs may require both short (e.g., length-1 or length-2) and long meta-paths simultaneously.

    To allow Graph Transformer Networks to learn meta-paths of arbitrary length up to l+1l + 1, the candidate adjacency set A\mathbb{A} is augmented to include the identity matrix I∈RN×NI \in \mathbb{R}^{N \times N} as an additional relation A0=IA_0 = I. When the soft selection mechanism assigns attention weight to A0A_0, multiplying an intermediate adjacency matrix by II maintains the existing path length without extending it. Consequently, a stack of ll GT layers can adaptively generate any weighted combination of meta-paths with lengths ranging from 11 to l+1l + 1.

  3. Knowl 3 — Graph Transformer Network Architecture

    model/method

    A Graph Transformer Network (GTN) generates multiple meta-path graph structures concurrently and learns node representations by applying graph convolutions over these learned structures.

    For an architecture with CC channels and ll stacked GT layers:

    1. The 1×11 \times 1 convolution produces CC output channels, yielding intermediate candidate tensors Q1(j),Q2(j)∈RN×N×CQ_1^{(j)}, Q_2^{(j)} \in \mathbb{R}^{N \times N \times C} at layer jj. At the first layer (j=1j=1), both Q1(1)Q_1^{(1)} and Q2(1)Q_2^{(1)} are computed from candidate adjacency tensor A\mathbb{A}. For subsequent layers j∈{2,…,l}j \in \{2, \dots, l\}, the intermediate output tensor A(j−1)A^{(j-1)} is reused as Q1(j)Q_1^{(j)}, and multiplied by Q2(j)=F(A;Wϕ(j))Q_2^{(j)} = F(\mathbb{A}; W_\phi^{(j)}).
    2. After ll GT layers, the network produces a meta-path adjacency tensor A(l)∈RN×N×CA^{(l)} \in \mathbb{R}^{N \times N \times C}.
    3. For each channel i∈{1,…,C}i \in \{1, \dots, C\}, self-loops are added: A~i(l)=Ai(l)+I\tilde{A}_i^{(l)} = A_i^{(l)} + I. A Graph Convolutional Network (GCN) layer is applied to each channel with shared parameter matrix W∈RD×dW \in \mathbb{R}^{D \times d}: Z=∥i=1Cσ(D~i−1A~i(l)XW)Z = \Vert_{i=1}^C \sigma\left(\tilde{D}_i^{-1} \tilde{A}_i^{(l)} X W\right) where ∥\Vert denotes feature concatenation across all CC channels, X∈RN×DX \in \mathbb{R}^{N \times D} is the input node feature matrix, D~i\tilde{D}_i is the diagonal degree matrix of A~i(l)\tilde{A}_i^{(l)}, and σ\sigma is an activation function.
    4. The concatenated representation Z∈RN×CdZ \in \mathbb{R}^{N \times Cd} is passed through two fully connected dense layers followed by a softmax function for node classification, trained end-to-end using standard cross-entropy loss on labeled nodes.
  4. Knowl 4 — Quantitative Interpretation of Meta-Path Importance in GTNs

    model/method

    The meta-path adjacency matrix A(l)A^{(l)} generated after ll Graph Transformer layers (considered for a single channel) can be expressed as a weighted sum over all possible meta-path compositions: A(l)=(D(l−1))−1…(D(1))−1(∑t0,t1,…,tl∈Te∪{0}αt0(0)αt1(1)…αtl(l)At0At1…Atl)A^{(l)} = (D^{(l-1)})^{-1} \dots (D^{(1)})^{-1} \left( \sum_{t_0, t_1, \dots, t_l \in \mathcal{T}^e \cup \{0\}} \alpha_{t_0}^{(0)} \alpha_{t_1}^{(1)} \dots \alpha_{t_l}^{(l)} A_{t_0} A_{t_1} \dots A_{t_l} \right) where Te∪{0}\mathcal{T}^e \cup \{0\} is the set of edge types including the identity relation A0=IA_0 = I, D(j)D^{(j)} is the degree normalization matrix at step jj, and αtj(j)=[softmax(Wϕ(j))]tj\alpha_{t_j}^{(j)} = [\text{softmax}(W_\phi^{(j)})]_{t_j} is the attention weight of edge type tjt_j at layer jj.

    The scalar product ∏j=0lαtj(j)\prod_{j=0}^l \alpha_{t_j}^{(j)} acts as a composite attention score quantifying the exact contribution and relative importance of the meta-path sequence (t0,t1,…,tl)(t_0, t_1, \dots, t_l) to the final prediction.

  5. Knowl 5 — Node Classification Performance on Heterogeneous Graphs

    data/table

    Node classification performance measured by Macro F1 score across three heterogeneous graph benchmarks (DBLP, ACM, IMDB). Models compared include homogeneous random walk (DeepWalk), heterogeneous random walk (metapath2vec), homogeneous GNNs (GCN, GAT), heterogeneous GNN with predefined meta-paths (HAN), and Graph Transformer Networks without identity matrix (GTN−IGTN_{-I}) and with identity matrix (GTN):

    Dataset DeepWalk metapath2vec GCN GAT HAN GTN
    DBLP 63.18 85.53 87.30 93.71 92.83 94.18
    ACM 67.42 87.61 91.60 92.33 90.96 92.68
    IMDB 32.08 35.21 56.89 58.14 56.77 60.92

    GTN outperforms all baseline methods on all three datasets without requiring hand-crafted meta-paths. Notably, GTN outperforms HAN (which uses predefined domain-expert meta-paths) and homogeneous GNNs (GCN, GAT) while using only a single GCN layer on top of the learned meta-path graph.

  6. Knowl 6 — Ablation on Identity Matrix in GTNs

    empirical result

    Comparing GTN with GTN−IGTN_{-I} (an identical architecture where the identity matrix II is excluded from candidate adjacency matrices A\mathbb{A}) demonstrates the necessity of learning variable-length meta-paths:

    • On DBLP: GTN scores 94.18 F1 vs 93.91 F1 for GTN−IGTN_{-I}.
    • On ACM: GTN scores 92.68 F1 vs 91.13 F1 for GTN−IGTN_{-I}.
    • On IMDB: GTN scores 60.92 F1 vs 52.33 F1 for GTN−IGTN_{-I} (a degradation of 8.59 F1 points).

    The performance drop is most pronounced on IMDB because a 3-layer GTN−IGTN_{-I} is restricted to constructing meta-paths of fixed length 4 (such as MDMDM), whereas shorter paths (such as MDM of length 2) are substantially more informative for movie genre classification. In full GTN, the model places high attention weight on the identity matrix in deeper layers, enabling it to dynamically learn and preserve shorter meta-paths.

  7. Knowl 7 — Consistency and Discovery of Learned Meta-Paths

    data/table

    Comparison between domain-expert predefined meta-paths and the top-ranked meta-paths learned automatically by GTNs based on path attention scores ∏i=0lαti(i)\prod_{i=0}^l \alpha_{t_i}^{(i)}:

    Dataset Predefined Meta-paths GTN Top-3 Meta-paths (Between Target Nodes)
    DBLP APCPA, APA APCPA, APAPA, APA
    ACM PAP, PSP PAP, PSP
    IMDB MAM, MDM MDM, MAM, MDMDM

    Across DBLP (node types: Author A, Paper P, Conference C), ACM (Paper P, Author A, Subject S), and IMDB (Movie M, Actor A, Director D), GTNs rank human-predefined meta-paths within the top-3 without any prior domain guidance. Additionally, GTNs discover novel effective meta-paths that were not part of the standard predefined sets, such as CPCPA in DBLP (connecting author research areas via conference publication profiles) and MDMDM in IMDB.

  8. Knowl 8 — Heterogeneous Graph Benchmark Dataset Specifications

    experimental setup

    The node classification evaluation uses three heterogeneous graph benchmark datasets:

    • DBLP: 18,405 nodes (Paper P, Author A, Conference C), 67,946 directed edges (PA, AP, PC, CP; 4 edge types), and 334-dimensional bag-of-words keyword features per node. The task is 4-class author research area classification with 800 training, 400 validation, and 2,857 test nodes.
    • ACM: 8,994 nodes (Paper P, Author A, Subject S), 25,922 directed edges (PA, AP, PS, SP; 4 edge types), and 1,902-dimensional bag-of-words keyword features per node. The task is 3-class paper category classification with 600 training, 300 validation, and 2,125 test nodes.
    • IMDB: 12,772 nodes (Movie M, Actor A, Director D), 37,288 directed edges (MA, AM, MD, DM; 4 edge types), and 1,256-dimensional bag-of-words plot features per node. The task is 3-class movie genre classification with 300 training, 300 validation, and 2,339 test nodes.

    All models use embedding dimension d=64d = 64 and are trained using Adam. GTN utilizes 3 GT layers for DBLP and IMDB, and 2 GT layers for ACM, with 1×11 \times 1 convolution weights initialized to a constant value.

Coverage note — None was omitted; all contributed models, mechanisms (GT layer, multi-channel GTN, identity trick, meta-path interpretability), experimental setups, empirical results, and ablation studies from the paper are fully covered.

References

  1. 1.R. v. d. Berg, T. N. Kipf, and M. Welling. Graph convolutional matrix completion. arXiv preprint arXiv:1706.02263, 2017.
  2. 2.S. Bhagat, G. Cormode, and S. Muthukrishnan. Node classification in social networks. In Social network data analytics, pages 115–148. Springer, 2011.
  3. 3.S. Bhagat, G. Cormode, and S. Muthukrishnan. Node classification in social networks. In Social network data analytics, pages 115–148. 2011.
  4. 4.M. M. Bronstein, J. Bruna, Y. LeCun, A. Szlam, and P. Vandergheynst. Geometric deep learning: going beyond euclidean data. IEEE Signal Processing Magazine, 34(4):18–42, 2017.
  5. 5.J. Bruna, W. Zaremba, A. Szlam, and Y. LeCun. Spectral networks and locally connected networks on graphs. arXiv preprint arXiv:1312.6203, 2013.
  6. 6.J. Chen, J. Zhu, and L. Song. Stochastic training of graph convolutional networks with variance reduction. arXiv preprint arXiv:1710.10568, 2017.
  7. 7.J. Chen, T. Ma, and C. Xiao. FastGCN: Fast learning with graph convolutional networks via importance sampling. In International Conference on Learning Representations, 2018.
  8. 8.Y. Chen, Y. Kalantidis, J. Li, S. Yan, and J. Feng. A2\text{A}^2-nets: Double attention networks. In Advances in Neural Information Processing Systems, pages 352–361, 2018.
  9. 9.M. Defferrard, X. Bresson, and P. Vandergheynst. Convolutional neural networks on graphs with fast localized spectral filtering. In Advances in neural information processing systems, pages 3844–3852, 2016.
  10. 10.Y. Dong, N. V. Chawla, and A. Swami. metapath2vec: Scalable representation learning for heterogeneous networks. In KDD '17, pages 135–144, 2017.
  11. 11.D. K. Duvenaud, D. Maclaurin, J. Iparraguirre, R. Bombarell, T. Hirzel, A. Aspuru-Guzik, and R. P. Adams. Convolutional networks on graphs for learning molecular fingerprints. In Advances in neural information processing systems, pages 2224–2232, 2015.
  12. 12.J. Gilmer, S. S. Schoenholz, P. F. Riley, O. Vinyals, and G. E. Dahl. Neural message passing for quantum chemistry. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 1263–1272, 2017.
  13. 13.A. Grover and J. Leskovec. node2vec: Scalable feature learning for networks. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016.
  14. 14.W. L. Hamilton, R. Ying, and J. Leskovec. Inductive representation learning on large graphs. CoRR, abs/1706.02216, 2017.
  15. 15.M. Henaff, J. Bruna, and Y. LeCun. Deep convolutional networks on graph-structured data. CoRR, abs/1506.05163, 2015.
  16. 16.M. Jaderberg, K. Simonyan, A. Zisserman, et al. Spatial transformer networks. In Advances in neural information processing systems, pages 2017–2025, 2015.
  17. 17.R. Kim, C. H. So, M. Jeong, S. Lee, J. Kim, and J. Kang. Hats: A hierarchical graph attention network for stock movement prediction, 2019.
  18. 18.T. N. Kipf and M. Welling. Variational graph auto-encoders. NIPS Workshop on Bayesian Deep Learning, 2016.
  19. 19.T. N. Kipf and M. Welling. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations (ICLR), 2017.
  20. 20.S. I. Ktena, S. Parisot, E. Ferrante, M. Rajchl, M. C. H. Lee, B. Glocker, and D. Rueckert. Distance metric learning using graph convolutional networks: Application to functional brain networks. CoRR, 2017.
  21. 21.J. Lee, I. Lee, and J. Kang. Self-attention graph pooling. CoRR, 2019.
  22. 22.T. Lei, W. Jin, R. Barzilay, and T. Jaakkola. Deriving neural architectures from sequence and graph kernels. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 2024–2033, 2017.
  23. 23.D. Liben-Nowell and J. Kleinberg. The link-prediction problem for social networks. Journal of the American society for information science and technology, 58(7):1019–1031, 2007.
  24. 24.H. Linmei, T. Yang, C. Shi, H. Ji, and X. Li. Heterogeneous graph attention networks for semi-supervised short text classification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2019.
  25. 25.Z. Liu, C. Chen, L. Li, J. Zhou, X. Li, L. Song, and Y. Qi. Geniepath: Graph neural networks with adaptive receptive paths. arXiv preprint arXiv:1802.00910, 2018.
  26. 26.F. Monti, D. Boscaini, J. Masci, E. Rodolà, J. Svoboda, and M. M. Bronstein. Geometric deep learning on graphs and manifolds using mixture model cnns. CoRR, abs/1611.08402, 2016.
  27. 27.F. Monti, M. Bronstein, and X. Bresson. Geometric matrix completion with recurrent multi-graph neural networks. In Advances in Neural Information Processing Systems, pages 3697–3707, 2017.
  28. 28.B. Perozzi, R. Al-Rfou, and S. Skiena. Deepwalk: Online learning of social representations. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 701–710, 2014.
  29. 29.F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini. The graph neural network model. IEEE Transactions on Neural Networks, 2009.
  30. 30.M. Schlichtkrull, T. N. Kipf, P. Bloem, R. Van Den Berg, I. Titov, and M. Welling. Modeling relational data with graph convolutional networks. In European Semantic Web Conference, pages 593–607, 2018.
  31. 31.C. Shi, Y. Li, J. Zhang, Y. Sun, and S. Y. Philip. A survey of heterogeneous information network analysis. IEEE Transactions on Knowledge and Data Engineering, 29(1):17–37, 2016.
  32. 32.J. Tang, M. Qu, M. Wang, M. Zhang, J. Yan, and Q. Mei. Line: Large-scale information network embedding. In WWW, 2015.
  33. 33.P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio. Graph attention networks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=rJXMpikCZ.
  34. 34.S. V. N. Vishwanathan, N. N. Schraudolph, R. Kondor, and K. M. Borgwardt. Graph kernels. Journal of Machine Learning Research, 11(Apr):1201–1242, 2010.
  35. 35.D. Wang, P. Cui, and W. Zhu. Structural deep network embedding. In Proceedings of the 22Nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1225–1234, 2016.
  36. 36.X. Wang, X. He, Y. Cao, M. Liu, and T.-S. Chua. Kgat: Knowledge graph attention network for recommendation. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD), pages 950–958, 2019.
  37. 37.X. Wang, H. Ji, C. Shi, B. Wang, P. Cui, P. Yu, and Y. Ye. Heterogeneous graph attention network. CoRR, abs/1903.07293, 2019.
  38. 38.B. Xu, H. Shen, Q. Cao, Y. Qiu, and X. Cheng. Graph wavelet neural network. In International Conference on Learning Representations, 2019.
  39. 39.R. Ying, R. He, K. Chen, P. Eksombatchai, W. L. Hamilton, and J. Leskovec. Graph convolutional neural networks for web-scale recommender systems. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 974–983, 2018.
  40. 40.R. Ying, J. You, C. Morris, X. Ren, W. L. Hamilton, and J. Leskovec. Hierarchical graph representation learning with differentiable pooling. CoRR, abs/1806.08804, 2018.
  41. 41.C. Zhang, D. Song, C. Huang, A. Swami, and N. V. Chawla. Heterogeneous graph neural network. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD), pages 793–803, 2019.
  42. 42.M. Zhang and Y. Chen. Link prediction based on graph neural networks. In Advances in Neural Information Processing Systems, pages 5165–5175, 2018.
  43. 43.Y. Zhang, Y. Xiong, X. Kong, S. Li, J. Mi, and Y. Zhu. Deep collective classification in heterogeneous information networks. In Proceedings of the 2018 World Wide Web Conference on World Wide Web, pages 399–408, 2018.

Citation

MLA
Yun, S., et al. “Graph Transformer Networks”. arXiv, 2019, http://arxiv.org/abs/1911.06455v2.
APA
Yun, S., Jeong, M., Kim, R., Kang, J., & Kim, H. J. (2019). Graph Transformer Networks. arXiv. http://arxiv.org/abs/1911.06455v2
Chicago
Yun, S., M. Jeong, R. Kim, J. Kang, and H. J. Kim. 2019. “Graph Transformer Networks”. arXiv. http://arxiv.org/abs/1911.06455v2.
Harvard
Yun, S. et al. (2019) “Graph Transformer Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1911.06455v2.
Vancouver
1. Yun S, Jeong M, Kim R, Kang J, Kim HJ (2019) Graph Transformer Networks. arXiv

BibTeX

@article{yun2019graph,
  title = {Graph Transformer Networks},
  author = {Yun, Seongjun and Jeong, Minbyul and Kim, Raehyun and Kang, Jaewoo and Kim, Hyunwoo J.},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1911.06455v2},
  eprint = {1911.06455}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors