Heterogeneous Graph Neural Network

Chuxu ZhangDongjin SongChao HuangA. SwamiN. Chawla

article2019KDD1,662 citations

Proposes HetGNN, a heterogeneous graph neural network architecture that integrates multimodal node contents with complex structural topologies through restart-based random walk sampling and hierarchical type-aware feature aggregation.

Listen

Modern data systems rely heavily on complex networks containing diverse entity types and relationships, such as academic networks connecting authors, papers, and venues, or e-commerce platforms linking users, items, and reviews. These networks also carry unstructured, multimodal information, including text descriptions, images, and user attributes. Effectively generating mathematical representations, or embeddings, for these entities is critical for downstream automated tasks like product recommendation, relationship inference, and categorization. However, existing techniques struggle because they either fail to handle multiple node types simultaneously, cannot deeply combine diverse content types, or treat all neighboring connections with equal importance.

To address these limitations, the article presents HetGNN, a heterogeneous graph neural network architecture. HetGNN evaluates how effectively a unified deep learning model can simultaneously encode complex multimodal content and heterogeneous structural relationships across both existing nodes and previously unseen entities.

The authors designed a three-part technical framework: first, a guided random-walk sampling strategy selects a fixed size of strongly correlated neighboring entities and groups them by entity type. Second, a bidirectional recurrent neural network encodes the deep interactions of multimodal content associated with each node. Third, another recurrent neural network aggregates neighbor embeddings by type, followed by an attention mechanism that assigns dynamic importance weights to different neighbor categories before optimizing via a graph context loss. The framework was evaluated across four large-scale real-world datasets—two academic datasets from AMiner spanning 1996 to 2015 and two e-commerce datasets from Amazon covering movies and CDs spanning 1996 to 2014—against five established baseline methods.

Across all evaluated tasks, HetGNN consistently demonstrated superior performance. In link prediction, HetGNN improved prediction accuracy over the strongest baselines by 1.5% to 5.6% on academic networks and by 3.4% to 10.5% on e-commerce review networks. In personalized recommendation benchmarks, the framework achieved performance gains ranging from 2.8% to 16.0% compared to competing models. Furthermore, when tested on inductive clustering tasks involving newly introduced nodes, HetGNN outperformed GraphSAGE by an average of 17.3% and Graph Attention Networks by 10.6% in clustering quality metrics. Ablation studies confirmed that bidirectional recurrent content encoding and attention-based neighbor weighting were essential contributors to these gains.

These findings demonstrate that organizations managing complex relational databases can significantly improve recommendation quality and entity classification by moving beyond shallow attribute concatenation. Implementing structured neighborhood sampling alongside deep multimodal encoding directly enhances automated decision systems, reducing the manual engineering effort required to extract domain-specific features. The framework also mitigates risks associated with data sparsity and 'cold-start' entities by reliably predicting properties for newly added users or products.

Organizations seeking to upgrade their graph analytics and recommendation pipelines should consider adopting heterogeneous aggregation architectures like HetGNN. When deploying such architectures, practitioners should tune the sampled neighbor size to moderate ranges, as the empirical analysis showed performance peaks between 20 and 30 sampled neighbors; exceeding this threshold introduces noise and degrades model accuracy. Similarly, embedding dimensions should be balanced around 128 to 256 dimensions to avoid overfitting.

While the empirical results provide high confidence in the model's effectiveness across academic and consumer review domains, the framework relies on multi-step pre-training pipelines for text and image features, which adds computational overhead. Practitioners should assess whether their operational infrastructure can accommodate the training requirements of bidirectional recurrent networks before enterprise-wide deployment.

  • Paper: Heterogeneous Graph Transformer, Ziniu Hu et al. (2020). This work advances heterogeneous graph representation learning beyond RNN- and meta-path-based aggregations like HetGNN by introducing a dedicated Heterogeneous Graph Transformer architecture.
  • Paper: Graph Neural Networks in Recommender Systems: A Survey, Shiwen Wu et al. (2020). This survey provides a comprehensive synthesis of graph neural network methodologies across recommendation domains, contextualizing architectures like HetGNN within broader industry applications.
  • Paper: LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation, Xiangnan He et al. (2020). This paper examines and simplifies graph neural architectures for recommendation, providing critical insights into the necessity of complex feature transformations introduced in models like HetGNN.
  • Paper: Self-supervised Graph Learning for Recommendation, Jiancan Wu et al. (2020). This study extends graph-based recommendation systems to self-supervised learning frameworks, addressing the data sparsity and cold-start problems tackled by HetGNN.
  • Paper: How Attentive are Graph Attention Networks?, Shaked Brody et al. (2021). This paper investigates the theoretical expressiveness and limitations of attention mechanisms in graph neural networks, building directly on the attention mechanisms used in HetGNN.
Cover for Heterogeneous Graph Neural Network

Abstract

Representation learning in heterogeneous graphs aims to pursue a meaningful vector representation for each node so as to facilitate downstream applications such as link prediction, personalized recommendation, node classification, etc. This task, however, is challenging not only because of the demand to incorporate heterogeneous structural (graph) information consisting of multiple types of nodes and edges, but also due to the need for considering heterogeneous attributes or contents (e.g., text or image) associated with each node. Despite a substantial amount of effort has been made to homogeneous (or heterogeneous) graph embedding, attributed graph embedding as well as graph neural networks, few of them can jointly consider heterogeneous structural (graph) information as well as heterogeneous contents information of each node effectively. In this paper, we propose HetGNN, a heterogeneous graph neural network model, to resolve this issue. Specifically, we first introduce a random walk with restart strategy to sample a fixed size of strongly correlated heterogeneous neighbors for each node and group them based upon node types. Next, we design a neural network architecture with two modules to aggregate feature information of those sampled neighboring nodes. The first module encodes "deep" feature interactions of heterogeneous contents and generates content embedding for each node. The second module aggregates content (attribute) embeddings of different neighboring groups (types) and further combines them by considering the impacts of different groups to obtain the optimal node embedding. Finally, we leverage a graph context loss and a mini-batch gradient descent procedure to train the model in an end-to-end manner. Extensive experiments on several datasets demonstrate that HetGNN can outperform state-of-the-art baselines in various graph mining tasks, i.e., link prediction, recommendation, node classification & clustering and inductive node classification & clustering.

Table of Contents

  • 1 INTRODUCTION
  • 2 PROBLEM DEFINITION
  • 3 HetGNN
  • 3.1 Sampling Heterogeneous Neighbors (C1)
  • 3.2 Encoding Heterogeneous Contents (C2)
  • 3.3 Aggregating Heterogeneous Neighbors (C3)
  • 3.3.1 Same Type Neighbors Aggregation.
  • 3.3.2 Types Combination.
  • 3.4 Objective and Model Training
  • 4 EXPERIMENTS
  • 4.1 Experiment Design
  • 4.1.1 Datasets.
  • 4.1.2 Baselines.
  • 4.1.3 Reproducibility.
  • 4.2 Applications
  • 4.2.1 Link Prediction (RQ1-1).
  • 4.2.2 Recommendation (RQ1-2).
  • 4.3 Analysis
  • 4.3.1 Ablation Study (RQ3).
  • 4.3.2 Hyper-parameters Sensitivity (RQ4).
  • 5 RELATED WORK
  • 6 CONCLUSION
  • ACKNOWLEDGMENTS
  • REFERENCES
  • A SUPPLEMENT
  • A.1 Pseudocode of HetGNN Training Procedure
  • A.2 Dataset Description
  • A.3 Baseline Description
  • A.4 Reproducibility Settings
  • A.5 Model Variants Description
  • A.6 Hyper-parameters Sensitivity Setup

Knowls

  1. Knowl 1 — Content-Associated Heterogeneous Graphs and Representation Learning

    definition

    A content-associated heterogeneous graph (C-HetG) is defined as a graph G=(V,E,OV,RE)G = (V, E, O_V, R_E), where VV is the set of nodes, EE is the set of edges, OVO_V denotes the set of node object types, and RER_E denotes the set of relation types, satisfying ∣OV∣+∣RE∣≥3|O_V| + |R_E| \ge 3. Each node v∈Vv \in V is associated with an unordered set of heterogeneous contents CvC_v, which may include structured attributes, textual content, or image data.

    The heterogeneous graph representation learning task seeks to learn a mapping function FΘF_\Theta parameterized by Θ\Theta that maps each node v∈Vv \in V to a dd-dimensional embedding vector Ev∈Rd\mathcal{E}_v \in \mathbb{R}^d (d≪∣V∣d \ll |V|), such that the embeddings preserve both the heterogeneous graph topological structure and the heterogeneous node content features.

  2. Knowl 2 — Heterogeneous Neighbor Sampling and Grouping via Random Walk with Restart

    model/method

    To address varying node degrees, lack of direct connections between arbitrary node types, and feature dimension heterogeneity, HetGNN extracts fixed-size, type-specific neighbor sets for each node v∈Vv \in V using Random Walk with Restart (RWR):

    1. Fixed-Length RWR Sampling: A random walk starts from node vv. At each step, the walk traverses to a randomly chosen neighbor of the current node with probability 1−p1 - p, or restarts at the root node vv with probability pp (set to p=0.5p = 0.5). The walk continues until a fixed total number of node visits is collected, denoted as RWR(v)\text{RWR}(v) (with walk length set to 100). The sampling process ensures that all node types in OVO_V are visited.

    2. Type-Based Grouping: For each node type t∈OVt \in O_V, the top ktk_t most frequently visited nodes in RWR(v)\text{RWR}(v) are selected to form the tt-type correlated neighbor set Nt(v)N_t(v) of node vv. For academic graphs with author, paper, and venue types, the group sizes are kauthor=10k_{\text{author}} = 10, kpaper=10k_{\text{paper}} = 10, and kvenue=3k_{\text{venue}} = 3. For user-item review graphs, the group sizes are kuser=10k_{\text{user}} = 10 and kitem=10k_{\text{item}} = 10.

  3. Knowl 3 — Bi-LSTM-Based Heterogeneous Content Encoder

    model/method

    For each node v∈Vv \in V with heterogeneous content set CvC_v, let xi∈Rdf×1x_i \in \mathbb{R}^{d_f \times 1} denote the pre-trained feature vector of the ii-th content item in CvC_v (e.g., Paragraph Vector embeddings for text, CNN features for images, or pre-trained node embeddings from DeepWalk). Each feature is first projected into a common subspace via a type-specific fully connected layer FCθx(xi)FC_{\theta_x}(x_i).

    To capture non-linear interactions across diverse content modalities, an unordered sequence of projected features is processed by a bidirectional LSTM (Bi-LSTM), and the hidden states are averaged across content items:

    f1(v)=∑i∈Cv[LSTM→(FCθx(xi))⊕LSTM←(FCθx(xi))]∣Cv∣f_1(v) = \frac{\sum_{i \in C_v} \left[ \overrightarrow{\text{LSTM}}\left(FC_{\theta_x}(x_i)\right) \oplus \overleftarrow{\text{LSTM}}\left(FC_{\theta_x}(x_i)\right) \right]}{|C_v|}

    where f1(v)∈Rd×1f_1(v) \in \mathbb{R}^{d \times 1} is the content embedding of node vv, ⊕\oplus denotes vector concatenation, and the forward and backward LSTM hidden state dimension is d/2d/2.

    The forward LSTM cell updates for content feature index ii are given by:

    zi=σ(UzFCθx(xi)+Wzhi−1+bz)fi=σ(UfFCθx(xi)+Wfhi−1+bf)oi=σ(UoFCθx(xi)+Wohi−1+bo)c^i=tanh⁡(UcFCθx(xi)+Wchi−1+bc)ci=fi∘ci−1+zi∘c^ihi=tanh⁡(ci)∘oi\begin{aligned} z_i &= \sigma(U_z FC_{\theta_x}(x_i) + W_z h_{i-1} + b_z) \\ f_i &= \sigma(U_f FC_{\theta_x}(x_i) + W_f h_{i-1} + b_f) \\ o_i &= \sigma(U_o FC_{\theta_x}(x_i) + W_o h_{i-1} + b_o) \\ \hat{c}_i &= \tanh(U_c FC_{\theta_x}(x_i) + W_c h_{i-1} + b_c) \\ c_i &= f_i \circ c_{i-1} + z_i \circ \hat{c}_i \\ h_i &= \tanh(c_i) \circ o_i \end{aligned}

    where hi∈R(d/2)×1h_i \in \mathbb{R}^{(d/2) \times 1}, Uj∈R(d/2)×dfU_j \in \mathbb{R}^{(d/2) \times d_f}, Wj∈R(d/2)×(d/2)W_j \in \mathbb{R}^{(d/2) \times (d/2)}, bj∈R(d/2)×1b_j \in \mathbb{R}^{(d/2) \times 1} for j∈{z,f,o,c}j \in \{z, f, o, c\}, σ(⋅)\sigma(\cdot) is the sigmoid function, and ∘\circ denotes element-wise (Hadamard) product.

  4. Knowl 4 — Same-Type Neighbor Aggregation via Bi-LSTM

    model/method

    For each node type t∈OVt \in O_V, the content embeddings f1(v′)f_1(v') of all sampled tt-type neighbors v′∈Nt(v)v' \in N_t(v) are aggregated to generate a type-specific neighborhood embedding f2t(v)∈Rd×1f_2^t(v) \in \mathbb{R}^{d \times 1} using a type-dedicated bidirectional LSTM:

    f2t(v)=∑v′∈Nt(v)[LSTM→t(f1(v′))⊕LSTM←t(f1(v′))]∣Nt(v)∣f_2^t(v) = \frac{\sum_{v' \in N_t(v)} \left[ \overrightarrow{\text{LSTM}}_t(f_1(v')) \oplus \overleftarrow{\text{LSTM}}_t(f_1(v')) \right]}{|N_t(v)|}

    where f1(v′)∈Rd×1f_1(v') \in \mathbb{R}^{d \times 1} is the content embedding of neighbor node v′v', ⊕\oplus is concatenation, and the hidden state dimensionality in each direction is d/2d/2. Distinct Bi-LSTM parameters are maintained for each node type t∈OVt \in O_V to accommodate differences across heterogeneous node types.

  5. Knowl 5 — Attention-Based Heterogeneous Node Types Combination

    equation

    To combine the node's self-content embedding f1(v)f_1(v) with the type-specific neighborhood embeddings {f2t(v)∣t∈OV}\{f_2^t(v) \mid t \in O_V\}, an attention mechanism assigns adaptive weights to the embedding set F(v)={f1(v)}∪{f2t(v)∣t∈OV}\mathcal{F}(v) = \{f_1(v)\} \cup \{f_2^t(v) \mid t \in O_V\}:

    Ev=∑fi∈F(v)αv,ifi\mathcal{E}_v = \sum_{f_i \in \mathcal{F}(v)} \alpha_{v,i} f_i

    αv,i=exp⁡(LeakyReLU(uT[fi⊕f1(v)]))∑fj∈F(v)exp⁡(LeakyReLU(uT[fj⊕f1(v)]))\alpha_{v,i} = \frac{\exp\left(\text{LeakyReLU}\left(u^T [f_i \oplus f_1(v)]\right)\right)}{\sum_{f_j \in \mathcal{F}(v)} \exp\left(\text{LeakyReLU}\left(u^T [f_j \oplus f_1(v)]\right)\right)}

    where Ev∈Rd×1\mathcal{E}_v \in \mathbb{R}^{d \times 1} is the final output representation of node vv, u∈R2d×1u \in \mathbb{R}^{2d \times 1} is a learnable attention parameter vector, ⊕\oplus denotes concatenation, and αv,i\alpha_{v,i} represents the normalized attention weight measuring the relative importance of embedding fif_i to node vv.

  6. Knowl 6 — Graph Context Loss and Optimization with Negative Sampling

    equation

    HetGNN is trained by maximizing the co-occurrence probability of nodes within a local graph context window across all node types OVO_V:

    o1=arg⁡max⁡Θ∏v∈V∏t∈OV∏vc∈CNvtp(vc∣v;Θ)o_1 = \arg \max_\Theta \prod_{v \in V} \prod_{t \in O_V} \prod_{v_c \in CN_v^t} p(v_c \mid v; \Theta)

    where CNvtCN_v^t denotes the set of tt-type context nodes appearing within a window of distance τ=5\tau = 5 along random walks rooted at vv, and p(vc∣v;Θ)=exp⁡(Evc⋅Ev)∑vk∈Vtexp⁡(Evk⋅Ev)p(v_c \mid v; \Theta) = \frac{\exp(\mathcal{E}_{v_c} \cdot \mathcal{E}_v)}{\sum_{v_k \in V_t} \exp(\mathcal{E}_{v_k} \cdot \mathcal{E}_v)}.

    Using negative sampling with sample size M=1M = 1, the objective is reformulated over a set of sampled triplets Twalk={⟨v,vc,vc′⟩}T_{\text{walk}} = \{\langle v, v_c, v_c' \rangle\} as:

    o2=∑⟨v,vc,vc′⟩∈Twalk[log⁡σ(Evc⋅Ev)+log⁡σ(−Evc′⋅Ev)]o_2 = \sum_{\langle v, v_c, v_c' \rangle \in T_{\text{walk}}} \left[ \log \sigma(\mathcal{E}_{v_c} \cdot \mathcal{E}_v) + \log \sigma(-\mathcal{E}_{v_c'} \cdot \mathcal{E}_v) \right]

    where σ(x)=11+exp⁡(−x)\sigma(x) = \frac{1}{1 + \exp(-x)}, and for each positive context node vcv_c of type tt, a negative node vc′v_c' of the same type tt is sampled from the noise distribution Pt(vc′)∝(dg,vc′)3/4P_t(v_c') \propto (d_{g, v_c'})^{3/4}, where dg,vc′d_{g, v_c'} is the visit frequency of vc′v_c' in the sampled walks.

  7. Knowl 7 — HetGNN End-to-End Training Algorithm

    algorithm
    Input: Pre-trained content features for each node v∈Vv \in V, triplet set TwalkT_{walk} sampled from random walks on graph GG
    Output: Optimized model parameters Θ\Theta for inferring node embeddings E\mathcal{E}
    while stopping criterion is not met do
        sample a mini-batch of triplets (v,vc,vc′)(v, v_c, v_c') from TwalkT_{walk}
        for each node u∈{v,vc,vc′}u \in \{v, v_c, v_c'\} do
            encode content embedding f1(u)f_1(u) using the content Bi-LSTM
            for each node type t∈OVt \in O_V do
                sample top ktk_t neighbors Nt(u)N_t(u) via RWR
                aggregate tt-type neighbor embeddings f2t(u)f_2^t(u) using type Bi-LSTM
            end for
            combine f1(u)f_1(u) and {f2t(u)}\{f_2^t(u)\} via attention mechanism to obtain Eu\mathcal{E}_u
        end for
        accumulate context loss: o2=∑(v,vc,vc′)[log⁡σ(Evc⋅Ev)+log⁡σ(−Evc′⋅Ev)]o_2 = \sum_{(v, v_c, v_c')} [\log \sigma(\mathcal{E}_{v_c} \cdot \mathcal{E}_v) + \log \sigma(-\mathcal{E}_{v_c'} \cdot \mathcal{E}_v)]
        update all neural network parameters Θ\Theta via Adam optimizer
    end while
    return optimized parameters Θ\Theta

    The training algorithm samples triplets ⟨v,vc,vc′⟩\langle v, v_c, v_c' \rangle generated from 10 random walks per node with walk length 30 and context window τ=5\tau = 5. Parameters Θ\Theta (including fully connected projection layers, content Bi-LSTMs, neighbor aggregation Bi-LSTMs, and attention vectors) are optimized end-to-end via the Adam optimizer until convergence.

  8. Knowl 8 — Transductive Link Prediction Performance Across Academic and Review Heterogeneous Graphs

    data/table

    Link prediction was evaluated using sequential temporal splits on academic data (AMiner A-I and A-II) and sequential transaction splits on Amazon review data (Movies R-I and CDs R-II). In academic graphs, Type-1 links denote co-authorship and Type-2 links denote author-paper citation. In review graphs, links denote user-item reviews. Binary logistic classifiers were trained on element-wise products of learned node embeddings and evaluated against an equal number of negative links.

    Dataset Split Metric MP2V ASNE SHNE GSAGE GAT HetGNN
    A-I 2003 (Type-1) AUC 0.636 0.683 0.696 0.694 0.701 0.714
    F1 0.435 0.584 0.597 0.586 0.606 0.620
    A-I 2003 (Type-2) AUC 0.790 0.794 0.781 0.790 0.821 0.837
    F1 0.743 0.774 0.755 0.746 0.792 0.815
    A-I 2002 (Type-1) AUC 0.626 0.667 0.688 0.681 0.691 0.710
    F1 0.412 0.554 0.590 0.567 0.589 0.615
    A-I 2002 (Type-2) AUC 0.808 0.782 0.795 0.806 0.837 0.851
    F1 0.770 0.753 0.761 0.772 0.816 0.828
    A-II 2013 (Type-1) AUC 0.596 0.689 0.683 0.695 0.678 0.717
    F1 0.348 0.643 0.639 0.615 0.613 0.669
    A-II 2013 (Type-2) AUC 0.712 0.721 0.695 0.714 0.732 0.767
    F1 0.647 0.713 0.674 0.664 0.705 0.754
    A-II 2012 (Type-1) AUC 0.586 0.671 0.672 0.676 0.655 0.701
    F1 0.318 0.615 0.612 0.573 0.560 0.642
    A-II 2012 (Type-2) AUC 0.724 0.726 0.706 0.739 0.750 0.775
    F1 0.664 0.737 0.692 0.706 0.715 0.757
    R-I 5:5 AUC 0.634 0.623 0.651 0.661 0.683 0.749
    F1 0.445 0.551 0.586 0.542 0.665 0.735
    R-I 7:3 AUC 0.701 0.656 0.695 0.716 0.706 0.787
    F1 0.595 0.613 0.660 0.688 0.702 0.776
    R-II 5:5 AUC 0.678 0.655 0.685 0.677 0.712 0.736
    F1 0.541 0.582 0.593 0.565 0.659 0.701
    R-II 7:3 AUC 0.737 0.695 0.728 0.721 0.742 0.772
    F1 0.660 0.648 0.685 0.653 0.713 0.749

    HetGNN consistently outperforms all baseline models, yielding relative improvements over the best-performing baselines of 1.5%–5.6% on academic graphs and 3.4%–10.5% on review graphs.

  9. Knowl 9 — Transductive and Inductive Node Classification and Clustering Performance

    empirical result

    HetGNN was evaluated on author node classification (logistic regression with 10% and 30% training ratios) and clustering (kk-means) across four research domains (Data Mining, Computer Vision, Natural Language Processing, and Database) on the AMiner A-II dataset, under both transductive and inductive settings (where test nodes were not present during training):

    1. Transductive Multi-label Classification (MC) and Node Clustering (NC):

      • MC (10% training): HetGNN achieves 0.978 Macro-F1 and 0.979 Micro-F1 (matching or exceeding GraphSAGE at 0.978 / 0.978 and GAT at 0.962 / 0.963).
      • MC (30% training): HetGNN achieves 0.981 Macro-F1 and 0.982 Micro-F1 (compared to GraphSAGE at 0.979 / 0.980 and GAT at 0.965 / 0.965).
      • NC: HetGNN obtains 0.901 NMI and 0.932 ARI (competitive with GraphSAGE at 0.914 NMI and 0.945 ARI, and outperforming GAT at 0.845 NMI and 0.882 ARI).
    2. Inductive Multi-label Classification (IMC) and Inductive Clustering (INC):

      • IMC (10% training): HetGNN reaches 0.962 Macro-F1 and 0.965 Micro-F1, versus GraphSAGE (0.938 / 0.945) and GAT (0.954 / 0.958).
      • IMC (30% training): HetGNN reaches 0.964 Macro-F1 and 0.968 Micro-F1, versus GraphSAGE (0.949 / 0.955) and GAT (0.956 / 0.960).
      • INC: HetGNN achieves 0.840 NMI and 0.894 ARI, substantially outperforming GraphSAGE (0.714 NMI / 0.764 ARI) and GAT (0.765 NMI / 0.803 ARI), with average relative improvements of 17.3% over GraphSAGE and 10.6% over GAT.
  10. Knowl 10 — Ablation Study and Hyperparameter Sensitivity in HetGNN

    empirical result

    Experimental analyses on the AMiner A-II dataset (2013 split) evaluate the contribution of individual architecture components and sensitivity to hyperparameters:

    1. Ablation of Architectural Components:

      • No-Neigh (removing neighborhood aggregation and relying solely on node content encoding f1(v)f_1(v)) causes substantial performance drops across link prediction and venue recommendation, confirming the necessity of structural neighborhood aggregation.
      • Content-FC (replacing the content Bi-LSTM encoder with a fully connected layer) performs worse than HetGNN, demonstrating that Bi-LSTM captures richer deep non-linear interactions across heterogeneous content modalities.
      • Type-FC (replacing the self-attention combination layer with a fully connected layer across type embeddings) is outperformed by HetGNN, verifying the benefit of attention-weighted combination of different neighbor types.
    2. Hyperparameter Sensitivity:

      • Embedding Dimension dd: As dd increases from 8 to 128, AUC, F1, Recall, and Precision increase consistently; performance stabilizes between d=128d = 128 and d=256d = 256, with slight degradation beyond 256 due to overfitting.
      • Sampled Neighbor Size: Evaluating neighbor set sizes from 6 to 34 (summed across author, paper, and venue partitions) shows initial performance gains as neighborhood context increases, peaking between 20 and 30 total neighbors (specifically 23: 10 authors, 10 papers, 3 venues). Beyond 30 neighbors, performance decays slightly due to the inclusion of weakly correlated noisy nodes.

Coverage note — Qualitative embedding projector visualizations (Figure 3) were omitted because their findings are quantitatively captured in the clustering and classification results.

References

  1. 1.Shiyu Chang, Wei Han, Jiliang Tang, Guo-Jun Qi, Charu C Aggarwal, and Thomas S Huang. 2015. Heterogeneous network embedding via deep architectures. In KDD. 119–128.
  2. 2.Ting Chen and Yizhou Sun. 2017. Task-Guided and Path-Augmented Heterogeneous Network Embedding for Author Identification. In WSDM. 295–304.
  3. 3.Peng Cui, Xiao Wang, Jian Pei, and Wenwu Zhu. 2018. A survey on network embedding. TKDE (2018).
  4. 4.Yuxiao Dong, Nitesh V Chawla, and Ananthram Swami. 2017. metapath2vec: Scalable Representation Learning for Heterogeneous Networks. In KDD. 135–144.
  5. 5.Hongyang Gao, Zhengyang Wang, and Shuiwang Ji. 2018. Large-scale learnable graph convolutional networks. In KDD. 1416–1424.
  6. 6.Aditya Grover and Jure Leskovec. 2016. node2vec: Scalable feature learning for networks. In KDD. 855–864.
  7. 7.Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. In NIPS. 1024–1034.
  8. 8.Ruining He and Julian McAuley. 2016. Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering. In WWW. 507–517.
  9. 9.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation 9, 8 (1997), 1735–1780.
  10. 10.Binbin Hu, Chuan Shi, Wayne Xin Zhao, and Philip S Yu. 2018. Leveraging meta-path based context for top-n recommendation with a neural co-attention model. In KDD. 1531–1540.
  11. 11.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014).
  12. 12.Thomas N Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In ICLR.
  13. 13.Quoc Le and Tomas Mikolov. 2014. Distributed representations of sentences and documents. In ICML. 1188–1196.
  14. 14.Jundong Li, Harsh Dani, Xia Hu, Jiliang Tang, Yi Chang, and Huan Liu. 2017. Attributed network embedding for learning in a dynamic environment. In CIKM. 387–396.
  15. 15.Lizi Liao, Xiangnan He, Hanwang Zhang, and Tat-Seng Chua. 2018. Attributed social network embedding. TKDE 30, 12 (2018), 2257–2270.
  16. 16.Ziqi Liu, Chaochao Chen, Xinxing Yang, Jun Zhou, Xiaolong Li, and Le Song. 2018. Heterogeneous Graph Neural Networks for Malicious Account Detection. In CIKM. 2077–2085.
  17. 17.Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015. Fully convolutional networks for semantic segmentation. In CVPR. 3431–3440.
  18. 18.Jianxin Ma, Peng Cui, Xiao Wang, and Wenwu Zhu. 2018. Hierarchical Taxonomy Aware Network Embedding. In KDD. 1920–1929.
  19. 19.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In NIPS. 3111–3119.
  20. 20.Bryan Perozzi, Rami Al-Rfou, and Steven Skiena. 2014. Deepwalk: Online learning of social representations. In KDD. 701–710.
  21. 21.Jiezhong Qiu, Yuxiao Dong, Hao Ma, Jian Li, Kuansan Wang, and Jie Tang. 2018. Network embedding as matrix factorization: Unifying deepwalk, line, pte, and node2vec. In WSDM. 459–467.
  22. 22.Meng Qu, Jian Tang, and Jiawei Han. 2018. Curriculum Learning for Heterogeneous Star Network Embedding via Deep Reinforcement Learning. In WSDM. 468–476.
  23. 23.Xiang Ren, Jialu Liu, Xiao Yu, Urvashi Khandelwal, Quanquan Gu, Lidan Wang, and Jiawei Han. 2014. Cluscite: Effective citation recommendation by information network-based clustering. In KDD. 821–830.
  24. 24.Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne Van Den Berg, Ivan Titov, and Max Welling. 2018. Modeling relational data with graph convolutional networks. In ESWC. 593–607.
  25. 25.Yizhou Sun, Jiawei Han, Charu C Aggarwal, and Nitesh V Chawla. 2012. When will it happen?: relationship prediction in heterogeneous information networks. In WSDM. 663–672.
  26. 26.Yizhou Sun, Jiawei Han, Xifeng Yan, Philip S Yu, and Tianyi Wu. 2011. Pathsim: Meta path-based top-k similarity search in heterogeneous information networks. VLDB 4, 11 (2011), 992–1003.
  27. 27.Yizhou Sun, Brandon Norick, Jaiwei Han, Xifeng Yan, Philip Yu, and Xiao Yu. 2012. PathSelClus: Integrating Meta-Path Selection with User-Guided Object Clustering in Heterogeneous Information Networks. In KDD. 1348–1356.
  28. 28.Jian Tang, Meng Qu, and Qiaozhu Mei. 2015. Pte: Predictive text embedding through large-scale heterogeneous text networks. In KDD. 1165–1174.
  29. 29.Jian Tang, Meng Qu, Mingzhe Wang, Ming Zhang, Jun Yan, and Qiaozhu Mei. 2015. Line: Large-scale information network embedding. In WWW. 1067–1077.
  30. 30.Jie Tang, Jing Zhang, Limin Yao, Juanzi Li, Li Zhang, and Zhong Su. 2008. Arnetminer: extraction and mining of academic social networks. In KDD. 990–998.
  31. 31.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2018. Graph attention networks. In ICLR.
  32. 32.Wenchao Yu, Cheng Zheng, Wei Cheng, Charu C Aggarwal, Dongjin Song, Bo Zong, Haifeng Chen, and Wei Wang. 2018. Learning Deep Network Representations with Adversarially Regularized Autoencoders. In KDD. 2663–2671.
  33. 33.Chuxu Zhang, Chao Huang, Lu Yu, Xiangliang Zhang, and Nitesh V Chawla. 2018. Camel: Content-Aware and Meta-path Augmented Metric Learning for Author Identification. In WWW. 709–718.
  34. 34.Chuxu Zhang, Ananthram Swami, and Nitesh V Chawla. 2019. SHNE: Representation Learning for Semantic-Associated Heterogeneous Networks. In WSDM. 690–698.
  35. 35.Chuxu Zhang, Lu Yu, Xiangliang Zhang, and Nitesh V Chawla. 2018. Task-Guided and Semantic-Aware Ranking for Academic Author-Paper Correlation Inference.. In IJCAI. 3641–3647.
  36. 36.Yizhou Zhang, Yun Xiong, Xiangnan Kong, Shanshan Li, Jinhong Mi, and Yangyong Zhu. 2018. Deep Collective Classification in Heterogeneous Information Networks. In WWW. 399–408.

Citation

MLA
Zhang, C., et al. “Heterogeneous Graph Neural Network”. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019, pp. 793–803, https://doi.org/10.1145/3292500.3330961.
APA
Zhang, C., Song, D., Huang, C., Swami, A., & Chawla, N. V. (2019). Heterogeneous Graph Neural Network. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 793–803. https://doi.org/10.1145/3292500.3330961
Chicago
Zhang, C., D. Song, C. Huang, A. Swami, and N. V. Chawla. 2019. “Heterogeneous Graph Neural Network”. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 793–803. https://doi.org/10.1145/3292500.3330961.
Harvard
Zhang, C. et al. (2019) “Heterogeneous Graph Neural Network”, Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. ACM, pp. 793–803. Available at: https://doi.org/10.1145/3292500.3330961.
Vancouver
1. Zhang C, Song D, Huang C, Swami A, Chawla NV (2019) Heterogeneous Graph Neural Network. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. ACM, pp 793–803

BibTeX

@inproceedings{Zhang_2019, series={KDD ’19}, title={Heterogeneous Graph Neural Network}, url={http://dx.doi.org/10.1145/3292500.3330961}, DOI={10.1145/3292500.3330961}, booktitle={Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining}, publisher={ACM}, author={Zhang, Chuxu and Song, Dongjin and Huang, Chao and Swami, Ananthram and Chawla, Nitesh V.}, year={2019}, month=July, pages={793–803}, collection={KDD ’19} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF