Attribute and Structure Preserving Graph Contrastive Learning

Jialu ChenGang Kou

article2023AAAI72 citations

Proposes an attribute and structure preserving graph contrastive learning framework that overcomes standard homophily assumptions by jointly contrasting original, attribute similarity, and higher-order structural views.

Listen

Organizations increasingly rely on automated graph analysis to extract insights from interconnected systems such as citation networks, web links, and social interactions. A persistent challenge in this domain is the scarcity of labeled training data, which has led to widespread adoption of self-supervised learning techniques that learn representations without human annotation. However, prevailing graph learning methods depend heavily on the assumption of homophily—the expectation that connected entities share similar characteristics or labels. When applied to real-world networks where connected items are dissimilar (heterophily), standard models often fail to capture critical relationships or lose essential node-level characteristics.

The article introduces and evaluates a self-supervised framework called Attribute and Structure Preserving Graph Contrastive Learning (ASP). The primary objective is to demonstrate that simultaneously preserving feature similarities and broader topological patterns allows an unsupervised model to generate high-performing node representations across networks regardless of their homophily level.

To evaluate the system, the authors conducted empirical experiments on seven real-world network datasets, encompassing three standard homophilous citation datasets and four non-homophilous web and co-occurrence datasets. The framework constructs three distinct views of the network data—an original graph view, an attribute similarity view built via nearest neighbors, and a higher-order structure view capturing multi-hop connections. ASP contrasts these views using a simplified graph convolutional encoder, combining node features to bridge wide disparities between graph topology and attributes, and aligns the representations through a cross-module objective. The learned representations were tested on node classification tasks against multiple state-of-the-art supervised and self-supervised models.

The experimental findings demonstrate significant performance advantages. First, ASP consistently outperformed all tested supervised and unsupervised baselines across all four non-homophilous datasets. Notably, on the Texas and Cornell web datasets, ASP exceeded the strongest competing unsupervised models by 6.8 and 3.9 percentage points, achieving classification accuracies of 81.6% and 78.2%, respectively. Second, the framework retained competitive, state-of-the-art accuracy on homophilous datasets, reaching 84.7% on Cora and 73.0% on Citeseer. Third, ablation studies confirmed that the attribute-preserving contrastive module provided the most substantial performance boost, while structural alignment further refined predictive accuracy. Sensitivity analyses showed that the model maintains stable performance across varying neighborhood sizes and hop lengths.

These findings indicate that organizations can deploy a single self-supervised architecture across diverse data structures without needing prior knowledge of whether a network is homophilous or heterophilous. By eliminating reliance on extensive labeled data and removing the need for complex, computationally heavy diffusion matrices, ASP lowers deployment costs, shortens development timelines, and mitigates the risk of model failure on non-standard graph topologies.

Decision-makers should consider adopting ASP for enterprise graph representation pipelines, particularly when ground-truth labels are scarce and network connectivity patterns are unknown or varied. When deploying the system, practitioners should select appropriate attribute distance metrics (such as Jaccard or Cosine) tailored to the underlying data characteristics to maximize accuracy. While the current results provide high confidence across the seven evaluated benchmarks, future validation should explore performance on ultra-large industrial datasets and dynamically changing graphs to confirm scalability prior to large-scale production rollout.

arXiv: 2302.09532
Cover for Attribute and Structure Preserving Graph Contrastive Learning

Abstract

Graph Contrastive Learning (GCL) has drawn much research interest due to its strong ability to capture both graph structure and node attribute information in a self-supervised manner. Current GCL methods usually adopt Graph Neural Networks (GNNs) as the base encoder, which typically relies on the homophily assumption of networks and overlooks node similarity in the attribute space. There are many scenarios where such assumption cannot be satisfied, or node similarity plays a crucial role. In order to design a more robust mechanism, we develop a novel attribute and structure preserving graph contrastive learning framework, named ASP, which comprehensively and efficiently preserves node attributes while exploiting graph structure. Specifically, we consider three different graph views in our framework, i.e., original view, attribute view, and global structure view. Then, we perform contrastive learning across three views in a joint fashion, mining comprehensive graph information. We validate the effectiveness of the proposed framework on various real-world networks with different levels of homophily. The results demonstrate the superior performance of our model over the representative baselines.

Table of Contents

  • Introduction
  • Related Work
  • Notations and Preliminaries
  • Notations
  • Homophily
  • Graph Contrastive Learning
  • Proposed Method
  • View Generation
  • Attribute Preserving Contrastive Learning
  • Structure Preserving Contrastive Learning
  • Model Training
  • Experimental Results
  • Experimental Setup
  • Performance Comparison
  • Ablation Study
  • Parameter Sensitivity
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — ASP jointly preserves attribute and structural information

    model/method

    ASP is a self-supervised graph contrastive learning framework designed for attributed graphs whose homophily may be high or low. It constructs three views of the same graph: the unmodified original graph, an attribute-similarity graph, and a higher-order global-structure graph. Two contrastive modules are trained jointly: an attribute-preserving module contrasts the original view with a view that adds attribute-derived representations to the original representations, while a structure-preserving module contrasts the original view with the higher-order structure view. A third cross-module objective aligns the original-view representations learned by the two modules.

  2. Knowl 2 — Attribute and global-structure view construction

    model/method

    For an attributed graph with node-feature vectors xi∈Rfx_i\in\mathbb{R}^{f}, ASP uses the unmodified graph as its original view. Its attribute view is a kk-nearest-neighbor graph constructed from node attributes: each node is connected to its kk closest nodes under either cosine distance or the paper's Jaccard distance. For nodes viv_i and vjv_j, these distances are

    dcos(vi,vj)=1−xiTxj∥xi∥2∥xj∥2,d_{\mathrm{cos}}(v_i,v_j)=1-\frac{x_i^{\mathsf T}x_j}{\lVert x_i\rVert_2\lVert x_j\rVert_2}, dJ(vi,vj)=#{dimensions in which xi and xj differ}#{nonzero dimensions of xi and xj}.d_J(v_i,v_j)=\frac{\#\{\text{dimensions in which }x_i\text{ and }x_j\text{ differ}\}}{\#\{\text{nonzero dimensions of }x_i\text{ and }x_j\}}.

    The resulting attribute adjacency and degree matrices are denoted by AFA_F and DFD_F. The global-structure view is a higher-order propagation view in which each node receives information from nodes ll hops away, with ll chosen substantially larger than the propagation depth used for the ordinary view. The higher-order operator can be based on either the original normalized adjacency or the attribute-graph normalized adjacency; ASP uses the attribute-graph operator for non-homophilous datasets.

  3. Knowl 3 — Simplified graph convolution encoder and homophily-dependent propagation

    model/method

    ASP uses Simplified Graph Convolution (SGC) as its default encoder. For an adjacency matrix A∈Rn×nA\in\mathbb{R}^{n\times n}, let A~=A+In\widetilde A=A+I_n add self-loops and let D~\widetilde D be its diagonal degree matrix. With node-feature matrix X∈Rn×fX\in\mathbb{R}^{n\times f}, trainable weights Θ∈Rf×d\Theta\in\mathbb{R}^{f\times d}, output dimension dd, and propagation depth P∈Z≥0P\in\mathbb{Z}_{\ge 0}, the encoder is

    H=SPXΘ,S=D~−1/2A~D~−1/2.H=S^P X\Theta,\qquad S=\widetilde D^{-1/2}\widetilde A\widetilde D^{-1/2}.

    Here H∈Rn×dH\in\mathbb{R}^{n\times d} is the node-representation matrix. ASP permits P=0P=0, which reduces the encoder to a one-layer MLP; the authors use this option because it is more suitable for non-homophilous graphs. For homophilous graphs, they use positive propagation depth so that information from nearby, likely similar nodes is aggregated.

  4. Knowl 4 — Attribute-preserving contrastive module

    model/method

    Let SS be the normalized original-graph operator and let SF=DF−1/2AFDF−1/2S_F=D_F^{-1/2}A_FD_F^{-1/2} be the normalized attribute-graph operator. ASP forms two representations with a shared weight matrix Θa∈Rf×d\Theta_a\in\mathbb{R}^{f\times d}:

    Ho=SpXΘa,Ha=SpXΘa+SFXΘa.H^o=S^pX\Theta_a, \qquad H^a=S^pX\Theta_a+S_FX\Theta_a.

    Ho∈Rn×dH^o\in\mathbb{R}^{n\times d} is the original-view representation and Ha∈Rn×dH^a\in\mathbb{R}^{n\times d} is the attribute-complemented representation. Thus, ASP does not contrast two entirely different encoders; it adds attribute-similarity information to the original representation before contrasting. For node viv_i, the corresponding rows hioh_i^o and hiah_i^a are treated as a positive pair, while representations of other nodes are negatives. The module uses an InfoNCE loss with temperature τ>0\tau>0 and cosine agreement score D(⋅,⋅)D(\cdot,\cdot):

    Lattr(vi)=−log⁡exp⁡ ⁣(D(hio,hia)/τ)∑j=1nexp⁡ ⁣(D(hio,hja)/τ)+∑v∈{o,a}∑j=1n1[j≠i]exp⁡ ⁣(D(hiv,hjv)/τ).\mathcal{L}_{\mathrm{attr}}(v_i)=-\log\frac{\exp\!\left(D(h_i^o,h_i^a)/\tau\right)}{\displaystyle\sum_{j=1}^{n}\exp\!\left(D(h_i^o,h_j^a)/\tau\right)+\displaystyle\sum_{v\in\{o,a\}}\sum_{j=1}^{n}\mathbf{1}[j\ne i]\exp\!\left(D(h_i^v,h_j^v)/\tau\right)}.

    Here nn is the number of nodes, hivh_i^v is the representation of node ii from view vv, and 1[j≠i]\mathbf{1}[j\ne i] equals one when j≠ij\ne i and zero otherwise. This objective encourages the original and attribute-complemented representations of the same node to agree while separating them from other-node representations.

  5. Knowl 5 — Structure-preserving contrastive module

    model/method

    ASP uses a second SGC-based module to preserve global structure. Let SS be the normalized original-graph operator, SFS_F the normalized attribute-graph operator, and SG∈{S,SF}S_G\in\{S,S_F\} the operator selected for the higher-order view. With distinct trainable matrices Θo~\Theta_{\widetilde o} and Θs\Theta_s, propagation depths p≥0p\ge 0 and l>pl>p, the two representations are

    Ho~=SpXΘo~,Hs=(SG)lXΘs.H^{\widetilde o}=S^pX\Theta_{\widetilde o}, \qquad H^s=(S_G)^lX\Theta_s.

    Ho~H^{\widetilde o} represents the original view and HsH^s represents the global-structure view. For non-homophilous datasets, ASP sets SG=SFS_G=S_F so that the long-range propagation is based on attribute similarity rather than the potentially misleading original edges. For node viv_i, the structure loss treats hio~h_i^{\widetilde o} and hish_i^s as a positive pair and uses all other-node representations as negatives:

    Lstr(vi)=−log⁡exp⁡ ⁣(D(hio~,his)/τ)∑j=1nexp⁡ ⁣(D(hio~,hjs)/τ)+∑v∈{o~,s}∑j=1n1[j≠i]exp⁡ ⁣(D(hiv,hjv)/τ).\mathcal{L}_{\mathrm{str}}(v_i)=-\log\frac{\exp\!\left(D(h_i^{\widetilde o},h_i^s)/\tau\right)}{\displaystyle\sum_{j=1}^{n}\exp\!\left(D(h_i^{\widetilde o},h_j^s)/\tau\right)+\displaystyle\sum_{v\in\{\widetilde o,s\}}\sum_{j=1}^{n}\mathbf{1}[j\ne i]\exp\!\left(D(h_i^v,h_j^v)/\tau\right)}.

    The score DD is cosine similarity and τ>0\tau>0 is the contrastive temperature.

  6. Knowl 6 — Cross-module alignment and overall training objective

    equation

    ASP aligns the original-view representations from its two contrastive modules using a cross-module InfoNCE objective. For node viv_i, hioh_i^o and hio~h_i^{\widetilde o} form the positive pair, while representations of other nodes are negatives:

    Lcross(vi)=−log⁡exp⁡ ⁣(D(hio,hio~)/τ)∑j=1nexp⁡ ⁣(D(hio,hjo~)/τ)+∑v∈{o,o~}∑j=1n1[j≠i]exp⁡ ⁣(D(hiv,hjv)/τ).\mathcal{L}_{\mathrm{cross}}(v_i)=-\log\frac{\exp\!\left(D(h_i^o,h_i^{\widetilde o})/\tau\right)}{\displaystyle\sum_{j=1}^{n}\exp\!\left(D(h_i^o,h_j^{\widetilde o})/\tau\right)+\displaystyle\sum_{v\in\{o,\widetilde o\}}\sum_{j=1}^{n}\mathbf{1}[j\ne i]\exp\!\left(D(h_i^v,h_j^v)/\tau\right)}.

    The complete objective averages the three per-node losses and weights the structure and cross-module terms by tunable coefficients λ1\lambda_1 and λ2\lambda_2:

    L=1n∑i=1n[Lattr(vi)+λ1Lstr(vi)+λ2Lcross(vi)].\mathcal{L}=\frac{1}{n}\sum_{i=1}^{n}\left[\mathcal{L}_{\mathrm{attr}}(v_i)+\lambda_1\mathcal{L}_{\mathrm{str}}(v_i)+\lambda_2\mathcal{L}_{\mathrm{cross}}(v_i)\right].

    Here nn is the number of nodes, DD is cosine similarity, and τ\tau is the shared temperature parameter. The cross-module term makes the two modules agree on the same original graph while retaining their complementary attribute and global-structure signals.

  7. Knowl 7 — ASP training procedure

    algorithm

    The ASP training algorithm takes an unlabeled attributed graph G=(V,E,X)G=(V,E,X) as input and returns the optimal encoder weights for the attribute, structure, and cross-module representations. In each epoch it computes the four representation matrices, evaluates the three contrastive losses, forms the weighted total objective, and updates all trainable weights by gradient-based optimization.

    Input: Unlabeled attributed graph G=(V,E,X)G=(V,E,X)
    Output: Encoder weights Θa\Theta_a, Θo~\Theta_{\widetilde o}, and Θs\Theta_s
    for each training epoch do
        Compute HoH^o and HaH^a for the attribute-preserving module
        Compute Ho~H^{\widetilde o} and HsH^s for the structure-preserving module
        Compute Lattr\mathcal{L}_{\mathrm{attr}} from pairs (Ho,Ha)(H^o,H^a)
        Compute Lstr\mathcal{L}_{\mathrm{str}} from pairs (Ho~,Hs)(H^{\widetilde o},H^s)
        Compute Lcross\mathcal{L}_{\mathrm{cross}} from pairs (Ho,Ho~)(H^o,H^{\widetilde o})
        Form L=Lattr+λ1Lstr+λ2Lcross\mathcal{L}=\mathcal{L}_{\mathrm{attr}}+\lambda_1\mathcal{L}_{\mathrm{str}}+\lambda_2\mathcal{L}_{\mathrm{cross}}
        Update all trainable encoder weights to reduce L\mathcal{L}
    end for
    return the optimized encoder weights
  8. Knowl 8 — Datasets, evaluation protocol, and implementation settings

    experimental setup

    ASP was evaluated on seven real-world node-classification datasets spanning homophilous citation networks and non-homophilous webpage or actor-co-occurrence networks. The model was pretrained without labels on (A,X)(A,X), and the resulting node embeddings were evaluated by training an ℓ2\ell_2-regularized logistic-regression classifier. Results are means and standard deviations over five runs. Cora, Citeseer, and Pubmed use the public fixed split; Actor, Cornell, Texas, and Wisconsin use per-class 60%/20%/20% train/validation/test splits. ASP-cos uses cosine distance for the attribute kNN graph, whereas ASP-J uses the paper's Jaccard distance.

    Could not parse LaTeX table

    The tuned learning-rate candidates were {10−1,10−2,10−3,10−4}\{10^{-1},10^{-2},10^{-3},10^{-4}\}, the kNN values were k∈{20,30,40,50,60,70}k\in\{20,30,40,50,60,70\}, and the higher-order hop values were l∈{10,20,30}l\in\{10,20,30\}. The encoder output dimension was fixed at 256 for homophilous datasets and selected from {24,32,64,128}\{24,32,64,128\} for non-homophilous datasets. Training used patience 20 and a maximum of 500 epochs.

  9. Knowl 9 — Node-classification performance across homophily levels

    data/table

    The evaluation compares supervised GNN baselines, unsupervised graph-contrastive baselines, and the two ASP variants. On homophilous graphs, ASP-J is the strongest method on Cora and Citeseer, while RoSA is strongest on Pubmed among the listed methods. On non-homophilous graphs, ASP is best on every dataset, including methods using MLP encoders rather than GNNs. The results support the paper's claim that combining attribute similarity with graph structure is particularly useful when ordinary GNN propagation is unreliable.

    Could not parse LaTeX table
    Could not parse LaTeX table

    On Cornell and Texas, ASP-cos exceeds the strongest competing MLP-based baseline by 3.9 and 6.8 percentage points, respectively. GRACE and RoSA improve substantially when their GNN encoders are replaced by MLPs on non-homophilous graphs, whereas ASP remains stronger than both variants. The better ASP distance metric depends on the dataset: ASP-J is better on all three homophilous datasets and Actor, while ASP-cos is better on Cornell, Texas, and Wisconsin.

  10. Knowl 10 — Ablation evidence for the three ASP components

    empirical result

    Removing any of ASP's three components reduces node-classification accuracy on the non-homophilous datasets. The attribute-preserving module is the most important: removing it causes the largest deterioration on Cornell, Texas, and Wisconsin, while the structure-preserving module and cross-module alignment provide additional gains. The reported values use ASP-J for Actor and ASP-cos for the other three datasets.

    Could not parse LaTeX table
  11. Knowl 11 — Sensitivity to attribute-neighborhood size and global hop count

    empirical result

    ASP was tested with k∈{20,30,40,50,60,70}k\in\{20,30,40,50,60,70\} nearest neighbors for the attribute graph and l∈{10,15,20,25,30}l\in\{10,15,20,25,30\} global propagation hops on Cora and Texas. The best observed kNN sizes were k=40k=40 for Cora and k=70k=70 for Texas, suggesting that relatively large attribute neighborhoods can be useful. Performance remained competitive across the tested kk values, including the worst settings. The best global hop counts were l=10l=10 for Cora and l=30l=30 for Texas, and performance was likewise described as robust across the tested hop values. These results also show that the preferred propagation range depends on homophily level.

Coverage note — No substantial contributed material was omitted; background, related work, references, and acknowledgements were excluded.

References

  1. 1.Belghazi, M. I.; Baratin, A.; Rajeshwar, S.; Ozair, S.; Bengio, Y.; Courville, A.; and Hjelm, D. 2018. Mutual Information Neural Estimation. In ICML, 531–540.
  2. 2.Bruna, J.; Zaremba, W.; Szlam, A.; and LeCun, Y. 2014. Spectral Networks and Locally Connected Networks on Graphs. In Bengio, Y.; and LeCun, Y., eds., ICLR.
  3. 3.Chen, M.; Wei, Z.; Huang, Z.; Ding, B.; and Li, Y. 2020a. Simple and deep graph convolutional networks. In International conference on machine learning, 1725–1735.
  4. 4.Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020b. A Simple Framework for Contrastive Learning of Visual Representations. In ICML.
  5. 5.Defferrard, M.; Bresson, X.; and Vandergheynst, P. 2016. Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering. In Proceedings of the 30th International Conference on Neural Information Processing Systems.
  6. 6.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805.
  7. 7.Do, K.; Tran, T.; and Venkatesh, S. 2018. Graph Transformation Policy Network for Chemical Reaction Prediction. arXiv:1812.09441.
  8. 8.Fey, M.; and Lenssen, J. E. 2019. Fast Graph Representation Learning with PyTorch Geometric. In ICLR.
  9. 9.Gasteiger, J.; Bojchevski, A.; and Gunnemann, S. 2019. Predict then Propagate: Graph Neural Networks meet Personalized PageRank. In ICLR.
  10. 10.Gilmer, J.; Schoenholz, S. S.; Riley, P. F.; Vinyals, O.; and Dahl, G. E. 2017. Neural Message Passing for Quantum Chemistry. In ICML, 1263–1272.
  11. 11.Gutmann, M.; and Hyvarinen, A. 2010. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In International Conference on Artificial Intelligence and Statistics, 297–304.
  12. 12.Hamilton, W.; Ying, Z.; and Leskovec, J. 2017. Inductive Representation Learning on Large Graphs. In Guyon, I.; Luxburg, U. V.; Bengio, S.; Wallach, H.; Fergus, R.; Vishwanathan, S.; and Garnett, R., eds., Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  13. 13.Hassani, K.; and Khasahmadi, A. H. 2020. Contrastive Multi-View Representation Learning on Graphs. In ICML.
  14. 14.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9726–9735.
  15. 15.Hu, W.; Liu, B.; Gomes, J.; Zitnik, M.; Liang, P.; Pande, V.; and Leskovec, J. 2020. Strategies for Pre-training Graph Neural Networks. In International Conference on Learning Representations.
  16. 16.Jiao, Y.; Xiong, Y.; Zhang, J.; Zhang, Y.; Zhang, T.; and Zhu, Y. 2020. Sub-graph Contrast for Scalable Self-Supervised Graph Representation Learning. arXiv:2009.10273.
  17. 17.Jin, D.; Yu, Z.; Huo, C.; Wang, R.; Wang, X.; He, D.; and Han, J. 2021a. Universal Graph Convolutional Networks. In Beygelzimer, A.; Dauphin, Y.; Liang, P.; and Vaughan, J. W., eds., Advances in Neural Information Processing Systems.
  18. 18.Jin, W.; Derr, T.; Liu, H.; Wang, Y.; Wang, S.; Liu, Z.; and Tang, J. 2020. Self-supervised Learning on Graphs: Deep Insights and New Direction. arXiv:2006.10141.
  19. 19.Jin, W.; Derr, T.; Wang, Y.; Ma, Y.; Liu, Z.; and Tang, J. 2021b. Node Similarity Preserving Graph Convolutional Networks. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining. ACM.
  20. 20.Kipf, T. N.; and Welling, M. 2017. Semi-Supervised Classification with Graph Convolutional Networks. In ICLR.
  21. 21.Kondor, R. I.; and Lafferty, J. 2002. Diffusion kernels on graphs and other discrete structures. In ICML, 315–322.
  22. 22.Li, Q.; Han, Z.; and Wu, X.-M. 2018. Deeper insights into graph convolutional networks for semi-supervised learning. In Thirty-Second AAAI conference on artificial intelligence.
  23. 23.Li, X.; Zhu, R.; Cheng, Y.; Shan, C.; Luo, S.; Li, D.; and Qian, W. 2022. Finding Global Homophily in Graph Neural Networks When Meeting Heterophily. arXiv:2205.07308.
  24. 24.Li, Y.; Gu, C.; Dullien, T.; Vinyals, O.; and Kohli, P. 2019. Graph Matching Networks for Learning the Similarity of Graph Structured Objects. In ICML, 3835–3845.
  25. 25.Lim, D.; Hohne, F. M.; Li, X.; Huang, S. L.; Gupta, V.; Bhalerao, O. P.; and Lim, S.-N. 2021. Large Scale Learning on Non-Homophilous Graphs: New Benchmarks and Strong Simple Methods. In Advances in Neural Information Processing Systems.
  26. 26.Liu, Y.; Jin, M.; Pan, S.; Zhou, C.; Zheng, Y.; Xia, F.; and Yu, P. 2022. Graph Self-Supervised Learning: A Survey. IEEE Transactions on Knowledge and Data Engineering.
  27. 27.Nowozin, S.; Cseke, B.; and Tomioka, R. 2016. F-GAN: Training Generative Neural Samplers Using Variational Divergence Minimization. In International Conference on Neural Information Processing Systems, 271–279.
  28. 28.Page, L.; Brin, S.; Motwani, R.; and Winograd, T. 1999. The PageRank Citation Ranking: Bringing Order to the Web. Technical Report 1999-66, Stanford InfoLab.
  29. 29.Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; and Lerer, A. 2017. Automatic Differentiation in PyTorch. In NIPS 2017 Workshop on Autodiff.
  30. 30.Pei, H.; Wei, B.; Chang, K. C.; Lei, Y.; and Yang, B. 2020. Geom-GCN: Geometric Graph Convolutional Networks. In ICLR.
  31. 31.Peng, Z.; Dong, Y.; Luo, M.; Wu, X.-M.; and Zheng, Q. 2020a. Self-Supervised Graph Representation Learning via Global Context Prediction. arXiv:2003.01604.
  32. 32.Peng, Z.; Huang, W.; Luo, M.; Zheng, Q.; Rong, Y.; Xu, T.; and Huang, J. 2020b. Graph Representation Learning via Graphical Mutual Information Maximization. In Proceedings of The Web Conference.
  33. 33.Shlomi, J.; Battaglia, P.; and Vlimant, J.-R. 2021. Graph neural networks in particle physics. Machine Learning: Science and Technology.
  34. 34.Shuman, D.; Narang, S. K.; Frossard, P.; Ortega, A.; and Vandergheynst, P. 2012. The Emerging Field of Signal Processing on Graphs: Extending High-Dimensional Data Analysis to Networks and Other Irregular Domains. IEEE Signal Processing Magazine, 30.
  35. 35.Sun, F.-Y.; Hoffman, J.; Verma, V.; and Tang, J. 2019. InfoGraph: Unsupervised and Semi-supervised Graph-Level Representation Learning via Mutual Information Maximization. In International Conference on Learning Representations.
  36. 36.Thakoor, S.; Tallec, C.; Azar, M. G.; Azabou, M.; Dyer, E. L.; Munos, R.; Velickovic, P.; and Valko, M. 2021. Large-Scale Representation Learning on Graphs via Bootstrapping. arXiv:2102.06514.
  37. 37.Tian, Y.; Sun, C.; Poole, B.; Krishnan, D.; Schmid, C.; and Isola, P. 2020. What Makes for Good Views for Contrastive Learning? In Proceedings of the 34th International Conference on Neural Information Processing Systems.
  38. 38.Velickovic, P.; Cucurull, G.; Casanova, A.; Romero, A.; Lio, P.; and Bengio, Y. 2018. Graph Attention Networks. ICLR. Accepted as poster.
  39. 39.Velickovic, P.; Fedus, W.; Hamilton, W. L.; Lio, P.; Bengio, Y.; and Hjelm, R. D. 2019. Deep Graph Infomax. In ICLR.
  40. 40.Wu, F.; Souza, A.; Zhang, T.; Fifty, C.; Yu, T.; and Weinberger, K. 2019. Simplifying Graph Convolutional Networks. In ICML, 6861–6871.
  41. 41.Wu, L.; Lin, H.; Tan, C.; Gao, Z.; and Li, S. Z. 2021. Self-supervised Learning on Graphs: Contrastive, Generative, or Predictive. IEEE Transactions on Knowledge and Data Engineering.
  42. 42.Xia, J.; Wu, L.; Chen, J.; Hu, B.; and Li, S. Z. 2022. SimGRACE: A Simple Framework for Graph Contrastive Learning without Data Augmentation. In Proceedings of the ACM Web Conference 2022.
  43. 43.Xie, Y.; Xu, Z.; Zhang, J.; Wang, Z.; and Ji, S. 2021. Self-Supervised Learning of Graph Neural Networks: A Unified Review. arXiv:2102.10757.
  44. 44.Yang, Z.; Cohen, W. W.; and Salakhutdinov, R. 2016. Revisiting Semi-Supervised Learning with Graph Embeddings. In ICML, 40–48.
  45. 45.You, Y.; Chen, T.; Shen, Y.; and Wang, Z. 2021. Graph Contrastive Learning Automated. arXiv:2106.07594.
  46. 46.You, Y.; Chen, T.; Sui, Y.; Chen, T.; Wang, Z.; and Shen, Y. 2020a. Graph Contrastive Learning with Augmentations. In Advances in Neural Information Processing Systems, 5812–5823.
  47. 47.You, Y.; Chen, T.; Wang, Z.; and Shen, Y. 2020b. When Does Self-Supervision Help Graph Convolutional Networks? In ICML.
  48. 48.Yu, J.; Yin, H.; Xia, X.; Chen, T.; Cui, L.; and Nguyen, Q. V. H. 2022. Are Graph Augmentations Necessary? Simple Graph Contrastive Learning for Recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval.
  49. 49.Zeng, D.; Liu, W.; Chen, W.; Zhou, L.; Zhang, M.; and Qu, H. 2023. Substructure Aware Graph Neural Networks. In Proc. of AAAI.
  50. 50.Zeng, D.; Zhou, L.; Liu, W.; Qu, H.; and Chen, W. 2022. A Simple Graph Neural Network via Layer Sniffer. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5687–5691. IEEE.
  51. 51.Zhang, Y.; Zhu, H.; Song, Z.; Koniusz, P.; and King, I. 2022. COSTA: Covariance-Preserving Feature Augmentation for Graph Contrastive Learning. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining.
  52. 52.Zhu, J.; Rossi, R. A.; Rao, A.; Mai, T.; Lipka, N.; Ahmed, N. K.; and Koutra, D. 2021a. Graph Neural Networks with Heterophily. Proceedings of the AAAI Conference on Artificial Intelligence, 35(12).
  53. 53.Zhu, J.; Yan, Y.; Zhao, L.; Heimann, M.; Akoglu, L.; and Koutra, D. 2020a. Beyond Homophily in Graph Neural Networks: Current Limitations and Effective Designs. In Advances in Neural Information Processing Systems, 7793–7804.
  54. 54.Zhu, Y.; Guo, J.; Wu, F.; and Tang, S. 2022. RoSA: A Robust Self-Aligned Framework for Node-Node Graph Contrastive Learning. arXiv:2204.13846.
  55. 55.Zhu, Y.; Xu, Y.; Yu, F.; Liu, Q.; Wu, S.; and Wang, L. 2020b. Deep Graph Contrastive Representation Learning. In ICML.
  56. 56.Zhu, Y.; Xu, Y.; Yu, F.; Liu, Q.; Wu, S.; and Wang, L. 2021b. Graph Contrastive Learning with Adaptive Augmentation. In Proceedings of the Web Conference 2021, WWW ’21, 2069–2080. Association for Computing Machinery.
  57. 57.Zitnik, M.; Agrawal, M.; and Leskovec, J. 2018. Modeling polypharmacy side effects with graph convolutional networks. Bioinformatics.

Citation

MLA
Lu, W., et al. “Pseudo Contrastive Learning for Graph-based Semi-supervised Learning”. arXiv, 2023, http://arxiv.org/abs/2302.09532v3.
APA
Lu, W., Guan, Z., Zhao, W., Yang, Y., Lv, Y., Xing, L., Yu, B., & Tao, D. (2023). Pseudo Contrastive Learning for Graph-based Semi-supervised Learning. arXiv. http://arxiv.org/abs/2302.09532v3
Chicago
Lu, W., Z. Guan, W. Zhao, et al. 2023. “Pseudo Contrastive Learning for Graph-based Semi-supervised Learning”. arXiv. http://arxiv.org/abs/2302.09532v3.
Harvard
Lu, W. et al. (2023) “Pseudo Contrastive Learning for Graph-based Semi-supervised Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.09532v3.
Vancouver
1. Lu W, Guan Z, Zhao W, Yang Y, Lv Y, Xing L, Yu B, Tao D (2023) Pseudo Contrastive Learning for Graph-based Semi-supervised Learning. arXiv

BibTeX

@article{lu2023pseudo,
  title = {Pseudo Contrastive Learning for Graph-based Semi-supervised Learning},
  author = {Lu, Weigang and Guan, Ziyu and Zhao, Wei and Yang, Yaming and Lv, Yuanhai and Xing, Lining and Yu, Baosheng and Tao, Dacheng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.09532v3},
  eprint = {2302.09532}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF