Attribute and Structure Preserving Graph Contrastive Learning
Jialu ChenGang Kou
Proposes an attribute and structure preserving graph contrastive learning framework that overcomes standard homophily assumptions by jointly contrasting original, attribute similarity, and higher-order structural views.
Organizations increasingly rely on automated graph analysis to extract insights from interconnected systems such as citation networks, web links, and social interactions. A persistent challenge in this domain is the scarcity of labeled training data, which has led to widespread adoption of self-supervised learning techniques that learn representations without human annotation. However, prevailing graph learning methods depend heavily on the assumption of homophily—the expectation that connected entities share similar characteristics or labels. When applied to real-world networks where connected items are dissimilar (heterophily), standard models often fail to capture critical relationships or lose essential node-level characteristics.
The article introduces and evaluates a self-supervised framework called Attribute and Structure Preserving Graph Contrastive Learning (ASP). The primary objective is to demonstrate that simultaneously preserving feature similarities and broader topological patterns allows an unsupervised model to generate high-performing node representations across networks regardless of their homophily level.
To evaluate the system, the authors conducted empirical experiments on seven real-world network datasets, encompassing three standard homophilous citation datasets and four non-homophilous web and co-occurrence datasets. The framework constructs three distinct views of the network data—an original graph view, an attribute similarity view built via nearest neighbors, and a higher-order structure view capturing multi-hop connections. ASP contrasts these views using a simplified graph convolutional encoder, combining node features to bridge wide disparities between graph topology and attributes, and aligns the representations through a cross-module objective. The learned representations were tested on node classification tasks against multiple state-of-the-art supervised and self-supervised models.
The experimental findings demonstrate significant performance advantages. First, ASP consistently outperformed all tested supervised and unsupervised baselines across all four non-homophilous datasets. Notably, on the Texas and Cornell web datasets, ASP exceeded the strongest competing unsupervised models by 6.8 and 3.9 percentage points, achieving classification accuracies of 81.6% and 78.2%, respectively. Second, the framework retained competitive, state-of-the-art accuracy on homophilous datasets, reaching 84.7% on Cora and 73.0% on Citeseer. Third, ablation studies confirmed that the attribute-preserving contrastive module provided the most substantial performance boost, while structural alignment further refined predictive accuracy. Sensitivity analyses showed that the model maintains stable performance across varying neighborhood sizes and hop lengths.
These findings indicate that organizations can deploy a single self-supervised architecture across diverse data structures without needing prior knowledge of whether a network is homophilous or heterophilous. By eliminating reliance on extensive labeled data and removing the need for complex, computationally heavy diffusion matrices, ASP lowers deployment costs, shortens development timelines, and mitigates the risk of model failure on non-standard graph topologies.
Decision-makers should consider adopting ASP for enterprise graph representation pipelines, particularly when ground-truth labels are scarce and network connectivity patterns are unknown or varied. When deploying the system, practitioners should select appropriate attribute distance metrics (such as Jaccard or Cosine) tailored to the underlying data characteristics to maximize accuracy. While the current results provide high confidence across the seven evaluated benchmarks, future validation should explore performance on ultra-large industrial datasets and dynamically changing graphs to confirm scalability prior to large-scale production rollout.
- Paper: Graph Contrastive Learning with Adaptive Augmentation, Yanqiao Zhu et al. (2020). This paper establishes adaptive graph contrastive learning that preserves critical topological and attribute information, providing foundational context for ASP's multi-view contrastive framework.
- Paper: Beyond Homophily in Graph Neural Networks: Current Limitations and Effective Designs, Jiong Zhu et al. (2020). This work analyzes the core limitations of standard message-passing architectures under heterophily and introduces higher-order design principles that directly motivate ASP's structure and attribute preserving views.
- Paper: Contrastive Multi-View Representation Learning on Graphs, Kaveh Hassani et al. (2020). This paper introduces multi-view graph contrastive learning using local and diffusion-based global views, serving as key background for multi-view construction and alignment in graph representation learning.
- Paper: Finding Global Homophily in Graph Neural Networks When Meeting Heterophily, Xiang Li et al. (2022). This work demonstrates how combining feature similarity and network topology bridges the gap between homophily and heterophily, laying the groundwork for ASP's attribute similarity view.
- Paper: Graph Contrastive Learning with Augmentations, Yuning You et al. (2020). This foundational text introduces standard graph contrastive augmentation techniques that ASP generalizes to handle heterophilic and attribute-driven graphs.
- Paper: Neural Sheaf Diffusion: A Topological Perspective on Heterophily and Oversmoothing in GNNs, Cristian Bodnar et al. (2022). This study provides a rigorous geometric and topological perspective on how standard graph neural networks break down on heterophilic networks.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). This seminal work introduces the graph convolutional network architecture used as the baseline encoder across contrastive views in ASP.
- Paper: MA-GCL: Model Augmentation Tricks for Graph Contrastive Learning, Xumeng Gong et al. (2023). Explore this paper to see how model-level architectural augmentations can enhance graph contrastive learning without relying exclusively on topological and attribute perturbations.
- Paper: Label-free Node Classification on Graphs with Large Language Models (LLMs), Zhikai Chen et al. (2024). Read this paper to explore how zero-shot large language model annotations can be integrated with graph neural networks as an alternative label-free node classification paradigm.
- Paper: PRODIGY: Enabling In-context Learning Over Graphs, Qian Huang et al. (2023). Consult this study to learn how self-supervised pretraining objectives over graph structures can be generalized into in-context, zero-parameter-update graph learning systems.
