Graph Contrastive Learning with Adaptive Augmentation

Yanqiao ZhuYichen XuFeng YuQiang LiuShu WuLiang Wang

article2020WWW1,514 citations

Proposes an adaptive graph contrastive learning framework that preserves critical topology and node semantics by selectively perturbing unimportant structures and features based on graph priors, consistently outperforming standard uniform augmentation methods across benchmark datasets.

Listen

Modern data applications across e-commerce, social networks, and academic citation analysis rely heavily on graph-structured data. While deep learning models known as Graph Neural Networks have shown strong performance on such data, most existing solutions require massive amounts of manually labeled data for training, which is costly, slow, and often impractical. Unsupervised contrastive learning has emerged as a promising alternative by training models to recognize core features without labels, yet existing techniques rely on uniform and random data corruption. Randomly removing links or masking features frequently damages vital structures, which degrades the overall quality of the learned representations.

The article introduces and evaluates a novel framework called Graph Contrastive Learning with Adaptive Augmentation (GCA), which aims to learn high-quality node embeddings without human supervision by adaptively preserving essential graph structure and node attributes during data perturbation.

The authors designed a contrastive representation learning framework that calculates network centrality metrics—such as degree, eigenvector, and PageRank centrality—to identify the most critical connections and features in a network. In this framework, stochastic corruption is applied adaptively: unimportant connections and feature dimensions receive higher probabilities of being removed or masked, while influential structures remain intact. The authors evaluated the approach across five standard benchmark datasets spanning Wikipedia articles, e-commerce co-purchase graphs, and co-authorship networks, benchmarking performance against traditional unsupervised baselines, recent deep contrastive models, and fully supervised neural networks.

The evaluation produced several key findings. First, GCA consistently outperformed all existing unsupervised baseline methods across all five datasets in node classification tasks, demonstrating higher accuracy margins. Second, the unsupervised GCA model matched or even exceeded the performance of fully supervised models, including standard Graph Convolutional Networks and Graph Attention Networks, across transductive tasks. Third, ablation experiments showed that combining adaptive strategies at both the network topology level and the feature attribute level produces superior results compared to using uniform corruption or adapting only one level, delivering up to a 1.5% absolute gain on e-commerce benchmarks. Finally, sensitivity analyses revealed that model performance remains stable across a wide range of corruption probabilities, provided the perturbation does not excessively dismantle the graph (below 0.5 probability).

These findings demonstrate that adaptive, structure-aware data augmentation resolves a critical bottleneck in self-supervised graph analysis. By eliminating the dependence on manual labels while preserving predictive accuracy, organizations can substantially reduce data preparation costs and development timelines. The framework provides an effective, plug-in solution for practical applications such as cold-start recommendation engines and community detection, without incurring heavy computational overhead since the centrality metrics need only be computed once upfront.

Organizations handling large interconnected datasets should consider adopting adaptive contrastive learning frameworks to reduce labeling expenses and enhance representation quality. When deploying these models, practitioners should maintain balanced perturbation rates and select standard degree or PageRank metrics for simplicity and efficiency. Stakeholders must also recognize that underlying data biases—such as historical demographic or behavioral skews—can persist through self-supervised training and should be monitored before deploying models into production systems.

  • Paper: Graph Contrastive Learning with Augmentations, Yuning You et al. (2020). This work establishes the foundational GraphCL framework for graph contrastive learning using standard stochastic data augmentations, which adaptive augmentation directly seeks to refine and improve upon.
  • Paper: Deep Graph Infomax, Petar Veličković et al. (2019). Deep Graph Infomax introduced contrastive mutual information maximization between local node embeddings and global graph summaries, establishing the core paradigm underlying unsupervised graph contrastive learning.
  • Paper: Contrastive Multi-View Representation Learning on Graphs, Kaveh Hassani et al. (2020). This paper establishes multi-view contrastive learning on graphs using localized and diffusion views, forming a direct precursor to contrasting multiple augmented graph views.
  • Paper: DropEdge: Towards Deep Graph Convolutional Networks on Node Classification, Yu Rong et al. (2019). This study introduces DropEdge for randomly perturbing graph connectivity during training, providing the baseline uniform edge-dropping strategy that adaptive augmentation explicitly improves with centrality-based priors.
  • Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). This paper presents the standard Graph Convolutional Network architecture widely adopted as the underlying encoder backbone for graph contrastive learning models.
Cover for Graph Contrastive Learning with Adaptive Augmentation

Abstract

Recently, contrastive learning (CL) has emerged as a successful method for unsupervised graph representation learning. Most graph CL methods first perform stochastic augmentation on the input graph to obtain two graph views and maximize the agreement of representations in the two views. Despite the prosperous development of graph CL methods, the design of graph augmentation schemes -- a crucial component in CL -- remains rarely explored. We argue that the data augmentation schemes should preserve intrinsic structures and attributes of graphs, which will force the model to learn representations that are insensitive to perturbation on unimportant nodes and edges. However, most existing methods adopt uniform data augmentation schemes, like uniformly dropping edges and uniformly shuffling features, leading to suboptimal performance. In this paper, we propose a novel graph contrastive representation learning method with adaptive augmentation that incorporates various priors for topological and semantic aspects of the graph. Specifically, on the topology level, we design augmentation schemes based on node centrality measures to highlight important connective structures. On the node attribute level, we corrupt node features by adding more noise to unimportant node features, to enforce the model to recognize underlying semantic information. We perform extensive experiments of node classification on a variety of real-world datasets. Experimental results demonstrate that our proposed method consistently outperforms existing state-of-the-art baselines and even surpasses some supervised counterparts, which validates the effectiveness of the proposed contrastive framework with adaptive augmentation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Contrastive Representation Learning
  • 2.2 Graph Representation Learning
  • 3 The Proposed Method
  • 3.1 Preliminaries
  • 3.2 The Contrastive Learning Framework
  • 3.3 Adaptive Graph Augmentation
  • 3.3.1 Topology-level augmentation.
  • 3.3.2 Node-attribute-level augmentation.
  • 3.4 Theoretical Justification
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.1.1 Datasets
  • 4.1.2 Evaluation protocol.
  • 4.1.3 Baselines.
  • 4.1.4 Implementation details.
  • 4.2 Performance on Node Classification (RQ1)
  • 4.3 Ablation Studies (RQ2)
  • 4.4 Sensitivity Analysis (RQ3)
  • 5 Conclusion
  • A Implementation Details
  • A.1 Computing Infrastructures
  • A.2 Hyperparameter Specifications
  • B Detailed Proofs
  • B.1 Proof of Theorem 1
  • B.2 Proof of Theorem 2
  • References

Knowls

  1. Knowl 1 — GCA Framework for Unsupervised Graph Representation Learning

    model/method

    Graph Contrastive learning with Adaptive augmentation (GCA) is an unsupervised framework for learning node representations on an unlabelled graph G=(V,E)G = (\mathcal{V}, \mathcal{E}) with node feature matrix X∈RN×F\mathbf{X} \in \mathbb{R}^{N \times F} and adjacency matrix A∈{0,1}N×N\mathbf{A} \in \{0, 1\}^{N \times N}, where N=∣V∣N = |\mathcal{V}| is the number of nodes and FF is the input feature dimension.

    The framework operates in three main steps:

    1. Generate two corrupted graph views G~1=(X~1,A~1)\widetilde{G}_1 = (\widetilde{\mathbf{X}}_1, \widetilde{\mathbf{A}}_1) and G~2=(X~2,A~2)\widetilde{G}_2 = (\widetilde{\mathbf{X}}_2, \widetilde{\mathbf{A}}_2) using stochastic, structure- and attribute-adaptive graph augmentations that selectively preserve influential topological links and feature dimensions.

    2. Feed both views into a shared Graph Neural Network (GNN) encoder f(X,A)f(\mathbf{X}, \mathbf{A}) to produce low-dimensional node representation matrices U=f(X~1,A~1)∈RN×F′\mathbf{U} = f(\widetilde{\mathbf{X}}_1, \widetilde{\mathbf{A}}_1) \in \mathbb{R}^{N \times F'} and V=f(X~2,A~2)∈RN×F′\mathbf{V} = f(\widetilde{\mathbf{X}}_2, \widetilde{\mathbf{A}}_2) \in \mathbb{R}^{N \times F'}, where F′≪FF' \ll F.

    3. Train the encoder by optimizing a node-level contrastive loss that pulls together representations of the same node across the two views (ui,vi)(\mathbf{u}_i, \mathbf{v}_i) while pushing them apart from representations of all other nodes in both views (intra-view and inter-view negative samples).

  2. Knowl 2 — Adaptive Topology-Level Augmentation via Edge Centrality

    model/method

    In GCA, topology-level data augmentation corrupts the input graph by sampling a modified edge set E~\widetilde{\mathcal{E}} from the original edge set E\mathcal{E}, where each edge (u,v)∈E(u, v) \in \mathcal{E} is retained with probability P{(u,v)∈E~}=1−puveP\{(u, v) \in \widetilde{\mathcal{E}}\} = 1 - p_{uv}^e. The edge removal probability puvep_{uv}^e is inversely related to the structural importance of edge (u,v)(u, v), ensuring that peripheral edges are more likely to be dropped while influential connective structures remain intact.

    Given a node centrality measure φc:V→R+\varphi_c: \mathcal{V} \rightarrow \mathbb{R}^+, edge centrality wuvew_{uv}^e is computed as:

    • For undirected graphs: wuve=φc(u)+φc(v)2w_{uv}^e = \frac{\varphi_c(u) + \varphi_c(v)}{2}
    • For directed graphs: wuve=φc(v)w_{uv}^e = \varphi_c(v) (the centrality of the target node).

    Supported node centrality functions φc\varphi_c include:

    1. Degree Centrality (in-degree for directed graphs).
    2. Eigenvector Centrality: the leading right eigenvector of the adjacency matrix A\mathbf{A}.
    3. PageRank Centrality: σ=αAD−1σ+1\boldsymbol{\sigma} = \alpha \mathbf{A} \mathbf{D}^{-1} \boldsymbol{\sigma} + \mathbf{1}, where D\mathbf{D} is the degree matrix and α=0.85\alpha = 0.85 is the damping factor.

    To compress heavy-tailed centrality variations across orders of magnitude, edge weights are log-transformed: suve=log⁡wuves_{uv}^e = \log w_{uv}^e. The edge removal probability is normalized and truncated via:

    puve=min⁡(smax⁡e−suvesmax⁡e−μse⋅pe, pτ)p_{uv}^e = \min\left( \frac{s_{\max}^e - s_{uv}^e}{s_{\max}^e - \mu_s^e} \cdot p_e, \, p_\tau \right)

    where smax⁡es_{\max}^e and μse\mu_s^e denote the maximum and mean of {suve}(u,v)∈E\{s_{uv}^e\}_{(u,v) \in \mathcal{E}}, pe∈[0,1)p_e \in [0, 1) is a hyperparameter governing the base edge removal rate, and pτ<1p_\tau < 1 is a truncation threshold to prevent graph over-corruption.

  3. Knowl 3 — Adaptive Node-Attribute-Level Augmentation via Feature Importance

    model/method

    In GCA, attribute-level data augmentation adds noise to the node feature matrix X∈RN×F\mathbf{X} \in \mathbb{R}^{N \times F} by masking feature dimensions with zeros based on global importance. A binary masking vector m~∈{0,1}F\widetilde{\mathbf{m}} \in \{0, 1\}^F is sampled where each dimension i∈{1,…,F}i \in \{1, \dots, F\} is drawn independently as m~i∼Bernoulli(1−pif)\widetilde{m}_i \sim \text{Bernoulli}(1 - p_i^f). The generated feature matrix X~\widetilde{\mathbf{X}} is computed via row-wise element-wise multiplication:

    X~=[x1∘m~; x2∘m~; … ; xN∘m~]⊤\widetilde{\mathbf{X}} = [\mathbf{x}_1 \circ \widetilde{\mathbf{m}};\, \mathbf{x}_2 \circ \widetilde{\mathbf{m}};\, \dots;\, \mathbf{x}_N \circ \widetilde{\mathbf{m}}]^\top

    The masking probability pifp_i^f reflects the importance weight wifw_i^f of feature dimension ii, computed using node centrality φc(u)\varphi_c(u):

    • For sparse binary features (xui∈{0,1}x_{ui} \in \{0, 1\}):

    wif=∑u∈Vxui⋅φc(u)w_i^f = \sum_{u \in \mathcal{V}} x_{ui} \cdot \varphi_c(u)

    • For dense continuous features (xui∈Rx_{ui} \in \mathbb{R}):

    wif=∑u∈V∣xui∣⋅φc(u)w_i^f = \sum_{u \in \mathcal{V}} |x_{ui}| \cdot \varphi_c(u)

    Feature log-weights sif=log⁡wifs_i^f = \log w_i^f are normalized to obtain feature masking probabilities:

    pif=min⁡(smax⁡f−sifsmax⁡f−μsf⋅pf, pτ)p_i^f = \min\left( \frac{s_{\max}^f - s_i^f}{s_{\max}^f - \mu_s^f} \cdot p_f, \, p_\tau \right)

    where smax⁡fs_{\max}^f and μsf\mu_s^f are the maximum and mean of {sif}i=1F\{s_i^f\}_{i=1}^F, pf∈[0,1)p_f \in [0, 1) is the base feature masking rate, and pτ<1p_\tau < 1 is a cut-off threshold. Feature dimensions frequently appearing in high-centrality nodes receive smaller masking probabilities.

  4. Knowl 4 — Node-Level Multi-View Graph Contrastive Objective

    equation

    Let ui,vi∈RF′\mathbf{u}_i, \mathbf{v}_i \in \mathbb{R}^{F'} denote the learned representations of node vi∈Vv_i \in \mathcal{V} (i∈{1,…,N}i \in \{1, \dots, N\}) from two corrupted graph views G~1\widetilde{G}_1 and G~2\widetilde{G}_2. For an anchor embedding ui\mathbf{u}_i, its positive counterpart is vi\mathbf{v}_i. Negative samples consist of all N−1N-1 other nodes in the alternate view (inter-view negatives) and all N−1N-1 other nodes in the same view (intra-view negatives).

    The pairwise contrastive loss for positive pair (ui,vi)(\mathbf{u}_i, \mathbf{v}_i) is:

    ℓ(ui,vi)=−log⁡exp⁡(θ(ui,vi)/τ)exp⁡(θ(ui,vi)/τ)+∑k≠iexp⁡(θ(ui,vk)/τ)+∑k≠iexp⁡(θ(ui,uk)/τ)\ell(\mathbf{u}_i, \mathbf{v}_i) = -\log \frac{\exp(\theta(\mathbf{u}_i, \mathbf{v}_i) / \tau)}{\exp(\theta(\mathbf{u}_i, \mathbf{v}_i) / \tau) + \sum_{k \neq i} \exp(\theta(\mathbf{u}_i, \mathbf{v}_k) / \tau) + \sum_{k \neq i} \exp(\theta(\mathbf{u}_i, \mathbf{u}_k) / \tau)}

    where τ\tau is a temperature hyperparameter and θ(u,v)=s(g(u),g(v))\theta(\mathbf{u}, \mathbf{v}) = s(g(\mathbf{u}), g(\mathbf{v})) is a critic function with s(⋅,⋅)s(\cdot, \cdot) denoting cosine similarity and g(⋅)g(\cdot) denoting a non-linear two-layer MLP projection head.

    The overall objective J\mathcal{J} to be maximized across all NN nodes over both symmetric view directions is:

    J=12N∑i=1N[ℓ(ui,vi)+ℓ(vi,ui)]\mathcal{J} = \frac{1}{2N} \sum_{i=1}^N \left[ \ell(\mathbf{u}_i, \mathbf{v}_i) + \ell(\mathbf{v}_i, \mathbf{u}_i) \right]

  5. Knowl 5 — GCA Model Training Algorithm

    algorithm

    The training procedure of Graph Contrastive learning with Adaptive augmentation (GCA) maximizes agreement between node representations across two distinct adaptively corrupted views over multiple epochs.

    Input: Input graph G=(V,E)G = (\mathcal{V}, \mathcal{E}), node feature matrix X\mathbf{X}, node centrality function φc\varphi_c, base drop probabilities pe,1,pf,1p_{e,1}, p_{f,1} for view 1 and pe,2,pf,2p_{e,2}, p_{f,2} for view 2, cut-off pτp_\tau, temperature τ\tau, GNN encoder ff, projection head gg
    Output: Trained GNN encoder parameters
    Compute edge removal probabilities {puve,1}\{p_{uv}^{e,1}\} and {puve,2}\{p_{uv}^{e,2}\} for each view based on edge centralities
    Compute feature masking probabilities {pif,1}\{p_i^{f,1}\} and {pif,2}\{p_i^{f,2}\} for each view based on feature weights
    for epoch = 1, 2, ... do
        Sample stochastic augmentation t1t_1: remove edges with probability puve,1p_{uv}^{e,1} and mask features with probability pif,1p_i^{f,1} to obtain view G~1=(X~1,A~1)\widetilde{G}_1 = (\widetilde{\mathbf{X}}_1, \widetilde{\mathbf{A}}_1)
        Sample stochastic augmentation t2t_2: remove edges with probability puve,2p_{uv}^{e,2} and mask features with probability pif,2p_i^{f,2} to obtain view G~2=(X~2,A~2)\widetilde{G}_2 = (\widetilde{\mathbf{X}}_2, \widetilde{\mathbf{A}}_2)
        Compute node representations U=f(X~1,A~1)\mathbf{U} = f(\widetilde{\mathbf{X}}_1, \widetilde{\mathbf{A}}_1)
        Compute node representations V=f(X~2,A~2)\mathbf{V} = f(\widetilde{\mathbf{X}}_2, \widetilde{\mathbf{A}}_2)
        Compute pairwise contrastive losses ℓ(ui,vi)\ell(\mathbf{u}_i, \mathbf{v}_i) and ℓ(vi,ui)\ell(\mathbf{v}_i, \mathbf{u}_i) for all i∈{1,…,N}i \in \{1, \dots, N\}
        Evaluate total contrastive objective J=12N∑i=1N[ℓ(ui,vi)+ℓ(vi,ui)]\mathcal{J} = \frac{1}{2N} \sum_{i=1}^N [\ell(\mathbf{u}_i, \mathbf{v}_i) + \ell(\mathbf{v}_i, \mathbf{u}_i)]
        Update parameters of ff and gg via stochastic gradient ascent to maximize J\mathcal{J}
    return GNN encoder ff
  6. Knowl 6 — Lower Bound on Input-Representation Mutual Information

    theoretical result

    Let Xi={xk}k∈N(i)\mathbf{X}_i = \{\mathbf{x}_k\}_{k \in \mathcal{N}(i)} denote the neighborhood features of node viv_i mapping to its output representation under a Graph Neural Network architecture, and let X\mathbf{X} be the random variable over neighborhoods with uniform distribution p(Xi)=1/Np(\mathbf{X}_i) = 1/N. Let U,V∈RF′\mathbf{U}, \mathbf{V} \in \mathbb{R}^{F'} be random variables representing node embeddings in two augmented graph views with joint distribution p(U,V)p(\mathbf{U}, \mathbf{V}).

    The GCA contrastive objective J\mathcal{J} satisfies:

    J≤I(X;U,V)\mathcal{J} \le I(\mathbf{X}; \mathbf{U}, \mathbf{V})

    where I(⋅;⋅)I(\cdot; \cdot) is the mutual information.

    Preconditions and Scope:

    1. The two augmented views satisfy the Markov relation U←X→V\mathbf{U} \leftarrow \mathbf{X} \rightarrow \mathbf{V}, meaning U\mathbf{U} and V\mathbf{V} are conditionally independent given X\mathbf{X}.
    2. By the InfoNCE lower bound property and the data processing inequality, J≤I(U;V)≤I(X;U,V)\mathcal{J} \le I(\mathbf{U}; \mathbf{V}) \le I(\mathbf{X}; \mathbf{U}, \mathbf{V}). Thus, maximizing the GCA objective directly maximizes a lower bound on the mutual information between the input graph data and the multi-view representations.
  7. Knowl 7 — Equivalence of GCA Contrastive Loss to Triplet Loss Maximization

    theoretical result

    When the projection head gg is the identity function (g(u)=ug(\mathbf{u}) = \mathbf{u}) and embedding similarity is measured by the inner product s(u,v)=u⊤vs(\mathbf{u}, \mathbf{v}) = \mathbf{u}^\top \mathbf{v}, and under the assumption that positive pairs are far more aligned than negative pairs (i.e., ui⊤vk≪ui⊤vi\mathbf{u}_i^\top \mathbf{v}_k \ll \mathbf{u}_i^\top \mathbf{v}_i and ui⊤uk≪ui⊤vi\mathbf{u}_i^\top \mathbf{u}_k \ll \mathbf{u}_i^\top \mathbf{v}_i for all k≠ik \neq i), minimizing the pairwise contrastive objective −ℓ(ui,vi)-\ell(\mathbf{u}_i, \mathbf{v}_i) is asymptotically proportional to maximizing a sum of triplet losses:

    −ℓ(ui,vi)∝4Nτ+∑j≠i(∥ui−vi∥2−∥ui−vj∥2+∥ui−vi∥2−∥ui−uj∥2)-\ell(\mathbf{u}_i, \mathbf{v}_i) \propto 4N\tau + \sum_{j \neq i} \left( \|\mathbf{u}_i - \mathbf{v}_i\|^2 - \|\mathbf{u}_i - \mathbf{v}_j\|^2 + \|\mathbf{u}_i - \mathbf{v}_i\|^2 - \|\mathbf{u}_i - \mathbf{u}_j\|^2 \right)

    where NN is the number of nodes and τ\tau is the temperature parameter.

    This equivalence demonstrates that the contrastive loss acts in metric space to minimize the distance between the two representations of the same node while maximizing the distance to both inter-view negatives (vj\mathbf{v}_j) and intra-view negatives (uj\mathbf{u}_j).

  8. Knowl 8 — Transductive Node Classification Performance of GCA

    data/table

    Node classification accuracy (in percentage ±\pm standard deviation) is measured across five benchmark graph datasets following the standard linear evaluation protocol (unsupervised GNN training followed by an ℓ2\ell_2-regularized logistic regression classifier evaluated over 20 random runs; 10%/10%/80% train/val/test splits for Amazon and Coauthor datasets, public splits for Wiki-CS):

    Method Training Data Wiki-CS Amazon-Computers Amazon-Photo Coauthor-CS Coauthor-Physics
    Raw features X\mathbf{X} 71.98±0.0071.98 \pm 0.00 73.81±0.0073.81 \pm 0.00 78.53±0.0078.53 \pm 0.00 90.37±0.0090.37 \pm 0.00 93.58±0.0093.58 \pm 0.00
    node2vec A\mathbf{A} 71.79±0.0571.79 \pm 0.05 84.39±0.0884.39 \pm 0.08 89.67±0.1289.67 \pm 0.12 85.08±0.0385.08 \pm 0.03 91.19±0.0491.19 \pm 0.04
    DeepWalk A\mathbf{A} 74.35±0.0674.35 \pm 0.06 85.68±0.0685.68 \pm 0.06 89.44±0.1189.44 \pm 0.11 84.61±0.2284.61 \pm 0.22 91.77±0.1591.77 \pm 0.15
    DeepWalk + features X,A\mathbf{X}, \mathbf{A} 77.21±0.0377.21 \pm 0.03 86.28±0.0786.28 \pm 0.07 90.05±0.0890.05 \pm 0.08 87.70±0.0487.70 \pm 0.04 94.90±0.0994.90 \pm 0.09
    GAE X,A\mathbf{X}, \mathbf{A} 70.15±0.0170.15 \pm 0.01 85.27±0.1985.27 \pm 0.19 91.62±0.1391.62 \pm 0.13 90.01±0.7190.01 \pm 0.71 94.92±0.0794.92 \pm 0.07
    VGAE X,A\mathbf{X}, \mathbf{A} 75.63±0.1975.63 \pm 0.19 86.37±0.2186.37 \pm 0.21 92.20±0.1192.20 \pm 0.11 92.11±0.0992.11 \pm 0.09 94.52±0.0094.52 \pm 0.00
    DGI X,A\mathbf{X}, \mathbf{A} 75.35±0.1475.35 \pm 0.14 83.95±0.4783.95 \pm 0.47 91.61±0.2291.61 \pm 0.22 92.15±0.6392.15 \pm 0.63 94.51±0.5294.51 \pm 0.52
    GMI X,A\mathbf{X}, \mathbf{A} 74.85±0.0874.85 \pm 0.08 82.21±0.3182.21 \pm 0.31 90.68±0.1790.68 \pm 0.17 OOM OOM
    MVGRL X,A\mathbf{X}, \mathbf{A} 77.52±0.0877.52 \pm 0.08 87.52±0.1187.52 \pm 0.11 91.74±0.0791.74 \pm 0.07 92.11±0.1292.11 \pm 0.12 95.33±0.0395.33 \pm 0.03
    GCA-DE X,A\mathbf{X}, \mathbf{A} 78.30±0.0078.30 \pm 0.00 87.85±0.31\mathbf{87.85 \pm 0.31} 92.49±0.0992.49 \pm 0.09 93.10±0.01\mathbf{93.10 \pm 0.01} 95.68±0.0595.68 \pm 0.05
    GCA-PR X,A\mathbf{X}, \mathbf{A} 78.35±0.05\mathbf{78.35 \pm 0.05} 87.80±0.2387.80 \pm 0.23 92.53±0.16\mathbf{92.53 \pm 0.16} 93.06±0.0393.06 \pm 0.03 95.72±0.0395.72 \pm 0.03
    GCA-EV X,A\mathbf{X}, \mathbf{A} 78.23±0.0478.23 \pm 0.04 87.54±0.4987.54 \pm 0.49 92.24±0.2192.24 \pm 0.21 92.95±0.1392.95 \pm 0.13 95.73±0.03\mathbf{95.73 \pm 0.03}
    GCN (supervised) X,A,Y\mathbf{X}, \mathbf{A}, \mathbf{Y} 77.19±0.1277.19 \pm 0.12 86.51±0.5486.51 \pm 0.54 92.42±0.2292.42 \pm 0.22 93.03±0.3193.03 \pm 0.31 95.65±0.1695.65 \pm 0.16
    GAT (supervised) X,A,Y\mathbf{X}, \mathbf{A}, \mathbf{Y} 77.65±0.1177.65 \pm 0.11 86.93±0.2986.93 \pm 0.29 92.56±0.3592.56 \pm 0.35 92.31±0.2492.31 \pm 0.24 95.47±0.1595.47 \pm 0.15

    OOM indicates Out-Of-Memory on a 32GB GPU. The GCA variants (degree centrality GCA-DE, PageRank centrality GCA-PR, and eigenvector centrality GCA-EV) consistently outperform all unsupervised baselines across all five datasets and exceed the performance of fully supervised GCN and GAT baselines on Wiki-CS, Amazon-Computers, Coauthor-CS, and Coauthor-Physics.

  9. Knowl 9 — Ablation Study of Topology and Attribute Adaptive Augmentations

    data/table

    An ablation study isolates the contributions of adaptive topology-level and adaptive attribute-level augmentations using degree centrality. The evaluated variants are:

    • GCA-T-A: Uniform edge dropping and uniform feature masking (baseline without adaptive augmentation, equivalent to GRACE).
    • GCA-T: Uniform edge dropping and adaptive feature masking.
    • GCA-A: Adaptive edge dropping and uniform feature masking.
    • GCA: Joint adaptive edge dropping and adaptive feature masking.

    Classification accuracy in percentage (±\pm standard deviation) across twenty runs:

    Variant Topology Attribute Wiki-CS Amazon-Computers Amazon-Photo Coauthor-CS Coauthor-Physics
    GCA-T-A Uniform Uniform 78.19±0.0178.19 \pm 0.01 86.25±0.2586.25 \pm 0.25 92.15±0.2492.15 \pm 0.24 92.93±0.0192.93 \pm 0.01 95.26±0.0295.26 \pm 0.02
    GCA-T Uniform Adaptive 78.23±0.0278.23 \pm 0.02 86.72±0.4986.72 \pm 0.49 92.20±0.2692.20 \pm 0.26 93.07±0.0193.07 \pm 0.01 95.59±0.0495.59 \pm 0.04
    GCA-A Adaptive Uniform 78.25±0.0278.25 \pm 0.02 87.66±0.3087.66 \pm 0.30 92.23±0.2092.23 \pm 0.20 93.02±0.0193.02 \pm 0.01 95.54±0.0295.54 \pm 0.02
    GCA Adaptive Adaptive 78.30±0.01\mathbf{78.30 \pm 0.01} 87.85±0.31\mathbf{87.85 \pm 0.31} 92.49±0.09\mathbf{92.49 \pm 0.09} 93.10±0.01\mathbf{93.10 \pm 0.01} 95.68±0.05\mathbf{95.68 \pm 0.05}

    Enabling adaptive augmentation on either topology alone (GCA-A) or attributes alone (GCA-T) consistently improves accuracy compared to uniform augmentation (GCA-T-A) across all datasets. Combining adaptive augmentation at both levels (GCA) achieves the highest performance across all benchmarks, providing up to a 1.60% absolute accuracy improvement on Amazon-Computers.

  10. Knowl 10 — Hyperparameter Sensitivity to Augmentation Probabilities

    empirical result

    Empirical sensitivity analysis of GCA with respect to the base edge removal probability pep_e and feature masking probability pfp_f on the Amazon-Photo dataset shows that node classification accuracy remains stable and high across a broad range of moderate values (0.1≤pe,pf≤0.50.1 \le p_e, p_f \le 0.5).

    When perturbation probabilities become excessively large (pe>0.5p_e > 0.5 or pf>0.5p_f > 0.5), accuracy degrades substantially (falling towards 80.0% at pe=0.9,pf=0.9p_e = 0.9, p_f = 0.9). Extreme edge removal rates disconnect nodes into isolated components in the generated views, preventing the GNN encoder from aggregating neighboring structural information and hindering the optimization of the contrastive objective.

Coverage note — Standard GCN layer formulations and dataset-specific hyperparameter grid specifications from Appendix A were omitted as standard baseline configurations.

References

  1. 1.Philip Bachman, R. Devon Hjelm, and William Buchwalter. 2019. Learning Representations by Maximizing Mutual Information Across Views. In Advances in Neural Information Processing Systems 32. 15509–15519.
  2. 2.Phillip Bonacich. 1987. Power and Centrality: A Family of Measures. Amer. J. Sociology 92, 5 (March 1987), 1170–1182.
  3. 3.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning, Vol. 119. PMLR, 10709–10719.
  4. 4.Ronan Collobert and Jason Weston. 2008. A Unified Architecture for Natural Language Processing: Deep Neural Networks with Multitask Learning. In Proceedings of the 25th International Conference on Machine Learning. ACM Press, 160–167.
  5. 5.Thomas M. Cover and Joy A. Thomas. 2006. Elements of Information Theory (Second Edition). Wiley-Interscience, USA.
  6. 6.William Falcon and Kyunghyun Cho. 2020. A Framework For Contrastive Self-Supervised Learning and Designing A New Approach. arXiv.org (Sept. 2020). arXiv:2009.00104v1 [cs.CV]
  7. 7.Matthias Fey and Jan Eric Lenssen. 2019. Fast Graph Representation Learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds.
  8. 8.Spyros Gidaris, Praveer Singh, and Nikos Komodakis. 2018. Unsupervised Representation Learning by Predicting Image Rotations. In Proceedings of the 6th International Conference on Learning Representations.
  9. 9.Xavier Glorot and Yoshua Bengio. 2010. Understanding the Difficulty of Training Deep Feedforward Neural Networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics. JMLR.org, 249–256.
  10. 10.Rafael C. Gonzalez and Richard E. Woods. 2018. Digital Image Processing (Fourth Edition). Pearson, USA.
  11. 11.Aditya Grover and Jure Leskovec. 2016. node2vec: Scalable Feature Learning for Networks. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 855–864.
  12. 12.Michael Gutmann and Aapo Hyvärinen. 2012. Noise-Contrastive Estimation of Unnormalized Statistical Models, with Applications to Natural Image Statistics. Journal of Machine Learning Research 13 (2012), 307–361.
  13. 13.Aric A. Hagberg, Daniel A. Schult, and Pieter J. Swart. 2008. Exploring Network Structure, Dynamics, and Function using NetworkX. In Proceedings of the 7th Python in Science Conference. 11–15.
  14. 14.William L. Hamilton, Rex Ying, and Jure Leskovec. 2017. Representation Learning on Graphs: Methods and Applications. Bulletin of the IEEE Computer Society Technical Committee on Data Engineering 40, 3 (2017), 52–74.
  15. 15.William L. Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive Representation Learning on Large Graphs. In Advances in Neural Information Processing Systems 30. 1024–1034.
  16. 16.Kaveh Hassani and Amir Hosein Khasahmadi. 2020. Contrastive Multi-View Representation Learning on Graphs. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119). PMLR, 3451–3461.
  17. 17.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 9726–9735.
  18. 18.Olivier J. Hénaff, Aravind Srinivas, Jeffrey De Fauw, Ali Razavi, Carl Doersch, S. M. Ali Eslami, and Aäron van den Oord. 2020. Data-Efficient Image Recognition with Contrastive Predictive Coding. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119). PMLR, 4182–4192.
  19. 19.R. Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Philip Bachman, Adam Trischler, and Yoshua Bengio. 2019. Learning Deep Representations by Mutual Information Estimation and Maximization. In Proceedings of the 7th International Conference on Learning Representations.
  20. 20.Fenyu Hu, Yanqiao Zhu, Shu Wu, Liang Wang, and Tieniu Tan. 2019. Hierarchical Graph Convolutional Networks for Semi-supervised Node Classification. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence. IJCAI.org, 4532–4539.
  21. 21.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In Proceedings of the 3rd International Conference on Learning Representations.
  22. 22.Thomas N. Kipf and Max Welling. 2016. Variational Graph Auto-Encoders. In Bayesian Deep Learning Workshop@NIPS.
  23. 23.Thomas N. Kipf and Max Welling. 2017. Semi-Supervised Classification with Graph Convolutional Networks. In Proceedings of the 5th International Conference on Learning Representations.
  24. 24.Johannes Klicpera, Stefan Weißenberger, and Stephan Günnemann. 2019. Diffusion Improves Graph Learning. In Advances in Neural Information Processing Systems 32. 13333–13345.
  25. 25.Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. 2017. Colorization as a Proxy Task for Visual Understanding. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 840–849.
  26. 26.Ralph Linsker. 1988. Self-Organization in a Perceptual Network. IEEE Computer 21, 3 (1988), 105–117.
  27. 27.Péter Mernyei and Catalina Cangea. 2020. Wiki-CS: A Wikipedia-Based Benchmark for Graph Neural Networks. In ICML Workshop on Graph Representation Learning and Beyond.
  28. 28.Andriy Mnih and Koray Kavukcuoglu. 2013. Learning Word Embeddings Efficiently with Noise-Contrastive Estimation. In Advances in Neural Information Processing Systems 26. 2265–2273.
  29. 29.Mark E. J. Newman. 2018. Networks: An Introduction (Second Edition). Oxford University Press.
  30. 30.Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. 1999. The PageRank Citation Ranking: Bringing Order to the Web. Technical Report. Stanford InfoLab.
  31. 31.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32. 8024–8035.
  32. 32.Zhen Peng, Wenbing Huang, Minnan Luo, Qinghua Zheng, Yu Rong, Tingyang Xu, and Junzhou Huang. 2020. Graph Representation Learning via Graphical Mutual Information Maximization. In Proceedings of the Web Conference 2020. ACM, 259–270.
  33. 33.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. GloVe: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing. ACL, 1532–1543.
  34. 34.Bryan Perozzi, Rami Al-Rfou, and Steven Skiena. 2014. DeepWalk: Online Learning of Social Representations. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 701–710.
  35. 35.Ben Poole, Sherjil Ozair, Aäron van den Oord, Alexander A. Alemi, and George Tucker. 2019. On Variational Bounds of Mutual Information. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97). PMLR, 5171–5180.
  36. 36.Jiezhong Qiu, Qibin Chen, Yuxiao Dong, Jing Zhang, Hongxia Yang, Ming Ding, Kuansan Wang, and Jie Tang. 2020. GCC: Graph Contrastive Coding for Graph Neural Network Pre-Training. In Proceedings of the 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. ACM, 1150–1160.
  37. 37.Jiezhong Qiu, Yuxiao Dong, Hao Ma, Jian Li, Kuansan Wang, and Jie Tang. 2018. Network Embedding as Matrix Factorization: Unifying DeepWalk, LINE, PTE, and node2vec. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining. ACM, 459–467.
  38. 38.Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015. FaceNet: A Unified Embedding for Face Recognition and Clustering. In Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 815–823.
  39. 39.Oleksandr Shchur, Maximilian Mumme, Aleksandar Bojchevski, and Stephan Günnemann. 2018. Pitfalls of Graph Neural Network Evaluation. arXiv.org (Nov. 2018). arXiv:1811.05868v2 [cs.LG]
  40. 40.Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan R. Salakhutdinov. 2014. Dropout: A Simple Way to Prevent Neural Networks From Overfitting. Journal of Machine Learning Research 15, 1 (2014), 1929–1958.
  41. 41.Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2019. Contrastive Multiview Coding. arXiv.org (June 2019). arXiv:1906.05849v4 [cs.CV]
  42. 42.Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. 2020. What Makes for Good Views for Contrastive Learning. arXiv.org (May 2020). arXiv:2005.10243v1 [cs.CV]
  43. 43.Michael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly, and Mario Lucic. 2020. On Mutual Information Maximization for Representation Learning. In Proceedings of the 8th International Conference on Learning Representations.
  44. 44.Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation Learning with Contrastive Predictive Coding. arXiv.org (2018). arXiv:1807.03748v2 [cs.LG]
  45. 45.Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018. Graph Attention Networks. In Proceedings of the 6th International Conference on Learning Representations.
  46. 46.Petar Veličković, William Fedus, William L. Hamilton, Pietro Liò, Yoshua Bengio, and R. Devon Hjelm. 2019. Deep Graph Infomax. In Proceedings of the 7th International Conference on Learning Representations.
  47. 47.Felix Wu, Tianyi Zhang, Amauri Holanda de Souza Jr., Christopher Fifty, Tao Yu, and Kilian Q. Weinberger. 2019. Simplifying Graph Convolutional Networks. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97). PMLR, 6861–6871.
  48. 48.Mike Wu, Chengxu Zhuang, Milan Mosse, Daniel Yamins, and Noah Goodman. 2020. On Mutual Information in Contrastive Learning for Visual Representations. arXiv.org (May 2020). arXiv:2005.13149v2 [cs.LG]
  49. 49.Zhirong Wu, Yuanjun Xiong, Stella X. Yu, and Dahua Lin. 2018. Unsupervised Feature Learning via Non-Parametric Instance Discrimination. In Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 3733–3742.
  50. 50.Tete Xiao, Xiaolong Wang, Alexei A. Efros, and Trevor Darrell. 2020. What Should Not Be Contrastive in Contrastive Learning. arXiv.org (Aug. 2020). arXiv:2008.05659v1 [cs.CV]
  51. 51.Mang Ye, Xu Zhang, Pong C. Yuen, and Shih-Fu Chang. 2019. Unsupervised Embedding Learning via Invariant and Spreading Instance Feature. In Proceedings of the 2019 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 6210–6219.
  52. 52.Wayne W. Zachary. 1977. An Information Flow Model for Conflict and Fission in Small Groups. Journal of Anthropological Research 33, 4 (1977), 452–473.
  53. 53.Yanqiao Zhu, Yichen Xu, Feng Yu, Qiang Liu, Shu Wu, and Liang Wang. 2020. Deep Graph Contrastive Representation Learning. In ICML Workshop on Graph Representation Learning and Beyond.

Citation

MLA
Zhu, Y., et al. “Graph Contrastive Learning with Adaptive Augmentation”. Proceedings of the Web Conference 2021, 2021, pp. 2069–80, https://doi.org/10.1145/3442381.3449802.
APA
Zhu, Y., Xu, Y., Yu, F., Liu, Q., Wu, S., & Wang, L. (2021). Graph Contrastive Learning with Adaptive Augmentation. Proceedings of the Web Conference 2021, 2069–2080. https://doi.org/10.1145/3442381.3449802
Chicago
Zhu, Y., Y. Xu, F. Yu, Q. Liu, S. Wu, and L. Wang. 2021. “Graph Contrastive Learning with Adaptive Augmentation”. Proceedings of the Web Conference 2021, 2069–80. https://doi.org/10.1145/3442381.3449802.
Harvard
Zhu, Y. et al. (2021) “Graph Contrastive Learning with Adaptive Augmentation”, Proceedings of the Web Conference 2021. ACM, pp. 2069–2080. Available at: https://doi.org/10.1145/3442381.3449802.
Vancouver
1. Zhu Y, Xu Y, Yu F, Liu Q, Wu S, Wang L (2021) Graph Contrastive Learning with Adaptive Augmentation. In: Proceedings of the Web Conference 2021. ACM, pp 2069–2080

BibTeX

@inproceedings{Zhu_2021, series={WWW ’21}, title={Graph Contrastive Learning with Adaptive Augmentation}, url={http://dx.doi.org/10.1145/3442381.3449802}, DOI={10.1145/3442381.3449802}, booktitle={Proceedings of the Web Conference 2021}, publisher={ACM}, author={Zhu, Yanqiao and Xu, Yichen and Yu, Feng and Liu, Qiang and Wu, Shu and Wang, Liang}, year={2021}, month=Apr, pages={2069–2080}, collection={WWW ’21} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/