MA-GCL: Model Augmentation Tricks for Graph Contrastive Learning

Xumeng GongCheng YangChuan Shi

article2023AAAI67 citations

Proposes a model augmentation paradigm for graph contrastive learning that perturbs the neural architectures of view encoders via layer asymmetry, randomized depth, and operator shuffling, generating diverse contrastive views without damaging underlying graph semantics.

Listen

Graph contrastive learning has emerged as a leading self-supervised method to extract meaningful representations from graph-structured data without requiring manual annotations. Contrastive learning relies on comparing two distinct perspectives, or views, of the same input to extract underlying shared patterns while discarding irrelevant noise. However, generating suitable views for graphs is notoriously difficult. Unlike images, where transformations such as cropping or rotating preserve underlying meaning, minor perturbations to graph topology—such as removing edges—risk destroying critical task-specific information. Furthermore, existing methods rely on identical neural network architectures for both views, producing outputs that are too similar to filter out noise effectively, while directly injecting parameter noise risks corrupting semantic representations.

The main objective of the article is to introduce and evaluate a new framework, Model Augmented Graph Contrastive Learning, which manipulates the internal architecture of the neural network encoders rather than perturbing input graphs or model parameters. The article demonstrates how this architectural variation produces diverse, noise-resilient representations across standard node classification benchmarks.

To evaluate this framework, the authors implemented three architecture modification techniques on top of a standard base graph neural network. First, the asymmetric strategy uses encoders with differing propagation depths to eliminate high-frequency noise. Second, the random strategy varies propagation depth from epoch to epoch to expand training data diversity. Third, the shuffling strategy alters the execution order of propagation and transformation layers to generate safe view variations. The approach was evaluated across six established benchmark datasets, encompassing academic citation and co-purchase networks, using both standard public data splits and randomized experimental splits.

The evaluation produced several key findings. First, applying the three architectural techniques allowed the base model to outperform recent state-of-the-art baselines on five of the six benchmark datasets, achieving relative performance gains of up to 2.7%. Second, ablation experiments confirmed that all three techniques provide measurable improvements: the asymmetric strategy yielded the largest individual gain, boosting accuracy by an average of 0.86%, while the random and shuffling strategies contributed average improvements of 0.65% and 0.54%, respectively. Third, mutual information analysis confirmed that the asymmetric strategy successfully pushed representations apart, reducing unneeded noise while preserving core semantic features.

These findings indicate that architectural diversity inside graph encoders can overcome the long-standing limitation of graph data augmentation without adding complex heuristic or adversarial pipelines. Practitioners can achieve competitive or superior representation quality simply by modifying layer arrangements in basic encoders, avoiding the computational overhead and tuning complexity associated with heavier frameworks. The results also caution against simple parameter-noise injections, which experimental evidence shows can degrade semantic integrity.

Organizations deploying graph neural networks should adopt architectural model augmentation as a plug-and-play enhancement for existing contrastive learning pipelines. While the reported performance gains are consistent across standard benchmark topologies, decision-makers should note that the current evaluation focuses primarily on node classification across homogeneous network structures. As recommended by the article, future validation should explore graphs characterized by heterophily, where connected nodes exhibit differing labels and characteristics, before universal deployment across diverse domain networks.

  • Paper: Graph Contrastive Learning with Augmentations, Yuning You et al. (2020). This foundational work establishes the core GraphCL framework of applying parameterized input graph augmentations and shared GNN encoders, which MA-GCL directly targets and replaces with architectural model augmentations.
  • Paper: Graph Contrastive Learning with Adaptive Augmentation, Yanqiao Zhu et al. (2020). This paper presents adaptive data-level augmentations for graph contrastive learning, representing the standard data-perturbation paradigm that MA-GCL critiques for limited diversity and noise.
  • Paper: Augmentation-Free Self-Supervised Learning on Graphs, Namkyeong Lee et al. (2022). This study analyzes how graph data augmentations arbitrarily alter underlying semantics, directly motivating MA-GCL's shift toward model-level view generation.
  • Paper: Contrastive Multi-View Representation Learning on Graphs, Kaveh Hassani et al. (2020). This work introduces multi-view graph contrastive learning using contrasting structural perspectives, providing key conceptual foundations for view generation in GCL.
  • Paper: What makes for good views for contrastive learning, Yonglong Tian et al. (2020). This seminal paper introduces the InfoMin principle for creating optimal contrastive views by minimizing task-irrelevant mutual information, which underpins the motivation for MA-GCL's view diversity tricks.
  • Paper: DropEdge: Towards Deep Graph Convolutional Networks on Node Classification, Yu Rong et al. (2019). This paper introduces edge dropping as a primary graph perturbation technique, serving as a primary baseline and comparison point for data augmentation in MA-GCL.
  • Paper: Deep Graph Infomax, Petar Veličković et al. (2019). This foundational work establishes unsupervised representation learning on graphs via mutual information maximization between contrasting views.
  • Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). This foundational text introduces the Graph Convolutional Network architecture that serves as the standard encoder backbone modified by MA-GCL's model augmentation tricks.

No sufficiently relevant recommendations were found.

Cover for MA-GCL: Model Augmentation Tricks for Graph Contrastive Learning

Abstract

Contrastive learning (CL), which can extract the information shared between different contrastive views, has become a popular paradigm for vision representation learning. Inspired by the success in computer vision, recent work introduces CL into graph modeling, dubbed as graph contrastive learning (GCL). However, generating contrastive views in graphs is more challenging than that in images, since we have little prior knowledge on how to significantly augment a graph without changing its labels. We argue that typical data augmentation techniques (e.g., edge dropping) in GCL cannot generate diverse enough contrastive views to filter out noises. Moreover, previous GCL methods employ two view encoders with exactly the same neural architecture and tied parameters, which further harms the diversity of augmented views. To address this limitation, we propose a novel paradigm named model augmented GCL (MA-GCL), which will focus on manipulating the architectures of view encoders instead of perturbing graph inputs. Specifically, we present three easy-to-implement model augmentation tricks for GCL, namely asymmetric, random and shuffling, which can respectively help alleviate high-frequency noises, enrich training instances and bring safer augmentations. All three tricks are compatible with typical data augmentations. Experimental results show that MA-GCL can achieve state-of-the-art performance on node classification benchmarks by applying the three tricks on a simple base model. Extensive studies also validate our motivation and the effectiveness of each trick. (Code, data and appendix are available at https://github.com/GXM1141/MA-GCL.)

Table of Contents

  • Introduction
  • Related Works
  • Notations and Preliminaries
  • Methodology
  • Asymmetric Strategy
  • Random Strategy
  • Shuffling Strategy
  • Implementation of MA-GCL
  • Experiments
  • Experimental Setup
  • Comparison with Baseline Methods
  • Ablation Studies
  • Motivation Verification
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Model-augmented graph contrastive learning

    model/method

    MA-GCL (Model Augmented Graph Contrastive Learning) generates contrastive views by changing the architectures of their GNN encoders, rather than relying only on graph-input perturbations or model-parameter noise. For an augmented graph with adjacency matrix ArA_r, feature matrix XrX_r, and degree matrix DrD_r, the propagation operator is gr(Z)=FrZg_r(Z)=F_rZ, where

    Fr=(1−π)I+πDr−1/2ArDr−1/2,π=0.5.F_r=(1-\pi)I+\pi D_r^{-1/2}A_rD_r^{-1/2},\qquad \pi=0.5.

    Here II is the identity matrix and ZZ is a node-embedding matrix. A transformation operator is hi(Z)=σ(ZWi)h_i(Z)=\sigma(ZW_i), where WiW_i is trainable and σ\sigma is a nonlinearity. The two view encoders share the transformation parameters but use different numbers and/or orderings of propagation operators:

    fr(Xr)=hN∘gr[Kr,N]∘hN−1∘⋯∘h1∘gr[Kr,1](Xr),f_r(X_r)=h_N\circ g_r^{[K_{r,N}]}\circ h_{N-1}\circ\cdots\circ h_1\circ g_r^{[K_{r,1}]}(X_r),

    where gr[K]g_r^{[K]} denotes KK repeated propagations and each Kr,iK_{r,i} is a nonnegative integer. The resulting view representations are optimized with a contrastive objective so that representations of two augmentations of the same node or graph are close while representations from different observations are separated. MA-GCL combines three architecture perturbations: asymmetric propagation depths, randomly sampled propagation depths, and shuffled operator permutations.

  2. Knowl 2 — Asymmetric encoders suppress high-frequency graph noise

    theoretical result

    The asymmetric strategy uses two view encoders with shared parameters, the same number of transformation operators, and different numbers L<L′L<L' of propagation operators. Under the paper's analysis, nonlinearities are omitted, node features are one-hot (X=I∣V∣X=I_{|V|}), and the two embeddings are Z=FLXWZ=F^LXW and Z′=FL′XWZ'=F^{L'}XW, where F=UΛUTF=U\Lambda U^T is a symmetric graph filter, UU is its orthonormal eigenvector matrix, Λ=diag⁡(λ1,…,λ∣V∣)\Lambda=\operatorname{diag}(\lambda_1,\ldots,\lambda_{|V|}) with λ1≥⋯≥λ∣V∣\lambda_1\geq\cdots\geq\lambda_{|V|}, and W∈R∣V∣×dOW\in\mathbb{R}^{|V|\times d_O} is constrained by WTW=IdOW^TW=I_{d_O}. The contrastive positive-pair term is

    min⁡Wtr⁡ ⁣((Z−Z′)(Z−Z′)T).\min_W\operatorname{tr}\!\left((Z-Z')(Z-Z')^T\right).

    The optimal transformation matrix is

    W∗=U[k1,k2,…,kdO],W^*=U[k_1,k_2,\ldots,k_{d_O}],

    where 1≤k1<⋯<kdO≤∣V∣1\leq k_1<\cdots<k_{d_O}\leq |V| are the indices that minimize (λkL−λkL′)2(\lambda_k^L-\lambda_k^{L'})^2. Thus, when the graph filter has eigenvalues in [0,1][0,1] and its largest eigenvalues correspond to task-relevant low-frequency components, the selected columns are likely to be the leading eigenvectors. In that case the asymmetric contrastive objective preferentially retains low-frequency information and suppresses high-frequency components. With identical encoders (L=L′L=L'), the unperturbed positive-pair difference is zero; the learned directions are then determined by graph-data-augmentation perturbations, which can cause high-frequency noise to enter the representation.

  3. Knowl 3 — Random propagation depths enlarge the effective training set

    model/method

    The random strategy samples the propagation depth independently in every training epoch instead of fixing it. For an SGC-style encoder f(X)=h∘g[L](X)f(X)=h\circ g^{[L]}(X), the paper samples an integer L^∼Uniform⁡(lower,upper)\hat L\sim\operatorname{Uniform}(\mathrm{lower},\mathrm{upper}) and uses L^\hat L propagations for that epoch. A node representation produced with LL propagations depends on the LL-hop computation tree rooted at that node. If S~L\widetilde{S}_L denotes the set of these computation trees for all training nodes, randomizing the depth exposes the learner to the union

    S~=⋃L^∈[lower,upper]S~L^,\widetilde{S}=\bigcup_{\hat L\in[\mathrm{lower},\mathrm{upper}]}\widetilde{S}_{\hat L},

    rather than only one fixed-depth set. The paper hypothesizes that this enlarges the diversity of training instances and improves downstream prediction. When graph data augmentation is also used, the effective instances include computation trees and their subtrees, so the difference between random-depth and fixed-depth training can become substantially larger, potentially exponential in the neighborhood structure.

  4. Knowl 4 — Operator shuffling provides topology-preserving view perturbations

    model/method

    The shuffling strategy changes the order in which propagation and nonlinear transformation operators are applied while leaving the input graph and encoder parameters unchanged. If an encoder has LL propagation operators and NN transformation operators, one ordering is

    f(X)=hN∘g[KN]∘hN−1∘⋯∘h1∘g[K1](X),f(X)=h_N\circ g^{[K_N]}\circ h_{N-1}\circ\cdots\circ h_1\circ g^{[K_1]}(X),

    where each Ki≥0K_i\geq 0 and ∑i=1NKi=L\sum_{i=1}^{N}K_i=L. The second view uses a different allocation K1′,…,KN′K'_1,\ldots,K'_N satisfying Ki′≥0K'_i\geq0 and ∑i=1NKi′=L\sum_{i=1}^{N}K'_i=L. The paper's design rationale is that reordering propagation and transformation operators does not alter the graph itself or destroy an edge, while nonlinear transformations still perturb the encoded representations. It therefore supplies a safer augmentation than random topology changes, and avoids the semantic risk of adding noise directly to model parameters.

  5. Knowl 5 — MA-GCL training and evaluation procedure

    algorithm

    The implemented MA-GCL model uses random edge and feature dropping, the filter F=12I+12D−1/2AD−1/2F=\tfrac12I+\tfrac12D^{-1/2}AD^{-1/2}, two shared transformation operators, a two-layer embedding projector, and an InfoNCE-style contrastive loss. The training procedure is:

    Input: Adjacency matrix AA; feature matrix XX; graph-augmentation distribution TT; integer propagation range [low,high][low, high]; transformation operators h1,h2,…,hNh_1, h_2, \ldots, h_N; projector projproj
    Output: Learned transformation operators h1,h2,…,hNh_1, h_2, \ldots, h_N
    for every training epoch do
        Draw two augmentations a,a′∼Ta,a' \sim T
        Set (A1,X1)=a(A,X)(A_1,X_1)=a(A,X) and (A2,X2)=a′(A,X)(A_2,X_2)=a'(A,X)
        For each i=1,…,Ni=1,\ldots,N, sample nonnegative integer depths Ki,Ki′K_i,K'_i from Uniform(low,high)Uniform(low,high)
        Require ∑iKi≠∑iKi′\sum_i K_i \neq \sum_i K'_i and Ki≠Ki′K_i \neq K'_i for every ii
        Compute F1=12I+12D1−1/2A1D1−1/2F_1=\frac12I+\frac12D_1^{-1/2}A_1D_1^{-1/2} and F2=12I+12D2−1/2A2D2−1/2F_2=\frac12I+\frac12D_2^{-1/2}A_2D_2^{-1/2}
        Construct f1=hN∘g1[KN]∘⋯∘h1∘g1[K1]f_1=h_N\circ g_1^{[K_N]}\circ\cdots\circ h_1\circ g_1^{[K_1]}
        Construct f2=hN∘g2[KN′]∘⋯∘h1∘g2[K1′]f_2=h_N\circ g_2^{[K'_N]}\circ\cdots\circ h_1\circ g_2^{[K'_1]}
        Compute Z1=f1(X1)Z_1=f_1(X_1) and Z2=f2(X2)Z_2=f_2(X_2)
        Update the shared transformation parameters by minimizing the contrastive loss L(proj(Z1),proj(Z2))\mathcal{L}(proj(Z_1),proj(Z_2))
    end for
    After training, discard the projector and use a fixed evaluation architecture with K1=⋯=KN=KK_1=\cdots=K_N=K, where the experiments use K∈{1,2}K\in\{1,2\}.
  6. Knowl 6 — Node-classification evaluation protocol

    experimental setup

    MA-GCL was evaluated for unsupervised node representation learning on six benchmark graphs: Cora, CiteSeer, PubMed, Coauthor-CS, Amazon-Computers, and Amazon-Photo. Cora, CiteSeer, and PubMed use their public train/validation/test splits. Coauthor-CS, Amazon-Computers, and Amazon-Photo use random splits containing 10% training nodes, 10% validation nodes, and 80% test nodes. After unsupervised representation learning, the same linear classifier is trained as a post-processing step, and classification accuracy is reported. Each result is the mean and standard deviation over five runs; for the randomly split datasets, the random seeds also change the splits. MA-GCL uses two transformation operators, with evaluation propagation depths K1=K2=2K_1=K_2=2 on Cora and CiteSeer and K1=K2=1K_1=K_2=1 on the other datasets.

  7. Knowl 7 — MA-GCL outperforms unsupervised GCL baselines on public splits

    data/table

    The following results compare node-classification accuracy on the six benchmark datasets using public splits for Cora, CiteSeer, and PubMed. Values are mean ±\pm standard deviation over five runs. The unsupervised methods are compared under the same linear-evaluation protocol; GCN, GAT, and InfoGCL are label-trained reference methods. MA-GCL is the best unsupervised method on five of six datasets, has the highest average accuracy and best average rank among the reported methods, and improves over the best baseline by as much as 2.7% according to the paper.

    Could not parse LaTeX table
  8. Knowl 8 — MA-GCL remains stronger under random splits

    empirical result

    On Cora, CiteSeer, and PubMed with random rather than public splits, MA-GCL again outperforms every listed baseline under the same linear-evaluation protocol. The reported five-run accuracies are:

    Could not parse LaTeX table

    The result supports the paper's claim that the architecture-based augmentation is not dependent on the particular public splits.

  9. Knowl 9 — All three architecture tricks contribute, with asymmetry strongest

    data/table

    The ablation study compares the base model with every one-strategy variant, every two-strategy variant, and the full MA-GCL model. Here A denotes asymmetric depths, R denotes random depths across epochs, and S denotes shuffled operator order. The full model has the highest average accuracy, while removing A causes the largest average degradation among the two-strategy variants; the paper reports average gains of 0.86%, 0.65%, and 0.54% when adding A, R, and S, respectively, to an ablated model.

    Could not parse LaTeX table
  10. Knowl 10 — View-information measurements support the proposed motivation

    empirical result

    The paper estimates task-relevant information and view-shared information for five variants of the same backbone on CiteSeer and Amazon-Photo. Downstream classification accuracy is used as an estimate of task-relevant information, while MINE-estimated mutual information between the two learned views is used as an estimate of task-relevant plus task-irrelevant shared information. The compared variants are the base model with random graph augmentations, the base model with parameter perturbations, Base Model+A, MA-GCL without A, and full MA-GCL.

    The base model consistently has the highest mutual information but the lowest classification accuracies, supporting the claim that ordinary graph augmentations make the two views too similar and preserve excessive noise. Parameter perturbation produces both low mutual information and low accuracy, indicating that it can damage encoded semantics. Adding asymmetric propagation depths lowers the mutual information while improving downstream accuracy, both when added to the base model and when added to the other two MA-GCL strategies. These measurements support the intended trade-off of keeping task-relevant information while reducing irrelevant shared noise.

Coverage note — Appendix-only graph-classification and node-clustering results, hyperparameter-sensitivity analyses, efficiency measurements, and additional dataset-specific motivation plots were omitted because the main paper identifies them but does not provide their detailed results in the supplied text.

References

  1. 1.Belghazi, M. I.; Baratin, A.; Rajeshwar, S.; Ozair, S.; Bengio, Y.; Courville, A.; and Hjelm, D. 2018. Mutual information neural estimation. In ICML.
  2. 2.Chen, M.; Wei, Z.; Huang, Z.; Ding, B.; and Li, Y. 2020. Simple and deep graph convolutional networks. In ICML.
  3. 3.Cui, G.; Zhou, J.; Yang, C.; and Liu, Z. 2020. Adaptive graph encoder for attributed graph embedding. In KDD.
  4. 4.Duan, H.; Vaezipoor, P.; Paulus, M. B.; Ruan, Y.; and Maddison, C. 2022. Augment with Care: Contrastive Learning for Combinatorial Problems. In ICML.
  5. 5.Feng, S.; Jing, B.; Zhu, Y.; and Tong, H. 2022a. Adversarial graph contrastive learning with information regularization. In WWW.
  6. 6.Feng, W.; Dong, Y.; Huang, T.; Yin, Z.; Cheng, X.; Kharlamov, E.; and Tang, J. 2022b. GRAND+: Scalable Graph Random Neural Networks. In WWW.
  7. 7.Feng, W.; Zhang, J.; Dong, Y.; Han, Y.; Luan, H.; Xu, Q.; Yang, Q.; Kharlamov, E.; and Tang, J. 2020. Graph random neural networks for semi-supervised learning on graphs. NeurIPS.
  8. 8.Grill, J.-B.; Strub, F.; Altche, F.; Tallec, C.; Richemond, ´P.; Buchatskaya, E.; Doersch, C.; Avila Pires, B.; Guo, Z.; Gheshlaghi Azar, M.; et al. 2020. Bootstrap your own latent-a new approach to self-supervised learning. NeurIPS.
  9. 9.Hassani, K.; and Khasahmadi, A. H. 2020. Contrastive multi-view representation learning on graphs. In ICML.
  10. 10.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020. Momentum contrast for unsupervised visual representation learning. In CVPR, 9729–9738.
  11. 11.Hjelm, R. D.; Fedorov, A.; Lavoie-Marchildon, S.; Grewal, K.; Bachman, P.; Trischler, A.; and Bengio, Y. 2019. Learning deep representations by mutual information estimation and maximization. ICLR.
  12. 12.Jing, L.; Vincent, P.; LeCun, Y.; and Tian, Y. 2021. Understanding dimensional collapse in contrastive self-supervised learning. ICLR.
  13. 13.Kipf, T. N.; and Welling, M. 2016. Semi-supervised classification with graph convolutional networks. ICLR.
  14. 14.Li, S.; Wang, X.; Zhang, A.; Wu, Y.; He, X.; and Chua, T.-S. 2022. Let Invariant Rationale Discovery Inspire Graph Contrastive Learning. In ICML.
  15. 15.Nt, H.; and Maehara, T. 2019. Revisiting graph neural networks: All we have is low-pass filters. NeurIPS.
  16. 16.Qiu, J.; Chen, Q.; Dong, Y.; Zhang, J.; Yang, H.; Ding, M.; Wang, K.; and Tang, J. 2020. Gcc: Graph contrastive coding for graph neural network pre-training. In KDD.
  17. 17.Shchur, O.; Mumme, M.; Bojchevski, A.; and Gunnemann, ¨S. 2018. Pitfalls of graph neural network evaluation. arXiv preprint arXiv:1811.05868.
  18. 18.Soatto, S.; and Chiuso, A. 2016. Modeling Visual representations: Defining properties and deep approximations. ICLR.
  19. 19.Sun, F.-Y.; Hoffman, J.; Verma, V.; and Tang, J. 2020. InfoGraph: Unsupervised and Semi-supervised Graph-Level Representation Learning via Mutual Information Maximization. In ICLR.
  20. 20.Suresh, S.; Li, P.; Hao, C.; and Neville, J. 2021. Adversarial graph augmentation to improve graph contrastive learning. NeurIPS.
  21. 21.Thakoor, S.; Tallec, C.; Azar, M. G.; Munos, R.; Velickovi ˇ c,´ P.; and Valko, M. 2021. Bootstrapped representation learning on graphs. In ICLR 2021 Workshop on Geometrical and Topological Representation Learning.
  22. 22.Tian, Y.; Sun, C.; Poole, B.; Krishnan, D.; Schmid, C.; and Isola, P. 2020. What makes for good views for contrastive learning? NeurIPS.
  23. 23.TISHBY, N. 1999. The information bottleneck method. In Proc. 37th Annual Allerton Conference on Communications, Control and Computing.
  24. 24.Tong, Z.; Liang, Y.; Ding, H.; Dai, Y.; Li, X.; and Wang, C. 2021. Directed Graph Contrastive Learning. NeurIPS.
  25. 25.Van den Oord, A.; Li, Y.; and Vinyals, O. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748.
  26. 26.Velickovic, P.; Cucurull, G.; Casanova, A.; Romero, A.; Lio, P.; and Bengio, Y. 2017. Graph attention networks. ICLR.
  27. 27.Velickovic, P.; Fedus, W.; Hamilton, W. L.; Lio, P.; Bengio, `Y.; and Hjelm, R. D. 2019. Deep Graph Infomax. ICLR.
  28. 28.Wang, M.; Yu, L.; Da Zheng, Q. G.; Gai, Y.; Ye, Z.; Li, M.; Zhou, J.; Huang, Q.; Ma, C.; et al. 2019. Deep Graph Library: Towards efficient and scalable deep learning on graphs.(2019). arXiv preprint arXiv:1909.01315.
  29. 29.Wang, T.; and Isola, P. 2020. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In ICML.
  30. 30.Wu, F.; Souza, A.; Zhang, T.; Fifty, C.; Yu, T.; and Weinberger, K. 2019. Simplifying graph convolutional networks. In ICML.
  31. 31.Xia, J.; Wu, L.; Chen, J.; Hu, B.; and Li, S. Z. 2022. SimGRACE: A Simple Framework for Graph Contrastive Learning without Data Augmentation. In WWW.
  32. 32.Xu, D.; Cheng, W.; Luo, D.; Chen, H.; and Zhang, X. 2021. Infogcl: Information-aware graph contrastive learning. NeurIPS.
  33. 33.Xu, K.; Hu, W.; Leskovec, J.; and Jegelka, S. 2018. How Powerful are Graph Neural Networks? In ICLR.
  34. 34.Yang, H.; Chen, H.; Pan, S.; Li, L.; Yu, P. S.; and Xu, G. 2022. Dual Space Graph Contrastive Learning. WWW.
  35. 35.Yang, L.; Zhang, L.; and Yang, W. 2021. Graph Adversarial Self-Supervised Learning. NeurIPS.
  36. 36.You, Y.; Chen, T.; Shen, Y.; and Wang, Z. 2021. Graph contrastive learning automated. In ICML.
  37. 37.You, Y.; Chen, T.; Sui, Y.; Chen, T.; Wang, Z.; and Shen, Y. 2020. Graph contrastive learning with augmentations. NeurIPS.
  38. 38.Yu, J.; Yin, H.; Xia, X.; Chen, T.; Cui, L.; and Nguyen, Q. V. H. 2022. Are graph augmentations necessary? Simple graph contrastive learning for recommendation. In SIGIR.
  39. 39.Yuan, J.; Yu, H.; Cao, M.; Xu, M.; Xie, J.; and Wang, C. 2021. Semi-Supervised and Self-Supervised Classification with Multi-View Graph Neural Networks. In CIKM.
  40. 40.Zhang, H.; Wu, Q.; Yan, J.; Wipf, D.; and Yu, P. S. 2021. From canonical correlation analysis to self-supervised graph neural networks. NeurIPS.
  41. 41.Zhou, J.; Cui, G.; Hu, S.; Zhang, Z.; Yang, C.; Liu, Z.; Wang, L.; Li, C.; and Sun, M. 2020. Graph neural networks: A review of methods and applications. AI Open.
  42. 42.Zhu, H.; Sun, K.; and Koniusz, P. 2021. Contrastive laplacian eigenmaps. NeurIPS.
  43. 43.Zhu, Y.; Xu, Y.; Yu, F.; Liu, Q.; Wu, S.; and Wang, L. 2020. Deep graph contrastive representation learning. arXiv preprint arXiv:2006.04131.
  44. 44.Zhu, Y.; Xu, Y.; Yu, F.; Liu, Q.; Wu, S.; and Wang, L. 2021. Graph contrastive learning with adaptive augmentation. In WWW.

Citation

MLA
Gong, X., et al. “MA-GCL: Model Augmentation Tricks for Graph Contrastive Learning”. arXiv, 2022, http://arxiv.org/abs/2212.07035v1.
APA
Gong, X., Yang, C., & Shi, C. (2022). MA-GCL: Model Augmentation Tricks for Graph Contrastive Learning. arXiv. http://arxiv.org/abs/2212.07035v1
Chicago
Gong, X., C. Yang, and C. Shi. 2022. “MA-GCL: Model Augmentation Tricks for Graph Contrastive Learning”. arXiv. http://arxiv.org/abs/2212.07035v1.
Harvard
Gong, X., Yang, C. and Shi, C. (2022) “MA-GCL: Model Augmentation Tricks for Graph Contrastive Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.07035v1.
Vancouver
1. Gong X, Yang C, Shi C (2022) MA-GCL: Model Augmentation Tricks for Graph Contrastive Learning. arXiv

BibTeX

@article{gong2022gcl,
  title = {MA-GCL: Model Augmentation Tricks for Graph Contrastive Learning},
  author = {Gong, Xumeng and Yang, Cheng and Shi, Chuan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.07035v1},
  eprint = {2212.07035}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF