Robust Optimization as Data Augmentation for Large-scale Graphs

Kezhi KongGuohao LiMucong DingZuxuan WuChen ZhuBernard GhanemGavin TaylorTom Goldstein

article2022CVPR109 citations

Proposes FLAG, a model-free adversarial data augmentation method that perturbs node features during training to improve graph neural network generalization across node classification, link prediction, and graph property tasks with minimal computational overhead.

Listen

Graph Neural Networks have become essential tools for learning from networked data in areas such as recommender systems, social media analysis, and molecular modeling. However, training these models on large-scale datasets frequently causes severe overfitting, and real-world graphs present high volumes of out-of-distribution samples that degrade predictive accuracy. Traditional data augmentation techniques, such as cropping or flipping in computer vision, cannot be easily applied to graph node features because those features are typically discrete embeddings like bag-of-words or categorical codes. While existing graph regularization methods attempt to overcome this by altering network topology—such as adding or removing edges—they are often computationally expensive and limited to specific tasks.

The article introduces and evaluates Free Large-scale Adversarial Augmentation on Graphs (FLAG), a scalable, general-purpose data augmentation technique designed to improve Graph Neural Network performance and generalization. The method demonstrates that adding iteratively calculated adversarial perturbations directly to node features, while keeping graph topological structures intact, reliably enhances model generalization across diverse prediction tasks.

To evaluate this approach, the researchers conducted extensive empirical testing on competitive baseline graph architectures using the standardized Open Graph Benchmark across node classification, link prediction, and graph classification tasks. FLAG generates multi-scale gradient-based feature perturbations during the standard backward training passes. By accumulating parameter gradients across iterative batch passes rather than updating model weights sequentially, the method produces diverse perturbation scales without requiring significant extra graphics memory or extensive computational time.

The experimental findings show that FLAG consistently boosts clean accuracy across diverse benchmarks. In node classification on large-scale benchmark datasets, adding FLAG increased Graph Attention Network test accuracy by an absolute 2.31%. In link prediction tasks, the method provided massive accuracy gains, boosting baseline hit rates on biological network datasets from 37.07% to 51.41% for Graph Convolutional Networks and from 53.90% to 63.31% for GraphSAGE. In molecular graph classification, it delivered solid improvements, including a 1.31% increase in average precision for Graph Isomorphism Networks. Furthermore, ablation studies demonstrated that FLAG is fully complementary with other standard regularizers like dropout, virtual nodes, and batch normalization.

These findings indicate that adversarial training can serve as a highly effective data regularizer rather than an accuracy-degrading defensive measure on graph data. The authors trace this benefit to the discrete nature of input node features, where adversarial noise bridges sparse gaps in the input distribution without altering core semantic labels. Because FLAG can be implemented in standard deep learning frameworks with roughly a dozen lines of code and incurs virtually no additional training overhead when paired with reduced epoch schedules, it delivers immediate performance improvements at minimal computational cost.

Organizations developing graph-based machine learning systems should integrate FLAG into their training pipelines across node, link, and graph classification workflows. When adopting the technique, practitioners should apply larger perturbation magnitudes to unlabeled nodes than labeled nodes during semi-supervised learning to maximize multi-scale diversity. Teams using mini-batch graph training should pair the augmentation with neighborhood sampling or sub-graph sampling techniques rather than cluster-based partitioning, as cluster partitioning degraded performance in testing.

While confidence in the empirical results is high across standard benchmark datasets, the approach carries specific limitations. For graphs that lack initial node features, the effectiveness of the perturbation depends heavily on how artificial features are engineered; sum-aggregated features benefited from FLAG, whereas mean-aggregated features did not. Additionally, the theoretical foundations explaining why adversarial perturbations benefit discrete data distributions while harming continuous ones remain an open research question requiring further mathematical formalization.

arXiv: 2010.09891
Cover for Robust Optimization as Data Augmentation for Large-scale Graphs

Abstract

Data augmentation helps neural networks generalize better by enlarging the training set, but it remains an open question how to effectively augment graph data to enhance the performance of GNNs (Graph Neural Networks). While most existing graph regularizers focus on manipulating graph topological structures by adding/removing edges, we offer a method to augment node features for better performance. We propose FLAG (Free Large-scale Adversarial Augmentation on Graphs), which iteratively augments node features with gradient-based adversarial perturbations during training. By making the model invariant to small fluctuations in input data, our method helps models generalize to out-of-distribution samples and boosts model performance at test time. FLAG is a general-purpose approach for graph data, which universally works in node classification, link prediction, and graph classification tasks. FLAG is also highly flexible and scalable, and is deployable with arbitrary GNN backbones and large-scale datasets. We demonstrate the efficacy and stability of our method through extensive experiments and ablation studies. We also provide intuitive observations for a deeper understanding of our method. We open source our implementation at https://github.com/devnkong/FLAG.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries and Related Work
  • 3. Proposed Method
  • 4. Experiments
  • 5. Ablation Studies and Discussions
  • References
  • 6. Conclusion

Knowls

  1. Knowl 1 — FLAG Training Algorithm for Semi-Supervised Node Classification

    algorithm

    Free Large-scale Adversarial Augmentation on Graphs (FLAG) performs data augmentation in the node feature space of Graph Neural Networks (GNNs) by iteratively crafting gradient-based adversarial perturbations across MM ascent steps, accumulating model parameter gradients on the backward passes, and updating the model once per minibatch/epoch.

    Input: Graph G=(V,E)\mathcal{G} = (\mathcal{V}, \mathcal{E}), labeled node subset Vl⊂V\mathcal{V}_l \subset \mathcal{V}, learning rate τ\tau, number of ascent steps MM, labeled node step size αv\alpha_v, unlabeled node step size αu\alpha_u, objective loss function L(⋅)L(\cdot), GNN aggregation AGGREGATEθ(k)\text{AGGREGATE}_{\theta}^{(k)} and combine COMBINEϕ(k)\text{COMBINE}_{\phi}^{(k)} operations for k∈{1,…,K}k \in \{1, \dots, K\}
    Output: Trained model parameters (θ,ϕ)(\boldsymbol{\theta}, \boldsymbol{\phi})
    Initialize parameters (θ,ϕ)(\boldsymbol{\theta}, \boldsymbol{\phi})
    for each labeled node batch v∈Vlv \in \mathcal{V}_l do
        Initialize noise δv(0)∼U(−αv,αv)\boldsymbol{\delta}_v^{(0)} \sim \mathcal{U}(-\alpha_v, \alpha_v) and δu(0)∼U(−αu,αu)\boldsymbol{\delta}_u^{(0)} \sim \mathcal{U}(-\alpha_u, \alpha_u) for unlabeled neighbors u∈N(v)u \in \mathcal{N}(v)
        Initialize parameter gradient accumulator gθ,ϕ(0)←0\boldsymbol{g}_{\boldsymbol{\theta},\boldsymbol{\phi}}^{(0)} \leftarrow \boldsymbol{0}
        for t=1t = 1 to MM do
            Set initial layer representations hv(0)←xv+δv(t−1)h_v^{(0)} \leftarrow x_v + \boldsymbol{\delta}_v^{(t-1)} and hu(0)←xu+δu(t−1)h_u^{(0)} \leftarrow x_u + \boldsymbol{\delta}_u^{(t-1)}
            for k=1k = 1 to KK do
                msgv(k)←AGGREGATEθ(k)({(hv(k−1),hu(k−1),euv):u∈N(v)})msg_v^{(k)} \leftarrow \text{AGGREGATE}_{\boldsymbol{\theta}}^{(k)}\left( \left\{ (h_v^{(k-1)}, h_u^{(k-1)}, e_{uv}) : u \in \mathcal{N}(v) \right\} \right)
                hv(k)←COMBINEϕ(k)(hv(k−1),msgv(k))h_v^{(k)} \leftarrow \text{COMBINE}_{\boldsymbol{\phi}}^{(k)}\left( h_v^{(k-1)}, msg_v^{(k)} \right)
            end for
            Compute loss L(hv(K),y)L(h_v^{(K)}, y)
            Execute backpropagation to obtain gradients ∇θ,ϕL\nabla_{\boldsymbol{\theta},\boldsymbol{\phi}} L and ∇δL\nabla_{\boldsymbol{\delta}} L
            Accumulate parameter gradients: gθ,ϕ(t)←gθ,ϕ(t−1)+1M∇θ,ϕL\boldsymbol{g}_{\boldsymbol{\theta},\boldsymbol{\phi}}^{(t)} \leftarrow \boldsymbol{g}_{\boldsymbol{\theta},\boldsymbol{\phi}}^{(t-1)} + \frac{1}{M} \nabla_{\boldsymbol{\theta},\boldsymbol{\phi}} L
            Update perturbations:
                δv(t)←δv(t−1)+αv⋅sign(∇δvL)\boldsymbol{\delta}_v^{(t)} \leftarrow \boldsymbol{\delta}_v^{(t-1)} + \alpha_v \cdot \text{sign}\left(\nabla_{\boldsymbol{\delta}_v} L\right)
                δu(t)←δu(t−1)+αu⋅sign(∇δuL)\boldsymbol{\delta}_u^{(t)} \leftarrow \boldsymbol{\delta}_u^{(t-1)} + \alpha_u \cdot \text{sign}\left(\nabla_{\boldsymbol{\delta}_u} L\right)
        end for
        Update parameters: (θ,ϕ)←(θ,ϕ)−τ⋅gθ,ϕ(M)(\boldsymbol{\theta}, \boldsymbol{\phi}) \leftarrow (\boldsymbol{\theta}, \boldsymbol{\phi}) - \tau \cdot \boldsymbol{g}_{\boldsymbol{\theta},\boldsymbol{\phi}}^{(M)}
    end for

    In standard implementations, the ascent step count is typically set to M=3M = 3.

  2. Knowl 2 — Multi-Scale Gradient Accumulation in FLAG

    model/method

    In robust optimization, standard Projected Gradient Descent (PGD) runs MM ascent steps to find a single worst-case perturbation δM\boldsymbol{\delta}_M and updates model parameters θ\boldsymbol{\theta} only once using δM\boldsymbol{\delta}_M, discarding intermediate noise scales. Standard "Free" adversarial training updates model parameters θt+1\boldsymbol{\theta}_{t+1} during step tt using the gradient evaluated at θt\boldsymbol{\theta}_t, leading to suboptimal asynchronous optimization.

    FLAG resolves both issues by freezing the model parameters θi\boldsymbol{\theta}_i during the MM batch replay steps, computing model weight gradients ∇θL(fθi(x+δt),y)\nabla_{\boldsymbol{\theta}} L(f_{\boldsymbol{\theta}_i}(x + \boldsymbol{\delta}_t), y) across all intermediate perturbation scales δt\boldsymbol{\delta}_t (t∈{1,…,M}t \in \{1, \dots, M\}), and applying a unified parameter update step:

    θi+1=θi−τM∑t=1M∇θL(fθi(x+δt),y)\boldsymbol{\theta}_{i+1} = \boldsymbol{\theta}_i - \frac{\tau}{M} \sum_{t=1}^{M} \nabla_{\boldsymbol{\theta}} L \left( f_{\boldsymbol{\theta}_i}(x + \boldsymbol{\delta}_t), y \right)

    where τ\tau is the learning rate, xx is the clean input feature vector, yy is the ground-truth label, and δ1∼U(−α,α)\boldsymbol{\delta}_1 \sim \mathcal{U}(-\alpha, \alpha). Because each perturbation δt\boldsymbol{\delta}_t reaches an effective scale bounded by t⋅αt \cdot \alpha, this accumulation exposes the model to multi-scale feature augmentations without increasing GPU memory usage beyond standard backward gradient accumulation.

  3. Knowl 3 — Weighted Perturbation Scheme for Labeled vs. Unlabeled Nodes

    model/method

    In semi-supervised node classification on graphs, the recursive message-passing formulation aggregates representations from a target node's kk-hop neighborhood. Because distant and unlabeled neighbor nodes have a more diffuse influence on the target node's classification decision than the target node itself, FLAG applies a larger perturbation magnitude to unlabeled nodes than to labeled nodes.

    Let αv\alpha_v denote the ascent step size for target labeled nodes v∈Vlv \in \mathcal{V}_l, and αu\alpha_u denote the ascent step size for unlabeled neighbor nodes u∈V∖Vlu \in \mathcal{V} \setminus \mathcal{V}_l. The perturbation ratio is parameterized by log⁡2(αu/αv)\log_2(\alpha_u / \alpha_v). When log⁡2(αu/αv)>0\log_2(\alpha_u / \alpha_v) > 0 (i.e., αu>αv\alpha_u > \alpha_v), the model enforces stronger smoothness constraints on neighborhood context. This weighting yields more pronounced accuracy gains on datasets with high label sparsity (e.g., ogbn-products, where the label rate is 8%) than on datasets with dense labels (e.g., ogbn-arxiv, where the label rate is 54%).

  4. Knowl 4 — Feature-Space Perturbation and Unprojected Ascent Strategy

    model/method

    Unlike standard adversarial training in computer vision which enforces an explicit ℓ∞\ell_\infty-norm projection constraint Π∥δ∥∞≤ϵ\Pi_{\|\boldsymbol{\delta}\|_{\infty} \le \epsilon} to maintain human visual imperceptibility, FLAG discards explicit projection bounds on node feature perturbations for graph data. In graph learning, distance thresholds for imperceptibility lack a natural perceptual meaning, and larger perturbations have been observed to aid out-of-distribution generalization. In FLAG, the maximum perturbation is implicitly bounded by the step size α\alpha multiplied by the number of ascent steps MM (i.e., ∥δM∥∞≤Mα\|\boldsymbol{\delta}_M\|_\infty \le M\alpha).

    When input node features are categorical or discrete (e.g., one-hot tokens, atom types), gradient ascent directly on discrete inputs is undefined. In such cases, FLAG projects the discrete node features into a continuous hidden embedding space via an initial embedding layer, and applies the adversarial perturbations δ\boldsymbol{\delta} directly to those continuous hidden embeddings.

  5. Knowl 5 — Evaluation of FLAG on Large-Scale OGB Node Property Prediction Benchmarks

    data/table

    FLAG was evaluated across multiple standard GNN backbones (GCN, GraphSAGE, GAT, DeeperGCN, and R-GCN) on large-scale Open Graph Benchmark (OGB) node classification datasets: ogbn-products, ogbn-proteins, ogbn-arxiv, and the heterogeneous graph ogbn-mag. Results report the mean and standard deviation over 10 independent runs.

    Backbone ogbn-products (Test Acc %) ogbn-proteins (Test ROC-AUC %) ogbn-arxiv (Test Acc %)
    GCN - 72.51 ±\pm 0.35 71.74 ±\pm 0.29
    +FLAG - 71.71 ±\pm 0.50 72.04 ±\pm 0.20
    GraphSAGE 78.70 ±\pm 0.36 77.68 ±\pm 0.20 71.49 ±\pm 0.27
    +FLAG 79.36 ±\pm 0.57 76.57 ±\pm 0.75 72.19 ±\pm 0.21
    GAT 79.45 ±\pm 0.59 - 73.65 ±\pm 0.11
    +FLAG 81.76 ±\pm 0.45 - 73.71 ±\pm 0.13
    DeeperGCN 80.98 ±\pm 0.20 85.80 ±\pm 0.17 71.92 ±\pm 0.16
    +FLAG 81.93 ±\pm 0.31 85.96 ±\pm 0.27 72.14 ±\pm 0.19

    On the heterogeneous graph dataset ogbn-mag, applying FLAG to R-GCN increases test accuracy from 46.78±0.67%46.78 \pm 0.67\% to 47.37±0.48%47.37 \pm 0.48\%. On ogbn-products, FLAG improves GAT by +2.31%+2.31\% absolute accuracy. For ogbn-proteins (which has no raw input node features), DeeperGCN using summed incoming edge features improves from 85.80%85.80\% to 85.96%85.96\%, whereas GCN/GraphSAGE using mean edge features experience minor degradation.

  6. Knowl 6 — Evaluation of FLAG on Large-Scale OGB Graph Property Prediction Benchmarks

    data/table

    FLAG was evaluated across molecular and biological graph classification benchmarks from the Open Graph Benchmark: ogbg-molhiv, ogbg-molpcba, ogbg-ppa, and ogbg-code. Backbones include GCN, GIN, and DeeperGCN, with and without virtual nodes (denoted by "Virtual"). Discrete features were mapped to continuous hidden spaces before perturbation. Mean and standard deviations over 10 runs are reported below.

    Backbone ogbg-molhiv (ROC-AUC %) ogbg-molpcba (AP %) ogbg-ppa (Acc %) ogbg-code (F1 %)
    GCN 76.06 ±\pm 0.97 20.20 ±\pm 0.24 68.39 ±\pm 0.34 31.63 ±\pm 0.18
    +FLAG 76.83 ±\pm 1.02 21.16 ±\pm 0.17 68.38 ±\pm 0.47 32.09 ±\pm 0.19
    GCN-Virtual 75.99 ±\pm 1.19 24.24 ±\pm 0.34 68.57 ±\pm 0.61 32.63 ±\pm 0.13
    +FLAG 75.45 ±\pm 1.58 24.83 ±\pm 0.37 69.44 ±\pm 0.52 33.16 ±\pm 0.25
    GIN 75.58 ±\pm 1.40 22.66 ±\pm 0.28 68.92 ±\pm 1.00 31.63 ±\pm 0.20
    +FLAG 76.54 ±\pm 1.14 23.95 ±\pm 0.40 69.05 ±\pm 0.92 32.41 ±\pm 0.40
    GIN-Virtual 77.07 ±\pm 1.49 27.03 ±\pm 0.23 70.37 ±\pm 1.07 32.04 ±\pm 0.18
    +FLAG 77.48 ±\pm 0.96 28.34 ±\pm 0.38 72.45 ±\pm 1.14 32.96 ±\pm 0.36
    DeeperGCN 78.58 ±\pm 1.17 27.81 ±\pm 0.38 77.12 ±\pm 0.71 -
    +FLAG 79.42 ±\pm 1.20 28.42 ±\pm 0.43 77.52 ±\pm 0.69 -

    FLAG consistently enhances prediction metrics across backbones, with notable gains on ogbg-molpcba (e.g., GIN-Virtual +1.31%+1.31\% AP) and ogbg-ppa (GIN-Virtual +2.08%+2.08\% Acc).

  7. Knowl 7 — Evaluation of FLAG on Large-Scale OGB Link Property Prediction Benchmarks

    data/table

    FLAG was tested on two link prediction datasets from the Open Graph Benchmark: ogbl-ddi and ogbl-collab, evaluating GCN and GraphSAGE backbones under full-batch training. Performance is measured via Hits@20 for ogbl-ddi and Hits@50 for ogbl-collab across 10 random seeds.

    Backbone ogbl-ddi (Hits@20 %) ogbl-collab (Hits@50 %)
    GCN 37.07 ±\pm 5.07 44.75 ±\pm 1.07
    +FLAG 51.41 ±\pm 3.76 46.22 ±\pm 0.81
    GraphSAGE 53.90 ±\pm 4.74 48.10 ±\pm 0.81
    +FLAG 63.31 ±\pm 6.06 48.44 ±\pm 0.40

    On ogbl-ddi, adding FLAG produces absolute test performance improvements of +14.34%+14.34\% Hits@20 for GCN and +9.41%+9.41\% Hits@20 for GraphSAGE.

  8. Knowl 8 — Empirical Comparison of FLAG Against PGD and Free Adversarial Training

    data/table

    Performance of GNN baselines (GAT on ogbn-products, GraphSAGE on ogbl-ddi, and GIN on ogbg-molhiv) was compared when augmented using standard PGD (with M=8M=8 ascent steps), standard "Free" adversarial training (with M=8M=8 steps), FLAG (M=3M=3 steps), and FLAG (fast, with reduced total training epochs).

    Method ogbn-products (Test Acc %) ogbl-ddi (Hits@20 %) ogbg-molhiv (Test ROC-AUC %)
    Baseline 79.45 ±\pm 0.59 53.90 ±\pm 4.74 75.58 ±\pm 1.40
    +PGD (M=8M=8) 80.96 ±\pm 0.41 62.02 ±\pm 6.56 76.14 ±\pm 1.62
    +"Free" (M=8M=8) 79.42 ±\pm 0.84 58.61 ±\pm 6.00 74.93 ±\pm 1.29
    +FLAG (M=3M=3) 81.76 ±\pm 0.45 63.31 ±\pm 6.06 76.54 ±\pm 1.14
    +FLAG (fast) 80.64 ±\pm 0.74 - -

    FLAG with M=3M=3 steps outperforms both PGD and Free training with M=8M=8 steps across all three tasks. Under equalized runtime conditions where epoch counts are reduced (FLAG fast on 100-epoch GAT takes 91 minutes on an Nvidia RTX 2080Ti vs. 88 minutes for vanilla GAT), FLAG retains a +1.19%+1.19\% test accuracy advantage over the baseline.

  9. Knowl 9 — Synergy of FLAG with Topological Augmentations, Dropout, and Batch Normalization

    empirical result

    Ablation experiments show that FLAG operates orthogonally and synergistically with standard structural regularizers and architectural components:

    1. Graph Topological Regularizers: On ogbn-products, full-batch GraphSAGE achieves 78.50±0.14%78.50 \pm 0.14\% test accuracy; neighbor sampling (NS) alone achieves 78.70±0.36%78.70 \pm 0.36\%; combining NS with FLAG reaches 79.36±0.57%79.36 \pm 0.57\%. On ogbg-ppa, vanilla GIN obtains 68.92±1.00%68.92 \pm 1.00\%; adding a virtual node increases accuracy to 70.37±1.07%70.37 \pm 1.07\%; adding both virtual nodes and FLAG raises it to 72.45±1.14%72.45 \pm 1.14\%.

    2. Dropout Regularization: On ogbn-products, GAT without dropout scores 75.67±0.27%75.67 \pm 0.27\%; GAT with dropout scores 79.45±0.59%79.45 \pm 0.59\%; adding FLAG on top of dropout further lifts performance to 81.76±0.45%81.76 \pm 0.45\%.

    3. Batch Normalization (BN) and Dual BN: On ogbn-arxiv, GCN scores 71.09±0.22%71.09 \pm 0.22\% without BN, 71.74±0.29%71.74 \pm 0.29\% with standard BN, 72.04±0.20%72.04 \pm 0.20\% with standard BN + FLAG, and 72.11±0.23%72.11 \pm 0.23\% with Dual BN (separate BN statistics for clean and adversarial features) + FLAG. GraphSAGE shows an identical pattern (69.58%69.58\% w/o BN →\rightarrow 71.49%71.49\% w/ BN →\rightarrow 72.19%72.19\% w/ BN + FLAG →\rightarrow 72.21%72.21\% w/ Dual BN + FLAG).

    4. Graph Mini-batching: When evaluated with GraphSAGE on ogbn-products, FLAG improves Neighbor Sampling (78.70%→79.36%78.70\% \rightarrow 79.36\%) and GraphSAINT (79.08%→79.60%79.08\% \rightarrow 79.60\%), while Cluster-GCN shows a slight regression (78.97%→78.60%78.97\% \rightarrow 78.60\%).

  10. Knowl 10 — Empirical Evidence on Discrete Feature Distributions as the Determinant of Adversarial Augmentation Efficacy

    empirical result

    While adversarial training frequently harms clean accuracy in continuous-input domains (e.g., standard natural image classification), it improves generalization on graph and NLP benchmarks. It is hypothesized that this discrepancy is driven primarily by the discrete versus continuous nature of the input feature distribution rather than specific model architectures:

    1. Model Agnosticism on Graph Features: Applying FLAG to a Multi-Layer Perceptron (MLP)—an architecture where adversarial training degrades clean image accuracy—on graph datasets improves test accuracy from 61.06±0.08%61.06 \pm 0.08\% to 62.41±0.16%62.41 \pm 0.16\% on ogbn-products and from 55.50±0.23%55.50 \pm 0.23\% to 56.02±0.19%56.02 \pm 0.19\% on ogbn-arxiv.

    2. Continuous Noise Injection on Cora: When FGSM adversarial feature augmentation is applied to a GCN trained on Cora with original discrete bag-of-words node features (Gaussian noise standard deviation σ=0\sigma = 0), the relative test accuracy improves by ≈1.5%\approx 1.5\%. When artificial Gaussian noise of increasing standard deviation σ∈[0,0.20]\sigma \in [0, 0.20] is added to make the node feature distribution continuous with broad support, FGSM augmentation progressively degrades clean test accuracy (dropping to a relative decrease of −2.0%-2.0\% at σ=0.20\sigma = 0.20).

Coverage note — Preliminary experiments evaluating GCN with FLAG on the MNIST superpixel vision dataset (87.83% baseline vs 89.10% with FLAG) and explicit line-plot coordinate traces for hyperparameter sweeps (step size and layer depth) were omitted in favor of the primary benchmark results and self-contained analytical findings.

References

  1. 1.Yogesh Balaji, Tom Goldstein, and Judy Hoffman. Instance adaptive adversarial training: Improved accuracy tradeoffs in neural nets. arXiv preprint arXiv:1910.08051, 2019. 1
  2. 2.Jie Chen, Tengfei Ma, and Cao Xiao. Fastgcn: fast learning with graph convolutional networks via importance sampling. arXiv preprint arXiv:1801.10247, 2018. 2
  3. 3.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020. 1, 3
  4. 4.Wei-Lin Chiang, Xuanqing Liu, Si Si, Yang Li, Samy Bengio, and Cho-Jui Hsieh. Cluster-gcn: An efficient algorithm for training deep and large graph convolutional networks. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 257–266, 2019. 7
  5. 5.Zhijie Deng, Yinpeng Dong, and Jun Zhu. Batch virtual adversarial training for graph convolutional networks. arXiv preprint arXiv:1902.09192, 2019. 2
  6. 6.Vijay Prakash Dwivedi, Chaitanya K Joshi, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. Benchmarking graph neural networks. arXiv preprint arXiv:2003.00982, 2020. 4
  7. 7.Federico Errica, Marco Podda, Davide Bacciu, and Alessio Micheli. A fair comparison of graph neural networks for graph classification. arXiv preprint arXiv:1912.09893, 2019. 4
  8. 8.Fuli Feng, Xiangnan He, Jie Tang, and Tat-Seng Chua. Graph adversarial training: Dynamically regularizing based on graph structure. IEEE Transactions on Knowledge and Data Engineering, 2019. 2
  9. 9.Matthias Fey, Jan Eric Lenssen, Frank Weichert, and Heinrich Müller. Splinecnn: Fast geometric deep learning with continuous b-spline kernels. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 869–877, 2018. 7
  10. 10.Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. Large-scale adversarial training for vision-and-language representation learning. arXiv preprint arXiv:2006.06195, 2020. 1, 3
  11. 11.Victor Garcia and Joan Bruna. Few-shot learning with graph neural networks. arXiv preprint arXiv:1711.04043, 2017. 1
  12. 12.Lise Getoor. Link-based classification. In Advanced methods for knowledge discovery from complex data, pages 189–207. Springer, 2005. 7
  13. 13.Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. Neural message passing for quantum chemistry. arXiv preprint arXiv:1704.01212, 2017. 1, 5
  14. 14.Jonathan Godwin, Michael Schaarschmidt, Alexander L Gaunt, Alvaro Sanchez-Gonzalez, Yulia Rubanova, Petar Veličković, James Kirkpatrick, and Peter Battaglia. Simple GNN regularisation for 3d molecular property prediction and beyond. In International Conference on Learning Representations, 2022. 1
  15. 15.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014. 1, 3, 5
  16. 16.Will Hamilton, Zhitao Ying, and Jure Leskovec. Inductive representation learning on large graphs. In Advances in neural information processing systems, pages 1024–1034, 2017. 1, 2, 5, 7
  17. 17.Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. Open graph benchmark: Datasets for machine learning on graphs. arXiv preprint arXiv:2005.00687, 2020. 1, 2, 3, 4, 5, 7
  18. 18.Weihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay Pande, and Jure Leskovec. Strategies for pre-training graph neural networks. arXiv preprint arXiv:1905.12265, 2019. 2
  19. 19.Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao. Smart: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization. arXiv preprint arXiv:1911.03437, 2019. 1
  20. 20.Hongwei Jin and Xinhua Zhang. Latent adversarial training of graph convolution networks. In ICML Workshop on Learning and Reasoning with Graph-Structured Representations, 2019. 2
  21. 21.Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016. 1
  22. 22.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012. 1
  23. 23.Chang Li and Dan Goldwasser. Encoding social information with graph convolutional networks forpolitical perspective detection in news media. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2594–2604, 2019. 1
  24. 24.Guohao Li, Chenxin Xiong, Ali Thabet, and Bernard Ghanem. Deepergcn: All you need to train deeper gcns. arXiv preprint arXiv:2006.07739, 2020. 7
  25. 25.Junying Li, Deng Cai, and Xiaofei He. Learning graph-level representation for drug discovery. arXiv preprint arXiv:1709.03741, 2017. 5
  26. 26.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083, 2017. 1, 3, 5
  27. 27.Takeru Miyato, Andrew M Dai, and Ian Goodfellow. Adversarial training methods for semi-supervised text classification. arXiv preprint arXiv:1605.07725, 2016. 1
  28. 28.Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE transactions on pattern analysis and machine intelligence, 41(8):1979–1993, 2018. 3
  29. 29.Jiezhong Qiu, Jian Tang, Hao Ma, Yuxiao Dong, Kuansan Wang, and Jie Tang. Deepinf: Social influence prediction with deep learning. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 2110–2119, 2018. 1
  30. 30.Yu Rong, Wenbing Huang, Tingyang Xu, and Junzhou Huang. Dropedge: Towards deep graph convolutional networks on node classification. In International Conference on Learning Representations, 2019. 1, 2, 5
  31. 31.Ali Shafahi, Mahyar Najibi, Mohammad Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein. Adversarial training for free! In Advances in Neural Information Processing Systems, pages 3358–3369, 2019. 2, 3
  32. 32.Oleksandr Shchur, Maximilian Mumme, Aleksandar Bojchevski, and Stephan Günnemann. Pitfalls of graph neural network evaluation. arXiv preprint arXiv:1811.05868, 2018. 4
  33. 33.Yantao Shen, Hongsheng Li, Shuai Yi, Dapeng Chen, and Xiaogang Wang. Person re-identification with deep similarity-guided graph neural network. In Proceedings of the European conference on computer vision (ECCV), pages 486–504, 2018. 1
  34. 34.Manli Shu, Zuxuan Wu, Micah Goldblum, and Tom Goldstein. Prepare for the worst: Generalizing across domain shifts with adversarial batch normalization. arXiv e-prints, pages arXiv–2009, 2020. 1
  35. 35.Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. Robustness may be at odds with accuracy. arXiv preprint arXiv:1805.12152, 2018. 1, 3, 7
  36. 36.Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John C Duchi, Vittorio Murino, and Silvio Savarese. Generalizing to unseen domains via adversarial data augmentation. In Advances in neural information processing systems, pages 5334–5344, 2018. 1, 3, 4
  37. 37.Yiwei Wang, Wei Wang, Yuxuan Liang, Yujun Cai, Juncheng Liu, and Bryan Hooi. Nodeaug: Semi-supervised node classification with data augmentation. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 207–217, 2020. 1
  38. 38.Jason Wei and Kai Zou. Eda: Easy data augmentation techniques for boosting performance on text classification tasks. arXiv preprint arXiv:1901.11196, 2019. 1
  39. 39.Eric Wong, Leslie Rice, and J Zico Kolter. Fast is better than free: Revisiting adversarial training. arXiv preprint arXiv:2001.03994, 2020. 7
  40. 40.Cihang Xie, Mingxing Tan, Boqing Gong, Jiang Wang, Alan L Yuille, and Quoc V Le. Adversarial examples improve image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 819–828, 2020. 1, 6
  41. 41.Rex Ying, Ruining He, Kaifeng Chen, Pong Eksombatchai, William L Hamilton, and Jure Leskovec. Graph convolutional neural networks for web-scale recommender systems. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 974–983, 2018. 1
  42. 42.Yuning You, Tianlong Chen, Yongduo Sui, Ting Chen, Zhangyang Wang, and Yang Shen. Graph contrastive learning with augmentations. Advances in Neural Information Processing Systems, 33, 2020. 1
  43. 43.Hanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan, and Viktor Prasanna. Graphsaint: Graph sampling based inductive learning method. arXiv preprint arXiv:1907.04931, 2019. 7
  44. 44.Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Thomas Goldstein, and Jingjing Liu. Freelb: Enhanced adversarial training for language understanding. arXiv preprint arXiv:1909.11764, 2019. 1
  45. 45.Daniel Zügner, Amir Akbarnejad, and Stephan Günnemann. Adversarial attacks on neural networks for graph data. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 2847–2856, 2018. 5

Citation

MLA
Kong, K., et al. “Robust Optimization as Data Augmentation for Large-scale Graphs”. arXiv, 2020, http://arxiv.org/abs/2010.09891v3.
APA
Kong, K., Li, G., Ding, M., Wu, Z., Zhu, C., Ghanem, B., Taylor, G., & Goldstein, T. (2020). Robust Optimization as Data Augmentation for Large-scale Graphs. arXiv. http://arxiv.org/abs/2010.09891v3
Chicago
Kong, K., G. Li, M. Ding, et al. 2020. “Robust Optimization as Data Augmentation for Large-scale Graphs”. arXiv. http://arxiv.org/abs/2010.09891v3.
Harvard
Kong, K. et al. (2020) “Robust Optimization as Data Augmentation for Large-scale Graphs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2010.09891v3.
Vancouver
1. Kong K, Li G, Ding M, Wu Z, Zhu C, Ghanem B, Taylor G, Goldstein T (2020) Robust Optimization as Data Augmentation for Large-scale Graphs. arXiv

BibTeX

@article{kong2020robust,
  title = {Robust Optimization as Data Augmentation for Large-scale Graphs},
  author = {Kong, Kezhi and Li, Guohao and Ding, Mucong and Wu, Zuxuan and Zhu, Chen and Ghanem, Bernard and Taylor, Gavin and Goldstein, Tom},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2010.09891v3},
  eprint = {2010.09891}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE