DropEdge: Towards Deep Graph Convolutional Networks on Node Classification

Yu RongWen-bing HuangTingyang XuJunzhou Huang

article2019ICLR1,698 citations

Proposes DropEdge, a flexible data augmentation technique that randomly removes graph edges during training to prevent over-smoothing and over-fitting in deep Graph Convolutional Networks.

Listen

Graph neural networks are vital tools for analyzing interconnected data, supporting real-world applications in social networks, recommendation systems, and citation mapping. Despite their value, standard models face severe performance bottlenecks as they scale. Deepening these architectures leads to two primary points of failure: over-fitting on small datasets and over-smoothing, a phenomenon where repetitive message passing between connected nodes causes internal representations to blend together until they lose meaningful information. Because of these constraints, practitioners have historically been restricted to very shallow networks, limiting the complexity of patterns they can extract from network data.

The main objective of the article is to demonstrate and evaluate DropEdge, a flexible training technique designed to alleviate both over-fitting and over-smoothing in deep graph convolutional networks for node classification tasks.

To evaluate this technique, the authors designed a training mechanism that randomly removes a specified fraction of edges from the input graph during each training iteration and re-normalizes the remaining connections. They evaluated this approach using extensive experiments across four standard benchmark datasets—three citation networks (Cora, Citeseer, and Pubmed) and one large social network (Reddit)—spanning both small transductive tasks and large inductive learning settings. The evaluation tested multiple depths (ranging from 2 to 64 layers) across five prominent baseline network architectures: standard Graph Convolutional Networks, Residual GCNs, Inception GCNs, Jumping Knowledge Networks, and GraphSAGE.

The findings establish that DropEdge consistently enhances node classification performance across all backbones and datasets, delivering the greatest benefits to deeper networks. On the Citeseer dataset, the technique produced a modest average absolute gain of 0.9% on 2-layer models but achieved a 13.5% average improvement on 64-layer models. DropEdge set new state-of-the-art benchmarks on all tested datasets, achieving a notable 97.02% accuracy on the Reddit network. Mathematical derivations and distance analyses confirmed that DropEdge significantly delays over-smoothing and preserves node representation differences after training. Furthermore, making the network matrix sparser reduced memory consumption, allowing deep 32-layer models to train successfully without encountering out-of-memory errors that previously caused complete execution failure.

These results demonstrate that DropEdge functions simultaneously as an unbiased data augmentor that curbs over-fitting and as a communication reducer that prevents over-smoothing. For technical and operational leaders, these findings show that model depth is no longer a major bottleneck in graph analytics. Teams can deploy deeper architectures with improved accuracy while reducing computing overhead and hardware memory costs. The technique is also fully complementary with standard feature dropout, providing greater combined stability than using feature-level regularization alone.

Technical leaders should integrate DropEdge into existing graph representation workflows as a standard regularizer. The global DropEdge variant is recommended over layer-wise edge dropping, as it delivers comparable validation accuracy with significantly lower computational complexity. Future efforts should test DropEdge on broader graph structures and edge-level tasks, establish automated methods for tuning the edge drop rate, and pilot the framework in production-scale enterprise pipelines.

  • Paper: Simple and Deep Graph Convolutional Networks, Ming Chen et al. (2020). It tackles the over-smoothing bottleneck highlighted in DropEdge by proposing GCNII with initial residual connections and identity mapping to scale networks up to 64 layers.
  • Paper: Self-supervised Graph Learning for Recommendation, Jiancan Wu et al. (2020). It builds directly on edge dropout techniques like DropEdge as data augmentations to create multi-view self-supervised contrastive learning for graph-based recommendations.
  • Paper: Graph Contrastive Learning with Augmentations, Yuning You et al. (2020). It incorporates edge perturbation and node dropping into a generalized graph contrastive learning framework (GraphCL) for self-supervised pre-training.
  • Paper: Open Graph Benchmark: Datasets for Machine Learning on Graphs, Weihua Hu et al. (2020). It provides large-scale, standardized benchmarks for evaluating graph neural network regularizations and scalability techniques like DropEdge across diverse domains.
Cover for DropEdge: Towards Deep Graph Convolutional Networks on Node Classification

Abstract

\emph{Over-fitting} and \emph{over-smoothing} are two main obstacles of developing deep Graph Convolutional Networks (GCNs) for node classification. In particular, over-fitting weakens the generalization ability on small dataset, while over-smoothing impedes model training by isolating output representations from the input features with the increase in network depth. This paper proposes DropEdge, a novel and flexible technique to alleviate both issues. At its core, DropEdge randomly removes a certain number of edges from the input graph at each training epoch, acting like a data augmenter and also a message passing reducer. Furthermore, we theoretically demonstrate that DropEdge either reduces the convergence speed of over-smoothing or relieves the information loss caused by it. More importantly, our DropEdge is a general skill that can be equipped with many other backbone models (e.g. GCN, ResGCN, GraphSAGE, and JKNet) for enhanced performance. Extensive experiments on several benchmarks verify that DropEdge consistently improves the performance on a variety of both shallow and deep GCNs. The effect of DropEdge on preventing over-smoothing is empirically visualized and validated as well. Codes are released on~\url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Notations and Preliminaries
  • 4 Our Method: DropEdge
  • 4.1 Methodology
  • 4.2 Towards preventing over-smoothing
  • 4.3 discussions
  • 5 Experiments
  • 5.1 Can DropEdge generally improve the performance of deep GCNs?
  • 5.2 How does DropEdge help?
  • 5.2.1 On preventing over-smoothing
  • 5.2.2 On Compatibility with Dropout
  • 5.2.3 On layer-wise DropEdge
  • 6 Conclusion
  • 7 Acknowledgements
  • References
  • A Appendix: Proof of Theorem 1
  • B Appendix: More Details in Experiments
  • B.1 Datasets Statistics
  • B.2 Models and Backbones
  • B.3 The Validation Loss on Different Backbones w and w/o DropEdge.
  • B.4 The Ablation Study on Citeseer

Knowls

  1. Knowl 1 — DropEdge Mechanism for Graph Convolutional Networks

    model/method

    DropEdge is a regularisation and message-passing modification technique for training Graph Convolutional Networks (GCNs). At each training epoch, a fraction p∈[0,1)p \in [0, 1) of the edges in the input graph G=(V,E)G = (V, \mathcal{E}) is randomly removed.

    Given the original adjacency matrix A∈RN×NA \in \mathbb{R}^{N \times N} with total edge count ∣E∣|\mathcal{E}|, DropEdge sets ∣E∣⋅p|\mathcal{E}| \cdot p non-zero elements of AA to zero at random. The resulting modified adjacency matrix AdropA_{\text{drop}} is: Adrop=A−A′,A_{\text{drop}} = A - A', where A′A' is a sparse matrix containing the randomly sampled ∣E∣⋅p|\mathcal{E}| \cdot p directed edges. A normalized adjacency matrix A^drop\hat{A}_{\text{drop}} is subsequently computed from AdropA_{\text{drop}} (for instance, via the re-normalization trick A^drop=D~drop−1/2(Adrop+I)D~drop−1/2\hat{A}_{\text{drop}} = \tilde{D}_{\text{drop}}^{-1/2}(A_{\text{drop}} + I)\tilde{D}_{\text{drop}}^{-1/2}, where D~drop\tilde{D}_{\text{drop}} is the diagonal degree matrix of Adrop+IA_{\text{drop}} + I).

    Forward propagation in each graph convolutional layer during training is given by: H(l+1)=σ(A^dropH(l)W(l)),H^{(l+1)} = \sigma\left(\hat{A}_{\text{drop}} H^{(l)} W^{(l)}\right), where H(l)∈RN×ClH^{(l)} \in \mathbb{R}^{N \times C_l} is the feature matrix at layer ll, W(l)∈RCl−1×ClW^{(l)} \in \mathbb{R}^{C_{l-1} \times C_l} is the trainable weight parameter matrix, and σ(⋅)\sigma(\cdot) is an element-wise activation function such as ReLU. During validation and testing, DropEdge is turned off and the original, complete adjacency matrix A^\hat{A} is utilized.

    DropEdge acts simultaneously as an unbiased data augmentation technique (maintaining the expectation of neighbor aggregation up to normalization) to counter over-fitting and as a message-passing reducer that sparsifies connectivity to prevent over-smoothing in deep architectures.

  2. Knowl 2 — Theoretical Effect of DropEdge on Over-Smoothing and Information Loss

    theoretical result

    Let GG be an input graph and G′G' be the perturbed graph obtained after removing a sufficient number of edges via DropEdge. Assume that for a given distance threshold ϵ>0\epsilon > 0, deep graph convolution on GG and G′G' incurs ϵ\epsilon-smoothing towards the subspaces MM and M′M', respectively, where M,M′⊂RN×CM, M' \subset \mathbb{R}^{N \times C}.

    Under the condition that the filter singular value supremum across all layers satisfies s=sup⁡l≥1sl≤1s = \sup_{l \ge 1} s_l \le 1, dropping sufficient edges from GG guarantees that at least one of the following two properties holds:

    1. Deceleration of Over-Smoothing: The relaxed ϵ\epsilon-smoothing layer depth strictly does not decrease: l^(M,ϵ)≤l^(M′,ϵ),\hat{l}(M, \epsilon) \le \hat{l}(M', \epsilon), meaning a greater network depth is required for node representations to converge within an ϵ\epsilon-distance of the invariant subspace.

    2. Reduction of Information Loss: The dimensional information loss between the full node representation space RN\mathbb{R}^N and the invariant convergence subspace decreases: N−dim⁡(M)>N−dim⁡(M′).N - \dim(M) > N - \dim(M').

    This occurs because removing edges increases the effective electrical resistance RstR_{st} between nodes in connected components, which increases the second largest eigenvalue λ\lambda of the normalized adjacency matrix, and eventually disconnects components into separate subgraphs, directly increasing the dimensionality of the invariant subspace dim⁡(M)\dim(M) by 1 per newly formed disconnected component.

  3. Knowl 3 — Definitions of Subspace Convergence and the Relaxed epsilon-Smoothing Layer

    definition

    Let NN denote the number of graph nodes and CC the node feature dimensionality.

    1. Subspace: An MM-dimensional subspace M⊂RN×CM \subset \mathbb{R}^{N \times C} is defined as: M:={EC∣C∈RM×C},\mathcal{M} := \{EC \mid C \in \mathbb{R}^{M \times C}\}, where E∈RN×ME \in \mathbb{R}^{N \times M} has orthogonal columns (ETE=IME^T E = I_M) with M≤NM \le N.

    2. ϵ\epsilon-Smoothing: A Graph Convolutional Network suffers from ϵ\epsilon-smoothing if, for a threshold ϵ>0\epsilon > 0, all hidden representations H(l)H^{(l)} beyond a certain layer LL lie within an ϵ\epsilon-distance of an invariant subspace M\mathcal{M} that is independent of the input node features XX: dM(H(l))<ϵ,∀l≥L,d_{\mathcal{M}}(H^{(l)}) < \epsilon, \quad \forall l \ge L, where dM(X):=inf⁡Y∈M∥X−Y∥Fd_{\mathcal{M}}(X) := \inf_{Y \in \mathcal{M}} \|X - Y\|_F.

    3. ϵ\epsilon-Smoothing Layer: The exact ϵ\epsilon-smoothing layer is the minimum layer index satisfying ϵ\epsilon-smoothing: l∗(M,ϵ):=min⁡l{dM(H(l))<ϵ}.l^*(M, \epsilon) := \min_{l} \{ d_{\mathcal{M}}(H^{(l)}) < \epsilon \}.

    4. Relaxed ϵ\epsilon-Smoothing Layer: The relaxed ϵ\epsilon-smoothing layer l^(M,ϵ)\hat{l}(M, \epsilon) is an upper bound on l∗(M,ϵ)l^*(M, \epsilon) defined as: l^(M,ϵ):=⌈log⁡(ϵ/dM(X))log⁡(sλ)⌉,\hat{l}(M, \epsilon) := \left\lceil \frac{\log\left(\epsilon / d_{\mathcal{M}}(X)\right)}{\log(s \lambda)} \right\rceil, where ⌈⋅⌉\lceil \cdot \rceil is the ceiling operator, s:=sup⁡l∈N+sl≤1s := \sup_{l \in \mathbb{N}^+} s_l \le 1 is the supremum of the filter matrices' maximum singular values across all layers, and λ\lambda is the second-largest eigenvalue modulus of the normalized adjacency matrix A^\hat{A}.

  4. Knowl 4 — Layer-Wise DropEdge

    model/method

    Layer-Wise DropEdge (LW DropEdge) is an extension of DropEdge where edge removal is performed independently at each individual graph convolutional layer rather than sharing a single perturbed adjacency matrix across the entire network.

    For each layer l∈{1,…,L}l \in \{1, \dots, L\}, an independent edge mask is sampled with edge dropping rate pp, producing a layer-specific dropped adjacency matrix Adrop(l)=A−A′(l)A_{\text{drop}}^{(l)} = A - A'^{(l)} and its normalized version A^drop(l)\hat{A}_{\text{drop}}^{(l)}. The layer propagation step is computed as: H(l+1)=σ(A^drop(l)H(l)W(l)).H^{(l+1)} = \sigma\left(\hat{A}_{\text{drop}}^{(l)} H^{(l)} W^{(l)}\right).

    Layer-Wise DropEdge generates distinct random graph deformations across layers within the same training step, leading to lower training loss compared to standard one-shot DropEdge. However, it incurs higher per-epoch computational and sampling complexity.

  5. Knowl 5 — Adjacency Normalization and Propagation Variants Compatible with DropEdge

    model/method

    DropEdge supports multiple graph propagation and adjacency matrix normalizations applied to the edge-dropped adjacency matrix AdropA_{\text{drop}} (where DD is the diagonal degree matrix of AdropA_{\text{drop}} and II is the identity matrix):

    • AugNormAdj (Augmented Normalized Adjacency): A^=(D+I)−1/2(Adrop+I)(D+I)−1/2\hat{A} = (D + I)^{-1/2}(A_{\text{drop}} + I)(D + I)^{-1/2}
    • BingGeNormAdj (Augmented Normalized Adjacency with Self-Loop): A^=I+(D+I)−1/2(Adrop+I)(D+I)−1/2\hat{A} = I + (D + I)^{-1/2}(A_{\text{drop}} + I)(D + I)^{-1/2}
    • FirstOrderGCN (First-Order GCN): A^=I+D−1/2AdropD−1/2\hat{A} = I + D^{-1/2}A_{\text{drop}}D^{-1/2}
    • AugRWalk (Augmented Random Walk): A^=(D+I)−1(Adrop+I)\hat{A} = (D + I)^{-1}(A_{\text{drop}} + I)

    Additionally, DropEdge can be combined with a self-feature modeling layer: H(l+1)=σ(A^dropH(l)W(l)+H(l)Wself(l)),H^{(l+1)} = \sigma\left(\hat{A}_{\text{drop}} H^{(l)} W^{(l)} + H^{(l)} W_{\text{self}}^{(l)}\right), where Wself(l)∈RCl−1×ClW_{\text{self}}^{(l)} \in \mathbb{R}^{C_{l-1} \times C_l} is a dedicated weight matrix that directly transforms a node's own hidden features from the preceding layer.

  6. Knowl 6 — Empirical Classification Accuracy Across Depths and Backbones With DropEdge

    data/table

    DropEdge consistently improves node classification accuracy across varying depths (2, 8, and 32 layers) on four standard benchmarks: transductive citation networks (Cora, Citeseer, Pubmed) and the inductive Reddit social network. It prevents catastrophic performance drops in deep architectures (up to 32/64 layers) and reduces memory usage, enabling 32-layer IncepGCN on Pubmed to run without Out-Of-Memory (OOM) errors.

    Dataset Backbone 2 layers 8 layers 32 layers
    Original DropEdge Original DropEdge Original DropEdge
    Cora GCN 86.10% 86.50% 78.70% 85.80% 71.60% 74.60%
    ResGCN - - 85.40% 86.90% 85.10% 86.80%
    JKNet - - 86.70% 87.80% 87.10% 87.60%
    IncepGCN - - 86.70% 88.20% 87.40% 87.70%
    GraphSAGE 87.80% 88.10% 84.30% 87.10% 31.90% 32.20%
    Citeseer GCN 75.90% 78.70% 74.60% 77.20% 59.20% 61.40%
    ResGCN - - 77.80% 78.80% 74.40% 77.90%
    JKNet - - 79.20% 80.20% 71.70% 80.00%
    IncepGCN - - 79.60% 80.50% 72.60% 80.30%
    GraphSAGE 78.40% 80.00% 74.10% 77.10% 37.00% 53.60%
    Pubmed GCN 90.20% 91.20% 90.10% 90.90% 84.60% 86.20%
    ResGCN - - 89.60% 90.50% 90.20% 91.10%
    JKNet - - 90.60% 91.20% 89.20% 91.30%
    IncepGCN - - 90.20% 91.50% OOM 90.50%
    GraphSAGE 90.10% 90.70% 90.20% 91.70% 41.30% 47.90%
    Reddit GCN 96.11% 96.13% 96.17% 96.48% 45.55% 50.51%
    ResGCN - - 96.37% 96.46% 93.93% 94.27%
    JKNet - - 96.82% 97.02% OOM OOM
    IncepGCN - - 96.43% 96.87% OOM OOM
    GraphSAGE 96.22% 96.28% 96.38% 96.42% 96.43% 96.47%
  7. Knowl 7 — State-of-the-Art Comparison of DropEdge-Equipped Models

    data/table

    When evaluated across transductive and inductive node classification tasks, DropEdge-equipped backbones outperform both standard baseline architectures and node-sampling methods (FastGCN, ASGCN). The best performing configuration for each architecture typically occurs at network depths strictly greater than 2.

    Method Transductive Inductive
    Cora Citeseer Pubmed Reddit
    GCN 86.64% 79.34% 90.22% 95.68%
    FastGCN 85.00% 77.60% 88.00% 93.70%
    ASGCN 87.44% 79.66% 90.60% 96.27%
    GraphSAGE 82.20% 71.40% 87.10% 94.32%
    GCN + DropEdge 87.60% (4) 79.20% (4) 91.30% (4) 96.71% (4)
    ResGCN + DropEdge 87.00% (4) 79.40% (16) 91.10% (32) 96.48% (16)
    JKNet + DropEdge 88.00% (16) 80.20% (8) 91.60% (64) 97.02% (8)
    IncepGCN + DropEdge 88.20% (8) 80.50% (8) 91.60% (4) 96.87% (8)
    GraphSAGE + DropEdge 88.10% (4) 80.00% (2) 91.70% (8) 96.54% (4)

    Note: Numbers in parentheses indicate the optimal layer depth.

  8. Knowl 8 — Empirical Validation of Over-Smoothing Alleviation via Inter-Layer Distance

    empirical result

    Over-smoothing can be empirically diagnosed by measuring the Euclidean distance between consecutive layer outputs, ∥H(l)−H(l−1)∥2\|H^{(l)} - H^{(l-1)}\|_2. A distance converging to zero indicates representation collapse into an invariant stationary subspace.

    In an 8-layer plain GCN on the Cora dataset:

    1. Before training: As layer depth increases from layer 2 to layer 6, the inter-layer distance decays monotonically. Applying DropEdge with edge dropping probability p=0.8p = 0.8 maintains substantially higher inter-layer distances and slows down the decay compared to standard training (p=0p = 0).
    2. After training (150 epochs): For GCN without DropEdge (p=0p = 0), the inter-layer distance between layers 5 and 6 falls to zero (10−1710^{-17} scale / numerical zero), confirming total feature collapse to an uninformative stationary point. For GCN with DropEdge (p=0.8p = 0.8), the inter-layer distance remains strictly positive (>10−2> 10^{-2}) across all layers, demonstrating that meaningful node representations are preserved and gradient vanishing is prevented.
  9. Knowl 9 — Compatibility and Orthogonality Between DropEdge and Feature Dropout

    empirical result

    Standard Dropout randomly zeroes node feature dimensions during training to reduce feature co-adaptation and mitigate over-fitting, but it leaves the adjacency matrix unaltered and cannot impede over-smoothing. DropEdge randomly removes edges, modifying message-passing pathways and graph connectivity.

    On a 4-layer GCN trained on the Cora and Citeseer datasets, combining Dropout and DropEdge simultaneously yields lower validation loss and better convergence than applying either Dropout alone or DropEdge alone. This confirms that edge-level perturbation (DropEdge) and feature-level perturbation (Dropout) provide distinct, complementary regularizations.

Coverage note — None was omitted. All key methodological contributions, theoretical definitions, formal theorems, experimental tables, ablation studies, and variant architectures from the paper are included.

References

  1. 1.Smriti Bhagat, Graham Cormode, and S Muthukrishnan. Node classification in social networks. In Social network data analytics, pp. 115–148. Springer, 2011.
  2. 2.Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann LeCun. Spectral networks and locally connected networks on graphs. In Proceedings of International Conference on Learning Representations, 2013.
  3. 3.Jie Chen, Tengfei Ma, and Cao Xiao. Fastgcn: Fast learning with graph convolutional networks via importance sampling. In Proceedings of the 6th International Conference on Learning Representations, 2018.
  4. 4.Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst. Convolutional neural networks on graphs with fast localized spectral filtering. In Advances in Neural Information Processing Systems, pp. 3844–3852, 2016.
  5. 5.David Eppstein, Zvi Galil, Giuseppe F Italiano, and Amnon Nissenzweig. Sparsification—a technique for speeding up dynamic graph algorithms. Journal of the ACM (JACM), 44(5):669–696, 1997.
  6. 6.Alex Fout, Jonathon Byrd, Basir Shariat, and Asa Ben-Hur. Protein interface prediction using graph convolutional networks. In Advances in Neural Information Processing Systems, pp. 6530–6539, 2017.
  7. 7.Linton C Freeman. Visualizing social networks. Journal of social structure, 1(1):4, 2000.
  8. 8.Hongyang Gao, Zhengyang Wang, and Shuiwang Ji. Large-scale learnable graph convolutional networks. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 1416–1424. ACM, 2018.
  9. 9.Will Hamilton, Zhitao Ying, and Jure Leskovec. Inductive representation learning on large graphs. In Advances in Neural Information Processing Systems, pp. 1025–1035, 2017.
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  11. 11.Mikael Henaff, Joan Bruna, and Yann LeCun. Deep convolutional networks on graph-structured data. arXiv preprint arXiv:1506.05163, 2015.
  12. 12.Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov. Improving neural networks by preventing co-adaptation of feature detectors. arXiv preprint arXiv:1207.0580, 2012.
  13. 13.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700–4708, 2017.
  14. 14.Wenbing Huang, Tong Zhang, Yu Rong, and Junzhou Huang. Adaptive sampling towards fast graph representation learning. In Advances in Neural Information Processing Systems, pp. 4558–4567, 2018.
  15. 15.Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In Proceedings of the International Conference on Learning Representations, 2017.
  16. 16.Johannes Klicpera, Aleksandar Bojchevski, and Stephan Günnemann. Predict then propagate: Graph neural networks meet personalized pagerank. In Proceedings of the 7th International Conference on Learning Representations, 2019.
  17. 17.Ron Levie, Federico Monti, Xavier Bresson, and Michael M Bronstein. Cayleynets: Graph convolutional neural networks with complex rational spectral filters. IEEE Transactions on Signal Processing, 67(1):97–109, 2017.
  18. 18.Guohao Li, Matthias Müller, Ali Thabet, and Bernard Ghanem. Deepgcns: Can gcns go as deep as cnns? In International Conference on Computer Vision, 2019.
  19. 19.Qimai Li, Zhichao Han, and Xiao-Ming Wu. Deeper insights into graph convolutional networks for semi-supervised learning. In Thirty-Second AAAI Conference on Artificial Intelligence, 2018a.
  20. 20.Ruoyu Li, Sheng Wang, Feiyun Zhu, and Junzhou Huang. Adaptive graph convolutional neural networks. In Thirty-Second AAAI Conference on Artificial Intelligence, 2018b.
  21. 21.David Liben-Nowell and Jon Kleinberg. The link-prediction problem for social networks. Journal of the American society for information science and technology, 58(7):1019–1031, 2007.
  22. 22.László Lovász et al. Random walks on graphs: A survey. Combinatorics, Paul erdos is eighty, 2(1): 1–46, 1993.
  23. 23.Federico Monti, Davide Boscaini, Jonathan Masci, Emanuele Rodola, Jan Svoboda, and Michael M Bronstein. Geometric deep learning on graphs and manifolds using mixture model cnns. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5115–5124, 2017.
  24. 24.Mathias Niepert, Mohamed Ahmed, and Konstantin Kutzkov. Learning convolutional neural networks for graphs. In International conference on machine learning, pp. 2014–2023, 2016.
  25. 25.Kenta Oono and Taiji Suzuki. On asymptotic behaviors of graph cnns from dynamical systems perspective. arXiv preprint arXiv:1905.10947, 2019.
  26. 26.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in PyTorch. In NIPS Autodiff Workshop, 2017.
  27. 27.Bryan Perozzi, Rami Al-Rfou, and Steven Skiena. Deepwalk: Online learning of social representations. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 701–710. ACM, 2014.
  28. 28.Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad. Collective classification in network data. AI magazine, 29(3):93, 2008.
  29. 29.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818–2826, 2016.
  30. 30.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. In ICLR, 2018.
  31. 31.Minjie Wang, Lingfan Yu, Da Zheng, Quan Gan, Yu Gai, Zihao Ye, Mufei Li, Jinjing Zhou, Qi Huang, Chao Ma, Ziyue Huang, Qipeng Guo, Hao Zhang, Haibin Lin, Junbo Zhao, Jinyang Li, Alexander J Smola, and Zheng Zhang. Deep graph library: Towards efficient and scalable deep learning on graphs. ICLR Workshop on Representation Learning on Graphs and Manifolds, 2019. URL https://arxiv.org/abs/1909.01315.
  32. 32.Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and Philip S Yu. A comprehensive survey on graph neural networks. arXiv preprint arXiv:1901.00596, 2019.
  33. 33.Keyulu Xu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie Jegelka. Representation learning on graphs with jumping knowledge networks. In Proceedings of the 35th International Conference on Machine Learning, 2018a.
  34. 34.Keyulu Xu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie Jegelka. Representation learning on graphs with jumping knowledge networks. arXiv preprint arXiv:1806.03536, 2018b.
  35. 35.Muhan Zhang, Zhicheng Cui, Marion Neumann, and Yixin Chen. An end-to-end deep learning architecture for graph classification. In Thirty-Second AAAI Conference on Artificial Intelligence, 2018.

Citation

MLA
Rong, Y., et al. “DropEdge: Towards Deep Graph Convolutional Networks on Node Classification”. arXiv, 2019, http://arxiv.org/abs/1907.10903v4.
APA
Rong, Y., Huang, W., Xu, T., & Huang, J. (2019). DropEdge: Towards Deep Graph Convolutional Networks on Node Classification. arXiv. http://arxiv.org/abs/1907.10903v4
Chicago
Rong, Y., W. Huang, T. Xu, and J. Huang. 2019. “DropEdge: Towards Deep Graph Convolutional Networks on Node Classification”. arXiv. http://arxiv.org/abs/1907.10903v4.
Harvard
Rong, Y. et al. (2019) “DropEdge: Towards Deep Graph Convolutional Networks on Node Classification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1907.10903v4.
Vancouver
1. Rong Y, Huang W, Xu T, Huang J (2019) DropEdge: Towards Deep Graph Convolutional Networks on Node Classification. arXiv

BibTeX

@article{rong2019dropedge,
  title = {DropEdge: Towards Deep Graph Convolutional Networks on Node Classification},
  author = {Rong, Yu and Huang, Wenbing and Xu, Tingyang and Huang, Junzhou},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1907.10903v4},
  eprint = {1907.10903}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors