DropEdge: Towards Deep Graph Convolutional Networks on Node Classification
Yu RongWen-bing HuangTingyang XuJunzhou Huang
Proposes DropEdge, a flexible data augmentation technique that randomly removes graph edges during training to prevent over-smoothing and over-fitting in deep Graph Convolutional Networks.
Graph neural networks are vital tools for analyzing interconnected data, supporting real-world applications in social networks, recommendation systems, and citation mapping. Despite their value, standard models face severe performance bottlenecks as they scale. Deepening these architectures leads to two primary points of failure: over-fitting on small datasets and over-smoothing, a phenomenon where repetitive message passing between connected nodes causes internal representations to blend together until they lose meaningful information. Because of these constraints, practitioners have historically been restricted to very shallow networks, limiting the complexity of patterns they can extract from network data.
The main objective of the article is to demonstrate and evaluate DropEdge, a flexible training technique designed to alleviate both over-fitting and over-smoothing in deep graph convolutional networks for node classification tasks.
To evaluate this technique, the authors designed a training mechanism that randomly removes a specified fraction of edges from the input graph during each training iteration and re-normalizes the remaining connections. They evaluated this approach using extensive experiments across four standard benchmark datasets—three citation networks (Cora, Citeseer, and Pubmed) and one large social network (Reddit)—spanning both small transductive tasks and large inductive learning settings. The evaluation tested multiple depths (ranging from 2 to 64 layers) across five prominent baseline network architectures: standard Graph Convolutional Networks, Residual GCNs, Inception GCNs, Jumping Knowledge Networks, and GraphSAGE.
The findings establish that DropEdge consistently enhances node classification performance across all backbones and datasets, delivering the greatest benefits to deeper networks. On the Citeseer dataset, the technique produced a modest average absolute gain of 0.9% on 2-layer models but achieved a 13.5% average improvement on 64-layer models. DropEdge set new state-of-the-art benchmarks on all tested datasets, achieving a notable 97.02% accuracy on the Reddit network. Mathematical derivations and distance analyses confirmed that DropEdge significantly delays over-smoothing and preserves node representation differences after training. Furthermore, making the network matrix sparser reduced memory consumption, allowing deep 32-layer models to train successfully without encountering out-of-memory errors that previously caused complete execution failure.
These results demonstrate that DropEdge functions simultaneously as an unbiased data augmentor that curbs over-fitting and as a communication reducer that prevents over-smoothing. For technical and operational leaders, these findings show that model depth is no longer a major bottleneck in graph analytics. Teams can deploy deeper architectures with improved accuracy while reducing computing overhead and hardware memory costs. The technique is also fully complementary with standard feature dropout, providing greater combined stability than using feature-level regularization alone.
Technical leaders should integrate DropEdge into existing graph representation workflows as a standard regularizer. The global DropEdge variant is recommended over layer-wise edge dropping, as it delivers comparable validation accuracy with significantly lower computational complexity. Future efforts should test DropEdge on broader graph structures and edge-level tasks, establish automated methods for tuning the edge drop rate, and pilot the framework in production-scale enterprise pipelines.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). It introduces the standard Graph Convolutional Network architecture whose vulnerability to over-smoothing and over-fitting under deep stacking is the core problem DropEdge solves.
- Paper: Deeper Insights into Graph Convolutional Networks for Semi-Supervised Learning, Qimai Li et al. (2018). It establishes the theoretical connection between graph convolutions and Laplacian smoothing, explaining why deeper GCNs suffer from over-smoothing.
- Paper: Inductive Representation Learning on Large Graphs, William L. Hamilton et al. (2017). It introduces the GraphSAGE architecture, which serves as one of the key backbone models evaluated and enhanced with DropEdge.
- Paper: Representation Learning on Graphs with Jumping Knowledge Networks, Keyulu Xu et al. (2018). It introduces Jumping Knowledge Networks, another principal multi-layer architecture integrated with and compared against DropEdge for deep graph representation learning.
- Paper: Dropout: a simple way to prevent neural networks from overfitting, Nitish Srivastava et al. (2014). It introduces the foundational concept of node and connection dropout for regularizing neural networks that DropEdge generalizes to graph edges.
- Paper: Predict then Propagate: Graph Neural Networks meet Personalized PageRank, Johannes Gasteiger et al. (2019). It presents an alternative approach to decoupling feature propagation and addressing over-smoothing in deep GCN architectures.
- Paper: Simple and Deep Graph Convolutional Networks, Ming Chen et al. (2020). It tackles the over-smoothing bottleneck highlighted in DropEdge by proposing GCNII with initial residual connections and identity mapping to scale networks up to 64 layers.
- Paper: Self-supervised Graph Learning for Recommendation, Jiancan Wu et al. (2020). It builds directly on edge dropout techniques like DropEdge as data augmentations to create multi-view self-supervised contrastive learning for graph-based recommendations.
- Paper: Graph Contrastive Learning with Augmentations, Yuning You et al. (2020). It incorporates edge perturbation and node dropping into a generalized graph contrastive learning framework (GraphCL) for self-supervised pre-training.
- Paper: Open Graph Benchmark: Datasets for Machine Learning on Graphs, Weihua Hu et al. (2020). It provides large-scale, standardized benchmarks for evaluating graph neural network regularizations and scalability techniques like DropEdge across diverse domains.
