Label-free Node Classification on Graphs with Large Language Models (LLMs)
Zhikai ChenHaitao MaoHongzhi WenHaoyu HanWei JinHaiyang ZhangHui LiuJiliang Tang
Proposes LLM-GNN, an active annotation framework that prompts large language models on an informative subset of graph nodes to train graph neural networks without human supervision, achieving high classification accuracy on massive text-attributed graphs for under one dollar.
Graph Neural Networks (GNNs) achieve strong performance in categorizing nodes within interconnected networks, such as academic citation maps or web-scale product catalogs. However, these systems traditionally rely on vast quantities of high-quality human annotations, which are slow, expensive, and difficult to acquire at scale. Large Language Models (LLMs) can categorize text-rich data without prior manual labeling, but they struggle to process structural connections directly and incur high computational and financial costs if applied to every node in a large network.
The article introduces and evaluates LLM-GNN, a label-free classification framework that combines the zero-shot capabilities of language models with the structural modeling efficiency of graph neural networks. The objective of the article is to demonstrate how strategically using language models to annotate a tiny, carefully selected subset of nodes enables training an accurate, low-cost GNN model to predict the remaining network without any human supervision.
The framework operates across four main steps. First, the system selects candidate nodes using active learning heuristics, including a difficulty-aware metric (C-Density) that identifies nodes closer to feature cluster centers, which are statistically easier for language models to classify accurately. Second, the language model annotates this small subset using prompts designed to output both class predictions and calibrated confidence scores. Third, a post-filtering step discards low-confidence annotations while balancing overall label diversity. Finally, a standard GNN is trained on these filtered, confidence-weighted annotations to classify the rest of the network. The authors validated the method across six benchmark datasets, including large-scale graphs containing millions of connections.
The findings show that LLM-GNN achieves high classification accuracy at a fraction of standard operational costs. On the massive OGBN-PRODUCTS dataset with over 2.4 million nodes, LLM-GNN reached an accuracy of 74.9%, matching the performance of a model trained on 400 human-labeled nodes while costing less than one dollar in model queries. In comparison, using language models directly to classify all nodes on that same dataset cost over $1,570—more than 2,000 times higher—for a nearly identical accuracy of 75.3%. In addition, the study revealed that combining active node selection with post-filtering and confidence-weighted loss functions consistently produced the most resilient models, and that errors from language model annotations were less disruptive to GNN training than synthetic random noise.
These results demonstrate that organizations can deploy high-performing graph models on large, unannotated datasets without investing in expensive manual labeling workflows. The framework eliminates the risk of excessive query costs by restricting language model calls to a tiny sample, while retaining the GNN's ability to propagate structural context across millions of items. For practical deployments, the authors recommend combining feature propagation selection with post-filtering and confidence-weighted training loss, as this configuration delivers the best balance of speed, accuracy, and budget control.
Readers should note that while the method performs robustly, the accuracy ceiling is inherently tied to the initial quality of the language model annotations. Highly ambiguous categories and extreme class imbalances can degrade performance if the active selection parameters are poorly calibrated. Nevertheless, the experimental results provide high confidence that selective language model annotation provides a cost-effective, scalable foundation for graph classification.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). Establishes the foundational semi-supervised node classification framework with Graph Convolutional Networks that LLM-GNN builds upon to train downstream classifiers with limited node annotations.
- Paper: Combining active learning and semi-supervised learning using Gaussian fields and harmonic functions, Xiaojin Zhu et al. (2003). Introduces the core paradigm of combining active learning node selection with semi-supervised graph label propagation, directly motivating LLM-GNN's active selection of nodes for annotation.
- Paper: Graph Convolutional Networks for Text Classification, Liang Yao et al. (2018). Provides the foundational methodology for applying graph neural networks to text-attributed networks and document classification tasks.
- Paper: Inductive Representation Learning on Large Graphs, William L. Hamilton et al. (2017). Develops inductive neighborhood aggregation for large-scale graphs, which underlies the GNN training and scalable inference pipeline evaluated in LLM-GNN.
- Paper: Deeper Insights into Graph Convolutional Networks for Semi-Supervised Learning, Qimai Li et al. (2018). Analyzes the behavior and limits of graph convolutions under extremely low label rates, contextualizing why active node selection is crucial when relying on scarce LLM annotations.
No sufficiently relevant recommendations were found.
