Label-free Node Classification on Graphs with Large Language Models (LLMs)

Zhikai ChenHaitao MaoHongzhi WenHaoyu HanWei JinHaiyang ZhangHui LiuJiliang Tang

article2024ICLR100 citations

Proposes LLM-GNN, an active annotation framework that prompts large language models on an informative subset of graph nodes to train graph neural networks without human supervision, achieving high classification accuracy on massive text-attributed graphs for under one dollar.

Listen

Graph Neural Networks (GNNs) achieve strong performance in categorizing nodes within interconnected networks, such as academic citation maps or web-scale product catalogs. However, these systems traditionally rely on vast quantities of high-quality human annotations, which are slow, expensive, and difficult to acquire at scale. Large Language Models (LLMs) can categorize text-rich data without prior manual labeling, but they struggle to process structural connections directly and incur high computational and financial costs if applied to every node in a large network.

The article introduces and evaluates LLM-GNN, a label-free classification framework that combines the zero-shot capabilities of language models with the structural modeling efficiency of graph neural networks. The objective of the article is to demonstrate how strategically using language models to annotate a tiny, carefully selected subset of nodes enables training an accurate, low-cost GNN model to predict the remaining network without any human supervision.

The framework operates across four main steps. First, the system selects candidate nodes using active learning heuristics, including a difficulty-aware metric (C-Density) that identifies nodes closer to feature cluster centers, which are statistically easier for language models to classify accurately. Second, the language model annotates this small subset using prompts designed to output both class predictions and calibrated confidence scores. Third, a post-filtering step discards low-confidence annotations while balancing overall label diversity. Finally, a standard GNN is trained on these filtered, confidence-weighted annotations to classify the rest of the network. The authors validated the method across six benchmark datasets, including large-scale graphs containing millions of connections.

The findings show that LLM-GNN achieves high classification accuracy at a fraction of standard operational costs. On the massive OGBN-PRODUCTS dataset with over 2.4 million nodes, LLM-GNN reached an accuracy of 74.9%, matching the performance of a model trained on 400 human-labeled nodes while costing less than one dollar in model queries. In comparison, using language models directly to classify all nodes on that same dataset cost over $1,570—more than 2,000 times higher—for a nearly identical accuracy of 75.3%. In addition, the study revealed that combining active node selection with post-filtering and confidence-weighted loss functions consistently produced the most resilient models, and that errors from language model annotations were less disruptive to GNN training than synthetic random noise.

These results demonstrate that organizations can deploy high-performing graph models on large, unannotated datasets without investing in expensive manual labeling workflows. The framework eliminates the risk of excessive query costs by restricting language model calls to a tiny sample, while retaining the GNN's ability to propagate structural context across millions of items. For practical deployments, the authors recommend combining feature propagation selection with post-filtering and confidence-weighted training loss, as this configuration delivers the best balance of speed, accuracy, and budget control.

Readers should note that while the method performs robustly, the accuracy ceiling is inherently tied to the initial quality of the language model annotations. Highly ambiguous categories and extreme class imbalances can degrade performance if the active selection parameters are poorly calibrated. Nevertheless, the experimental results provide high confidence that selective language model annotation provides a cost-effective, scalable foundation for graph classification.

No sufficiently relevant recommendations were found.

Cover for Label-free Node Classification on Graphs with Large Language Models (LLMs)

Abstract

In recent years, there have been remarkable advancements in node classification achieved by Graph Neural Networks (GNNs). However, they necessitate abundant high-quality labels to ensure promising performance. In contrast, Large Language Models (LLMs) exhibit impressive zero-shot proficiency on text-attributed graphs. Yet, they face challenges in efficiently processing structural data and suffer from high inference costs. In light of these observations, this work introduces a label-free node classification on graphs with LLMs pipeline, LLM-GNN. It amalgamates the strengths of both GNNs and LLMs while mitigating their limitations. Specifically, LLMs are leveraged to annotate a small portion of nodes and then GNNs are trained on LLMs' annotations to make predictions for the remaining large portion of nodes. The implementation of LLM-GNN faces a unique challenge: how can we actively select nodes for LLMs to annotate and consequently enhance the GNN training? How can we leverage LLMs to obtain annotations of high quality, representativeness, and diversity, thereby enhancing GNN performance with less cost? To tackle this challenge, we develop an annotation quality heuristic and leverage the confidence scores derived from LLMs to advanced node selection. Comprehensive experimental results validate the effectiveness of LLM-GNN. In particular, LLM-GNN can achieve an accuracy of 74.9% on a vast-scale dataset \products with a cost less than 1 dollar.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Method
  • 3.1 An overview of LLM-GNN
  • 3.2 Difficulty-aware active node selection
  • 3.3 Confidence-aware annotations
  • 3.4 Post-filtering
  • 3.5 GNN training and prediction
  • 4 Experiment
  • 4.1 Experimental Settings
  • 4.2 RQ1. Impact of different active selection strategies
  • 4.3 (RQ2.) Comparison with other label-free node classification methods
  • 4.4 (RQ3.) How do different budgets affect the performance of our pipelines?
  • 4.5 (RQ4.) Characteristics of LLMs’ annotations
  • 5 Related Works
  • 6 Conclusion
  • 7 Acknowledgements
  • References
  • A Detailed Introductions of Baseline Models
  • A.1 Graph Active Learning Methods
  • A.2 Zero-shot classification baseline models
  • B More Related Works
  • C Datasets
  • D Prompts
  • E Complete results for the preliminary study on effectiveness of confidence-aware prompts
  • F Preliminary Observations of LLM’s annotations
  • F.1 Label-level observations
  • F.2 Relationship between annotation quality and C-density
  • G Effectiveness of confidence generated by LLMs
  • H Hyperparamters
  • I Comparing LLMs annotations to synthetic noisy labels
  • J Extra results for the comparative study in Table
  • K Theoretical Motivation
  • L Design philosophy behind LLMGNN

Knowls

  1. Knowl 1 — LLM-GNN Pipeline for Label-Free Node Classification

    model/method

    The LLM-GNN pipeline performs label-free node classification on text-attributed graphs (TAGs), represented as GT=(V,A,T,X)G_T = (V, A, T, X), where V={v1,…,vn}V = \{v_1, \dots, v_n\} is the node set, A∈{0,1}n×nA \in \{0, 1\}^{n \times n} is the adjacency matrix, T={t1,…,tn}T = \{t_1, \dots, t_n\} contains raw textual attributes for each node, and X={x1,…,xn}X = \{x_1, \dots, x_n\} contains text embeddings generated by SentenceBERT. The pipeline eliminates the need for human-annotated labels by combining Large Language Models (LLMs) as zero-shot annotators and Graph Neural Networks (GNNs) as structural learners through four stages:

    1. Difficulty-Aware Active Node Selection: A small subset of nodes Vanno⊂VV_{\text{anno}} \subset V is selected using a ranking aggregation that balances traditional active learning criteria (such as graph coverage and diversity) with node annotation difficulty.
    2. Confidence-Aware Annotation: An LLM (e.g., GPT-3.5-turbo) annotates the selected nodes using text attributes and outputs predicted class labels along with calibrated confidence scores fconf(vi)∈[0,100]f_{\text{conf}}(v_i) \in [0, 100].
    3. Post-Filtering: Low-quality or miscalibrated annotations are filtered out while preserving class diversity by evaluating confidence scores, class entropy changes, and clustering density.
    4. GNN Training and Inference: A GNN (e.g., Graph Convolutional Network, GCN) is trained on the filtered pseudo-labeled nodes using standard cross-entropy or confidence-weighted cross-entropy loss, where confidence scores serve as sample weights. The trained GNN then predicts labels for the remaining unlabeled nodes V∖VannoV \setminus V_{\text{anno}}.
  2. Knowl 2 — Difficulty-Aware Active Node Selection and C-Density Metric

    model/method

    To identify candidate nodes for LLM annotation that balance informativeness with annotation reliability, difficulty-aware active selection combines traditional active learning criteria with a clustering density metric called C-Density.

    kk-means clustering is applied directly to the SentenceBERT node feature space XX, where the number of clusters kk is set to the total number of distinct classes CC. For a node viv_i with feature vector xvix_{v_i} and closest cluster center CCviCC_{v_i}, the C-Density score is defined as:

    C-Density(vi)=11+∥xvi−xCCvi∥\text{C-Density}(v_i) = \frac{1}{1 + \|x_{v_i} - x_{CC_{v_i}}\|}

    To integrate C-Density into an arbitrary graph active selection scoring function fact(vi)f_{\text{act}}(v_i) (e.g., PageRank centrality, feature density, or partition-based diversity) without scale distortion, ranking aggregation is applied. Scores are converted to descending percentile ranks rfact(vi)∈[0,1]r_{f_{\text{act}}}(v_i) \in [0, 1] and rC-Density(vi)∈[0,1]r_{\text{C-Density}}(v_i) \in [0, 1], yielding the difficulty-aware score:

    fDA-act(vi)=α0×rfact(vi)+α1×rC-Density(vi)f_{\text{DA-act}}(v_i) = \alpha_0 \times r_{f_{\text{act}}}(v_i) + \alpha_1 \times r_{\text{C-Density}}(v_i)

    where α0\alpha_0 and α1\alpha_1 are weighting hyperparameters. Nodes with the highest fDA-act(vi)f_{\text{DA-act}}(v_i) are selected into the candidate annotation set VannoV_{\text{anno}} under a budget BB (commonly set to 20×C20 \times C).

  3. Knowl 3 — Confidence-Aware Post-Filtering via Change of Entropy

    model/method

    Directly discarding low-confidence LLM annotations can induce severe label distribution shift and reduce class diversity. To refine annotations while preserving balanced class representation, post-filtering uses a Change of Entropy (COE) metric alongside LLM confidence.

    Let VselV_{\text{sel}} be the current set of annotated candidate nodes and y~Vsel\tilde{y}_{V_{\text{sel}}} denote the LLM-generated pseudo-labels. The COE of removing node viv_i is computed as:

    COE(vi)=H(y~Vsel∖{vi})−H(y~Vsel)\text{COE}(v_i) = H(\tilde{y}_{V_{\text{sel}} \setminus \{v_i\}}) - H(\tilde{y}_{V_{\text{sel}}})

    where H(⋅)H(\cdot) represents the Shannon entropy of the empirical label distribution. A low or negative COE indicates that removing viv_i hurts the overall label diversity.

    The overall filtering score combines confidence, diversity, and feature density using ranking aggregation:

    ffilter(vi)=β0×rfconf(vi)+β1×rCOE(vi)+β2×rC-Density(vi)f_{\text{filter}}(v_i) = \beta_0 \times r_{f_{\text{conf}}}(v_i) + \beta_1 \times r_{\text{COE}}(v_i) + \beta_2 \times r_{\text{C-Density}}(v_i)

    where rfconf(vi)r_{f_{\text{conf}}}(v_i), rCOE(vi)r_{\text{COE}}(v_i), and rC-Density(vi)r_{\text{C-Density}}(v_i) are the descending percentile ranks of LLM confidence fconf(vi)f_{\text{conf}}(v_i), the entropy delta COE(vi)\text{COE}(v_i), and the cluster density C-Density(vi)\text{C-Density}(v_i), respectively, weighted by hyperparameters β0,β1,β2\beta_0, \beta_1, \beta_2. The filtering process iteratively removes the node with the lowest ffilter(vi)f_{\text{filter}}(v_i) and updates entropy until a target training budget is reached.

  4. Knowl 4 — Theoretical Justification for C-Density in LLM Annotation Reliability

    theoretical result

    Let QQ denote the parameter distribution of the LLM annotator, PP denote the parameter distribution of the sentence encoder (e.g., SentenceBERT), YL∈RN×MY_L \in \mathbb{R}^{N \times M} denote the true one-hot labels for NN nodes across MM classes, and Y∈RN×MY \in \mathbb{R}^{N \times M} denote the LLM pseudo-labels. Let X∈RN×dX \in \mathbb{R}^{N \times d} be the node embeddings produced by the encoder and XL∈RN×dX_L \in \mathbb{R}^{N \times d} be the latent ground-truth representations.

    For any annotation yy generated by the LLM, the discrepancy function f(y)=ℓ(y,yL)=ln⁡(1+∥x−xL∥2)f(y) = \ell(y, y_L) = \ln(1 + \|x - x_L\|^2) satisfies the upper bound:

    Eθ∼Q[f(y)]≤log⁡Eθ′∼P[exp⁡(f(y))]+KL(Q∥P)\mathbb{E}_{\theta \sim Q}[f(y)] \le \log \mathbb{E}_{\theta' \sim P}[\exp(f(y))] + \text{KL}(Q \parallel P)

    Under Gaussian assumptions for the encoder representation X∼N(μi,σi2I)X \sim \mathcal{N}(\mu_i, \sigma_i^2 I) and the true latent representation XL∼N(μj,σj2I)X_L \sim \mathcal{N}(\mu_j, \sigma_j^2 I) for class ii, the expected squared distance expands to:

    E[∥X−XL∥2]=nσi2+nσj2+∥μi−μj∥2\mathbb{E}[\|X - X_L\|^2] = n\sigma_i^2 + n\sigma_j^2 + \|\mu_i - \mu_j\|^2

    Because (μj,σj)(\mu_j, \sigma_j) is unknown and fixed, minimizing the upper bound on LLM annotation error Eθ∼Q[f(y)]\mathbb{E}_{\theta \sim Q}[f(y)] reduces to minimizing the local feature variance σi2\sigma_i^2. For an arbitrary node, σi\sigma_i is approximated by its Euclidean distance to its cluster center ∥xvi−xCCvi∥\|x_{v_i} - x_{CC_{v_i}}\|. Consequently, selecting nodes with higher C-Density (closer to cluster centroids) minimizes the upper bound on expected annotation error.

  5. Knowl 5 — Comparison of Prompt Strategies for Confidence-Calibrated Annotation

    empirical result

    Different LLM prompting strategies produce varying levels of zero-shot annotation accuracy and confidence calibration on text-attributed graph nodes:

    1. Vanilla (zero-shot): Asks the LLM for a label and confidence score directly. Achieves competitive zero-shot accuracy (e.g., 68.33%68.33\% on Cora, 75.33%75.33\% on ogbn-products, 68.33%68.33\% on WikiCS) with baseline token cost (1.0x).
    2. Few-shot / One-shot: Providing a 1-shot in-context demonstration improves accuracy modestly (e.g., +1.34%+1.34\% on Cora, +3.34%+3.34\% on ogbn-products, +3.67%+3.67\% on WikiCS) but approximately doubles token cost (1.8x to 2.4x).
    3. Reasoning-based Prompts (Chain-of-Thought): Frequently fail strict output format compliance and significantly increase query latency and cost.
    4. Top-K & Self-Consistency (Most Voting): Querying multiple outputs or generating top candidates improves reliability with minimal overhead (1.1x to 1.4x cost).
    5. Zero-shot Hybrid Prompt: Combining Top-K choice with multi-query consistency produces the best calibrated confidence scores: sorting candidate nodes by hybrid prompt confidence yields a monotonic relationship where top-confidence subsets reach over 85%–90%85\%\text{--}90\% accuracy across datasets at a modest cost overhead (1.2x to 1.5x of vanilla zero-shot).
  6. Knowl 6 — Cost and Accuracy Benchmarking of Label-Free Node Classification

    data/table

    The LLM-GNN pipeline was evaluated against zero-shot graph classification baselines (SES, TAG-Z), a zero-shot text classification model (BART-large-MNLI), and direct LLM prediction (LLMs-as-Predictors) on the large-scale benchmarks OGBN-ARXIV and OGBN-PRODUCTS using GPT-3.5-turbo.

    OGBN-ARXIV OGBN-PRODUCTS
    Methods Acc (%) Cost ($) Acc (%) Cost ($)
    SES 13.08 N/A 6.67 N/A
    TAG-Z 37.08 N/A 47.08 N/A
    BART-large-MNLI 13.20 N/A 28.80 N/A
    LLMs-as-Predictors 73.33 79.00 75.33 1572.00
    LLM-GNN 66.32 0.63 74.91 0.74

    On OGBN-PRODUCTS (2,449,0292,449,029 nodes), LLM-GNN achieves 74.91%74.91\% classification accuracy at a financial cost of $0.74\$0.74, performing on par with direct LLM inference (75.33%75.33\%) while reducing monetary cost by 2,124×2,124\times (down from $1572.00\$1572.00). Classical zero-shot node classification methods (SES: 6.67%6.67\%, TAG-Z: 47.08%47.08\%) fall substantially behind.

  7. Knowl 7 — Performance of Active Selection and Post-Filtering with GCN

    empirical result

    Integrating post-filtering (PS) and difficulty-aware selection (DA) with traditional graph active learning methods to train a 2-layer GCN yields consistent accuracy improvements across small, medium, and large-scale graph benchmarks under a fixed budget of 20×C20 \times C nodes:

    • Feature Propagation (FeatProp): Applying post-filtering with weighted cross-entropy loss (PS-FeatProp-W) achieves top-tier performance on Cora (76.23±0.07%76.23 \pm 0.07\%, vs. Random 70.48%70.48\%), CiteSeer (68.64±0.71%68.64 \pm 0.71\%, vs. Random 65.11%65.11\%), PubMed (78.84±1.05%78.84 \pm 1.05\%, vs. Random 72.98%72.98\%), WikiCS (64.72±0.19%64.72 \pm 0.19\%, vs. Random 60.69%60.69\%), and OGBN-PRODUCTS (74.54±0.24%74.54 \pm 0.24\%, vs. Random 70.40%70.40\%).
    • Scalability: Advanced graph active selection algorithms like RIM and GraphPart fail with out-of-time (OOT) errors on million-scale graphs (OGBN-ARXIV and OGBN-PRODUCTS), whereas FeatProp-based methods scale efficiently.
    • Loss Function: Using LLM confidence scores as sample weights in a weighted cross-entropy loss (-W) consistently outperforms standard cross-entropy across nearly all active selection variants.
    • Combining DA and PS: Applying Difficulty-Aware selection and Post-Filtering simultaneously often causes performance degradation due to over-constraining the candidate pool and worsening class imbalance; using either DA or PS individually with weighted loss provides superior results.
  8. Knowl 8 — Asymmetric Semantic Structure and Benign Dynamics of LLM Label Noise

    empirical result

    Empirical analysis of LLM-generated pseudo-labels reveals distinct noise characteristics compared to synthetic uniform label noise:

    1. Class-Dependent Variance: Annotation accuracy varies sharply across classes. On WikiCS, zero-shot LLM predictions reach 100%100\% accuracy on class c0c_0 (Computational Linguistics) but only 31%31\% on class c7c_7 (Distributed Computing Architecture).
    2. Asymmetric Class Confusion: When LLMs misclassify a node, the errors do not distribute uniformly across alternative classes. Instead, they flip asymmetrically into semantically related categories. For example, WikiCS class c7c_7 frequently flips to class c8c_8 (Web Technology, 47%47\% flip rate), but nodes from class c8c_8 almost never flip into class c7c_7.
    3. Resistance to Overfitting: When training GNNs, LLM-generated label noise exhibits much more benign training dynamics than synthetic random noise with an equivalent error rate: GNN test accuracy remains stable across training epochs without the severe test-set performance collapse typical of synthetic uniform label noise.
  9. Knowl 9 — Ineffectiveness of Structure-Aware Prompts for LLM Graph Annotation

    limitation

    Incorporating graph neighborhood textual descriptions directly into LLM prompts (structure-aware prompting) without ground-truth neighbor labels degrades annotation performance compared to structure-free prompts based solely on ego-node text:

    • On Cora, FeatProp selection accuracy drops from 72.82±0.08%72.82 \pm 0.08\% with structure-free prompts to 69.08±0.39%69.08 \pm 0.39\% with structure-aware prompts. With post-filtering (FeatProp + PS), accuracy drops from 75.54±0.34%75.54 \pm 0.34\% to 67.55±0.70%67.55 \pm 0.70\%.
    • On CiteSeer, FeatProp selection accuracy drops from 66.61±0.55%66.61 \pm 0.55\% (structure-free) to 65.92±0.43%65.92 \pm 0.43\% (structure-aware), and FeatProp + PS drops from 69.06±0.32%69.06 \pm 0.32\% to 67.80±0.45%67.80 \pm 0.45\%.

    Because LLMs have limited native capacity to reason over complex graph topology and summarize neighborhood context in the absence of neighbor ground-truth labels, injecting raw neighbor texts adds noise. Decoupling the pipeline—using LLMs exclusively for text-based annotation and relying on GNN message passing to model graph structure—is empirically superior.

Coverage note — None omitted; all core contributions, theoretical bounds, empirical algorithms, prompt evaluation results, comparative benchmarks, and noise analyses have been captured as self-contained knowls.

References

  1. 1.Yingbin Bai, Erkun Yang, Bo Han, Yanhua Yang, Jiatong Li, Yinian Mao, Gang Niu, and Tongliang Liu. Understanding and improving early stopping for learning with noisy labels. Advances in Neural Information Processing Systems, 34:24392–24403, 2021.
  2. 2.Parikshit Bansal and Amit Sharma. Large language models as annotators: Enhancing generalization of nlp models at minimal cost. arXiv preprint arXiv:2306.15766, 2023.
  3. 3.Hongyun Cai, Vincent W Zheng, and Kevin Chen-Chuan Chang. Active learning for graph embedding. arXiv preprint arXiv:1705.05085, 2017.
  4. 4.Zhikai Chen, Haitao Mao, Hang Li, Wei Jin, Hongzhi Wen, Xiaochi Wei, Shuaiqiang Wang, Dawei Yin, Wenqi Fan, Hui Liu, et al. Exploring the potential of large language models (llms) in learning on graphs. arXiv preprint arXiv:2307.03393, 2023.
  5. 5.Bosheng Ding, Chengwei Qin, Linlin Liu, Lidong Bing, Shafiq Joty, and Boyang Li. Is gpt-3 a good data annotator? arXiv preprint arXiv:2212.10450, 2022.
  6. 6.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. A survey for in-context learning. arXiv preprint arXiv:2301.00234, 2022.
  7. 7.Li Gao, Hong Yang, Chuan Zhou, Jia Wu, Shirui Pan, and Yue Hu. Active discriminative network representation learning. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18, pp. 2142–2148. International Joint Conferences on Artificial Intelligence Organization, 7 2018. doi: 10.24963/ijcai.2018/296. URL https://doi.org/10.24963/ijcai.2018/296.
  8. 8.Fabrizio Gilardi, Meysam Alizadeh, and Maèl Kubli. Chatgpt outperforms crowd-workers for text-annotation tasks. arXiv preprint arXiv:2303.15056, 2023.
  9. 9.C. Lee Giles, Kurt D. Bollacker, and Steve Lawrence. Citeseer: An automatic citation indexing system. In Proceedings of the Third ACM Conference on Digital Libraries, DL ’98, pp. 89–98, New York, NY, USA, 1998. ACM. ISBN 0-89791-965-3. doi: 10.1145/276675.276685. URL http://doi.acm.org/10.1145/276675.276685.
  10. 10.Jiayan Guo, Lun Du, and Hengyu Liu. Gpt4graph: Can large language models understand graph structured data? an empirical evaluation and benchmarking. arXiv preprint arXiv:2305.15066, 2023.
  11. 11.Will Hamilton, Zhitao Ying, and Jure Leskovec. Inductive representation learning on large graphs. Advances in neural information processing systems, 30, 2017.
  12. 12.Haibo He and Edwardo A. Garcia. Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 21(9):1263–1284, 2009. doi: 10.1109/TKDE.2008.239.
  13. 13.Xiaoxin He, Xavier Bresson, Thomas Laurent, and Bryan Hooi. Explanations as features: Llm-based features for text-attributed graphs. arXiv preprint arXiv:2305.19523, 2023a.
  14. 14.Xingwei He, Zhenghao Lin, Yeyun Gong, Alex Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, Weizhu Chen, et al. Annollm: Making large language models to be better crowd-sourced annotators. arXiv preprint arXiv:2303.16854, 2023b.
  15. 15.Shengding Hu, Zheng Xiong, Meng Qu, Xingdi Yuan, Marc-Alexandre Côté, Zhiyuan Liu, and Jian Tang. Graph policy network for transferable active learning on graphs. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 10174–10185. Curran Associates, Inc., 2020a. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/73740ea85c4ec25f00f9acbd859f861d-Paper.pdf.
  16. 16.Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. Open graph benchmark: Datasets for machine learning on graphs. Advances in neural information processing systems, 33:22118–22133, 2020b.
  17. 17.Jin Huang, Xingjian Zhang, Qiaozhu Mei, and Jiaqi Ma. Can llms effectively leverage graph structural information: When and why. arXiv preprint arXiv:2309.16595, 2023.
  18. 18.Sheng-jun Huang, Rong Jin, and Zhi-Hua Zhou. Active learning by querying informative and representative examples. In J. Lafferty, C. Williams, J. Shawe-Taylor, R. Zemel, and A. Culotta (eds.), Advances in Neural Information Processing Systems, volume 23. Curran Associates, Inc., 2010. URL https://proceedings.neurips.cc/paper_files/paper/2010/file/5487315b1286f907165907aa8fc96619-Paper.pdf.
  19. 19.Wei Ju, Yifang Qin, Siyu Yi, Zhengyang Mao, Kangjie Zheng, Luchen Liu, Xiao Luo, and Ming Zhang. Zero-shot node classification with graph contrastive embedding network. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=8wGXnjRLSy.
  20. 20.Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016.
  21. 21.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. CoRR, abs/1910.13461, 2019. URL http://arxiv.org/abs/1910.13461.
  22. 22.Yuexin Li and Bryan Hooi. Prompt-based zero-and few-shot node classification: A multimodal approach. arXiv preprint arXiv:2307.11572, 2023.
  23. 23.Hao Liu, Jiarui Feng, Lecheng Kong, Ningyue Liang, Dacheng Tao, Yixin Chen, and Muhan Zhang. One for all: Towards training one graph model for all classification tasks. arXiv preprint arXiv:2310.00149, 2023.
  24. 24.Jiaqi Ma, Ziqiao Ma, Joyce Chai, and Qiaozhu Mei. Partition-based active learning for graph neural networks. arXiv preprint arXiv:2201.09391, 2022.
  25. 25.Yao Ma and Jiliang Tang. Deep learning on graphs. Cambridge University Press, 2021.
  26. 26.Andrew McCallum, Kamal Nigam, Jason D. M. Rennie, and Kristie Seymore. Automating the construction of internet portals with machine learning. Information Retrieval, 3:127–163, 2000.
  27. 27.Péter Mernyei and Cătălina Cangea. Wiki-cs: A wikipedia-based benchmark for graph neural networks. arXiv preprint arXiv:2007.02901, 2020.
  28. 28.Nicholas Pangakis, Samuel Wolken, and Neil Fasching. Automated annotation with generative ai requires validation. arXiv preprint arXiv:2306.00176, 2023.
  29. 29.Xiaohuan Pei, Yanxi Li, and Chang Xu. Gpt self-supervision for a better data annotator. arXiv preprint arXiv:2306.04349, 2023.
  30. 30.Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. In Conference on Empirical Methods in Natural Language Processing, 2019. URL https://api.semanticscholar.org/CorpusID:201646309.
  31. 31.Zhicheng Ren, Yifu Yuan, Yuxin Wu, Xiaxuan Gao, Yewen Wang, and Yizhou Sun. Dissimilar nodes improve graph active learning. arXiv preprint arXiv:2212.01968, 2022.
  32. 32.Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad. Collective classification in network data. AI Magazine, 29(3):93, Sep. 2008. doi: 10.1609/aimag.v29i3.2157. URL https://ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/2157.
  33. 33.C. E. Shannon. A mathematical theory of communication. The Bell System Technical Journal, 27(3):379–423, 1948. doi: 10.1002/j.1538-7305.1948.tb01338.x.
  34. 34.Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee. Learning from noisy labels with deep neural networks: A survey. IEEE Transactions on Neural Networks and Learning Systems, 2022.
  35. 35.Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv preprint arXiv:2305.14975, 2023.
  36. 36.Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017.
  37. 37.Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan, Xiaochuang Han, and Yulia Tsvetkov. Can language models solve graph problems in natural language? arXiv preprint arXiv:2305.10037, 2023a.
  38. 38.Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. A survey on large language model based autonomous agents. arXiv preprint arXiv:2308.11432, 2023b.
  39. 39.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
  40. 40.Zheng Wang, Jialong Wang, Yuchen Guo, and Zhiguo Gong. Zero-shot node classification with decomposed graph prototype network. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp. 1769–1779, 2021.
  41. 41.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022.
  42. 42.Jiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu, Gang Niu, and Yang Liu. Learning with noisy labels revisited: A study using real-world human annotations. arXiv preprint arXiv:2110.12088, 2021.
  43. 43.Yuexin Wu, Yichong Xu, Aarti Singh, Yiming Yang, and Artur Dubrawski. Active learning for graph neural networks via node feature propagation. arXiv preprint arXiv:1910.07567, 2019.
  44. 44.Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. arXiv preprint arXiv:2306.13063, 2023.
  45. 45.Zhilin Yang, William W. Cohen, and Ruslan Salakhutdinov. Revisiting semi-supervised learning with graph embeddings. ArXiv, abs/1603.08861, 2016. URL https://api.semanticscholar.org/CorpusID:7008752.
  46. 46.Ruosong Ye, Caiqi Zhang, Runhui Wang, Shuyuan Xu, and Yongfeng Zhang. Natural language is all a graph needs. arXiv preprint arXiv:2308.07134, 2023.
  47. 47.Qin Yue, Jiye Liang, Junbiao Cui, and Liang Bai. Dual bidirectional graph convolutional networks for zero-shot node classification. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’22, pp. 2408–2417, New York, NY, USA, 2022. Association for Computing Machinery. ISBN 9781450393850. doi: 10.1145/3534678.3539316. URL https://doi.org/10.1145/3534678.3539316.
  48. 48.Wentao Zhang, Yu Shen, Yang Li, Lei Chen, Zhi Yang, and Bin Cui. Alg: Fast and accurate active learning framework for graph convolutional networks. In Proceedings of the 2021 International Conference on Management of Data, SIGMOD ’21, pp. 2366–2374, New York, NY, USA, 2021a. Association for Computing Machinery. ISBN 9781450383431. doi: 10.1145/3448016.3457325. URL https://doi.org/10.1145/3448016.3457325.
  49. 49.Wentao Zhang, Yexin Wang, Zhenbang You, Meng Cao, Ping Huang, Jiulong Shan, Zhi Yang, and Bin Cui. Rim: Reliable influence-based active learning on graphs. Advances in Neural Information Processing Systems, 34:27978–27990, 2021b.
  50. 50.Wentao Zhang, Zhi Yang, Yexin Wang, Yu Shen, Yang Li, Liang Wang, and Bin Cui. Grain: Improving data efficiency of graph neural networks via diversified influence maximization. arXiv preprint arXiv:2108.00219, 2021c.
  51. 51.Yuheng Zhang, Hanghang Tong, Yinglong Xia, Yan Zhu, Yuejie Chi, and Lei Ying. Batch active learning with graph neural networks via multi-agent deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 9118–9126, 2022.
  52. 52.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. Automatic chain of thought prompting in large language models. In The Eleventh International Conference on Learning Representations (ICLR 2023), 2023.
  53. 53.Jianan Zhao, Le Zhuo, Yikang Shen, Meng Qu, Kai Liu, Michael Bronstein, Zhaocheng Zhu, and Jian Tang. Graphtext: Graph reasoning in text space. arXiv preprint arXiv:2310.01089, 2023.
  54. 54.Qi Zhu, Natalia Ponomareva, Jiawei Han, and Bryan Perozzi. Shift-robust gnns: Overcoming the limitations of localized graph training data. Advances in Neural Information Processing Systems, 34:27965–27977, 2021a.
  55. 55.Zhaowei Zhu, Yiwen Song, and Yang Liu. Clusterability as an alternative to anchor points when learning with noisy labels. In International Conference on Machine Learning, pp. 12912–12923. PMLR, 2021b.

Citation

MLA
Chen, Z., et al. “Label-free Node Classification on Graphs with Large Language Models (LLMS)”. arXiv, 2023, http://arxiv.org/abs/2310.04668v3.
APA
Chen, Z., Mao, H., Wen, H., Han, H., Jin, W., Zhang, H., Liu, H., & Tang, J. (2023). Label-free Node Classification on Graphs with Large Language Models (LLMS). arXiv. http://arxiv.org/abs/2310.04668v3
Chicago
Chen, Z., H. Mao, H. Wen, et al. 2023. “Label-free Node Classification on Graphs with Large Language Models (LLMS)”. arXiv. http://arxiv.org/abs/2310.04668v3.
Harvard
Chen, Z. et al. (2023) “Label-free Node Classification on Graphs with Large Language Models (LLMS)”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.04668v3.
Vancouver
1. Chen Z, Mao H, Wen H, Han H, Jin W, Zhang H, Liu H, Tang J (2023) Label-free Node Classification on Graphs with Large Language Models (LLMS). arXiv

BibTeX

@article{chen2023label,
  title = {Label-free Node Classification on Graphs with Large Language Models (LLMS)},
  author = {Chen, Zhikai and Mao, Haitao and Wen, Hongzhi and Han, Haoyu and Jin, Wei and Zhang, Haiyang and Liu, Hui and Tang, Jiliang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.04668v3},
  eprint = {2310.04668}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors