Universal Prompt Tuning for Graph Neural Networks

Taoran FangYunchao ZhangYang YangChunping WangLei Chen

article2023NeurIPS119 citations

Proposes Graph Prompt Feature, a universal feature-space prompting method for graph neural networks that theoretically unifies diverse pre-training strategies and consistently outperforms fine-tuning across full-shot and few-shot downstream tasks.

Listen

Graph neural networks are critical tools for analyzing complex relational data, such as molecular structures and biological networks. However, adapting pre-trained graph models to specific tasks typically relies on fine-tuning, which updates all model parameters. This traditional approach creates major challenges: it risks catastrophic forgetting, demands significant computational resources, performs poorly when labeled data is scarce, and suffers from task misalignment because graphs use diverse pre-training strategies. Prior attempts to implement prompt tuning—a method that modifies input data rather than internal model weights—were heavily specialized for single pre-training objectives, such as edge prediction, leaving no universal solution across different graph architectures.

The article develops and evaluates a universal framework called Graph Prompt Feature (GPF) and its advanced variant (GPF-plus) to adapt pre-trained graph neural networks across any pre-training strategy. The primary objective is to demonstrate that directly modifying the input graph’s feature space can theoretically and empirically match or exceed the performance of traditional fine-tuning and specialized prompt methods while updating only a minute fraction of parameters.

To evaluate this framework, the authors conducted extensive empirical experiments using a standard five-layer Graph Isomorphism Network across chemistry and biology benchmarks, including molecular classification and protein function prediction tasks, as well as social network datasets. The methodology tested five distinct pre-training strategies across full-data scenarios, few-shot environments (50 and 100 labeled samples), and comparative baselines such as linear probing, partial-layer tuning, and existing prompt methods. These empirical evaluations were supported by formal mathematical proofs establishing the universal representational power and loss bounds of the proposed feature-prompting technique.

The investigation produced four central findings. First, GPF and GPF-plus consistently outperform traditional fine-tuning across all pre-training strategies, achieving average performance improvements of approximately 1.4% in full-data settings and 3.2% to 3.4% in few-shot settings. Second, the method delivers extreme parameter efficiency, requiring over 99% fewer tunable parameters than fine-tuning (using approximately 0.02% of parameters for GPF and under 0.7% for GPF-plus). Third, against specialized graph prompting baselines designed specifically for edge prediction, the proposed methods demonstrated substantial advantages, outperforming them by 3% to 13%. Finally, GPF-plus slightly outperforms standard GPF across most benchmark tests due to its ability to assign node-specific prompt features using an attentive basis mechanism.

These findings indicate that organizations can significantly cut computational and storage costs by freezing pre-trained graph backbones and training only lightweight feature prompts. This strategy mitigates overfitting and preserves generalization, making it exceptionally valuable for high-stakes domains with limited labeled samples, such as drug discovery and rare disease research. Organizations deploying graph models should consider adopting universal prompt tuning as a direct replacement for full fine-tuning, choosing standard GPF for simplicity or GPF-plus when maximum task flexibility is required.

While the theoretical proofs and empirical tests provide high confidence across standard graph classification tasks, current evaluations primarily reflect benchmark datasets in chemistry and biology under standard regression and classification assumptions. Future efforts should focus on validating the framework in production environments and exploring automated selection mechanisms for prompt hyper-parameters across broader graph-level and node-level operational tasks.

  • Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). Introduces parameter-efficient prompt tuning with continuous soft prompt vectors on frozen backbones, establishing the foundational tuning paradigm adapted to graph feature spaces.
  • Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). Pioneers visual prompt tuning by modifying the input space of frozen vision models, providing the non-text prompt-tuning precursor for input-space graph prompt design.
  • Paper: GPT Understands, Too, Xiao Liu et al. (2021). Develops continuous prompt optimization via neural prompt encoders, which informs continuous and basis-attentive prompt feature modeling.
  • Paper: Strategies for Pre-training Graph Neural Networks, Weihua Hu et al. (2020). Establishes pre-training and downstream adaptation benchmarks on Graph Isomorphism Networks across chemistry and biology, providing the pre-training paradigms and evaluation setup utilized by the source.
  • Paper: How Powerful are Graph Neural Networks?, Keyulu Xu et al. (2019). Introduces the Graph Isomorphism Network (GIN) architecture and multiset representational theory that serves as the backbone model and theoretical basis in the source.
  • Paper: Graph Contrastive Learning with Augmentations, Yuning You et al. (2020). Defines self-supervised graph contrastive pre-training strategies that represent key pre-trained baselines tested under universal prompt adaptation.
  • Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). Provides the foundational message-passing formulation for semi-supervised graph representation learning underpinning standard GNN backbones.

No sufficiently relevant recommendations were found.

Cover for Universal Prompt Tuning for Graph Neural Networks

Abstract

In recent years, prompt tuning has sparked a research surge in adapting pre-trained models. Unlike the unified pre-training strategy employed in the language field, the graph field exhibits diverse pre-training strategies, posing challenges in designing appropriate prompt-based tuning methods for graph neural networks. While some pioneering work has devised specialized prompting functions for models that employ edge prediction as their pre-training tasks, these methods are limited to specific pre-trained GNN models and lack broader applicability. In this paper, we introduce a universal prompt-based tuning method called Graph Prompt Feature (GPF) for pre-trained GNN models under any pre-training strategy. GPF operates on the input graph’s feature space and can theoretically achieve an equivalent effect to any form of prompting function. Consequently, we no longer need to illustrate the prompting function corresponding to each pre-training strategy explicitly. Instead, we employ GPF to obtain the prompted graph for the downstream task in an adaptive manner. We provide rigorous derivations to demonstrate the universality of GPF and make guarantee its effectiveness. The experimental results under various pre-training strategies indicate that our method performs better than fine-tuning, with an average improvement of about 1.4% in full-shot scenarios and about 3.2% in few-shot scenarios. Moreover, our method significantly outperforms existing specialized prompt-based tuning methods when applied to models utilizing the pre-training strategy they specialize in. These numerous advantages position our method as a compelling alternative to fine-tuning for downstream adaptations. Our code is available at: https://github.com/zjunet/GPF.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Methodology
  • 3.1 Preliminaries
  • 3.2 Graph Prompt Tuning
  • 3.3 Universal Graph Prompt Design
  • 3.4 Theoretical Analysis
  • 4 Experiments
  • 4.1 Experiment Setup
  • 4.2 Main Results
  • 4.3 Comparison with Existing Graph Prompt-based Methods
  • 4.4 Additional Experiments
  • 5 Conclusion
  • 6 Acknowledgements
  • References
  • A Extra Materials for Section 3
  • A.1 Extension to node-wise tasks
  • A.2 Proof for Theorem 1
  • A.3 Proof for Theorem 2
  • B More Information on Experiments
  • B.1 Details of the datasets
  • B.2 Details of pre-training strategies
  • B.3 Results of few-shot graph classification
  • B.4 Parameter efficiency analysis
  • B.5 Comparison with linear probing
  • B.6 Comparison with other tuning methods
  • B.7 Extra results on GCC
  • B.8 Hyper-parameter settings

Knowls

  1. Knowl 1 — Graph Prompt Feature (GPF)

    model/method

    For a graph G=(A,X)G=(A,X) with adjacency matrix A∈{0,1}N×NA\in\{0,1\}^{N\times N} and node-feature matrix X∈RN×FX\in\mathbb{R}^{N\times F}, Graph Prompt Feature (GPF) freezes the parameters of a pre-trained GNN ff and adds one shared learnable vector p∈RFp\in\mathbb{R}^{F} to every node feature. If xi∈RFx_i\in\mathbb{R}^{F} is the feature of node viv_i, the prompted feature matrix is

    X∗={x1+p,x2+p,…,xN+p}.X^*=\{x_1+p,x_2+p,\ldots,x_N+p\}.

    A downstream projection head θ\theta and the prompt vector pp are optimized using downstream labeled graphs, while ff remains frozen; equivalently, the downstream objective maximizes Pf,θ(y∣(A,X∗))P_{f,\theta}(y\mid(A,X^*)) over pp and θ\theta. At test time, the learned vector is added to every node feature before the graph is passed through the frozen GNN. Because GPF modifies only the feature space and does not depend on the pre-training objective or graph architecture, the same construction is intended for models pre-trained with any graph pre-training strategy.

  2. Knowl 2 — Graph Prompt Feature-Plus (GPF-plus)

    model/method

    GPF-plus increases the expressiveness of GPF by assigning a node-dependent prompt vector. For node viv_i with original feature xi∈RFx_i\in\mathbb{R}^{F}, the prompted feature is xi+pix_i+p_i, where pi∈RFp_i\in\mathbb{R}^{F}. To avoid storing a separate vector for every node, GPF-plus learns kk basis vectors p1b,…,pkb∈RFp^{b}_1,\ldots,p^{b}_k\in\mathbb{R}^{F} and kk projection vectors a1,…,ak∈RFa_1,\ldots,a_k\in\mathbb{R}^{F}. The node-specific prompt is

    pi=∑j=1kαi,jpjb,αi,j=exp⁡(ajTxi)∑l=1kexp⁡(alTxi).p_i=\sum_{j=1}^{k}\alpha_{i,j}p^{b}_j, \qquad \alpha_{i,j}=\frac{\exp(a_j^{\mathsf T}x_i)}{\sum_{l=1}^{k}\exp(a_l^{\mathsf T}x_i)}.

    The resulting prompted feature matrix has rows xi+pix_i+p_i. The basis-and-attention parameterization supports graphs with different numbers of nodes and uses O(kF)O(kF) learned prompt parameters rather than O(NF)O(NF) independently stored node prompts, where NN is the graph's node count. Setting k=1k=1 makes all nodes use the same prompt and reduces GPF-plus to GPF.

  3. Knowl 3 — Universal capability of GPF

    theoretical result

    Let ff be a fixed pre-trained GNN, let G=(A,X)G=(A,X) be an input graph, and let ψt\psi_t be any prompting function associated with any pre-training task tt. Suppose ψt(G)\psi_t(G) produces a candidate graph template whose adjacency and feature matrices lie in candidate spaces A\mathcal{A} and X\mathcal{X}. For every candidate prompted graph (A^,X^)∈A×X(\widehat A,\widehat X)\in\mathcal{A}\times\mathcal{X}, the paper establishes the existence of a shared GPF vector p^∈RF\widehat p\in\mathbb{R}^{F} such that

    f(A,X+p^)=f(A^,X^).f(A,X+\widehat p)=f(\widehat A,\widehat X).

    Thus, within the theorem's scope, optimizing GPF can reproduce the representation produced by any explicitly designed graph-prompting function, including prompts that alter node features, graph links, or isolated components. GPF therefore has the same theoretical representation-performance upper bound as an arbitrary task-specific prompting template, without requiring a manually specified template for each pre-training strategy. The paper states the result for arbitrary pre-trained GNNs, with the constructive analysis established for the linear message-passing architectures described separately.

  4. Knowl 4 — Universality for common message-passing architectures

    theoretical result

    The universality analysis is extended beyond the single-layer linear GIN used for the simplest argument. A broad linear message-passing layer can be written as

    H=SXW,H=S X W,

    where H∈RN×F′H\in\mathbb{R}^{N\times F'} is the node-representation matrix, S∈RN×NS\in\mathbb{R}^{N\times N} is the fixed diffusion matrix, X∈RN×FX\in\mathbb{R}^{N\times F} is the node-feature matrix, and W∈RF×F′W\in\mathbb{R}^{F\times F'} is a frozen linear projection. Under this form, a shared feature prompt produces an additive node-representation change proportional to the row sums of SS multiplied by pWpW. The paper shows that the same representation-equivalence arguments cover feature transformations, link transformations, and additions or removals of isolated components when graph representations use an additive readout.

    For a KK-layer linear GNN without nonlinear activations between layers, with layer-specific diffusion matrices S(r)S^{(r)} and projections W(r)W^{(r)}, the node representations reduce to

    H(K)=(∏r=1KS(r))X(∏r=1KW(r))=S′XW′.H^{(K)}=\left(\prod_{r=1}^{K}S^{(r)}\right)X\left(\prod_{r=1}^{K}W^{(r)}\right)=S'XW'.

    Consequently, the same equivalence applies after replacing SS and WW by the effective matrices S′S' and W′W'. This theoretical extension covers the linearized multi-layer message-passing architectures used in the paper's analysis, rather than relying only on the one-layer GIN example.

  5. Knowl 5 — Effectiveness guarantee relative to fine-tuning

    theoretical result

    The paper gives a strict theoretical comparison between GPF and full fine-tuning for a restricted regression setting. For graph GiG_i, let SiS_i be its fixed diffusion matrix, XiX_i its node-feature matrix, and let the frozen GNN representation be hi=Sum⁡(SiXiW)h_i=\operatorname{Sum}(S_iX_iW), where W∈RF×F′W\in\mathbb{R}^{F\times F'} is the pre-trained linear projection and Sum⁡\operatorname{Sum} sums node representations. Let θ∈RF′×1\theta\in\mathbb{R}^{F'\times1} be a linear downstream head and let the scalar target be yiy_i. The squared loss is

    ℓ=∑i=1m(hiθ−yi)2.\ell=\sum_{i=1}^{m}(h_i\theta-y_i)^2.

    Under the paper's non-degeneracy condition—using a column-full-rank graph coefficient matrix CC, a unique-feature matrix XSXX_{SX} for the dataset such that XSXv=1mX_{SX}v=\mathbf{1}_m has no solution, and a nonzero prompt/head contribution—the paper constructs target labels y1′,…,ym′y'_1,\ldots,y'_m for which yi=yi′y_i=y'_i and

    ℓGPF=min⁡p,θ∑i=1m(f(Ai,Xi+p)θ−yi)2  <ℓFT=min⁡W,θ∑i=1m(fW(Ai,Xi)θ−yi)2.\ell_{\mathrm{GPF}}=\min_{p,\theta}\sum_{i=1}^{m}\bigl(f(A_i,X_i+p)\theta-y_i\bigr)^2 \;< \ell_{\mathrm{FT}}=\min_{W,\theta}\sum_{i=1}^{m}\bigl(f_W(A_i,X_i)\theta-y_i\bigr)^2.

    Here ℓGPF\ell_{\mathrm{GPF}} optimizes only the shared feature prompt and head with the pre-trained GNN fixed, whereas ℓFT\ell_{\mathrm{FT}} also optimizes the GNN projection parameters WW. The result does not claim that GPF strictly beats fine-tuning for every possible dataset or target; it proves that fine-tuning is not a universal theoretical upper bound because there are non-degenerate downstream settings where GPF attains a lower optimum.

  6. Knowl 6 — Extension from graph classification to node-wise tasks

    model/method

    The paper extends graph prompting to node classification and link prediction by using a Subgraph GNN. For each node viv_i in an input graph, let GiG_i be the induced subgraph associated with that node. A frozen Subgraph GNN ff produces the node representation

    hi=f(Gi).h_i=f(G_i).

    Graph prompting replaces each induced subgraph by its prompted version:

    hi=f(gϕ(Gi)),h_i=f(g_{\phi}(G_i)),

    where gϕg_{\phi} is a learnable graph prompt, instantiated by GPF or GPF-plus. The resulting node representations are then supplied to a downstream node-classification or link-prediction head. Because node representations are obtained from graph representations of induced subgraphs, the feature-space prompting mechanism does not require a separate prompt design for node-wise tasks.

  7. Knowl 7 — Experimental protocol across pre-training strategies

    experimental setup

    The experiments use a five-layer Graph Isomorphism Network (GIN) as the pre-trained backbone. Models are pre-trained with five strategies: Deep Graph Infomax (Infomax), Edge Prediction (EdgePred), Attribute Masking (AttrMasking), Context Prediction (ContextPred), and Graph Contrastive Learning (GCL). The chemistry domain uses approximately 2 million unlabeled ZINC15 molecules and ChEMBL graph-property data; the biology domain uses 395K unlabeled protein ego-networks and 88K labeled protein ego-networks. Downstream evaluation covers eight molecular binary-classification benchmarks—BBBP, Tox21, ToxCast, SIDER, ClinTox, MUV, HIV, and BACE—and 40 protein-function prediction tasks represented by PPI.

    Fine-tuning updates the pre-trained GNN and the task head. GPF freezes the GNN and learns the shared vector pp plus the projection head, while GPF-plus freezes the GNN and learns its basis vectors, attention projections, and projection head. Each setting is run with five random seeds and reported as mean test ROC-AUC. The projection head is selected from one-, two-, or three-layer equal-width MLPs, and the GPF-plus basis count kk is selected from {5,10,20}\{5,10,20\}.

  8. Knowl 8 — Full-shot performance and parameter efficiency

    empirical result

    On the nine downstream benchmarks, GPF and GPF-plus generally outperform full fine-tuning while updating far fewer parameters. For the four non-edge pre-training strategies, the mean test ROC-AUC (%) across the nine tasks is:

    • Infomax: fine-tuning 72.9372.93, GPF 74.3674.36, GPF-plus 74.7974.79.
    • Attribute Masking: fine-tuning 73.9773.97, GPF 75.7675.76, GPF-plus 75.7675.76.
    • Context Prediction: fine-tuning 74.5374.53, GPF 75.8675.86, GPF-plus 76.1576.15.
    • Graph Contrastive Learning: fine-tuning 70.2770.27, GPF 70.2870.28, GPF-plus 71.3971.39.
    • Edge Prediction: fine-tuning 72.7272.72, GPF 74.5174.51, GPF-plus 74.5474.54.

    Across the 36 experiments in the first four strategies, GPF beats fine-tuning in 28 cases and GPF-plus beats it in 29 cases. Averaged across those strategies, GPF improves over fine-tuning by 1.14 percentage points and GPF-plus by 1.60 percentage points. The parameter counts excluding the task head are approximately 1.81.8M for chemistry fine-tuning versus 0.30.3K for GPF and 33–1212K for GPF-plus; for biology they are approximately 2.72.7M, 0.30.3K, and 33–1212K, respectively. Thus, GPF uses at most 0.02%0.02\% and GPF-plus at most 0.7%0.7\% of the fine-tuning parameters in these experiments.

  9. Knowl 9 — Comparison with specialized edge-prediction prompts

    empirical result

    The paper compares GPF and GPF-plus against methods designed specifically for Edge Prediction pre-training: GPPT, GPPT without its orthogonal-prompt constraint loss, and GraphPrompt. On the same nine downstream benchmarks, the mean test ROC-AUC (%) is 72.72 for fine-tuning, 61.80 for GPPT, 70.97 for GPPT without the constraint loss, 61.20 for GraphPrompt, 74.51 for GPF, and 74.54 for GPF-plus.

    The universal feature prompts therefore outperform the specialized methods even when those methods are used with the pre-training strategy they target. Relative to GPPT, GPPT without the constraint loss, and GraphPrompt, the paper reports average gains of approximately 12%, 3%, and 13%, respectively. GPF and GPF-plus are also the only prompt-based methods in this comparison that exceed full fine-tuning.

  10. Knowl 10 — Few-shot robustness and preservation of generalization

    empirical result

    When each downstream task is restricted to 50 labeled training examples, GPF improves mean ROC-AUC over fine-tuning by an average of 2.95 percentage points and GPF-plus by 3.42 percentage points across the five pre-training strategies. With 100 labeled examples, GPF or GPF-plus gives the best result in 42 of 45 reported comparisons: GPF is best in 14 and GPF-plus in 28, and both methods have higher average performance than fine-tuning for every pre-training strategy.

    Training-curve analysis on biology tasks pre-trained with Attribute Masking and Context Prediction shows that all methods continue improving on training ROC-AUC, but fine-tuning's test ROC-AUC fluctuates and declines after an initial increase. In contrast, GPF and GPF-plus maintain high test ROC-AUC that continues to improve during adaptation. These observations support the paper's claim that freezing the pre-trained GNN and adapting the input can reduce the loss of downstream generalization in low-data settings.

Coverage note — Detailed per-dataset tables, linear-probing and partial-layer/MLP comparisons, the auxiliary GCC experiments, and proof derivations were omitted because they support the central method and headline results rather than constituting additional load-bearing contributions.

References

  1. 1.Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, and Phillip Isola. Exploring visual prompts for adapting large-scale models. 2022.
  2. 2.Beatrice Bevilacqua, Fabrizio Frasca, Derek Lim, Balasubramaniam Srinivasan, Chen Cai, G. Balamurugan, Michael M. Bronstein, and Haggai Maron. Equivariant subgraph aggregation networks. ICLR, 2022.
  3. 3.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. NeurIPS, 2020.
  4. 4.Leonardo Cotta, Christopher Morris, and Bruno Ribeiro. Reconstruction for powerful graph representations. NeurIPS, 2021.
  5. 5.Ning Ding, Yujia Qin, Guang Yang, Fu Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, Jing Yi, Weilin Zhao, Xiaozhi Wang, Zhiyuan Liu, Haitao Zheng, Jianfei Chen, Yang Liu, Jie Tang, Juan Li, and Maosong Sun. Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models. ArXiv, abs/2203.06904, 2022.
  6. 6.Simon Shaolei Du, Wei Hu, Sham M. Kakade, J. Lee, and Qi Lei. Few-shot learning via learning the representation, provably. ICLR, 2020.
  7. 7.Fabrizio Frasca, Beatrice Bevilacqua, Michael Bronstein, and Haggai Maron. Understanding and extending subgraph gnns by rethinking their symmetries. NeurIPS, 2022.
  8. 8.Anna Gaulton, Louisa J. Bellis, A. Patrícia Bento, Jon Chambers, Mark Davies, Anne Hersey, Yvonne Light, Shaun McGlinchey, David Michalovich, Bissan Al-Lazikani, and John P. Overington. Chembl: a large-scale bioactivity database for drug discovery. Nucleic Acids Research, 40:D1100 – D1107, 2012.
  9. 9.Will Hamilton, Zhitao Ying, and Jure Leskovec. Inductive representation learning on large graphs. In NeurIPS, pages 1024–1034, 2017.
  10. 10.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll’ar, and Ross B. Girshick. Masked autoencoders are scalable vision learners. ArXiv, abs/2111.06377, 2021.
  11. 11.Weihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay S. Pande, and Jure Leskovec. Strategies for pre-training graph neural networks. ICLR, 2020a.
  12. 12.Ziniu Hu, Yuxiao Dong, Kuansan Wang, Kai-Wei Chang, and Yizhou Sun. Gpt-gnn: Generative pre-training of graph neural networks. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2020b.
  13. 13.Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. ECCV, 2022.
  14. 14.Wei Jin, Tyler Derr, Haochen Liu, Yiqi Wang, Suhang Wang, Zitao Liu, and Jiliang Tang. Self-supervised learning on graphs: Deep insights and new direction. ArXiv, abs/2006.10141, 2020.
  15. 15.Thomas Kipf and Max Welling. Variational graph auto-encoders. ArXiv, abs/1611.07308, 2016a.
  16. 16.Thomas N Kipf and Max Welling. Variational graph auto-encoders. In ArXiv, volume abs/1611.07308, 2016b.
  17. 17.Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations, 2017.
  18. 18.Johannes Klicpera, Stefan Weißenberger, and Stephan Günnemann. Diffusion improves graph learning. In Neural Information Processing Systems, 2019.
  19. 19.Boris Knyazev, Graham W. Taylor, and Mohamed R. Amer. Understanding attention and generalization in graph neural networks. In NeurIPS, 2019.
  20. 20.Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution. ICLR, 2022.
  21. 21.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. EMNLP, 2021a.
  22. 22.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021b.
  23. 23.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), abs/2101.00190, 2021.
  24. 24.Huihui Liu, Yiding Yang, and Xinchao Wang. Overcoming catastrophic forgetting in graph neural networks. In AAAI Conference on Artificial Intelligence, 2020.
  25. 25.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586, 2021a.
  26. 26.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys (CSUR), 2022a.
  27. 27.Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. ArXiv, abs/2110.07602, 2021b.
  28. 28.Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602, 2021c.
  29. 29.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. Gpt understands, too. ArXiv, abs/2103.10385, 2021d.
  30. 30.Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In ACL, 2022b.
  31. 31.Yixin Liu, Shirui Pan, Ming Jin, Chuan Zhou, Feng Xia, and Philip S. Yu. Graph self-supervised learning: A survey. IEEE Transactions on Knowledge and Data Engineering, 35:5879–5900, 2021e.
  32. 32.Zemin Liu, Xingtong Yu, Yuan Fang, and Xinming Zhang. Graphprompt: Unifying pre-training and downstream tasks for graph neural networks. Proceedings of the ACM Web Conference 2023, 2023.
  33. 33.Siqu Long, Feiqi Cao, Soyeon Caren Han, and Haiqing Yang. Vision-and-language pretrained models: A survey. IJCAI, 2022.
  34. 34.Yuanfu Lu, Xunqiang Jiang, Yuan Fang, and Chuan Shi. Learning to pre-train graph neural networks. In AAAI, 2021.
  35. 35.Andreas Mayr, Günter Klambauer, Thomas Unterthiner, Marvin N. Steijaert, Jörg Kurt Wegner, Hugo Ceulemans, Djork-Arné Clevert, and Sepp Hochreiter. Large-scale comparison of machine learning methods for drug target prediction on chembl† †electronic supplementary information (esi) available: Overview, data collection and clustering, methods, results, appendix. see doi: 10.1039/c8sc00148k. Chemical Science, 9:5441 – 5451, 2018.
  36. 36.Diego Mesquita, Amauri H. de Souza, and Samuel Kaski. Rethinking pooling in graph neural networks. ArXiv, abs/2010.11418, 2020.
  37. 37.Christopher Morris, Martin Ritzert, Matthias Fey, William L. Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe. Weisfeiler and leman go neural: Higher-order graph neural networks. AAAI, 2019.
  38. 38.Jiezhong Qiu, Qibin Chen, Yuxiao Dong, Jing Zhang, Hongxia Yang, Ming Ding, Kuansan Wang, and Jie Tang. Gcc: Graph contrastive coding for graph neural network pre-training. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2020a.
  39. 39.Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. Pre-trained models for natural language processing: A survey. Science China Technological Sciences, 63:1872 – 1897, 2020b.
  40. 40.Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying Wei, Wenbing Huang, and Junzhou Huang. Self-supervised graph transformer on large-scale molecular data. Advances in Neural Information Processing Systems, 33:12559–12571, 2020.
  41. 41.Timo Schick and Hinrich Schütze. Exploiting cloze-questions for few-shot text classification and natural language inference. In Conference of the European Chapter of the Association for Computational Linguistics, 2020a.
  42. 42.Timo Schick and Hinrich Schütze. It’s not just size that matters: Small language models are also few-shot learners. NAACL, 2020b.
  43. 43.T. Sterling and John J. Irwin. Zinc 15 – ligand discovery for everyone. Journal of Chemical Information and Modeling, 55:2324 – 2337, 2015.
  44. 44.Fan-Yun Sun, Jordan Hoffmann, Vikas Verma, and Jian Tang. Infograph: Unsupervised and semi-supervised graph-level representation learning via mutual information maximization. arXiv preprint arXiv:1908.01000, 2019.
  45. 45.Mingchen Sun, Kaixiong Zhou, Xingbo He, Ying Wang, and Xin Wang. Gppt: Graph pre-training and prompt tuning to generalize graph neural networks. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022.
  46. 46.Susheel Suresh, Pan Li, Cong Hao, and Jennifer Neville. Adversarial graph augmentation to improve graph contrastive learning. In Neural Information Processing Systems, 2021.
  47. 47.Junjiao Tian, Xiaoliang Dai, Chih-Yao Ma, Zecheng He, Yen-Cheng Liu, and Zsolt Kira. Trainable projected gradient method for robust fine-tuning. ArXiv, abs/2303.10720, 2023.
  48. 48.Nilesh Tripuraneni, Michael I. Jordan, and Chi Jin. On the theory of transfer learning: The importance of task diversity. NeurIPS, 2020.
  49. 49.Petar Velickovic, William Fedus, William L. Hamilton, Pietro Lio’, Yoshua Bengio, and R. Devon Hjelm. Deep graph infomax. ICLR, 2019a.
  50. 50.Petar Velickovic, William Fedus, William L. Hamilton, Pietro Lio’, Yoshua Bengio, and R. Devon Hjelm. Deep graph infomax. ICLR, 2019b.
  51. 51.Colin Wei, Sang Michael Xie, and Tengyu Ma. Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning. ArXiv, abs/2106.09226, 2021.
  52. 52.Junyang Wu, Xianhang Li, Chen Wei, Huiyu Wang, Alan Loddon Yuille, Yuyin Zhou, and Cihang Xie. Unleashing the power of visual prompting at the pixel level. ArXiv, abs/2212.10556, 2022.
  53. 53.Sen Wu, Hongyang Zhang, and Christopher Ré. Understanding and improving information transfer in multi-task learning. ICLR, 2020.
  54. 54.Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S. Pappu, Karl Leswing, and Vijay S. Pande. Moleculenet: A benchmark for molecular machine learning. arXiv: Learning, 2017.
  55. 55.Jun Xia, Lirong Wu, Jintao Chen, Bozhen Hu, and Stan Z. Li. Simgrace: A simple framework for graph contrastive learning without data augmentation. Proceedings of the ACM Web Conference 2022, 2022a.
  56. 56.Jun Xia, Yanqiao Zhu, Yuanqi Du, and Stan Z Li. A survey of pretraining on graphs: Taxonomy, methods, and applications. arXiv preprint arXiv:2202.07893, 2022b.
  57. 57.Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang, and Yanning Zhang. Class-aware visual prompt tuning for vision-language pre-trained model. ArXiv, abs/2208.08340, 2022.
  58. 58.Dongkuan Xu, Ian En-Hsu Yen, Jinxi Zhao, and Zhibin Xiao. Rethinking network pruning – under the pre-train and fine-tune paradigm. NAACL, 2021a.
  59. 59.Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks? ICLR, 2019.
  60. 60.Minghao Xu, Hang Wang, Bingbing Ni, Hongyu Guo, and Jian Tang. Self-supervised graph-level representation learning with local and global structure. In International Conference on Machine Learning, 2021b.
  61. 61.Pinar Yanardag and S. V. N. Vishwanathan. Deep graph kernels. Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2015.
  62. 62.Gilad Yehudai, Ethan Fetaya, Eli A. Meirom, Gal Chechik, and Haggai Maron. From local structures to size generalization in graph neural networks. In ICML, 2021.
  63. 63.Yuning You, Tianlong Chen, Yongduo Sui, Ting Chen, Zhangyang Wang, and Yang Shen. Graph contrastive learning with augmentations. NeurIPS, 2020.
  64. 64.Yuning You, Tianlong Chen, Yang Shen, and Zhangyang Wang. Graph contrastive learning automated. In International Conference on Machine Learning, pages 12121–12132. PMLR, 2021.
  65. 65.Yuning You, Tianlong Chen, Zhangyang Wang, and Yang Shen. Bringing your own view: Graph contrastive learning without prefabricated data augmentations. Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, 2022.
  66. 66.Chuxu Zhang, Kaize Ding, Jundong Li, Xiangliang Zhang, Yanfang Ye, N. Chawla, and Huan Liu. Few-shot learning on graphs. In International Joint Conference on Artificial Intelligence, 2022.
  67. 67.Muhan Zhang and Pan Li. Nested graph neural networks. NeurIPS, 2021.
  68. 68.Richard Zhang, Phillip Isola, and Alexei A. Efros. Colorful image colorization. In ECCV, 2016.
  69. 69.Z. Zhang, Q. Liu, H. Wang, C. Lu, and C. K. Lee. Motif-based graph self-supervised learning for molecular property prediction. 2021a.
  70. 70.Zaixin Zhang, Qi Liu, Hao Wang, Chengqiang Lu, and Chee-Kong Lee. Motif-based graph self-supervised learning for molecular property prediction. In Neural Information Processing Systems, 2021b.
  71. 71.Lingxiao Zhao, Wei Jin, Leman Akoglu, and Neil Shah. From stars to subgraphs: Uplifting any gnn with local structure awareness. ICLR, 2022.
  72. 72.Fan Zhou and Chengtai Cao. Overcoming catastrophic forgetting in graph neural networks with experience replay. In AAAI Conference on Artificial Intelligence, 2021.
  73. 73.Marinka Zitnik, Rok Sosi, and Jure Leskovec. Prioritizing network communities. Nature Communications, 9, 2018.

Citation

MLA
Fang, T., et al. “Universal Prompt Tuning for Graph Neural Networks”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 52464–89, https://proceedings.neurips.cc/paper_files/paper/2023/file/a4a1ee071ce0fe63b83bce507c9dc4d7-Paper-Conference.pdf.
APA
Fang, T., Zhang, Y., YANG, Y., Wang, C., & Chen, L. (2023). Universal Prompt Tuning for Graph Neural Networks. Advances in Neural Information Processing Systems, 36, 52464–52489. https://proceedings.neurips.cc/paper_files/paper/2023/file/a4a1ee071ce0fe63b83bce507c9dc4d7-Paper-Conference.pdf
Chicago
Fang, T., Y. Zhang, Y. YANG, C. Wang, and L. Chen. 2023. “Universal Prompt Tuning for Graph Neural Networks”. Advances in Neural Information Processing Systems 36: 52464–89. https://proceedings.neurips.cc/paper_files/paper/2023/file/a4a1ee071ce0fe63b83bce507c9dc4d7-Paper-Conference.pdf.
Harvard
Fang, T. et al. (2023) “Universal Prompt Tuning for Graph Neural Networks”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 52464–52489. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/a4a1ee071ce0fe63b83bce507c9dc4d7-Paper-Conference.pdf.
Vancouver
1. Fang T, Zhang Y, YANG Y, Wang C, Chen L (2023) Universal Prompt Tuning for Graph Neural Networks. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 52464–52489

BibTeX

@inproceedings{fang2023universal,
  title = {Universal Prompt Tuning for Graph Neural Networks},
  author = {Fang, Taoran and Zhang, Yunchao and YANG, YANG and Wang, Chunping and Chen, Lei},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {52464-52489},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/a4a1ee071ce0fe63b83bce507c9dc4d7-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors