Position: Graph Foundation Models Are Already Here

Haitao MaoZhikai ChenWenzhuo TangJianan ZhaoYao MaTong ZhaoNeil ShahMikhail GalkinJiliang Tang

article2024ICML109 citations

Presents a unifying graph vocabulary perspective that explains how current primitive models achieve cross-dataset transferability and provides concrete principles to guide the design of general-purpose Graph Foundation Models.

Listen

Graph learning traditionally relies on training separate Graph Neural Networks from scratch for individual tasks and datasets. While foundational AI models have transformed fields like computer vision and natural language processing through broad pre-training, graph-based applications continue to suffer from redundant development costs and poor cross-dataset generalization. Developing universal Graph Foundation Models is challenging because graphs represent abstract, highly heterogeneous, non-Euclidean structures without native standard representations. Addressing this challenge is critical for scaling machine learning applications across drug discovery, molecular chemistry, knowledge management, and fraud detection.

The article systematically evaluates the current landscape of Graph Foundation Models and investigates how to achieve positive cross-dataset transferability. It proposes a foundational framework centered on constructing a transferable "graph vocabulary" based on network analysis, mathematical expressiveness, and structural stability.

To conduct this assessment, the authors analyzed foundational principles and empirical results across multiple model paradigms, benchmark suites, and tasks including node classification, link prediction, and graph classification. They examined leading specialized architectures (such as ULTRA and DiG), cross-task unified models (such as OneForAll), and methods bridging graph structures with Large Language Models.

The analysis establishes several key findings. First, while general-purpose models spanning all domains do not yet exist, specialized task-specific and domain-specific models have achieved practical zero-shot generalization to unseen graphs. Second, successful knowledge transfer depends on establishing an invariant vocabulary that avoids compressing non-isomorphic structures into identical representations. Third, distinct network properties require specialized modeling choices; for instance, modeling homophilic graphs ("birds of a feather") and heterophilic graphs ("opposites attract") requires separate aggregation mechanisms. Fourth, Large Language Models excel as feature encoders for text-rich graphs but demonstrate fundamental limitations when performing complex graph structural reasoning on their own. Finally, graph models can exhibit neural scaling laws when architectures incorporate suitable inductive structural priors, such as geometric invariants or discrete tokenization.

These findings indicate that organizations can move beyond single-purpose graph models toward reusable, domain-tailored foundation models. Adopting this paradigm can reduce training resource consumption and decrease dependence on expensive domain-expert labeling. However, moving toward a single universal graph model carries operational risks. Forcing completely unaligned domains into a shared architecture risks negative transfer, where model performance degrades across non-isomorphic or conflicting relational patterns.

Organizations developing graph-based AI should prioritize task- and domain-specific foundation models, particularly in structured areas like molecular chemistry and knowledge graphs. Architecture choices must pair expressive graph neural network backbones or graph transformers with separate encoding pipelines for structural and feature proximity. Rather than relying entirely on text-based language models for topological reasoning, engineering teams should leverage language models primarily to align disparate text features while employing dedicated graph tokenizers for structural modeling. Further research and standardized benchmarking are required to determine whether a truly universal structural representation space exists across divergent graph domains.

Confidence in specialized domain-specific graph models is strong, supported by robust empirical evidence in knowledge graphs and chemistry. However, substantial uncertainties remain regarding cross-domain transferability, severe data scarcity relative to text and image datasets, and noisy graph construction standards (such as mislabeling rates exceeding 15% in standard benchmarks). Decision-makers should approach claims of universal cross-domain graph capabilities with caution until standardized, robust scaling laws are further demonstrated.

  • Paper: Cooperative Graph Neural Networks, Ben Finkelshtein et al. (2024). It develops adaptive, entity-specific communication policies that extend the source’s case for specialized graph models to settings with heterophily and long-range information flow.
Cover for Position: Graph Foundation Models Are Already Here

Table of Contents

  • 1. Introduction
  • 2. Existing GFMs and Key Designs
  • 2.1. Existing GFM Categories
  • 2.2. The Key to A Successful GFM Design.
  • 3. Graph Transferability Principles with Actionable Steps
  • 3.1. An overview on Graph transferability principles
  • 3.2. Transferability Principles in Node Classification
  • 3.3. Transferability Principles in Link Prediction
  • 3.4. Transferability Principles in Graph Classification
  • 3.5. Transferability Principles across Tasks
  • 4. Neural Scaling Law on GFMs
  • 4.1. When Neural Scaling Law Happens
  • 4.2. Data Scaling
  • 4.3. Model Scaling
  • 4.4. Leveraging Large-scale LMs for Graphs
  • 5. Insights & Open Questions
  • 5.1. Potential Redundancy on Pretext Task and Architecture Design
  • 5.2. The Feasibility of GFMs
  • 5.3. Broader Usage of GFM
  • 6. Conclusion
  • Acknowledgement
  • Impact Statement
  • References
  • A. A Collection of Datasets to Support Pre-training
  • B. Existing GFMs
  • C. Practical Recipes for GFM Applications
  • C.1. Tackling the Feature Heterogeneity Issue
  • C.2. Pretext Task Design
  • C.3. Efficiency Issues in Subgraph-based Methods.
  • D. Additional Principles
  • D.1. Principles on deeper GNN design
  • D.2. Additional Description on the Relational Graph Vocabulary of ULTRA
  • E. Discussions & Open questions
  • E.1. More discussions on LLMs and Graphs
  • E.2. Whether There exists a General Graph Vocabulary?
  • E.3. Deeper GNNs as the Backbone of GFMs
  • E.4. Is Invariance a Necessity for Building GFM?
  • E.5. Comparison with Past Relevant Literature

Knowls

  1. Knowl 1 — A graph vocabulary is the basis for transferable graph foundation models

    definition

    A graph vocabulary is a set of basic transferable units—or a representation mechanism—that maps graphs from different datasets into a shared space while preserving the invariances relevant to downstream tasks. It need not be a discrete tokenizer or embedding layer. The paper argues that a useful vocabulary must be both discriminative, so structurally distinct cases are not collapsed into the same representation, and inclusive, so previously unseen structures or relation types can connect to learned representations. This shared representation is treated as a prerequisite for positive transfer and, consequently, for scaling training data.

    Knowledge-graph completion illustrates the design criteria. ULTRA combines expressive, query-conditioned message passing with a graph of relations that represents interactions independently of graph-specific relation names. The first component helps distinguish structurally different entity pairs; the second uses double permutation equivariance—equivariance to permutations of both entities and relation types—to relate unseen relation types to the learned relational vocabulary.

  2. Knowl 2 — Network analysis, expressiveness, and stability provide complementary transfer principles

    model/method

    The paper organizes graph transferability principles into three complementary perspectives. Network analysis identifies recurring patterns and domain-level regularities, such as motifs, triadic closure, and homophily; these offer practical design guidance but may depend on expert knowledge and do not by themselves provide formal guarantees. Expressiveness characterizes which structural patterns a model can distinguish. For a most-expressive structural representation, two structures should receive the same representation if and only if they are equivalent under the relevant node permutation. Stability concerns how much predictions change under small graph perturbations: similar structures should have similar representations or predictions, rather than merely separating isomorphic from non-isomorphic cases. Together, the principles guide vocabulary construction toward representations that capture useful patterns, distinguish relevant structures, and remain robust to minor changes.

  3. Knowl 3 — Existing graph foundation models differ by the scope of their transfer

    definition

    The paper distinguishes three levels of graph foundation model (GFM) transferability. A task-specific GFM transfers across datasets for a particular task; ULTRA is an example for knowledge-graph completion and reasoning. A domain-specific GFM transfers across tasks within a domain; DiG is an example for chemical tasks. A primitive GFM generalizes across a limited collection of datasets and tasks; OFA is an example trained across citation, molecular, and knowledge graphs using unified task formulations. These categories describe different scopes of transfer, not a single model that works across all graph tasks and domains; the paper notes that existing GFMs do not yet provide that universal capability.

  4. Knowl 4 — Node-classification transfer depends on the graph’s homophily regime and perturbation stability

    model/method

    For node classification, homophily—the tendency of linked nodes to have similar features—supports transfer among homophilous graphs and motivates many successful GNN designs. It is not universal: heterophilous graphs connect dissimilar nodes, and the variety of their interaction patterns makes transfer less dependable, except where consistent “good heterophily” patterns can be identified. The paper therefore recommends against relying on one fixed GNN aggregation scheme for both regimes. Candidate designs include adaptive GNNs with different filters for homophilic and heterophilic patterns, or graph transformers without a fixed aggregation process.

    Stability provides an additional constraint. The reviewed theory associates greater spectral smoothness of a graph filter with robustness to edge perturbations, and a smaller maximum frequency response with robustness to feature perturbations; these properties are reported to improve transferability. The paper suggests adapting spectral regularization for GFM design.

  5. Knowl 5 — Link prediction requires expressive pair representations and separate treatment of feature and structural proximity

    model/method

    The paper identifies three useful link-prediction principles: local structural proximity, associated with triadic closure; global structural proximity, associated with the greater likelihood of connection when more short paths link two nodes; and feature proximity, associated with homophily. These signals are not interchangeable: pairs with high feature proximity may have low local structural proximity, and vice versa. The paper therefore recommends encoding pairwise structural proximity separately from feature proximity.

    A standard GNN that produces only single-node representations can miss distinctions between candidate links. In the paper’s illustrative featureless graph, two structurally symmetric nodes receive identical representations, so the GNN gives identical predictions to links involving them even though the candidate pairs have different structural proximity. An expressive link representation should distinguish non-equivalent node pairs while remaining invariant to permutations that preserve the pair’s structure. Node-labeling approaches can provide this representation when they distinguish the source and target nodes from other nodes and are permutation equivariant; double-radius node labeling (DRNL) and zero-one labeling are examples.

  6. Knowl 6 — Graph motifs are a candidate shared vocabulary for graph classification

    theoretical result

    For graph classification, the paper proposes recurring small subgraphs, or network motifs, as potential vocabulary units because they can function as interpretable building blocks and may recur across domains. Graph kernels use motif counts and other predefined structural features for classification. The paper cites evidence of positive transfer using shared motif sets across neuronal connectivity networks, food webs, and electronic circuits, and conjectures that transferable motifs could therefore support a graph-classification vocabulary.

    The paper further conjectures that more expressive GNNs, which can detect a wider range of substructures, may be better able to identify motifs shared across datasets and thus transfer more effectively. Stability also matters: the reviewed stable positional encoding method uses a weighted sum of eigenvectors and is reported to perform well on out-of-distribution molecular prediction, suggesting that stable encodings may be useful in GFMs.

  7. Knowl 7 — Task transfer can use unified formulations, but unified supervised co-training is not yet guaranteed to help

    model/method

    A unified task formulation can make datasets for different graph tasks usable by one model and can enlarge the pool of training examples. Examples include converting node classification into link prediction between a target node and label nodes; converting node classification into ego-graph classification and link prediction into binary classification on the target pair’s induced subgraph; and adding a virtual prompt node connected to the nodes relevant to node-, link-, or graph-level prediction. The paper notes that directly using link prediction as a pretext task can harm node classification, whereas reformulating node classification as link prediction has produced positive transfer in reported work.

    Unifying tasks makes co-training possible but does not establish that co-training avoids negative transfer. Nor is a unified supervised formulation necessarily required: a GFM could be self-supervised-pretrained and then fine-tuned for downstream tasks. The paper points to shared principles—including feature homophily, global structural proximity, and motifs—as possible bases for transfer across tasks, while emphasizing that this direction remains under-studied.

  8. Knowl 8 — Graph data scaling requires transferable data and reliable graph construction

    empirical result

    The paper treats graph data scaling as conditional rather than automatic: adding pretraining data is expected to help when that data shares relevant properties with downstream data and follows transferable graph principles. It reviews reported data-scaling behavior for supervised and self-supervised GNNs on molecular property prediction and for node classification on text-attributed graphs. It also reports guidance for selecting pretraining data using graphon signal analysis and network entropy.

    The paper cautions that graph construction can introduce uncertainty because edges may depend on expert decisions, and labels can be erroneous or encode conflicting principles even when graph structure is identical. It identifies expert effort in defining relationships and intellectual-property restrictions as barriers to obtaining large graph corpora. Synthetic graphs are proposed as a potential response: traditional generators can reproduce selected statistical properties, while deep generative models can provide richer graph-distribution samples. The benefit of high-quality synthetic graph pretraining is presented as a prospect, not as a result established by this paper.

  9. Knowl 9 — Model scaling depends on architecture and vocabulary, not parameter count alone

    empirical result

    The paper argues that increasing GFM size will not necessarily improve performance unless the architecture captures transferable graph structure. It reviews a counterexample in which a larger graph attention network underperforms smaller versions on graph-regression tasks, while geometric GNNs with an appropriate geometric prior show scaling behavior for atomic-potential prediction. Graph transformers that encode geometric priors through a GNN encoder or positional encoding are also reported to scale positively on molecular tasks.

    A separate route serializes graphs as lossless Eulerian-path token sequences and trains a vanilla transformer with next-token prediction. The reviewed work reports model-scaling behavior and promising fine-tuned results for protein association and molecular property prediction. Whether this transformer approach scales effectively on other graph tasks remains unclear.

  10. Knowl 10 — Large language models can align graph features, but their role as structural foundation models remains uncertain

    limitation

    The paper identifies two main uses of large language models (LLMs) for graphs. As feature encoders, LLMs can map heterogeneous node attributes into a shared textual embedding space; the reviewed OFA approach also uses multimodal models to produce text descriptions for attributes such as molecules. This alignment can make different graph domains usable by a common model and can support satisfactory performance with a vanilla GCN. However, LLM feature encoding has not shown a scaling law in the cited work, and its performance can depend strongly on prompts.

    As predictors, LLMs can handle graph tasks presented in language, but simply flattening graph structure into prompts has not supplied the structural information needed to match well-trained GNNs. Reviewed approaches add GNN or graph-transformer structure encoders, non-parametric aggregation, or structure-preserving prompts; graph question-answering systems also combine LLMs with graph tokenizers or retrieval-augmented generation. The paper leaves open whether LLMs can reliably understand essential graph structure and serve as the core of a GFM, rather than primarily as textual feature encoders. Efficiency and task-specific fine-tuning are additional concerns.

Coverage note — The paper’s deeper-GNN failure modes, detailed subgraph-extraction efficiency discussion, dataset catalog, and broader application vignettes are omitted because they are secondary technical or illustrative material rather than the central transferability and scaling framework.

References

  1. 1.Abramson, J., Adler, J., Dunger, J., Evans, R., Green, T., Pritzel, A., Ronneberger, O., Willmore, L., Ballard, A. J., Bambrick, J., et al. Accurate structure prediction of biomolecular interactions with alphafold 3. Nature, pp. 1–3, 2024.
  2. 2.AbuOda, G., De Francisci Morales, G., and Aboulnaga, A. Link prediction via higher-order motif features. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2019, Wurzburg, Germany, September 16–20, 2019, Proceedings, Part I, pp. 412–429. Springer, 2020.
  3. 3.Adamic, L. A. and Adar, E. Friends and neighbors on the web. Social networks, 25(3):211–230, 2003.
  4. 4.Airoldi, E. M., Blei, D., Fienberg, S., and Xing, E. Mixed membership stochastic blockmodels. Advances in neural information processing systems, 21, 2008.
  5. 5.Albert, R. and Barabasi, A.-L. Statistical mechanics of complex networks. Reviews of modern physics, 74(1):47, 2002.
  6. 6.Bai, Y., Geng, X., Mangalam, K., Bar, A., Yuille, A., Darrell, T., Malik, J., and Efros, A. A. Sequential modeling enables scalable learning for large vision models. arXiv preprint arXiv:2312.00785, 2023.
  7. 7.Barcelo, P., Galkin, M., Morris, C., and Orth, M. R. Weisfeiler and leman go relational. In Learning on Graphs Conference, pp. 46–1. PMLR, 2022.
  8. 8.Barcelo, P., Kostylev, E. V., Monet, M., P´erez, J., Reutter, J., and Silva, J. P. The logical expressiveness of graph neural networks. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=r1lZ7AEKvB.
  9. 9.Batatia, I., Benner, P., Chiang, Y., Elena, A. M., Kovacs, D. P., Riebesell, J., Advincula, X. R., Asta, M., Baldwin, W. J., Bernstein, N., Bhowmik, A., Blau, S. M., Carare, V., Darby, J. P., De, S., Pia, F. D., Deringer, V. L., Elijosius, R., El-Machachi, Z., Fako, E., Ferrari, A. C., Genreith-Schriever, A., George, J., Goodall, R. E. A., Grey, C. P., Han, S., Handley, W., Heenen, H. H., Hermansson, K., Holm, C., Jaafar, J., Hofmann, S., Jakob, K. S., Jung, H., Kapil, V., Kaplan, A. D., Karimitari, N., Kroupa, N., Kullgren, J., Kuner, M. C., Kuryla, D., Liepuoniute, G., Margraf, J. T., Magdau, I.-B., Michaelides, A., Moore, J. H., Naik, A. A., Niblett, S. P., Norwood, S. W., O’Neill, N., Ortner, C., Persson, K. A., Reuter, K., Rosen, A. S., Schaaf, L. L., Schran, C., Sivonxay, E., Stenczel, T. K., Svahn, V., Sutton, C., van der Oord, C., Varga-Umbrich, E., Vegge, T., Vondrak, M., Wang, Y., Witt, W. C., Zills, F., and Csanyi, G. A foundation model for atomistic materials chemistry, 2023.
  10. 10.Battiston, F., Cencetti, G., Iacopini, I., Latora, V., Lucas, M., Patania, A., Young, J.-G., and Petri, G. Networks beyond pairwise interactions: Structure and dynamics. Physics Reports, 874:1–92, 2020.
  11. 11.Beaini, D., Huang, S., Cunha, J. A., Moisescu-Pareja, G., Dymov, O., Maddrell-Mander, S., McLean, C., Wenkel, F., Muller, L., Mohamud, J. H., et al. Towards foundational models for molecular learning on large-scale multi-task datasets. arXiv preprint arXiv:2310.04292, 2023.
  12. 12.BehnamGhader, P., Adlakha, V., Mosbach, M., Bahdanau, D., Chapados, N., and Reddy, S. Llm2vec: Large language models are secretly powerful text encoders. arXiv preprint arXiv:2404.05961, 2024.
  13. 13.Benson, A. R., Gleich, D. F., and Leskovec, J. Higher-order organization of complex networks. Science, 353(6295):163–166, 2016.
  14. 14.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  15. 15.Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A. Video generation models as world simulators. 2024.
  16. 16.Brugere, I., Gallagher, B., and Berger-Wolf, T. Y. Network structure inference, a survey: Motivations, methods, and applications. ACM Computing Surveys (CSUR), 51(2):1–39, 2018.
  17. 17.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  18. 18.Cai, C. and Wang, Y. A note on over-smoothing for graph neural networks. arXiv preprint arXiv:2006.13318, 2020.
  19. 19.Cao, Y., Xu, J., Yang, C., Wang, J., Zhang, Y., Wang, C., Chen, L., and Yang, Y. When to pre-train graph neural networks? an answer from data generation perspective! arXiv preprint arXiv:2303.16458, 2023.
  20. 20.Chai, Z., Zhang, T., Wu, L., Han, K., Hu, X., Huang, X., and Yang, Y. Graphllm: Boosting graph reasoning ability of large language model. arXiv preprint arXiv:2310.05845, 2023.
  21. 21.Chamberlain, B. P., Shirobokov, S., Rossi, E., Frasca, F., Markovich, T., Hammerla, N., Bronstein, M. M., and Hansmire, M. Graph neural networks for link prediction with subgraph sketching. arXiv preprint arXiv:2209.15486, 2022.
  22. 22.Chawla, N. V. and Karakoulas, G. Learning from labeled and unlabeled data: An empirical study across techniques and domains. Journal of Artificial Intelligence Research, 23:331–366, 2005.
  23. 23.Chen, D., Zhu, Y., Zhang, J., Du, Y., Li, Z., Liu, Q., Wu, S., and Wang, L. Uncovering neural scaling laws in molecular representation learning. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023a. URL https://openreview.net/forum?id=Ys8RmfF9w1.
  24. 24.Chen, N., Li, Y., Tang, J., and Li, J. Graphwiz: An instruction-following language model for graph problems. arXiv preprint arXiv:2402.16029, 2024a.
  25. 25.Chen, R., Zhao, T., Jaiswal, A., Shah, N., and Wang, Z. Llaga: Large language and graph assistant. arXiv preprint arXiv:2402.08170, 2024b.
  26. 26.Chen, Z., Liu, J., Wang, X., and Yin, W. On representing linear programs by graph neural networks. In The Eleventh International Conference on Learning Representations, 2022.
  27. 27.Chen, Z., Mao, H., Li, H., Jin, W., Wen, H., Wei, X., Wang, S., Yin, D., Fan, W., Liu, H., and Tang, J. Exploring the potential of large language models (llms) in learning on graphs. ArXiv, abs/2307.03393, 2023b.
  28. 28.Chien, E., Peng, J., Li, P., and Milenkovic, O. Adaptive universal generalized pagerank graph neural network. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=n6jl7fLxrP.
  29. 29.Dess´ı, D., Osborne, F., Recupero, D. R., Buscaldi, D., and Motta, E. Cs-kg: A large-scale knowledge graph of research entities and claims in computer science. In International Workshop on the Semantic Web, 2022. URL https://api.semanticscholar.org/CorpusID:253021556.
  30. 30.Di Giovanni, F., Rusch, T. K., Bronstein, M. M., Deac, A., Lackenby, M., Mishra, S., and Velickoviˇc, P. How does over-squashing affect the power of gnns? arXiv preprint arXiv:2306.03589, 2023.
  31. 31.Dong, K., Mao, H., Guo, Z., and Chawla, N. V. Universal link predictor by in-context learning. arXiv preprint arXiv:2402.07738, 2024.
  32. 32.Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., and Sui, Z. A survey for in-context learning. arXiv preprint arXiv:2301.00234, 2022.
  33. 33.Dong, Y., Johnson, R. A., Xu, J., and Chawla, N. V. Structural diversity and homophily: A study across more than one hundred big networks. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 807–816, 2017.
  34. 34.Dziri, N., Lu, X., Sclar, M., Li, X. L., Jian, L., Lin, B. Y., West, P., Bhagavatula, C., Bras, R. L., Hwang, J. D., et al. Faith and fate: Limits of transformers on compositionality. arXiv preprint arXiv:2305.18654, 2023.
  35. 35.Edmonds, J. and Johnson, E. L. Matching, euler tours and the chinese postman. Mathematical Programming, 5:88–124, 1973. URL https://api.semanticscholar.org/CorpusID:15249924.
  36. 36.Fatemi, B., Halcrow, J., and Perozzi, B. Talk like a graph: Encoding graphs for large language models. arXiv preprint arXiv:2310.04560, 2023.
  37. 37.Fey, M. and Lenssen, J. E. Fast graph representation learning with pytorch geometric. arXiv preprint arXiv:1903.02428, 2019.
  38. 38.Fey, M., Lenssen, J. E., Weichert, F., and Leskovec, J. Gnnautoscale: Scalable and expressive graph neural networks via historical embeddings. In International conference on machine learning, pp. 3294–3304. PMLR, 2021.
  39. 39.Freitas, S., Duggal, R., and Chau, D. H. Malnet: A large-scale image database of malicious software. arXiv preprint arXiv:2102.01072, 2021.
  40. 40.Galkin, M., Yuan, X., Mostafa, H., Tang, J., and Zhu, Z. Towards foundation models for knowledge graph reasoning. arXiv preprint arXiv:2310.04562, 2023.
  41. 41.Galkin, M., Zhou, J., Ribeiro, B., Tang, J., and Zhu, Z. Zero-shot logical query reasoning on any knowledge graph. arXiv preprint arXiv:2404.07198, 2024.
  42. 42.Gao, J., Zhou, Y., and Ribeiro, B. Double permutation equivariance for knowledge graph completion. arXiv preprint arXiv:2302.01313, 2023a.
  43. 43.Gao, J., Zhou, Y., Zhou, J., and Ribeiro, B. Double equivariance for inductive link prediction for both new nodes and new relation types. In NeurIPS 2023 Workshop: New Frontiers in Graph Learning, 2023b.
  44. 44.Granovetter, M. S. The strength of weak ties. American journal of sociology, 78(6):1360–1380, 1973.
  45. 45.Gruver, N., Finzi, M., Qiu, S., and Wilson, A. G. Large language models are zero-shot time series forecasters. Advances in Neural Information Processing Systems, 36, 2024a.
  46. 46.Gruver, N., Sriram, A., Madotto, A., Wilson, A. G., Zitnick, C. L., and Ulissi, Z. Fine-tuned language models generate stable inorganic materials as text. arXiv preprint arXiv:2402.04379, 2024b.
  47. 47.Gupta, S., Manchanda, S., Ranu, S., and Bedathur, S. J. Grafenne: learning on graphs with heterogeneous and dynamic feature sets. In International Conference on Machine Learning, pp. 12165–12181. PMLR, 2023.
  48. 48.Han, X., Zhang, Z., Ding, N., Gu, Y., Liu, X., Huo, Y., Qiu, J., Yao, Y., Zhang, A., Zhang, L., et al. Pre-trained models: Past, present and future. AI Open, 2:225–250, 2021.
  49. 49.Hassani, K. and Khasahmadi, A. H. Contrastive multi-view representation learning on graphs. In Proceedings of International Conference on Machine Learning, pp. 3451–3461. 2020.
  50. 50.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  51. 51.He, X., Bresson, X., Laurent, T., Perold, A., LeCun, Y., and Hooi, B. Harnessing explanations: Llm-to-lm interpreter for enhanced text-attributed graph representation learning, 2023.
  52. 52.He, X., Tian, Y., Sun, Y., Chawla, N. V., Laurent, T., LeCun, Y., Bresson, X., and Hooi, B. G-retriever: Retrieval-augmented generation for textual graph understanding and question answering. arXiv preprint arXiv:2402.07630, 2024.
  53. 53.Hibshman, J. I., Gonzalez, D., Sikdar, S., and Weninger, T. Joint subgraph-to-subgraph transitions: Generalizing triadic closure for powerful and interpretable graph modeling. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining, pp. 815–823, 2021.
  54. 54.Hocevar, T. and Demˇsar, J. A combinatorial approach to graphlet counting. Bioinformatics, 30(4):559–565, 2014.
  55. 55.Hou, Z., Liu, X., Cen, Y., Dong, Y., Yang, H., Wang, C., and Tang, J. Graphmae: Self-supervised masked graph autoencoders. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 594–604, 2022.
  56. 56.Hu, W., Liu, B., Gomes, J., Zitnik, M., Liang, P., Pande, V., and Leskovec, J. Strategies for pre-training graph neural networks. arXiv preprint arXiv:1905.12265, 2019.
  57. 57.Hu, W., Fey, M., Zitnik, M., Dong, Y., Ren, H., Liu, B., Catasta, M., and Leskovec, J. Open graph benchmark: Datasets for machine learning on graphs. Advances in neural information processing systems, 33:22118–22133, 2020.
  58. 58.Huang, H., Tang, J., Liu, L., Luo, J., and Fu, X. Triadic closure pattern analysis and prediction in social networks. IEEE Transactions on Knowledge and Data Engineering, 27(12):3374–3389, 2015.
  59. 59.Huang, Q., Ren, H., Chen, P., Krzmanc, G., Zeng, D., Liang, P., and Leskovec, J. Prodigy: Enabling in-context learning over graphs. arXiv preprint arXiv:2305.12600, 2023a.
  60. 60.Huang, S., Poursafaei, F., Danovitch, J., Fey, M., Hu, W., Rossi, E., Leskovec, J., Bronstein, M., Rabusseau, G., and Rabbany, R. Temporal graph benchmark for machine learning on temporal graphs. arXiv preprint arXiv:2307.01026, 2023b.
  61. 61.Huang, X., Orth, M. R., Ceylan, ˙I. ˙I., and Barcelo, P. A theory of link prediction via relational weisfeiler-leman. arXiv preprint arXiv:2302.02209, 2023c.
  62. 62.Huang, Y., Lu, W., Robinson, J., Yang, Y., Zhang, M., Jegelka, S., and Li, P. On the stability of expressive positional encodings for graph neural networks. arXiv preprint arXiv:2310.02579, 2023d.
  63. 63.Ibarz, B., Kurin, V., Papamakarios, G., Nikiforou, K., Bennani, M., Csordas, R., Dudzik, A. J., Boˇsnjak, M., Vitvitskyi, A., Rubanova, Y., et al. A generalist neural algorithmic learner. In Learning on graphs conference, pp. 2–1. PMLR, 2022.
  64. 64.Jeh, G. and Widom, J. Simrank: a measure of structural-context similarity. In Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 538–543, 2002.
  65. 65.Jin, B., Liu, G., Han, C., Jiang, M., Ji, H., and Han, J. Large language models on graphs: A comprehensive survey. arXiv preprint arXiv:2312.02783, 2023a.
  66. 66.Jin, B., Zhang, W., Zhang, Y., Meng, Y., Zhang, X., Zhu, Q., and Han, J. Patton: Language model pretraining on text-rich networks. arXiv preprint arXiv:2305.12268, 2023b.
  67. 67.Jin, W., Barzilay, R., and Jaakkola, T. Hierarchical generation of molecular graphs using structural motifs. In International conference on machine learning, pp. 4839–4848. PMLR, 2020a.
  68. 68.Jin, W., Derr, T., Liu, H., Wang, Y., Wang, S., Liu, Z., and Tang, J. Self-supervised learning on graphs: Deep insights and new direction. arXiv preprint arXiv:2006.10141, 2020b.
  69. 69.Jing, Y., Yuan, C., Ju, L., Yang, Y., Wang, X., and Tao, D. Deep graph reprogramming. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 24345–24354, 2023.
  70. 70.Jo, J., Lee, S., and Hwang, S. J. Score-based generative modeling of graphs via the system of stochastic differential equations. In International Conference on Machine Learning, pp. 10362–10383. PMLR, 2022.
  71. 71.Ju, M., Zhao, T., Wen, Q., Yu, W., Shah, N., Ye, Y., and Zhang, C. Multi-task self-supervised graph neural networks enable stronger task generalization. 2023.
  72. 72.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  73. 73.Katz, L. A new status index derived from sociometric analysis. Psychometrika, 18(1):39–43, 1953.
  74. 74.Khanam, K. Z., Srivastava, G., and Mago, V. The homophily principle in social network analysis. arXiv preprint arXiv:2008.10383, 2020.
  75. 75.Kim, J., Nguyen, T. D., Min, S., Cho, S., Lee, M., Lee, H., and Hong, S. Pure transformers are powerful graph learners. arXiv, abs/2207.02505, 2022. URL https://arxiv.org/abs/2207.02505.
  76. 76.Kipf, T. N. and Welling, M. Variational graph auto-encoders. arXiv preprint arXiv:1611.07308, 2016.
  77. 77.Klicpera, J., Bojchevski, A., and Gunnemann, S. Predict then propagate: Graph neural networks meet personalized pagerank. In International Conference on Learning Representations, 2018.
  78. 78.Kreuzer, D., Beaini, D., Hamilton, W., Letourneau, V., and Tossou, P. Rethinking graph transformers with spectral attention. Advances in Neural Information Processing Systems, 34:21618–21629, 2021.
  79. 79.Kriege, N. M., Johansson, F. D., and Morris, C. A survey on graph kernels. Applied Network Science, 5(1):1–42, 2020.
  80. 80.Leskovec, J. and Krevl, A. SNAP Datasets: Stanford large network dataset collection. http://snap.stanford.edu/data, June 2014.
  81. 81.Leskovec, J., Chakrabarti, D., Kleinberg, J., Faloutsos, C., and Ghahramani, Z. Kronecker graphs: an approach to modeling networks. Journal of Machine Learning Research, 11(2), 2010.
  82. 82.Li, J., Shomer, H., Ding, J., Wang, Y., Ma, Y., Shah, N., Tang, J., and Yin, D. Are graph neural networks really helpful for knowledge graph completion? arXiv preprint arXiv:2205.10652, 2022.
  83. 83.Li, J., Shomer, H., Mao, H., Zeng, S., Ma, Y., Shah, N., Tang, J., and Yin, D. Evaluating graph neural networks for link prediction: Current pitfalls and new benchmarking. arXiv preprint arXiv:2306.10453, 2023a.
  84. 84.Li, Y., Li, Z., Wang, P., Li, J., Sun, X., Cheng, H., and Yu, J. X. A survey of graph meets large language model: Progress and future directions. arXiv preprint arXiv:2311.12399, 2023b.
  85. 85.Li, Y., Xiong, M., and Hooi, B. Graphcleaner: Detecting mislabelled samples in popular graph learning benchmarks. arXiv preprint arXiv:2306.00015, 2023c.
  86. 86.Li, Y., Wang, P., Li, Z., Yu, J. X., and Li, J. Zerog: Investigating cross-dataset zero-shot transferability in graphs. arXiv preprint arXiv:2402.11235, 2024.
  87. 87.Lim, D., Hohne, F., Li, X., Huang, S. L., Gupta, V., Bhalerao, O., and Lim, S. N. Large scale learning on non-homophilous graphs: New benchmarks and strong simple methods. Advances in Neural Information Processing Systems, 34:20887–20902, 2021.
  88. 88.Lim, D., Robinson, J. D., Zhao, L., Smidt, T., Sra, S., Maron, H., and Jegelka, S. Sign and basis invariant networks for spectral graph representation learning. In The Eleventh International Conference on Learning Representations, 2022.
  89. 89.Ling, X., Wu, L., Wang, S., Pan, G., Ma, T., Xu, F., Liu, A. X., Wu, C., and Ji, S. Deep graph matching and searching for semantic code retrieval. ACM Transactions on Knowledge Discovery from Data (TKDD), 15(5):1–21, 2021.
  90. 90.Liu, G., Inae, E., Zhao, T., Xu, J., Luo, T., and Jiang, M. Data-centric learning from unlabeled graphs with diffusion model. arXiv preprint arXiv:2303.10108, 2023a.
  91. 91.Liu, H., Feng, J., Kong, L., Liang, N., Tao, D., Chen, Y., and Zhang, M. One for all: Towards training one graph model for all classification tasks. arXiv preprint arXiv:2310.00149, 2023b.
  92. 92.Liu, H., Liao, N., and Luo, S. Simga: A simple and effective heterophilous graph neural network with efficient global aggregation. arXiv preprint arXiv:2305.09958, 2023c.
  93. 93.Liu, J., Yang, C., Lu, Z., Chen, J., Li, Y., Zhang, M., Bai, T., Fang, Y., Sun, L., Yu, P. S., et al. Towards graph foundation models: A survey and beyond. arXiv preprint arXiv:2310.11829, 2023d.
  94. 94.Liu, J., Mao, H., Chen, Z., Zhao, T., Shah, N., and Tang, J. Neural scaling laws on graphs. 2024a.
  95. 95.Liu, N., Wang, X., Bo, D., Shi, C., and Pei, J. Revisiting graph contrastive learning from the perspective of graph spectrum. Advances in Neural Information Processing Systems, 35:2972–2983, 2022.
  96. 96.Liu, R., Wang, Y., Xu, H., Liu, B., Sun, J., Guo, Z., and Ma, W. Source code vulnerability detection: Combining code language models and code property graphs. arXiv preprint arXiv:2404.14719, 2024b.
  97. 97.Liu, Z., Shi, Y., Zhang, A., Zhang, E., Kawaguchi, K., Wang, X., and Chua, T.-S. Rethinking tokenizer and decoder in masked graph modeling for molecules. In NeurIPS, 2023e. URL https://openreview.net/forum?id=fWLf8DV0fI.
  98. 98.Liu, Z., Yu, X., Fang, Y., and Zhang, X. Graphprompt: Unifying pre-training and downstream tasks for graph neural networks. In Proceedings of the ACM Web Conference 2023, pp. 417–428, 2023f.
  99. 99.Lu, S., Gao, Z., He, D., Zhang, L., and Ke, G. Highly accurate quantum chemical property prediction with unimol+. arXiv preprint arXiv:2303.16982, 2023.
  100. 100.Luan, S., Hua, C., Lu, Q., Zhu, J., Zhao, M., Zhang, S., Chang, X.-W., and Precup, D. Is heterophily a real nightmare for graph neural networks to do node classification? arXiv preprint arXiv:2109.05641, 2021.
  101. 101.Luan, S., Hua, C., Xu, M., Lu, Q., Zhu, J., Chang, X.-W., Fu, J., Leskovec, J., and Precup, D. When do graph neural networks help with node classification: Investigating the homophily principle on node distinguishability. arXiv preprint arXiv:2304.14274, 2023.
  102. 102.Luo, Y., Yan, K., and Ji, S. Graphdf: A discrete flow model for molecular graph generation. In International Conference on Machine Learning, pp. 7192–7203. PMLR, 2021.
  103. 103.Luo, Z., Song, X., Huang, H., Lian, J., Zhang, C., Jiang, J., Xie, X., and Jin, H. Graphinstruct: Empowering large language models with graph understanding and reasoning capability. arXiv preprint arXiv:2403.04483, 2024.
  104. 104.Ma, Y., Liu, X., Shah, N., and Tang, J. Is homophily a necessity for graph neural networks? arXiv preprint arXiv:2106.06134, 2021.
  105. 105.Mao, H., Chen, Z., Jin, W., Han, H., Ma, Y., Zhao, T., Shah, N., and Tang, J. Demystifying structural disparity in graph neural networks: Can one size fit all? arXiv preprint arXiv:2306.01323, 2023a.
  106. 106.Mao, H., Li, J., Shomer, H., Li, B., Fan, W., Ma, Y., Zhao, T., Shah, N., and Tang, J. Revisiting link prediction: A data perspective. arXiv preprint arXiv:2310.00793, 2023b.
  107. 107.Mao, H., Liu, G., Ma, Y., Wang, R., and Tang, J. A data generation perspective to the mechanism of in-context learning. arXiv preprint arXiv:2402.02212, 2024.
  108. 108.Masters, D., Dean, J., Klaser, K., Li, Z., Maddrell-Mander, S., Sanders, A., Helal, H., Beker, D., Rampa´sek, L., and Beaini, D. Gps++: An optimised hybrid mpnn/transformer for molecular property prediction. arXiv preprint arXiv:2212.02229, 2022.
  109. 109.McCoy, R. T., Yao, S., Friedman, D., Hardy, M., and Griffiths, T. L. Embers of autoregression: Understanding large language models through the problem they are trained to solve. arXiv preprint arXiv:2309.13638, 2023.
  110. 110.Menczer, F., Fortunato, S., and Davis, C. A. A First Course in Network Science. Cambridge University Press, 2020.
  111. 111.Milo, R., Shen-Orr, S., Itzkovitz, S., Kashtan, N., Chklovskii, D., and Alon, U. Network motifs: simple building blocks of complex networks. Science, 298(5594):824–827, 2002.
  112. 112.Mishra, S., Panda, R., Phoo, C. P., Chen, C.-F. R., Karlinsky, L., Saenko, K., Saligrama, V., and Feris, R. S. Task2sim: Towards effective pre-training and transfer from synthetic data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9194–9204, 2022.
  113. 113.Morris, C., Ritzert, M., Fey, M., Hamilton, W. L., Lenssen, J. E., Rattan, G., and Grohe, M. Weisfeiler and leman go neural: Higher-order graph neural networks. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pp. 4602–4609, 2019.
  114. 114.Morris, C., Kriege, N. M., Bause, F., Kersting, K., Mutzel, P., and Neumann, M. Tudataset: A collection of benchmark datasets for learning with graphs. In ICML 2020 Workshop on Graph Representation Learning and Beyond (GRL+ 2020), 2020. URL www.graphlearning.io.
  115. 115.Morris, C., Lipman, Y., Maron, H., Rieck, B., Kriege, N. M., Grohe, M., Fey, M., and Borgwardt, K. Weisfeiler and leman go machine learning: The story so far. Journal of Machine Learning Research, 24(333):1–59, 2023. URL http://jmlr.org/papers/v24/22-0240.html.
  116. 116.Muller, L., Galkin, M., Morris, C., and Ramp´aˇsek, L. Attending to graph transformers. arXiv preprint arXiv:2302.04181, 2023.
  117. 117.Murase, Y., Jo, H. H., Tor¨ok, J., Kert´esz, J., Kaski, K., et al. Structural transition in social networks. 2019.
  118. 118.Oono, K. and Suzuki, T. Graph neural networks exponentially lose expressive power for node classification. In International Conference on Learning Representations, 2019.
  119. 119.Panwar, M., Ahuja, K., and Goyal, N. In-context learning through the bayesian prism. In The Twelfth International Conference on Learning Representations, 2023.
  120. 120.Perozzi, B., Fatemi, B., Zelle, D., Tsitsulin, A., Kazemi, M., Al-Rfou, R., and Halcrow, J. Let your graph do the talking: Encoding structured data for llms. arXiv preprint arXiv:2402.05862, 2024.
  121. 121.Project, U. C. R. Recommender systems and personalization datasets. URL https://cseweb.ucsd.edu/˜jmcauley/datasets.html.
  122. 122.Prystawski, B. and Goodman, N. D. Why think step-by-step? reasoning emerges from the locality of experience. arXiv preprint arXiv:2304.03843, 2023.
  123. 123.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  124. 124.Rafi, M. N., Kim, D. J., Chen, A. R., Chen, T.-H., and Wang, S. Towards better graph neural neural network-based fault localization through enhanced code representation. arXiv preprint arXiv:2404.04496, 2024.
  125. 125.Ramsundar, B., Eastman, P., Walters, P., Pande, V., Leswing, K., and Wu, Z. Deep Learning for the Life Sciences. O’Reilly Media, 2019. https://www.amazon.com/Deep-Learning-Life-Sciences-Microscopy/dp/1492039837.
  126. 126.Ribeiro, P., Silva, F., and Kaiser, M. Strategies for network motifs discovery. In 2009 Fifth IEEE International Conference on e-Science, pp. 80–87. IEEE, 2009.
  127. 127.Ribeiro, P., Paredes, P., Silva, M. E., Aparicio, D., and Silva, F. A survey on subgraph counting: concepts, algorithms, and applications to network motifs and graphlets. ACM Computing Surveys (CSUR), 54(2):1–36, 2021.
  128. 128.Robins, G., Pattison, P., Kalish, Y., and Lusher, D. An introduction to exponential random graph (p*) models for social networks. Social networks, 29(2):173–191, 2007.
  129. 129.Rossi, R. A. and Ahmed, N. K. The network data repository with interactive graph analytics and visualization. In AAAI, 2015. URL http://networkrepository.com.
  130. 130.Ruiz, L., Chamon, L. F. O., and Ribeiro, A. Transferability properties of graph neural networks. IEEE Transactions on Signal Processing, 71:3474–3489, 2023. doi: 10.1109/TSP.2023.3297848.
  131. 131.Saparov, A. and He, H. Language models are greedy reasoners: A systematic formal analysis of chain-of-thought. arXiv preprint arXiv:2210.01240, 2022.
  132. 132.Schlichtkrull, M., Kipf, T. N., Bloem, P., Van Den Berg, R., Titov, I., and Welling, M. Modeling relational data with graph convolutional networks. In The Semantic Web: 15th International Conference, ESWC 2018, Heraklion, Crete, Greece, June 3–7, 2018, Proceedings 15, pp. 593–607. Springer, 2018.
  133. 133.Shi, H., Ding, J., Cao, Y., Liu, L., Li, Y., et al. Learning symbolic models for graph-structured physical mechanism. In The Eleventh International Conference on Learning Representations, 2022.
  134. 134.Shoghi, N., Kolluru, A., Kitchin, J. R., Ulissi, Z. W., Zitnick, C. L., and Wood, B. M. From molecules to materials: Pre-training large generalizable models for atomic property prediction, 2023.
  135. 135.Srinivasan, B. and Ribeiro, B. On the equivalence between positional node embeddings and structural graph representations. arXiv preprint arXiv:1910.00452, 2019.
  136. 136.Stechly, K., Marquez, M., and Kambhampati, S. Gpt-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems. In NeurIPS 2023 Foundation Models for Decision Making Workshop, 2023.
  137. 137.Sun, F.-Y., Hoffmann, J., Verma, V., and Tang, J. Infograph: Unsupervised and semi-supervised graph-level representation learning via mutual information maximization. arXiv preprint arXiv:1908.01000, 2019.
  138. 138.Sun, M., Zhou, K., He, X., Wang, Y., and Wang, X. Gppt: Graph pre-training and prompt tuning to generalize graph neural networks. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’22, pp. 1717–1727, New York, NY, USA, 2022. Association for Computing Machinery. ISBN 9781450393850. doi: 10.1145/3534678.3539249. URL https://doi.org/10.1145/3534678.3539249.
  139. 139.Sun, X., Cheng, H., Li, J., Liu, B., and Guan, J. All in one: Multi-task prompting for graph neural networks. 2023.
  140. 140.Taguchi, H., Liu, X., and Murata, T. Graph convolutional networks for graphs containing missing features. Future Generation Computer Systems, 117:155–168, 2021.
  141. 141.Tang, J. Aminer: Toward understanding big scholar data. In Proceedings of the Ninth ACM International Conference on Web Search and Data Mining, WSDM ’16, pp. 467, New York, NY, USA, 2016. Association for Computing Machinery. ISBN 9781450337168. doi: 10.1145/2835776.2835849. URL https://doi.org/10.1145/2835776.2835849.
  142. 142.Tang, J., Yang, Y., Wei, W., Shi, L., Su, L., Cheng, S., Yin, D., and Huang, C. Graphgpt: Graph instruction tuning for large language models. arXiv preprint arXiv:2310.13023, 2023.
  143. 143.Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085, 2022.
  144. 144.Topping, J., Di Giovanni, F., Chamberlain, B. P., Dong, X., and Bronstein, M. M. Understanding over-squashing and bottlenecks on graphs via curvature. In International Conference on Learning Representations, 2021.
  145. 145.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  146. 146.Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T. Solving olympiad geometry without human demonstrations. Nature, 625(7995):476–482, 2024.
  147. 147.Um, D., Park, J., Park, S., and young Choi, J. Confidence-based feature imputation for graphs with partially known features. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=YPKBIILy-Kt.
  148. 148.Van Den Oord, A., Vinyals, O., et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017.
  149. 149.Vashishth, S., Sanyal, S., Nitin, V., and Talukdar, P. Composition-based multi-relational graph convolutional networks. In International Conference on Learning Representations, 2019.
  150. 150.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  151. 151.Velickoviˇc, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017.
  152. 152.Velickoviˇc, P., Fedus, W., Hamilton, W. L., Lio, P., Bengio, Y., and Hjelm, R. D. Deep graph infomax. arXiv preprint arXiv:1809.10341, 2018.
  153. 153.Velickoviˇc, P., Badia, A. P., Budden, D., Pascanu, R., Banino, A., Dashevskiy, M., Hadsell, R., and Blundell, C. The clrs algorithmic reasoning benchmark. In International Conference on Machine Learning, pp. 22084–22102. PMLR, 2022.
  154. 154.Vignac, C., Krawczuk, I., Siraudin, A., Wang, B., Cevher, V., and Frossard, P. Digress: Discrete denoising diffusion for graph generation. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=UaAD-Nu86WX.
  155. 155.Vishwanathan, S. V. N., Schraudolph, N. N., Kondor, R., and Borgwardt, K. M. Graph kernels. Journal of Machine Learning Research, 11:1201–1242, 2010.
  156. 156.Wang, H., Yin, H., Zhang, M., and Li, P. Equivariant and stable positional encoding for more powerful graph neural networks. In International Conference on Learning Representations, 2021.
  157. 157.Wang, H., Feng, S., He, T., Tan, Z., Han, X., and Tsvetkov, Y. Can language models solve graph problems in natural language? In Thirty-seventh Conference on Neural Information Processing Systems, 2023a. URL https://openreview.net/forum?id=UDqHhbqYJV.
  158. 158.Wang, J., Guo, Y., Yang, L., and Wang, Y. Understanding heterophily for graph neural networks. arXiv preprint arXiv:2401.09125, 2024a.
  159. 159.Wang, J., Wu, J., Hou, Y., Liu, Y., Gao, M., and McAuley, J. Instructgraph: Boosting large language models via graph-centric instruction tuning and preference alignment. arXiv preprint arXiv:2402.08785, 2024b.
  160. 160.Wang, X., Yang, H., and Zhang, M. Neural common neighbor with completion for link prediction. arXiv preprint arXiv:2302.00890, 2023b.
  161. 161.Wang, Y., Elhag, A. A., Jaitly, N., Susskind, J. M., and Bautista, M. A. Generating molecular conformer fields. arXiv preprint arXiv:2311.17932, 2023c.
  162. 162.Wang, Y., Cui, H., and Kleinberg, J. Microstructures and accuracy of graph recall by large language models. arXiv preprint arXiv:2402.11821, 2024c.
  163. 163.Wu, Q., Zhao, W., Li, Z., Wipf, D. P., and Yan, J. Node-former: A scalable graph structure learning transformer for node classification. Advances in Neural Information Processing Systems, 35:27387–27401, 2022.
  164. 164.Wu, X., Ajorlou, A., Wu, Z., and Jadbabaie, A. Demystifying oversmoothing in attention-based graph neural networks. arXiv preprint arXiv:2305.16102, 2023.
  165. 165.Xia, J., Zhao, C., Hu, B., Gao, Z., Tan, C., Liu, Y., Li, S., and Li, S. Z. Mole-BERT: Rethinking pre-training graph neural networks for molecules. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=jevY-DtiZTR.
  166. 166.Xie, H., Zheng, D., Ma, J., Zhang, H., Ioannidis, V. N., Song, X., Ping, Q., Wang, S., Yang, C., Xu, Y., et al. Graph-aware language model pre-training on a large graph corpus can help multiple graph applications. arXiv preprint arXiv:2306.02592, 2023.
  167. 167.Xu, K., Hu, W., Leskovec, J., and Jegelka, S. How powerful are graph neural networks? In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=ryGs6iA5Km.
  168. 168.Xu, K., Li, J., Zhang, M., Du, S. S., Kawarabayashi, K.-i., and Jegelka, S. What can neural networks reason about? ICLR 2020, 2020.
  169. 169.Yang, L., Tian, Y., Xu, M., Liu, Z., Hong, S., Qu, W., Zhang, W., Cui, B., Zhang, M., and Leskovec, J. Vqgraph: Graph vector-quantization for bridging gnns and mlps. arXiv preprint arXiv:2308.02117, 2023.
  170. 170.Yang, Y., Liu, T., Wang, Y., Zhou, J., Gan, Q., Wei, Z., Zhang, Z., Huang, Z., and Wipf, D. Graph neural networks inspired by classical iterative algorithms. In International Conference on Machine Learning, pp. 11773–11783. PMLR, 2021.
  171. 171.Yao, Y., Wang, X., Zhang, Z., Qin, Y., Zhang, Z., Chu, X., Yang, Y., Zhu, W., and Mei, H. Exploring the potential of large language models in graph generation. arXiv preprint arXiv:2403.14358, 2024.
  172. 172.Yasunaga, M., Bosselut, A., Ren, H., Zhang, X., Manning, C. D., Liang, P., and Leskovec, J. Deep bidirectional language-knowledge graph pretraining. In Neural Information Processing Systems (NeurIPS), 2022a.
  173. 173.Yasunaga, M., Leskovec, J., and Liang, P. Linkbert: Pre-training language models with document links. arXiv preprint arXiv:2203.15827, 2022b.
  174. 174.Ye, H., Zhang, N., Chen, H., and Chen, H. Generative knowledge graph construction: A review. arXiv preprint arXiv:2210.12714, 2022.
  175. 175.Ye, R., Zhang, C., Wang, R., Xu, S., and Zhang, Y. Language is all a graph needs. In Findings of the Association for Computational Linguistics: EACL 2024, pp. 1955–1973, 2024.
  176. 176.Yin, H., Zhang, M., Wang, Y., Wang, J., and Li, P. Algorithm and system co-design for efficient subgraph-based graph representation learning. arXiv preprint arXiv:2202.13538, 2022.
  177. 177.Ying, R., He, R., Chen, K., Eksombatchai, P., Hamilton, W. L., and Leskovec, J. Graph convolutional neural networks for web-scale recommender systems. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pp. 974–983, 2018.
  178. 178.You, J., Ying, R., and Leskovec, J. Position-aware graph neural networks. In International conference on machine learning, pp. 7134–7143. PMLR, 2019.
  179. 179.You, J., Gomes-Selman, J. M., Ying, R., and Leskovec, J. Identity-aware graph neural networks. In Proceedings of the AAAI conference on artificial intelligence, volume 35, pp. 10737–10745, 2021.
  180. 180.You, Y., Chen, T., Sui, Y., Chen, T., Wang, Z., and Shen, Y. Graph contrastive learning with augmentations. Advances in neural information processing systems, 33:5812–5823, 2020.
  181. 181.You, Y., Chen, T., Wang, Z., and Shen, Y. Graph domain adaptation via theory-grounded spectral regularization. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=OysfLgrk8mk.
  182. 182.Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Gupta, A., Gu, X., Hauptmann, A. G., et al. Language model beats diffusion–tokenizer is key to visual generation. arXiv preprint arXiv:2310.05737, 2023.
  183. 183.Yue, Z., Rabhi, S., Moreira, G. d. S. P., Wang, D., and Oldridge, E. Llamarec: Two-stage recommendation using large language models for ranking. arXiv preprint arXiv:2311.02089, 2023.
  184. 184.Zeng, H., Zhou, H., Srivastava, A., Kannan, R., and Prasanna, V. Graphsaint: Graph sampling based inductive learning method. arXiv preprint arXiv:1907.04931, 2019.
  185. 185.Zeng, H., Zhang, M., Xia, Y., Srivastava, A., Malevich, A., Kannan, R., Prasanna, V., Jin, L., and Chen, R. Decoupling the depth and scope of graph neural networks. In Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, 2021. URL https://openreview.net/forum?id=d0MtHWY0NZ.
  186. 186.Zhai, G., Ornek, E. P., Wu, S.-C., Di, Y., Tombari, F., Navab, N., and Busam, B. Commonscenes: Generating commonsense 3d indoor scenes with scene graphs. arXiv preprint arXiv:2305.16283, 2023.
  187. 187.Zhai, G., Ornek, E. P., Wu, S.-C., Di, Y., Tombari, F., Navab, N., and Busam, B. Commonscenes: Generating commonsense 3d indoor scenes with scene graphs. Advances in Neural Information Processing Systems, 36, 2024.
  188. 188.Zhang, B., Gai, J., Du, Y., Ye, Q., He, D., and Wang, L. Beyond weisfeiler-lehman: A quantitative framework for gnn expressiveness. arXiv preprint arXiv:2401.08514, 2024.
  189. 189.Zhang, D., Liu, X., Zhang, X., Zhang, C., Cai, C., Bi, H., Du, Y., Qin, X., Huang, J., Li, B., Shan, Y., Zeng, J., Zhang, Y., Liu, S., Li, Y., Chang, J., Wang, X., Zhou, S., Liu, J., Luo, X., Wang, Z., Jiang, W., Wu, J., Yang, Y., Yang, J., Yang, M., Gong, F.-Q., Zhang, L., Shi, M., Dai, F.-Z., York, D. M., Liu, S., Zhu, T., Zhong, Z., Lv, J., Cheng, J., Jia, W., Chen, M., Ke, G., E, W., Zhang, L., and Wang, H. Dpa-2: Towards a universal large atomic model for molecular and material simulation. arXiv preprint arXiv:2312.15492, 2023a.
  190. 190.Zhang, F., Liu, X., Tang, J., Dong, Y., Yao, P., Zhang, J., Gu, X., Wang, Y., Kharlamov, E., Shao, B., Li, R., and Wang, K. Oag: Linking entities across large-scale heterogeneous knowledge graphs. IEEE Transactions on Knowledge and Data Engineering, 35(9):9225–9239, 2023b. doi: 10.1109/TKDE.2022.3222168.
  191. 191.Zhang, M. and Chen, Y. Link prediction based on graph neural networks. Advances in neural information processing systems, 31, 2018.
  192. 192.Zhang, M., Li, P., Xia, Y., Wang, K., and Jin, L. Labeling trick: A theory of using graph neural networks for multi-node representation learning. Advances in Neural Information Processing Systems, 34:9061–9073, 2021.
  193. 193.Zhang, Y., Zhang, F., Yang, Z., and Wang, Z. What and how does in-context learning learn? bayesian model averaging, parameterization, and generalization. arXiv preprint arXiv:2305.19420, 2023c.
  194. 194.Zhang, Z., Li, H., Zhang, Z., Qin, Y., Wang, X., and Zhu, W. Graph meets llms: Towards large graph models. In NeurIPS 2023 Workshop: New Frontiers in Graph Learning, 2023d.
  195. 195.Zhang, Z., Luo, B., Lu, S., and He, B. Live graph lab: Towards open, dynamic and real transaction graphs with nft. arXiv preprint arXiv:2310.11709, 2023e.
  196. 196.Zhao, H., Liu, S., Chang, M., Xu, H., Fu, J., Deng, Z., Kong, L., and Liu, Q. Gimlet: A unified graph-text model for instruction-based molecule zero-shot learning. Advances in Neural Information Processing Systems, 36, 2024.
  197. 197.Zhao, J., Zhuo, L., Shen, Y., Qu, M., Liu, K., Bronstein, M., Zhu, Z., and Tang, J. Graphtext: Graph reasoning in text space. arXiv preprint arXiv:2310.01089, 2023a.
  198. 198.Zhao, Q., Ren, W., Li, T., Xu, X., and Liu, H. Graphgpt: Graph learning with generative pre-trained transformers. arXiv preprint arXiv:2401.00529, 2023b.
  199. 199.Zheng, S., He, J., Liu, C., Shi, Y., Lu, Z., Feng, W., Ju, F., Wang, J., Zhu, J., Min, Y., Zhang, H., Tang, S., Hao, H., Jin, P., Chen, C., Noe, F., Liu, H., and Liu, T.-Y. Towards predicting equilibrium distributions for molecular systems with deep learning. arXiv preprint arXiv:2306.05445, 2023a.
  200. 200.Zheng, W., Huang, E. W., Rao, N., Wang, Z., and Subbian, K. You only transfer what you share: Intersection-induced graph transfer learning for link prediction. arXiv preprint arXiv:2302.14189, 2023b.
  201. 201.Zhong, Y., Shi, J., Yang, J., Xu, C., and Li, Y. Learning to generate scene graph from natural language supervision. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1823–1834, 2021.
  202. 202.Zhu, Y., Xu, Y., Liu, Q., and Wu, S. An empirical study of graph contrastive learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021a.
  203. 203.Zhu, Z., Zhang, Z., Xhonneux, L.-P., and Tang, J. Neural bellman-ford networks: A general graph neural network framework for link prediction. Advances in Neural Information Processing Systems, 34:29476–29490, 2021b.

Citation

MLA
Mao, H., et al. “Position: Graph Foundation Models Are Already Here”. arXiv, 2024, http://arxiv.org/abs/2402.02216v3.
APA
Mao, H., Chen, Z., Tang, W., Zhao, J., Ma, Y., Zhao, T., Shah, N., Galkin, M., & Tang, J. (2024). Position: Graph Foundation Models are Already Here. arXiv. http://arxiv.org/abs/2402.02216v3
Chicago
Mao, H., Z. Chen, W. Tang, et al. 2024. “Position: Graph Foundation Models Are Already Here”. arXiv. http://arxiv.org/abs/2402.02216v3.
Harvard
Mao, H. et al. (2024) “Position: Graph Foundation Models are Already Here”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.02216v3.
Vancouver
1. Mao H, Chen Z, Tang W, Zhao J, Ma Y, Zhao T, Shah N, Galkin M, Tang J (2024) Position: Graph Foundation Models are Already Here. arXiv

BibTeX

@article{mao2024position,
  title = {Position: Graph Foundation Models are Already Here},
  author = {Mao, Haitao and Chen, Zhikai and Tang, Wenzhuo and Zhao, Jianan and Ma, Yao and Zhao, Tong and Shah, Neil and Galkin, Mikhail and Tang, Jiliang},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.02216v3},
  eprint = {2402.02216}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/