Position: Graph Foundation Models Are Already Here
Haitao MaoZhikai ChenWenzhuo TangJianan ZhaoYao MaTong ZhaoNeil ShahMikhail GalkinJiliang Tang
Presents a unifying graph vocabulary perspective that explains how current primitive models achieve cross-dataset transferability and provides concrete principles to guide the design of general-purpose Graph Foundation Models.
Graph learning traditionally relies on training separate Graph Neural Networks from scratch for individual tasks and datasets. While foundational AI models have transformed fields like computer vision and natural language processing through broad pre-training, graph-based applications continue to suffer from redundant development costs and poor cross-dataset generalization. Developing universal Graph Foundation Models is challenging because graphs represent abstract, highly heterogeneous, non-Euclidean structures without native standard representations. Addressing this challenge is critical for scaling machine learning applications across drug discovery, molecular chemistry, knowledge management, and fraud detection.
The article systematically evaluates the current landscape of Graph Foundation Models and investigates how to achieve positive cross-dataset transferability. It proposes a foundational framework centered on constructing a transferable "graph vocabulary" based on network analysis, mathematical expressiveness, and structural stability.
To conduct this assessment, the authors analyzed foundational principles and empirical results across multiple model paradigms, benchmark suites, and tasks including node classification, link prediction, and graph classification. They examined leading specialized architectures (such as ULTRA and DiG), cross-task unified models (such as OneForAll), and methods bridging graph structures with Large Language Models.
The analysis establishes several key findings. First, while general-purpose models spanning all domains do not yet exist, specialized task-specific and domain-specific models have achieved practical zero-shot generalization to unseen graphs. Second, successful knowledge transfer depends on establishing an invariant vocabulary that avoids compressing non-isomorphic structures into identical representations. Third, distinct network properties require specialized modeling choices; for instance, modeling homophilic graphs ("birds of a feather") and heterophilic graphs ("opposites attract") requires separate aggregation mechanisms. Fourth, Large Language Models excel as feature encoders for text-rich graphs but demonstrate fundamental limitations when performing complex graph structural reasoning on their own. Finally, graph models can exhibit neural scaling laws when architectures incorporate suitable inductive structural priors, such as geometric invariants or discrete tokenization.
These findings indicate that organizations can move beyond single-purpose graph models toward reusable, domain-tailored foundation models. Adopting this paradigm can reduce training resource consumption and decrease dependence on expensive domain-expert labeling. However, moving toward a single universal graph model carries operational risks. Forcing completely unaligned domains into a shared architecture risks negative transfer, where model performance degrades across non-isomorphic or conflicting relational patterns.
Organizations developing graph-based AI should prioritize task- and domain-specific foundation models, particularly in structured areas like molecular chemistry and knowledge graphs. Architecture choices must pair expressive graph neural network backbones or graph transformers with separate encoding pipelines for structural and feature proximity. Rather than relying entirely on text-based language models for topological reasoning, engineering teams should leverage language models primarily to align disparate text features while employing dedicated graph tokenizers for structural modeling. Further research and standardized benchmarking are required to determine whether a truly universal structural representation space exists across divergent graph domains.
Confidence in specialized domain-specific graph models is strong, supported by robust empirical evidence in knowledge graphs and chemistry. However, substantial uncertainties remain regarding cross-domain transferability, severe data scarcity relative to text and image datasets, and noisy graph construction standards (such as mislabeling rates exceeding 15% in standard benchmarks). Decision-makers should approach claims of universal cross-domain graph capabilities with caution until standardized, robust scaling laws are further demonstrated.
- Paper: Strategies for Pre-training Graph Neural Networks, Weihua Hu et al. (2020). Its comparison of graph pre-training strategies and transfer across chemistry and biology supplies the direct foundation for assessing whether graph models can generalize across datasets.
- Paper: Weisfeiler and Leman go Machine Learning: The Story so far, Christopher Morris et al. (2023). Its account of Weisfeiler–Leman expressiveness clarifies why preserving structural distinctions is central to the source’s proposed transferable graph vocabulary.
- Paper: Representation Learning on Graphs: Methods and Applications, William L. Hamilton et al. (2017). Its encoder–decoder framework and treatment of inductive generalization provide useful grounding for the source’s discussion of transferable graph representations.
- Paper: Relational inductive biases, deep learning, and graph networks, Peter W. Battaglia et al. (2018). Its framework for relational inductive biases explains why graph models need structure-aware priors, a key premise behind the source’s account of transfer and scaling.
- Paper: Cooperative Graph Neural Networks, Ben Finkelshtein et al. (2024). It develops adaptive, entity-specific communication policies that extend the source’s case for specialized graph models to settings with heterophily and long-range information flow.
