keyword
graph Transformer
A graph Transformer is a deep learning architecture that adapts the self-attention mechanism of standard Transformers to learn representations from graph-structured data. Unlike conventional message-passing neural networks that are constrained to local neighborhood aggregation, graph Transformers can model long-range dependencies and global interactions across all nodes in a network. To preserve and utilize non-Euclidean topology, these architectures incorporate graph-specific structural and positional encodings, such as shortest-path distances, node centrality measures, Laplacian eigenvectors, and edge attributes, directly into node embeddings or attention matrices. This design helps alleviate common limitations of standard graph neural networks, such as information bottlenecks and over-smoothing, making graph Transformers effective for node classification, link prediction, and whole-graph property prediction across domains such as molecular chemistry, bioinformatics, and network analysis.
9 items

DiGress: Discrete Denoising diffusion for graph generation
Clement Vignac, Igor Krawczuk, Antoine Siraudin, Bohan Wang, Volkan Cevher, Pascal Frossard
Why you should read this
Introduces a discrete denoising diffusion model that generates graphs by directly editing categorical node and edge attributes, scaling diffusion-based generation to million-molecule benchmarks while outperforming existing methods across molecular and non-molecular structures.
This work introduces DiGress, a discrete denoising diffusion model for generating graphs with categorical node and edge attributes. Our model utilizes a discrete diffusion process that progressively edits graphs with noise, through the process of adding or removing edges and changing the categories. A graph transformer network is trained to revert this process, simplifying the problem of distribution learning over graphs into a sequence of node and edge classification tasks. We further improve sample quality by introducing a Markovian noise model that preserves the marginal distribution of node and edge types during diffusion, and by incorporating auxiliary graph-theoretic features. A procedure for conditioning the generation on graph-level features is also proposed. DiGress achieves state-of-the-art performance on molecular and non-molecular datasets, with up to 3x validity improvement on a planar graph dataset. It is also the first model to scale to the large GuacaMol dataset containing 1.3M drug-like molecules without the use of molecule-specific representations.
Added
2026-10-05

Graph Generation with Diffusion Mixture
Jaehyeong Jo, Dongki Kim, Sung Ju Hwang
Why you should read this
Presents GruM, a graph generative framework that formulates the generation process as a mixture of endpoint-conditioned Ornstein-Uhlenbeck bridge diffusions to directly predict final graph topologies rather than denoise step-by-step, achieving faster convergence and superior generation across general graph and 2D/3D molecule benchmarks.
Generation of graphs is a major challenge for real-world tasks that require understanding the complex nature of their non-Euclidean structures. Although diffusion models have achieved notable success in graph generation recently, they are ill-suited for modeling the topological properties of graphs since learning to denoise the noisy samples does not explicitly learn the graph structures to be generated. To tackle this limitation, we propose a generative framework that models the topology of graphs by explicitly learning the final graph structures of the diffusion process. Specifically, we design the generative process as a mixture of endpoint-conditioned diffusion processes which is driven toward the predicted graph that results in rapid convergence. We further introduce a simple parameterization of the mixture process and develop an objective for learning the final graph structure, which enables maximum likelihood training. Through extensive experimental validation on general graph and 2D/3D molecule generation tasks, we show that our method outperforms previous generative models, generating graphs with correct topology with both continuous (e.g. 3D coordinates) and discrete (e.g. atom types) features. Our code is available at https://github.com/harryjo97/GruM.
Added
2026-10-04


On the Connection Between MPNN and Graph Transformer
Chen Cai, Truong Son Hy, Rose Yu, Yusu Wang
Why you should read this
Proves that Message Passing Neural Networks augmented with a single virtual node can theoretically approximate graph transformer self-attention layers under explicit depth-width trade-offs, bridging the expressiveness gap between local message passing and global attention architectures.
Graph Neural Networks (GNNs) and Transformers have recently become prominent models for graph representation learning. In particular, Message Passing Neural Networks (MPNNs), a leading class of GNNs, enjoy great expressive power by alleviating the bottle- neck caused by separating two procedures within each layer? One defined so-called degle model fe unit sort, the schemes gave birth increasing connectivity among graphson transition. njajaja and Lamberti et al.
Added
2026-10-02

LLaGA: Large Language and Graph Assistant
Runjin Chen, Tong Zhao, Ajay Kumar Jaiswal, Neil Shah, Zhangyang Wang
Why you should read this
Introduces a general-purpose framework that reorganizes graph topologies into structure-aware node sequences and maps them directly into large language model token embeddings, outperforming specialized graph neural networks across multiple tasks and unseen datasets without modifying the base model parameters.
Graph Neural Networks (GNNs) have empowered the advance in graph-structured data analysis. Recently, the rise of Large Language Models (LLMs) like GPT-4 has heralded a new era in deep learning. However, their application to graph data poses distinct challenges due to the inherent difficulty of translating graph structures to language. To this end, we introduce the Large Language and Graph Assistant (LLaGA), an innovative model that effectively integrates LLM capabilities to handle the complexities of graph-structured data. LLaGA retains the general-purpose nature of LLMs while adapting graph data into a format compatible with LLM input. LLaGA achieves this by reorganizing graph nodes to structure-aware sequences and then mapping these into the token embedding space through a versatile projector. LLaGA excels in versatility, generalizability and interpretability, allowing it to perform consistently well across different datasets and tasks, extend its ability to unseen datasets or tasks, and provide explanations for graphs. Our extensive experiments across popular graph benchmarks show that LLaGA delivers outstanding performance across four datasets and three tasks using one single model, surpassing state-of-the-art graph models in both supervised and zero-shot scenarios. Our code is available at https://github.com/VITA-Group/LLaGA
Added
2026-09-30

Brain Network Transformer
Xuan Kan, Wei Dai, Hejie Cui, Zilong Zhang, Ying Guo, Carl Yang
Why you should read this
Proposes the Brain Network Transformer, a specialized graph architecture that leverages ROI connection profiles for natural positional encoding and introduces an orthonormal clustering readout to identify functional brain modules, achieving superior predictive accuracy on standard fMRI benchmarks.
Human brains are commonly modeled as networks of Regions of Interest (ROIs) and their connections for the understanding of brain functions and mental disorders. Recently, Transformer-based models have been studied over different types of data, including graphs, shown to bring performance gains widely. In this work, we study Transformer-based models for brain network analysis. Driven by the unique properties of data, we model brain networks as graphs with nodes of fixed size and order, which allows us to (1) use connection profiles as node features to provide natural and low-cost positional information and (2) learn pair-wise connection strengths among ROIs with efficient attention weights across individuals that are predictive towards downstream analysis tasks. Moreover, we propose an ORTHONORMAL CLUSTERING READOUT operation based on self-supervised soft clustering and orthonormal projection. This design accounts for the underlying functional modules that determine similar behaviors among groups of ROIs, leading to distinguishable cluster-aware node embeddings and informative graph embeddings. Finally, we re-standardize the evaluation pipeline on the only one publicly available large-scale brain network dataset of ABIDE, to enable meaningful comparison of different models. Experiment results show clear improvements of our proposed BRAIN NETWORK TRANSFORMER on both the public ABIDE and our restricted ABCD datasets. The implementation is available at https://github.com/Wayfear/BrainNetworkTransformer.
Added
2026-09-26

Molformer: Motif-Based Transformer on 3D Heterogeneous Molecular Graphs
Fang Wu, Dragomir Radev, Stan Z. Li
Why you should read this
Proposes Molformer, a geometric Transformer that integrates multi-level molecular motifs and 3D atomic coordinates into heterogeneous graphs via reinforcement learning-based motif mining and heterogeneous self-attention to improve molecular property prediction across small molecules and proteins.
Procuring expressive molecular representations underpins AI-driven molecule design and scientific discovery. The research mainly focuses on atom-level homogeneous molecular graphs, ignoring the rich information in subgraphs or motifs. However, it has been widely accepted that substructures play a dominant role in identifying and determining molecular properties. To address such issues, we formulate heterogeneous molecular graphs (HMGs), and introduce a novel architecture to exploit both molecular motifs and 3D geometry. Precisely, we extract functional groups as motifs for small molecules and employ reinforcement learning to adaptively select quaternary amino acids as motif candidates for proteins. Then HMGs are constructed with both atom-level and motif-level nodes. To better accommodate those HMGs, we introduce a variant of the Transformer named Molformer, which adopts a heterogeneous self-attention layer to distinguish the interactions between multi-level nodes. Besides, it is also coupled with a multi-scale mechanism to capture fine-grained local patterns with increasing contextual scales. An attentive farthest point sampling algorithm is also proposed to obtain the molecular representations. We validate Molformer across a broad range of domains, including quantum chemistry, physiology, and biophysics. Extensive experiments show that Molformer outperforms or achieves the comparable performance of several state-of-the-art baselines. Our work provides a promising way to utilize informative motifs from the perspective of multi-level graph construction. The code is available at https://github.com/smiles724/Molformer.
Added
2026-09-26
Representational Strengths and Limitations of Transformers
Clayton Sanford, Daniel J. Hsu, Matus Telgarsky
Added
2026-09-26

Do Transformers Really Perform Badly for Graph Representation?
Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, Tie-Yan Liu
Why you should read this
Introduces Graphormer, a standard Transformer-based architecture that integrates centrality, spatial, and edge encodings to outperform conventional graph neural networks across major graph representation benchmarks.
The Transformer architecture has become a dominant choice in many domains, such as natural language processing and computer vision. Yet, it has not achieved competitive performance on popular leaderboards of graph-level prediction compared to mainstream GNN variants. Therefore, it remains a mystery how Transformers could perform well for graph representation learning. In this paper, we solve this mystery by presenting Graphormer, which is built upon the standard Transformer architecture, and could attain excellent results on a broad range of graph representation learning tasks, especially on the recent OGB Large-Scale Challenge. Our key insight to utilizing Transformer in the graph is the necessity of effectively encoding the structural information of a graph into the model. To this end, we propose several simple yet effective structural encoding methods to help Graphormer better model graph-structured data. Besides, we mathematically characterize the expressive power of Graphormer and exhibit that with our ways of encoding the structural information of graphs, many popular GNN variants could be covered as the special cases of Graphormer. The code and models of Graphormer will be made publicly available at https://github.com/microsoft/Graphormer.
Added
2026-09-24

Transformers in Time Series: A Survey
Qingsong Wen, Tian Zhou, Chao Zhang, Weiqiu Chen, Ziqing Ma, Junchi Yan, Liang Sun
Why you should read this
Presents a systematic taxonomy of Transformer adaptations for time series forecasting, anomaly detection, and classification alongside empirical analyses of model size and seasonal decomposition to guide future architectural designs.
Transformers have achieved superior performances in many tasks in natural language processing and computer vision, which also triggered great interest in the time series community. Among multiple advantages of Transformers, the ability to capture long-range dependencies and interactions is especially attractive for time series modeling, leading to exciting progress in various time series applications. In this paper, we systematically review Transformer schemes for time series modeling by highlighting their strengths as well as limitations. In particular, we examine the development of time series Transformers in two perspectives. From the perspective of network structure, we summarize the adaptations and modifications that have been made to Transformers in order to accommodate the challenges in time series analysis. From the perspective of applications, we categorize time series Transformers based on common tasks including forecasting, anomaly detection, and classification. Empirically, we perform robust analysis, model size analysis, and seasonal-trend decomposition analysis to study how Transformers perform in time series. Finally, we discuss and suggest future directions to provide useful research guidance. To the best of our knowledge, this paper is the first work to comprehensively and systematically summarize the recent advances of Transformers for modeling time series data. We hope this survey will ignite further research interests in time series Transformers.
Added
2026-09-24
