Brain Network Transformer

Xuan KanWei DaiHejie CuiZilong ZhangYing GuoCarl Yang

article2022NeurIPS208 citations

Proposes the Brain Network Transformer, a specialized graph architecture that leverages ROI connection profiles for natural positional encoding and introduces an orthonormal clustering readout to identify functional brain modules, achieving superior predictive accuracy on standard fMRI benchmarks.

Listen

Analyzing brain networks derived from neuroimaging is vital for understanding brain organization, cognitive functions, and mental disorders. While attention-based Transformer models have achieved major breakthroughs across artificial intelligence, adapting them to brain network analysis remains difficult. Standard graph Transformers are designed for sparse graphs and rely on computationally expensive positional encodings, whereas brain networks are dense, fully connected structures where traditional positional encodings and edge-weight features are either computationally prohibitive or uninformative.

The article develops and evaluates the Brain Network Transformer, a specialized architecture tailored to the unique properties of brain networks. The system aims to provide accurate diagnostic classifications and brain trait predictions while aligning learned network interactions with underlying modular functional brain systems.

The researchers evaluated the proposed model on two large-scale neuroimaging datasets: the publicly available Autism Brain Imaging Data Exchange (ABIDE) dataset comprising 1,009 subjects for autism spectrum disorder diagnosis, and the restricted Adolescent Brain Cognitive Development (ABCD) dataset comprising 7,901 subjects for biological sex prediction. To overcome cross-site scanner inconsistencies in the public ABIDE data, the article introduced a standardized stratified data splitting strategy based on collection sites. The model architecture uses full connection profiles as low-cost natural positional features, learns fully pairwise attention weights via multi-head self-attention, and pools node representations into graph-level embeddings using a novel Orthonormal Clustering Readout module.

The evaluation produced four core findings. First, the proposed model consistently outperformed state-of-the-art graph neural networks, convolutional baselines, and existing graph Transformers, achieving diagnostic prediction improvements of up to 6 percentage points in area under the receiver operating characteristic curve across both datasets (reaching 80.2% on ABIDE and 96.2% on ABCD). Second, the Orthonormal Clustering Readout demonstrated clear advantages over standard pooling functions, outperforming methods such as Mean, Sum, and DiffPool. Third, cluster initialization via orthonormal bases significantly improved representation quality, achieving higher accuracy and lower variance when using a relatively small number of clusters (around 4 to 10), which aligns with the true biological scale of brain functional modules. Fourth, the attention scores learned by the model accurately matched recognized neurological functional modules, confirming biological interpretability.

These results demonstrate that deep learning models can achieve higher diagnostic accuracy and computational efficiency when customized to neuroimaging structures rather than relying on generic graph architectures. By eliminating redundant mathematical encodings and automating the grouping of brain regions into functional clusters, the framework provides an interpretable tool for clinical research without requiring costly manual module labeling.

Stakeholders and research teams should adopt the proposed architecture as a foundational backbone for automated connectome analysis and incorporate the standardized stratified evaluation pipeline on multi-site datasets. Future work should prioritize extending the clustering framework to structural connectivity datasets, evaluating broader neurodegenerative conditions, and integrating explicit explainability modules for clinical validation.

The primary limitations include reliance on functional connectivity data derived from specific anatomical atlases and the presence of multi-site scanner variations, which can induce instability if stratified sampling is not carefully applied. Despite these challenges, there is high confidence in the architectural performance gains given the rigorous theoretical proofs and empirical validations across large participant cohorts.

  • Paper: Do Transformers Really Perform Badly for Graph Representation?, Chengxuan Ying et al. (2021). Establishes foundational structural and positional encoding principles for graph Transformers, providing the sparse-graph baseline architecture that Brain Network Transformer adapts for dense connectomes.
  • Paper: Graph Attention Networks, Petar Veličković et al. (2018). Introduces self-attention mechanisms tailored to graph-structured data, serving as the conceptual foundation for learning attention-based interactions over network nodes.
  • Paper: Graph Transformer Networks, Seongjun Yun et al. (2019). Pioneers graph transformer architectures to learn composite relationships directly from network structures, laying groundwork for attention-driven graph representations.
  • Paper: Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering, Michaël Defferrard et al. (2016). Formulates graph clustering and localized spectral operations that underpin hierarchical pooling and modular node grouping in brain network analysis.
  • Paper: Beyond Mind-Reading: Multi-Voxel Pattern Analysis of fMRI Data, Kenneth A. Norman et al. (2006). Presents the principles of multi-voxel pattern analysis on functional neuroimaging data for decoding cognitive states and clinical phenotypes from distributed brain activations.

No sufficiently relevant recommendations were found.

Cover for Brain Network Transformer

Abstract

Human brains are commonly modeled as networks of Regions of Interest (ROIs) and their connections for the understanding of brain functions and mental disorders. Recently, Transformer-based models have been studied over different types of data, including graphs, shown to bring performance gains widely. In this work, we study Transformer-based models for brain network analysis. Driven by the unique properties of data, we model brain networks as graphs with nodes of fixed size and order, which allows us to (1) use connection profiles as node features to provide natural and low-cost positional information and (2) learn pair-wise connection strengths among ROIs with efficient attention weights across individuals that are predictive towards downstream analysis tasks. Moreover, we propose an ORTHONORMAL CLUSTERING READOUT operation based on self-supervised soft clustering and orthonormal projection. This design accounts for the underlying functional modules that determine similar behaviors among groups of ROIs, leading to distinguishable cluster-aware node embeddings and informative graph embeddings. Finally, we re-standardize the evaluation pipeline on the only one publicly available large-scale brain network dataset of ABIDE, to enable meaningful comparison of different models. Experiment results show clear improvements of our proposed BRAIN NETWORK TRANSFORMER on both the public ABIDE and our restricted ABCD datasets. The implementation is available at https://github.com/Wayfear/BrainNetworkTransformer.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 2.1 GNNs for Brain Network Analysis
  • 2.2 Graph Transformer
  • 3 BRAIN NETWORK TRANSFORMER
  • 3.1 Problem Definition
  • 3.2 Multi-Head Self-Attention Module (MHSA)
  • 3.3 ORTHONORMAL CLUSTERING READOUT (OCREAD)
  • 3.3.1 Theoretical Justifications
  • 3.4 Generalizing OCREAD to Other Graph Tasks and Domains
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Performance Analysis (RQ1)
  • 4.3 Ablation Studies on the OCREAD Module (RQ2)
  • 4.3.1 OCREAD with varying readout functions
  • 4.3.2 OCREAD with varying cluster initializations
  • 4.4 In-depth Analysis of Attention Scores and Cluster Assignments (RQ3)
  • 5 Discussion and Conclusion
  • 6 Acknowledgments
  • References

Knowls

  1. Knowl 1 — Brain Network Transformer Architecture

    model/method

    The Brain Network Transformer (BrainNetTF) is a specialized graph Transformer model designed to analyze brain networks represented as complete weighted graphs X∈RV×VX \in \mathbb{R}^{V \times V}, where VV denotes the number of Regions of Interest (ROIs). The framework consists of two core components:

    1. A Multi-Head Self-Attention (MHSA) module with LL layers that maps the initial node features (the connection profile matrix XX) to contextualized node embeddings ZL∈RV×VZ^L \in \mathbb{R}^{V \times V}.
    2. An Orthonormal Clustering Readout (OCREAD) module that maps the node embeddings ZLZ^L into modular graph-level representations ZG∈RK×VZ_G \in \mathbb{R}^{K \times V} using KK orthonormal cluster centers.

    The pooled graph embedding ZGZ_G is flattened into a vector of length KVKV and processed by a multi-layer perceptron (MLP) to output downstream classification predictions, trained end-to-end with the cross-entropy loss. The overall computational complexity of BrainNetTF is O(LMV2+KV)=O(V2)O(LMV^2 + KV) = O(V^2), matching standard graph neural networks for brain networks.

  2. Knowl 2 — Multi-Head Self-Attention over Connection Profiles

    model/method

    BrainNetTF uses the connection profile Xi⋅∈RVX_{i\cdot} \in \mathbb{R}^V (the ii-th row of the brain network correlation matrix XX) as the initial feature vector for node ii. Because the self-connection Xii=1X_{ii} = 1 naturally encodes the unique position of each ROI in a fixed-atlas brain graph, explicit positional encodings (such as Laplacian eigenvectors) and edge-weight embeddings are omitted.

    For an LL-layer network with MM attention heads, setting Z0=XZ^0 = X, each layer l∈{1,…,L}l \in \{1, \dots, L\} computes updated node embeddings Zl∈RV×VZ^l \in \mathbb{R}^{V \times V} via:

    Zl=(∥m=1Mhl,m)WOlZ^l = \left(\Vert_{m=1}^M h^{l,m}\right) W_O^l

    hl,m=Softmax(WQl,mZl−1(WKl,mZl−1)⊤dKl,m)WVl,mZl−1h^{l,m} = \text{Softmax}\left(\frac{W_Q^{l,m} Z^{l-1} (W_K^{l,m} Z^{l-1})^\top}{\sqrt{d_K^{l,m}}}\right) W_V^{l,m} Z^{l-1}

    where ∥\Vert denotes head concatenation, WOl∈RMdV×VW_O^l \in \mathbb{R}^{M d_V \times V} is the output projection parameter matrix, WQl,m,WKl,m∈RdKl,m×VW_Q^{l,m}, W_K^{l,m} \in \mathbb{R}^{d_K^{l,m} \times V} and WVl,m∈RdV×VW_V^{l,m} \in \mathbb{R}^{d_V \times V} are head-specific linear transformation matrices, and dKl,md_K^{l,m} is the projection dimension of the query and key vectors.

  3. Knowl 3 — Orthonormal Clustering Readout Operator

    model/method

    The Orthonormal Clustering Readout (OCREAD) is a global graph pooling function that aggregates node embeddings into functional module embeddings via soft clustering. Given KK cluster centers E∈RK×VE \in \mathbb{R}^{K \times V} with mutually orthonormal rows and the node embedding matrix ZL∈RV×VZ^L \in \mathbb{R}^{V \times V} from the final MHSA layer, OCREAD computes the soft assignment probability PikP_{ik} of assigning node ii to cluster kk using a Softmax projection:

    Pik=exp⁡(⟨Zi⋅L,Ek⋅⟩)∑k′=1Kexp⁡(⟨Zi⋅L,Ek′⋅⟩)P_{ik} = \frac{\exp(\langle Z^L_{i\cdot}, E_{k\cdot} \rangle)}{\sum_{k'=1}^K \exp(\langle Z^L_{i\cdot}, E_{k'\cdot} \rangle)}

    where ⟨⋅,⋅⟩\langle \cdot, \cdot \rangle denotes the inner product. Arranging the probabilities into an assignment matrix P∈RV×KP \in \mathbb{R}^{V \times K}, the graph-level embedding ZG∈RK×VZ_G \in \mathbb{R}^{K \times V} is computed by:

    ZG=P⊤ZLZ_G = P^\top Z^L

    This readout groups functionally similar brain regions in an unsupervised manner while preserving module-level distinctions.

  4. Knowl 4 — Gram-Schmidt Orthonormal Cluster Center Initialization

    algorithm

    To initialize the KK cluster centers in OCREAD, a random candidate matrix is generated and orthogonalized via the Gram-Schmidt process. This ensures initial cluster centers are mutually perpendicular unit vectors, preventing degenerate cluster assignments during unsupervised modular pooling.

    Input: Number of clusters KK, node feature dimension VV
    Output: Orthonormal cluster center matrix E∈RK×VE \in \mathbb{R}^{K \times V}
    Initialize candidate matrix C∈RK×VC \in \mathbb{R}^{K \times V} using Xavier uniform initialization
    for k=1k = 1 to KK do
        uk←Ck⋅−∑j=1k−1⟨uj,Ck⋅⟩⟨uj,uj⟩uju_k \leftarrow C_{k\cdot} - \sum_{j=1}^{k-1} \frac{\langle u_j, C_{k\cdot}\rangle}{\langle u_j, u_j\rangle} u_j
        Ek⋅←uk∥uk∥E_{k\cdot} \leftarrow \frac{u_k}{\|u_k\|}
    end for
    return EE
  5. Knowl 5 — Variance Maximization of Softmax Projections on Orthonormal Bases

    theoretical result

    Let Br={Z∈RV:∥Z∥≤r}B_r = \{Z \in \mathbb{R}^V : \|Z\| \le r\} denote the closed ball centered at the origin with radius r>0r > 0, and let VrV_r denote the volume of BrB_r. For cluster center vectors E=[E1⋅⊤,…,EK⋅⊤]⊤∈RK×VE = [E_{1\cdot}^\top, \dots, E_{K\cdot}^\top]^\top \in \mathbb{R}^{K \times V}, the average variance of the Softmax projection across BrB_r is defined as:

    1Vr∫Br∑k=1K(exp⁡(⟨Z,Ek⋅⟩)∑k′=1Kexp⁡(⟨Z,Ek′⋅⟩)−1K)2dZ\frac{1}{V_r} \int_{B_r} \sum_{k=1}^K \left( \frac{\exp(\langle Z, E_{k\cdot}\rangle)}{\sum_{k'=1}^K \exp(\langle Z, E_{k'\cdot}\rangle)} - \frac{1}{K} \right)^2 dZ

    This average variance attains its maximum value when the matrix of cluster centers EE is orthonormal. Maximizing this variance maximizes the discrepancy among cluster assignment probabilities, preventing embeddings of distinct functional modules from collapsing toward the uniform distribution Pˉ=1/K\bar{P} = 1/K.

  6. Knowl 6 — Hypothesis Testing and Error Minimization of Orthonormal Readout

    theoretical result

    Let PT(Zi⋅,Ek⋅)P_T(Z_{i\cdot}, E_{k\cdot}) be the true node-to-cluster assignment probability, modeled as the computed Softmax projection P(Zi⋅,Ek⋅)P(Z_{i\cdot}, E_{k\cdot}) perturbed by additive stochastic Gaussian error:

    PT(Zi⋅,Ek⋅)=P(Zi⋅,Ek⋅)+ϵi,ϵi∼N(0,σ2),E(ϵi)=0,Var(ϵi)=σ2P_T(Z_{i\cdot}, E_{k\cdot}) = P(Z_{i\cdot}, E_{k\cdot}) + \epsilon_i, \quad \epsilon_i \sim \mathcal{N}(0, \sigma^2), \quad \mathbb{E}(\epsilon_i) = 0, \quad \text{Var}(\epsilon_i) = \sigma^2

    Under this regression formulation, the variance inflation factor of the estimated readout probabilities is minimized when the cluster center vectors {Ek⋅}k=1K\{E_{k\cdot}\}_{k=1}^K are mutually orthogonal. Consequently, in statistical hypothesis testing, the significance level αEk⋅\alpha_{E_{k\cdot}} (which reflects the probability of falsely rejecting a well-estimated pooling configuration) is strictly lower when sampling from orthonormal cluster centers than from non-orthonormal centers.

  7. Knowl 7 — Classification Performance of BrainNetTF versus Baselines

    data/table

    BrainNetTF was evaluated against graph transformers (SAN, Graphormer, VanillaTF), fixed-network GNNs (BrainGNN, BrainGB, BrainNetCNN), and learnable-network GNNs (FBNETGEN, BrainNetGNN, DGM) on two fMRI datasets: ABIDE (Autism Spectrum Disorder diagnosis, 1009 subjects, Craddock 200 atlas) and ABCD (biological sex prediction, 7901 subjects, HCP 360 atlas). Values represent the mean ±\pm standard deviation over 5 random runs.

    Dataset: ABIDE Dataset: ABCD
    Type Method AUROC Accuracy Sensitivity Specificity AUROC Accuracy Sensitivity Specificity
    Graph Transformer SAN 71.3±\pm2.1 65.3±\pm2.9 55.4±\pm9.2 68.3±\pm7.5 90.1±\pm1.2 81.0±\pm1.3 84.9±\pm3.5 77.5±\pm4.1
    Graphormer 63.5±\pm3.7 60.8±\pm2.7 78.7±\pm22.3 36.7±\pm23.5 89.0±\pm1.4 80.2±\pm1.3 81.8±\pm11.6 82.4±\pm7.4
    VanillaTF 76.4±\pm1.2 65.2±\pm1.2 66.4±\pm11.4 71.1±\pm12.0 94.3±\pm0.7 85.9±\pm1.4 87.7±\pm2.4 82.6±\pm3.9
    Fixed Network BrainGNN 62.4±\pm3.5 59.4±\pm2.3 36.7±\pm24.0 70.7±\pm19.3 OOM OOM OOM OOM
    BrainGB 69.7±\pm3.3 63.6±\pm1.9 63.7±\pm8.3 60.4±\pm10.1 91.9±\pm0.3 83.1±\pm0.5 84.6±\pm4.3 81.5±\pm3.9
    BrainNetCNN 74.9±\pm2.4 67.8±\pm2.7 63.8±\pm9.7 71.0±\pm10.2 93.5±\pm0.3 85.7±\pm0.8 87.9±\pm3.4 83.0±\pm4.4
    Learnable Network FBNETGEN 75.6±\pm1.2 68.0±\pm1.4 64.7±\pm8.7 62.4±\pm9.2 94.5±\pm0.7 87.2±\pm1.2 87.0±\pm2.5 86.7±\pm2.8
    BrainNetGNN 55.3±\pm1.9 51.2±\pm5.4 67.7±\pm37.5 33.9±\pm34.2 75.3±\pm5.2 67.5±\pm4.7 67.7±\pm5.7 68.0±\pm6.5
    DGM 52.7±\pm3.8 60.7±\pm12.6 53.8±\pm41.2 51.1±\pm40.9 76.8±\pm19.0 68.6±\pm8.1 40.5±\pm29.7 95.6±\pm4.2
    Ours BRAINNETTF 80.2±\pm1.0 71.0±\pm1.2 72.5±\pm5.2 69.3±\pm6.5 96.2±\pm0.3 88.4±\pm0.4 89.4±\pm2.6 88.4±\pm1.5

    BrainNetTF achieves the best AUROC and accuracy across both datasets, outperforming the strongest baseline by 4.6% AUROC on ABIDE and 1.7% AUROC on ABCD (all performance gains passed two-sided t-tests with p<0.03p < 0.03).

  8. Knowl 8 — Readout Mechanism Ablation across Graph Transformers

    data/table

    The impact of the readout function was evaluated by substituting OCREAD with six standard graph pooling operations across three Transformer architectures (SAN, Graphormer, and VanillaTF) on the ABIDE and ABCD datasets. Performance is measured by AUROC (%).

    Dataset: ABIDE Dataset: ABCD
    Readout SAN Graphormer VanillaTF SAN Graphormer VanillaTF
    MEAN 63.7±\pm2.4 50.1±\pm1.1 73.4±\pm1.4 88.5±\pm0.9 87.6±\pm1.3 91.3±\pm0.7
    MAX 61.9±\pm2.5 54.5±\pm3.6 75.6±\pm1.4 87.4±\pm1.1 81.6±\pm0.8 94.4±\pm0.6
    SUM 62.0±\pm2.3 54.1±\pm1.3 70.3±\pm1.6 84.2±\pm0.8 71.5±\pm0.9 91.6±\pm0.6
    SortPooling 68.7±\pm2.3 51.3±\pm2.2 72.4±\pm1.3 84.6±\pm1.1 86.7±\pm1.0 89.9±\pm0.6
    DiffPool 57.4±\pm5.2 50.5±\pm4.7 62.9±\pm7.3 78.1±\pm1.5 70.0±\pm1.9 83.9±\pm1.3
    CONCAT 71.3±\pm2.1 63.5±\pm3.7 76.4±\pm1.2 90.1±\pm1.2 89.0±\pm1.4 94.3±\pm0.7
    OCREAD 70.6±\pm2.4 64.9±\pm2.7 80.2±\pm1.0 91.2±\pm0.7 90.2±\pm0.7 96.2±\pm0.4

    OCREAD consistently outperforms standard flat poolings (MEAN, MAX, SUM), hierarchical clustering (DiffPool), and sorting-based pooling across all architectures, yielding up to 3.8% AUROC improvement on VanillaTF.

  9. Knowl 9 — Impact of Cluster Count and Center Initialization on BrainNetTF

    empirical result

    The performance of BrainNetTF under different cluster initialization methods (Random, Learnable via gradient descent, and Orthonormal via Gram-Schmidt) and varying cluster numbers K∈{2,3,4,5,10,50,100}K \in \{2, 3, 4, 5, 10, 50, 100\} reveals three main characteristics:

    1. With orthonormal centers, performance increases as KK increases from 2 to 10 and then decreases as KK increases from 10 to 100. The optimal cluster range (K≤10K \le 10) aligns with neuroscience evidence that typical functional brain networks comprise fewer than 25 functional modules.
    2. Orthonormal initialization consistently yields lower standard deviations and higher AUROC than Random or Learnable initializations at small cluster counts (K≤10K \le 10).
    3. While Random, Learnable, and Orthonormal initializations converge toward comparable performance at very large KK (K≥50K \ge 50), larger KK substantially increases model parameters and computational overhead without improving accuracy.
  10. Knowl 10 — Site-Stratified Splitting Protocol for ABIDE

    experimental setup

    The ABIDE dataset contains rs-fMRI data pooled across 17 international scanning sites using varying scanner hardware and acquisition parameters. Simple random dataset splitting creates site distribution mismatches between validation and test partitions, resulting in unstable training curves and severe performance drops.

    To standardize evaluation, a site-stratified splitting protocol was designed: subjects within each acquisition site are partitioned into 70% training, 10% validation, and 20% testing sets while maintaining balanced positive (ASD) and negative (control) label distributions. This protocol stabilizes validation-to-test performance tracking and eliminates non-neural site-specific bias.

Coverage note — All primary methodological, theoretical, algorithmic, and empirical contributions of the paper have been represented as standalone knowls. Specific qualitative visualizations of attention maps and detailed baseline hyperparameter tables were summarized within the empirical knowls rather than extracted as independent knowls.

References

  1. 1.G. A. F. and C. J. Wild. Nonlinear Regression: Seber/Nonlinear Regression. 1989.
  2. 2.David Ahmedt-Aristizabal, Mohammad Ali Armin, Simon Denman, Clinton Fookes, and Lars Petersson. Graph-based deep learning for medical diagnosis and analysis: past, present and future. Sensors, 21:4758, 2021.
  3. 3.Teddy J Akiki and Chadi G Abdallah. Determining the hierarchical architecture of the human brain using subject-level clustering of functional networks. Scientific Reports, 9:1–15, 2019.
  4. 4.Parinaz Babaeeghazvini, Laura M. Rueda-Delgado, Jolien Gooijers, Stephan P. Swinnen, and Andreas Daffertshofer. Brain structural and functional connectivity: A review of combined works of diffusion magnetic resonance imaging and electro-encephalography. Frontiers in Human Neuroscience, 15, 2021.
  5. 5.Ed Bullmore and Olaf Sporns. Complex brain networks: graph theoretical analysis of structural and functional systems. Nature Reviews Neuroscience, 10:186–198, 2009.
  6. 6.Craddock Cameron, Benhajali Yassine, Chu Carlton, Chouinard Francois, E. Aykan Alan, Jakab András, Khundrakpam Budhachandra, Lewis John, Liub Qingyang, Milham Michael, Yan Chaogan, and Bellec Pierre. The neuro bureau preprocessing initiative: open sharing of preprocessed neuroimaging data and derivatives. Frontiers in Neuroinformatics, 2013.
  7. 7.Alfonso Caramazza and Max Coltheart. Cognitive neuropsychology twenty years on. Cognitive Neuropsychology, 2006.
  8. 8.B.J. Casey, Tariq Cannonier, and May I. Conley et al. The adolescent brain cognitive development (abcd) study: Imaging acquisition across 21 sites. Developmental Cognitive Neuroscience, 32:43–54, 2018.
  9. 9.Gemai Chen and N. Balakrishnan. A General Purpose Approximate Goodness-of-Fit Test. Journal of Quality Technology, 27:154–161, 1995.
  10. 10.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. In NeurIPS, 2021.
  11. 11.A. Colin Cameron and Frank A.G. Windmeijer. An R-squared measure of goodness of fit for some common nonlinear regression models. Journal of Econometrics, 77:329–342, 1997.
  12. 12.R Cameron Craddock, G Andrew James, Paul E Holtzheimer III, Xiaoping P Hu, and Helen S Mayberg. A whole brain fmri atlas generated via spatially constrained spectral clustering. Human Brain Mapping, 33:1914–1928, 2012.
  13. 13.Hejie Cui, Wei Dai, Yanqiao Zhu, Xuan Kan, Antonio Aodong Chen Gu, Joshua Lukemire, Liang Zhan, Lifang He, Ying Guo, and Carl Yang. Braingb: A benchmark for brain network analysis with graph neural networks, 2022.
  14. 14.Hejie Cui, Wei Dai, Yanqiao Zhu, Xiaoxiao Li, Lifang He, and Carl Yang. Interpretable graph neural networks for connectome-based brain disorder analysis. In MICCAI, 2022.
  15. 15.Hejie Cui, Zijie Lu, Pan Li, and Carl Yang. On positional and structural node features for graph neural networks on non-attributed graphs. CIKM, 2022.
  16. 16.Tian Dai, Ying Guo, Alzheimer’s Disease Neuroimaging Initiative, et al. Predicting individual brain functional connectivity using a bayesian hierarchical model. NeuroImage, 147:772–787, 2017.
  17. 17.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. In ACL, 2019.
  18. 18.Gustavo Deco, Viktor K. Jirsa, and Anthony R. McIntosh. Emerging concepts for the dynamical organization of resting-state activity in the brain. Nature Reviews Neuroscience, 12:43–56, 2011.
  19. 19.Morris H. DeGroot and Mark J. Schervish. Probability and Statistics. 4th ed edition, 2012.
  20. 20.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  21. 21.Vijay Prakash Dwivedi and Xavier Bresson. A generalization of transformer networks to graphs. AAAI Workshop on Deep Learning on Graphs: Methods and Applications, 2021.
  22. 22.Vijay Prakash Dwivedi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. Graph neural networks with learnable structural and positional representations. In ICLR, 2022.
  23. 23.Matthias Fey and Jan E. Lenssen. Fast graph representation learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds, 2019.
  24. 24.Matthew F. Glasser, Stamatios N. Sotiropoulos, J. Anthony Wilson, Timothy S. Coalson, Bruce Fischl, Jesper L. Andersson, Junqian Xu, Saad Jbabdi, Matthew Webster, Jonathan R. Polimeni, David C. Van Essen, and Mark Jenkinson. The minimal preprocessing pipelines for the human connectome project. NeuroImage, 80:105–124, 2013.
  25. 25.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In AISTATS, 2010.
  26. 26.Shupeng Gui, Xiangliang Zhang, Pan Zhong, Shuang Qiu, Mingrui Wu, Jieping Ye, Zhengdao Wang, and Ji Liu. Pine: Universal deep embedding for graph nodes via partial permutation invariant set functions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44:770–782, 2022.
  27. 27.Ying Guo and Giuseppe Pagnoni. A unified framework for group independent component analysis for multi-subject fmri data. NeuroImage, 42(3):1078–1093, 2008.
  28. 28.Ixavier A Higgins, Suprateek Kundu, Ki Sueng Choi, Helen S Mayberg, and Ying Guo. A difference degree test for comparing brain networks. Human brain mapping, pages 4518–4536, 2019.
  29. 29.Ixavier A Higgins, Suprateek Kundu, and Ying Guo. Integrative bayesian analysis of brain functional networks incorporating anatomical knowledge. Neuroimage, 181:263–278, 2018.
  30. 30.Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. Open graph benchmark: Datasets for machine learning on graphs. In NeurIPS, 2020.
  31. 31.Yingtian Hu, Mahmoud Zeydabadinezhad, Longchuan Li, and Ying Guo. A multimodal multilevel neuroimaging model for investigating brain connectome development. Journal of the American Statistical Association, pages 1–15, 2022.
  32. 32.Ziniu Hu, Yuxiao Dong, Kuansan Wang, and Yizhou Sun. Heterogeneous graph transformer. In WWW, pages 2704–2710, 2020.
  33. 33.Md Shamim Hussain, Mohammed J Zaki, and Dharmashankar Subramanian. Edge-augmented graph transformers: Global self-attention is enough for graphs. arXiv, 2021.
  34. 34.Mingxuan Ju, Shifu Hou, Yujie Fan, Jianan Zhao, Liang Zhao, and Yanfang Ye. Adaptive kernel graph neural network. AAAI, 2022.
  35. 35.Xuan Kan, Hejie Cui, Joshua Lukemire, Ying Guo, and Carl Yang. FBNETGEN: Task-aware GNN-based fMRI analysis via functional brain network generation. In Medical Imaging with Deep Learning, 2022.
  36. 36.Minoru Kanehisa and Susumu Goto. Kegg: kyoto encyclopedia of genes and genomes. Nucleic Acids Research, 28:27–30, 2000.
  37. 37.Jeremy Kawahara, Colin J. Brown, Steven P. Miller, Brian G. Booth, Vann Chau, Ruth E. Grunau, Jill G. Zwicker, and Ghassan Hamarneh. BrainNetCNN: Convolutional neural networks for brain networks; towards predicting neurodevelopment. NeuroImage, 146:1038–1049, 2017.
  38. 38.Anees Kazi, Luca Cosmo, Seyed-Ahmad Ahmadi, Nassir Navab, and Michael Bronstein. Differentiable graph module (dgm) for graph convolutional networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 1–1, 2022.
  39. 39.Byung-Hoon Kim, Jong Chul Ye, and Jae-Jin Kim. Learning dynamic graph representation of brain connectome with spatio-temporal attention. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, NeurIPS, 2021.
  40. 40.Devin Kreuzer, Dominique Beaini, William L. Hamilton, Vincent Létourneau, and Prudencio Tossou. Rethinking graph transformers with spectral attention. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021.
  41. 41.Suprateek Kundu, Joshua Lukemire, Yikai Wang, and Ying Guo. A novel joint brain network analysis using longitudinal alzheimer’s disease data. Scientific reports, 9(1):1–18, 2019.
  42. 42.Xiaoxiao Li, Nicha C. Dvornek, Yuan Zhou, Juntang Zhuang, Pamela Ventola, and James S. Duncan. Graph neural network for interpreting task-fmri biomarkers. In MICCAI, 2019.
  43. 43.Xiaoxiao Li, Yuan Zhou, Siyuan Gao, Nicha Dvornek, Muhan Zhang, Juntang Zhuang, Shi Gu, Dustin Scheinost, Lawrence Staib, Pamela Ventola, et al. Braingnn: Interpretable brain graph neural network for fmri analysis. Medical Image Analysis, 2021.
  44. 44.Joshua Lukemire, Suprateek Kundu, Giuseppe Pagnoni, and Ying Guo. Bayesian joint modeling of multiple brain functional networks. Journal of the American Statistical Association, 116:518–530, 2021.
  45. 45.Usman Mahmood, Zening Fu, Vince D. Calhoun, and Sergey Plis. A deep learning model for data-driven discovery of functional connectivity. Algorithms, 14, 2021.
  46. 46.Haitao Mao, Xu Chen, Qiang Fu, Lun Du, Shi Han, and Dongmei Zhang. Neuron Campaign for Initialization Guided by Information Bottleneck Theory. 2021.
  47. 47.F. H. C. Marriott, J. Neter, W. Wasserman, and M. H. Kutner. Applied Linear Regression Models. Biometrics, 1985.
  48. 48.Jaina Mistry, Sara Chuguransky, Lowri Williams, Matloob Qureshi, Gustavo A Salazar, Erik L L Sonnhammer, Silvio C E Tosatto, Lisanna Paladin, Shriya Raj, Lorna J Richardson, Robert D Finn, and Alex Bateman. Pfam: The protein families database in 2021. Nucleic Acids Research, 49:D412–D419, 2020.
  49. 49.Wonpyo Park, Woonggi Chang, Donggeon Lee, Juntae Kim, and Seung-won Hwang. Grpe: Relative positional encoding for graph transformer, 2022.
  50. 50.Theodore D. Satterthwaite, Daniel H. Wolf, David R. Roalf, Kosha Ruparel, Guray Erus, Simon Vandekar, Efstathios D. Gennatas, Mark A. Elliott, Alex Smith, Hakon Hakonarson, Ragini Verma, Christos Davatzikos, Raquel E. Gur, and Ruben C. Gur. Linked Sex Differences in Cognition and Functional Connectivity in Youth. Cerebral Cortex, 25:2383–2394, 2015.
  51. 51.Andrew M. Saxe, James L. McClelland, and Surya Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. 2014.
  52. 52.Ran Shi and Ying Guo. Investigating differences in brain functional networks using hierarchical covariate-adjusted independent component analysis. The annals of applied statistics, page 1930, 2016.
  53. 53.Sean L Simpson, F DuBois Bowman, and Paul J Laurienti. Analyzing complex functional brain networks: fusing statistics and network science to understand the brain. Statistics Surveys, 7:1, 2013.
  54. 54.Stephen M. Smith, Karla L. Miller, Gholamreza Salimi-Khorshidi, Matthew Webster, Christian F. Beckmann, Thomas E. Nichols, Joseph D. Ramsey, and Mark W. Woolrich. Network modelling methods for FMRI. NeuroImage, 54, 2011.
  55. 55.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxvit: Multi-axis vision transformer. ECCV, 2022.
  56. 56.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, NeurIPS, 2017.
  57. 57.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. In ICLR, 2018.
  58. 58.Yikai Wang and Ying Guo. A hierarchical independent component analysis model for longitudinal neuroimaging studies. NeuroImage, 189:380–400, 2019.
  59. 59.Yikai Wang, Jian Kang, Phebe B. Kemmer, and Ying Guo. An Efficient and Reliable Statistical Method for Estimating Functional Connectivity in Large Scale Brain Networks Using Partial Correlation. Frontiers in Neuroscience, 10:123, 2016.
  60. 60.Junyuan Xie, Ross Girshick, and Ali Farhadi. Unsupervised deep embedding for clustering analysis. In ICML, 2016.
  61. 61.Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks? In ICLR, 2019.
  62. 62.Yujun Yan, Jiong Zhu, Marlena Duda, Eric Solarz, Chandra Sripada, and Danai Koutra. Groupinn: Grouping-based interpretable neural network-based classification of limited, noisy brain data. In KDD, 2019.
  63. 63.Yi Yang, Yanqiao Zhu, Hejie Cui, Xuan Kan, Lifang He, Ying Guo, and Carl Yang. Dataefficient brain connectome analysis via multi-task meta-learning. KDD, 2022.
  64. 64.Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. Do transformers really perform badly for graph representation? In NeurIPS, 2021.
  65. 65.Muhan Zhang and Yixin Chen. Link prediction based on graph neural networks. In NeurIPS, 2018.
  66. 66.Muhan Zhang, Zhicheng Cui, Marion Neumann, and Yixin Chen. An end-to-end deep learning architecture for graph classification. In AAAI, 2018.
  67. 67.Muhan Zhang and Pan Li. Nested graph neural networks. In NeurIPS, 2021.
  68. 68.Jianan Zhao, Chaozhuo Li, Qianlong Wen, Yiqi Wang, Yuming Liu, Hao Sun, Xing Xie, and Yanfang Ye. Gophormer: Ego-graph transformer for node classification, 2021.
  69. 69.Yanqiao Zhu, Hejie Cui, Lifang He, Lichao Sun, and Carl Yang. Joint embedding of structural and functional brain networks with graph neural networks for mental illness diagnosis. In 2022 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), pages 272–276, 2022.

Citation

MLA
Kan, X., et al. “Brain Network Transformer”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 25586–99, https://proceedings.neurips.cc/paper_files/paper/2022/file/a408234a9b80604a9cf6ca518e474550-Paper-Conference.pdf.
APA
Kan, X., Dai, W., Cui, H., Zhang, Z., Guo, Y., & Yang, C. (2022). Brain Network Transformer. Advances in Neural Information Processing Systems, 35, 25586–25599. https://proceedings.neurips.cc/paper_files/paper/2022/file/a408234a9b80604a9cf6ca518e474550-Paper-Conference.pdf
Chicago
Kan, X., W. Dai, H. Cui, Z. Zhang, Y. Guo, and C. Yang. 2022. “Brain Network Transformer”. Advances in Neural Information Processing Systems 35: 25586–99. https://proceedings.neurips.cc/paper_files/paper/2022/file/a408234a9b80604a9cf6ca518e474550-Paper-Conference.pdf.
Harvard
Kan, X. et al. (2022) “Brain Network Transformer”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 25586–25599. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/a408234a9b80604a9cf6ca518e474550-Paper-Conference.pdf.
Vancouver
1. Kan X, Dai W, Cui H, Zhang Z, Guo Y, Yang C (2022) Brain Network Transformer. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 25586–25599

BibTeX

@inproceedings{kan2022brain,
  title = {Brain Network Transformer},
  author = {Kan, Xuan and Dai, Wei and Cui, Hejie and Zhang, Zilong and Guo, Ying and Yang, Carl},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {25586-25599},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/a408234a9b80604a9cf6ca518e474550-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors