Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning

Xiaoxin HeXavier BressonThomas LaurentAdam PeroldYann LeCunBryan Hooi

article2024ICLR142 citations

Proposes an LLM-to-LM interpreter that converts zero-shot large language model explanations into informative node features for graph neural networks, achieving state-of-the-art accuracy and a near-threefold training speedup on text-attributed graph benchmarks.

Listen

Real-world data networks, such as academic citation webs and e-commerce product catalogs, frequently combine complex relational structures with rich textual information. While graph neural networks (GNNs) excel at modeling network connectivity, they traditionally rely on shallow text embeddings or complex, computationally intensive language model integrations that cannot easily leverage the sophisticated reasoning of modern large language models (LLMs). This computational bottleneck limits the ability of organizations to harness high-level textual reasoning for network classification tasks.

The article demonstrates a novel framework called TAPE (Title, Abstract, Prediction, and Explanation) that integrates the broad reasoning of LLMs with the structural learning of GNNs. The primary objective is to evaluate whether LLM-generated explanations and ranked predictions can be translated into compact, high-performance features for downstream graph classification using standard, budget-friendly infrastructure.

The authors evaluated this approach across five benchmark datasets, including citation networks (Cora, PubMed, ogbn-arxiv, and the newly introduced tape-arxiv23) and an e-commerce network (ogbn-products). The method queries an LLM—such as GPT-3.5 or open-source Llama-2—in a zero-shot manner to generate predictions along with natural language rationales. A smaller language model (such as DeBERTa) is then fine-tuned to interpret these texts and produce fixed vectorial representations. These features are kept frozen while training standard GNN architectures (such as GCN, GraphSAGE, and RevGAT), entirely decoupling text encoding from graph training.

The evaluation yielded several key findings. First, the framework established new state-of-the-art accuracy across all tested benchmarks, achieving 77.50% test accuracy on the ogbn-arxiv benchmark and outperforming existing iterative language-graph models. Second, decoupling feature extraction from graph learning reduced total training computation time by 2.88 times compared to the closest state-of-the-art baseline (GLEM). Third, the approach maintained strong generalization on contemporary data (tape-arxiv23), reaching 84.23% accuracy on papers published after the LLM's training cutoff. Finally, using open-source models like Llama-2 delivered competitive performance (76.19% on ogbn-arxiv), confirming that organizations can avoid proprietary API costs without substantial performance degradation.

These findings indicate that organizations can significantly enhance classification accuracy in relational text systems while lowering computational expense and development overhead. Because the framework operates via standard application programming interfaces (APIs) and freezes extracted features prior to graph training, it avoids expensive end-to-end retraining and scales efficiently within standard hardware budgets. The results demonstrate that generating intermediate text explanations captures vital contextual nuances that raw text alone fails to convey to smaller models.

Decision-makers should consider adopting this decoupled interpreter architecture when upgrading graph analytics pipelines, particularly for document routing, recommendation systems, or knowledge management. Teams can balance budget constraints by choosing between hosted commercial LLM APIs and free open-source models depending on latency and hosting capabilities. Future technical work should focus on automating dataset-specific prompt generation and evaluating the pipeline on dynamic, continuously evolving graphs.

The primary limitation of the study is its reliance on manually engineered prompts, which can cause slight fluctuations in extraction quality across different domains. Nonetheless, extensive ablation experiments and consistent results across multiple models provide strong confidence in the stability, modularity, and practical value of the proposed architecture.

Cover for Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning

Abstract

Representation learning on text-attributed graphs (TAGs) has become a critical research problem in recent years. A typical example of a TAG is a paper citation graph, where the text of each paper serves as node attributes. Initial graph neural network (GNN) pipelines handled these text attributes by transforming them into shallow or hand-crafted features, such as skip-gram or bag-of-words features. Recent efforts have focused on enhancing these pipelines with language models (LMs), which typically demand intricate designs and substantial computational resources. With the advent of powerful large language models (LLMs) such as GPT or Llama2, which demonstrate an ability to reason and to utilize general knowledge, there is a growing need for techniques which combine the textual modelling abilities of LLMs with the structural learning capabilities of GNNs. Hence, in this work, we focus on leveraging LLMs to capture textual information as features, which can be used to boost GNN performance on downstream tasks. A key innovation is our use of explanations as features: we prompt an LLM to perform zero-shot classification, request textual explanations for its decision-making process, and design an LLM-to-LM interpreter to translate these explanations into informative features for downstream GNNs. Our experiments demonstrate that our method achieves state-of-the-art results on well-established TAG datasets, including Cora, PubMed, ogbn-arxiv, as well as our newly introduced dataset, tape-arxiv23. Furthermore, our method significantly speeds up training, achieving a 2.88 times improvement over the closest baseline on ogbn-arxiv. Lastly, we believe the versatility of the proposed method extends beyond TAGs and holds the potential to enhance other tasks involving graph-text data. Our codes and datasets are available at: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Formalization
  • 4 Proposed Method
  • 4.1 Generating Predictions and Explanations with LLMs
  • 4.2 Fine-Tuning LM Interpreter and Node Feature Extraction
  • 4.3 GNN Training on Enriched Features
  • 4.4 Theoretical Analysis
  • 5 Experiments
  • 5.1 Main Results
  • 5.2 Scalability
  • 5.3 Ablation Study
  • 6 Conclusion
  • References
  • A Theoretical Analysis
  • B time analysis and money estimation
  • C Addressing Label Leakage Concerns with a New Dataset
  • D Llama as a cost-efficient alternative
  • E Case Study
  • F Prompt Design
  • G Dataset
  • G.1 Dataset Description
  • G.2 Dataset splits and random seeds
  • G.3 Shallow Embedding Methods for Node Feature Extraction
  • H Experiment Details
  • H.1 Computing Environment and Resources
  • H.2 Hyperparameters
  • H.3 Detailed Ablation Study
  • I Effect of LM Finetuning
  • J Effect of different LMs
  • K Memory Utilization
  • L GLEM

Knowls

  1. Knowl 1 — TAPE Framework for Text-Attributed Graph Representation Learning

    model/method

    The TAPE (Title, Abstract, Prediction, and Explanation) framework improves representation learning on text-attributed graphs (TAGs) by using a Large Language Model (LLM) to generate semantic predictions and natural language explanations, followed by a smaller fine-tuned Language Model (LM) that acts as an interpreter for downstream Graph Neural Networks (GNNs).

    Formally, a TAG is defined as G=(V,A,{sn}n∈V)G = (V, A, \{s_n\}_{n \in V}), where VV denotes the set of NN nodes, A∈RN×NA \in \mathbb{R}^{N \times N} is the adjacency matrix, and sn∈DLns_n \in \mathcal{D}^{L_n} is the text sequence (e.g., paper title and abstract) of length LnL_n associated with node nn from token vocabulary D\mathcal{D}.

    TAPE executes across three decoupled stages:

    1. LLM Querying (LMaaS-compatible): For each node i∈Vi \in V, its text sis_i is formatted into a prompt querying a frozen LLM (e.g., GPT-3.5 or LLaMA-2) via text input/output only. The LLM outputs a top-kk ranked class prediction list and a natural language explanation describing its reasoning.
    2. LM Interpreter Fine-Tuning and Feature Extraction: A smaller pre-trained language model (e.g., DeBERTa) is fine-tuned on the labeled training nodes to interpret the original text sorigs^{\text{orig}} and the LLM-generated explanations sexpls^{\text{expl}}, producing dense feature vectors horigh_{\text{orig}} and hexplh_{\text{expl}}. The top-kk predictions are one-hot encoded and linearly transformed into hpredh_{\text{pred}}.
    3. Decoupled GNN Training: The enriched node features hTAPE={horig,hexpl,hpred}h_{\text{TAPE}} = \{h_{\text{orig}}, h_{\text{expl}}, h_{\text{pred}}\} are frozen and used to train independent GNN classifiers whose predictions are fused via simple averaging, completely eliminating the LM from the graph training loop.
  2. Knowl 2 — Theoretical Condition for Information Gain from LLM Explanations

    theoretical result

    Let yy be the target node classification label, Z∈RdZ \in \mathbb{R}^d be the vectorial text representation generated by a small LM, ZL∈RdLZ_L \in \mathbb{R}^{d_L} be the internal representation of the text modeled by an LLM, and EE be the textual explanation generated by the LLM. Let H(⋅∣⋅)H(\cdot \mid \cdot) denote conditional entropy.

    Assume the following two conditions hold:

    1. Fidelity: The explanation EE captures the LLM representation ZLZ_L such that: H(ZL∣E)=ϵ,with ϵ>0H(Z_L \mid E) = \epsilon, \quad \text{with } \epsilon > 0

    2. Non-redundancy: The LLM representation ZLZ_L contains predictive information about yy not captured in ZZ, such that: H(y∣Z,ZL)=H(y∣Z)−ϵ′,with ϵ′>ϵH(y \mid Z, Z_L) = H(y \mid Z) - \epsilon', \quad \text{with } \epsilon' > \epsilon

    Under these conditions, conditioning on both the small LM representation ZZ and the LLM explanation EE strictly reduces the conditional entropy of label yy compared to conditioning on ZZ alone: H(y∣Z,E)<H(y∣Z)H(y \mid Z, E) < H(y \mid Z)

    This result guarantees that faithful textual explanations from an LLM possessing non-redundant knowledge strictly decrease label uncertainty when provided to a smaller LM.

  3. Knowl 3 — Node Feature Extraction and Ensemble GNN Inference in TAPE

    model/method

    TAPE constructs three distinct node feature representations for each node n∈Vn \in V from the original text sorigs^{\text{orig}}, generated textual explanation sexpls^{\text{expl}}, and ranked predictions:

    1. Text and Explanation Encodings: Two pre-trained language models, LMorig\text{LM}_{\text{orig}} and LMexpl\text{LM}_{\text{expl}}, encode the original text and explanation texts into embeddings: horig=LMorig(sorig)∈RN×d,hexpl=LMexpl(sexpl)∈RN×dh_{\text{orig}} = \text{LM}_{\text{orig}}(s^{\text{orig}}) \in \mathbb{R}^{N \times d}, \quad h_{\text{expl}} = \text{LM}_{\text{expl}}(s^{\text{expl}}) \in \mathbb{R}^{N \times d} where dd is the LM hidden dimension. The LMs and accompanying classification heads MLPorig\text{MLP}_{\text{orig}} and MLPexpl\text{MLP}_{\text{expl}} are fine-tuned on the labeled training set L⊂V\mathcal{L} \subset V using cross-entropy loss over the predicted logits: yorig=MLPorig(horig)∈RN×C,yexpl=MLPexpl(hexpl)∈RN×Cy_{\text{orig}} = \text{MLP}_{\text{orig}}(h_{\text{orig}}) \in \mathbb{R}^{N \times C}, \quad y_{\text{expl}} = \text{MLP}_{\text{expl}}(h_{\text{expl}}) \in \mathbb{R}^{N \times C} for CC target classes.

    2. Ranked Prediction Features: The LLM's top-kk ranked category predictions for node ii are one-hot encoded as pi,1,…,pi,k∈{0,1}Cp_{i,1}, \dots, p_{i,k} \in \{0, 1\}^C, concatenated into a kCkC-dimensional vector, and mapped through a linear projection layer to dimension dPd_P, yielding prediction matrix hpred∈RN×dPh_{\text{pred}} \in \mathbb{R}^{N \times d_P}.

    3. GNN Ensemble Inference: Three independent GNNs (forig,fexpl,fpredf_{\text{orig}}, f_{\text{expl}}, f_{\text{pred}}) are trained on the frozen features along with graph adjacency AA: y^source=fsource(hsource,A)∈RN×C,for source∈{orig,expl,pred}\hat{y}_{\text{source}} = f_{\text{source}}(h_{\text{source}}, A) \in \mathbb{R}^{N \times C}, \quad \text{for } \text{source} \in \{\text{orig}, \text{expl}, \text{pred}\}

    The final prediction y^\hat{y} is obtained via arithmetic averaging: y^=mean(y^orig,y^expl,y^pred)∈RN×C\hat{y} = \text{mean}(\hat{y}_{\text{orig}}, \hat{y}_{\text{expl}}, \hat{y}_{\text{pred}}) \in \mathbb{R}^{N \times C}

  4. Knowl 4 — Node Classification Performance of TAPE Across TAG Benchmarks

    data/table

    The node classification accuracy of TAPE was evaluated across five text-attributed graph datasets using four downstream classifiers: Multi-Layer Perceptron (MLP), Graph Convolutional Network (GCN), GraphSAGE (SAGE), and Reversible Graph Attention Network (RevGAT). TAPE features (hTAPEh_{\text{TAPE}}) were compared against shallow features (hshallowh_{\text{shallow}}), GIANT features (hGIANTh_{\text{GIANT}}), zero-shot GPT-3.5 (LLM), and DeBERTa-base fine-tuned on labeled nodes (LMfinetune\text{LM}_{\text{finetune}}). Results represent the mean and standard deviation over 4 runs:

    Dataset Method GNN LM Ours
    hshallowh_{\text{shallow}} hGIANTh_{\text{GIANT}} G↑G \uparrow LLM LMfinetune\text{LM}_{\text{finetune}} hTAPEh_{\text{TAPE}}
    Cora MLP 0.6388 ±\pm 0.0213 0.7196 ±\pm 0.0000 37.41% 0.6769 0.7606 ±\pm 0.0378 0.8778 ±\pm 0.0485
    GCN 0.8911 ±\pm 0.0015 0.8423 ±\pm 0.0053 2.33% 0.6769 0.7606 ±\pm 0.0378 0.9119 ±\pm 0.0158
    SAGE 0.8824 ±\pm 0.0009 0.8455 ±\pm 0.0028 5.28% 0.6769 0.7606 ±\pm 0.0378 0.9290 ±\pm 0.0307
    RevGAT 0.8911 ±\pm 0.0000 0.8353 ±\pm 0.0038 4.14% 0.6769 0.7606 ±\pm 0.0378 0.9280 ±\pm 0.0275
    PubMed MLP 0.8635 ±\pm 0.0032 0.8175 ±\pm 0.0059 10.77% 0.9342 0.9494 ±\pm 0.0046 0.9565 ±\pm 0.0060
    GCN 0.8031 ±\pm 0.0425 0.8419 ±\pm 0.0050 17.43% 0.9342 0.9494 ±\pm 0.0046 0.9431 ±\pm 0.0043
    SAGE 0.8881 ±\pm 0.0002 0.8372 ±\pm 0.0082 8.30% 0.9342 0.9494 ±\pm 0.0046 0.9618 ±\pm 0.0053
    RevGAT 0.8850 ±\pm 0.0005 0.8502 ±\pm 0.0048 8.52% 0.9342 0.9494 ±\pm 0.0046 0.9604 ±\pm 0.0047
    ogbn-arxiv MLP 0.5336 ±\pm 0.0038 0.7308 ±\pm 0.0006 42.19% 0.7350 0.7361 ±\pm 0.0004 0.7587 ±\pm 0.0015
    GCN 0.7182 ±\pm 0.0027 0.7329 ±\pm 0.0010 4.71% 0.7350 0.7361 ±\pm 0.0004 0.7520 ±\pm 0.0003
    SAGE 0.7171 ±\pm 0.0017 0.7435 ±\pm 0.0014 6.98% 0.7350 0.7361 ±\pm 0.0004 0.7672 ±\pm 0.0007
    RevGAT 0.7083 ±\pm 0.0017 0.7590 ±\pm 0.0019 9.42% 0.7350 0.7361 ±\pm 0.0004 0.7750 ±\pm 0.0012
    ogbn-products MLP 0.5385 ±\pm 0.0017 0.6125 ±\pm 0.0078 46.30% 0.7440 0.7297 ±\pm 0.0023 0.7878 ±\pm 0.0082
    (subset) GCN 0.7052 ±\pm 0.0051 0.6977 ±\pm 0.0042 13.39% 0.7440 0.7297 ±\pm 0.0023 0.7996 ±\pm 0.0041
    SAGE 0.6913 ±\pm 0.0026 0.6869 ±\pm 0.0119 17.71% 0.7440 0.7297 ±\pm 0.0023 0.8137 ±\pm 0.0043
    RevGAT 0.6964 ±\pm 0.0017 0.7189 ±\pm 0.0030 18.24% 0.7440 0.7297 ±\pm 0.0023 0.8234 ±\pm 0.0036
    tape-arxiv23 MLP 0.6202 ±\pm 0.0064 0.5574 ±\pm 0.0032 35.20% 0.7356 0.7358 ±\pm 0.0006 0.8385 ±\pm 0.0246
    GCN 0.6341 ±\pm 0.0062 0.5672 ±\pm 0.0061 27.42% 0.7356 0.7358 ±\pm 0.0006 0.8080 ±\pm 0.0215
    SAGE 0.6430 ±\pm 0.0037 0.5665 ±\pm 0.0032 30.45% 0.7356 0.7358 ±\pm 0.0006 0.8388 ±\pm 0.0264
    RevGAT 0.6563 ±\pm 0.0062 0.5834 ±\pm 0.0038 28.34% 0.7356 0.7358 ±\pm 0.0006 0.8423 ±\pm 0.0256

    TAPE consistently outperforms all baselines across all datasets and GNN backbones, achieving top accuracies on ogbn-arxiv (77.50%), Cora (92.90% with SAGE, 92.80% with RevGAT), PubMed (96.18%), ogbn-products subset (82.34%), and tape-arxiv23 (84.23%).

  5. Knowl 5 — tape-arxiv23 Dataset for Leakage-Free TAG Benchmarking

    definition

    The tape-arxiv23 dataset is a text-attributed citation network constructed specifically to evaluate representation learning without the risk of LLM training data memorization or label leakage.

    • Temporal Window: Contains papers published in the Computer Science category on arXiv between January 2023 and September 2023, strictly after the pre-training knowledge cutoff of GPT-3.5 (November 2022).
    • Graph Structure: Nodes represent CS arXiv papers, and directed edges represent citation relationships retrieved via the Semantic Scholar API.
    • Scale: Comprises 46,198 nodes and 78,548 directed edges.
    • Task and Label Space: 40-class multi-class node classification, predicting the arXiv CS primary subject areas (e.g., cs.AI, cs.CV, cs.LG, cs.OS).
    • Splits: Standardized random partition of 60% training, 20% validation, and 20% testing.
  6. Knowl 6 — Computational Efficiency and Memory Comparison: TAPE vs. GLEM

    data/table

    On the ogbn-arxiv dataset (169,343 nodes, 1,166,243 edges), using DeBERTa-base (139M parameters) as the LM backbone and RevGAT (1.8M parameters) as the GNN backbone on 4 NVIDIA RTX A5000 24GB GPUs with batch size 36, TAPE achieves superior accuracy while reducing total training time by a factor of 2.88×\times compared to the iterative GLEM method:

    Method Val Acc. Test Acc. Params Max bsz. Total Time
    LMorig\text{LM}_{\text{orig}} 0.7503 ±\pm 0.0008 0.7361 ±\pm 0.0004 139,223,080 36 1.73h
    GNN-hshallow\text{GNN-}h_{\text{shallow}} 0.7144 ±\pm 0.0021 0.7083 ±\pm 0.0017 427,728 all nodes 1.80min
    GLEM-G-Step 0.7761 ±\pm 0.0005 0.7657 ±\pm 0.0029 1,837,136 all nodes 9.18h
    GLEM-L-Step 0.7548 ±\pm 0.0039 0.7495 ±\pm 0.0037 138,632,488 36
    TAPE-LMorig\text{LM}_{\text{orig}}-Step 0.7503 ±\pm 0.0008 0.7361 ±\pm 0.0004 139,223,080 36 1.73h
    TAPE-LMexpl\text{LM}_{\text{expl}}-Step 0.7506 ±\pm 0.0008 0.7432 ±\pm 0.0012 139,223,080 36 1.40h
    TAPE-GNN-hTAPE\text{GNN-}h_{\text{TAPE}}-Step 0.7785 ±\pm 0.0016 0.7750 ±\pm 0.0012 1,837,136 all nodes 3.76min

    Total wall-clock training time is 192 minutes for TAPE (3.13h sequential, or 1.73h if LM training is parallelized) compared to 551 minutes (9.18h) for GLEM.

    Memory Footprint: During the GNN phase, TAPE requires 4,430 MB of GPU memory (identical to training on shallow features), whereas GLEM requires 8,112 MB for the GNN step and 11,064 MB for the LM step.

  7. Knowl 7 — Evaluation of Open-Source LLaMA-2 as the LLM Generator in TAPE

    data/table

    To evaluate open-source, cost-free local models as alternatives to proprietary APIs, TAPE was evaluated using llama-2-13b-chat in place of GPT-3.5 across four datasets. Results reflect classification accuracy over 4 runs:

    Dataset Method llama2-13b-chat GPT-3.5
    LLM LMfinetune\text{LM}_{\text{finetune}} hTAPEh_{\text{TAPE}} LLM LMfinetune\text{LM}_{\text{finetune}} hTAPEh_{\text{TAPE}}
    Cora GCN 0.5746 0.6845 ±\pm 0.0194 0.9045 ±\pm 0.0231 0.6769 0.7606 ±\pm 0.0378 0.9119 ±\pm 0.0158
    SAGE 0.5746 0.6845 ±\pm 0.0194 0.9170 ±\pm 0.0337 0.6769 0.7606 ±\pm 0.0378 0.9290 ±\pm 0.0307
    RevGAT 0.5746 0.6845 ±\pm 0.0194 0.9313 ±\pm 0.0237 0.6769 0.7606 ±\pm 0.0378 0.9280 ±\pm 0.0275
    PubMed GCN 0.3958 0.9121 ±\pm 0.0026 0.9362 ±\pm 0.0050 0.9342 0.9494 ±\pm 0.0046 0.9431 ±\pm 0.0043
    SAGE 0.3958 0.9121 ±\pm 0.0026 0.9581 ±\pm 0.0073 0.9342 0.9494 ±\pm 0.0046 0.9618 ±\pm 0.0053
    RevGAT 0.3958 0.9121 ±\pm 0.0026 0.9561 ±\pm 0.0068 0.9342 0.9494 ±\pm 0.0046 0.9604 ±\pm 0.0047
    ogbn-arxiv GCN 0.4423 0.6941 ±\pm 0.0020 0.7418 ±\pm 0.0031 0.7350 0.7361 ±\pm 0.0004 0.7520 ±\pm 0.0003
    SAGE 0.4423 0.6941 ±\pm 0.0020 0.7536 ±\pm 0.0028 0.7350 0.7361 ±\pm 0.0004 0.7672 ±\pm 0.0007
    RevGAT 0.4423 0.6941 ±\pm 0.0020 0.7619 ±\pm 0.0027 0.7350 0.7361 ±\pm 0.0004 0.7750 ±\pm 0.0012
    tape-arxiv23 GCN 0.4452 0.7677 ±\pm 0.0042 0.8045 ±\pm 0.0264 0.7356 0.7832 ±\pm 0.0052 0.8080 ±\pm 0.0215
    SAGE 0.4452 0.7677 ±\pm 0.0042 0.8378 ±\pm 0.0302 0.7356 0.7832 ±\pm 0.0052 0.8388 ±\pm 0.0264
    RevGAT 0.4452 0.7677 ±\pm 0.0042 0.8407 ±\pm 0.0308 0.7356 0.7832 ±\pm 0.0052 0.8423 ±\pm 0.0256

    Even though LLaMA-2-13b-chat achieves significantly lower zero-shot accuracy than GPT-3.5 (e.g., 44.23% vs. 73.50% on ogbn-arxiv), the TAPE interpreter transforms its generated explanations and predictions into highly effective representations, reaching 76.19% on ogbn-arxiv and 93.13% on Cora with RevGAT.

  8. Knowl 8 — Ablation of Feature Modalities and LM Fine-Tuning Protocols

    empirical result

    Ablation experiments on the ogbn-arxiv dataset demonstrate the individual impact of each feature modality and the necessity of task-specific LM fine-tuning:

    1. Feature Component Ablation (Test Accuracy with RevGAT):

      • Full hTAPEh_{\text{TAPE}} features: 0.7750±0.00120.7750 \pm 0.0012
      • Without original text features (−horig-h_{\text{orig}}): 0.7656±0.00380.7656 \pm 0.0038
      • Without explanation features (−hexpl-h_{\text{expl}}): 0.7693±0.00330.7693 \pm 0.0033
      • Without prediction features (−hpred-h_{\text{pred}}): 0.7686±0.00510.7686 \pm 0.0051
      • Individual feature sets alone achieve 0.7504±0.00200.7504 \pm 0.0020 (horigh_{\text{orig}}), 0.7529±0.00520.7529 \pm 0.0052 (hexplh_{\text{expl}}), and 0.7519±0.00310.7519 \pm 0.0031 (hpredh_{\text{pred}}).
    2. LM Fine-Tuning Impact:

      • Without Fine-Tuning (pre-trained DeBERTa embeddings used directly without tuning on labels): Accuracies drop severely across architectures: MLP (0.5797±0.02170.5797 \pm 0.0217), GCN (0.4178±0.11480.4178 \pm 0.1148), SAGE (0.4507±0.05290.4507 \pm 0.0529), RevGAT (0.7507±0.01890.7507 \pm 0.0189).
      • Fine-Tuning a Single LM (shared LM for both sorigs^{\text{orig}} and sexpls^{\text{expl}}): RevGAT reaches 0.7728±0.00140.7728 \pm 0.0014.
      • Fine-Tuning Separate LMs (distinct LMorig\text{LM}_{\text{orig}} and LMexpl\text{LM}_{\text{expl}}): RevGAT reaches 0.7750±0.00120.7750 \pm 0.0012.

    These results demonstrate that LLM explanations provide orthogonal predictive signal to raw text, and that fine-tuning the intermediate LM interpreter is essential for downstream GNN performance.

  9. Knowl 9 — Invariance of TAPE to Language Model Backbone Selection

    data/table

    The performance of TAPE was evaluated across three distinct language model backbones on ogbn-arxiv to determine sensitivity to the choice of the underlying LM interpreter:

    LM Backbone MLP GCN SAGE RevGAT
    deberta-base 0.7587 ±\pm 0.0015 0.7520 ±\pm 0.0003 0.7672 ±\pm 0.0007 0.7750 ±\pm 0.0012
    all-roberta-large-v1 0.7587 ±\pm 0.0003 0.7412 ±\pm 0.0015 0.7695 ±\pm 0.0008 0.7737 ±\pm 0.0004
    e5-large 0.7595 ±\pm 0.0015 0.7443 ±\pm 0.0021 0.7688 ±\pm 0.0010 0.7730 ±\pm 0.0006

    The classification accuracy of the downstream GNN remains stable across different LM families and sizes (varying by less than 0.20% on RevGAT), showing that TAPE's performance improvements are robust to the specific architecture of the small LM interpreter.

  10. Knowl 10 — Dependency on Hand-Crafted Prompt Templates

    limitation

    A key limitation of the TAPE framework is its dependency on manually crafted, dataset-specific prompts. For each new TAG dataset or classification task, the user must explicitly write instructions specifying the class label domain (e.g., listing all 40 arXiv categories or 47 Amazon product categories) and the required output format.

    Empirical testing on ogbn-arxiv demonstrates that minor prompt variations influence LLM zero-shot accuracy (ranging from 69.5% for title-first or content-focused prompts up to 72.0% for default prompts placing the title after the abstract), which in turn propagates to downstream GNN test accuracy (ranging between 76.60% and 77.50% on RevGAT). TAPE does not provide an automated mechanism for prompt optimization or generation.

Coverage note — None was omitted; all primary methodology, theoretical bounds, empirical results across benchmarks and backbones, ablation studies, and dataset contributions have been captured.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  2. 2.Wei-Lin Chiang, Xuanqing Liu, Si Si, Yang Li, Samy Bengio, and Cho-Jui Hsieh. Cluster-gcn: An efficient algorithm for training deep and large graph convolutional networks. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, pp. 257–266, 2019.
  3. 3.Eli Chien, Wei-Cheng Chang, Cho-Jui Hsieh, Hsiang-Fu Yu, Jiong Zhang, Olgica Milenkovic, and Inderjit S Dhillon. Node feature extraction by self-supervised multi-scale neighborhood prediction. arXiv preprint arXiv:2111.00064, 2021.
  4. 4.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  6. 6.Tu Anh Dinh, Jeroen den Boef, Joran Cornelisse, and Paul Groth. E2eg: End-to-end node classification using graph topology and text-based node attributes. arXiv preprint arXiv:2208.04609, 2022.
  7. 7.Keyu Duan, Qian Liu, Tat-Seng Chua, Shuicheng Yan, Wei Tsang Ooi, Qizhe Xie, and Junxian He. Simteg: A frustratingly simple approach improves textual graph learning. arXiv preprint arXiv:2308.02565, 2023.
  8. 8.Matthias Fey and Jan Eric Lenssen. Fast graph representation learning with pytorch geometric. arXiv preprint arXiv:1903.02428, 2019.
  9. 9.Jiayan Guo, Lun Du, and Hengyu Liu. Gpt4graph: Can large language models understand graph structured data? an empirical evaluation and benchmarking. arXiv preprint arXiv:2305.15066, 2023.
  10. 10.Will Hamilton, Zhitao Ying, and Jure Leskovec. Inductive representation learning on large graphs. Advances in neural information processing systems, 30, 2017.
  11. 11.Zellig Harris. Distributional structure. The philosophy of linguistics, 1985.
  12. 12.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=XPZIaotutsD.
  13. 13.Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. Open graph benchmark: Datasets for machine learning on graphs. Advances in neural information processing systems, 33:22118–22133, 2020a.
  14. 14.Ziniu Hu, Yuxiao Dong, Kuansan Wang, Kai-Wei Chang, and Yizhou Sun. Gpt-gnn: Generative pre-training of graph neural networks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 1857–1867, 2020b.
  15. 15.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  16. 16.Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016.
  17. 17.Guohao Li, Matthias Müller, Bernard Ghanem, and Vladlen Koltun. Training graph neural networks with 1000 layers. In International conference on machine learning, pp. 6437–6449. PMLR, 2021.
  18. 18.Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958, 2021.
  19. 19.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1–35, 2023.
  20. 20.Zhenghao Liu, Chenyan Xiong, Maosong Sun, and Zhiyuan Liu. Fine-grained fact verification with kernel graph attention network. arXiv preprint arXiv:1910.09796, 2019.
  21. 21.Andrew Kachites McCallum, Kamal Nigam, Jason Rennie, and Kristie Seymore. Automating the construction of internet portals with machine learning. Information Retrieval, 3:127–163, 2000.
  22. 22.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. Advances in neural information processing systems, 26, 2013.
  23. 23.Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. Mteb: Massive text embedding benchmark. arXiv preprint arXiv:2210.07316, 2022.
  24. 24.Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084, 2019.
  25. 25.Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad. Collective classification in network data. AI magazine, 29(3):93–93, 2008.
  26. 26.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.
  27. 27.Chuxiong Sun, Hongming Gu, and Jie Hu. Scalable and adaptive graph neural networks with self-label-enhanced training. arXiv preprint arXiv:2104.09376, 2021.
  28. 28.Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang, and Xipeng Qiu. Black-box tuning for language-model-as-a-service. In International Conference on Machine Learning, pp. 20841–20855. PMLR, 2022.
  29. 29.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  30. 30.Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems, 34:200–212, 2021.
  31. 31.Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017.
  32. 32.Petar Velickovic, William Fedus, William L Hamilton, Pietro Liò, Yoshua Bengio, and R Devon Hjelm. Deep graph infomax. ICLR (Poster), 2(3):4, 2019.
  33. 33.Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan, Xiaochuang Han, and Yulia Tsvetkov. Can language models solve graph problems in natural language? arXiv preprint arXiv:2305.10037, 2023.
  34. 34.Kuansan Wang, Zhihong Shen, Chiyuan Huang, Chieh-Han Wu, Yuxiao Dong, and Anshul Kanakia. Microsoft academic graph: When experts are not enough. Quantitative Science Studies, 1(1):396–413, 2020.
  35. 35.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533, 2022.
  36. 36.Minjie Wang, Da Zheng, Zihao Ye, Quan Gan, Mufei Li, Xiang Song, Jinjing Zhou, Chao Ma, Lingfan Yu, Yu Gai, et al. Deep graph library: A graph-centric, highly-performant package for graph neural networks. arXiv preprint arXiv:1909.01315, 2019.
  37. 37.Suhang Wang, Jiliang Tang, Charu Aggarwal, and Huan Liu. Linked document embedding for classification. In Proceedings of the 25th ACM international on conference on information and knowledge management, pp. 115–124, 2016.
  38. 38.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022.
  39. 39.Cheng Yang, Zhiyuan Liu, Deli Zhao, Maosong Sun, and Edward Y Chang. Network representation learning with rich text information. In IJCAI, volume 2015, pp. 2111–2117, 2015.
  40. 40.Junhan Yang, Zheng Liu, Shitao Xiao, Chaozhuo Li, Defu Lian, Sanjay Agrawal, Amit Singh, Guangzhong Sun, and Xing Xie. Graphformers: Gnn-nested transformers for representation learning on textual graph. Advances in Neural Information Processing Systems, 34:28798–28810, 2021.
  41. 41.Zhilin Yang, William Cohen, and Ruslan Salakhudinov. Revisiting semi-supervised learning with graph embeddings. In International conference on machine learning, pp. 40–48. PMLR, 2016.
  42. 42.Michihiro Yasunaga, Rui Zhang, Kshitijh Meelu, Ayush Pareek, Krishnan Srinivasan, and Dragomir Radev. Graph-based neural multi-document summarization. arXiv preprint arXiv:1706.06681, 2017.
  43. 43.Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. Qa-gnn: Reasoning with language models and knowledge graphs for question answering. arXiv preprint arXiv:2104.06378, 2021.
  44. 44.Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christopher D Manning, Percy S Liang, and Jure Leskovec. Deep bidirectional language-knowledge graph pretraining. Advances in Neural Information Processing Systems, 35:37309–37323, 2022.
  45. 45.Jiawei Zhang. Graph-toolformer: To empower llms with graph reasoning ability via prompt augmented by chatgpt. arXiv preprint arXiv:2304.11116, 2023.
  46. 46.Shichang Zhang, Yozen Liu, Yizhou Sun, and Neil Shah. Graph-less neural networks: Teaching old mlps new tricks via distillation. arXiv preprint arXiv:2110.08727, 2021.
  47. 47.Xikun Zhang, Antoine Bosselut, Michihiro Yasunaga, Hongyu Ren, Percy Liang, Christopher D Manning, and Jure Leskovec. Greaselm: Graph reasoning enhanced language models. In International conference on learning representations, 2022.
  48. 48.Jianan Zhao, Meng Qu, Chaozhuo Li, Hao Yan, Qian Liu, Rui Li, Xing Xie, and Jian Tang. Learning on large-scale text-attributed graphs via variational inference. arXiv preprint arXiv:2210.14709, 2022.
  49. 49.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pp. 12697–12706. PMLR, 2021.
  50. 50.Jie Zhou, Xu Han, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, and Maosong Sun. Gear: Graph-based evidence aggregating and reasoning for fact verification. arXiv preprint arXiv:1908.01843, 2019.
  51. 51.Jason Zhu, Yanling Cui, Yuming Liu, Hao Sun, Xue Li, Markus Pelger, Tianqi Yang, Liangjie Zhang, Ruofei Zhang, and Huasha Zhao. Textgnn: Improving text encoder via graph neural network in sponsored search. In Proceedings of the Web Conference 2021, pp. 2848–2857, 2021.

Citation

MLA
He, X., et al. “Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning”. arXiv, 2023, http://arxiv.org/abs/2305.19523v5.
APA
He, X., Bresson, X., Laurent, T., Perold, A., LeCun, Y., & Hooi, B. (2023). Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning. arXiv. http://arxiv.org/abs/2305.19523v5
Chicago
He, X., X. Bresson, T. Laurent, A. Perold, Y. LeCun, and B. Hooi. 2023. “Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning”. arXiv. http://arxiv.org/abs/2305.19523v5.
Harvard
He, X. et al. (2023) “Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.19523v5.
Vancouver
1. He X, Bresson X, Laurent T, Perold A, LeCun Y, Hooi B (2023) Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning. arXiv

BibTeX

@article{he2023harnessing,
  title = {Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning},
  author = {He, Xiaoxin and Bresson, Xavier and Laurent, Thomas and Perold, Adam and LeCun, Yann and Hooi, Bryan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.19523v5},
  eprint = {2305.19523}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors