Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning
Xiaoxin HeXavier BressonThomas LaurentAdam PeroldYann LeCunBryan Hooi
Proposes an LLM-to-LM interpreter that converts zero-shot large language model explanations into informative node features for graph neural networks, achieving state-of-the-art accuracy and a near-threefold training speedup on text-attributed graph benchmarks.
Real-world data networks, such as academic citation webs and e-commerce product catalogs, frequently combine complex relational structures with rich textual information. While graph neural networks (GNNs) excel at modeling network connectivity, they traditionally rely on shallow text embeddings or complex, computationally intensive language model integrations that cannot easily leverage the sophisticated reasoning of modern large language models (LLMs). This computational bottleneck limits the ability of organizations to harness high-level textual reasoning for network classification tasks.
The article demonstrates a novel framework called TAPE (Title, Abstract, Prediction, and Explanation) that integrates the broad reasoning of LLMs with the structural learning of GNNs. The primary objective is to evaluate whether LLM-generated explanations and ranked predictions can be translated into compact, high-performance features for downstream graph classification using standard, budget-friendly infrastructure.
The authors evaluated this approach across five benchmark datasets, including citation networks (Cora, PubMed, ogbn-arxiv, and the newly introduced tape-arxiv23) and an e-commerce network (ogbn-products). The method queries an LLM—such as GPT-3.5 or open-source Llama-2—in a zero-shot manner to generate predictions along with natural language rationales. A smaller language model (such as DeBERTa) is then fine-tuned to interpret these texts and produce fixed vectorial representations. These features are kept frozen while training standard GNN architectures (such as GCN, GraphSAGE, and RevGAT), entirely decoupling text encoding from graph training.
The evaluation yielded several key findings. First, the framework established new state-of-the-art accuracy across all tested benchmarks, achieving 77.50% test accuracy on the ogbn-arxiv benchmark and outperforming existing iterative language-graph models. Second, decoupling feature extraction from graph learning reduced total training computation time by 2.88 times compared to the closest state-of-the-art baseline (GLEM). Third, the approach maintained strong generalization on contemporary data (tape-arxiv23), reaching 84.23% accuracy on papers published after the LLM's training cutoff. Finally, using open-source models like Llama-2 delivered competitive performance (76.19% on ogbn-arxiv), confirming that organizations can avoid proprietary API costs without substantial performance degradation.
These findings indicate that organizations can significantly enhance classification accuracy in relational text systems while lowering computational expense and development overhead. Because the framework operates via standard application programming interfaces (APIs) and freezes extracted features prior to graph training, it avoids expensive end-to-end retraining and scales efficiently within standard hardware budgets. The results demonstrate that generating intermediate text explanations captures vital contextual nuances that raw text alone fails to convey to smaller models.
Decision-makers should consider adopting this decoupled interpreter architecture when upgrading graph analytics pipelines, particularly for document routing, recommendation systems, or knowledge management. Teams can balance budget constraints by choosing between hosted commercial LLM APIs and free open-source models depending on latency and hosting capabilities. Future technical work should focus on automating dataset-specific prompt generation and evaluating the pipeline on dynamic, continuously evolving graphs.
The primary limitation of the study is its reliance on manually engineered prompts, which can cause slight fluctuations in extraction quality across different domains. Nonetheless, extensive ablation experiments and consistent results across multiple models provide strong confidence in the stability, modularity, and practical value of the proposed architecture.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). Provides the foundational semi-supervised Graph Convolutional Network (GCN) architecture evaluated and trained on text-attributed citation graphs in the TAPE framework.
- Paper: Graph Attention Networks, Petar Veličković et al. (2018). Introduces Graph Attention Networks (GAT), establishing the attention-based graph learning baseline that TAPE utilizes for downstream classification.
- Paper: Graph Convolutional Networks for Text Classification, Liang Yao et al. (2018). Establishes the paradigm of using graph convolutional networks directly on text-rich node networks for document classification.
- Paper: Simplifying Graph Convolutional Networks, Felix Wu et al. (2019). Demonstrates the benefits of simplifying graph convolutional operations, motivating decoupled and computationally efficient feature learning pipelines on graph datasets.
- Paper: Predict then Propagate: Graph Neural Networks meet Personalized PageRank, Johannes Gasteiger et al. (2019). Pioneers the 'predict then propagate' paradigm that underpins decoupled feature extraction and downstream propagation in text-attributed graph learning.
- Paper: Pitfalls of Graph Neural Network Evaluation, Oleksandr Shchur et al. (2018). Highlights key benchmarking pitfalls and standard evaluation practices across citation datasets like Cora and PubMed that TAPE adopts.
- Paper: Benchmarking Graph Neural Networks, Vijay Prakash Dwivedi et al. (2023). Establishes standardized, reproducible benchmarking protocols and parameter-constrained evaluations for graph neural networks across diverse graph tasks.
- Paper: Relational inductive biases, deep learning, and graph networks, Peter W. Battaglia et al. (2018). Formulates the foundational conceptual framework for relational inductive biases and graph network representations.
- Paper: Label-free Node Classification on Graphs with Large Language Models (LLMs), Zhikai Chen et al. (2024). Extends the integration of LLM zero-shot capabilities with GNNs by developing active-learning node selection and confidence filtering for label-free graph classification.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). Builds upon LLM-graph interaction by proposing adaptive, self-correcting planning mechanisms over knowledge graphs rather than static feature interpreter pipelines.
- Paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization, Darren Edge et al. (2024). Applies the synergy of language model reasoning and graph representations to hierarchical query-focused summarization and global corpus retrieval.
