Unifying Large Language Models and Knowledge Graphs: A Roadmap

Shirui PanLinhao LuoYufei WangChen ChenJiapu WangXindong Wu

article2023TKDE1,765 citations

Presents a structured framework for integrating large language models and knowledge graphs, detailing concrete methods to resolve factual hallucinations, automate graph construction, and enable bidirectional reasoning.

Listen

Modern artificial intelligence systems increasingly depend on large language models that excel at natural language processing and general reasoning across diverse domains. However, these models operate largely as black-box systems prone to generating factually inaccurate statementsoften termed hallucinationsand struggle to access fresh or domain-specific data. In contrast, knowledge graphs offer structured, decisive, and fully interpretable factual repositories, but they are expensive to construct, frequently incomplete, and incapable of flexible linguistic reasoning. Because these two technologies have complementary strengths and weaknesses, integrating them represents an important opportunity to create more dependable and explainable artificial intelligence systems.

The article establishes a comprehensive roadmap for unifying language models and structured knowledge bases. It categorizes current integration paradigms, evaluates their operational mechanisms, and demonstrates how combining explicit symbolic data with statistical neural networks enhances factual accuracy, reasoning capability, and interpretability across downstream applications.

To map this landscape, the article conducts an extensive, high-level review of recent literature spanning natural language processing and graph neural architectures. It organizes existing methodologies into three core frameworks: knowledge-enhanced language models, language-model-augmented knowledge graphs, and fully synergized systems where both components operate as equal partners in bidirectional reasoning.

The review yields several critical findings regarding system performance and design trade-offs. First, integrating structured knowledge during model pre-training deeply embeds factual representations into neural parameters, but it prevents subsequent knowledge updates without costly re-training. Second, inference-time knowledge retrieval and structured prompting provide dynamic access to up-to-date facts without requiring model re-training, though they depend heavily on manual prompt engineering and retriever performance. Third, using language models as sequence-to-sequence generators directly completes missing graph facts more efficiently and generalizes better to unseen entities than classification-based scoring methods. Fourth, synergized frameworks that merge bidirectional attention across text tokens and graph entities achieve superior multi-hop reasoning and explainability compared to isolated systems.

These findings indicate that unifying language models with structured knowledge bases directly mitigates operational risks and compliance challenges in high-stakes fields like medical diagnosis and legal decision-making. By grounding neural text generation in verified facts, organizations can curtail factual errors, reduce computational expenses associated with continuous re-training, and provide auditable reasoning paths to satisfy regulatory demands.

Decision-makers should choose integration architectures according to their operational requirements. When handling static, domain-specific foundational knowledge, teams should adopt pre-training injection methods, whereas fast-evolving operational data requires retrieval-augmented inference pipelines. Furthermore, organizations should deploy language models to automate the extraction and construction of proprietary knowledge bases, while preparing workflows to test synergized agent-based reasoning frameworks on complex tasks.

Nevertheless, several limitations remain that warrant cautious deployment. Most existing integration techniques require access to internal model weights, making them difficult to implement on closed-source, API-only models. Additionally, transforming large-scale graph structures into text sequences risks information loss due to token input constraints. Future progress depends on establishing robust multimodal alignment, automated hallucination detection pipelines, and standardized methods for live knowledge updating.

Cover for Unifying Large Language Models and Knowledge Graphs: A Roadmap

Abstract

Large language models (LLMs), such as ChatGPT and GPT4, are making new waves in the field of natural language processing and artificial intelligence, due to their emergent ability and generalizability. However, LLMs are black-box models, which often fall short of capturing and accessing factual knowledge. In contrast, Knowledge Graphs (KGs), Wikipedia and Huapu for example, are structured knowledge models that explicitly store rich factual knowledge. KGs can enhance LLMs by providing external knowledge for inference and interpretability. Meanwhile, KGs are difficult to construct and evolving by nature, which challenges the existing methods in KGs to generate new facts and represent unseen knowledge. Therefore, it is complementary to unify LLMs and KGs together and simultaneously leverage their advantages. In this article, we present a forward-looking roadmap for the unification of LLMs and KGs. Our roadmap consists of three general frameworks, namely, 1) KG-enhanced LLMs, which incorporate KGs during the pre-training and inference phases of LLMs, or for the purpose of enhancing understanding of the knowledge learned by LLMs; 2) LLM-augmented KGs, that leverage LLMs for different KG tasks such as embedding, completion, construction, graph-to-text generation, and question answering; and 3) Synergized LLMs + KGs, in which LLMs and KGs play equal roles and work in a mutually beneficial way to enhance both LLMs and KGs for bidirectional reasoning driven by both data and knowledge. We review and summarize existing efforts within these three frameworks in our roadmap and pinpoint their future research directions.

Table of Contents

  • I Introduction
  • II Background
  • II-A Large Language models (LLMs)
  • II-A1 Encoder-only LLMs.
  • II-A2 Encoder-decoder LLMs.
  • II-A3 Decoder-only LLMs.
  • II-A4 Prompt Engineering
  • II-B Knowledge Graphs (KGs)
  • II-B1 Encyclopedic Knowledge Graphs.
  • II-B2 Commonsense Knowledge Graphs.
  • II-B3 Domain-specific Knowledge Graphs
  • II-B4 Multi-modal Knowledge Graphs.
  • II-C Applications
  • III Roadmap & Categorization
  • III-A Roadmap
  • III-A1 KG-enhanced LLMs
  • III-A2 LLM-augmented KGs
  • III-A3 Synergized LLMs + KGs
  • III-B Categorization
  • IV KG-enhanced LLMs
  • IV-A KG-enhanced LLM Pre-training
  • IV-A1 Integrating KGs into Training Objective
  • IV-A2 Integrating KGs into LLM Inputs
  • IV-A3 KGs Instruction-tuning
  • IV-B KG-enhanced LLM Inference
  • IV-B1 Retrieval-Augmented Knowledge Fusion
  • IV-B2 KGs Prompting
  • IV-C Comparison between KG-enhanced LLM Pre-training and Inference
  • IV-D KG-enhanced LLM Interpretability
  • IV-D1 KGs for LLM Probing
  • IV-D2 KGs for LLM Analysis
  • V LLM-augmented KGs
  • V-A LLM-augmented KG Embedding
  • V-A1 LLMs as Text Encoders
  • V-A2 LLMs for Joint Text and KG Embedding
  • V-B LLM-augmented KG Completion
  • V-B1 LLM as Encoders (PaE).
  • V-B2 LLM as Generators (PaG).
  • V-B3 Model Analysis
  • V-C LLM-augmented KG Construction
  • V-C1 Entity Discovery
  • V-C2 Coreference Resolution (CR)
  • V-C3 Relation Extraction (RE)
  • V-C4 Distilling Knowledge Graphs from LLMs
  • V-D LLM-augmented KG-to-text Generation
  • V-D1 Leveraging Knowledge from LLMs
  • V-D2 Constructing large weakly KG-text aligned Corpus
  • V-E LLM-augmented KG Question Answering
  • V-E1 LLMs as Entity/relation Extractors
  • V-E2 LLMs as Answer Reasoners
  • VI Synergized LLMs + KGs
  • VI-A Synergized Knowledge Representation
  • VI-B Synergized Reasoning
  • VII Future Directions and Milestones
  • VII-A KGs for Hallucination Detection in LLMs
  • VII-B KGs for Editing Knowledge in LLMs
  • VII-C KGs for Black-box LLMs Knowledge Injection
  • VII-D Multi-Modal LLMs for KGs
  • VII-E LLMs for Understanding KG Structure
  • VII-F Synergized LLMs and KGs for Birectional Reasoning
  • VIII Conclusion
  • References
  • A Pros and Cons for LLMs and KGs

Knowls

  1. Knowl 1 — Tripartite Roadmap and Milestones for Unifying Large Language Models and Knowledge Graphs

    model/method

    The unification of Large Language Models (LLMs) and Knowledge Graphs (KGs) is conceptualized as a three-framework roadmap designed to leverage their complementary strengths—the general language understanding and generalization of LLMs and the structured, decisive, and accurate factual representation of KGs:

    1. KG-Enhanced LLMs: Incorporates structured KGs into LLMs during pre-training, fine-tuning, or inference, or utilizes KGs to probe, analyze, and interpret the representations and reasoning mechanisms of LLMs.
    2. LLM-Augmented KGs: Applies LLMs to augment core KG tasks, including KG embedding, KG completion, KG construction, KG-to-text generation, and KG question answering (KGQA).
    3. Synergized LLMs + KGs: Unifies LLMs and KGs into an integrated framework where both components operate bidirectionally on equal footing for joint representation learning and dual data-and-knowledge-driven reasoning.

    The development of this unified domain is projected across three progressive milestones:

    • Stage 1: Parallel and separate advancements in KG-enhanced LLMs and LLM-augmented KGs.
    • Stage 2: Deep synergy between LLMs and KGs for joint multi-hop reasoning and representation learning.
    • Stage 3: Advanced graph structure comprehension by LLMs, multimodal KG integration, and continuous knowledge updating without full model retraining.
  2. Knowl 2 — Taxonomy and Strategies of KG-Enhanced Large Language Models

    model/method

    KG-Enhanced LLMs inject structured factual knowledge into language models to alleviate hallucinations, lack of domain knowledge, and black-box opacity. The methodology is categorized into three stages:

    1. KG-Enhanced Pre-Training:
      • Knowledge-Aware Objectives: Designs pre-training loss functions that mask entities reachable within specific graph hops or enforce token-entity alignment (e.g., entity prediction or contrastive losses).
      • Knowledge-Augmented Inputs: Linearizes relevant KG subgraphs and concatenates them with text, utilizing visible attention matrices or unified word-knowledge graphs to prevent irrelevant triples from causing knowledge noise.
      • KG Instruction Tuning: Fine-tunes LLMs on instruction datasets derived from KG structures, relation paths, or logical query formats to enhance graph reasoning.
    2. KG-Enhanced Inference:
      • Retrieval-Augmented Knowledge Fusion: Retrieves relevant entities, documents, or subgraphs from an external KG during inference and fuses them as non-parametric memory into the LLM decoder without modifying model weights.
      • KG Prompting: Converts triples or subgraphs into prompt structures, mind maps, or chain-of-knowledge paths to guide LLM inference.
    3. KG-Enhanced Interpretability:
      • KG Probing: Converts factual triples (h,r,t)(h, r, t) into cloze prompts (e.g., masking the tail entity tt) to assess and quantify latent factual memorization in LLM parameters.
      • KG Analysis: Grounding multi-step internal representations and neuron activations onto explicit KG subgraphs to trace causal attribution and verify decision pathways.
  3. Knowl 3 — Trade-Off Analysis: KG-Enhanced LLM Pre-Training versus KG-Enhanced LLM Inference

    theoretical result

    The choice between integrating knowledge graphs into LLMs during pre-training versus during inference is governed by trade-offs among semantic alignment, update frequency, and computational overhead:

    • KG-Enhanced Pre-Training: Pre-training directly aligns explicit entity and relation representations with language embeddings from scratch. This delivers optimal empirical performance on static and knowledge-intensive downstream tasks within specific domains (such as commonsense reasoning). However, it is computationally intensive and static: incorporating newly emerging or edited facts requires retraining the model, causing poor adaptation to dynamically evolving real-world knowledge.
    • KG-Enhanced Inference: Inference-stage methods maintain separate text and knowledge spaces, dynamically querying non-parametric KG stores or prepending retrieved triples via structured prompts. This enables immediate knowledge updating and zero-shot adaptation to new domains without model retraining. However, because the underlying base LLM is not explicitly trained from scratch to digest graph representations, it can exhibit suboptimal fusion efficiency and is constrained by context window length limits.
  4. Knowl 4 — Taxonomy of LLM-Augmented Knowledge Graph Tasks

    model/method

    LLM-Augmented KGs utilize the linguistic, contextual, and generative capabilities of LLMs to overcome standard KG bottlenecks, including graph incompleteness, lack of textual context, and poor generalization to unseen entities. The taxonomy covers five primary tasks:

    1. KG Embedding (KGE): Uses LLMs to encode textual descriptions of entities and relations into dense vectors, or jointly trains LLM token representations with structural loss functions.
    2. KG Completion (KGC): Infers missing triples (h,r,?)(h, r, ?) or (?,r,t)(?, r, t) by using LLMs as discriminative encoders (ranking candidates via classification heads) or as sequence-to-sequence generators (directly decoding the string of the missing entity).
    3. KG Construction: Deploys LLMs across the entire extraction pipeline, comprising:
      • Named Entity Recognition (NER): Identifying flat, nested, and discontinuous entities.
      • Entity Typing (ET): Assigning fine-grained hierarchical type labels to context mentions.
      • Entity Linking (EL): Disambiguating and mapping text mentions to canonical KG nodes.
      • Coreference Resolution (CR): Clustering coreferent mentions within and across documents.
      • Relation Extraction (RE): Detecting relational links between entity pairs at the sentence or document level.
      • End-to-End Extraction & Distillation: Generating complete triples directly from unstructured text or harvesting symbolic KGs from LLMs.
    4. KG-to-Text Generation: Generates coherent natural language descriptions corresponding to input subgraphs through linearized graph traversals or structure-aware attention.
    5. KG Question Answering (KGQA): Bridges natural language questions and KG facts by using LLMs as entity/relation extractors or as answer reasoners over candidate graph paths.
  5. Knowl 5 — Mathematical Formulations for LLM-Augmented Knowledge Graph Embedding

    model/method

    Methods augmenting Knowledge Graph Embedding (KGE) via LLMs fall into two primary formulation categories:

    1. LLMs as Text Encoders: For a triple (h,r,t)E×R×E(h, r, t) ⊆ \mathcal{E} \times \mathcal{R} \times \mathcal{E} with corresponding textual descriptions Texth,Textt,Textr\text{Text}_h, \text{Text}_t, \text{Text}_r, initial semantic embeddings are generated by an LLM encoder: eh=LLM(Texth),et=LLM(Textt),er=LLM(Textr)e_h = \text{LLM}(\text{Text}_h), \quad e_t = \text{LLM}(\text{Text}_t), \quad e_r = \text{LLM}(\text{Text}_r) These embeddings are mapped to graph representations vh,vr,vtv_h, v_r, v_t and optimized using a margin-based structural ranking loss: L=[γ+f(vh,vr,vt)f(vh,vr,vt)]+\mathcal{L} = [\gamma + f(v_h, v_r, v_t) - f(v_h', v_r', v_t')]_+ where f()f(\cdot) is a KGE scoring function (such as TransE or DisMult), γ>0\gamma > 0 is a margin hyperparameter, and (vh,vr,vt)(v_h', v_r', v_t') denotes negative triple samples.

    2. LLMs for Joint Text and KG Embedding: The graph structure and descriptive text are fused directly in the token sequence input to the LLM. For instance, replacing the tail entity with a mask token [MASK]\text{[MASK]} yields: x=[CLS]  h  Texth  [SEP]  r  [SEP]  [MASK]  Textt  [SEP]x = \text{[CLS]} \; h \; \text{Text}_h \; \text{[SEP]} \; r \; \text{[SEP]} \; \text{[MASK]} \; \text{Text}_t \; \text{[SEP]} The LLM parameters Θ\Theta are optimized to maximize the probability of recovering the correct entity: PLLM(th,r)=P([MASK]=tx,Θ)P_{\text{LLM}}(t \mid h, r) = P(\text{[MASK]} = t \mid x, \Theta) The resulting hidden states of the entity and relation tokens serve as the unified embeddings.

  6. Knowl 6 — Paradigms for LLM-Augmented Knowledge Graph Completion: Encoders versus Generators

    model/method

    LLM-augmented Knowledge Graph Completion (KGC) addresses missing triple prediction via two distinct paradigms:

    1. LLM as Encoders (PaE):

      • Joint Encoding: Encodes the full triple string x=[CLS]Texth[SEP]Textr[SEP]Textt[SEP]x = \text{[CLS]} \, \text{Text}_h \, \text{[SEP]} \, \text{Text}_r \, \text{[SEP]} \, \text{Text}_t \, \text{[SEP]} and calculates a validity score s=σ(MLP(e[CLS]))s = \sigma(\text{MLP}(e_{\text{[CLS]}})), where σ\sigma is the sigmoid function.
      • Masked Language Model (MLM) Encoding: Inputs x=[CLS]Texth[SEP]Textr[SEP][MASK][SEP]x = \text{[CLS]} \, \text{Text}_h \, \text{[SEP]} \, \text{Text}_r \, \text{[SEP]} \, \text{[MASK]} \, \text{[SEP]} to predict masked entity tokens via the LLM classification head.
      • Separated Encoding: Encodes the query pair x(h,r)=[CLS]Texth[SEP]Textr[SEP]x_{(h,r)} = \text{[CLS]} \, \text{Text}_h \, \text{[SEP]} \, \text{Text}_r \, \text{[SEP]} and tail entity xt=[CLS]Textt[SEP]x_t = \text{[CLS]} \, \text{Text}_t \, \text{[SEP]} through Siamese encoders and evaluates s=fscore(e(h,r),et)s = f_{\text{score}}(e_{(h,r)}, e_t).
      • Properties: Enables parameter-efficient fine-tuning (freezing the LLM backbone and tuning shallow prediction heads), but requires scoring all candidates at inference time, which cannot generalize efficiently to unseen entities.
    2. LLM as Generators (PaG):

      • Uses encoder-decoder or autoregressive decoder-only LLMs fed with the query prompt (e.g., Head: h, Relation: r, Tail:) to directly generate the text name of tail entity tt.
      • Properties: Requires no candidate ranking, supports zero-shot generation of unseen entities, and handles closed-source models via prompting. However, autoregressive decoding has higher inference latency and risks generating invalid or out-of-KG entities.
  7. Knowl 7 — Mechanisms for LLM-Augmented Knowledge Graph Question Answering

    model/method

    Knowledge Graph Question Answering (KGQA) uses LLMs to bridge natural language questions qq and structured KG facts via two main functional roles:

    1. LLMs as Entity/Relation Extractors: LLMs identify question mentions and score relations rRr \in \mathcal{R} via semantic similarity: s(r,q)=LLM(r)LLM(q)s(r, q) = \text{LLM}(r)^\top \text{LLM}(q) For multi-hop reasoning paths p=(r1,r2,,rp)p = (r_1, r_2, \dots, r_{|p|}), the path probability given qq is factorized as: P(pq)=t=1ps(rt,q)P(p \mid q) = \prod_{t=1}^{|p|} s(r_t, q)

    2. LLMs as Answer Reasoners: Retrieved relations or subgraphs are verbalized into text paths p1,,pnp_1, \dots, p_n and concatenated with candidate answer aa and question qq: x=[CLS]q[SEP]p1[SEP][SEP]pn[SEP]x = \text{[CLS]} \, q \, \text{[SEP]} \, p_1 \, \text{[SEP]} \dots \text{[SEP]} \, p_n \, \text{[SEP]} The contextual representation e[CLS]=LLM(x)e_{\text{[CLS]}} = \text{LLM}(x) is passed to a classification head s=σ(MLP(e[CLS]))s = \sigma(\text{MLP}(e_{\text{[CLS]}})) to verify if aa answers qq. The overall answer probability marginalized over candidate paths P\mathcal{P} is: P(aq)=pPP(ap)P(pq)P(a \mid q) = \sum_{p \in \mathcal{P}} P(a \mid p) P(p \mid q)

  8. Knowl 8 — Four-Layer Architecture of Synergized LLMs and Knowledge Graphs

    model/method

    The unified Synergized LLMs + KGs framework organizes mutual integration between LLMs and KGs into four hierarchical layers:

    1. Data Layer: Ingests and organizes multimodal source data. LLMs process unstructured text corpora, while KGs manage structured factual triples (h,r,t)(h, r, t). Extensions accommodate multimodal data (images, videos, audio).
    2. Synergized Model Layer: Fuses LLMs (providing general linguistic processing, implicit knowledge, and open-ended generalization) and KGs (providing explicit, decisive, domain-specific, and interpretable relational knowledge).
    3. Technique Layer: Integrates foundational methods from both domains, including prompt engineering, in-context learning, graph neural networks (GNNs), representation learning, neural-symbolic reasoning, and few-shot learning.
    4. Application Layer: Applies the unified capabilities to complex downstream real-world tasks, including search engines, conversational dialogue systems, intelligent recommendation engines, and domain AI assistants.
  9. Knowl 9 — Synergized Reasoning Paradigms: LLM-KG Fusion Reasoning versus LLMs as Agents

    model/method

    Synergized reasoning combines text representations and graph structures through two foundational paradigms:

    1. LLM-KG Fusion Reasoning:

      • Employs separate LLM and KG encoders trained jointly end-to-end.
      • Uses bidirectional cross-attention layers (LM-to-KG and KG-to-LM attention) to calculate pairwise interactions between all text tokens and graph entities.
      • Applies dynamic graph pruning based on cross-attention weights to discard irrelevant subgraphs at deeper layers.
      • Trade-off: Achieves deep semantic fusion between text and structural contexts, but introduces substantial architectural complexity, extra model parameters, and opaque multi-layer interactions.
    2. LLMs as Agents for KG Reasoning:

      • Uses the LLM as an autonomous agent that directly navigates external KGs via structured APIs, beam searches over graph paths, or iterative chain-of-thought verification (e.g., Think-on-Graph, StructGPT).
      • Trade-off: Fully modular and plug-and-play without requiring end-to-end retraining, providing human-interpretable reasoning paths. However, performance depends heavily on agent prompt design, token budget constraints, and the robustness of the search action space.
  10. Knowl 10 — Open Research Challenges for the Unification of LLMs and KGs

    limitation

    Six core open challenges and future directions define the frontier of unifying LLMs and KGs:

    1. KG-Based Hallucination Detection: Utilizing structured KGs as dynamic, verifiable ground-truth references to construct cross-domain fact-checking classifiers for AI-generated text.
    2. Knowledge Editing and Ripple Effects: Developing algorithms that update obsolete or erroneous facts in LLM parameters via KGs without degrading adjacent knowledge or causing unpredictable downstream inference errors.
    3. Black-Box LLM Knowledge Injection: Finding parameter-free mechanisms to reliably inject complex, structured KG subgraphs into closed-source, API-only models (e.g., GPT-4) within context length limits.
    4. Multimodal KG and LLM Alignment: Bridging modalities by aligning image, audio, and video entities within multimodal KGs with multimodal foundation models.
    5. Native Graph Structure Comprehension: Designing LLM architectures and pre-training objectives that natively understand graph topology, structural hierarchies, and edge semantics directly, moving beyond lossy 1D text linearization.
    6. Bidirectional Dual-Wheel Reasoning: Creating closed-loop architectures where knowledge-driven graph search (exploring unseen and structural constraints) and text-driven LLM inference continuously validate and guide each other.

Coverage note — None was omitted; all primary taxonomies, mathematical formulations, comparison frameworks, architectural layers, and research roadmaps contributed by the paper are fully captured.

References

  1. 1.J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, 2018.
  2. 2.Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “Roberta: A robustly optimized bert pretraining approach,” arXiv preprint arXiv:1907.11692, 2019.
  3. 3.C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research, vol. 21, no. 1, pp. 5485–5551, 2020.
  4. 4.D. Su, Y. Xu, G. I. Winata, P. Xu, H. Kim, Z. Liu, and P. Fung, “Generalizing question answering system with pre-trained language model fine-tuning,” in Proceedings of the 2nd Workshop on Machine Reading for Question Answering, 2019, pp. 203–211.
  5. 5.M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in ACL, 2020, pp. 7871–7880.
  6. 6.J. Li, T. Tang, W. X. Zhao, and J.-R. Wen, “Pretrained language models for text generation: A survey,” arXiv preprint arXiv:2105.10311, 2021.
  7. 7.J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler et al., “Emergent abilities of large language models,” Transactions on Machine Learning Research.
  8. 8.K. Malinka, M. Peresˇ´ıni, A. Firc, O. Hujnˇak, and F. Janu´sˇ, “On the educational impact of chatgpt: Is artificial intelligence ready to obtain a university degree?” arXiv preprint arXiv:2303.11146, 2023.
  9. 9.Z. Li, C. Wang, Z. Liu, H. Wang, S. Wang, and C. Gao, “Cctest: Testing and repairing code completion systems,” ICSE, 2023.
  10. 10.J. Liu, C. Liu, R. Lv, K. Zhou, and Y. Zhang, “Is chatgpt a good recommender? a preliminary study,” arXiv preprint arXiv:2304.10149, 2023.
  11. 11.W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong et al., “A survey of large language models,” arXiv preprint arXiv:2303.18223, 2023.
  12. 12.X. Qiu, T. Sun, Y. Xu, Y. Shao, N. Dai, and X. Huang, “Pre-trained models for natural language processing: A survey,” Science China Technological Sciences, vol. 63, no. 10, pp. 1872–1897, 2020.
  13. 13.J. Yang, H. Jin, R. Tang, X. Han, Q. Feng, H. Jiang, B. Yin, and X. Hu, “Harnessing the power of llms in practice: A survey on chatgpt and beyond,” arXiv preprint arXiv:2304.13712, 2023.
  14. 14.F. Petroni, T. Rocktaschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu,¨ and A. Miller, “Language models as knowledge bases?” in EMNLP-IJCNLP, 2019, pp. 2463–2473.
  15. 15.Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys, vol. 55, no. 12, pp. 1–38, 2023.
  16. 16.H. Zhang, H. Song, S. Li, M. Zhou, and D. Song, “A survey of controllable text generation using transformer-based pre-trained language models,” arXiv preprint arXiv:2201.05337, 2022.
  17. 17.M. Danilevsky, K. Qian, R. Aharonov, Y. Katsis, B. Kawas, and P. Sen, “A survey of the state of explainable ai for natural language processing,” arXiv preprint arXiv:2010.00711, 2020.
  18. 18.J. Wang, X. Hu, W. Hou, H. Chen, R. Zheng, Y. Wang, L. Yang, H. Huang, W. Ye, X. Geng et al., “On the robustness of chatgpt: An adversarial and out-of-distribution perspective,” arXiv preprint arXiv:2302.12095, 2023.
  19. 19.S. Ji, S. Pan, E. Cambria, P. Marttinen, and S. Y. Philip, “A survey on knowledge graphs: Representation, acquisition, and applications,” IEEE TNNLS, vol. 33, no. 2, pp. 494–514, 2021.
  20. 20.D. Vrandeciˇc and M. Kr´otzsch, “Wikidata: a free collaborative¨ knowledgebase,” Communications of the ACM, vol. 57, no. 10, pp. 78–85, 2014.
  21. 21.S. Hu, L. Zou, and X. Zhang, “A state-transition framework to answer complex questions over knowledge base,” in EMNLP, 2018, pp. 2098–2108.
  22. 22.J. Zhang, B. Chen, L. Zhang, X. Ke, and H. Ding, “Neural, symbolic and neural-symbolic reasoning on knowledge graphs,” AI Open, vol. 2, pp. 14–35, 2021.
  23. 23.B. Abu-Salih, “Domain-specific knowledge graphs: A survey,” Journal of Network and Computer Applications, vol. 185, p. 103076, 2021.
  24. 24.T. Mitchell, W. Cohen, E. Hruschka, P. Talukdar, B. Yang, J. Betteridge, A. Carlson, B. Dalvi, M. Gardner, B. Kisiel, K. Jayant, L. Ni, M. Kathryn, M. Thahir, N. Ndapandula, P. Emmanouil, R. Alan, S. Mehdi, S. Burr, W. Derry, G. Abhinav, C. Xi, S. Abulhair, and W. Joel, “Never-ending learning,” Communications of the ACM, vol. 61, no. 5, pp. 103–115, 2018.
  25. 25.L. Zhong, J. Wu, Q. Li, H. Peng, and X. Wu, “A comprehensive survey on automatic knowledge graph construction,” arXiv preprint arXiv:2302.05019, 2023.
  26. 26.L. Yao, C. Mao, and Y. Luo, “Kg-bert: Bert for knowledge graph completion,” arXiv preprint arXiv:1909.03193, 2019.
  27. 27.L. Luo, Y.-F. Li, G. Haffari, and S. Pan, “Normalizing flow-based neural process for few-shot knowledge graph completion,” SIGIR, 2023.
  28. 28.Y. Bang, S. Cahyawijaya, N. Lee, W. Dai, D. Su, B. Wilie, H. Lovenia, Z. Ji, T. Yu, W. Chung et al., “A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity,” arXiv preprint arXiv:2302.04023, 2023.
  29. 29.X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” arXiv preprint arXiv:2203.11171, 2022.
  30. 30.O. Golovneva, M. Chen, S. Poff, M. Corredor, L. Zettlemoyer, M. Fazel-Zarandi, and A. Celikyilmaz, “Roscoe: A suite of metrics for scoring step-by-step reasoning,” ICLR, 2023.
  31. 31.F. M. Suchanek, G. Kasneci, and G. Weikum, “Yago: a core of semantic knowledge,” in WWW, 2007, pp. 697–706.
  32. 32.A. Carlson, J. Betteridge, B. Kisiel, B. Settles, E. Hruschka, and T. Mitchell, “Toward an architecture for never-ending language learning,” in Proceedings of the AAAI conference on artificial intelligence, vol. 24, no. 1, 2010, pp. 1306–1313.
  33. 33.A. Bordes, N. Usunier, A. Garcia-Duran, J. Weston, and O. Yakhnenko, “Translating embeddings for modeling multi-relational data,” NeurIPS, vol. 26, 2013.
  34. 34.G. Wan, S. Pan, C. Gong, C. Zhou, and G. Haffari, “Reasoning like human: Hierarchical reinforcement learning for knowledge graph reasoning,” in AAAI, 2021, pp. 1926–1932.
  35. 35.Z. Zhang, X. Han, Z. Liu, X. Jiang, M. Sun, and Q. Liu, “ERNIE: Enhanced language representation with informative entities,” in ACL, 2019, pp. 1441–1451.
  36. 36.W. Liu, P. Zhou, Z. Zhao, Z. Wang, Q. Ju, H. Deng, and P. Wang, “K-BERT: enabling language representation with knowledge graph,” in AAAI, 2020, pp. 2901–2908.
  37. 37.Y. Liu, Y. Wan, L. He, H. Peng, and P. S. Yu, “KG-BART: knowledge graph-augmented BART for generative commonsense reasoning,” in AAAI, 2021, pp. 6418–6425.
  38. 38.B. Y. Lin, X. Chen, J. Chen, and X. Ren, “KagNet: Knowledge-aware graph networks for commonsense reasoning,” in EMNLP-IJCNLP, 2019, pp. 2829–2839.
  39. 39.D. Dai, L. Dong, Y. Hao, Z. Sui, B. Chang, and F. Wei, “Knowledge neurons in pretrained transformers,” arXiv preprint arXiv:2104.08696, 2021.
  40. 40.X. Wang, T. Gao, Z. Zhu, Z. Zhang, Z. Liu, J. Li, and J. Tang, “KEPLER: A unified model for knowledge embedding and pre-trained language representation,” Transactions of the Association for Computational Linguistics, vol. 9, pp. 176–194, 2021.
  41. 41.I. Melnyk, P. Dognin, and P. Das, “Grapher: Multi-stage knowledge graph construction using pretrained language models,” in NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021.
  42. 42.P. Ke, H. Ji, Y. Ran, X. Cui, L. Wang, L. Song, X. Zhu, and M. Huang, “JointGT: Graph-text joint representation learning for text generation from knowledge graphs,” in ACL Finding, 2021, pp. 2526–2538.
  43. 43.J. Jiang, K. Zhou, W. X. Zhao, and J.-R. Wen, “Unikgqa: Unified retrieval and reasoning for solving multi-hop question answering over knowledge graph,” ICLR 2023, 2023.
  44. 44.M. Yasunaga, A. Bosselut, H. Ren, X. Zhang, C. D. Manning, P. S. Liang, and J. Leskovec, “Deep bidirectional language-knowledge graph pretraining,” NeurIPS, vol. 35, pp. 37 309–37 323, 2022.
  45. 45.N. Choudhary and C. K. Reddy, “Complex logical reasoning over knowledge graphs using large language models,” arXiv preprint arXiv:2305.01157, 2023.
  46. 46.S. Wang, Z. Wei, J. Xu, and Z. Fan, “Unifying structure reasoning and language model pre-training for complex reasoning,” arXiv preprint arXiv:2301.08913, 2023.
  47. 47.C. Zhen, Y. Shang, X. Liu, Y. Li, Y. Chen, and D. Zhang, “A survey on knowledge-enhanced pre-trained language models,” arXiv preprint arXiv:2212.13428, 2022.
  48. 48.X. Wei, S. Wang, D. Zhang, P. Bhatia, and A. Arnold, “Knowledge enhanced pretrained language models: A compreshensive survey,” arXiv preprint arXiv:2110.08455, 2021.
  49. 49.D. Yin, L. Dong, H. Cheng, X. Liu, K.-W. Chang, F. Wei, and J. Gao, “A survey of knowledge-intensive nlp with pre-trained language models,” arXiv preprint arXiv:2202.08772, 2022.
  50. 50.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS, vol. 30, 2017.
  51. 51.Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “Albert: A lite bert for self-supervised learning of language representations,” in ICLR, 2019.
  52. 52.K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning, “Electra: Pre-training text encoders as discriminators rather than generators,” arXiv preprint arXiv:2003.10555, 2020.
  53. 53.K. Hakala and S. Pyysalo, “Biomedical named entity recognition with multilingual bert,” in Proceedings of the 5th workshop on BioNLP open shared tasks, 2019, pp. 56–61.
  54. 54.Y. Tay, M. Dehghani, V. Q. Tran, X. Garcia, J. Wei, X. Wang, H. W. Chung, D. Bahri, T. Schuster, S. Zheng et al., “Ul2: Unifying language learning paradigms,” in ICLR, 2022.
  55. 55.V. Sanh, A. Webson, C. Raffel, S. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, A. Raja, M. Dey et al., “Multitask prompted training enables zero-shot task generalization,” in ICLR, 2022.
  56. 56.B. Zoph, I. Bello, S. Kumar, N. Du, Y. Huang, J. Dean, N. Shazeer, and W. Fedus, “St-moe: Designing stable and transferable sparse expert models,” URL https://arxiv.org/abs/2202.08906, 2022.
  57. 57.A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y. Xu, W. Zheng, X. Xia, W. L. Tam, Z. Ma, Y. Xue, J. Zhai, W. Chen, Z. Liu, P. Zhang, Y. Dong, and J. Tang, “GLM-130b: An open bilingual pre-trained model,” in ICLR, 2023.
  58. 58.L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mt5: A massively multilingual pre-trained text-to-text transformer,” in NAACL, 2021, pp. 483–498.
  59. 59.T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language models are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020.
  60. 60.L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al., “Training language models to follow instructions with human feedback,” NeurIPS, vol. 35, pp. 27 730–27 744, 2022.
  61. 61.H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Roziere, N. Goyal, E. Hambro, F. Azhar ` et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023.
  62. 62.E. Saravia, “Prompt Engineering Guide,” https://github.com/dair-ai/Prompt-Engineering-Guide, 2022, accessed: 2022-12.
  63. 63.J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. H. Chi, Q. V. Le, D. Zhou et al., “Chain-of-thought prompting elicits reasoning in large language models,” in NeurIPS.
  64. 64.S. Li, Y. Gao, H. Jiang, Q. Yin, Z. Li, X. Yan, C. Zhang, and B. Yin, “Graph reasoning for question answering with triplet retrieval,” in ACL, 2023.
  65. 65.Y. Wen, Z. Wang, and J. Sun, “Mindmap: Knowledge graph prompting sparks graph of thoughts in large language models,” arXiv preprint arXiv:2308.09729, 2023.
  66. 66.K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor, “Freebase: A collaboratively created graph database for structuring human knowledge,” in SIGMOD, 2008, pp. 1247–1250.
  67. 67.S. Auer, C. Bizer, G. Kobilarov, J. Lehmann, R. Cyganiak, and Z. Ives, “Dbpedia: A nucleus for a web of open data,” in The Semantic Web: 6th International Semantic Web Conference. Springer, 2007, pp. 722–735.
  68. 68.B. Xu, Y. Xu, J. Liang, C. Xie, B. Liang, W. Cui, and Y. Xiao, “Cndbpedia: A never-ending chinese knowledge extraction system,” in 30th International Conference on Industrial Engineering and Other Applications of Applied Intelligent Systems. Springer, 2017, pp. 428–438.
  69. 69.P. Hai-Nyzhnyk, “Vikidia as a universal multilingual online encyclopedia for children,” The Encyclopedia Herald of Ukraine, vol. 14, 2022.
  70. 70.F. Ilievski, P. Szekely, and B. Zhang, “Cskg: The commonsense knowledge graph,” Extended Semantic Web Conference (ESWC), 2021.
  71. 71.R. Speer, J. Chin, and C. Havasi, “Conceptnet 5.5: An open multilingual graph of general knowledge,” in Proceedings of the AAAI conference on artificial intelligence, vol. 31, no. 1, 2017.
  72. 72.H. Ji, P. Ke, S. Huang, F. Wei, X. Zhu, and M. Huang, “Language generation with multi-hop reasoning on commonsense knowledge graph,” in EMNLP, 2020, pp. 725–736.
  73. 73.J. D. Hwang, C. Bhagavatula, R. Le Bras, J. Da, K. Sakaguchi, A. Bosselut, and Y. Choi, “(comet-) atomic 2020: On symbolic and neural commonsense knowledge graphs,” in AAAI, vol. 35, no. 7, 2021, pp. 6384–6392.
  74. 74.H. Zhang, X. Liu, H. Pan, Y. Song, and C. W.-K. Leung, “Aser: A large-scale eventuality knowledge graph,” in Proceedings of the web conference 2020, 2020, pp. 201–211.
  75. 75.H. Zhang, D. Khashabi, Y. Song, and D. Roth, “Transomcs: from linguistic graphs to commonsense knowledge,” in IJCAI, 2021, pp. 4004–4010.
  76. 76.Z. Li, X. Ding, T. Liu, J. E. Hu, and B. Van Durme, “Guided generation of cause and effect,” in IJCAI, 2020.
  77. 77.O. Bodenreider, “The unified medical language system (umls): integrating biomedical terminology,” Nucleic acids research, vol. 32, no. suppl 1, pp. D267–D270, 2004.
  78. 78.Y. Liu, Q. Zeng, J. Ordieres Mere, and H. Yang, “Anticipating stock market of the renowned companies: a knowledge graph approach,” Complexity, vol. 2019, 2019.
  79. 79.Y. Zhu, W. Zhou, Y. Xu, J. Liu, Y. Tan et al., “Intelligent learning for knowledge graph towards geological data,” Scientific Programming, vol. 2017, 2017.
  80. 80.W. Choi and H. Lee, “Inference of biomedical relations among chemicals, genes, diseases, and symptoms using knowledge representation learning,” IEEE Access, vol. 7, pp. 179 373–179 384, 2019.
  81. 81.F. Farazi, M. Salamanca, S. Mosbach, J. Akroyd, A. Eibeck, L. K. Aditya, A. Chadzynski, K. Pan, X. Zhou, S. Zhang et al., “Knowledge graph approach to combustion chemistry and interoperability,” ACS omega, vol. 5, no. 29, pp. 18 342–18 348, 2020.
  82. 82.X. Wu, T. Jiang, Y. Zhu, and C. Bu, “Knowledge graph for china’s genealogy,” IEEE TKDE, vol. 35, no. 1, pp. 634–646, 2023.
  83. 83.X. Zhu, Z. Li, X. Wang, X. Jiang, P. Sun, X. Wang, Y. Xiao, and N. J. Yuan, “Multi-modal knowledge graph construction and application: A survey,” IEEE TKDE, 2022.
  84. 84.S. Ferrada, B. Bustos, and A. Hogan, “Imgpedia: a linked dataset with content-based analysis of wikimedia images,” in The Semantic Web–ISWC 2017. Springer, 2017, pp. 84–93.
  85. 85.Y. Liu, H. Li, A. Garcia-Duran, M. Niepert, D. Onoro-Rubio, and D. S. Rosenblum, “Mmkg: multi-modal knowledge graphs,” in The Semantic Web: 16th International Conference, ESWC 2019, Portoroˇz, Slovenia, June 2–6, 2019, Proceedings 16. Springer, 2019, pp. 459–474.
  86. 86.M. Wang, H. Wang, G. Qi, and Q. Zheng, “Richpedia: a largescale, comprehensive multi-modal knowledge graph,” Big Data Research, vol. 22, p. 100159, 2020.
  87. 87.B. Shi, L. Ji, P. Lu, Z. Niu, and N. Duan, “Knowledge aware semantic concept expansion for image-text matching.” in IJCAI, vol. 1, 2019, p. 2.
  88. 88.S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar, “Kvqa: Knowledge-aware visual question answering,” in AAAI, vol. 33, no. 01, 2019, pp. 8876–8884.
  89. 89.R. Sun, X. Cao, Y. Zhao, J. Wan, K. Zhou, F. Zhang, Z. Wang, and K. Zheng, “Multi-modal knowledge graphs for recommender systems,” in CIKM, 2020, pp. 1405–1414.
  90. 90.S. Deng, C. Wang, Z. Li, N. Zhang, Z. Dai, H. Chen, F. Xiong, M. Yan, Q. Chen, M. Chen, J. Chen, J. Z. Pan, B. Hooi, and H. Chen, “Construction and applications of billion-scale pretrained multimodal business knowledge graph,” in ICDE, 2023.
  91. 91.C. Rosset, C. Xiong, M. Phan, X. Song, P. Bennett, and S. Tiwary, “Knowledge-aware language model pretraining,” arXiv preprint arXiv:2007.00655, 2020.
  92. 92.P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Kuttler, M. Lewis, W.-t. Yih, T. Rockt¨aschel, S. Riedel,¨ and D. Kiela, “Retrieval-augmented generation for knowledgeintensive nlp tasks,” in NeurIPS, vol. 33, 2020, pp. 9459–9474.
  93. 93.Y. Zhu, X. Wang, J. Chen, S. Qiao, Y. Ou, Y. Yao, S. Deng, H. Chen, and N. Zhang, “Llms for knowledge graph construction and reasoning: Recent capabilities and future opportunities,” arXiv preprint arXiv:2305.13168, 2023.
  94. 94.Z. Zhang, X. Liu, Y. Zhang, Q. Su, X. Sun, and B. He, “Pretrainkge: learning knowledge representation from pretrained language models,” in EMNLP Finding, 2020, pp. 259–266.
  95. 95.A. Kumar, A. Pandey, R. Gadia, and M. Mishra, “Building knowledge graph using pre-trained language model for learning entity-aware relationships,” in 2020 IEEE International Conference on Computing, Power and Communication Technologies (GUCON). IEEE, 2020, pp. 310–315.
  96. 96.X. Xie, N. Zhang, Z. Li, S. Deng, H. Chen, F. Xiong, M. Chen, and H. Chen, “From discrimination to generation: Knowledge graph completion with generative transformer,” in WWW, 2022, pp. 162–165.
  97. 97.Z. Chen, C. Xu, F. Su, Z. Huang, and Y. Dou, “Incorporating structured sentences with time-enhanced bert for fully-inductive temporal relation prediction,” SIGIR, 2023.
  98. 98.D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with advanced large language models,” arXiv preprint arXiv:2304.10592, 2023.
  99. 99.M. Warren, D. A. Shamma, and P. J. Hayes, “Knowledge engineering with image data in real-world settings,” in AAAI, ser. CEUR Workshop Proceedings, vol. 2846, 2021.
  100. 100.R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du et al., “Lamda: Language models for dialog applications,” arXiv preprint arXiv:2201.08239, 2022.
  101. 101.Y. Sun, S. Wang, S. Feng, S. Ding, C. Pang, J. Shang, J. Liu, X. Chen, Y. Zhao, Y. Lu et al., “Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation,” arXiv preprint arXiv:2107.02137, 2021.
  102. 102.T. Shen, Y. Mao, P. He, G. Long, A. Trischler, and W. Chen, “Exploiting structured knowledge in text via graph-guided representation learning,” in EMNLP, 2020, pp. 8980–8994.
  103. 103.D. Zhang, Z. Yuan, Y. Liu, F. Zhuang, H. Chen, and H. Xiong, “E-bert: A phrase and product knowledge enhanced language model for e-commerce,” arXiv preprint arXiv:2009.02835, 2020.
  104. 104.S. Li, X. Li, L. Shang, C. Sun, B. Liu, Z. Ji, X. Jiang, and Q. Liu, “Pre-training language models with deterministic factual knowledge,” in EMNLP, 2022, pp. 11 118–11 131.
  105. 105.M. Kang, J. Baek, and S. J. Hwang, “Kala: Knowledge-augmented language model adaptation,” in NAACL, 2022, pp. 5144–5167.
  106. 106.W. Xiong, J. Du, W. Y. Wang, and V. Stoyanov, “Pretrained encyclopedia: Weakly supervised knowledge-pretrained language model,” in ICLR, 2020.
  107. 107.T. Sun, Y. Shao, X. Qiu, Q. Guo, Y. Hu, X. Huang, and Z. Zhang, “CoLAKE: Contextualized language and knowledge embedding,” in Proceedings of the 28th International Conference on Computational Linguistics, 2020, pp. 3660–3670.
  108. 108.T. Zhang, C. Wang, N. Hu, M. Qiu, C. Tang, X. He, and J. Huang, “DKPLM: decomposable knowledge-enhanced pre-trained language model for natural language understanding,” in AAAI, 2022, pp. 11 703–11 711.
  109. 109.J. Wang, W. Huang, M. Qiu, Q. Shi, H. Wang, X. Li, and M. Gao, “Knowledge prompting in pre-trained language model for natural language understanding,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3164–3177.
  110. 110.H. Ye, N. Zhang, S. Deng, X. Chen, H. Chen, F. Xiong, X. Chen, and H. Chen, “Ontology-enhanced prompt-tuning for few-shot learning,” in Proceedings of the ACM Web Conference 2022, 2022, pp. 778–787.
  111. 111.H. Luo, Z. Tang, S. Peng, Y. Guo, W. Zhang, C. Ma, G. Dong, M. Song, W. Lin et al., “Chatkbqa: A generate-then-retrieve framework for knowledge base question answering with fine-tuned large language models,” arXiv preprint arXiv:2310.08975, 2023.
  112. 112.L. Luo, Y.-F. Li, G. Haffari, and S. Pan, “Reasoning on graphs: Faithful and interpretable large language model reasoning,” arXiv preprint arxiv:2310.01061, 2023.
  113. 113.R. Logan, N. F. Liu, M. E. Peters, M. Gardner, and S. Singh, “Barack’s wife hillary: Using knowledge graphs for fact-aware language modeling,” in ACL, 2019, pp. 5962–5971.
  114. 114.K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang, “Realm: Retrieval-augmented language model pre-training,” in ICML, 2020.
  115. 115.Y. Wu, Y. Zhao, B. Hu, P. Minervini, P. Stenetorp, and S. Riedel, “An efficient memory-augmented transformer for knowledgeintensive NLP tasks,” in EMNLP, 2022, pp. 5184–5196.
  116. 116.L. Luo, J. Ju, B. Xiong, Y.-F. Li, G. Haffari, and S. Pan, “Chatrule: Mining logical rules with large language models for knowledge graph reasoning,” arXiv preprint arXiv:2309.01538, 2023.
  117. 117.J. Wang, Q. Sun, N. Chen, X. Li, and M. Gao, “Boosting language models reasoning with chain-of-knowledge prompting,” arXiv preprint arXiv:2306.06427, 2023.
  118. 118.Z. Jiang, F. F. Xu, J. Araki, and G. Neubig, “How can we know what language models know?” Transactions of the Association for Computational Linguistics, vol. 8, pp. 423–438, 2020.
  119. 119.T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh, “Autoprompt: Eliciting knowledge from language models with automatically generated prompts,” arXiv preprint arXiv:2010.15980, 2020.
  120. 120.Z. Meng, F. Liu, E. Shareghi, Y. Su, C. Collins, and N. Collier, “Rewire-then-probe: A contrastive recipe for probing biomedical knowledge of pre-trained language models,” arXiv preprint arXiv:2110.08173, 2021.
  121. 121.L. Luo, T.-T. Vu, D. Phung, and G. Haffari, “Systematic assessment of factual knowledge in large language models,” in EMNLP, 2023.
  122. 122.V. Swamy, A. Romanou, and M. Jaggi, “Interpreting language models through knowledge graph extraction,” arXiv preprint arXiv:2111.08546, 2021.
  123. 123.S. Li, X. Li, L. Shang, Z. Dong, C. Sun, B. Liu, Z. Ji, X. Jiang, and Q. Liu, “How pre-trained language models capture factual knowledge? a causal-inspired analysis,” arXiv preprint arXiv:2203.16747, 2022.
  124. 124.H. Tian, C. Gao, X. Xiao, H. Liu, B. He, H. Wu, H. Wang, and F. Wu, “SKEP: Sentiment knowledge enhanced pre-training for sentiment analysis,” in ACL, 2020, pp. 4067–4076.
  125. 125.W. Yu, C. Zhu, Y. Fang, D. Yu, S. Wang, Y. Xu, M. Zeng, and M. Jiang, “Dict-BERT: Enhancing language model pre-training with dictionary,” in ACL, 2022, pp. 1907–1918.
  126. 126.T. McCoy, E. Pavlick, and T. Linzen, “Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference,” in ACL, 2019, pp. 3428–3448.
  127. 127.D. Wilmot and F. Keller, “Memory and knowledge augmented language models for inferring salience in long-form stories,” in EMNLP, 2021, pp. 851–865.
  128. 128.L. Adolphs, S. Dhuliawala, and T. Hofmann, “How to query language models?” arXiv preprint arXiv:2108.01928, 2021.
  129. 129.M. Sung, J. Lee, S. Yi, M. Jeon, S. Kim, and J. Kang, “Can language models be biomedical knowledge bases?” in EMNLP, 2021, pp. 4723–4734.
  130. 130.A. Mallen, A. Asai, V. Zhong, R. Das, H. Hajishirzi, and D. Khashabi, “When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories,” arXiv preprint arXiv:2212.10511, 2022.
  131. 131.M. Yasunaga, H. Ren, A. Bosselut, P. Liang, and J. Leskovec, “QAGNN: Reasoning with language models and knowledge graphs for question answering,” in NAACL, 2021, pp. 535–546.
  132. 132.M. Nayyeri, Z. Wang, M. Akter, M. M. Alam, M. R. A. H. Rony, J. Lehmann, S. Staab et al., “Integrating knowledge graph embedding and pretrained language models in hypercomplex spaces,” arXiv preprint arXiv:2208.02743, 2022.
  133. 133.N. Huang, Y. R. Deshpande, Y. Liu, H. Alberts, K. Cho, C. Vania, and I. Calixto, “Endowing language models with multimodal knowledge graph representations,” arXiv preprint arXiv:2206.13163, 2022.
  134. 134.M. M. Alam, M. R. A. H. Rony, M. Nayyeri, K. Mohiuddin, M. M. Akter, S. Vahdati, and J. Lehmann, “Language model guided knowledge graph embeddings,” IEEE Access, vol. 10, pp. 76 008–76 020, 2022.
  135. 135.X. Wang, Q. He, J. Liang, and Y. Xiao, “Language models as knowledge embeddings,” arXiv preprint arXiv:2206.12617, 2022.
  136. 136.N. Zhang, X. Xie, X. Chen, S. Deng, C. Tan, F. Huang, X. Cheng, and H. Chen, “Reasoning through memorization: Nearest neighbor knowledge graph embeddings,” arXiv preprint arXiv:2201.05575, 2022.
  137. 137.X. Xie, Z. Li, X. Wang, Y. Zhu, N. Zhang, J. Zhang, S. Cheng, B. Tian, S. Deng, F. Xiong, and H. Chen, “Lambdakg: A library for pre-trained language model-based knowledge graph embeddings,” 2022.
  138. 138.B. Kim, T. Hong, Y. Ko, and J. Seo, “Multi-task learning for knowledge graph completion with pre-trained language models,” in COLING, 2020, pp. 1737–1743.
  139. 139.X. Lv, Y. Lin, Y. Cao, L. Hou, J. Li, Z. Liu, P. Li, and J. Zhou, “Do pre-trained models benefit knowledge graph completion? A reliable evaluation and a reasonable approach,” in ACL, 2022, pp. 3570–3581.
  140. 140.J. Shen, C. Wang, L. Gong, and D. Song, “Joint language semantic and structure embedding for knowledge graph completion,” in COLING, 2022, pp. 1965–1978.
  141. 141.B. Choi, D. Jang, and Y. Ko, “MEM-KGC: masked entity model for knowledge graph completion with pre-trained language model,” IEEE Access, vol. 9, pp. 132 025–132 032, 2021.
  142. 142.B. Choi and Y. Ko, “Knowledge graph extension with a pretrained language model via unified learning method,” Knowl. Based Syst., vol. 262, p. 110245, 2023.
  143. 143.B. Wang, T. Shen, G. Long, T. Zhou, Y. Wang, and Y. Chang, “Structure-augmented text representation learning for efficient knowledge graph completion,” in WWW, 2021, pp. 1737–1748.
  144. 144.L. Wang, W. Zhao, Z. Wei, and J. Liu, “Simkgc: Simple contrastive knowledge graph completion with pre-trained language models,” in ACL, 2022, pp. 4281–4294.
  145. 145.D. Li, M. Yi, and Y. He, “Lp-bert: Multi-task pre-training knowledge graph bert for link prediction,” arXiv preprint arXiv:2201.04843, 2022.
  146. 146.A. Saxena, A. Kochsiek, and R. Gemulla, “Sequence-to-sequence knowledge graph completion and question answering,” in ACL, 2022, pp. 2814–2828.
  147. 147.C. Chen, Y. Wang, B. Li, and K. Lam, “Knowledge is flat: A seq2seq generative framework for various knowledge graph completion,” in COLING, 2022, pp. 4005–4017.
  148. 148.M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in NAACL, 2018, pp. 2227–2237.
  149. 149.H. Yan, T. Gui, J. Dai, Q. Guo, Z. Zhang, and X. Qiu, “A unified generative framework for various NER subtasks,” in ACL, 2021, pp. 5808–5822.
  150. 150.Y. Onoe and G. Durrett, “Learning to denoise distantly-labeled data for entity typing,” in NAACL, 2019, pp. 2407–2417.
  151. 151.Y. Onoe, M. Boratko, A. McCallum, and G. Durrett, “Modeling fine-grained entity types with box embeddings,” in ACL, 2021, pp. 2051–2064.
  152. 152.B. Z. Li, S. Min, S. Iyer, Y. Mehdad, and W. Yih, “Efficient onepass end-to-end entity linking for questions,” in EMNLP, 2020, pp. 6433–6441.
  153. 153.T. Ayoola, S. Tyagi, J. Fisher, C. Christodoulopoulos, and A. Pierleoni, “Refined: An efficient zero-shot-capable approach to endto-end entity linking,” in NAACL, 2022, pp. 209–220.
  154. 154.M. Joshi, O. Levy, L. Zettlemoyer, and D. S. Weld, “BERT for coreference resolution: Baselines and analysis,” in EMNLP, 2019, pp. 5802–5807.
  155. 155.M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy, “Spanbert: Improving pre-training by representing and predicting spans,” Trans. Assoc. Comput. Linguistics, vol. 8, pp. 64–77, 2020.
  156. 156.A. Caciularu, A. Cohan, I. Beltagy, M. E. Peters, A. Cattan, and I. Dagan, “CDLM: cross-document language modeling,” in EMNLP, 2021, pp. 2648–2662.
  157. 157.A. Cattan, A. Eirew, G. Stanovsky, M. Joshi, and I. Dagan, “Crossdocument coreference resolution over predicted mentions,” in ACL, 2021, pp. 5100–5107.
  158. 158.Y. Wang, Y. Shen, and H. Jin, “An end-to-end actor-critic-based neural coreference resolution system,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2021, Toronto, ON, Canada, June 6-11, 2021, 2021, pp. 7848–7852.
  159. 159.P. Shi and J. Lin, “Simple BERT models for relation extraction and semantic role labeling,” CoRR, vol. abs/1904.05255, 2019.
  160. 160.S. Park and H. Kim, “Improving sentence-level relation extraction through curriculum learning,” CoRR, vol. abs/2107.09332, 2021.
  161. 161.Y. Ma, A. Wang, and N. Okazaki, “DREEAM: guiding attention with evidence for improving document-level relation extraction,” in EACL, 2023, pp. 1963–1975.
  162. 162.Q. Guo, Y. Sun, G. Liu, Z. Wang, Z. Ji, Y. Shen, and X. Wang, “Constructing chinese historical literature knowledge graph based on bert,” in Web Information Systems and Applications: 18th International Conference, WISA 2021, Kaifeng, China, September 24–26, 2021, Proceedings 18. Springer, 2021, pp. 323–334.
  163. 163.J. Han, N. Collier, W. Buntine, and E. Shareghi, “Pive: Prompting with iterative verification improving graph-based generative capability of llms,” arXiv preprint arXiv:2305.12392, 2023.
  164. 164.A. Bosselut, H. Rashkin, M. Sap, C. Malaviya, A. Celikyilmaz, and Y. Choi, “Comet: Commonsense transformers for knowledge graph construction,” in ACL, 2019.
  165. 165.S. Hao, B. Tan, K. Tang, H. Zhang, E. P. Xing, and Z. Hu, “Bertnet: Harvesting knowledge graphs from pretrained language models,” arXiv preprint arXiv:2206.14268, 2022.
  166. 166.P. West, C. Bhagavatula, J. Hessel, J. Hwang, L. Jiang, R. Le Bras, X. Lu, S. Welleck, and Y. Choi, “Symbolic knowledge distillation: from general language models to commonsense models,” in NAACL, 2022, pp. 4602–4625.
  167. 167.L. F. R. Ribeiro, M. Schmitt, H. Schutze, and I. Gurevych, “Investi-¨ gating pretrained language models for graph-to-text generation,” in Proceedings of the 3rd Workshop on Natural Language Processing for Conversational AI, 2021, pp. 211–227.
  168. 168.J. Li, T. Tang, W. X. Zhao, Z. Wei, N. J. Yuan, and J.-R. Wen, “Few-shot knowledge graph-to-text generation with pretrained language models,” in ACL, 2021, pp. 1558–1568.
  169. 169.A. Colas, M. Alvandipour, and D. Z. Wang, “GAP: A graphaware language model framework for knowledge graph-to-text generation,” in Proceedings of the 29th International Conference on Computational Linguistics, 2022, pp. 5755–5769.
  170. 170.Z. Jin, Q. Guo, X. Qiu, and Z. Zhang, “GenWiki: A dataset of 1.3 million content-sharing text and graphs for unsupervised graph-to-text generation,” in Proceedings of the 28th International Conference on Computational Linguistics, 2020, pp. 2398–2409.
  171. 171.W. Chen, Y. Su, X. Yan, and W. Y. Wang, “KGPT: Knowledgegrounded pre-training for data-to-text generation,” in EMNLP, 2020, pp. 8635–8648.
  172. 172.D. Lukovnikov, A. Fischer, and J. Lehmann, “Pretrained transformers for simple question answering over knowledge graphs,” in The Semantic Web–ISWC 2019: 18th International Semantic Web Conference, Auckland, New Zealand, October 26–30, 2019, Proceedings, Part I 18. Springer, 2019, pp. 470–486.
  173. 173.D. Luo, J. Su, and S. Yu, “A bert-based approach with relationaware attention for knowledge base question answering,” in IJCNN. IEEE, 2020, pp. 1–8.
  174. 174.N. Hu, Y. Wu, G. Qi, D. Min, J. Chen, J. Z. Pan, and Z. Ali, “An empirical study of pre-trained language models in simple knowledge graph question answering,” arXiv preprint arXiv:2303.10368, 2023.
  175. 175.Y. Xu, C. Zhu, R. Xu, Y. Liu, M. Zeng, and X. Huang, “Fusing context into knowledge graph for commonsense question answering,” in ACL, 2021, pp. 1201–1207.
  176. 176.M. Zhang, R. Dai, M. Dong, and T. He, “Drlk: Dynamic hierarchical reasoning with language model and knowledge graph for question answering,” in EMNLP, 2022, pp. 5123–5133.
  177. 177.Z. Hu, Y. Xu, W. Yu, S. Wang, Z. Yang, C. Zhu, K.-W. Chang, and Y. Sun, “Empowering language models with knowledge graph reasoning for open-domain question answering,” in EMNLP, 2022, pp. 9562–9581.
  178. 178.X. Zhang, A. Bosselut, M. Yasunaga, H. Ren, P. Liang, C. D. Manning, and J. Leskovec, “Greaselm: Graph reasoning enhanced language models,” in ICLR, 2022.
  179. 179.X. Cao and Y. Liu, “Relmkg: reasoning with pre-trained language models and knowledge graphs for complex question answering,” Applied Intelligence, pp. 1–15, 2022.
  180. 180.X. Huang, J. Zhang, D. Li, and P. Li, “Knowledge graph embedding based question answering,” in WSDM, 2019, pp. 105–113.
  181. 181.H. Wang, F. Zhang, X. Xie, and M. Guo, “Dkn: Deep knowledgeaware network for news recommendation,” in WWW, 2018, pp. 1835–1844.
  182. 182.B. Yang, S. W.-t. Yih, X. He, J. Gao, and L. Deng, “Embedding entities and relations for learning and inference in knowledge bases,” in ICLR, 2015.
  183. 183.W. Xiong, M. Yu, S. Chang, X. Guo, and W. Y. Wang, “One-shot relational learning for knowledge graphs,” in EMNLP, 2018, pp. 1980–1990.
  184. 184.P. Wang, J. Han, C. Li, and R. Pan, “Logic attention based neighborhood aggregation for inductive knowledge graph embedding,” in AAAI, vol. 33, no. 01, 2019, pp. 7152–7159.
  185. 185.Y. Lin, Z. Liu, M. Sun, Y. Liu, and X. Zhu, “Learning entity and relation embeddings for knowledge graph completion,” in Proceedings of the AAAI conference on artificial intelligence, vol. 29, no. 1, 2015.
  186. 186.C. Chen, Y. Wang, A. Sun, B. Li, and L. Kwok-Yan, “Dipping plms sauce: Bridging structure and text for effective knowledge graph completion via conditional soft prompting,” in ACL, 2023.
  187. 187.J. Lovelace and C. P. Rose, “A framework for adapting pre-´ trained language models to knowledge graph completion,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, 2022, pp. 5937–5955.
  188. 188.J. Fu, L. Feng, Q. Zhang, X. Huang, and P. Liu, “Larger-context tagging: When and why does it work?” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACLHLT 2021, Online, June 6-11, 2021, 2021, pp. 1463–1475.
  189. 189.X. Liu, K. Ji, Y. Fu, Z. Du, Z. Yang, and J. Tang, “P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks,” CoRR, vol. abs/2110.07602, 2021.
  190. 190.J. Yu, B. Bohnet, and M. Poesio, “Named entity recognition as dependency parsing,” in ACL, 2020, pp. 6470–6476.
  191. 191.F. Li, Z. Lin, M. Zhang, and D. Ji, “A span-based model for joint overlapped and discontinuous named entity recognition,” in ACL, 2021, pp. 4814–4828.
  192. 192.C. Tan, W. Qiu, M. Chen, R. Wang, and F. Huang, “Boundary enhanced neural span classification for nested named entity recognition,” in The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, 2020, pp. 9016–9023.
  193. 193.Y. Xu, H. Huang, C. Feng, and Y. Hu, “A supervised multi-head self-attention network for nested named entity recognition,” in Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, 2021, pp. 14 185–14 193.
  194. 194.J. Yu, B. Ji, S. Li, J. Ma, H. Liu, and H. Xu, “S-NER: A concise and efficient span-based model for named entity recognition,” Sensors, vol. 22, no. 8, p. 2852, 2022.
  195. 195.Y. Fu, C. Tan, M. Chen, S. Huang, and F. Huang, “Nested named entity recognition with partially-observed treecrfs,” in AAAI, 2021, pp. 12 839–12 847.
  196. 196.C. Lou, S. Yang, and K. Tu, “Nested named entity recognition as latent lexicalized constituency parsing,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, 2022, pp. 6183–6198.
  197. 197.S. Yang and K. Tu, “Bottom-up constituency parsing and nested named entity recognition with pointer networks,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, 2022, pp. 2403–2416.
  198. 198.F. Li, Z. Lin, M. Zhang, and D. Ji, “A span-based model for joint overlapped and discontinuous named entity recognition,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, 2021, pp. 4814–4828.
  199. 199.Q. Liu, H. Lin, X. Xiao, X. Han, L. Sun, and H. Wu, “Fine-grained entity typing via label reasoning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, 2021, pp. 4611–4622.
  200. 200.H. Dai, Y. Song, and H. Wang, “Ultra-fine entity typing with weak supervision from a masked language model,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, 2021, pp. 1790–1799.
  201. 201.N. Ding, Y. Chen, X. Han, G. Xu, X. Wang, P. Xie, H. Zheng, Z. Liu, J. Li, and H. Kim, “Prompt-learning for fine-grained entity typing,” in Findings of the Association for Computational Linguistics: EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, 2022, pp. 6888–6901.
  202. 202.W. Pan, W. Wei, and F. Zhu, “Automatic noisy label correction for fine-grained entity typing,” in Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI 2022, Vienna, Austria, 23-29 July 2022, 2022, pp. 4317–4323.
  203. 203.B. Li, W. Yin, and M. Chen, “Ultra-fine entity typing with indirect supervision from natural language inference,” Trans. Assoc. Comput. Linguistics, vol. 10, pp. 607–622, 2022.
  204. 204.S. Broscheit, “Investigating entity knowledge in BERT with simple neural end-to-end entity linking,” CoRR, vol. abs/2003.05473, 2020.
  205. 205.N. D. Cao, G. Izacard, S. Riedel, and F. Petroni, “Autoregressive entity retrieval,” in 9th ICLR, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, 2021.
  206. 206.N. D. Cao, L. Wu, K. Popat, M. Artetxe, N. Goyal, M. Plekhanov, L. Zettlemoyer, N. Cancedda, S. Riedel, and F. Petroni, “Multilingual autoregressive entity linking,” Trans. Assoc. Comput. Linguistics, vol. 10, pp. 274–290, 2022.
  207. 207.N. D. Cao, W. Aziz, and I. Titov, “Highly parallel autoregressive entity linking with discriminative correction,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, 2021, pp. 7662–7669.
  208. 208.K. Lee, L. He, and L. Zettlemoyer, “Higher-order coreference resolution with coarse-to-fine inference,” in NAACL, 2018, pp. 687–692.
  209. 209.T. M. Lai, T. Bui, and D. S. Kim, “End-to-end neural coreference resolution revisited: A simple yet effective baseline,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Singapore, 23-27 May 2022, 2022, pp. 8147–8151.
  210. 210.W. Wu, F. Wang, A. Yuan, F. Wu, and J. Li, “Corefqa: Coreference resolution as query-based span prediction,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, 2020, pp. 6953–6963.
  211. 211.T. M. Lai, H. Ji, T. Bui, Q. H. Tran, F. Dernoncourt, and W. Chang, “A context-dependent gated module for incorporating symbolic semantics into event coreference resolution,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACLHLT 2021, Online, June 6-11, 2021, 2021, pp. 3491–3499.
  212. 212.Y. Kirstain, O. Ram, and O. Levy, “Coreference resolution without span representations,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 2: Short Papers), Virtual Event, August 1-6, 2021, 2021, pp. 14–19.
  213. 213.R. Thirukovalluru, N. Monath, K. Shridhar, M. Zaheer, M. Sachan, and A. McCallum, “Scaling within document coreference to long texts,” in Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, ser. Findings of ACL, vol. ACL/IJCNLP 2021, 2021, pp. 3921–3931.
  214. 214.I. Beltagy, M. E. Peters, and A. Cohan, “Longformer: The longdocument transformer,” CoRR, vol. abs/2004.05150, 2020.
  215. 215.C. Alt, M. Hubner, and L. Hennig, “Improving relation extraction¨ by pre-trained language representations,” in 1st Conference on Automated Knowledge Base Construction, AKBC 2019, Amherst, MA, USA, May 20-22, 2019, 2019.
  216. 216.L. B. Soares, N. FitzGerald, J. Ling, and T. Kwiatkowski, “Matching the blanks: Distributional similarity for relation learning,” in ACL, 2019, pp. 2895–2905.
  217. 217.S. Lyu and H. Chen, “Relation classification with entity type restriction,” in Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, ser. Findings of ACL, vol. ACL/IJCNLP 2021, 2021, pp. 390–395.
  218. 218.J. Zheng and Z. Chen, “Sentence-level relation extraction via contrastive learning with descriptive relation prompts,” CoRR, vol. abs/2304.04935, 2023.
  219. 219.H. Wang, C. Focke, R. Sylvester, N. Mishra, and W. Y. Wang, “Fine-tune bert for docred with two-step process,” CoRR, vol. abs/1909.11898, 2019.
  220. 220.H. Tang, Y. Cao, Z. Zhang, J. Cao, F. Fang, S. Wang, and P. Yin, “HIN: hierarchical inference network for document-level relation extraction,” in PAKDD, ser. Lecture Notes in Computer Science, vol. 12084, 2020, pp. 197–209.
  221. 221.D. Wang, W. Hu, E. Cao, and W. Sun, “Global-to-local neural networks for document-level relation extraction,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, 2020, pp. 3711–3721.
  222. 222.S. Zeng, Y. Wu, and B. Chang, “SIRE: separate intra- and inter-sentential reasoning for document-level relation extraction,” in Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, ser. Findings of ACL, vol. ACL/IJCNLP 2021, 2021, pp. 524–534.
  223. 223.G. Nan, Z. Guo, I. Sekulic, and W. Lu, “Reasoning with latent structure refinement for document-level relation extraction,” in ACL, 2020, pp. 1546–1557.
  224. 224.S. Zeng, R. Xu, B. Chang, and L. Li, “Double graph based reasoning for document-level relation extraction,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, 2020, pp. 1630–1640.
  225. 225.N. Zhang, X. Chen, X. Xie, S. Deng, C. Tan, M. Chen, F. Huang, L. Si, and H. Chen, “Document-level relation extraction as semantic segmentation,” in IJCAI, 2021, pp. 3999–4006.
  226. 226.O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015 - 18th International Conference Munich, Germany, October 5 - 9, 2015, Proceedings, Part III, ser. Lecture Notes in Computer Science, vol. 9351, 2015, pp. 234–241.
  227. 227.W. Zhou, K. Huang, T. Ma, and J. Huang, “Document-level relation extraction with adaptive thresholding and localized context pooling,” in AAAI, 2021, pp. 14 612–14 620.
  228. 228.C. Gardent, A. Shimorina, S. Narayan, and L. Perez-Beltrachini, “The WebNLG challenge: Generating text from RDF data,” in Proceedings of the 10th International Conference on Natural Language Generation, 2017, pp. 124–133.
  229. 229.J. Guan, Y. Wang, and M. Huang, “Story ending generation with incremental encoding and commonsense knowledge,” in AAAI, 2019, pp. 6473–6480.
  230. 230.H. Zhou, T. Young, M. Huang, H. Zhao, J. Xu, and X. Zhu, “Commonsense knowledge aware conversation generation with graph attention,” in IJCAI, 2018, pp. 4623–4629.
  231. 231.M. Kale and A. Rastogi, “Text-to-text pre-training for data-to-text tasks,” in Proceedings of the 13th International Conference on Natural Language Generation, 2020, pp. 97–102.
  232. 232.M. Mintz, S. Bills, R. Snow, and D. Jurafsky, “Distant supervision for relation extraction without labeled data,” in ACL, 2009, pp. 1003–1011.
  233. 233.A. Saxena, A. Tripathi, and P. Talukdar, “Improving multi-hop question answering over knowledge graphs using knowledge base embeddings,” in ACL, 2020, pp. 4498–4507.
  234. 234.Y. Feng, X. Chen, B. Y. Lin, P. Wang, J. Yan, and X. Ren, “Scalable multi-hop relational reasoning for knowledge-aware question answering,” in EMNLP, 2020, pp. 1295–1309.
  235. 235.Y. Yan, R. Li, S. Wang, H. Zhang, Z. Daoguang, F. Zhang, W. Wu, and W. Xu, “Large-scale relation learning for question answering over knowledge bases with pre-trained language models,” in EMNLP, 2021, pp. 3653–3660.
  236. 236.J. Zhang, X. Zhang, J. Yu, J. Tang, J. Tang, C. Li, and H. Chen, “Subgraph retrieval enhanced model for multi-hop knowledge base question answering,” in ACL (Volume 1: Long Papers), 2022, pp. 5773–5784.
  237. 237.J. Jiang, K. Zhou, Z. Dong, K. Ye, W. X. Zhao, and J.-R. Wen, “Structgpt: A general framework for large language model to reason over structured data,” arXiv preprint arXiv:2305.09645, 2023.
  238. 238.H. Zhu, H. Peng, Z. Lyu, L. Hou, J. Li, and J. Xiao, “Pre-training language model incorporating domain-specific heterogeneous knowledge into a unified representation,” Expert Systems with Applications, vol. 215, p. 119369, 2023.
  239. 239.C. Feng, X. Zhang, and Z. Fei, “Knowledge solver: Teaching llms to search for domain knowledge from knowledge graphs,” arXiv preprint arXiv:2309.03118, 2023.
  240. 240.J. Sun, C. Xu, L. Tang, S. Wang, C. Lin, Y. Gong, H.-Y. Shum, and J. Guo, “Think-on-graph: Deep and responsible reasoning of large language model with knowledge graph,” arXiv preprint arXiv:2307.07697, 2023.
  241. 241.B. He, D. Zhou, J. Xiao, X. Jiang, Q. Liu, N. J. Yuan, and T. Xu, “BERT-MK: Integrating graph contextualized knowledge into pre-trained language models,” in EMNLP, 2020, pp. 2281–2290.
  242. 242.Y. Su, X. Han, Z. Zhang, Y. Lin, P. Li, Z. Liu, J. Zhou, and M. Sun, “Cokebert: Contextual knowledge selection and embedding towards enhanced pre-trained language models,” AI Open, vol. 2, pp. 127–134, 2021.
  243. 243.D. Yu, C. Zhu, Y. Yang, and M. Zeng, “JAKET: joint pre-training of knowledge graph and language understanding,” in AAAI, 2022, pp. 11 630–11 638.
  244. 244.X. Wang, P. Kapanipathi, R. Musa, M. Yu, K. Talamadupula, I. Abdelaziz, M. Chang, A. Fokoue, B. Makni, N. Mattei, and M. Witbrock, “Improving natural language inference using external knowledge in the science questions domain,” in AAAI, 2019, pp. 7208–7215.
  245. 245.Y. Sun, Q. Shi, L. Qi, and Y. Zhang, “JointLK: Joint reasoning with language models and knowledge graphs for commonsense question answering,” in NAACL, 2022, pp. 5049–5060.
  246. 246.X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang et al., “Agentbench: Evaluating llms as agents,” arXiv preprint arXiv:2308.03688, 2023.
  247. 247.Y. Wang, N. Lipka, R. A. Rossi, A. Siu, R. Zhang, and T. Derr, “Knowledge graph prompting for multi-document question answering,” arXiv preprint arXiv:2308.11730, 2023.
  248. 248.A. Zeng, M. Liu, R. Lu, B. Wang, X. Liu, Y. Dong, and J. Tang, “Agenttuning: Enabling generalized agent abilities for llms,” 2023.
  249. 249.W. Krysci´nski, B. McCann, C. Xiong, and R. Socher, “Evaluating´ the factual consistency of abstractive text summarization,” arXiv preprint arXiv:1910.12840, 2019.
  250. 250.Z. Ji, Z. Liu, N. Lee, T. Yu, B. Wilie, M. Zeng, and P. Fung, “Rho (\ρ): Reducing hallucination in open-domain dialogues with knowledge grounding,” arXiv preprint arXiv:2212.01588, 2022.
  251. 251.S. Feng, V. Balachandran, Y. Bai, and Y. Tsvetkov, “Factkb: Generalizable factuality evaluation using language models enhanced with factual knowledge,” arXiv preprint arXiv:2305.08281, 2023.
  252. 252.Y. Yao, P. Wang, B. Tian, S. Cheng, Z. Li, S. Deng, H. Chen, and N. Zhang, “Editing large language models: Problems, methods, and opportunities,” arXiv preprint arXiv:2305.13172, 2023.
  253. 253.Z. Li, N. Zhang, Y. Yao, M. Wang, X. Chen, and H. Chen, “Unveiling the pitfalls of knowledge editing for large language models,” arXiv preprint arXiv:2310.02129, 2023.
  254. 254.R. Cohen, E. Biran, O. Yoran, A. Globerson, and M. Geva, “Evaluating the ripple effects of knowledge editing in language models,” arXiv preprint arXiv:2307.12976, 2023.
  255. 255.S. Diao, Z. Huang, R. Xu, X. Li, Y. Lin, X. Zhou, and T. Zhang, “Black-box prompt learning for pre-trained language models,” arXiv preprint arXiv:2201.08531, 2022.
  256. 256.T. Sun, Y. Shao, H. Qian, X. Huang, and X. Qiu, “Black-box tuning for language-model-as-a-service,” in International Conference on Machine Learning. PMLR, 2022, pp. 20 841–20 855.
  257. 257.X. Chen, A. Shrivastava, and A. Gupta, “NEIL: extracting visual knowledge from web data,” in IEEE International Conference on Computer Vision, ICCV 2013, Sydney, Australia, December 1-8, 2013, 2013, pp. 1409–1416.
  258. 258.M. Warren and P. J. Hayes, “Bounding ambiguity: Experiences with an image annotation system,” in Proceedings of the 1st Workshop on Subjectivity, Ambiguity and Disagreement in Crowdsourcing, ser. CEUR Workshop Proceedings, vol. 2276, 2018, pp. 41–54.
  259. 259.Z. Chen, Y. Huang, J. Chen, Y. Geng, Y. Fang, J. Z. Pan, N. Zhang, and W. Zhang, “Lako: Knowledge-driven visual estion answering via late knowledge-to-text injection,” 2022.
  260. 260.R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” in ICCV, 2023, pp. 15 180–15 190.
  261. 261.J. Zhang, Z. Yin, P. Chen, and S. Nichele, “Emotion recognition using multi-modal data and machine learning techniques: A tutorial and review,” Information Fusion, vol. 59, pp. 103–126, 2020.
  262. 262.H. Zhang, B. Wu, X. Yuan, S. Pan, H. Tong, and J. Pei, “Trustworthy graph neural networks: Aspects, methods and trends,” arXiv:2205.07424, 2022.
  263. 263.T. Wu, M. Caccia, Z. Li, Y.-F. Li, G. Qi, and G. Haffari, “Pretrained language model in continual learning: A comparative study,” in ICLR, 2022.
  264. 264.X. L. Li, A. Kuncoro, J. Hoffmann, C. de Masson d’Autume, P. Blunsom, and A. Nematzadeh, “A systematic investigation of commonsense knowledge in large language models,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 11 838–11 855.
  265. 265.Y. Zheng, H. Y. Koh, J. Ju, A. T. Nguyen, L. T. May, G. I. Webb, and S. Pan, “Large language models for scientific synthesis, inference and explanation,” arXiv preprint arXiv:2310.07984, 2023.
  266. 266.B. Min, H. Ross, E. Sulem, A. P. B. Veyseh, T. H. Nguyen, O. Sainz, E. Agirre, I. Heintz, and D. Roth, “Recent advances in natural language processing via large pre-trained language models: A survey,” ACM Computing Surveys, vol. 56, no. 2, pp. 1–40, 2023.
  267. 267.J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” in International Conference on Learning Representations, 2021.
  268. 268.Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen, L. Wang, A. T. Luu, W. Bi, F. Shi, and S. Shi, “Siren’s song in the ai ocean: A survey on hallucination in large language models,” arXiv preprint arXiv:2309.01219, 2023.

Citation

MLA
Pan, S., et al. “Unifying Large Language Models and Knowledge Graphs: A Roadmap”. IEEE Transactions on Knowledge and Data Engineering, vol. 36, no. 7, 2024, pp. 3580–99, https://doi.org/10.1109/TKDE.2024.3352100.
APA
Pan, S., Luo, L., Wang, Y., Chen, C., Wang, J., & Wu, X. (2024). Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE Transactions on Knowledge and Data Engineering, 36(7), 3580–3599. https://doi.org/10.1109/TKDE.2024.3352100
Chicago
Pan, S., L. Luo, Y. Wang, C. Chen, J. Wang, and X. Wu. 2024. “Unifying Large Language Models and Knowledge Graphs: A Roadmap”. IEEE Transactions on Knowledge and Data Engineering 36 (7): 3580–99. https://doi.org/10.1109/TKDE.2024.3352100.
Harvard
Pan, S. et al. (2024) “Unifying Large Language Models and Knowledge Graphs: A Roadmap”, IEEE Transactions on Knowledge and Data Engineering, 36(7), pp. 3580–3599. Available at: https://doi.org/10.1109/TKDE.2024.3352100.
Vancouver
1. Pan S, Luo L, Wang Y, Chen C, Wang J, Wu X (2024) Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE Transactions on Knowledge and Data Engineering 36:3580–3599

BibTeX

@article{Pan_2024, title={Unifying Large Language Models and Knowledge Graphs: A Roadmap}, volume={36}, ISSN={2326-3865}, url={http://dx.doi.org/10.1109/TKDE.2024.3352100}, DOI={10.1109/tkde.2024.3352100}, number={7}, journal={IEEE Transactions on Knowledge and Data Engineering}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Pan, Shirui and Luo, Linhao and Wang, Yufei and Chen, Chen and Wang, Jiapu and Wu, Xindong}, year={2024}, month=July, pages={3580–3599} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF