Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey

Garima AgrawalTharindu KumarageZeyad AlghamdiHuan Liu

article2024NAACL228 citations

Presents a systematic taxonomy and evaluation of knowledge graph augmentation techniques across inference, training, and validation stages to effectively curb hallucinations and improve factual reasoning in large language models.

Listen

Large language models often generate plausible yet incorrect or fabricated information, commonly referred to as hallucinations. These errors stem from training data gaps, ambiguous context, and the models' statistical text-generation mechanisms. As organizations increasingly deploy these models for decision-making and operational tasks, hallucinations introduce significant reliability, safety, and compliance risks. The article comprehensively evaluates how integrating structured external databases, known as knowledge graphs, can mitigate hallucinations and enhance the reasoning capabilities of large language models.

The article conducts a systematic literature review of studies published between 2019 and 2023 across leading natural language processing venues. It categorizes knowledge graph augmentation methodologies into three distinct lifecycle stages: inference, training, and validation. In knowledge-aware inference, structured facts are retrieved at run time, integrated into step-by-step reasoning prompts, or used to enforce operational constraints without retraining model parameters. Knowledge-aware training incorporates structured relational facts directly into models via pre-training or fine-tuning. Knowledge-aware validation leverages graphs as post-generation reference sources to verify claims and guide fact-checking.

The analysis reveals several key findings regarding performance and operational trade-offs. First, for question-answering tasks, augmenting smaller models with retrieved knowledge graph facts improves answer correctness by more than 80%, outperforming the baseline performance gains achieved merely by increasing model size. Second, integrating structured graph pathways into multi-step reasoning significantly enhances larger models; for example, knowledge graph-augmented reasoning boosted baseline accuracy from 66.8% to 85.7% on complex reasoning tasks and reached 88.2% accuracy in clinical diagnosis benchmarks. Third, while training-stage integrations effectively specialize models, pre-training and fine-tuning demand substantial computational resources and produce rigid, task-specific systems with limited transferability. Consequently, the research landscape has shifted away from resource-intensive pre-training toward inference-time retrieval, reasoning, and validation frameworks that avoid additional training costs.

These findings suggest that organizations can achieve higher accuracy and reduce hallucination risks without undertaking costly model training from scratch. Implementing inference-time retrieval and reasoning guardrails allows enterprises to maintain smaller, cost-effective models while enforcing factual consistency through structured business data. However, post-generation validation frameworks introduce additional computational overhead and may still fail to catch subtle errors if the underlying knowledge base is incomplete or biased.

Decision-makers should prioritize modular, inference-time knowledge graph integrations for general and knowledge-intensive workflows, reserving parameter fine-tuning only for narrow domains with stable data. Future development should focus on building dynamic, multi-modal, and bias-resistant knowledge graphs, exploring causality-aware representations, and establishing bidirectional synergies where models and graphs refine each other. Readers should maintain cautious confidence in these evaluations, as baseline benchmarks and model capabilities evolve rapidly, and retrieval-based performance remains strictly bounded by the scope and quality of the underlying graph data.

arXiv: 2311.07914
Cover for Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey

Abstract

The contemporary LLMs are prone to produ-cing hallucinations, stemming mainly from the knowledge gaps within the models. To address this critical limitation, researchers employ di-verse strategies to augment the LLMs by incor-porating external knowledge, aiming to reduce hallucinations and enhance reasoning accuracy. Among these strategies, leveraging knowledge graphs as a source of external information has demonstrated promising results. In this survey, we comprehensively review these knowledge-graph-based augmentation techniques in LLMs, focusing on their efficacy in mitigating hallu-cinations. We systematically categorize these methods into three overarching groups, offering methodological comparisons and performance evaluations. Lastly, this survey explores the current trends and challenges associated with these techniques and outlines potential avenues for future research in this emerging field.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Large Language Models
  • 2.2 Knowledge Graphs
  • 3 Knowledge Graph-Enhanced LLMs
  • 3.1 Knowledge-Aware Inference
  • 3.1.1 KG-Augmented Retrieval
  • 3.1.2 KG-Augmented Reasoning
  • 3.1.3 Knowledge-Controlled Generation
  • 3.2 Knowledge-Aware Training
  • 3.2.1 Knowledge-Aware Pre-Training
  • 3.2.2 Knowledge-Aware Fine-Tuning
  • 3.3 Knowledge-Aware Validation
  • 4 Discussion, Challenges and Future
  • 4.1 Resources
  • 4.2 Evaluation Metrics
  • 4.3 Performance Analysis
  • 4.4 Trend Analysis
  • 4.5 Future Directions
  • 5 Conclusion
  • 6 Limitations
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Taxonomy of Knowledge Graph-Augmented Large Language Models for Hallucination Mitigation

    model/method

    Large language models (LLMs) frequently generate plausible yet factually inaccurate text (hallucinations) due to knowledge gaps, training data biases, and stochastic decoding. Structured external knowledge from Knowledge Graphs (KGs)—which store relational facts as entity-relation-entity triples (h,r,t)(h, r, t)—can be integrated into LLMs at various stages of the machine learning lifecycle to mitigate hallucinations. The taxonomy categorizes KG augmentation into three distinct paradigms:

    1. Knowledge-Aware Inference: Incorporates structured KG data at test time without modifying model parameters. This paradigm comprises:

      • KG-Augmented Retrieval: Querying and injecting relevant KG triples or subgraphs into the context prompt to provide ground-truth factual support.
      • KG-Augmented Reasoning: Leveraging graph-structured relational paths to guide step-by-step reasoning processes (e.g., Chain of Thought or Tree of Thoughts).
      • KG-Controlled Generation: Using LLMs to construct structured queries (e.g., SPARQL, API calls) or applying ontological guardrails to enforce generation boundaries.
    2. Knowledge-Aware Training: Modifies model representations and parameters via structured KG signals. This paradigm comprises:

      • Knowledge-Aware Pre-Training: Enriching language representations from scratch or continuing pre-training using knowledge-enhanced objectives, knowledge-guided entity masking, text-graph joint fusion, or knowledge probing.
      • Knowledge-Aware Fine-Tuning: Adapting pre-trained models to domain-specific tasks using linearized KG triples, entity embeddings, or synthetic instruction data derived from KGs.
    3. Knowledge-Aware Validation: Utilizing KGs as post-generation or decoding-time verification mechanisms (e.g., first-order logic claim verification, critic-guided generation, or sub-graph fact-checking) to detect, prevent, and explain factual inconsistencies.

  2. Knowl 2 — Knowledge-Aware Inference Mechanisms for Hallucination Reduction

    model/method

    Knowledge-aware inference injects structured symbolic facts directly into the prompt context at inference time, improving factual accuracy and multi-step reasoning without requiring model parameter updates:

    • KG-Augmented Retrieval: Matches question entities to knowledge graphs to retrieve factual triples (h,r,t)(h, r, t). Approaches include textualizing triples into natural language statements to improve LLM comprehension, training dedicated neural retriever modules over KG question answering (KGQA) datasets for complex multi-hop queries, and structured query generation over heterogeneous databases.

    • KG-Augmented Reasoning: Extends linear reasoning prompting (such as Chain-of-Thought) by establishing faithful reasoning paths over graph topologies. Techniques include interleaving step generation with graph retrieval (e.g., IRCoT), generating explicit relation paths over KGs to serve as interpretable reasoning plans (e.g., Reasoning on Graphs), and constructing graph-of-thought representations (e.g., MindMap) for domain-specific multi-step deduction.

    • Knowledge-Controlled Generation: Directs LLM generation within predefined logical and semantic constraints. This includes using LLMs (e.g., Codex) to generate API or SPARQL queries that fetch auxiliary facts, tuning cloze prompts for relation extraction, and applying programmatic enterprise guardrails (e.g., NeMo Guardrails) based on KG ontologies to enforce hard boundary conditions on model outputs.

  3. Knowl 3 — Knowledge-Aware Pre-Training and Fine-Tuning Strategies

    model/method

    Knowledge-aware training embeds structured knowledge graph facts directly into the parametric memory of language models through pre-training or parameter-efficient fine-tuning:

    • Pre-Training Approaches:

      • Knowledge-Enhanced Representation: Enriches text corpora with explicit KG entity embeddings or token dictionary mappings (e.g., ERNIE, KALM) via dual auto-regressive and auto-encoding objectives.
      • Knowledge-Guided Masking: Uses KG entity links to selectively mask interconnected entity spans and relation words in the corpus (e.g., SKEP, GLM), compelling the model to learn structural and relational dependencies.
      • Knowledge-Fusion: Connects language model encoders with Graph Neural Network (GNN) encoders (e.g., JointLK, LKPNR) to jointly update text and graph representations via cross-attention and message-passing.
      • Knowledge-Probing: Probes and re-aligns pre-trained weights through contrastive objectives over domain KGs (e.g., Rewire-then-Probe over biomedical taxonomies).
    • Fine-Tuning Approaches:

      • Synthetic Corpus Fine-Tuning: Converts KG triples into verbalized synthetic corpora (e.g., KELM, SKILL) and fine-tunes checkpoints via sequence-to-sequence objectives.
      • Structural Adapter Tuning: Integrates entity-relation embedding layers or low-rank adapters aligned with KG triples for domain-specific tasks such as link prediction, clinical diagnosis, and financial customer service.
  4. Knowl 4 — Knowledge-Aware Post-Generation and Decoding Validation

    model/method

    Knowledge-aware validation employs knowledge graphs as objective, structured reference bases to verify the factual consistency of LLM generations and eliminate hallucinations before or during final output emission:

    • Sub-Graph Contextual Checking: Extracts relevant sub-graphs based on the input context and cross-checks whether newly generated entities and attributes are topologically connected and consistent with the underlying KG (e.g., SURGE for dialogue generation, KGLM for fact-aware language modeling).

    • Critic-Driven Decoding: Utilizes a trained text-critic classifier to evaluate the semantic and factual correspondence between input structured data and partially generated token sequences, guiding beam search or sampling towards fact-aligned sequences.

    • First-Order Logic (FOL) Misinformation Verification: Formulates claim verification by translating claims into first-order logic predicates validated against knowledge graphs (e.g., FOLK). In addition to binary truth determination, this produces explicit, human-interpretable reasoning chains and explanations of why a generated statement is false or unsupported.

  5. Knowl 5 — Comparative Summary of Representative KG-Augmented LLM Frameworks

    data/table

    Representative methods across knowledge-aware inference, training, and validation differ significantly in downstream task targeting, underlying KG sources, base LLM architectures, and training resource demands.

    Category Representative Method Downstream Task KG Dataset Training Cost / Paradigm
    KG-Augmented Retrieval KAPING Question-Answering Mintaka, WebQSP Training-Free
    KG-Augmented Retrieval Rigel Facts Question-Answering WebQuestions, Mintaka, LC-QuAD Training-Free
    KG-Augmented Retrieval Retrieve-Rewrite-Answer Question-Answering MetaQA, WebQSP, ZJQA Training-Free
    KG-Augmented Reasoning IRCoT Multi-step Reasoning QA HotpotQA, 2WikiMultihopQA, MusiQue Training-Free
    KG-Augmented Reasoning MindMap Medical Diagnosis GenMedGPT-5k, CMCQA, ExplainCPE Training-Free
    KG-Augmented Reasoning RoG Reasoning WebQSP, ComplexWebQuestions Training-Free
    Controlled Generation KnowPrompt Relation Extraction SemEval, DialogRE, TACRED Few-shot Prompt Tuning
    Controlled Generation BINDER Information Extraction WikiTableQuestions, TabFact Few-shot In-context Learning
    Controlled Generation BeamQA Question Generation MetaQA, WebQSP Fine-tuned 4 epochs
    Pre-Training SKEP Sentiment Analysis SST, Amazon, MPQA Pre-trained on 3.2M data
    Pre-Training JointLK Commonsense QA CommonSenseQA, OpenBookQA Joint LM/GNN on 20 GPU hours
    Pre-Training LKPNR News Recommendation MIND 200K click logs on GPU
    Fine-Tuning SKILL Closed-book QA Wikidata, KELM, MetaQA T5 fine-tuned 50k steps
    Fine-Tuning KGLM Link Prediction WN18RR, FB15k-237, UMLS Tuned 5 epochs
    Fine-Tuning Neurosymbolic Enterprise Banking Query Chase Enterprise KG T5-large tuned 10 epochs
    Validation Fact-aware LM Fact Generation Linked WikiText-2 Trained on 256-d embeddings
    Validation SURGE Dialogue Generation OpenDialKG Training-Free
    Validation FOLK Claim Verification HoVER, FEVEROUS, SciFact Training-Free

    The comparison shows that inference-based and validation-based techniques (e.g., KAPING, MindMap, RoG, FOLK) achieve hallucination mitigation in large foundational models (such as GPT-3.5/4 and Llama-2) without parameter retraining, whereas pre-training and fine-tuning methods require substantial GPU resources and task-specific datasets.

  6. Knowl 6 — Evaluation Metrics for Hallucination Mitigation in KG-LLMs

    definition

    The effectiveness of knowledge graph augmentation in mitigating hallucinations and enhancing reasoning in LLMs is quantitatively and qualitatively evaluated using the following core metrics:

    • Factual Accuracy: Direct comparison of answer accuracy generated with versus without KG-augmented knowledge.

    • Top-KK Retrieval Accuracy and Mean Reciprocal Rank (MRR): Measures the capability of the retriever to place answer-containing KG triples within the top-KK candidates, defined as: MRR=1∣Q∣∑i=1∣Q∣1ranki\text{MRR} = \frac{1}{|Q|} \sum_{i=1}^{|Q|} \frac{1}{\text{rank}_i} where ∣Q∣|Q| is the total number of queries and ranki\text{rank}_i is the rank position of the first relevant factual triple for the ii-th query. Triples are also categorized as "Helpful" or "Harmful" relative to the no-knowledge baseline.

    • Hits@1: Measures the percentage of instances where the top-ranked candidate prediction or retrieved triple is strictly correct.

    • Execution Accuracy (EA): Evaluates knowledge-controlled generation models by measuring the exact semantic parsing correctness, valid syntax of generated API/SPARQL queries, and the success rate of code/query execution on underlying databases.

    • Exact Match (EM): The percentage of model-generated answers that match the ground-truth target text strings exactly on benchmark test sets.

    • Multi-Dimensional Human Evaluation: Qualitative scoring of generated outputs along dimensions of explanation quality, factual coverage, logical soundness, grammatical fluency, and presence of fabricated or ungrounded assertions.

  7. Knowl 7 — Empirical Performance Findings Across KG-Augmented LLM Paradigms

    empirical result

    Empirical analyses across various KG augmentation frameworks reveal distinct performance characteristics:

    1. Knowledge Retrieval Enhances Small LLMs: Augmenting compact language models (e.g., T5, T0, OPT) with factual KG triples boosts question-answering correctness by more than 80%80\% compared to unaugmented baselines, overcoming parameter constraints without requiring architectural scaling. However, retrieval accuracy remains bounded by the completeness of the KG and the semantic precision of the retriever module.

    2. Graph-Guided Reasoning Scales with Model Size: Step-wise reasoning grounded in KGs provides substantial gains for large models. For example, Reasoning on Graphs (RoG) improves ChatGPT reasoning accuracy on multi-hop question answering from 66.8%66.8\% to 85.7%85.7\%. Similarly, the MindMap framework achieves 88.2%88.2\% diagnostic accuracy in medical disease and drug recommendation tasks using clinical reasoning graphs.

    3. Knowledge-Controlled Generation Trade-offs: Generating structured queries (e.g., SPARQL or API calls via Codex) improves contextual relevance and parsing accuracy, but performance varies widely if the generated queries suffer from syntax errors or malformed entities.

    4. Computational Costs and Generalization in Training-Based Approaches: Joint pre-training and fine-tuning significantly enhance domain-specific factual retention, but they require heavy GPU compute (e.g., 20+ GPU hours for JointLK; 50k steps for SKILL) and create data dependencies that limit out-of-domain transferability.

  8. Knowl 8 — Historical Research Trends in KG-Augmented LLMs (2019–2023)

    empirical result

    Analysis of literature between 2019 and 2023 highlights a clear structural shift in how knowledge graphs are integrated with language models:

    • 2019–2021 (Pre-Training Dominance): Early research concentrated heavily on knowledge-aware pre-training and fine-tuning (e.g., ERNIE, K-BERT, KEPLER). Language models were small enough to allow full pre-training or extensive fine-tuning alongside graph encoders on combined text-graph objectives.

    • 2022–2023 (Inference and Validation Shift): With the rise of modern LLMs exceeding tens or hundreds of billions of parameters (e.g., GPT-3, GPT-4, LLaMA), full re-training and large-scale fine-tuning became computationally prohibitive. Consequently, research shifted predominantly toward training-free paradigms: KG-augmented retrieval, graph-based reasoning prompting (e.g., IRCoT, MindMap, RoG), controlled generation via in-context learning, and post-generation knowledge-aware validation.

  9. Knowl 9 — Key Future Research Directions for Knowledge Graph-Enhanced LLMs

    model/method

    To advance the integration of KGs and LLMs for hallucination mitigation, the survey identifies several future research frontiers:

    1. Dynamic Context-Aware and Fair KGs: Developing dynamically updating KGs that continuously incorporate real-time world events while implementing fairness-aware algorithms to prevent propagating social biases or misinformation into LLM prompts.

    2. Cross-Domain and Multimodal KGs: Unifying multi-domain knowledge (science, law, medicine, arts) and multimodal entities (incorporating images, audio, video) into unified knowledge graphs to enrich context for multimodal foundation models.

    3. Mixture of Experts (MoE) Integration: Coupling KG structures with MoE architectures to dynamically route tokens or sub-tasks to context-specialized experts, enhancing interpretability and efficiency without scaling compute.

    4. Symbolic-Subsymbolic Unification: Developing unified knowledge fabrics that tightly integrate symbolic KG structures with continuous subsymbolic vector spaces, enabling dual-process reasoning.

    5. Bidirectional LLM-KG Synergy: Establishing closed-loop frameworks where KGs ground LLM outputs in facts, while LLMs concurrently assist in automated KG construction, entity linking, and link prediction.

    6. Causality-Aware Knowledge Graphs: Encoding causal relationships rather than pure statistical correlations into graph edges, allowing LLMs to perform causal inference and counterfactual reasoning.

Coverage note — None. All primary contributions of the survey—including the structural taxonomy, inference, training, and validation methods, comparative attributes (Table 1), evaluation metrics, empirical performance analyses, historical trend evolution, and future directions—have been extracted as standalone knowls.

References

  1. 1.Oshin Agarwal, Heming Ge, Siamak Shakeri, and Rami Al-Rfou. 2020. Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training. arXiv preprint arXiv:2010.12688.
  2. 2.Garima Agrawal, Dimitri Bertsekas, and Huan Liu. 2023a. Auction-based learning for question answering over knowledge graphs. Information, 14(6):336.
  3. 3.Garima Agrawal, Yuli Deng, Jongchan Park, Huan Liu, and Ying-Chih Chen. 2022. Building knowledge graphs from unstructured texts: Applications and impact analyses in cybersecurity education. Information, 13(11):526.
  4. 4.Garima Agrawal, Kuntal Pal, Yuli Deng, Huan Liu, and Chitta Baral. 2023b. Aiseckg: Knowledge graph dataset for cybersecurity education. AAAI-MAKE 2023: Challenges Requiring the Combination of Machine Learning 2023.
  5. 5.Badr AlKhamissi, Millicent Li, Asli Celikyilmaz, Mona Diab, and Marjan Ghazvininejad. 2022. A review on language models as knowledge bases. arXiv preprint arXiv:2204.06031.
  6. 6.Farah Atif, Ola El Khatib, and Djellel Difallah. 2023. Beamqa: Multi-hop knowledge graph question answering with sequence-to-sequence prediction and beam search. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 781–790.
  7. 7.Jinheon Baek, Alham Fikri Aji, and Amir Saffari. 2023. Knowledge-augmented language model prompting for zero-shot knowledge graph question answering. arXiv preprint arXiv:2306.04136.
  8. 8.Teodoro Baldazzi, Luigi Bellomarini, Stefano Ceri, Andrea Colombo, Andrea Gentili, and Emanuel Sallinger. 2023. Fine-tuning large enterprise language models via ontological reasoning. arXiv preprint arXiv:2306.10723.
  9. 9.Parishad BehnamGhader, Santiago Miret, and Siva Reddy. 2022. Can retriever-augmented language models reason? the blame game between the retriever and the language model. arXiv preprint arXiv:2212.09146.
  10. 10.Yoshua Bengio, Rejean Ducharme, and Pascal Vincent. 2000. A neural probabilistic language model. Advances in neural information processing systems, 13.
  11. 11.Ryan Brate, Minh-Hoang Dang, Fabian Hoppe, Yuan He, Albert Merono-Penuela, and Vijay Sadashivaiah. 2022. Improving language model predictions via prompts enriched with knowledge graphs. In DL4KG@ ISWC2022.
  12. 12.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  13. 13.Harrison Chase. 2022. LangChain.
  14. 14.Xiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2022. Knowprompt: Knowledge-aware prompt-tuning with synergistic optimization for relation extraction. In Proceedings of the ACM Web conference 2022, pages 2778–2788.
  15. 15.Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, et al. 2022. Binding language models in symbolic languages. arXiv preprint arXiv:2210.02875.
  16. 16.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  17. 17.Jianfeng Deng, Chong Chen, Xinyi Huang, Wenyan Chen, and Lianglun Cheng. 2023a. Research on the construction of event logic knowledge graph of supply chain management. Advanced Engineering Informatics, 56:101921.
  18. 18.Shumin Deng, Chengming Wang, Zhoubo Li, Ningyu Zhang, Zelin Dai, Hehong Chen, Feiyu Xiong, Ming Yan, Qiang Chen, Mosha Chen, et al. 2023b. Construction and applications of billion-scale pre-trained multimodal business knowledge graph. In 2023 IEEE 39th International Conference on Data Engineering (ICDE), pages 2988–3002. IEEE.
  19. 19.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314.
  20. 20.Yue Dong, John Wieting, and Pat Verga. 2022. Faithful to the document or to the world? mitigating hallucinations via entity-linked knowledge in abstractive summarization. arXiv preprint arXiv:2204.13761.
  21. 21.Dieter Fensel, Umutcan Șimșek, Kevin Angele, Elwin Huaman, Elias Karle, Oleksandra Panasiuk, Ioan Toma, Jurgen Umbrich, Alexander Wahler, Dieter Fensel, et al. 2020. Why we need knowledge graphs: Applications. Knowledge Graphs: Methodology, Tools and Selected Use Cases, pages 95–112.
  22. 22.Negar Foroutan, Mohammadreza Banaei, Karl Aberer, and Antoine Bosselut. 2023. Breaking the language barrier: Improving cross-lingual reasoning with structured self-attention. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9422–9442.
  23. 23.Peng Fu, Yiming Zhang, Haobo Wang, Weikang Qiu, and Junbo Zhao. 2023. Revisiting the knowledge injection frameworks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 10983–10997.
  24. 24.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Pal: Program-aided language models. In International Conference on Machine Learning, pages 10764–10799. PMLR.
  25. 25.Almog Gueta, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen. 2023. Knowledge is a region in weight space for fine-tuned language models. arXiv preprint arXiv:2302.04863.
  26. 26.Qingyu Guo, Fuzhen Zhuang, Chuan Qin, Hengshu Zhu, Xing Xie, Hui Xiong, and Qing He. 2020. A survey on knowledge graph-based recommender systems. IEEE Transactions on Knowledge and Data Engineering, 34(8):3549–3568.
  27. 27.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929–3938. PMLR.
  28. 28.Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. 2023. Reasoning with language model is planning with world model. arXiv preprint arXiv:2305.14992.
  29. 29.Bin He, Di Zhou, Jinghui Xiao, Qun Liu, Nicholas Jing Yuan, Tong Xu, et al. 2019. Integrating graph contextualized knowledge into pre-trained language models. arXiv preprint arXiv:1912.00147.
  30. 30.Hangfeng He, Hongming Zhang, and Dan Roth. 2022. Rethinking with retrieval: Faithful large language model inference. arXiv preprint arXiv:2301.00303.
  31. 31.Aidan Hogan, Eva Blomqvist, Michael Cochez, Claudia d’Amato, Gerard De Melo, Claudio Gutierrez, Sabrina Kirrane, Jose Emilio Labra Gayo, Roberto Navigli, Sebastian Neumaier, et al. 2021. Knowledge graphs. ACM Computing Surveys (CSUR), 54(4):1–37.
  32. 32.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  33. 33.Linmei Hu, Zeyi Liu, Ziwang Zhao, Lei Hou, Liqiang Nie, and Juanzi Li. 2023. A survey of knowledge enhanced pre-trained language models. IEEE Transactions on Knowledge and Data Engineering.
  34. 34.Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2022. Large language models can self-improve. arXiv preprint arXiv:2210.11610.
  35. 35.Jie Huang and Kevin Chen-Chuan Chang. 2022. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403.
  36. 36.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  37. 37.Jinhao Jiang, Kun Zhou, Zican Dong, Keming Ye, Wayne Xin Zhao, and Ji-Rong Wen. 2023. Structgpt: A general framework for large language model to reason over structured data. arXiv preprint arXiv:2305.09645.
  38. 38.Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438.
  39. 39.Zhixue Jiang, Chengying Chi, and Yunyun Zhan. 2021. Research on medical question answering system based on knowledge graph. IEEE Access, 9:21094–21101.
  40. 40.Minki Kang, Jinheon Baek, and Sung Ju Hwang. 2022a. Kala: knowledge-augmented language model adaptation. arXiv preprint arXiv:2204.10555.
  41. 41.Minki Kang, Jin Myung Kwak, Jinheon Baek, and Sung Ju Hwang. 2022b. Knowledge-consistent dialogue generation with knowledge graphs. In ICML 2022 Workshop on Knowledge Retrieval and Language Models.
  42. 42.Nora Kassner, Philipp Dufter, and Hinrich Schutze. 2021. Multilingual lama: Investigating knowledge in multilingual pretrained language models. arXiv preprint arXiv:2102.00894.
  43. 43.Pei Ke, Haozhe Ji, Yu Ran, Xin Cui, Liwei Wang, Linfeng Song, Xiaoyan Zhu, and Minlie Huang. 2021. Jointgt: Graph-text joint representation learning for text generation from knowledge graphs. arXiv preprint arXiv:2106.10502.
  44. 44.Seungone Kim, Se June Joo, Doyoung Kim, Joel Jang, Seonghyeon Ye, Jamin Shin, and Minjoon Seo. 2023. The cot collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning. arXiv preprint arXiv:2305.14045.
  45. 45.Mateusz Lango and OndȘrej DuȘsek. 2023. Critic-driven decoding for mitigating hallucinations in data-to-text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2853–2862.
  46. 46.Doug Lenat and Gary Marcus. 2023. Getting from generative ai to trustworthy ai: What llms might learn from cyc. arXiv preprint arXiv:2308.04445.
  47. 47.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen-tau Yih, Tim Rocktaschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  48. 48.Tianle Li, Xueguang Ma, Alex Zhuang, Yu Gu, Yu Su, and Wenhu Chen. 2023. Few-shot in-context learning for knowledge base question answering. arXiv preprint arXiv:2305.01750.
  49. 49.Xiaonan Li and Xipeng Qiu. 2023. Mot: Memory-of-thought enables chatgpt to self-improve. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6354–6374.
  50. 50.Ke Liang, Lingyuan Meng, Meng Liu, Yue Liu, Wenxuan Tu, Siwei Wang, Sihang Zhou, Xinwang Liu, and Fuchun Sun. 2022. Reasoning over different types of knowledge graphs: Static, temporal and multi-modal. arXiv preprint arXiv:2212.05767.
  51. 51.Jerry Liu. 2022. LlamaIndex.
  52. 52.Jiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras, Yejin Choi, and Hannaneh Hajishirzi. 2021. Generated knowledge prompting for commonsense reasoning. arXiv preprint arXiv:2110.08387.
  53. 53.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1–35.
  54. 54.Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Qi Ju, Haotang Deng, and Ping Wang. 2020. K-bert: Enabling language representation with knowledge graph. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 2901–2908.
  55. 55.Robert L Logan IV, Nelson F Liu, Matthew E Peters, Matt Gardner, and Sameer Singh. 2019. Barack’s wife hillary: Using knowledge-graphs for fact-aware language modeling. arXiv preprint arXiv:1906.07241.
  56. 56.Pan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, and Ashwin Kalyan. 2022. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. arXiv preprint arXiv:2209.14610.
  57. 57.Linhao Luo, Yuan-Fang Li, Gholamreza Haffari, and Shirui Pan. 2023. Reasoning on graphs: Faithful and interpretable large language model reasoning. arXiv preprint arXiv:2310.01061.
  58. 58.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9802–9822.
  59. 59.Xuting Mao, Hao Sun, Xiaoqian Zhu, and Jianping Li. 2022. Financial fraud detection using the related-party transaction knowledge graph. Procedia Computer Science, 199:733–740.
  60. 60.Ariana Martino, Michael Iannelli, and Coleen Truong. 2023. Knowledge injection to counter large language model (llm) hallucination. In European Semantic Web Conference, pages 182–185. Springer.
  61. 61.Zaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su, Charlotte Collins, and Nigel Collier. 2021. Rewire-then-probe: A contrastive recipe for probing biomedical knowledge of pre-trained language models. arXiv preprint arXiv:2110.08173.
  62. 62.Gregoire Mialon, Roberto Dessı, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Roziere, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, et al. 2023. Augmented language models: a survey. arXiv preprint arXiv:2302.07842.
  63. 63.Fedor Moiseev, Zhe Dong, Enrique Alfonseca, and Martin Jaggi. 2022. Skill: structured knowledge infusion for large language models. arXiv preprint arXiv:2205.08184.
  64. 64.Vishwas Mruthyunjaya, Pouya Pezeshkpour, Estevam Hruschka, and Nikita Bhutani. 2023. Rethinking language models as symbolic knowledge graphs. arXiv preprint arXiv:2308.13676.
  65. 65.Carlos Nunez-Molina, Pablo Mesejo, and Juan Fernandez-Olivares. 2023. A review of symbolic, subsymbolic and hybrid methods for sequential decision making. arXiv preprint arXiv:2304.10590.
  66. 66.Reham Omar, Ishika Dhall, Panos Kalnis, and Essam Mansour. 2023. A universal question-answering platform for knowledge graphs. Proceedings of the ACM on Management of Data, 1(1):1–25.
  67. 67.Yasumasa Onoe, Michael JQ Zhang, Shankar Padmanabhan, Greg Durrett, and Eunsol Choi. 2023. Can lms learn new entities from descriptions? challenges in propagating injected knowledge. arXiv preprint arXiv:2305.01651.
  68. 68.OpenAI. 2023. Gpt-4 technical report.
  69. 69.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  70. 70.Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2023. Unifying large language models and knowledge graphs: A roadmap. arXiv preprint arXiv:2306.08302.
  71. 71.Matthew E Peters, Mark Neumann, Robert L Logan IV, Roy Schwartz, Vidur Joshi, Sameer Singh, and Noah A Smith. 2019. Knowledge enhanced contextual word representations. arXiv preprint arXiv:1909.04164.
  72. 72.Fabio Petroni, Tim Rocktaschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019. Language models as knowledge bases? arXiv preprint arXiv:1909.01066.
  73. 73.Nina Poerner, Ulli Waltinger, and Hinrich Schutze. 2019. E-bert: Efficient-yet-effective entity embeddings for bert. arXiv preprint arXiv:1911.03681.
  74. 74.Archiki Prasad, Swarnadeep Saha, Xiang Zhou, and Mohit Bansal. 2023. Receval: Evaluating reasoning chains via correctness and informativeness. arXiv preprint arXiv:2304.10703.
  75. 75.Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2022. Reasoning with language model prompting: A survey. arXiv preprint arXiv:2212.09597.
  76. 76.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  77. 77.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented language models. arXiv preprint arXiv:2302.00083.
  78. 78.Traian Rebedea, Razvan Dinu, Makesh Sreedhar, Christopher Parisien, and Jonathan Cohen. 2023. Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails. arXiv preprint arXiv:2310.10501.
  79. 79.Corby Rosset, Chenyan Xiong, Minh Phan, Xia Song, Paul Bennett, and Saurabh Tiwary. 2020. Knowledge-aware language model pretraining. arXiv preprint arXiv:2007.00655.
  80. 80.Xie Runfeng, Cui Xiangyang, Yan Zhou, Wang Xin, Xuan Zhanwei, Zhang Kai, et al. 2023. Lkpnr: Llm and kg for personalized news recommendation framework. arXiv preprint arXiv:2308.12028.
  81. 81.Knowledge Graphs Seminar, Nahor Gebretensae, and Heiko Paulheim. 2019. Wikidata: A free collaborative knowledge graph.
  82. 82.Priyanka Sen, Sandeep Mavadia, and Amir Saffari. 2023. Knowledge graph-augmented language models for complex question answering.
  83. 83.Tao Shen, Yi Mao, Pengcheng He, Guodong Long, Adam Trischler, and Weizhu Chen. 2020. Exploiting structured knowledge in text via graph-guided representation learning. arXiv preprint arXiv:2004.14224.
  84. 84.Xiao Shi, Zhengyuan Zhu, Zeyu Zhang, and Chengkai Li. 2023. Hallucination mitigation in natural language generation from large-scale open-domain knowledge graphs. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12506–12521.
  85. 85.Noah Shinn, Beck Labash, and Ashwin Gopinath. 2023. Reflexion: an autonomous agent with dynamic memory and self-reflection. arXiv preprint arXiv:2303.11366.
  86. 86.Chandan Singh, John Morris, Alexander M Rush, Jianfeng Gao, and Yuntian Deng. 2023. Tree prompting: Efficient task adaptation without fine-tuning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6253–6267.
  87. 87.Amit Singhal. 2012. Introducing the knowledge graph: things, not strings, may 2012. URL http://googleblog.blogspot. ie/2012/05/introducing-knowledgegraph-things-not. html.
  88. 88.Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, et al. 2021a. Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2107.02137.
  89. 89.Yueqing Sun, Qi Shi, Le Qi, and Yu Zhang. 2021b. Jointlk: Joint reasoning with language models and knowledge graphs for commonsense question answering. arXiv preprint arXiv:2112.02732.
  90. 90.Vinitra Swamy, Angelika Romanou, and Martin Jaggi. 2021. Interpreting language models through knowledge graph extraction. arXiv preprint arXiv:2111.08546.
  91. 91.Hao Tian, Can Gao, Xinyan Xiao, Hao Liu, Bolei He, Hua Wu, Haifeng Wang, and Feng Wu. 2020. Skep: Sentiment knowledge enhanced pre-training for sentiment analysis. arXiv preprint arXiv:2005.05635.
  92. 92.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. arXiv preprint arXiv:2212.10509.
  93. 93.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  94. 94.Blerta Veseli, Simon Razniewski, Jan-Christoph Kalo, and Gerhard Weikum. 2023. Evaluating the knowledge base completion potential of gpt. Findings of EMNLP 2023.
  95. 95.Boxin Wang, Wei Ping, Peng Xu, Lawrence McAfee, Zihan Liu, Mohammad Shoeybi, Yi Dong, Oleksii Kuchaiev, Bo Li, Chaowei Xiao, et al. 2023a. Shall we pretrain autoregressive language models with retrieval? a comprehensive study. arXiv preprint arXiv:2304.06762.
  96. 96.Haoran Wang and Kai Shu. 2023. Explainable claim verification via knowledge-grounded reasoning with large language models. arXiv preprint arXiv:2310.05253.
  97. 97.Hongru Wang, Minda Hu, Yang Deng, Rui Wang, Fei Mi, Weichao Wang, Yasheng Wang, Wai-Chung Kwan, Irwin King, and Kam-Fai Wong. 2023b. Large language models as source planner for personalized knowledge-grounded dialogue. arXiv preprint arXiv:2310.08840.
  98. 98.Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2021. Kepler: A unified model for knowledge embedding and pre-trained language representation. Transactions of the Association for Computational Linguistics, 9:176–194.
  99. 99.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  100. 100.Zhaoyang Wang, Shaohan Huang, Yuxuan Liu, Jiahai Wang, Minghui Song, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, et al. 2023c. Democratizing reasoning ability: Tailored learning from large language model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 1948–1966.
  101. 101.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022a. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  102. 102.Xiaokai Wei, Shen Wang, Dejiao Zhang, Parminder Bhatia, and Andrew Arnold. 2021. Knowledge enhanced pretrained language models: A compreshensive survey. arXiv preprint arXiv:2110.08455.
  103. 103.Yanbin Wei, Qiushi Huang, Yu Zhang, and James Kwok. 2023. Kicgpt: Large language model with knowledge in context for knowledge graph completion. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 8667–8683.
  104. 104.Yinwei Wei, Xiang Wang, Liqiang Nie, Shaoyu Li, Dingxian Wang, and Tat-Seng Chua. 2022b. Causal inference for knowledge graph based recommendation. IEEE Transactions on Knowledge and Data Engineering.
  105. 105.Yilin Wen, Zifeng Wang, and Jimeng Sun. 2023. Mindmap: Knowledge graph prompting sparks graph of thoughts in large language models. arXiv preprint arXiv:2308.09729.
  106. 106.Yike Wu, Nan Hu, Guilin Qi, Sheng Bi, Jie Ren, Anhuan Xie, and Wei Song. 2023. Retrieve-rewrite-answer: A kg-to-text enhanced llms framework for knowledge graph question answering. arXiv preprint arXiv:2309.11206.
  107. 107.Zilin Xiao, Ming Gong, Jie Wu, Xingyao Zhang, Linjun Shou, and Daxin Jiang. 2023. Instructed language models with retrievers are powerful entity linkers. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2267–2282.
  108. 108.Haoyan Yang, Zhitao Li, Yong Zhang, Jianzong Wang, Ning Cheng, Ming Li, and Jing Xiao. 2023. Prca: Fitting black-box large language models for retrieval question answering via pluggable reward-driven contextual adapter. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5364–5375.
  109. 109.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601.
  110. 110.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
  111. 111.Hongbin Ye, Ningyu Zhang, Hui Chen, and Huajun Chen. 2022. Generative knowledge graph construction: A review. arXiv preprint arXiv:2210.12714.
  112. 112.Da Yin, Li Dong, Hao Cheng, Xiaodong Liu, Kai-Wei Chang, Furu Wei, and Jianfeng Gao. 2022. A survey of knowledge-intensive nlp with pre-trained language models. arXiv preprint arXiv:2202.08772.
  113. 113.Xunjian Yin, Baizhou Huang, and Xiaojun Wan. 2023a. Alcuna: Large language models meet new knowledge. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 1397–1414.
  114. 114.Zhangyue Yin, Qiushi Sun, Cheng Chang, Qipeng Guo, Junqi Dai, Xuan-Jing Huang, and Xipeng Qiu. 2023b. Exchange-of-thought: Enhancing large language model capabilities through cross-model communication. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 15135–15153.
  115. 115.Jason Youn and Ilias Tagkopoulos. 2022. Kglm: Integrating knowledge graph structure in language models for link prediction. arXiv preprint arXiv:2211.02744.
  116. 116.Mengxia Yu, Zhihan Zhang, Wenhao Yu, and Meng Jiang. 2023. Pre-training language models for comparative reasoning. arXiv preprint arXiv:2305.14457.
  117. 117.Wenhao Yu, Chenguang Zhu, Lianhui Qin, Zhihan Zhang, Tong Zhao, and Meng Jiang. 2022. Diversifying content generation for commonsense reasoning with mixture of knowledge graph experts. arXiv preprint arXiv:2203.07285.
  118. 118.Denghui Zhang, Zixuan Yuan, Yanchi Liu, Fuzhen Zhuang, and Hui Xiong. E-bert: Adapting bert to e-commerce with adaptive hybrid masking and neighbor product reconstruction.
  119. 119.Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A Smith. 2023a. How language model hallucinations can snowball. arXiv preprint arXiv:2305.13534.
  120. 120.Zhebin Zhang, Xinyu Zhang, Yuanhang Ren, Saijiang Shi, Meng Han, Yongkang Wu, Ruofei Lai, and Zhao Cao. 2023b. Iag: Induction-augmented generation framework for answering reasoning questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 1–14.
  121. 121.Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019. Ernie: Enhanced language representation with informative entities. arXiv preprint arXiv:1905.07129.
  122. 122.Zihan Zhang, Meng Fang, Ling Chen, Mohammad Reza Namazi Rad, and Jun Wang. 2023c. How do large language models capture the ever-changing world knowledge? a review of recent advances. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8289–8311.
  123. 123.Shen Zheng, Jie Huang, and Kevin Chen-Chuan Chang. 2023. Why does chatgpt fall short in answering questions faithfully? arXiv preprint arXiv:2304.10513.
  124. 124.Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew M Dai, Quoc V Le, James Laudon, et al. 2022. Mixture-of-experts with expert choice routing. Advances in Neural Information Processing Systems, 35:7103–7114.

Citation

MLA
Agrawal, G., et al. “Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 3947–60, https://doi.org/10.18653/v1/2024.naacl-long.219.
APA
Agrawal, G., Kumarage, T., Alghamdi, Z., & Liu, H. (2024). Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3947–3960. https://doi.org/10.18653/v1/2024.naacl-long.219
Chicago
Agrawal, G., T. Kumarage, Z. Alghamdi, and H. Liu. 2024. “Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3947–60. https://doi.org/10.18653/v1/2024.naacl-long.219.
Harvard
Agrawal, G. et al. (2024) “Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3947–3960. Available at: https://doi.org/10.18653/v1/2024.naacl-long.219.
Vancouver
1. Agrawal G, Kumarage T, Alghamdi Z, Liu H (2024) Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 3947–3960

BibTeX

@inproceedings{agrawal-etal-2024-knowledge,
    title = "Can Knowledge Graphs Reduce Hallucinations in {LLM}s? : A Survey",
    author = "Agrawal, Garima  and
      Kumarage, Tharindu  and
      Alghamdi, Zeyad  and
      Liu, Huan",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.219/",
    doi = "10.18653/v1/2024.naacl-long.219",
    pages = "3947--3960"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/