Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction

Bowen ZhangHarold Soh

article2024EMNLP102 citations

Proposes a three-phase knowledge graph construction framework that overcomes large language model context window limits by extracting open triplets, generating definitions, and canonicalizing relations with or without a pre-defined schema.

Listen

Organizations increasingly rely on knowledge graphs—structured networks of entities and their relationships—to drive search, recommendation systems, and decision-making. However, constructing knowledge graphs from unstructured text has traditionally required intensive manual labor or fine-tuned, domain-specific models. While large language models offer strong text understanding, existing extraction techniques struggle to scale. Standard methods require inserting an entire predefined schema directly into the prompt, which quickly exceeds context window limits for large schemas and fails when no predetermined schema exists.

The article evaluates a modular, three-phase framework called Extract-Define-Canonicalize (EDC) to automate knowledge graph construction across complex text. The objective is to demonstrate that decomposing the extraction process into free extraction, definition generation, and post-processing canonicalization enables high-quality graph construction for both large predefined schemas and self-generated schemas without fine-tuning the base language model.

The evaluated approach breaks the task into open information extraction, generating natural language definitions for extracted relations, and standardizing those relations through vector search and model verification. An enhanced variant, EDC+R, incorporates an iterative refinement step driven by a trained Schema Retriever that supplies contextually relevant schema hints. The authors evaluated EDC across three established benchmarks with schemas containing up to 200 distinct relation types—WebNLG, REBEL, and Wiki-NRE—as well as a synthetic dataset of fictional entities. Performance was assessed against specialized supervised baselines and leading clustering techniques using standard precision, recall, F1 metrics, and blinded human evaluation.

The analysis yielded several key findings. First, EDC matched or exceeded state-of-the-art specialized models across all datasets, achieving partial F1 scores of up to 0.820 on WebNLG when paired with refinement and advanced language models. Second, the framework substantially outperformed constrained generative baselines on complex datasets like REBEL (0.601 vs. 0.385 F1) and Wiki-NRE (0.713 vs. 0.484 F1), primarily because it successfully extracts numerical and date literals. Third, when constructing graphs without any predefined schema, human evaluators confirmed that EDC produced highly accurate graphs (87–96% precision) with significantly more concise schemas and lower redundancy than existing clustering baselines, avoiding improper grouping of distinct relations. Finally, ablation studies showed that the Schema Retriever provided critical contextual disambiguation, boosting F1 performance across all tested benchmarks.

These findings indicate that organizations can build high-accuracy knowledge graphs without the significant cost of training specialized extraction models or manually maintaining rigid schemas. By handling schemas modularly after extraction, EDC eliminates the constraint limitations of standard prompting and allows systems to dynamically discover and integrate new relations into existing knowledge bases.

Stakeholders adopting this architecture should implement the retrieval-augmented refinement loop to capture subtle, fine-grained relationships. When deploying in production, engineering teams should evaluate hybrid architectures—such as combining the initial extraction and definition steps or replacing intermediate modules with smaller, specialized classifiers—to manage latency and language model operational expenses (observed at approximately $0.009 per example using commercial models).

While confidence in the core extraction methodology is high, readers should note certain operational boundaries. The evaluation relied on sentence- and paragraph-level inputs, meaning whole-document deployments will require upstream chunking and coreference resolution modules. Additionally, the study focused primarily on relation canonicalization rather than full entity deduplication. Further validation through pilot implementations in production environments will help establish exact cost-accuracy trade-offs across different open-source and proprietary language models.

arXiv: 2404.03868
Cover for Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction

Abstract

In this work, we are interested in automated methods for knowledge graph creation (KGC) from input text. Progress on large language models (LLMs) has prompted a series of recent works applying them to KGC, e.g., via zero/few-shot prompting. Despite successes on small domain-specific datasets, these models face difficulties scaling up to text common in many real-world applications. A principal issue is that, in prior methods, the KG schema has to be included in the LLM prompt to generate valid triplets; larger and more complex schemas easily exceed the LLMs’ context window length. Furthermore, there are scenarios where a fixed pre-defined schema is not available and we would like the method to construct a high-quality KG with a succinct self-generated schema. To address these problems, we propose a three-phase framework named Extract-Define-Canonicalize (EDC): open information extraction followed by schema definition and post-hoc canonicalization. EDC is flexible in that it can be applied to settings where a pre-defined target schema is available and when it is not; in the latter case, it constructs a schema automatically and applies self-canonicalization. To further improve performance, we introduce a trained component that retrieves schema elements relevant to the input text; this improves the LLMs’ extraction performance in a retrieval-augmented generation-like manner. We demonstrate on three KGC benchmarks that EDC is able to extract high-quality triplets without any parameter tuning and with significantly larger schemas compared to prior works. Code for EDC is available at https://github.com/clear-nus/edc.

Table of Contents

  • 1 Introduction
  • EDC: Extract-Define-Canonicalize
  • 2 Background
  • 3 Method: EDC for KGC
  • 3.1 EDC: Extract-Define-Canonicalize
  • 3.2 EDC+R: iteratively refine EDC with Schema Retriever
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.1.1 Evaluation Criteria and Baselines
  • 4.2 Results and Analysis
  • 4.2.1 Target Alignment
  • 4.2.2 Self Canonicalization
  • 5 Conclusion
  • 6 Limitations and Future Directions
  • 7 Ethical Considerations
  • Acknowledgements
  • References
  • A Implementation Details
  • A.1 Models and Infrastructures Details
  • A.2 Prompting-related hyperparameters
  • A.3 Schema Retriever Training
  • A.4 Details of Refinement Hint
  • A.4.1 Obtaining Candidate Entities
  • A.4.2 Obtaining Candidate Relations
  • A.4.3 Usage of Hint for Refined OIE
  • B Annotation Instruction
  • C Detailed Results of Target Alignment
  • C.1 Complete Results
  • C.2 Effect of More Refinement Iterations
  • C.3 Ablation Study on Last-Round Extractions
  • C.4 Discussion on KGC Dataset Annotations
  • D Experiments on a Novel Dataset
  • E Comparison against previous LLM-based approaches
  • F Combine EDC with other IE tools
  • G Combining OIE and Schema Definition

Knowls

  1. Knowl 1 — Extract-Define-Canonicalize (EDC) Framework

    model/method

    The Extract-Define-Canonicalize (EDC) framework decomposes knowledge graph construction (KGC) from unstructured text into three sequential phases to avoid context-window bottlenecks associated with large schemas:

    1. Open Information Extraction (OIE): An open relation extraction prompt is issued to a large language model (LLM) using few-shot in-context examples (typically 6-shot). The LLM extracts relational triplets ⟨Subject,Relation,Object⟩\langle \text{Subject}, \text{Relation}, \text{Object} \rangle freely from the text without conditioning on a predefined schema, producing an unconstrained Open Knowledge Graph.

    2. Schema Definition: For every relation extracted in Phase 1, the LLM is prompted with the source text and the extracted triplet to generate a natural language definition justifying the semantic relation in its context (e.g., specifying the role of the subject and object entities).

    3. Schema Canonicalization: The natural language definition generated in Phase 2 is mapped to a vector embedding using a sentence transformer model (such as E5-Mistral-7b). A vector similarity search retrieves the top-kk (e.g., k=5k=5) semantically nearest relations from candidate schema definitions. The LLM is then presented with a multiple-choice prompt containing the retrieved definitions and a "None of the above" option to select the canonicalized relation or verify whether a valid mapping exists.

  2. Knowl 2 — Iterative Refinement with Schema Retriever (EDC+R)

    model/method

    EDC+R is an iterative extension of the Extract-Define-Canonicalize framework that refines extracted triplets using retrieval-augmented hints to improve extraction recall and precision on fine-grained relations.

    In EDC+R, a second extraction pass is conducted where the OIE prompt is augmented with a structured hint consisting of:

    • Candidate Entities: The union of entities extracted by EDC in the prior iteration and entities extracted from the text by a dedicated LLM named entity extraction prompt.
    • Candidate Relations: The union of relations identified by EDC in the previous iteration and the top-10 relations retrieved from the target or canonical schema using a trained dense Schema Retriever.

    The LLM executes the extraction prompt conditioned on these entity and relation hints (along with their canonical definitions), allowing it to discover fine-grained, schema-specific relations that were omitted in unguided open extraction.

  3. Knowl 3 — Schema Retriever Training Objective and Formulation

    equation

    The Schema Retriever is a dense retrieval model adapted from the E5-mistral-7b-instruct embedding backbone to score the contextual relevance between an input text and candidate relations in a target schema.

    Given a positive text-relation pair (t+,r+)(t^+, r^+), the input text t+t^+ is formatted using an instruction template: tinst+="Instruct: retrieve relations that are present in the given text \nQuery: {t+}"t^+_{\text{inst}} = \text{"Instruct: retrieve relations that are present in the given text \n Query: } \{t^+\}\text{"}

    The model is trained to maximize the similarity of (tinst+,r+)(t^+_{\text{inst}}, r^+) against a set N\mathcal{N} of non-relevant relation samples using the InfoNCE contrastive loss: L=−log⁡ϕ(tinst+,r+)ϕ(tinst+,r+)+∑ni∈Nϕ(tinst+,ni)\mathcal{L} = -\log \frac{\phi(t^+_{\text{inst}}, r^+)}{\phi(t^+_{\text{inst}}, r^+) + \sum_{n_i \in \mathcal{N}} \phi(t^+_{\text{inst}}, n_i)}

    where ϕ(u,v)=u⊤v∥u∥2∥v∥2\phi(u, v) = \frac{u^\top v}{\|u\|_2 \|v\|_2} denotes the cosine similarity function between text representation embeddings. The training set is synthesized from text-triplet alignments (e.g., TEKGEN) comprising 37,500 positive and negative text-relation pairs.

  4. Knowl 4 — Target Alignment versus Self-Canonicalization Modes in EDC

    definition

    The Schema Canonicalization phase of EDC operates in two distinct operational regimes depending on schema availability:

    • Target Alignment: Applied when a predefined target schema Starget\mathcal{S}_{\text{target}} exists. The generated relation definitions are matched against definitions in Starget\mathcal{S}_{\text{target}} via vector search. The LLM verifies whether any retrieved candidate in Starget\mathcal{S}_{\text{target}} preserves the semantic meaning of the extracted relation. If the LLM determines no match is feasible ("None of the above"), the triplet is discarded to ensure strict conformity with Starget\mathcal{S}_{\text{target}}.

    • Self-Canonicalization: Applied in an open-world setting where no target schema is provided. The system maintains an emergent canonical schema Scanon\mathcal{S}_{\text{canon}}, initialized as empty. For each extracted triplet, the relation definition is compared against existing relations in Scanon\mathcal{S}_{\text{canon}}. If the LLM verifies a match, the relation is mapped to that existing canonical relation; if no candidate matches, the novel relation and its LLM-generated definition are added into Scanon\mathcal{S}_{\text{canon}}, expanding the canonical schema dynamically while avoiding synonym redundancy.

  5. Knowl 5 — Performance of EDC and EDC+R on Target Alignment Benchmarks

    data/table

    Evaluation of EDC and EDC+R in the Target Alignment setting was conducted across WebNLG (159 relation types), REBEL (sub-sampled 1,000 pairs, 200 relation types), and Wiki-NRE (sub-sampled 1,000 pairs, 45 relation types). Scores are token-based Precision (PP), Recall (RR), and F1 under Partial, Strict, and Exact named-entity match criteria.

    Partial Strict Exact
    Dataset Method (LLM) PP RR F1 PP RR F1 PP RR F1
    WebNLG EDC (GPT-4) 0.776 0.796 0.783 0.729 0.741 0.733 0.751 0.765 0.756
    EDC+R (GPT-4) 0.814 0.831 0.820 0.782 0.794 0.786 0.796 0.808 0.800
    EDC (GPT-3.5) 0.739 0.760 0.746 0.684 0.697 0.688 0.708 0.722 0.713
    EDC+R (GPT-3.5) 0.788 0.806 0.794 0.749 0.761 0.753 0.768 0.781 0.772
    EDC (Mistral-7b) 0.723 0.739 0.728 0.668 0.679 0.672 0.692 0.703 0.696
    EDC+R (Mistral-7b) 0.756 0.775 0.762 0.716 0.727 0.720 0.735 0.747 0.739
    REGEN (Baseline) 0.755 0.788 0.767 0.713 0.735 0.720 0.714 0.738 0.723
    REBEL EDC (GPT-4) 0.543 0.552 0.546 0.498 0.503 0.500 0.511 0.517 0.514
    EDC+R (GPT-4) 0.599 0.606 0.601 0.557 0.561 0.559 0.572 0.576 0.574
    EDC (GPT-3.5) 0.503 0.512 0.506 0.448 0.453 0.449 0.471 0.476 0.473
    EDC+R (GPT-3.5) 0.556 0.565 0.559 0.513 0.519 0.516 0.527 0.533 0.529
    GenIE (Baseline) 0.381 0.391 0.385 0.353 0.361 0.356 0.362 0.369 0.364
    Wiki-NRE EDC (GPT-4) 0.682 0.686 0.683 0.675 0.679 0.677 0.676 0.680 0.678
    EDC+R (GPT-4) 0.712 0.715 0.713 0.708 0.710 0.709 0.708 0.711 0.709
    EDC (GPT-3.5) 0.645 0.651 0.647 0.636 0.640 0.638 0.638 0.643 0.640
    EDC+R (GPT-3.5) 0.691 0.696 0.693 0.684 0.688 0.685 0.685 0.689 0.687
    GenIE (Baseline) 0.482 0.486 0.484 0.462 0.464 0.463 0.477 0.479 0.478

    EDC outperforms baseline fine-tuned sequence-to-sequence models (REGEN and GenIE) across all criteria, and EDC+R consistently yields additional performance improvements.

  6. Knowl 6 — Self-Canonicalization Performance and Schema Metrics

    data/table

    In the Self-Canonicalization setting (where no target schema is provided), EDC was evaluated using GPT-3.5-turbo against the initial uncanonicalized Open KG and the clustering-based canonicalization method CESI. The evaluation measures human-evaluated precision (inter-annotator agreement κ=0.94\kappa = 0.94), schema size (number of distinct relation types, No. Rel.), and a schema redundancy score defined as the average cosine similarity between each canonicalized relation definition and its nearest counterpart.

    Dataset Method Precision (↑\uparrow) No. Rel. (↓\downarrow) Redundancy (↓\downarrow)
    WebNLG EDC 0.956 200 0.833
    CESI 0.724 280 0.893
    Open KG 0.982 529 0.927
    REBEL EDC 0.867 225 0.831
    CESI 0.504 307 0.854
    Open KG 0.903 667 0.895
    Wiki-NRE EDC 0.898 106 0.833
    CESI 0.753 114 0.849
    Open KG 0.909 204 0.881

    EDC achieves a substantial reduction in schema size and redundancy relative to the raw Open KG while retaining high precision (0.8670.867--0.9560.956). In contrast, CESI experiences severe precision degradation due to over-generalization and inappropriate clustering of semantically distinct relations.

  7. Knowl 7 — Ablation of Schema Retriever in EDC+R

    data/table

    An ablation study assessing the contribution of the Schema Retriever (S.R.) in the refinement stage was carried out using GPT-3.5-turbo across WebNLG, REBEL, and Wiki-NRE. In the ablated variant (EDC+R w/o S.R.), candidate relations retrieved from the schema were omitted from the refinement hint, retaining only candidate entities and relations extracted in the previous round.

    Dataset Method Partial F1 Strict F1 Exact F1
    WebNLG EDC+R 0.794 0.753 0.772
    EDC+R w/o S.R. 0.752 0.701 0.721
    EDC 0.746 0.688 0.713
    REBEL EDC+R 0.559 0.516 0.529
    EDC+R w/o S.R. 0.517 0.466 0.482
    EDC 0.506 0.449 0.473
    Wiki-NRE EDC+R 0.693 0.685 0.657
    EDC+R w/o S.R. 0.653 0.645 0.641
    EDC 0.647 0.638 0.640

    Removing the Schema Retriever from the refinement hint leads to a marked drop in F1 across all datasets, demonstrating that dense retrieval of schema relations enables the LLM to identify specific target-schema relations that are typically missed during unguided open extraction.

  8. Knowl 8 — Comparison of EDC with In-Prompt LLM-based KGC Baselines

    data/table

    EDC and EDC+R were compared against direct in-prompt LLM extraction methods (CodeKGC and ChatIE) on datasets where the full schema can fit within standard context windows: CoNLL04 (4 relations), SciERC (7 relations), and Wiki-NRE (45 relations), using GPT-3.5-turbo across all methods.

    Partial Strict Exact
    Dataset Method PP RR F1 PP RR F1 PP RR F1
    CoNLL04 EDC 0.536 0.552 0.543 0.481 0.491 0.485 0.503 0.515 0.509
    EDC+R 0.580 0.593 0.585 0.514 0.522 0.517 0.549 0.558 0.552
    CodeKGC 0.542 0.550 0.545 0.503 0.506 0.504 0.542 0.546 0.543
    ChatIE 0.463 0.477 0.468 0.360 0.366 0.363 0.418 0.427 0.421
    SciERC EDC 0.389 0.408 0.395 0.288 0.301 0.292 0.352 0.365 0.357
    EDC+R 0.447 0.461 0.451 0.340 0.349 0.343 0.406 0.416 0.410
    CodeKGC 0.389 0.398 0.392 0.277 0.283 0.279 0.346 0.353 0.349
    ChatIE 0.351 0.367 0.357 0.212 0.221 0.215 0.294 0.302 0.297
    Wiki-NRE EDC 0.645 0.651 0.647 0.636 0.640 0.638 0.638 0.643 0.640
    EDC+R 0.691 0.696 0.693 0.684 0.688 0.685 0.685 0.689 0.687
    CodeKGC 0.611 0.614 0.612 0.605 0.607 0.606 0.607 0.609 0.608
    ChatIE 0.569 0.574 0.571 0.541 0.545 0.543 0.553 0.557 0.555

    While baseline in-prompt methods struggle on moderate schemas (Wiki-NRE) and frequently generate out-of-schema relations, EDC+R achieves higher F1 scores by resolving contextual ambiguities (e.g., distinguishing homonymous relations and resolving head-tail directionality) via explicit natural language schema definitions.

  9. Knowl 9 — Document-Level Knowledge Graph Extraction via Coreference Resolution Integration

    empirical result

    When applied to document-level knowledge graph construction on the Re-DOCRED benchmark (where input texts exceed single LLM context windows), EDC was paired with LingMess (a coreference resolution model) and sentence-level chunking.

    On Re-DOCRED, direct zero/few-shot prompting of LLMs achieves a strict micro F1 score of 0.060. Applying EDC alone achieves an F1 of 0.132. Integrating EDC with LingMess coreference resolution and chunking increases the strict micro F1 score to 0.234, demonstrating modular composability with external information extraction pipelines.

  10. Knowl 10 — Computational Cost and Pipeline Latency of Multi-Stage EDC

    limitation

    Because EDC executes multiple decoupled LLM calls per document (initial open extraction, schema definition generation, canonical verification MCQ, and optional refinement passes), it incurs notable computational latency and monetary cost. When using GPT-3.5-turbo across all modules, the processing cost is approximately 0.009 USD per example.

    Combining Phase 1 (OIE) and Phase 2 (Schema Definition) into a single prompt reduces prompt token usage from approximately 3,000 to 2,000 tokens per sample and yields comparable extraction performance (0.516 to 0.518 Partial F1 on REBEL), offering a partial mitigation for API overhead.

Coverage note — None was omitted; all key architectural components, training formulations, benchmark results, ablations, baseline comparisons, modular extensions, and stated limitations were fully extracted.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Oshin Agarwal, Heming Ge, Siamak Shakeri, and Rami Al-Rfou. 2020. Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training. arXiv preprint arXiv:2010.12688.
  3. 3.Zhen Bi, Jing Chen, Yinuo Jiang, Feiyu Xiong, Wei Guo, Huajun Chen, and Ningyu Zhang. 2024. Codekgc: Code language model for generative knowledge graph construction. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(3):1–16.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Pere-Lluís Huguet Cabot and Roberto Navigli. 2021. Rebel: Relation extraction by end-to-end language generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2370–2381.
  6. 6.Eunsol Choi, Omer Levy, Yejin Choi, and Luke Zettlemoyer. 2018. Ultra-fine entity typing. arXiv preprint arXiv:1807.04905.
  7. 7.Sarthak Dash, Gaetano Rossiello, Nandana Mihindukulasooriya, Sugato Bagchi, and Alfio Gliozzo. 2020. Open knowledge graphs canonicalization using variational autoencoders. arXiv preprint arXiv:2012.04780.
  8. 8.Bayu Distiawan, Gerhard Weikum, Jianzhong Qi, and Rui Zhang. 2019. Neural relation extraction for knowledge base enrichment. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 229–240.
  9. 9.Pierre L Dognin, Inkit Padhi, Igor Melnyk, and Payel Das. 2021. Regen: Reinforcement learning for text and knowledge base generation using pretrained language models. arXiv preprint arXiv:2108.12472.
  10. 10.Thiago Castro Ferreira, Claire Gardent, Nikolai Ilinykh, Chris Van Der Lee, Simon Mille, Diego Moussallem, and Anastasia Shimorina. 2020. The 2020 bilingual, bi-directional webnlg+ shared task overview and evaluation results (webnlg+ 2020). In Proceedings of the 3rd International Workshop on Natural Language Generation from the Semantic Web (WebNLG+).
  11. 11.Debasis Ganguly, Dwaipayan Roy, Mandar Mitra, and Gareth JF Jones. 2015. Word embedding based generalized language model for information retrieval. In Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval, pages 795–798.
  12. 12.Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch. 2013. Ppdb: The paraphrase database. In Proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: Human language technologies, pages 758–764.
  13. 13.Liang Guo, Fu Yan, Yuqian Lu, Ming Zhou, and Tao Yang. 2021. An automatic machining process decision-making system based on knowledge graph. International journal of computer integrated manufacturing, 34(12):1348–1369.
  14. 14.Qingyu Guo, Fuzhen Zhuang, Chuan Qin, Hengshu Zhu, Xing Xie, Hui Xiong, and Qing He. 2020. A survey on knowledge graph-based recommender systems. IEEE Transactions on Knowledge and Data Engineering, 34(8):3549–3568.
  15. 15.Harsha Gurulingappa, Abdul Mateen Rajput, Angus Roberts, Juliane Fluck, Martin Hofmann-Apitius, and Luca Toldo. 2012. Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports. Journal of biomedical informatics, 45(5):885–892.
  16. 16.Ridong Han, Tao Peng, Chaohao Yang, Benyou Wang, Lu Liu, and Xiang Wan. 2023. Is information extraction solved by chatgpt? an analysis of performance, evaluation criteria, robustness and errors. arXiv preprint arXiv:2305.14450.
  17. 17.Xiao Huang, Jingyuan Zhang, Dingcheng Li, and Ping Li. 2019. Knowledge graph embedding based question answering. In Proceedings of the twelfth ACM international conference on web search and data mining, pages 105–113.
  18. 18.Shaoxiong Ji, Shirui Pan, Erik Cambria, Pekka Marttinen, and S Yu Philip. 2021. A survey on knowledge graphs: Representation, acquisition, and applications. IEEE transactions on neural networks and learning systems, 33(2):494–514.
  19. 19.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  20. 20.Martin Josifoski, Nicola De Cao, Maxime Peyrard, Fabio Petroni, and Robert West. 2022. GenIE: Generative information extraction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4626–4643, Seattle, United States. Association for Computational Linguistics.
  21. 21.Serafina Kamp, Morteza Fayazi, Zineb Benameur-El, Shuyan Yu, and Ronald Dreslinski. 2023. Open information extraction: A review of baseline techniques, approaches, and applications. arXiv preprint arXiv:2310.11644.
  22. 22.Keshav Kolluru, Vaibhav Adlakha, Samarth Aggarwal, Soumen Chakrabarti, et al. 2020. Openie6: Iterative grid labeling and coordination analysis for open information extraction. arXiv preprint arXiv:2010.03147.
  23. 23.Luong Thi Hong Lan, Tran Manh Tuan, Tran Thi Ngan, Nguyen Long Giang, Vo Truong Nhu Ngoc, Pham Van Hai, et al. 2020. A new complex fuzzy inference system with fuzzy knowledge graph and extensions in decision making. Ieee Access, 8:164899–164921.
  24. 24.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461.
  25. 25.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  26. 26.Bo Li, Gexiang Fang, Yang Yang, Quansen Wang, Wei Ye, Wen Zhao, and Shikun Zhang. 2023. Evaluating chatgpt’s information extraction capabilities: An assessment of performance, explainability, calibration, and faithfulness. arXiv preprint arXiv:2304.11633.
  27. 27.Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173.
  28. 28.Pai Liu, Wenyang Gao, Wenjie Dong, Songfang Huang, and Yue Zhang. 2022. Open information extraction from 2007 to 2022–a survey. arXiv preprint arXiv:2208.08690.
  29. 29.Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018. Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction. arXiv preprint arXiv:1808.09602.
  30. 30.Pedro Henrique Martins, Zita Marinho, and André FT Martins. 2019. Joint learning of named entity recognition and entity linking. arXiv preprint arXiv:1907.08243.
  31. 31.Igor Melnyk, Pierre Dognin, and Payel Das. 2022. Knowledge graph generation from text. arXiv preprint arXiv:2211.10511.
  32. 32.George A Miller. 1995. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39–41.
  33. 33.Yasumasa Onoe and Greg Durrett. 2020. Fine-grained entity typing for domain independent entity linking. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8576–8583.
  34. 34.Shon Otmazgin, Arie Cattan, and Yoav Goldberg. 2023. LingMess: Linguistically informed multi expert scorers for coreference resolution. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2752–2760, Dubrovnik, Croatia. Association for Computational Linguistics.
  35. 35.Rifki Afina Putri, Giwon Hong, and Sung-Hyon Myaeng. 2019. Aligning open ie relations and kb relations using a siamese network based on word embedding. In Proceedings of the 13th International Conference on Computational Semantics-Long Papers, pages 142–153.
  36. 36.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67.
  37. 37.Dan Roth and Wen-tau Yih. 2004. A linear programming formulation for global inference in natural language tasks. In Proceedings of the eighth conference on computational natural language learning (CoNLL-2004) at HLT-NAACL 2004, pages 1–8.
  38. 38.Alisa Smirnova and Philippe Cudré-Mauroux. 2018. Relation extraction using distant supervision: A survey. ACM Computing Surveys (CSUR), 51(5):1–35.
  39. 39.Rhea Sukthanker, Soujanya Poria, Erik Cambria, and Ramkumar Thirunavukarasu. 2020. Anaphora and coreference resolution: A review. Information Fusion, 59:139–162.
  40. 40.Qingyu Tan, Lu Xu, Lidong Bing, Hwee Tou Ng, and Sharifah Mahani Aljunied. 2022. Revisiting docred – addressing the false negative problem in relation extraction. In Proceedings of EMNLP.
  41. 41.Shikhar Vashishth, Prince Jain, and Partha Talukdar. 2018. Cesi: Canonicalizing open knowledge bases using embeddings and side information. In Proceedings of the 2018 World Wide Web Conference, pages 1317–1327.
  42. 42.Denny Vrandečić and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Communications of the ACM, 57(10):78–85.
  43. 43.Somin Wadhwa, Silvio Amir, and Byron C Wallace. 2023. Revisiting relation extraction in the era of large language models. In Proceedings of the conference. Association for Computational Linguistics. Meeting, volume 2023, page 15566. NIH Public Access.
  44. 44.Hongwei Wang, Miao Zhao, Xing Xie, Wenjie Li, and Minyi Guo. 2019. Knowledge graph convolutional networks for recommender systems. In The world wide web conference, pages 3307–3313.
  45. 45.Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2023. Improving text embeddings with large language models. arXiv preprint arXiv:2401.00368.
  46. 46.Xiang Wei, Xingyu Cui, Ning Cheng, Xiaobin Wang, Xin Zhang, Shen Huang, Pengjun Xie, Jinan Xu, Yufeng Chen, Meishan Zhang, et al. 2023. Zero-shot information extraction via chatting with chatgpt. arXiv preprint arXiv:2302.10205.
  47. 47.Yaqi Xie, Chen Yu, Tongyao Zhu, Jinbin Bai, Ze Gong, and Harold Soh. 2023. Translating natural language to planning goals with large-language models. arXiv preprint arXiv:2302.05128.
  48. 48.Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024. Hallucination is inevitable: An innate limitation of large language models. arXiv preprint arXiv:2401.11817.
  49. 49.Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. 2024. A survey on large language model (llm) security and privacy: The good, the bad, and the ugly. High-Confidence Computing, page 100211.
  50. 50.Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021. Qa-gnn: Reasoning with language models and knowledge graphs for question answering. arXiv preprint arXiv:2104.06378.
  51. 51.Hongbin Ye, Ningyu Zhang, Hui Chen, and Huajun Chen. 2022. Generative knowledge graph construction: A review. arXiv preprint arXiv:2210.12714.
  52. 52.Daojian Zeng, Kang Liu, Yubo Chen, and Jun Zhao. 2015. Distant supervision for relation extraction via piecewise convolutional neural networks. In Proceedings of the 2015 conference on empirical methods in natural language processing, pages 1753–1762.
  53. 53.Daojian Zeng, Kang Liu, Siwei Lai, Guangyou Zhou, and Jun Zhao. 2014. Relation classification via convolutional deep neural network. In Proceedings of COLING 2014, the 25th international conference on computational linguistics: technical papers, pages 2335–2344.
  54. 54.Bowen Zhang and Harold Soh. 2023. Large language models as zero-shot human models for human-robot interaction. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 7961–7968. IEEE.
  55. 55.Lingfeng Zhong, Jia Wu, Qian Li, Hao Peng, and Xindong Wu. 2023. A comprehensive survey on automatic knowledge graph construction. ACM Computing Surveys, 56(4):1–62.
  56. 56.Shaowen Zhou, Bowen Yu, Aixin Sun, Cheng Long, Jingyang Li, Haiyang Yu, Jian Sun, and Yongbin Li. 2022. A survey on neural open information extraction: Current status and future directions. arXiv preprint arXiv:2205.11725.
  57. 57.Andrej Žukov-Gregorič, Yoram Bachrach, and Sam Coope. 2018. Named entity recognition with parallel recurrent neural networks. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 69–74.

Citation

MLA
Zhang, B., and H. Soh. “Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 9820–36, https://doi.org/10.18653/v1/2024.emnlp-main.548.
APA
Zhang, B., & Soh, H. (2024). Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 9820–9836. https://doi.org/10.18653/v1/2024.emnlp-main.548
Chicago
Zhang, B., and H. Soh. 2024. “Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 9820–36. https://doi.org/10.18653/v1/2024.emnlp-main.548.
Harvard
Zhang, B. and Soh, H. (2024) “Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 9820–9836. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.548.
Vancouver
1. Zhang B, Soh H (2024) Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 9820–9836

BibTeX

@inproceedings{zhang-soh-2024-extract,
    title = "Extract, Define, Canonicalize: An {LLM}-based Framework for Knowledge Graph Construction",
    author = "Zhang, Bowen  and
      Soh, Harold",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.548/",
    doi = "10.18653/v1/2024.emnlp-main.548",
    pages = "9820--9836"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/