Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction
Bowen ZhangHarold Soh
Proposes a three-phase knowledge graph construction framework that overcomes large language model context window limits by extracting open triplets, generating definitions, and canonicalizing relations with or without a pre-defined schema.
Organizations increasingly rely on knowledge graphs—structured networks of entities and their relationships—to drive search, recommendation systems, and decision-making. However, constructing knowledge graphs from unstructured text has traditionally required intensive manual labor or fine-tuned, domain-specific models. While large language models offer strong text understanding, existing extraction techniques struggle to scale. Standard methods require inserting an entire predefined schema directly into the prompt, which quickly exceeds context window limits for large schemas and fails when no predetermined schema exists.
The article evaluates a modular, three-phase framework called Extract-Define-Canonicalize (EDC) to automate knowledge graph construction across complex text. The objective is to demonstrate that decomposing the extraction process into free extraction, definition generation, and post-processing canonicalization enables high-quality graph construction for both large predefined schemas and self-generated schemas without fine-tuning the base language model.
The evaluated approach breaks the task into open information extraction, generating natural language definitions for extracted relations, and standardizing those relations through vector search and model verification. An enhanced variant, EDC+R, incorporates an iterative refinement step driven by a trained Schema Retriever that supplies contextually relevant schema hints. The authors evaluated EDC across three established benchmarks with schemas containing up to 200 distinct relation types—WebNLG, REBEL, and Wiki-NRE—as well as a synthetic dataset of fictional entities. Performance was assessed against specialized supervised baselines and leading clustering techniques using standard precision, recall, F1 metrics, and blinded human evaluation.
The analysis yielded several key findings. First, EDC matched or exceeded state-of-the-art specialized models across all datasets, achieving partial F1 scores of up to 0.820 on WebNLG when paired with refinement and advanced language models. Second, the framework substantially outperformed constrained generative baselines on complex datasets like REBEL (0.601 vs. 0.385 F1) and Wiki-NRE (0.713 vs. 0.484 F1), primarily because it successfully extracts numerical and date literals. Third, when constructing graphs without any predefined schema, human evaluators confirmed that EDC produced highly accurate graphs (87–96% precision) with significantly more concise schemas and lower redundancy than existing clustering baselines, avoiding improper grouping of distinct relations. Finally, ablation studies showed that the Schema Retriever provided critical contextual disambiguation, boosting F1 performance across all tested benchmarks.
These findings indicate that organizations can build high-accuracy knowledge graphs without the significant cost of training specialized extraction models or manually maintaining rigid schemas. By handling schemas modularly after extraction, EDC eliminates the constraint limitations of standard prompting and allows systems to dynamically discover and integrate new relations into existing knowledge bases.
Stakeholders adopting this architecture should implement the retrieval-augmented refinement loop to capture subtle, fine-grained relationships. When deploying in production, engineering teams should evaluate hybrid architectures—such as combining the initial extraction and definition steps or replacing intermediate modules with smaller, specialized classifiers—to manage latency and language model operational expenses (observed at approximately $0.009 per example using commercial models).
While confidence in the core extraction methodology is high, readers should note certain operational boundaries. The evaluation relied on sentence- and paragraph-level inputs, meaning whole-document deployments will require upstream chunking and coreference resolution modules. Additionally, the study focused primarily on relation canonicalization rather than full entity deduplication. Further validation through pilot implementations in production environments will help establish exact cost-accuracy trade-offs across different open-source and proprietary language models.
- Paper: Generative Knowledge Graph Construction: A Review, Hongbin Ye et al. (2022). This review establishes the foundational sequence-to-sequence paradigms and output structures for generative knowledge graph construction that EDC builds upon and addresses schema scalability for.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). This survey provides the overarching framework for unifying large language models and knowledge graphs, detailing the trade-offs of generative extraction that motivate EDC's modular pipeline.
- Paper: Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning, Saibo Geng et al. (2023). It explores grammar-constrained decoding for structured extraction in LLMs, providing key context for the constrained generative baselines that EDC aims to outperform.
- Paper: Knowledge Graphs, Aidan Hogan et al. (2020). This comprehensive tutorial covers core knowledge graph concepts, relation schemas, and canonicalization principles essential for understanding the graph construction objectives evaluated in EDC.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). It demonstrates iterative prompting and interface retrieval for LLM reasoning over structured data, establishing key concepts underlying EDC's retrieval-augmented refinement loop.
- Paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization, Darren Edge et al. (2024). GraphRAG directly applies LLM-driven graph extraction over unstructured text documents to perform hierarchical community summarization and global question answering.
- Paper: Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based Retrofitting, Xinyan Guan et al. (2024). This work builds on LLM-extracted knowledge and graph retrieval to autonomously retrofit and mitigate hallucinations in generative model responses.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). It extends the integration of LLMs and structured knowledge bases by introducing dynamic, self-correcting path exploration and planning over knowledge graphs.
- Paper: GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer, Urchade Zaratiana et al. (2024). GLiNER presents a specialized, compact architecture for flexible, zero-shot entity extraction that addresses the latency and operational expense trade-offs highlighted in EDC.
