Revisiting Relation Extraction in the era of Large Language Models
Somin WadhwaSilvio AmirByron C. Wallace
Demonstrates that fine-tuning smaller models like Flan-T5 on Chain-of-Thought explanations generated by GPT-3 achieves state-of-the-art relation extraction performance while revealing through human evaluation that standard exact-match metrics substantially underestimate generative language model capabilities.
Extracting structured facts and semantic relationships from unstructured text is a foundational capability for modern enterprise information retrieval, automated knowledge curation, and biomedical intelligence. Traditional approaches rely on training specialized models to classify pre-identified entity tokens. While generative sequence-to-sequence models offer a more unified approach, the emergence of modern large language models presents an opportunity to reassess whether massive general-purpose models or smaller, adapted models are optimal for extracting structured relations.
The article evaluates whether generative large language models can perform relation extraction effectively across few-shot and supervised settings. It specifically assesses the performance of a proprietary massive model (GPT-3) against an open-source, smaller architecture (Flan-T5 Large) across standard benchmarks.
The authors conducted empirical evaluations across benchmark datasets spanning biomedical texts (ADE) and general news corpora (CoNLL04 and NYT). Because generative language models express equivalent answers using varied phrasing, standard exact-match metrics mischaracterize true performance. To address this, the authors recruited human annotators to rigorously re-assess ostensible model errors. They subsequently developed a training method that prompts GPT-3 to produce step-by-step reasoning explanations—known as Chain-of-Thought explanations—and uses these reasoning chains alongside gold labels to fine-tune the open-source Flan-T5 Large model.
The article reveals several critical findings. First, strict exact-match evaluation severely penalizes generative models; human re-evaluation showed that over 50% of ostensible false positives on ADE and CoNLL04 were actually correct facts omitted or mislabeled in the original benchmark references. Second, few-shot prompting of GPT-3 with only 12 to 20 examples matches or slightly exceeds traditional fully supervised state-of-the-art models on datasets with compact relation sets, achieving micro-F1 scores of 82.66 on ADE and 76.53 on CoNLL04. Third, few-shot prompting fails when schemas are complex; on the 24-relation NYT dataset, GPT-3's score dropped by roughly 30 F1 points due to prompt length constraints. Fourth, Flan-T5 Large performs poorly in few-shot settings (yielding high rates of out-of-domain relations), but fine-tuning it with GPT-3-generated Chain-of-Thought explanations achieves new state-of-the-art performance across all evaluated datasets, improving micro-F1 by 3.37 to 9.97 points over previous fully supervised methods.
These findings indicate that organizations do not need to rely permanently on expensive, opaque external APIs for high-accuracy relation extraction. Instead, teams can achieve superior accuracy, lower operational latency, and complete deployment control by leveraging a large model once to generate explanatory training data, and then distilling that reasoning into a smaller, privately hostable open-source model. The results also demonstrate that legacy benchmark evaluation pipelines underestimate generative extraction accuracy, meaning evaluation standards must be updated to avoid false rejections of viable systems.
Organizations should adopt open-source, instruction-tuned models like Flan-T5 fine-tuned with explanatory reasoning as the default architecture for production information extraction. When deploying generative extraction pipelines, engineering teams must implement semantic or model-based verification rather than rigid string-matching metrics. For long documents or wide relation schemas where in-context prompting fails, teams should prioritize fine-tuning over zero- or few-shot prompting.
These conclusions are bounded by specific conditions: the experiments examined only English-language texts with binary relationships, excluding complex multi-entity or cross-document extractions such as DocRED. Furthermore, the authors did not directly audit the factual accuracy of the GPT-3-generated explanations. Within these parameters, confidence remains high that explanatory distillation into compact models offers an optimal balance of cost, operational governance, and extraction performance.
- Paper: Generative Knowledge Graph Construction: A Review, Hongbin Ye et al. (2022). Provides a comprehensive foundational review of framing information and relation extraction as sequence-to-sequence generative modeling, which the source builds upon.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). Introduces the probing paradigm for evaluating pre-trained language models as implicit knowledge stores, establishing the foundational basis for prompt-based relational inference.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). Establishes essential methodology for mining and optimizing prompts to query factual relations from language models, directly informing few-shot prompting techniques for relation extraction.
- Paper: Relation Classification via Convolutional Deep Neural Network, Daojian Zeng et al. (2014). Presents standard neural benchmark formulations and baselines for supervised relation classification on SemEval datasets that contextualize the source's comparative evaluations.
- Paper: Attention-Based Bidirectional Long Short-Term Memory Networks for Relation Classification, Peng Zhou et al. (2016). Serves as a classical baseline for supervised relation extraction using neural attention networks before the shift toward generative large language models.
- Paper: GPT-RE: In-context Learning for Relation Extraction using Large Language Models, Zhen Wan et al. (2023). Directly extends LLM-based relation extraction by introducing task-aware demonstration retrieval and label-induced chain-of-thought explanations for in-context learning.
- Paper: Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction, Bowen Zhang et al. (2024). Generalizes LLM-driven relation extraction pipelines into an end-to-end knowledge graph construction framework using open extraction, definition generation, and canonicalization.
- Paper: Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning, Xiaoxin He et al. (2024). Applies the concept of distilling LLM-generated explanations and predictions into smaller language models for downstream relational graph learning tasks.
- Paper: Empirical Study of Zero-Shot NER with ChatGPT, Tingyu Xie et al. (2023). Explores zero-shot prompting and structured reasoning strategies for entity-level information extraction using ChatGPT across specialized domains.
