GPT-RE: In-context Learning for Relation Extraction using Large Language Models
Zhen WanFei ChengZhuoyuan MaoQianying LiuHaiyue SongJiwei LiSadao Kurohashi
Proposes GPT-RE, a framework that bridges the performance gap between large language models and fully supervised baselines on relation extraction benchmarks by combining task-aware demonstration retrieval with gold-label-induced reasoning.
Large language models have shown remarkable success across various natural language processing tasks through in-context learning, where a model makes predictions based on a small set of example demonstrations without updating its internal weights. Despite this progress, applying these models to relation extraction—the task of identifying target semantic relationships between specific entity pairs in text or flagging them as having no relation—has consistently underperformed compared to traditional, fully supervised fine-tuned models. This performance gap stems from two primary issues: standard retrieval methods select demonstration examples based on overall sentence similarity rather than entity-specific relevance, and prompts lack explanatory reasoning to help the model understand complex input-to-label relationships.
The article introduces GPT-RE, a framework designed to bridge this gap and demonstrate that large language models can match or exceed fine-tuned systems on relation extraction tasks. The objective is to evaluate how task-aware example retrieval and label-induced reasoning logic improve extraction accuracy and mitigate common failure modes such as false positives on unrelated entity pairs.
To evaluate this framework, the authors conducted experiments across four standard benchmarks: three general-domain datasets (SemEval-2010 Task 8, TACRED, and ACE05) and one scientific dataset (SciERC). The approach incorporates two core mechanisms. First, it uses task-aware demonstration selection via either entity-prompted sentence embeddings or representations extracted from a fine-tuned relation extraction model. Second, it enriches prompt demonstrations with intermediate reasoning explanations generated by the language model using ground-truth labels. Performance was assessed against standard random and sentence-similarity selection baselines as well as state-of-the-art fully supervised models.
The analysis yields four key findings. First, GPT-RE with fine-tuned representations consistently outperformed previous language model baselines and matched or exceeded fully supervised systems, achieving new top-tier performance on SemEval (91.90 micro-F1) and SciERC (69.00 micro-F1) while remaining competitive on TACRED (72.14 micro-F1) and ACE05 (68.73 micro-F1). Second, high-quality demonstration selection proved more impactful than quantity; providing just 5 highly relevant demonstrations outperformed 30 demonstrations retrieved via standard methods. Third, adding label-induced reasoning explanations improved accuracy by approximately 2% micro-F1 in settings with fewer demonstrations. Finally, the framework significantly reduced the tendency of language models to overpredict relations when no relation actually existed, particularly when using supervised representations that effectively recognize non-relation instances.
These findings indicate that large language models do not need extensive architectural fine-tuning to excel at structured extraction tasks, provided that demonstration examples are tightly aligned with task-specific entity features. This shift reduces the need to develop and maintain separate, dedicated models for distinct extraction tasks, lowering development timelines and operational complexity. The results challenge earlier assumptions that large language models are inherently ill-suited for relation extraction.
Organizations developing information extraction pipelines should consider adopting hybrid architectures that pair lightweight, specialized retrievers with general-purpose language models for inference. When prompt space is constrained, teams should prioritize demonstration quality and reasoning explanations over simply increasing the raw number of examples. In extremely low-resource settings with fewer than 650 labeled examples, relying on prompt-based large language models delivers superior accuracy compared to training purely supervised models.
These conclusions carry a few operational boundaries. While the framework substantially curtails the false-positive classification of unrelated pairs, recall on non-relation instances still lags behind fully supervised systems on datasets with extreme class imbalances where non-relations exceed 90%. Additionally, API costs required evaluating subset samples on larger datasets rather than full corpuses, and proprietary access constraints prevented using the large language model's native internal embeddings for the retrieval step. Overall confidence in the performance improvements remains high across evaluated standard benchmarks, provided the deployment distribution aligns with these operational parameters.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). This paper establishes the foundational understanding of how in-context learning mechanisms process demonstration mappings and prompt formats, motivating GPT-RE's exploration into enriching input-label reasoning logic.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). This work introduces prompt engineering and retrieval methods to extract relational facts from pretrained language models, providing essential background for formulating relation extraction prompts.
- Paper: Relation Classification via Convolutional Deep Neural Network, Daojian Zeng et al. (2014). This foundational paper establishes the standard benchmark SemEval-2010 Task 8 and core entity-distance modeling for relation classification evaluated directly in GPT-RE.
- Paper: Attention-Based Bidirectional Long Short-Term Memory Networks for Relation Classification, Peng Zhou et al. (2016). This paper presents standard neural baselines and context attention mechanisms for relation classification that modern in-context learning frameworks aim to match and surpass.
- Paper: ERNIE: Enhanced Language Representation with Informative Entities, Zhengyan Zhang et al. (2019). This study details entity-informed representations and the FewRel relation extraction setup that inform task-aware demonstration selection in GPT-RE.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). StructGPT expands upon in-context reasoning with structured and relational data by introducing an iterative reading-and-reasoning interface for black-box large language models.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). Plan-on-Graph builds on structured reasoning over entity-relation networks by enabling large language models to dynamically backtrack and self-correct across knowledge graphs.
- Paper: Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning, Xiaoxin He et al. (2024). This paper generalizes the use of generated textual reasoning explanations to enhance downstream graph representation learning.
- Paper: Making Text Embedders Few-Shot Learners, Chaofan Li et al. (2025). This work adapts in-context few-shot demonstration techniques to train task-aware text embedding models for improved retrieval performance.
