GPT-RE: In-context Learning for Relation Extraction using Large Language Models

Zhen WanFei ChengZhuoyuan MaoQianying LiuHaiyue SongJiwei LiSadao Kurohashi

article2023EMNLP184 citations

Proposes GPT-RE, a framework that bridges the performance gap between large language models and fully supervised baselines on relation extraction benchmarks by combining task-aware demonstration retrieval with gold-label-induced reasoning.

Listen

Large language models have shown remarkable success across various natural language processing tasks through in-context learning, where a model makes predictions based on a small set of example demonstrations without updating its internal weights. Despite this progress, applying these models to relation extraction—the task of identifying target semantic relationships between specific entity pairs in text or flagging them as having no relation—has consistently underperformed compared to traditional, fully supervised fine-tuned models. This performance gap stems from two primary issues: standard retrieval methods select demonstration examples based on overall sentence similarity rather than entity-specific relevance, and prompts lack explanatory reasoning to help the model understand complex input-to-label relationships.

The article introduces GPT-RE, a framework designed to bridge this gap and demonstrate that large language models can match or exceed fine-tuned systems on relation extraction tasks. The objective is to evaluate how task-aware example retrieval and label-induced reasoning logic improve extraction accuracy and mitigate common failure modes such as false positives on unrelated entity pairs.

To evaluate this framework, the authors conducted experiments across four standard benchmarks: three general-domain datasets (SemEval-2010 Task 8, TACRED, and ACE05) and one scientific dataset (SciERC). The approach incorporates two core mechanisms. First, it uses task-aware demonstration selection via either entity-prompted sentence embeddings or representations extracted from a fine-tuned relation extraction model. Second, it enriches prompt demonstrations with intermediate reasoning explanations generated by the language model using ground-truth labels. Performance was assessed against standard random and sentence-similarity selection baselines as well as state-of-the-art fully supervised models.

The analysis yields four key findings. First, GPT-RE with fine-tuned representations consistently outperformed previous language model baselines and matched or exceeded fully supervised systems, achieving new top-tier performance on SemEval (91.90 micro-F1) and SciERC (69.00 micro-F1) while remaining competitive on TACRED (72.14 micro-F1) and ACE05 (68.73 micro-F1). Second, high-quality demonstration selection proved more impactful than quantity; providing just 5 highly relevant demonstrations outperformed 30 demonstrations retrieved via standard methods. Third, adding label-induced reasoning explanations improved accuracy by approximately 2% micro-F1 in settings with fewer demonstrations. Finally, the framework significantly reduced the tendency of language models to overpredict relations when no relation actually existed, particularly when using supervised representations that effectively recognize non-relation instances.

These findings indicate that large language models do not need extensive architectural fine-tuning to excel at structured extraction tasks, provided that demonstration examples are tightly aligned with task-specific entity features. This shift reduces the need to develop and maintain separate, dedicated models for distinct extraction tasks, lowering development timelines and operational complexity. The results challenge earlier assumptions that large language models are inherently ill-suited for relation extraction.

Organizations developing information extraction pipelines should consider adopting hybrid architectures that pair lightweight, specialized retrievers with general-purpose language models for inference. When prompt space is constrained, teams should prioritize demonstration quality and reasoning explanations over simply increasing the raw number of examples. In extremely low-resource settings with fewer than 650 labeled examples, relying on prompt-based large language models delivers superior accuracy compared to training purely supervised models.

These conclusions carry a few operational boundaries. While the framework substantially curtails the false-positive classification of unrelated pairs, recall on non-relation instances still lags behind fully supervised systems on datasets with extreme class imbalances where non-relations exceed 90%. Additionally, API costs required evaluating subset samples on larger datasets rather than full corpuses, and proprietary access constraints prevented using the large language model's native internal embeddings for the retrieval step. Overall confidence in the performance improvements remains high across evaluated standard benchmarks, provided the deployment distribution aligns with these operational parameters.

arXiv: 2305.02105
Cover for GPT-RE: In-context Learning for Relation Extraction using Large Language Models

Abstract

In spite of the potential for ground-breaking achievements offered by large language models (LLMs) (e.g., GPT-3) via in-context learning (ICL), they still lag significantly behind fully-supervised baselines (e.g., fine-tuned BERT) in relation extraction (RE). This is due to the two major shortcomings of ICL for RE: (1) low relevance regarding entity and relation in existing sentence-level demonstration retrieval approaches for ICL; and (2) the lack of explaining input-label mappings of demonstrations leading to poor ICL effectiveness.

In this paper, we propose GPT-RE to successfully address the aforementioned issues by (1) incorporating task-aware representations in demonstration retrieval; and (2) enriching the demonstrations with gold label-induced reasoning logic. We evaluate GPT-RE on four widely-used RE datasets and observe that GPT-RE achieves improvements over not only existing GPT-3 baselines, but also fully-supervised baselines as in Figure 1. Specifically, GPT-RE achieves SOTA performances on the SemEval and SciERC datasets, and competitive performances on the TACRED and ACE05 datasets.

Additionally, a critical issue of LLMs revealed by previous work, the strong inclination to wrongly classify NULL examples into other pre-defined labels, is substantially alleviated by our method. We show an empirical analysis.1

Table of Contents

  • 1 Introduction
  • 2 Methodology: GPT-RE
  • 2.1 Task Definition
  • 2.2 Overview
  • 2.3 Prompt Construction
  • 2.4 Task-aware Demonstration Retrieval
  • 2.4.1 Entity-Prompted Sentence Embedding
  • 2.4.2 Fine-tuned Relation Representation
  • 2.5 Gold Label-induced Reasoning
  • 3 Experiment Setup
  • 3.1 Datasets
  • 3.2 Baseline Methods
  • 4 Experimental Results
  • 4.1 Main Results
  • 4.2 Ablation Study on Task-aware Retrieval
  • 4.3 Ablation Study on Reasoning Enhancing
  • 4.4 Low-resource Scenario
  • 5 Analysis
  • 5.1 The Issue of 'Overpredicting'
  • 5.2 Case Study of Demonstration Quality
  • 6 Related Work
  • 7 Conclusions
  • Limitations
  • References
  • A Hyperparameters
  • A.1 GPT-3 Hyperparameters
  • A.2 Fine-tuning Baseline PURE
  • A.3 Sentence Embedding Methods
  • B Case Study
  • C Subset

Knowls

  1. Knowl 1 — GPT-RE In-Context Learning Framework for Relation Extraction

    model/method

    Relation extraction (RE) seeks to determine the semantic relation y∈Ry \in \mathcal{R} holding between a subject entity esub∈Ce_{\text{sub}} \in C and an object entity eobj∈Ce_{\text{obj}} \in C within a context sentence CC, or to output y=NULLy = \text{NULL} if no relation from the predefined set R\mathcal{R} applies.

    GPT-RE formulates relation extraction as a generative in-context learning (ICL) problem using large language models (such as GPT-3 text-davinci-003). The generation distribution is defined as:

    p(ytest∈R∪{NULL}∣I,D,xtest)p(y_{\text{test}} \in \mathcal{R} \cup \{\text{NULL}\} \mid \mathcal{I}, \mathcal{D}, x_{\text{test}})

    where:

    • I\mathcal{I} is the task instruction outlining the relation extraction objective and explicitly listing all candidate classes R\mathcal{R}, directing the model to generate NULL\text{NULL} if no candidate relation fits.
    • D={(xi,yi,ri)}i=1k\mathcal{D} = \{(x_i, y_i, r_i)\}_{i=1}^k is a set of kk retrieved demonstrations. Each demonstration consists of an input sentence with identified entities xix_i, the gold relation label yi∈R∪{NULL}y_i \in \mathcal{R} \cup \{\text{NULL}\}, and a gold label-induced reasoning rationale rir_i.
    • xtestx_{\text{test}} is the test context specifying the candidate subject and object entity pair.
    • ytesty_{\text{test}} is the predicted relation string.
  2. Knowl 2 — Fine-Tuned Relation Representation Retrieval for In-Context Demonstrations

    model/method

    To prevent retrieving demonstrations that share overall sentence topics but mismatch the target entity relation, GPT-RE uses representations extracted from a fine-tuned relation extraction model (such as PURE) to perform kk-nearest neighbor (kNNk\text{NN}) demonstration retrieval over the training corpus.

    An input sentence with subject esube_{\text{sub}} and object eobje_{\text{obj}} is formatted with typed entity markers:

    [CLS] [SUB_TYPE] esub [/SUB_TYPE] … [OBJ_TYPE] eobj [/OBJ_TYPE] … [SEP]\text{[CLS]} \, [\text{SUB}\_\text{TYPE}] \, e_{\text{sub}} \, [/\text{SUB}\_\text{TYPE}] \, \dots \, [\text{OBJ}\_\text{TYPE}] \, e_{\text{obj}} \, [/\text{OBJ}\_\text{TYPE}] \, \dots \, \text{[SEP]}

    where TYPE\text{TYPE} indicates the entity type. Let hih_i and hjh_j represent the hidden state vectors from the fine-tuned BERT encoder corresponding to the start markers [SUB_TYPE][\text{SUB}\_\text{TYPE}] and [OBJ_TYPE][\text{OBJ}\_\text{TYPE}]. The relation representation Rel\text{Rel} is formed by concatenating these hidden states:

    Rel=hi⊕hj\text{Rel} = h_i \oplus h_j

    where ⊕\oplus denotes vector concatenation along the first dimension. During inference, the retriever computes cosine similarities between Reltest\text{Rel}_{\text{test}} and candidate representations Reltrain\text{Rel}_{\text{train}} from the entire training set (including negative NULL\text{NULL} instances), returning the top-kk most similar examples as in-context demonstrations.

  3. Knowl 3 — Gold Label-Induced Reasoning for In-Context Demonstration Enrichment

    model/method

    Standard in-context demonstrations present input-label pairs (xi,yi)(x_i, y_i) without explanation, which can encourage large language models to rely on surface word correlations. GPT-RE enriches demonstrations with explicit reasoning explanations rir_i generated by querying GPT-3 using the ground-truth relation label.

    For each retrieved demonstration containing context sentence CC, subject entity esube_{\text{sub}}, object entity eobje_{\text{obj}}, and gold relation label y∈R∪{NULL}y \in \mathcal{R} \cup \{\text{NULL}\}, GPT-3 is prompted with the query:

    "What are the clues that lead to the relation between "esub" and "eobj" to be "y" in the sentence "ˊC"?ˊ"\text{"What are the clues that lead to the relation between "} e_{\text{sub}} \text{" and "} e_{\text{obj}} \text{" to be "} y \text{" in the sentence \'"} C \text{"\'?"}

    GPT-3 is prompted to output an explanation starting with "It is because: ...", highlighting the lexical, semantic, and syntactic evidence supporting label yy (or explaining why no predefined relation holds if y=NULLy = \text{NULL}). The generated explanation rir_i is appended to the demonstration, transforming it into (xi,yi,ri)(x_i, y_i, r_i) in the prompt.

  4. Knowl 4 — Entity-Prompted Sentence Embedding for Demonstration Retrieval

    model/method

    Generic sentence embeddings encode global sentence semantics without isolating the specific entity pair under evaluation. To inject entity-level focus without task fine-tuning, GPT-RE employs entity-prompted sentence embedding.

    Given a context sentence CC with subject esube_{\text{sub}} and object eobje_{\text{obj}}, the input string is restructured prior to embedding as:

    "The relation between ’"esub"’ and ’"eobj"’ in the context: "C\text{"The relation between '"} e_{\text{sub}} \text{"' and '"} e_{\text{obj}} \text{"' in the context: "} C

    Dense representations are computed using a contrastive sentence encoder (such as sup-simcse-bert-base-uncased). kk-nearest neighbor search is then performed in the resulting embedding space between the reconstructed test sentence and candidate reconstructed training sentences.

  5. Knowl 5 — Comparative Evaluation of GPT-RE Against Fine-Tuning and GPT Baselines

    data/table

    Micro-F1 performance across four relation extraction benchmarks: Semeval 2010 task 8 (9 relation classes, 17.40% NULL), TACRED (41 classes, 79.40% NULL), SciERC (7 classes, 90.16% NULL), and ACE05 (6 classes, 95.60% NULL). Experiments evaluate GPT-3 (text-davinci-003, temperature 0.0) with best kk-shot demonstrations indicated in parentheses.

    Methods Retriever Semeval TACRED SciERC ACE05
    GPT-3 Baselines (Best kk-shot)
    GPT-Random - 70.04 (30) 32.49 (15) 17.92 (25) 9.04 (25)
    GPT-Sent SimCSE 79.94 (30) 33.45 (15) 20.96 (25) 6.31 (25)
    GPT-RE (Best kk-shot)
    GPT-RE_SimCSE SimCSE 81.02 (30) 37.44 (15) 26.46 (25) 8.67 (25)
    GPT-RE_SimCSE* SimCSE 77.49 (15) 31.58 (10) - -
    + Reasoning SimCSE 79.88 (15) 33.18 (10) - -
    GPT-RE_FT PURE 91.90 (25) 72.14 (15) 69.00 (30) 68.73 (25)
    GPT-RE_FT* PURE 91.11 (15) 70.38 (10) - -
    + Reasoning PURE 91.82 (15) 70.97 (10) - -
    Fine-tuned RE Baselines
    Cohen et al. (2020) - 91.90 - - -
    Wang et al. (2022a) - - 76.80 - -
    PURE (Zhong and Chen, 2021) - 89.90 69.72 68.45 70.09

    Note: Asterisk () marks identical kk-shot counts configured for direct ablation with "+ Reasoning". Subsets preserving original class proportions were sampled for TACRED (1,600 examples) and ACE05 (2,442 examples).

    GPT-RE utilizing fine-tuned relation representations (GPT-RE_FT) surpasses the fully supervised PURE baseline on Semeval (+2.00 F1), TACRED (+2.42 F1), and SciERC (+0.55 F1), reaching state-of-the-art results on Semeval and SciERC.

  6. Knowl 6 — Demonstration Quality Outweighs Quantity in In-Context Relation Extraction

    empirical result

    Ablation experiments on Semeval evaluating demonstration count k∈[5,30]k \in [5, 30] show that representation quality in demonstration retrieval is far more impactful than demonstration volume:

    1. Retrieval-based in-context methods (GPT-Sent, GPT-RE_SimCSE, GPT-RE_FT) exhibit steep performance increases as kk increases, whereas randomly selected demonstrations (GPT-Random) achieve much lower gains, indicating that GPT-3 requires high-relevance examples to generalize effectively in relation extraction.
    2. Incorporating entity prompting into sentence embeddings (GPT-RE_SimCSE) consistently outperforms generic sentence-level retrieval (GPT-Sent) across all shot counts kk.
    3. Retrieval using fine-tuned relation representations (GPT-RE_FT) outperforms all alternative retrieval strategies by a substantial margin. With only k=5k = 5 demonstrations, GPT-RE_FT achieves 83.43 Micro-F1, surpassing the 30-shot performance of GPT-RE_SimCSE (80.30 Micro-F1). At k>15k > 15, GPT-RE_FT exceeds the supervised PURE fine-tuning baseline (89.90 Micro-F1).
  7. Knowl 7 — Effectiveness of Gold Label-Induced Reasoning Across Few-Shot Budgets

    empirical result

    Augmenting demonstrations with gold label-induced reasoning rationale consistently improves Micro-F1 on relation extraction tasks:

    • On Semeval with entity-prompted SimCSE retrieval, adding reasoning improves performance from 77.49% to 79.88% at k=15k=15 (+2.39% Micro-F1).
    • On TACRED with entity-prompted SimCSE retrieval, adding reasoning improves performance from 31.58% to 33.18% at k=10k=10 (+1.60% Micro-F1).
    • On Semeval with fine-tuned PURE retrieval, adding reasoning improves performance from 91.11% to 91.82% at k=15k=15 (+0.71% Micro-F1).
    • On TACRED with fine-tuned PURE retrieval, adding reasoning improves performance from 70.38% to 70.97% at k=10k=10 (+0.59% Micro-F1).

    The relative performance gain provided by explicit reasoning is largest when few demonstrations are provided (k≤10k \le 10). When a large number of high-quality demonstrations are retrieved via fine-tuned representations, the marginal utility of explicit reasoning diminishes, as the language model can infer the mapping logic directly from the demonstration context.

  8. Knowl 8 — In-Context Learning vs Fine-Tuning in Low-Resource RE Scenarios

    empirical result

    Evaluating relation extraction on Semeval across varying amounts of training data (from 100 to 1,950 examples, representing up to 30% of the training set) identifies two regimes:

    1. In low-data regimes where training examples are ≤10%\le 10\% (≤650\le 650 examples), all GPT-3 in-context learning configurations (GPT-Random, GPT-Sent, GPT-RE_SimCSE, GPT-RE_FT) outperform the fully fine-tuned BERT baseline, leveraging GPT-3's pretrained prior knowledge.
    2. GPT-RE_FT delivers the highest Micro-F1 across all data fractions. Even when the underlying PURE model is fine-tuned on only 100 to 400 training examples (where standalone classifier accuracy is low), the extracted relation representations still provide sufficient discriminatory quality for GPT-3 to retrieve relevant in-context demonstrations.
    3. Entity-prompted sentence embeddings (GPT-RE_SimCSE) only show marked improvements over generic sentence embeddings (GPT-Sent) after the training corpus size exceeds 30%, indicating that embedding-based entity retrieval requires a sufficiently dense pool of candidate examples.
  9. Knowl 9 — The Overpredicting Phenomenon on NULL Instances in LLM Relation Extraction

    empirical result

    In relation extraction tasks with negative (NULL\text{NULL}) instances, large language models exhibit an "overpredicting" tendency, frequently assigning valid real-world relationships outside the target schema to predefined relation classes instead of predicting NULL\text{NULL}:

    • On SciERC (where 90.16% of instances are NULL\text{NULL}), standard GPT-Random and GPT-Sent achieve only 17.92% and 20.96% Micro-F1. When NULL\text{NULL} examples are completely removed from training and evaluation, their performance rises to ~35% and ~42% Micro-F1, demonstrating the destructive impact of overpredicting NULL\text{NULL}.
    • Fine-tuned demonstration retrieval (GPT-RE_FT) alleviates this issue by retrieving accurate negative demonstrations from the training set, sustaining 69.00% Micro-F1 on SciERC with NULL\text{NULL}.
    • Despite this improvement, LLM in-context methods continue to exhibit lower recall on NULL\text{NULL} instances than fully supervised models on datasets with extreme class imbalance (such as ACE05 with 95.60% NULL), due to entity pairs possessing natural semantic associations that conflict with the closed ontology requirement.
  10. Knowl 10 — Limitations of GPT-RE

    limitation

    GPT-RE has two notable limitations:

    1. Residual NULL Overprediction: While task-aware retrieval significantly reduces the overprediction of NULL\text{NULL} instances into predefined classes, the recall on NULL\text{NULL} instances still lags behind fully supervised fine-tuned baselines, particularly on heavily skewed datasets such as ACE05 (95.60% NULL).
    2. External PLM Dependency for Retrieval: Demonstration retrieval depends on representations extracted from separate smaller pretrained language models (such as BERT or SimCSE) rather than representations from GPT-3 itself, due to API limitations restricting access to GPT-3's internal hidden states.

Coverage note — None. All main contributions, including the GPT-RE architecture, retrieval strategies, reasoning module, empirical comparisons across 4 datasets, ablation analyses, low-resource experiments, NULL overprediction analyses, and stated limitations, are fully covered.

References

  1. 1.Livio Baldini Soares, Nicholas FitzGerald, Jeffrey Ling, and Tom Kwiatkowski. 2019. Matching the blanks: Distributional similarity for relation learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2895–2905, Florence, Italy. Association for Computational Linguistics.
  2. 2.Michele Banko and Oren Etzioni. 2008. The tradeoffs between open and traditional relation extraction. In Proceedings of ACL-08: HLT, pages 28–36, Columbus, Ohio. Association for Computational Linguistics.
  3. 3.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3615–3620, Hong Kong, China. Association for Computational Linguistics.
  4. 4.Terra Blevins, Hila Gonen, and Luke Zettlemoyer. 2022. Prompting language models for linguistic structure. CoRR, abs/2211.07830.
  5. 5.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. CoRR, abs/2204.02311.
  8. 8.Amir D. N. Cohen, Shachar Rosenman, and Yoav Goldberg. 2020. Relation extraction as two-way span-prediction. CoRR, abs/2010.04829.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  10. 10.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  11. 11.Xiaoqing Geng, Xiwen Chen, Kenny Q. Zhu, Libin Shen, and Yinggong Zhao. 2020. MICK: A meta-learning framework for few-shot relation classification with small training data. In CIKM '20: The 29th ACM International Conference on Information and Knowledge Management, Virtual Event, Ireland, October 19-23, 2020, pages 415–424. ACM.
  12. 12.Bernal Jiménez Gutiérrez, Nikolas McNeal, Clay Washington, You Chen, Lang Li, Huan Sun, and Yu Su. 2022. Thinking about GPT-3 in-context learning for biomedical ie? think again. CoRR, abs/2203.08410.
  13. 13.Xu Han, Hao Zhu, Pengfei Yu, Ziyun Wang, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2018. FewRel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4803–4809, Brussels, Belgium. Association for Computational Linguistics.
  14. 14.Iris Hendrickx, Su Nam Kim, Zornitsa Kozareva, Preslav Nakov, Diarmuid Ó Séaghdha, Sebastian Padó, Marco Pennacchiotti, Lorenza Romano, and Stan Szpakowicz. 2010. SemEval-2010 task 8: Multi-way classification of semantic relations between pairs of nominals. In Proceedings of the 5th International Workshop on Semantic Evaluation, pages 33–38, Uppsala, Sweden. Association for Computational Linguistics.
  15. 15.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022. Training compute-optimal large language models.
  16. 16.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. CoRR, abs/2205.11916.
  17. 17.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations.
  18. 18.Fangchao Liu, Hongyu Lin, Xianpei Han, Boxi Cao, and Le Sun. 2022a. Pre-training to match for unified low-shot relation extraction. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5785–5795, Dublin, Ireland. Association for Computational Linguistics.
  19. 19.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022b. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, Dublin, Ireland and Online. Association for Computational Linguistics.
  20. 20.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  21. 21.Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018. Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3219–3232, Brussels, Belgium. Association for Computational Linguistics.
  22. 22.Nikolay Malkin, Zhen Wang, and Nebojsa Jojic. 2022. Coherence boosting: When your pretrained language model is not paying enough attention. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8214–8236, Dublin, Ireland. Association for Computational Linguistics.
  23. 23.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022a. Rethinking the role of demonstrations: What makes in-context learning work?
  24. 24.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022b. Rethinking the role of demonstrations: What makes in-context learning work? CoRR, abs/2202.12837.
  25. 25.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 11054–11070.
  26. 26.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training gopher.
  27. 27.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer.
  28. 28.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  29. 29.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2671, Seattle, United States. Association for Computational Linguistics.
  30. 30.Richard Shin, Christopher Lin, Sam Thomson, Charles Chen, Subhro Roy, Emmanouil Antonios Platanios, Adam Pauls, Dan Klein, Jason Eisner, and Benjamin Van Durme. 2021. Constrained language models yield few-shot semantic parsers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7699–7715, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  31. 31.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. 2022. Lamda: Language models for dialog applications.
  32. 32.Zhen Wan, Qianying Liu, Zhuoyuan Mao, Fei Cheng, Sadao Kurohashi, and Jiwei Li. 2022. Rescue implicit and long-tail cases: Nearest neighbor relation extraction. CoRR, abs/2210.11800.
  33. 33.Chenguang Wang, Xiao Liu, Zui Chen, Haoyun Hong, Jie Tang, and Dawn Song. 2022a. DeepStruct: Pre-training of language models for structure prediction. In Findings of the Association for Computational Linguistics: ACL 2022, pages 803–823, Dublin, Ireland. Association for Computational Linguistics.
  34. 34.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, and Denny Zhou. 2022b. Self-consistency improves chain of thought reasoning in language models. CoRR, abs/2203.11171.
  35. 35.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. CoRR, abs/2201.11903.
  36. 36.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  37. 37.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  38. 38.Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D. Manning. 2017. Position-aware attention and supervised data improve slot filling. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 35–45, Copenhagen, Denmark. Association for Computational Linguistics.
  39. 39.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 12697–12706. PMLR.
  40. 40.Zexuan Zhong and Danqi Chen. 2021. A frustratingly easy approach for entity and relation extraction. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 50–61, Online. Association for Computational Linguistics.
  41. 41.Liu Zhuang, Lin Wayne, Shi Ya, and Zhao Jun. 2021. A robustly optimized BERT pre-training approach with post-training. In Proceedings of the 20th Chinese National Conference on Computational Linguistics, pages 1218–1227, Huhhot, China. Chinese Information Processing Society of China.

Citation

MLA
Wan, Z., et al. “GPT-RE: In-context Learning for Relation Extraction Using Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 3534–47, https://doi.org/10.18653/v1/2023.emnlp-main.214.
APA
Wan, Z., Cheng, F., Mao, Z., Liu, Q., Song, H., Li, J., & Kurohashi, S. (2023). GPT-RE: In-context Learning for Relation Extraction using Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 3534–3547. https://doi.org/10.18653/v1/2023.emnlp-main.214
Chicago
Wan, Z., F. Cheng, Z. Mao, et al. 2023. “GPT-RE: In-context Learning for Relation Extraction Using Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 3534–47. https://doi.org/10.18653/v1/2023.emnlp-main.214.
Harvard
Wan, Z. et al. (2023) “GPT-RE: In-context Learning for Relation Extraction using Large Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3534–3547. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.214.
Vancouver
1. Wan Z, Cheng F, Mao Z, Liu Q, Song H, Li J, Kurohashi S (2023) GPT-RE: In-context Learning for Relation Extraction using Large Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3534–3547

BibTeX

@inproceedings{wan-etal-2023-gpt,
    title = "{GPT}-{RE}: In-context Learning for Relation Extraction using Large Language Models",
    author = "Wan, Zhen  and
      Cheng, Fei  and
      Mao, Zhuoyuan  and
      Liu, Qianying  and
      Song, Haiyue  and
      Li, Jiwei  and
      Kurohashi, Sadao",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.214/",
    doi = "10.18653/v1/2023.emnlp-main.214",
    pages = "3534--3547"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/