Revisiting Relation Extraction in the era of Large Language Models

Somin WadhwaSilvio AmirByron C. Wallace

article2023ACL271 citations

Demonstrates that fine-tuning smaller models like Flan-T5 on Chain-of-Thought explanations generated by GPT-3 achieves state-of-the-art relation extraction performance while revealing through human evaluation that standard exact-match metrics substantially underestimate generative language model capabilities.

Listen

Extracting structured facts and semantic relationships from unstructured text is a foundational capability for modern enterprise information retrieval, automated knowledge curation, and biomedical intelligence. Traditional approaches rely on training specialized models to classify pre-identified entity tokens. While generative sequence-to-sequence models offer a more unified approach, the emergence of modern large language models presents an opportunity to reassess whether massive general-purpose models or smaller, adapted models are optimal for extracting structured relations.

The article evaluates whether generative large language models can perform relation extraction effectively across few-shot and supervised settings. It specifically assesses the performance of a proprietary massive model (GPT-3) against an open-source, smaller architecture (Flan-T5 Large) across standard benchmarks.

The authors conducted empirical evaluations across benchmark datasets spanning biomedical texts (ADE) and general news corpora (CoNLL04 and NYT). Because generative language models express equivalent answers using varied phrasing, standard exact-match metrics mischaracterize true performance. To address this, the authors recruited human annotators to rigorously re-assess ostensible model errors. They subsequently developed a training method that prompts GPT-3 to produce step-by-step reasoning explanations—known as Chain-of-Thought explanations—and uses these reasoning chains alongside gold labels to fine-tune the open-source Flan-T5 Large model.

The article reveals several critical findings. First, strict exact-match evaluation severely penalizes generative models; human re-evaluation showed that over 50% of ostensible false positives on ADE and CoNLL04 were actually correct facts omitted or mislabeled in the original benchmark references. Second, few-shot prompting of GPT-3 with only 12 to 20 examples matches or slightly exceeds traditional fully supervised state-of-the-art models on datasets with compact relation sets, achieving micro-F1 scores of 82.66 on ADE and 76.53 on CoNLL04. Third, few-shot prompting fails when schemas are complex; on the 24-relation NYT dataset, GPT-3's score dropped by roughly 30 F1 points due to prompt length constraints. Fourth, Flan-T5 Large performs poorly in few-shot settings (yielding high rates of out-of-domain relations), but fine-tuning it with GPT-3-generated Chain-of-Thought explanations achieves new state-of-the-art performance across all evaluated datasets, improving micro-F1 by 3.37 to 9.97 points over previous fully supervised methods.

These findings indicate that organizations do not need to rely permanently on expensive, opaque external APIs for high-accuracy relation extraction. Instead, teams can achieve superior accuracy, lower operational latency, and complete deployment control by leveraging a large model once to generate explanatory training data, and then distilling that reasoning into a smaller, privately hostable open-source model. The results also demonstrate that legacy benchmark evaluation pipelines underestimate generative extraction accuracy, meaning evaluation standards must be updated to avoid false rejections of viable systems.

Organizations should adopt open-source, instruction-tuned models like Flan-T5 fine-tuned with explanatory reasoning as the default architecture for production information extraction. When deploying generative extraction pipelines, engineering teams must implement semantic or model-based verification rather than rigid string-matching metrics. For long documents or wide relation schemas where in-context prompting fails, teams should prioritize fine-tuning over zero- or few-shot prompting.

These conclusions are bounded by specific conditions: the experiments examined only English-language texts with binary relationships, excluding complex multi-entity or cross-document extractions such as DocRED. Furthermore, the authors did not directly audit the factual accuracy of the GPT-3-generated explanations. Within these parameters, confidence remains high that explanatory distillation into compact models offers an optimal balance of cost, operational governance, and extraction performance.

arXiv: 2305.05003
Cover for Revisiting Relation Extraction in the era of Large Language Models

Abstract

Relation extraction (RE) is the core NLP task of inferring semantic relationships between entities from text. Standard supervised RE techniques entail training modules to tag tokens comprising entity spans and then predict the relationship between them. Recent work has instead treated the problem as a sequence-to-sequence task, linearizing relations between entities as target strings to be generated conditioned on the input. Here we push the limits of this approach, using larger language models (GPT-3 and Flan-T5 large) than considered in prior work and evaluating their performance on standard RE tasks under varying levels of supervision. We address issues inherent to evaluating generative approaches to RE by doing human evaluations, in lieu of relying on exact matching. Under this refined evaluation, we find that: (1) Few-shot prompting with GPT-3 achieves near SOTA performance, i.e., roughly equivalent to existing fully supervised models; (2) Flan-T5 is not as capable in the few-shot setting, but supervising and fine-tuning it with Chain-of-Thought (CoT) style explanations (generated via GPT-3) yields SOTA results. We release this model as a new baseline for RE tasks¹.

Table of Contents

  • 1 Introduction
  • 2 RE via Text Generation
  • 3 In-Context Few-Shot Learning with GPT-3 for RE
  • 3.1 Prompts
  • 3.2 Manually re-evaluating 'errors'
  • 3.3 Results
  • 4 SOTA RE Performance with Flan-T5
  • 4.1 Few-Shot RE with Flan-T5
  • 4.2 Fine-tuning Flan-T5 for RE
  • 4.2.1 Eliciting CoT reasoning for RE
  • 4.2.2 Fine-tuning Flan-T5 with CoT explanations
  • 4.2.3 'Fully Supervising' Flan with GPT-3
  • 5 Related work
  • 5.1 Relation extraction with pre-trained LMs
  • 5.2 Few Shot In-Context Learning
  • 6 Conclusions and Future Directions
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Datasets
  • B Models and Reproducibility
  • B.1 Costs ( $$$ )
  • C Prompts
  • D Learning to Identify False False Positives and Negatives
  • D.1 List of out-of-domain relation-types generated by Flan during Few-Shot Prompting with CoNLL

Knowls

  1. Knowl 1 — Formulating Relation Extraction as Conditional Text Generation with Target Linearization

    model/method

    Relation extraction (RE) is formulated as an end-to-end conditional sequence-to-sequence text generation task. Given an input text xx and an optional prompt context CC containing nn in-context exemplar pairs (xi,yi)(x_i, y_i) with n≪Nn \ll N (where NN is the total dataset size), the conditional probability of generating the linearized target relation sequence y=(y1,y2,…,yT)y = (y_1, y_2, \dots, y_T) is defined autoregressively as:

    pLM(y∣C,x)=∏t=1Tp(yt∣C,x,y<t)p_{\mathrm{LM}}(y \mid C, x) = \prod_{t=1}^{T} p(y_t \mid C, x, y_{<t})

    Target relation structures are linearized into structured string formats:

    1. Single-relation datasets (e.g., ADE): Linearized as a list of entity tuples: [(drug, effect), ..., (drug, effect)]

    2. Multi-relation datasets (e.g., CoNLL04, NYT): Linearized as a list of typed entity triplets comprising a subject, relation, and object with entity types, ordered by the appearance of the subject entity in the input text: [(entity_1:entity_1_type, relation_type, entity_2:entity_2_type), ...]

    This generative formulation removes the requirement for specialized span classification and relation classification heads, allowing standard causal and encoder-decoder language models to extract relations directly.

  2. Knowl 2 — Fine-Tuning Flan-T5 with GPT-3-Generated Chain-of-Thought Explanations

    model/method

    To enhance relation extraction performance in smaller open-source language models, Flan-T5 (Large, 760M parameters) is fine-tuned on training datasets where the target outputs are augmented with Chain-of-Thought (CoT) reasoning explanations generated by GPT-3 (text-davinci-002).

    The training pipeline consists of two stages:

    1. CoT Explanation Generation via GPT-3: GPT-3 is prompted with an instructional prefix and 12 human-authored exemplars containing input texts, gold relation triplets, and step-by-step explanatory natural language justifications describing how the relations are inferred. GPT-3 then generates explanatory reasoning chains for all labeled training instances conditioned on the input text and gold relation reference labels.

    2. Seq2Seq Supervision: Flan-T5 Large is fine-tuned to generate both the target linearized relation triplets and the GPT-3-generated CoT explanation given the input sentence and instructional prefix.

    Fine-tuning hyperparameters for Flan-T5 Large:

    • ADE: batch size 8, learning rate 3×10−53 \times 10^{-5}, linear warmup 10%, 6 epochs (evaluated across 10 folds).
    • CoNLL04: batch size 4, learning rate 3×10−53 \times 10^{-5}, linear warmup 12%, 10 epochs.
    • NYT: batch size 4, learning rate 2×10−52 \times 10^{-5}, linear warmup 12%, 4 epochs (trained on a subset of 25,000 GPT-3-annotated examples).
  3. Knowl 3 — Empirical Relation Extraction Performance Across Fully Supervised and Few-Shot LLMs

    data/table

    Comparison of relation extraction micro-F1 scores across CoNLL04, ADE, and NYT benchmarks. A predicted relation triplet or pair is counted as correct only if both entity types and the relation label match the ground truth.

    Method Parameters CoNLL04 ADE NYT
    Fully Supervised Baselines
    SpERT 110M 71.54 79.22 –
    TANL 220M 71.48 80.61 90.83
    TANL (MT) 220M 72.66 80.00 90.52
    REBEL 460M 75.44 82.21 92.00
    Flan-T5 (Large) 760M 75.28 83.15 91.03
    Flan-T5 (Large) + GPT-3 CoT 760M 80.76 92.17 95.23
    Few-Shot In-Context Prompting
    GPT-3 (12–20 shots) 175B 76.53 82.66 61.79
    GPT-3 + CoT Prompting 175B 78.18 – –
    Flan-T5 (Large) w/ GPT-3 pseudo-labels + CoT 760M 76.13 – –

    Key findings demonstrated by these data:

    • Fine-tuning Flan-T5 Large with GPT-3-generated CoT explanations achieves state-of-the-art results across all datasets, outperforming the previous best supervised model (REBEL) by +5.32+5.32 F1 on CoNLL04, +9.96+9.96 F1 on ADE (averaged across 10 folds), and +3.23+3.23 F1 on NYT.
    • Few-shot prompting with GPT-3 using only 12--20 exemplars matches or outperforms existing fully supervised models on CoNLL04 (76.53, rising to 78.18 with CoT) and ADE (82.66).
    • Standard fine-tuning of Flan-T5 Large without CoT explanations achieves performance competitive with prior supervised models (75.28 on CoNLL04, 83.15 on ADE, 91.03 on NYT) but does not exceed them.
  4. Knowl 4 — Exact-Matching Evaluation Failure Modes and Manual Re-evaluation Rates in Generative RE

    empirical result

    Evaluating generative language models on relation extraction using strict surface-form exact matching against gold references artificially deflates measured precision and recall due to surface variations, semantic equivalence, and reference label errors.

    Human evaluation conducted via Amazon Mechanical Turk (requiring unanimous agreement across 3 annotators, achieving Fleiss' κ=0.83\kappa = 0.83) on ostensible errors produced by few-shot GPT-3 and fine-tuned Flan-T5 revealed high proportions of incorrect error designations:

    • ADE (GPT-3 few-shot outputs): 51.67% of ostensible false positives (108 out of 209) are valid relation extractions derivable from the text. 32.61% of ostensible false negatives (136 out of 417) are accurate true negatives (e.g., using synonymous medical surface terms or active chemical ingredients).

    • CoNLL04 (GPT-3 few-shot outputs): 50.27% of ostensible false positives (92 out of 183) are valid relations. 36.60% of ostensible false negatives (56 out of 152) are factual relations correctly extracted from text (such as predicting Work_For instead of erroneous reference labels like Live_In for diplomats).

    • NYT (Flan-T5 fine-tuned outputs): 36.90% of ostensible false positives and 22.97% of ostensible false negatives are judged accurate under human evaluation.

  5. Knowl 5 — Distillation of Flan-T5 Using GPT-3 Few-Shot Pseudo-Labels and CoT Explanations

    model/method

    A semi-supervised distillation method trains a smaller model (Flan-T5 Large, 760M parameters) using entirely synthetic supervision from few-shot GPT-3 (text-davinci-002) without requiring human-annotated training targets:

    1. GPT-3 is prompted with an instructional prefix and 12 seed exemplar sentences paired with human-written relation triplets and CoT explanations.
    2. GPT-3 generates both pseudo-labels (linearized relation triplets) and CoT explanations for the rest of the unlabeled training instances.
    3. Flan-T5 Large is fine-tuned on these GPT-3-generated pseudo-labels and CoT rationales using sequence-to-sequence training (batch size 4, learning rate 3×10−53 \times 10^{-5}, 10 epochs).

    On the CoNLL04 test set, Flan-T5 Large trained entirely on GPT-3 pseudo-labels and CoT explanations achieves a Micro-F1 of 76.13 (Precision 76.41, Recall 75.85), outperforming the fully supervised SOTA model REBEL (75.44) and in-context GPT-3 (76.53) while requiring only 12 human-annotated seed shots.

  6. Knowl 6 — Automating Generative RE Error Verification with Fine-Tuned BERT Classifiers

    model/method

    To semi-automate the identification of false false-positives and false false-negatives resulting from strict exact-matching evaluation in generative relation extraction, binary BERT classifiers are fine-tuned on human-annotated verification judgments.

    The classifiers format inputs using the [CLS] and [SEP] tokens:

    1. False Positive (FP) Classifier: Predicts whether an ostensible FP relation pair or triplet is genuinely supported by the text:

    [CLS] Input Text [SEP] Potential FP\text{[CLS] Input Text [SEP] Potential FP}

    1. False Negative (FN) Classifier: Predicts whether an ostensible reference FN triplet is semantically captured by the generated relation set:

    [CLS] Input Text [SEP] Potential FN [SEP] Generated Labels\text{[CLS] Input Text [SEP] Potential FN [SEP] Generated Labels}

    Classification performance evaluated by Area Under the ROC Curve (AUC-ROC):

    • CoNLL04 False Positives: AUC=0.88\text{AUC} = 0.88
    • ADE False Positives: AUC=0.75\text{AUC} = 0.75
    • ADE False Negatives: AUC=0.77\text{AUC} = 0.77
    • CoNLL04 False Negatives: AUC=0.73\text{AUC} = 0.73
  7. Knowl 7 — Prompt Design and Execution Setup for Few-Shot In-Context RE

    experimental setup

    In-context few-shot relation extraction with GPT-3 (text-davinci-002, sampling temperature 0.5, maximum output token limit 256) uses task-specific instructional prompting:

    • ADE: The prompt consists of the instructional prefix "List all [drug, adverse effects] pairs in the TEXT provided below." followed by 12 randomly selected in-context exemplar pairs (755 total prompt tokens).

    • CoNLL04: The prompt includes the prefix "List the relations of the types [OrgBased In, Work For, Located In, Live In, Kill] among the entities [PERSON, LOCATION, ORGANIZATION, OTHER] in the given text and provide a reasonable explanation." along with 12 manually curated exemplars ensuring at least one instance of each entity and relation type (960 total prompt tokens).

    • NYT: Because the dataset contains 24 relation types, specific relation descriptions are omitted from the prefix due to context window constraints. The prompt contains 20 exemplars covering all entity and relation types (2095 total prompt tokens).

  8. Knowl 8 — In-Context Context-Window Bottlenecks and Small LLM Few-Shot Degeneracy in RE

    limitation

    Generative relation extraction with language models exhibits critical limitations depending on model scale and schema complexity:

    • Prompt Context-Window Bottlenecks on Large Schemas: For datasets containing numerous target relations or long documents (e.g., NYT with 24 relations, DocRED with 96 relations across multi-sentence documents), fitting complete schema definitions and multi-shot exemplars within the in-context window is infeasible. Omitting relation type descriptors in NYT causes GPT-3 few-shot Micro-F1 to drop to 61.79 (o∼30 o \sim 30 points below fully supervised SOTA) with 10.6%10.6\% invalid or empty outputs, and completely precludes in-context prompting on DocRED.

    • Few-Shot Degeneracy in Sub-Billion Parameter LLMs: When evaluated in few-shot prompting with 7 exemplars, Flan-T5 Large (760M) produces high rates of non-conforming outputs (13.9%13.9\% on ADE and 12.5%12.5\% on CoNLL04, featuring repetitive token loops and malformed tuples), as well as over 120 unique out-of-domain relations on CoNLL04, resulting in a ∼20\sim 20 point micro-F1 drop relative to GPT-3.

    • Unverified Synthetic CoT Quality: The CoT-style natural language reasoning chains generated automatically by GPT-3 to supervise Flan-T5 are unverified and not filtered for reasoning accuracy or factual hallucination prior to fine-tuning.

Coverage note — No substantial contributed material was omitted; all primary empirical results, fine-tuning and distillation methodologies with CoT, prompt designs, human evaluation error analyses, and BERT-based verification classifiers are covered.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020a. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020b. Language models are few-shot learners. ArXiv, abs/2005.14165.
  3. 3.Mingda Chen, Jingfei Du, Ramakanth Pasunuru, Todor Mihaylov, Srini Iyer, Veselin Stoyanov, and Zornitsa Kozareva. 2022. Improving in-context few-shot learning via self-supervised training. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3558–3573, Seattle, United States. Association for Computational Linguistics.
  4. 4.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  6. 6.Markus Eberts and Adrian Ulges. 2019a. Span-based joint entity and relation extraction with transformer pre-training. ArXiv, abs/1909.07755.
  7. 7.Markus Eberts and Adrian Ulges. 2019b. Span-based joint entity and relation extraction with transformer pre-training. CoRR, abs/1909.07755.
  8. 8.Markus Eberts and Adrian Ulges. 2021. An end-to-end model for entity-level relation extraction using multi-instance learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3650–3660, Online. Association for Computational Linguistics.
  9. 9.Harsha Gurulingappa, Abdul Mateen Rajput, Angus Roberts, Juliane Fluck, Martin Hofmann-Apitius, and Luca Toldo. 2012. Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports. Journal of biomedical informatics, 45 5:885–92.
  10. 10.Pere-Lluís Huguet Cabot and Roberto Navigli. 2021. REBEL: Relation extraction by end-to-end language generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2370–2381, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  11. 11.John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proceedings of the Eighteenth International Conference on Machine Learning, ICML ’01, page 282–289, San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.
  12. 12.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  13. 13.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, Dublin, Ireland and Online. Association for Computational Linguistics.
  14. 14.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022a. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  15. 15.Yaojie Lu, Qing Liu, Dai Dai, Xinyan Xiao, Hongyu Lin, Xianpei Han, Le Sun, and Hua Wu. 2022b. Unified structure generation for universal information extraction. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5755–5772, Dublin, Ireland. Association for Computational Linguistics.
  16. 16.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022. MetaICL: Learning to learn in context. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2791–2809, Seattle, United States. Association for Computational Linguistics.
  17. 17.Tapas Nayak and Hwee Tou Ng. 2020. Effective modeling of encoder-decoder architecture for joint entity and relation extraction. In AAAI Conference on Artificial Intelligence.
  18. 18.Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Jie Ma, Alessandro Achille, RISHITA ANUBHAI, Cícero Nogueira dos Santos, Bing Xiang, and Stefano Soatto. 2021. Structured prediction as translation between augmented natural languages. In International Conference on Learning Representations.
  19. 19.Alec Radford and Karthik Narasimhan. 2018. Improving language understanding by generative pre-training.
  20. 20.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  21. 21.Sebastian Riedel, Limin Yao, and Andrew McCallum. 2010. Modeling relations and their mentions without labeled text. In ECML/PKDD.
  22. 22.Dan Roth and Wen-tau Yih. 2004. A linear programming formulation for global inference in natural language tasks. In Proceedings of the Eighth Conference on Computational Natural Language Learning (CoNLL-2004) at HLT-NAACL 2004, pages 1–8, Boston, Massachusetts, USA. Association for Computational Linguistics.
  23. 23.Olivier Taboureau, Sonny Kim Nielsen, Karine Audouze, Nils Weinhold, Daniel Edsgärd, Francisco S. Roque, Irene Kouskoumvekaki, Alina Bora, Ramona Curpan, Thomas Skøt Jensen, Søren Brunak, and Tudor I. Oprea. 2010. ChemProt: a disease chemical biology database. Nucleic Acids Research, 39:D367–D372.
  24. 24.Bruno Taillé, Vincent Guigue, Geoffrey Scoutheeten, and Patrick Gallinari. 2020. Let’s Stop Incorrect Comparisons in End-to-end Relation Extraction! In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3689–3701, Online. Association for Computational Linguistics.
  25. 25.Ioannis Tsochantaridis, Thomas Hofmann, Thorsten Joachims, and Yasemin Altun. 2004. Support vector machine learning for interdependent and structured output spaces. In Proceedings of the Twenty-First International Conference on Machine Learning, ICML ’04, page 104, New York, NY, USA. Association for Computing Machinery.
  26. 26.Chenguang Wang, Xiao Liu, Zui Chen, Haoyun Hong, Jie Tang, and Dawn Song. 2022. DeepStruct: Pre-training of language models for structure prediction. In Findings of the Association for Computational Linguistics: ACL 2022, pages 803–823, Dublin, Ireland. Association for Computational Linguistics.
  27. 27.Jue Wang and Wei Lu. 2020. Two are better than one: Joint entity and relation extraction with table-sequence encoders. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1706–1721, Online. Association for Computational Linguistics.
  28. 28.Shuohang Wang, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021. Want to reduce labeling cost? GPT-3 can help. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4195–4205, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  29. 29.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022a. Finetuned language models are zero-shot learners. ArXiv, abs/2109.01652.
  30. 30.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  31. 31.Hanwei Xu, Yujun Chen, Yulun Du, Nan Shao, Yanggang Wang, Haiyu Li, and Zhilin Yang. 2022. Zeroprompt: Scaling prompt-based pretraining to 1, 000 tasks improves zero-shot generalization. CoRR, abs/2201.06910.
  32. 32.Yuan Yao, Jiaju Du, Yankai Lin, Peng Li, Zhiyuan Liu, Jie Zhou, and Maosong Sun. 2021. CodRED: A cross-document relation extraction dataset for acquiring knowledge in the wild. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4452–4472, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  33. 33.Yuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, and Maosong Sun. 2019. DocRED: A large-scale document-level relation extraction dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 764–777, Florence, Italy. Association for Computational Linguistics.
  34. 34.Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. 2021. CrossFit: A few-shot learning challenge for cross-task generalization in NLP. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7163–7189, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  35. 35.Daojian Zeng, Haoran Zhang, and Qianying Liu. 2020. Copymtl: Copy mechanism for joint extraction of entities and relations with multi-task learning. ArXiv, abs/1911.10438.
  36. 36.Xiangrong Zeng, Daojian Zeng, Shizhu He, Kang Liu, and Jun Zhao. 2018. Extracting relational facts by an end-to-end neural model with copy mechanism. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 506–514, Melbourne, Australia. Association for Computational Linguistics.

Citation

MLA
Wadhwa, S., et al. “Revisiting Relation Extraction in the Era of Large Language Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 15566–89, https://doi.org/10.18653/v1/2023.acl-long.868.
APA
Wadhwa, S., Amir, S., & Wallace, B. C. (2023). Revisiting Relation Extraction in the era of Large Language Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15566–15589. https://doi.org/10.18653/v1/2023.acl-long.868
Chicago
Wadhwa, S., S. Amir, and B. C. Wallace. 2023. “Revisiting Relation Extraction in the Era of Large Language Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15566–89. https://doi.org/10.18653/v1/2023.acl-long.868.
Harvard
Wadhwa, S., Amir, S. and Wallace, B.C. (2023) “Revisiting Relation Extraction in the era of Large Language Models”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 15566–15589. Available at: https://doi.org/10.18653/v1/2023.acl-long.868.
Vancouver
1. Wadhwa S, Amir S, Wallace BC (2023) Revisiting Relation Extraction in the era of Large Language Models. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 15566–15589

BibTeX

@inproceedings{wadhwa-etal-2023-revisiting,
    title = "Revisiting Relation Extraction in the era of Large Language Models",
    author = "Wadhwa, Somin  and
      Amir, Silvio  and
      Wallace, Byron",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.868/",
    doi = "10.18653/v1/2023.acl-long.868",
    pages = "15566--15589"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/