Instruct and Extract: Instruction Tuning for On-Demand Information Extraction

Yizhu JiaoMing ZhongSha LiRuining ZhaoSiru OuyangHeng JiJiawei Han

article2023EMNLP67 citations

Proposes an instruction-tuned extraction framework and benchmark, INSTRUCTIE, that enables language models to convert unstructured text into custom tabular formats based on either explicit user specifications or inferred contextual headers.

Listen

Traditional natural language processing models for information extraction rely on rigid, predefined categories and schemas. In practice, non-expert users frequently require custom, ad hoc data extraction from unstructured text that does not fit conventional templates. While modern instruction-tuned artificial intelligence models offer conversational flexibility, open-source models often fail to extract accurate, structured tabular data on demand.

The article introduces and evaluates a new task called On-Demand Information Extraction, which converts unstructured text into organized tables based on user instructions. The authors demonstrate that an open-source model specialized with targeted, high-quality synthetic training data can effectively handle diverse extraction queries across multiple domains.

To benchmark and train models for this task, the authors created a dataset named INSTRUCTIE. It comprises 14,579 automatically generated training pairs spanning 84 domains and a human-annotated test set of 150 diverse cases. The training data generation used multi-stage automated pipelines followed by strict quality filtering across four dimensions: validity, informativeness, consistency with instructions, and faithfulness to the source text. Using this dataset, the authors trained ODIE, an on-demand information extractor built on an open-source seven-billion-parameter foundation model using parameter-efficient fine-tuning, and evaluated it against existing open-source and commercial systems.

ODIE substantially outperformed leading open-source models of comparable size, achieving an overall header similarity score of 73.8% compared to 69.4% for the best-performing baseline, despite using only a small fraction of the training data volume. In content extraction, ODIE achieved a 45.9% ROUGE-L score, surpassing the best open-source alternative by over 5 percentage points. Quality filtering proved critical, boosting overall content extraction accuracy by at least 2.6 percentage points compared to unfiltered versions. While open instructions requiring inferred table headers remained challenging across all models, ODIE nearly matched commercial proprietary models in identifying correct table headers.

These findings indicate that organizations can achieve specialized, high-accuracy data extraction with smaller, cost-effective open-source language models through targeted data curation and filtering. This approach reduces the need to deploy costly commercial models or manually build rigid extraction pipelines for custom tasks. However, the results also show a persistent gap between open-source models and massive proprietary systems when extracting dense, complex table contents requiring advanced reasoning.

Organizations seeking to implement on-demand extraction should prioritize multi-faceted data validation and decide between direct extraction and step-by-step reasoning prompts based on the query type. Direct prompting works best for explicit instructions, whereas step-by-step reasoning better suits open-ended prompts where headers must be inferred. Before enterprise deployment, technical teams should test model scalability across larger parameter sizes and establish fine-grained evaluation metrics tailored to structural table accuracy.

Readers should note that the evaluation was limited to seven-billion-parameter foundation models and a 150-instance test set. The models also exhibited higher error rates when processing real retrieved web text compared to generated text, and frequently struggled with tasks requiring multi-step reasoning. Confidence is high regarding the effectiveness of instruction filtering, but cautious validation is advised for open-ended queries in specialized operational environments.

No sufficiently relevant recommendations were found.

Cover for Instruct and Extract: Instruction Tuning for On-Demand Information Extraction

Abstract

Large language models with instruction-following capabilities open the door to a wider group of users. However, when it comes to information extraction – a classic task in natural language processing – most task-specific systems cannot align well with long-tail ad hoc extraction use cases for non-expert users. To address this, we propose a novel paradigm, termed On-Demand Information Extraction, to fulfill the personalized demands of real-world users. Our task aims to follow the instructions to extract the desired content from the associated text and present it in a structured tabular format. The table headers can either be user-specified or inferred contextually by the model. To facilitate research in this emerging area, we present a benchmark named INSTRUCTIE, inclusive of both automatically generated training data, as well as the human-annotated test set. Building on INSTRUCTIE, we further develop an On-Demand Information Extractor, ODIE. Comprehensive evaluations on our benchmark reveal that ODIE substantially outperforms the existing open-source models of similar size. Our code and dataset are released on https://github.com/yzjiao/On-Demand-IE.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Information Extraction
  • 2.2 Instruction Tuning
  • 3 INSTRUCTIE Dataset
  • 3.1 Task Formulation
  • 3.2 Automatic Generation of Training Data
  • 3.3 Manual Annotation for Test Data
  • 3.4 Dataset Analysis
  • 4 Experiment
  • 4.1 Model Training - ODIE
  • 4.2 Evaluation Metrics
  • 4.3 Evaluated Models
  • 4.4 Table Header Evaluation
  • 4.5 Table Content Evaluation
  • 4.6 Human Evaluation
  • 4.7 Evaluation Metric Analysis
  • 5 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Domain of Training Data
  • B Prompting Templates for Data Generation
  • C Human Evaluation Details
  • C.1 Human Evaluation Setup
  • C.2 Human Evaluation Agreement
  • D Implementation Details
  • D.1 Training Data Generation
  • D.2 Model Training
  • E Case Study

Knowls

  1. Knowl 1 — On-demand information extraction is instruction-conditioned table generation

    definition

    On-demand information extraction takes a user instruction II and associated background text XX, and produces a structured table TT containing the requested information. The first row of TT is the header; later rows contain extracted content. A fixed instruction explicitly names the information types to extract, so the model should use those as table headers. An open instruction does not name the headers, requiring the model to infer suitable headers from the instruction and text before extracting their values. Thus, the output schema may be specified by the user or inferred contextually rather than selected from a predefined ontology.

  2. Knowl 2 — INSTRUCTIE training examples are generated through a six-stage instruction-to-table pipeline

    model/method

    INSTRUCTIE uses ChatGPT to synthesize varied instruction-conditioned extraction examples. The pipeline begins with five demonstrations and generates 10 fixed instructions with distinct domains per iteration; the fixed instructions specify both the information to extract and the text domain. For each fixed instruction, ChatGPT generates a corresponding background text containing the requested information. It then generates an open instruction from the text alone, rather than directly converting the fixed instruction, to encourage distinct open requests. Instructions are paraphrased in four styles—comprehensive query, casual interaction, direct command, and professional request—while preserving their key extraction requirements. For each instruction-text pair, the pipeline generates a Markdown table either directly or with an explanatory paragraph before the table (the Chain-of-Thought, or CoT, variant). The table-generation prompt asks for only specified headers for fixed instructions, but as many relevant columns as possible for open instructions.

    In the reported implementation, 500 iterations yield 5,000 fixed instructions and texts; a corresponding open instruction is generated for each text. The fixed and open requests produce 10,000 instruction-text pairs, which are paraphrased and used to generate direct and CoT table outputs. The final training set contains 7,483 direct and 7,096 CoT instances after filtering.

  3. Knowl 3 — Synthetic tables are filtered for format, information content, instruction alignment, and textual support

    model/method

    INSTRUCTIE filters generated tables using four checks. Validity removes outputs that do not conform to a table format. Informativeness requires more than one column, a sum of row and column counts greater than 3, and fewer than 4 cells containing N/A. Instruction consistency is assessed for fixed instructions by treating the instruction as a premise and a sentence of the form “extract HH from the text” as the hypothesis for each table header HH. Faithfulness to the text is assessed for each cell CC under header HH, using the hypothesis “The HH is CC” and the background text as premise. For the two NLI-based checks, a neural evaluator scores consistency or support, and instances are retained only when the relevant average score exceeds 0.5.

    The reported counts removed by validity, informativeness, consistency, and faithfulness checks, respectively, are 438, 653, 552, and 874 direct instances, and 961, 250, 696, and 997 CoT instances. These are the reported removals by check; the final dataset sizes are 7,483 direct and 7,096 CoT instances.

  4. Knowl 4 — INSTRUCTIE pairs broad synthetic training data with a manually curated test set

    data/table

    INSTRUCTIE contains model-generated training examples and 150 human-annotated test examples spanning diverse extraction requests. Training instructions and texts are synthetic and cover more than 80 domains. The test set covers 61 domains; its background texts were retrieved from the web in 119 cases and generated with GPT-4 in 31 privacy-sensitive cases. Test instructions are human-written, and the test tables are reviewed by annotators; GPT-4-generated tables serve as annotation references. The reported dataset statistics are:

    Statistic Direct CoT Train total Test
    Instructions 7,483 7,096 14,579 150
    Open instructions 3,773 3,676 7,449 36
    Fixed instructions 3,710 3,420 7,130 114
    Domains 82 82 83 61
    Mean instruction length (words) 20.4 20.3 20.4 26.8
    Mean text length (words) 281.2 277.8 279.6 310.8
    Mean table cells 20.8 23.2 22.0 17.1
    Mean table rows 3.6 3.7 3.6 4.7
    Mean table columns 6.2 6.7 6.5 4.1

    The table shows that the training set is almost evenly divided between fixed and open instructions, while the test set contains more fixed instructions. CoT training tables average 2.4 more cells than direct tables. Test instances were categorized across easy, medium, and hard levels, with the set described as approximately balanced.

  5. Knowl 5 — ODIE adapts LLaMA-7B to extraction tables with LoRA instruction tuning

    model/method

    ODIE is an instruction-tuned information extractor built by fine-tuning LLaMA-7B with LoRA on INSTRUCTIE. Examples use a chatbot-style sequence with system, user, and assistant tokens; training cross-entropy loss is applied only to tokens in the assistant response. Direct and CoT training use different system prompts, with the CoT prompt requesting an explanatory paragraph before the table and the direct prompt omitting that request.

    The reported LoRA rank is 16 and dropout is 0.05. Training uses batch size 64, maximum learning rate 3×10−43\times10^{-4}, and warmup ratio 0.03. ODIE is trained for 20 epochs; the comparison models use the same backbone, LoRA settings, and training paradigm, but different training data and epoch counts. At inference, the shared settings are temperature 0.1, top-pp 0.75, top-kk 40, 4 beams, and maximum generation length 2,048 tokens.

  6. Knowl 6 — Evaluation separates header understanding from extracted table content

    experimental setup

    Evaluation measures table headers and table contents separately because headers reflect interpretation of the user instruction, while cells reflect extraction quality. Header semantic similarity is computed by embedding headers with SentenceBERT and taking cosine similarity for soft matching; results are reported as F1. Content quality is principally reported as ROUGE-L F1 between generated and reference table output. The study also reports exact-match and semantic-similarity F1 scores for both headers and contents, and human ratings of header correctness and content accuracy. The principal header comparison uses soft-matching F1, and the principal content comparison uses ROUGE-L F1.

  7. Knowl 7 — ODIE-DIRECT approaches proprietary-model performance on header soft matching

    empirical result

    On the 150-example test set, ODIE-DIRECT achieves 73.82% header soft-matching F1 overall, 4.38 percentage points above the strongest open-source baseline, TÜLU at 69.44%, and close to CHATGPT at 74.49% and GPT-4 at 74.47%. Results by instruction category show that CoT performs better than direct training on open instructions, whereas direct training performs better on fixed instructions. Removing filtering changes the overall scores only slightly for this header metric.

    Model Overall Fixed Open
    ALPACA 59.80 65.89 45.69
    TÜLU 69.44 77.78 49.26
    ODIE-DIRECT 73.82 83.59 51.67
    ODIE-DIRECT, no filtering 73.61 83.54 51.37
    ODIE-COT 66.81 72.32 54.17
    ODIE-COT, no filtering 64.11 68.97 53.01
    CHATGPT 74.49 81.69 57.86
    GPT-4 74.47 82.06 57.78

    All entries are soft-matching F1 scores in percent. ODIE-COT’s open-instruction score is 2.50 points above ODIE-DIRECT’s, while its fixed-instruction score is 11.27 points lower. The authors suggest that CoT may help with the variability of open instructions but may encourage extra, instruction-inconsistent headers for fixed instructions.

  8. Knowl 8 — ODIE-DIRECT improves open-source baselines on content ROUGE-L, while GPT-4 remains strongest

    empirical result

    For table content, ODIE-DIRECT reaches 45.93% overall ROUGE-L F1, compared with 40.72% for TÜLU, the best open-source baseline. Filtering raises ODIE-DIRECT from 42.93% to 45.93% and ODIE-COT from 39.53% to 42.12%. GPT-4 scores 59.06%, so the open-source models, including ODIE, remain behind the proprietary models. The table gives the reported ROUGE-L F1 scores in percent across difficulty, instruction category, and text source; the overall column aggregates the test set.

    Model Easy Medium Hard Fixed Open Generated Retrieved Overall
    ALPACA 26.27 20.08 22.72 25.29 16.07 30.37 21.18 23.08
    TÜLU 43.69 39.15 38.68 42.55 34.94 45.08 39.59 40.72
    ODIE-DIRECT 48.01 45.38 43.71 47.19 41.92 49.49 45.00 45.93
    ODIE-DIRECT, no filtering 46.31 39.85 42.44 44.42 38.23 46.71 41.95 42.93
    ODIE-COT 44.47 39.99 41.73 43.02 39.25 49.01 40.32 42.12
    ODIE-COT, no filtering 41.26 36.59 41.20 40.08 38.93 45.92 37.87 39.53
    CHATGPT 52.45 50.56 51.07 53.21 45.66 55.97 50.21 51.40
    GPT-4 60.78 55.76 61.24 61.51 51.29 65.89 57.28 59.06

    Across models, fixed instructions score better than open instructions. Retrieved texts are generally more difficult than generated texts, which the authors attribute to the more structured format of generated contexts. Some models score higher on hard than medium examples; the authors associate this with medium cases requiring complete extraction of large, complex tables, while hard cases may allow reasonable partial answers to reasoning or summarization requests.

  9. Knowl 9 — Human ratings support ODIE's header gains, and semantic metrics track judgments better than content exact match

    empirical result

    Three evaluators rated anonymized outputs from ALPACA, TÜLU, ODIE, and GPT-4 on the test instructions. Header ratings used three levels—wrong, partly correct, and correct—and content ratings used four levels, from correct and satisfying to irrelevant or invalid. ODIE’s header ratings were comparable to GPT-4’s, with 9 completely wrong header votes for each, and better than ALPACA and TÜLU. For content, GPT-4 received the strongest ratings; ODIE showed clear gains over ALPACA and TÜLU, particularly in the correct/satisfying and acceptable-with-minor-imperfection categories.

    The study also correlated automatic metrics with human evaluations. Pearson, Spearman, and Kendall coefficients, respectively, were 0.640, 0.609, and 0.494 for header exact match; 0.817, 0.769, and 0.637 for header semantic similarity; 0.338, 0.375, and 0.272 for content exact match; 0.764, 0.705, and 0.558 for content semantic similarity; and 0.713, 0.704, and 0.554 for content ROUGE-L. Thus, semantic similarity correlates most strongly for headers, while content semantic similarity and ROUGE-L correlate substantially better than content exact match.

  10. Knowl 10 — The study leaves model scaling, combined training, fine-grained evaluation, and complex inference unresolved

    limitation

    The experiments focus on a 7B backbone and the available training-data scale, so they do not establish how larger models or larger datasets affect on-demand extraction performance or efficiency. Direct and CoT training are evaluated separately; whether combining them has complementary benefits is untested. The automatic metrics mainly capture output similarity and do not assess flexible table structure, organization, and cell-level correctness with fine granularity. Finally, the current model still has difficulty inferring table structures contextually and handling complex instructions, including requests that require reasoning or summarization.

Coverage note — No substantial contributed material was deliberately omitted; implementation details and per-example case studies were left out because they support, rather than add to, the central task, dataset, model, and evaluation findings.

References

  1. 1.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback.
  2. 2.Junwei Bao, Duyu Tang, Nan Duan, Zhao Yan, Yuanhua Lv, Ming Zhou, and Tiejun Zhao. 2018. Table-to-text: Describing table region with natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 32.
  3. 3.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  4. 4.Yixin Cao, Zikun Hu, Tat-Seng Chua, Zhiyuan Liu, and Heng Ji. 2019. Low-resource name tagging learned with weakly labeled data. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 261–270. Association for Computational Linguistics.
  5. 5.Sahil Chaudhary. 2023. Code alpaca: An instruction-following llama model for code generation. https://github.com/sahil280114/codealpaca.
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. Blog post.
  7. 7.Databricks. 2023. Databricks’ dolly, a large language model trained on the databricks machine learning platform. https://github.com/databrickslabs/dolly.
  8. 8.Shumin Deng, Ningyu Zhang, Jiaojian Kang, Yichi Zhang, Wei Zhang, and Huajun Chen. 2020. Meta-learning with dynamic-memory-based prototypical network for few-shot event detection. In WSDM ’20: The Thirteenth ACM International Conference on Web Search and Data Mining, Houston, TX, USA, February 3-7, 2020, pages 151–159. ACM.
  9. 9.Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpaca-farm: A simulation framework for methods that learn from human feedback. CoRR, abs/2305.14387.
  10. 10.Xinyang Geng, Arnav Gudibande, Hao Liu, Eric Wallace, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. Koala: A dialogue model for academic research. Blog post.
  11. 11.Xu Han, Hao Zhu, Pengfei Yu, Ziyun Wang, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2018. Fewrel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 4803–4809. Association for Computational Linguistics.
  12. 12.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022. Unnatural instructions: Tuning language models with (almost) no human labor. CoRR, abs/2212.09689.
  13. 13.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  14. 14.Jiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose, Shobana Balakrishnan, Weizhu Chen, Baolin Peng, Jianfeng Gao, and Jiawei Han. 2021. Few-shot named entity recognition: An empirical baseline study. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 10408–10423. Association for Computational Linguistics.
  15. 15.Heng Ji and Ralph Grishman. 2008. Refining event extraction through cross-document inference. In ACL 2008, Proceedings of the 46th Annual Meeting of the Association for Computational Linguistics, June 15-20, 2008, Columbus, Ohio, USA, pages 254–262. The Association for Computer Linguistics.
  16. 16.Yizhu Jiao, Sha Li, Yiqing Xie, Ming Zhong, Heng Ji, and Jiawei Han. 2022. Open-vocabulary argument role prediction for event extraction. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 5404–5418, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  17. 17.Yizhu Jiao, Ming Zhong, Jiaming Shen, Yunyi Zhang, Chao Zhang, and Jiawei Han. 2023. Unsupervised event chain mining from multiple documents. In Proceedings of the ACM Web Conference 2023, WWW 2023, Austin, TX, USA, 30 April 2023 - 4 May 2023, pages 1948–1959. ACM.
  18. 18.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. 2023. Openassistant conversations - democratizing large language model alignment. CoRR, abs/2304.07327.
  19. 19.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. In NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016, pages 260–270. The Association for Computational Linguistics.
  20. 20.Gina-Anne Levow. 2006. The third international chinese language processing bakeoff: Word segmentation and named entity recognition. In Proceedings of the Fifth Workshop on Chinese Language Processing, SIGHAN@COLING/ACL 2006, Sydney, Australia, July 22-23, 2006, pages 108–117. Association for Computational Linguistics.
  21. 21.Bo Li, Gexiang Fang, Yang Yang, Quansen Wang, Wei Ye, Wen Zhao, and Shikun Zhang. 2023. Evaluating chatgpt’s information extraction capabilities: An assessment of performance, explainability, calibration, and faithfulness. CoRR, abs/2304.11633.
  22. 22.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  23. 23.Ying Lin, Heng Ji, Fei Huang, and Lingfei Wu. 2020. A joint neural model for information extraction with global features. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 7999–8009. Association for Computational Linguistics.
  24. 24.Ying Lin, Liyuan Liu, Heng Ji, Dong Yu, and Jiawei Han. 2019. Reliability-aware dynamic feature composition for name tagging. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 165–174. Association for Computational Linguistics.
  25. 25.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts. 2023. The flan collection: Designing data and methods for effective instruction tuning. CoRR, abs/2301.13688.
  26. 26.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 3470–3487. Association for Computational Linguistics.
  27. 27.OpenAI. 2023. Gpt-4 technical report. ArXiv, abs/2303.08774.
  28. 28.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In NeurIPS.
  29. 29.Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, and Jiawei Han. 2023. The shifted and the overlooked: a task-oriented investigation of user-gpt interactions. In the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics.
  30. 30.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction tuning with GPT-4. CoRR, abs/2304.03277.
  31. 31.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 3980–3990. Association for Computational Linguistics.
  32. 32.Erik F. Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the conll-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning, CoNLL 2003, Held in cooperation with HLT-NAACL 2003, Edmonton, Canada, May 31 - June 1, 2003, pages 142–147. ACL.
  33. 33.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  34. 34.Dianbo Sui, Yubo Chen, Kang Liu, Jun Zhao, Xiangrong Zeng, and Shengping Liu. 2020. Joint entity and relation extraction with set prediction networks. CoRR, abs/2011.01675.
  35. 35.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  36. 36.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  37. 37.Xiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang, Rong Han, Zhiyuan Liu, Juanzi Li, Peng Li, Yankai Lin, and Jie Zhou. 2020. MAVEN: A massive general domain event detection dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 1652–1671. Association for Computational Linguistics.
  38. 38.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023. How far can camels go? exploring the state of instruction tuning on open resources. CoRR, abs/2306.04751.
  39. 39.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022a. Self-instruct: Aligning language model with self generated instructions. CoRR, abs/2212.10560.
  40. 40.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, and Xudong Shen. 2022b. Super-naturalinstructions: Generalization via declarative instructions on 1600+ NLP tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 5085–5109. Association for Computational Linguistics.
  41. 41.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS.
  42. 42.Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al. 2013. Ontonotes release 5.0 ldc2013t19. Linguistic Data Consortium, Philadelphia, PA, 23.
  43. 43.Xueqing Wu, Jiacheng Zhang, and Hang Li. 2022. Text-to-table: A new way of information extraction. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 2518–2533. Association for Computational Linguistics.
  44. 44.Yuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, and Maosong Sun. 2019. Docred: A large-scale document-level relation extraction dataset. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 764–777. Association for Computational Linguistics.
  45. 45.Dian Yu, Lifu Huang, and Heng Ji. 2017. Open relation extraction and grounding. In Proceedings of the Eighth International Joint Conference on Natural Language Processing, IJCNLP 2017, Taipei, Taiwan, November 27 - December 1, 2017 - Volume 1: Long Papers, pages 854–864. Asian Federation of Natural Language Processing.
  46. 46.Pengfei Yu and Heng Ji. 2023. Shorten the long tail for rare entity and event extraction. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May 2-6, 2023, pages 1331–1342. Association for Computational Linguistics.
  47. 47.Pengfei Yu, Heng Ji, and Prem Natarajan. 2021. Lifelong event detection with knowledge transfer. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 5278–5290. Association for Computational Linguistics.
  48. 48.Qiusi Zhan, Sha Li, Kathryn Conger, Martha Palmer, Heng Ji, and Jiawei Han. 2023. GLEN: general-purpose event detection for thousands of types. CoRR, abs/2303.09093.
  49. 49.Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D. Manning. 2017. Position-aware attention and supervised data improve slot filling. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017, pages 35–45. Association for Computational Linguistics.
  50. 50.Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022. Towards a unified multi-dimensional evaluator for text generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 2023–2038. Association for Computational Linguistics.

Citation

MLA
Jiao, Y., et al. “Instruct and Extract: Instruction Tuning for On-Demand Information Extraction”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 10030–51, https://doi.org/10.18653/v1/2023.emnlp-main.620.
APA
Jiao, Y., Zhong, M., Li, S., Zhao, R., Ouyang, S., Ji, H., & Han, J. (2023). Instruct and Extract: Instruction Tuning for On-Demand Information Extraction. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10030–10051. https://doi.org/10.18653/v1/2023.emnlp-main.620
Chicago
Jiao, Y., M. Zhong, S. Li, et al. 2023. “Instruct and Extract: Instruction Tuning for On-Demand Information Extraction”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10030–51. https://doi.org/10.18653/v1/2023.emnlp-main.620.
Harvard
Jiao, Y. et al. (2023) “Instruct and Extract: Instruction Tuning for On-Demand Information Extraction”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10030–10051. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.620.
Vancouver
1. Jiao Y, Zhong M, Li S, Zhao R, Ouyang S, Ji H, Han J (2023) Instruct and Extract: Instruction Tuning for On-Demand Information Extraction. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10030–10051

BibTeX

@inproceedings{jiao-etal-2023-instruct,
    title = "Instruct and Extract: Instruction Tuning for On-Demand Information Extraction",
    author = "Jiao, Yizhu  and
      Zhong, Ming  and
      Li, Sha  and
      Zhao, Ruining  and
      Ouyang, Siru  and
      Ji, Heng  and
      Han, Jiawei",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.620/",
    doi = "10.18653/v1/2023.emnlp-main.620",
    pages = "10030--10051"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/