Instruct and Extract: Instruction Tuning for On-Demand Information Extraction
Yizhu JiaoMing ZhongSha LiRuining ZhaoSiru OuyangHeng JiJiawei Han
Proposes an instruction-tuned extraction framework and benchmark, INSTRUCTIE, that enables language models to convert unstructured text into custom tabular formats based on either explicit user specifications or inferred contextual headers.
Traditional natural language processing models for information extraction rely on rigid, predefined categories and schemas. In practice, non-expert users frequently require custom, ad hoc data extraction from unstructured text that does not fit conventional templates. While modern instruction-tuned artificial intelligence models offer conversational flexibility, open-source models often fail to extract accurate, structured tabular data on demand.
The article introduces and evaluates a new task called On-Demand Information Extraction, which converts unstructured text into organized tables based on user instructions. The authors demonstrate that an open-source model specialized with targeted, high-quality synthetic training data can effectively handle diverse extraction queries across multiple domains.
To benchmark and train models for this task, the authors created a dataset named INSTRUCTIE. It comprises 14,579 automatically generated training pairs spanning 84 domains and a human-annotated test set of 150 diverse cases. The training data generation used multi-stage automated pipelines followed by strict quality filtering across four dimensions: validity, informativeness, consistency with instructions, and faithfulness to the source text. Using this dataset, the authors trained ODIE, an on-demand information extractor built on an open-source seven-billion-parameter foundation model using parameter-efficient fine-tuning, and evaluated it against existing open-source and commercial systems.
ODIE substantially outperformed leading open-source models of comparable size, achieving an overall header similarity score of 73.8% compared to 69.4% for the best-performing baseline, despite using only a small fraction of the training data volume. In content extraction, ODIE achieved a 45.9% ROUGE-L score, surpassing the best open-source alternative by over 5 percentage points. Quality filtering proved critical, boosting overall content extraction accuracy by at least 2.6 percentage points compared to unfiltered versions. While open instructions requiring inferred table headers remained challenging across all models, ODIE nearly matched commercial proprietary models in identifying correct table headers.
These findings indicate that organizations can achieve specialized, high-accuracy data extraction with smaller, cost-effective open-source language models through targeted data curation and filtering. This approach reduces the need to deploy costly commercial models or manually build rigid extraction pipelines for custom tasks. However, the results also show a persistent gap between open-source models and massive proprietary systems when extracting dense, complex table contents requiring advanced reasoning.
Organizations seeking to implement on-demand extraction should prioritize multi-faceted data validation and decide between direct extraction and step-by-step reasoning prompts based on the query type. Direct prompting works best for explicit instructions, whereas step-by-step reasoning better suits open-ended prompts where headers must be inferred. Before enterprise deployment, technical teams should test model scalability across larger parameter sizes and establish fine-grained evaluation metrics tailored to structural table accuracy.
Readers should note that the evaluation was limited to seven-billion-parameter foundation models and a 150-instance test set. The models also exhibited higher error rates when processing real retrieved web text compared to generated text, and frequently struggled with tasks requiring multi-step reasoning. Confidence is high regarding the effectiveness of instruction filtering, but cautious validation is advised for open-ended queries in specialized operational environments.
- Paper: Text-to-Table: A New Way of Information Extraction, Xueqing Wu et al. (2022). This text-to-table study establishes the schema-free table-extraction task that Instruct and Extract adapts to user-specified, on-demand instructions.
- Paper: Unified Structure Generation for Universal Information Extraction, Yaojie Lu et al. (2022). UIE shows how prompting can unify diverse extraction tasks, providing a foundation for understanding the source’s instruction-driven extraction setup.
- Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). Self-Instruct introduces synthetic instruction-data generation, a key prerequisite for understanding the source’s use of generated training examples.
- Paper: GenIE: Generative Information Extraction, Martin Josifoski et al. (2022). GenIE establishes generative extraction from unstructured text, clarifying the shift away from the rigid extraction pipelines the source addresses.
No sufficiently relevant recommendations were found.
