Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction
Martin JosifoskiMarija SakotaMaxime PeyrardRobert West
Demonstrates that generating input texts from sampled knowledge graph triplets exploits task asymmetry to produce balanced synthetic training datasets, enabling compact models to outperform prior closed information extraction methods by huge margins.
Extracting structured knowledge from natural language text—known as closed information extraction—is vital for building and populating knowledge bases. However, creating large-scale training datasets for this task is expensive and time-consuming because human annotators must master extensive catalogs of entities and relation types. Consequently, existing public datasets rely heavily on automated heuristics, resulting in noisy labels and severe class imbalances where rare relations are largely ignored. Furthermore, standard large language models cannot directly solve closed information extraction out of the box because they lack specific knowledge of predefined catalog schemas and identifiers.
The article demonstrates a reverse synthetic data generation strategy that leverages task asymmetry to solve this data bottleneck. Specifically, while asking a language model to extract rigid, structured facts from unstructured text is difficult, prompting the model to generate natural, fluent text from a provided structured set of facts is straightforward. The researchers evaluate whether this inverse approach can produce high-quality, balanced synthetic datasets and whether training compact language models on such data yields superior extraction performance.
To implement this approach, the authors constructed a graph from Wikidata covering 2.7 million entities and 888 relations. They developed a random-walk sampling technique with aggressive reweighting to extract 1.8 million coherent, balanced fact sets. They then prompted large language models to generate fluent natural text expressing only the facts in each set, creating a massive synthetic dataset. Using this generated data, they fine-tuned compact models with 220 million and 770 million parameters, called SynthIE, and benchmarked them against the baseline model, GenIE, using both human evaluations and standardized metrics.
The findings demonstrate the clear superiority of this synthetic generation framework. First, human evaluation revealed that the existing standard benchmark dataset is heavily flawed: approximately 70% of the information in its text is absent from its gold annotations, and 45% of its labeled facts are not actually expressed in the text. In contrast, the newly generated synthetic text accurately matched target facts with over 84% precision. Second, compact SynthIE models trained on the synthetic dataset dramatically outperformed the prior state of the art on high-quality test benchmarks, achieving a 57-point increase in micro-F1 and a 79-point increase in macro-F1 (reaching 93.0% F1). Third, while baseline models completely failed on the 46% least-frequent relation types, SynthIE maintained consistent, high-accuracy performance across all relations regardless of real-world frequency.
These results demonstrate that organizations can effectively bypass expensive, slow manual data annotation by exploiting task asymmetry with large generative models. This process enables small, computationally efficient models to achieve high-performance information extraction while drastically lowering serving costs, computational latency, and dependency on large proprietary APIs. It also provides a repeatable blueprint for generating clean, balanced training corpora across other structured language tasks.
Moving forward, practitioners and technical teams should adopt inverse synthetic generation pipelines for specialized data extraction tasks rather than relying on noisy heuristic collections. Organizations should also integrate the released synthetic datasets to benchmark existing internal extraction systems. For production deployments, teams must expand the framework to handle non-canonical text containing information that cannot be linked to existing knowledge bases, and implement validation checks to monitor and mitigate potential language model biases embedded during text generation.
- Paper: Unified Structure Generation for Universal Information Extraction, Yaojie Lu et al. (2022). UIE establishes text-to-structure generation as a unified approach to information extraction, making SynthIE’s structured-output task and its generative framing easier to follow.
- Paper: Generative Knowledge Graph Construction: A Review, Hongbin Ye et al. (2022). This review maps generative information-extraction methods and output formats, providing the IE landscape in which SynthIE’s closed-extraction approach is situated.
- Paper: Is GPT-3 a Good Data Annotator?, Bosheng Ding et al. (2023). Its comparison of LLM annotation and synthetic example generation for NER and relation extraction gives useful prior context for SynthIE’s strategy of creating training data for a smaller extractor.
No sufficiently relevant recommendations were found.
