Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction

Martin JosifoskiMarija SakotaMaxime PeyrardRobert West

article2023EMNLP139 citations

Demonstrates that generating input texts from sampled knowledge graph triplets exploits task asymmetry to produce balanced synthetic training datasets, enabling compact models to outperform prior closed information extraction methods by huge margins.

Listen

Extracting structured knowledge from natural language text—known as closed information extraction—is vital for building and populating knowledge bases. However, creating large-scale training datasets for this task is expensive and time-consuming because human annotators must master extensive catalogs of entities and relation types. Consequently, existing public datasets rely heavily on automated heuristics, resulting in noisy labels and severe class imbalances where rare relations are largely ignored. Furthermore, standard large language models cannot directly solve closed information extraction out of the box because they lack specific knowledge of predefined catalog schemas and identifiers.

The article demonstrates a reverse synthetic data generation strategy that leverages task asymmetry to solve this data bottleneck. Specifically, while asking a language model to extract rigid, structured facts from unstructured text is difficult, prompting the model to generate natural, fluent text from a provided structured set of facts is straightforward. The researchers evaluate whether this inverse approach can produce high-quality, balanced synthetic datasets and whether training compact language models on such data yields superior extraction performance.

To implement this approach, the authors constructed a graph from Wikidata covering 2.7 million entities and 888 relations. They developed a random-walk sampling technique with aggressive reweighting to extract 1.8 million coherent, balanced fact sets. They then prompted large language models to generate fluent natural text expressing only the facts in each set, creating a massive synthetic dataset. Using this generated data, they fine-tuned compact models with 220 million and 770 million parameters, called SynthIE, and benchmarked them against the baseline model, GenIE, using both human evaluations and standardized metrics.

The findings demonstrate the clear superiority of this synthetic generation framework. First, human evaluation revealed that the existing standard benchmark dataset is heavily flawed: approximately 70% of the information in its text is absent from its gold annotations, and 45% of its labeled facts are not actually expressed in the text. In contrast, the newly generated synthetic text accurately matched target facts with over 84% precision. Second, compact SynthIE models trained on the synthetic dataset dramatically outperformed the prior state of the art on high-quality test benchmarks, achieving a 57-point increase in micro-F1 and a 79-point increase in macro-F1 (reaching 93.0% F1). Third, while baseline models completely failed on the 46% least-frequent relation types, SynthIE maintained consistent, high-accuracy performance across all relations regardless of real-world frequency.

These results demonstrate that organizations can effectively bypass expensive, slow manual data annotation by exploiting task asymmetry with large generative models. This process enables small, computationally efficient models to achieve high-performance information extraction while drastically lowering serving costs, computational latency, and dependency on large proprietary APIs. It also provides a repeatable blueprint for generating clean, balanced training corpora across other structured language tasks.

Moving forward, practitioners and technical teams should adopt inverse synthetic generation pipelines for specialized data extraction tasks rather than relying on noisy heuristic collections. Organizations should also integrate the released synthetic datasets to benchmark existing internal extraction systems. For production deployments, teams must expand the framework to handle non-canonical text containing information that cannot be linked to existing knowledge bases, and implement validation checks to monitor and mitigate potential language model biases embedded during text generation.

arXiv: 2303.04132epfl-dlab/SynthIE
  • Paper: Unified Structure Generation for Universal Information Extraction, Yaojie Lu et al. (2022). UIE establishes text-to-structure generation as a unified approach to information extraction, making SynthIE’s structured-output task and its generative framing easier to follow.
  • Paper: Generative Knowledge Graph Construction: A Review, Hongbin Ye et al. (2022). This review maps generative information-extraction methods and output formats, providing the IE landscape in which SynthIE’s closed-extraction approach is situated.
  • Paper: Is GPT-3 a Good Data Annotator?, Bosheng Ding et al. (2023). Its comparison of LLM annotation and synthetic example generation for NER and relation extraction gives useful prior context for SynthIE’s strategy of creating training data for a smaller extractor.

No sufficiently relevant recommendations were found.

Cover for Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction

Abstract

Large language models (LLMs) have great potential for synthetic data generation. This work shows that useful data can be synthetically generated even for tasks that cannot be solved directly by LLMs: for problems with structured outputs, it is possible to prompt an LLM to perform the task in the reverse direction, by generating plausible input text for a target output structure. Leveraging this asymmetry in task difficulty makes it possible to produce large-scale, high-quality data for complex tasks. We demonstrate the effectiveness of this approach on closed information extraction, where collecting ground-truth data is challenging, and no satisfactory dataset exists to date. We synthetically generate a dataset of 1.8M data points, establish its superior quality compared to existing datasets in a human evaluation, and use it to finetune small models (220M and 770M parameters), termed SynthIE, that outperform the prior state of the art (with equal model size) by a substantial margin of 57 absolute points in micro-F1 and 79 points in macro-F1. Code, data, and models are available at https://github.com/epfl-dlab/SynthIE.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 2.1 Synthetic Data Generation
  • 2.2 Closed Information Extraction
  • 2.3 Data as a Core Limitation
  • 3 Exploiting Asymmetry for Synthetic Data Generation
  • 3.1 Knowledge Graph Construction
  • 3.2 Sampling Triplet Sets
  • 3.3 Triplet-Set-to-Text Generation
  • 3.4 Distributional Properties of Data
  • 3.5 Human Evaluation
  • 4 Synthetic Data in Action
  • 4.1 Modeling and Inference
  • 4.2 Experimental Setup
  • 4.3 Results
  • 4.3.1 Human Evaluation on REBEL
  • 4.3.2 Performance Evaluation
  • 4.3.3 Performance by Relation Frequency
  • 5 Discussion
  • Limitations
  • Acknowledgements
  • References
  • A. LLMcIE failure cases
  • B. Synthetic Data Generation
  • B.1 Details about the Knowledge Graph
  • B.2 Triplet Sampling
  • B.3 Triplet Set to Text
  • B.4 Distributional Properties of the Data
  • B.5 SDG Evaluation
  • B.5.1 Computing Performance Metrics
  • B.5.2 Human Annotation Task
  • B.5.3 Quality of annotations
  • B.5.4 Precision Annotation Procedure
  • C. Datasets
  • D. Performance Metrics
  • E. Experiment Implementation Details
  • F. Output Linearization
  • G. SynthIE in Action
  • G.1 Human Evaluation of REBEL
  • G.1.1 Constructing REBEL Clean
  • G.1.2 Precision on REBEL Clean
  • G.1.3 Recall on REBEL Clean

Knowls

  1. Knowl 1 — Reverse-direction generation turns structured prediction into synthetic-data creation

    model/method

    For a task that maps an input text xx to a structured output yy, the paper generates training examples in the reverse direction: first sample a valid target structure yy, then prompt a large language model (LLM) to produce a plausible text xx that expresses all and only that structure. This uses an asymmetry in task difficulty: an LLM may be poor at producing the required structured output from text while still being able to verbalize a supplied structure. In closed information extraction (cIE), yy is a set of knowledge-base (KB) triplets (s,r,o)(s,r,o), where subject ss and object oo are KB entities and relation rr is a KB relation. Sampling yy directly also gives control over output coverage and balance, rather than inheriting the skew of collected text annotations.

  2. Knowl 2 — Coherent triplet sets are sampled with anchor bias and coverage reweighting

    model/method

    The synthetic-data pipeline samples triplet sets from a Wikidata graph to obtain sets that are both expressible in short text and broadly cover entities and relations. A set starts from an entity or a graph edge; it is expanded by either starting a new triplet or selecting an object adjacent to the current subject and adding the corresponding relation edge. Sampling favors entities already present in the set, producing recurring anchor entities that make the triplets more likely to co-occur coherently. Set size is drawn from a Poisson distribution with mean 3, and the anchor bias factor is 7. To reduce overrepresentation of graph hubs, entity and relation sampling distributions are reweighted inversely to their observed frequencies in previously sampled sets, with dampening factor d=0.01d=0.01. Every 20,000 samples, the distributions are recomputed and the starting-point strategy switches between entity-centric sampling and relation-centric sampling. In the latter, a relation is selected first, then a triplet with that relation is selected using the reweighted subject-entity distribution.

  3. Knowl 3 — Wikidata filtering and LLM verbalization produce the Wiki-cIE datasets

    experimental setup

    The generation graph contains 2,715,483 Wikidata entities, 888 relations, and 17,655,864 edges. It is restricted to entities and relations occurring in REBEL training data, excludes literal-argument relations and entities without an associated Wikipedia page, and uses Wikipedia page titles and Wikidata relation labels as textual identifiers. The pipeline prompts OpenAI models to turn each sampled triplet set into fluent text expressing exactly those triplets. The code-davinci-002 model is used for the approximately 1.8-million-example Wiki-cIE Code dataset; its splits contain 1,815,378 training examples, 10,000 validation examples, and 50,286 test examples before experiment-time filtering. The text-davinci-003 model generates Wiki-cIE Text validation and test sets of 10,000 and 50,286 examples, using the same triplet sets as Wiki-cIE Code. The experiments filter examples to those whose input and linearized output are each at most 256 tokens; the resulting Wiki-cIE Code training split has 1,669,708 examples. The best-performing setup uses few-shot prompting for code-davinci-002 and instruction prompting for text-davinci-003. Both use temperature 0.7, top-pp 1, frequency penalty 0.2, presence penalty 0, and stop sequence newline; maximum generation lengths are 100 and 50 tokens, respectively. At the time of the study, reported generation costs were 0forWiki−cIECodeand0 for Wiki-cIE Code and 223.55 for Wiki-cIE Text.

  4. Knowl 4 — Human evaluation finds synthetic text substantially more faithful than REBEL

    empirical result

    For a human evaluation, 50 REBEL test triplet sets were each paired with the original REBEL text and with text generated by the Wiki-cIE Code and Wiki-cIE Text procedures. Annotators checked which target triplets were expressed in each text, and the authors also estimated the total triplets expressed so precision could be measured. The reported micro precision, micro recall, micro F1, and macro recall (in that order) are: REBEL, 29.35 ± 7.77%, 56.05 ± 10.40%, 39.87 ± 7.62%, and 24.20 ± 6.20%; Wiki-cIE Code, 57.40 ± 10.28%, 70.38 ± 7.83%, 65.08 ± 7.35%, and 50.70 ± 9.10%; Wiki-cIE Text, 84.78 ± 5.80%, 78.45 ± 8.20%, 82.97 ± 5.53%, and 72.14 ± 8.73%. Thus, in the evaluated sample, more than 70% of the information in REBEL text was absent from its target set, while about 44% of its target triplets were not expressed in the text. Wiki-cIE Text was the most faithful of the three. Its precision estimate is an upper bound in the sense that annotators could have missed expressed triplets; macro precision was not reported because the exhaustive annotation relaxed the relation-catalog constraint.

  5. Knowl 5 — SynthIE represents triplet sets as constrained autoregressive sequences

    model/method

    SynthIE is a Flan-T5 sequence-to-sequence model that generates a linearized representation of the exhaustive KB triplets expressed in an input text. It is trained by maximizing the conditional log-likelihood of the target sequence with teacher forcing, using cross-entropy loss, dropout, and label smoothing. The fully expanded representation serializes every triplet separately, marking subject, relation, object, and triplet boundaries with delimiters. The subject-collapsed representation groups triplets by subject, writing that subject once and then listing its relation-object pairs; this shortens sequences but creates more complex dependencies within the output. At inference, constrained beam search enforces both the chosen output structure and identifier validity: a precomputed trie restricts entity and relation generation to catalog entries, while structural constraints restrict which element can follow each prefix. The experiments use an entity catalog of about 2.6 million tokenizable entities and 888 relations. Subject-collapsed outputs are shorter and incur no performance cost on the synthetic test sets, but reduce performance on REBEL Clean.

  6. Knowl 6 — Training and evaluation use fixed sequence limits and shared optimization settings

    experimental setup

    Models are trained for 8,000 steps with Adam, learning rate 3×10−43\times10^{-4}, gradient clipping at 0.1, weight decay 0.05, batch size 2,568, and a polynomial learning-rate schedule with 1,000 warm-up steps and final learning rate 3×10−53\times10^{-5}. The same tuned settings are used for GenIE models trained on REBEL and SynthIE models trained on synthetic data, isolating training data as the principal comparison. Inputs and outputs are limited to 256 tokens. Inference uses constrained beam search with 10 beams and length-normalized log probabilities; the tuned length penalty is 0.8 for fully expanded output and 0.6 for subject-collapsed output. Reported model metrics are micro- and macro-precision, recall, and F1, with 95% confidence intervals constructed from 50 bootstrap samples.

  7. Knowl 7 — SynthIE greatly outperforms an equal-sized REBEL-trained model on Wiki-cIE Text

    empirical result

    On Wiki-cIE Text, the T5-base SynthIE model (220M parameters) achieves micro precision/recall/F1 of 92.08 ± 0.17% / 90.75 ± 0.21% / 91.41 ± 0.18%, and macro precision/recall/F1 of 94.10 ± 0.15% / 92.42 ± 0.17% / 93.05 ± 0.11%. The equal-sized T5-base GenIE model trained on REBEL scores 49.10 ± 0.33% / 26.69 ± 0.17% / 34.58 ± 0.20% micro and 29.82 ± 0.67% / 11.14 ± 0.15% / 13.94 ± 0.17% macro. The absolute F1 gains for SynthIE T5-base are 56.83 points micro and 79.11 points macro. SynthIE T5-large reaches 93.04 ± 0.13% micro-F1 and 94.99 ± 0.12% macro-F1 on Wiki-cIE Text. On Wiki-cIE Code, SynthIE T5-base scores 74.93 ± 0.27% micro-F1 and 77.91 ± 0.42% macro-F1, while SynthIE T5-large scores 77.59 ± 0.24% and 81.95 ± 0.22%, respectively.

  8. Knowl 8 — Balanced synthetic coverage supports performance across relation frequencies

    empirical result

    Relation counts are much less skewed in the synthetic datasets than in REBEL. The minimum, first quartile, median, third quartile, and maximum relation occurrence counts are, respectively: REBEL, 1, 4, 34, 432, and 716,679; Wiki-cIE Code, 65, 934, 1,380, 3,629, and 479,250; Wiki-cIE Text, 4, 42, 62, 136, and 14,323. The rarest Wiki-cIE Code relation therefore occurs more often than REBEL's median relation. On Wiki-cIE Text, GenIE has near-zero F1 for the 46% of relations occurring fewer than 32 times in REBEL training and reaches about 58% only for the most frequent relations. SynthIE trained on Wiki-cIE Code instead achieves about 93% F1 across relation-frequency buckets. This supports the finding that balanced, consistently annotated training data improves performance on rare as well as frequent relations.

  9. Knowl 9 — REBEL Clean exposes errors in original labels but is a selected evaluation sample

    empirical result

    The authors manually curated REBEL Clean from 360 REBEL test examples, selected from a random pool using criteria favoring substantial extractable information and excluding examples whose central facts required literal arguments. For each example, annotators reviewed a candidate set formed from the original REBEL triplets plus predictions from SynthIE T5-large, filtering out triplets not expressed in the text. Against the resulting annotations, REBEL's original target triplets achieve 73.76 ± 2.20% micro-F1 and 43.76 ± 4.62% macro-F1; SynthIE T5-large reaches 51.04 ± 4.76% macro-F1, exceeding the original labels on that metric. This comparison indicates that REBEL's gold labels are not a reliable ceiling on annotation quality. However, REBEL Clean remains imbalanced and can contain text facts that cannot be resolved to the KB, so the paper treats Wiki-cIE Text as the more reliable evaluation set.

  10. Knowl 10 — Exhaustive extraction remains limited by KB coverage, model expressivity, and generator bias

    limitation

    SynthIE is trained for exhaustive cIE, but the task can be ill-posed when important text information cannot be linked to an entity or relation in the allowed KB, or when the output format cannot express that information. In those cases, a model trained to extract exhaustively may try to encode unsupported content and make errors; the paper suggests that data covering such cases or broader catalogs and output formats could mitigate the problem, but does not establish a solution. Synthetic examples may also inherit biases from the LLM that generates their text. The pipeline further relies on access to specific OpenAI models: code-davinci-002 was not publicly available through the standard API at the time, although access was possible through OpenAI's Researcher Access Program.

Coverage note — No substantial contributed material was omitted; appendix-only prompt examples, full annotation instructions, and hardware timings are excluded as supporting reproducibility detail.

References

  1. 1.Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, and Naama Zwerdling. 2020. Do not have enough data? deep learning to the rescue! In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7383–7390.
  2. 2.Gabor Angeli, Victor Zhong, Danqi Chen, Arun Tejasvi Chaganty, Jason Bolton, Melvin Jose Johnson Premkumar, Panupong Pasupat, Sonal Gupta, and Christopher D. Manning. 2015. Bootstrapped self training for knowledge base population. In Proceedings of the 2015 Text Analysis Conference.
  3. 3.Laura Banarescu, Claire Bonial, Shu Cai, Madalina Georgescu, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Martha Palmer, and Nathan Schneider. 2013. Abstract meaning representation for sembanking. In Proceedings of the 7th linguistic annotation workshop and interoperability with discourse, pages 178–186.
  4. 4.Arun Chaganty, Ashwin Paranjape, Percy Liang, and Christopher D. Manning. 2017. Importance sampling for unbiased on-demand evaluation of knowledge base population. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1038–1048, Copenhagen, Denmark. Association for Computational Linguistics.
  5. 5.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. CoRR, abs/2210.11416.
  6. 6.Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frederique Laforest, and Elena Simperl. 2018. T-REx: A large scale alignment of natural language with knowledge base triples. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  7. 7.Luis Galárraga, Geremy Heitz, Kevin Murphy, and Fabian M. Suchanek. 2014. Canonicalizing open knowledge bases. In Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management, CIKM 2014, Shanghai, China, November 3-7, 2014, pages 1679–1688. ACM.
  8. 8.Jiahui Gao, Renjie Pi, Yong Lin, Hang Xu, Jiacheng Ye, Zhiyong Wu, Xiaodan Liang, Zhenguo Li, and Lingpeng Kong. 2022. Zerogen+: Self-guided high-quality data generation in efficient zero-shot learning.
  9. 9.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  10. 10.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022. Unnatural instructions: Tuning language models with (almost) no human labor. arXiv preprint arXiv:2212.09689.
  11. 11.Pere-Lluís Huguet Cabot and Roberto Navigli. 2021. REBEL: Relation extraction by end-to-end language generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2370–2381, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  12. 12.Martin Josifoski, Nicola De Cao, Maxime Peyrard, Fabio Petroni, and Robert West. 2022. GenIE: Generative information extraction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4626–4643, Seattle, United States. Association for Computational Linguistics.
  13. 13.Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020. Data augmentation using pre-trained transformer models. arXiv preprint arXiv:2003.02245.
  14. 14.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  15. 15.Junlong Li, Zhuosheng Zhang, and Hai Zhao. 2022. Self-prompting large language models for open-domain qa.
  16. 16.Yu Meng, Jiaxin Huang, Yu Zhang, and Jiawei Han. 2022a. Generating training data with language models: Towards zero-shot language understanding.
  17. 17.Yu Meng, Jiaxin Huang, Yu Zhang, and Jiawei Han. 2022b. Generating training data with language models: Towards zero-shot language understanding. arXiv preprint arXiv:2202.04538.
  18. 18.Yu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang, Tarek Abdelzaher, and Jiawei Han. 2022c. Tuning language models as training data generators for augmentation-enhanced few-shot learning.
  19. 19.Biswesh Mohapatra, Gaurav Pandey, Danish Contractor, and Sachindra Joshi. 2020. Simulated chats for building dialog systems: Learning to generate conversations from instructions. arXiv preprint arXiv:2010.10216.
  20. 20.Yannis Papanikolaou and Andrea Pierleoni. 2020. Dare: Data augmented relation extraction with gpt-2. arXiv preprint arXiv:2004.13845.
  21. 21.Gaurav Sahu, Pau Rodriguez, Issam H. Laradji, Parmida Atighehchian, David Vazquez, and Dzmitry Bahdanau. 2022. Data augmentation for intent classification with off-the-shelf large language models.
  22. 22.Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, and Sebastian Riedel. 2022. Peer: A collaborative language model. arXiv preprint arXiv:2208.11663.
  23. 23.Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Synthetic prompting: Generating chain-of-thought demonstrations for large language models.
  24. 24.Ryan Smith, Jason A. Fries, Braden Hancock, and Stephen H. Bach. 2022. Language models in the loop: Incorporating prompting into weak supervision.
  25. 25.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(56):1929–1958.
  26. 26.Ilya Sutskever, James Martens, and Geoffrey E. Hinton. 2011. Generating text with recurrent neural networks. In Proceedings of the 28th International Conference on Machine Learning, ICML 2011, Bellevue, Washington, USA, June 28 - July 2, 2011, pages 1017–1024. Omnipress.
  27. 27.Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada, pages 3104–3112.
  28. 28.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 2818–2826. IEEE Computer Society.
  29. 29.Denny Vrandevčić. 2012. Wikidata: A new platform for collaborative data collection. In Proceedings of the 21st international conference on world wide web, pages 1063–1064.
  30. 30.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461.
  31. 31.Zirui Wang, Adams Wei Yu, Orhan Firat, and Yuan Cao. 2021. Towards zero-label language learning. CoRR, abs/2109.09193.
  32. 32.Yiben Yang, Chaitanya Malaviya, Jared Fernandez, Swabha Swayamdipta, Ronan Le Bras, Ji-Ping Wang, Chandra Bhagavatula, Yejin Choi, and Doug Downey. 2020. Generative data augmentation for commonsense reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1008–1025, Online. Association for Computational Linguistics.
  33. 33.Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong. 2022. Zerogen: Efficient zero-shot learning via dataset generation.

Citation

MLA
Josifoski, M., et al. “Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1555–74, https://doi.org/10.18653/v1/2023.emnlp-main.96.
APA
Josifoski, M., Sakota, M., Peyrard, M., & West, R. (2023). Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1555–1574. https://doi.org/10.18653/v1/2023.emnlp-main.96
Chicago
Josifoski, M., M. Sakota, M. Peyrard, and R. West. 2023. “Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1555–74. https://doi.org/10.18653/v1/2023.emnlp-main.96.
Harvard
Josifoski, M. et al. (2023) “Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1555–1574. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.96.
Vancouver
1. Josifoski M, Sakota M, Peyrard M, West R (2023) Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1555–1574

BibTeX

@inproceedings{josifoski-etal-2023-exploiting,
    title = "Exploiting Asymmetry for Synthetic Training Data Generation: {S}ynth{IE} and the Case of Information Extraction",
    author = "Josifoski, Martin  and
      Sakota, Marija  and
      Peyrard, Maxime  and
      West, Robert",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.96/",
    doi = "10.18653/v1/2023.emnlp-main.96",
    pages = "1555--1574"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/