Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Linearization schema

A linearization schema is a formal specification or set of structural rules used to convert multi-dimensional structured data, such as knowledge graphs, relational triples, trees, or tables, into a flat, sequential format readable by sequence-based models. Because autoregressive architectures and sequence-to-sequence language models process information sequentially as strings of tokens, a linearization schema defines the precise ordering, syntax, and delimiter tokens needed to encode entities, attributes, and relational dependencies. By establishing consistent patterns, such as grouping facts around common entities or serializing relational expressions, the schema enables models to consume, generate, and accurately decode complex structured relationships back into their original graph or tabular format without loss of semantic integrity.

1 item

Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction

Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction

Martin Josifoski, Marija Sakota, Maxime Peyrard, Robert West

OrganizationsÉcole Polytechnique Fédérale de Lausanne

Why you should read this

Demonstrates that generating input texts from sampled knowledge graph triplets exploits task asymmetry to produce balanced synthetic training datasets, enabling compact models to outperform prior closed information extraction methods by huge margins.

Large language models (LLMs) have great potential for synthetic data generation. This work shows that useful data can be synthetically generated even for tasks that cannot be solved directly by LLMs: for problems with structured outputs, it is possible to prompt an LLM to perform the task in the reverse direction, by generating plausible input text for a target output structure. Leveraging this asymmetry in task difficulty makes it possible to produce large-scale, high-quality data for complex tasks. We demonstrate the effectiveness of this approach on closed information extraction, where collecting ground-truth data is challenging, and no satisfactory dataset exists to date. We synthetically generate a dataset of 1.8M data points, establish its superior quality compared to existing datasets in a human evaluation, and use it to finetune small models (220M and 770M parameters), termed SynthIE, that outperform the prior state of the art (with equal model size) by a substantial margin of 57 absolute points in micro-F1 and 79 points in macro-F1. Code, data, and models are available at https://github.com/epfl-dlab/SynthIE.

Added

2026-10-01