Tailor: Generating and Perturbing Text with Semantic Controls
Alexis RossTongshuang WuHao PengMatthew E. PetersMatt Gardner
Presents TAILOR, a semantically controlled text generation framework that perturbs sentences via composable semantic role representations, enabling automated contrast set creation across four NLP tasks and boosting model generalization on NLI challenge sets with minimal augmented data.
Controlled text perturbation—modifying text to exhibit specific attributes such as changes in verb tense, voice, or semantic roles—is vital for evaluating model robustness, diagnosing biases, and improving generalization. However, prevailing methods require training dedicated, task-specific models for each targeted transformation or relying on labor-intensive manual data annotation. This paradigm is computationally expensive, difficult to scale, and lacks flexibility when complex, multi-step linguistic transformations are required.
The article introduces Tailor, a semantically-controlled text generation system designed to perform fine-grained, application-agnostic, and compositional text perturbations using a single underlying model.
Tailor adapts a sequence-to-sequence neural network (specifically, fine-tuning T5-base on OntoNotes 5.0 data) conditioned on structured, human-readable control codes derived from Proposition Bank semantic structures. These codes specify core semantic roles (such as agent and patient), adjuncts (such as temporal and locative information), verb forms (tense and voice), and keyword content at varying levels of specificity. To force the generator to adhere strictly to these constraints, the authors implemented unlikelihood training, which penalizes outputs that deviate from target control codes. The system executes modular perturbation macros (such as swapping core arguments, modifying specificity, or altering verb forms) that can be combined to perform complex linguistic edits.
The article demonstrates several significant findings. First, Tailor delivers controllable and minimally invasive edits: it achieves a 64.3% closeness score and adheres to semantic role and content controls up to 81.6% of the time, substantially outperforming traditional maximum likelihood training. Second, Tailor successfully replicates contrast evaluation sets across four diverse language processing benchmarks, generating valid diagnostic instances with high accuracy (reaching 81% to 82% top-1 validity on question answering tasks) while cutting out spurious dataset artifacts and preserving lexical diversity. Third, in data augmentation experiments, adding Tailor-perturbed instances to just 2% of the training data yielded a 5.8-point overall gain on a difficult natural language inference challenge set, including a 29.2-point improvement on non-entailment cases, without degrading performance on original test sets.
These findings indicate that incorporating classical semantic frameworks with modern generative models enables scalable, automated stress-testing and data enrichment for natural language systems. Organizations can use this approach to lower the high manual annotation costs typically associated with auditing machine learning models and creating robust training pipelines. The ability to systematically modify specific linguistic components directly reduces the risk of models relying on superficial statistical heuristics.
Stakeholders and engineering teams should consider integrating modular, semantic perturbation tools into model validation pipelines rather than relying exclusively on manual red-teaming or single-purpose augmentation scripts. While the results demonstrate clear utility, practitioners should note limitations: Tailor relies on the accuracy of upstream semantic role labeling tools, was evaluated only on English text, and occasionally produces degenerate outputs that require straightforward heuristic or perplexity-based filtering.
- Paper: CTRL: A Conditional Transformer Language Model for Controllable Generation, Nitish Shirish Keskar et al. (2019). Introduces the foundational paradigm of conditioning transformer generation on explicit control codes, which Tailor adapts into modular semantic representations.
- Paper: Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment, Di Jin et al. (2019). Establishes standard methods for evaluating NLP model robustness via targeted semantic text perturbations, motivating Tailor's automated contrast-set generation.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). Provides key background on evaluating generation fidelity and semantic drift under synthetic variations and transformations.
- Paper: Composable Text Controls in Latent Space with ODEs, Guangyi Liu et al. (2023). Extends composable attribute control from explicit discrete representations to continuous latent space manipulation with ordinary differential equations.
- Paper: Controlled Text Generation with Natural Language Instructions, Wangchunshu Zhou et al. (2023). Generalizes fine-grained controlled text generation by expressing structural and linguistic constraints as natural language instructions rather than dedicated control codes.
- Paper: SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control, Xiaochuang Han et al. (2023). Applies modular, plug-and-play attribute control to semi-autoregressive diffusion language models.
- Paper: Diffusion-LM Improves Controllable Text Generation, Xiang Lisa Li et al. (2022). Explores classifier-guided diffusion in discrete text spaces as an alternative architecture for fine-grained semantic and syntactic control.
- Paper: Fine-Grained Controllable Text Generation Using Non-Residual Prompting, Fredrik Carlsson et al. (2022). Develops non-residual prompting to achieve fine-grained, dynamic steering of text generation across arbitrary positions without full retraining.
