Tailor: Generating and Perturbing Text with Semantic Controls

Alexis RossTongshuang WuHao PengMatthew E. PetersMatt Gardner

article2022ACL85 citations

Presents TAILOR, a semantically controlled text generation framework that perturbs sentences via composable semantic role representations, enabling automated contrast set creation across four NLP tasks and boosting model generalization on NLI challenge sets with minimal augmented data.

Listen

Controlled text perturbation—modifying text to exhibit specific attributes such as changes in verb tense, voice, or semantic roles—is vital for evaluating model robustness, diagnosing biases, and improving generalization. However, prevailing methods require training dedicated, task-specific models for each targeted transformation or relying on labor-intensive manual data annotation. This paradigm is computationally expensive, difficult to scale, and lacks flexibility when complex, multi-step linguistic transformations are required.

The article introduces Tailor, a semantically-controlled text generation system designed to perform fine-grained, application-agnostic, and compositional text perturbations using a single underlying model.

Tailor adapts a sequence-to-sequence neural network (specifically, fine-tuning T5-base on OntoNotes 5.0 data) conditioned on structured, human-readable control codes derived from Proposition Bank semantic structures. These codes specify core semantic roles (such as agent and patient), adjuncts (such as temporal and locative information), verb forms (tense and voice), and keyword content at varying levels of specificity. To force the generator to adhere strictly to these constraints, the authors implemented unlikelihood training, which penalizes outputs that deviate from target control codes. The system executes modular perturbation macros (such as swapping core arguments, modifying specificity, or altering verb forms) that can be combined to perform complex linguistic edits.

The article demonstrates several significant findings. First, Tailor delivers controllable and minimally invasive edits: it achieves a 64.3% closeness score and adheres to semantic role and content controls up to 81.6% of the time, substantially outperforming traditional maximum likelihood training. Second, Tailor successfully replicates contrast evaluation sets across four diverse language processing benchmarks, generating valid diagnostic instances with high accuracy (reaching 81% to 82% top-1 validity on question answering tasks) while cutting out spurious dataset artifacts and preserving lexical diversity. Third, in data augmentation experiments, adding Tailor-perturbed instances to just 2% of the training data yielded a 5.8-point overall gain on a difficult natural language inference challenge set, including a 29.2-point improvement on non-entailment cases, without degrading performance on original test sets.

These findings indicate that incorporating classical semantic frameworks with modern generative models enables scalable, automated stress-testing and data enrichment for natural language systems. Organizations can use this approach to lower the high manual annotation costs typically associated with auditing machine learning models and creating robust training pipelines. The ability to systematically modify specific linguistic components directly reduces the risk of models relying on superficial statistical heuristics.

Stakeholders and engineering teams should consider integrating modular, semantic perturbation tools into model validation pipelines rather than relying exclusively on manual red-teaming or single-purpose augmentation scripts. While the results demonstrate clear utility, practitioners should note limitations: Tailor relies on the accuracy of upstream semantic role labeling tools, was evaluated only on English text, and occasionally produces degenerate outputs that require straightforward heuristic or perplexity-based filtering.

arXiv: 2107.07150
Cover for Tailor: Generating and Perturbing Text with Semantic Controls

Abstract

Controlled text perturbation is useful for evaluating and improving model generalizability. However, current techniques rely on training a model for every target perturbation, which is expensive and hard to generalize. We present TAILOR, a semantically-controlled text generation system. TAILOR builds on a pretrained seq2seq model and produces textual outputs conditioned on control codes derived from semantic representations. We craft a set of operations to modify the control codes, which in turn steer generation towards targeted attributes. These operations can be further composed into higher-level ones, allowing for flexible perturbation strategies. We demonstrate the effectiveness of these perturbations in multiple applications. First, we use TAILOR to automatically create high-quality contrast sets for four distinct natural language processing (NLP) tasks. These contrast sets contain fewer spurious artifacts and are complementary to manually annotated ones in their lexical diversity. Second, we show that TAILOR perturbations can improve model generalization through data augmentation. Perturbing just ~2% of training data leads to a 5.8-point gain on an NLI challenge set measuring reliance on syntactic heuristics.

Table of Contents

  • 1 Introduction
  • 2 Tailor’s Controllable Generator
  • 2.1 Three Types of Controls
  • 2.2 Input Format Design
  • 2.3 Training
  • 3 Creating Perturbations with Tailor
  • 4 Intrinsic Evaluation
  • 4.1 Results
  • 5 Contrast Set Creation
  • 5.1 Replicating Contrast Sets with Tailor
  • 5.2 Measuring Contrast Set Quality
  • 5.3 Discussion
  • 6 Data Augmentation
  • 7 Related Work
  • 8 Discussion
  • Acknowledgements
  • References
  • Appendices
  • A Tailor Generator Details
  • A.1 Input and Output Formats
  • A.2 Training details
  • B Intrinsic Evaluation Details
  • C Degenerate Outputs
  • D Contrast Set Details (§5)
  • D.1 Perturbation Strategies
  • D.2 Predictor Performance Evaluation
  • E Data Augmentation Details (§6)
  • F Tailor’s fine-grained and compositional perturbations on StylePTB

Knowls

  1. Knowl 1 — TAILOR Semantic Control Representation and Input-Output Format

    model/method

    TAILOR is a semantically controlled sequence-to-sequence text generation and perturbation system based on fine-tuned T5-base. It conditions generation on structured semantic representations derived from Proposition Bank (PropBank) semantic role labeling (SRL).

    The input to the generator consists of two parts:

    1. Bracketed Header: An input-independent ordered list of abstract semantic control codes:
      • Predicate Controls: Formatted as control_code+voice+tense:lemma, specifying the verb lemma, voice (active or passive), and tense (past, present, or future) (e.g., [VERB+active+past: comfort]).
      • Argument Controls: Formatted as control_code+specificity:content, mapping PropBank roles to human-readable names (ARG0 →\to AGENT, ARG1 →\to PATIENT, and adjuncts such as LOCATIVE, TEMPORAL, MANNER, CAUSE, PURPOSE). Keyword specificity is categorized as complete (entire span), partial (all but up to 5 tokens), or sparse (* symbol, imposing no lexical constraints beyond the role).
    2. Context Sequence: Contains original text spans to preserve interspersed with mask tokens (<id_0>, <id_1>, etc.) denoting target generation locations. Extra empty blank tokens are randomly inserted between predicate subtrees to give the model freedom to place arguments fluently.

    The model generates full sentences containing role tags and bracket delimiters (e.g., [LOCATIVE: In the operating room], the doctor [VERB: comforted] the athlete.), which explicitly map input control codes to generated spans. These bracketed tags are stripped via regular expressions for downstream usage.

  2. Knowl 2 — Unlikelihood Training for Semantic Control Code Compliance

    model/method

    Standard Maximum Likelihood Estimation (MLE) allows models to ignore semantic control codes in favor of natural language priors (e.g., generating common collocations instead of swapping roles). To enforce strict control adherence, TAILOR combines standard MLE on positive examples with sequence-level unlikelihood training on negative examples.

    Training instances are constructed from gold PropBank annotations in OntoNotes 5.0 (yielding 223K positive instances and 541K negative instances). Negative examples (up to three per positive instance) are constructed by perturbing control codes while keeping target outputs fixed:

    1. Role and Predicate Swaps: Swapping AGENT and PATIENT labels, replacing adjunct role tags with uniformly sampled alternative adjuncts, or changing verb voice/tense.
    2. Keyword Content Corruption: Replacing verb lemmas and argument keywords with candidates sampled uniformly from the 15 most frequent keywords for that role and specificity in the dataset.
    3. Specificity Corruption: Uniformly selecting an incorrect keyword specificity tag.

    During training, unlikelihood loss penalizes generated sequences that conflict with the perturbed control headers, assigning an unlikelihood reward of −1-1 to negative samples and +1+1 to positive samples.

  3. Knowl 3 — Primitive Semantic and Syntactic Perturbation Macros in TAILOR

    model/method

    TAILOR provides a modular set of primitive perturbation operations that modify input headers and context blanks to execute targeted edits:

    1. Syntactically Controlled Rewriting:
      • CHANGE_VTENSE(target_tense): Updates the verb tense secondary control code (e.g., past →\to present).
      • CHANGE_VVOICE(target_voice): Updates the verb voice control code (e.g., active →\to passive).
      • CHANGE_IDX(source_idx:target_idx): Reorders blank position tokens in the context to shift constituent order.
      • CORE(SWAP_CORE): Swaps the keyword contents of core semantic arguments (e.g., exchanging AGENT and PATIENT).
    2. Sentence Expansion and Abstraction:
      • CHANGE_SPEC(target_specificity): Adjusts keyword specificity (e.g., complete →\to partial or sparse) to force argument deletion, elaboration, or abstraction.
      • DELETE: Removes an argument from the header and its tokens from the context sequence.
    3. Data Recombination:
      • CHANGE_CONTENT(new_text): Injects external context, phrases, or entities into an argument's keyword control.

    These primitive macros can be chained into higher-level strategies, such as altering prepositional phrase attachments or substituting passage entities in reading comprehension questions.

  4. Knowl 4 — Intrinsic Evaluation of TAILOR Controllability, Closeness, and Likelihood

    data/table

    Intrinsic generation quality was evaluated on 1,000 randomly perturbed sentences from the OntoNotes 5.0 validation set, comparing TAILOR with unlikelihood training against an MLE-only baseline (TAILORMLE\text{TAILOR}_{\text{MLE}}). Sentence likelihood was assessed via the ratio of GPT-2 language modeling loss on edited versus original text (TAILOR achieved 0.982, close to the ideal 1.0). Controllability was evaluated via cycle consistency with an external SRL predictor, and closeness was measured by length-weighted F1 on expected-to-change versus actually-changed spans (where a span is changed if ≥50%\ge 50\% of its tokens are altered).

    Closeness (%) Predicate Controllability (%) Argument Controllability (%)
    Generator F1 Precision Recall Lemma Tense Voice Role Content / Spec.
    TAILOR 64.3 66.5 73.4 74.3 80.3 81.6 70.5 64.5
    TAILORMLE\text{TAILOR}_{\text{MLE}} 58.5 59.5 68.6 72.2 70.2 76.1 60.3 45.1

    Unlikelihood training substantially improves controllability across all dimensions (e.g., an improvement of 19.4 percentage points in argument content/specificity control over pure MLE). Masking only the single targeted argument maximizes closeness F1 to 67.4%, whereas adding empty blanks modulates fluency by increasing the likelihood ratio from 0.93 to 0.95.

  5. Knowl 5 — Automated Contrast Set Replication Across NLP Benchmarks

    empirical result

    Using composed perturbation macros, TAILOR can automatically generate contrast sets across four distinct NLP tasks, reducing human annotation effort to post-hoc validation:

    1. Boolean Question Answering (BoolQ): Replicating entity changes by replacing question agent keywords with context paragraph entities sharing the same semantic role (AGENT:CHANGE_CONTENT), achieving 82% top-1 validity.
    2. Universal Dependencies (UD) Parsing: Replicating prepositional phrase (PP) attachment shifts between noun and verb attachments via CHANGE_CONTENT, CHANGE_SPEC(partial), and DELETE, achieving 65% top-10 validity (82% for noun →\to verb; 48% for verb →\to noun).
    3. MATRES Temporal Relation Extraction: Changing temporal ordering of events via verb tense manipulation (CHANGE_VFORM) and syntactic movement (PATIENT:MOVE), achieving 71% top-1 validity.
    4. SQuAD Question Implication: Converting question targets to alternate entities via WH-word mapping, achieving 81% top-1 validity.

    Evaluation of downstream models on TAILOR contrast sets showed performance drops consistent with human-authored contrast sets:

    • BoolQ (T5-base): 82.8% on original test →\to 64.8% on human contrast (-17.5) vs. 64.7% on TAILOR contrast (-17.6).
    • SQuAD (RoBERTa): 91.8% on original test →\to 66.1% on human contrast (-25.7) vs. 55.3% on TAILOR contrast (-36.5).
    • MATRES (T5-base): 70.3% on original test →\to 49.4% on human contrast (-20.9) vs. 42.3% on TAILOR contrast (-28.0).
  6. Knowl 6 — Lexical Diversity and Artifact Mitigation in TAILOR-Generated Contrast Sets

    empirical result

    Analysis of contrast sets created by TAILOR demonstrates high lexical diversity and mitigation of dataset-level spurious correlations:

    • Lexical Diversity: On the UD Parsing contrast benchmark (100 sampled instances), lexical diversity—measured as the ratio of unique to total new tokens in modified prepositional phrases—was 0.78 for TAILOR versus 0.99 for human annotators on noun →\to verb shifts, and 1.0 for both on verb →\to noun shifts. Unique tokens produced by TAILOR overlapped with human-authored tokens by <15%<15\% for verb →\to noun and ∼6%\sim 6\% for noun →\to verb.
    • Dataset Artifact Mitigation: When subjected to statistical hypothesis testing evaluating token-level spurious correlation with task labels (z=±2z = \pm 2), the original BoolQ validation set contains numerous tokens with statistically significant correlation to positive labels. In contrast, instances perturbed with TAILOR show token-label conditional distributions that fall almost entirely within the null-correlation confidence boundary, demonstrating that TAILOR generates evaluation data with fewer spurious artifacts.
  7. Knowl 7 — Data Augmentation with Core Argument Swapping for NLI Robustness

    data/table

    To reduce model reliance on syntactic heuristics in Natural Language Inference (NLI)—where high lexical overlap falsely triggers entailment predictions—TAILOR was used to augment the Stanford Natural Language Inference (SNLI) training corpus. Hypotheses were perturbed using the SWAP_CORE operation (swapping AGENT and PATIENT), using the original hypothesis as the premise and the perturbed hypothesis as a non-entailed hypothesis.

    RoBERTa-base classifiers were trained on SNLI (549,367 train instances), SNLI augmented with a rule-based syntactic perturbation baseline (10,987 instances, ∼2%\sim 2\%), and SNLI augmented with TAILOR (10,987 instances, ∼2%\sim 2\%), averaged over 20 random seeds:

    HANS Accuracy (%)
    Training Data SNLI Test Acc. (%) All Entailment Non-entailment
    SNLI Train 91.1 64.7 99.0 30.5
    + Syntactic Perturbation Baseline 91.0 67.5 95.8 39.2
    + TAILOR Perturbation 91.1 70.5 81.3 59.7

    Augmenting with TAILOR yielded a 5.81-point overall accuracy gain and a 29.2-point gain on the non-entailment subset of the HANS challenge set (t=−6.42,p<10−3t = -6.42, p < 10^{-3} via Student's tt-test) without degrading in-domain SNLI test accuracy (91.1%). TAILOR outperforms the dedicated syntactic perturbation baseline because PropBank SWAP_CORE generalizes across diverse syntactic configurations beyond rigid template structures.

  8. Knowl 8 — Zero-Shot Compositional Style Transfer on StylePTB

    data/table

    TAILOR was evaluated zero-shot (without transfer-specific fine-tuning or paired transfer examples) on the StylePTB fine-grained stylistic transfer benchmark. Using greedy beam search decoding (beam width 10, no repeated bigrams) with candidate selection via lowest GPT-2 perplexity, TAILOR was evaluated on single and multi-attribute compositional transfers against task-specific fine-tuned baselines using BLEU-1:

    Transfer Task Fine-Tuned Baseline (CS-GPT*) Multi-Single Baseline (CS-Sys-Gen*) TAILOR (Zero-Shot)
    Tense + Voice Composition
    ToPast + ActiveToPassive 40.9 33.7 66.0
    ToFuture + ActiveToPassive 49.6 41.9 46.8
    ToFuture + PassiveToActive 52.8 39.9 68.3
    ToPast + PassiveToActive 47.4 36.5 70.2
    ToPresent + PassiveToActive 52.3 42.4 69.9
    ToPresent + ActiveToPassive 50.3 44.5 31.5
    Tense + PP Removal Composition
    ToFuture + PPRemoval 73.8 46.5 74.3
    ToPast + PPRemoval 77.2 54.2 73.8
    ToPresent + PPRemoval 70.9 54.5 69.1

    A single general TAILOR model outperformed the multi-single transfer baseline (CS-Sys-Gen*) on 8 out of 9 compositional tasks and exceeded the compositional fine-tuned models (CS-GPT*) on 5 out of 9 tasks without any transfer-specific training.

  9. Knowl 9 — Degenerate Output Generation and Perplexity Filtering in TAILOR

    limitation

    A limitation of training TAILOR with sequence-level unlikelihood training is the occasional emergence of degenerate text generations: the generator learns to decrease the probability of negative training sequences by outputting highly unnatural or repetitive non-English tokens (such as repeating the Romanian token strings sanatate or pastra).

    These degenerate generations exhibit substantially lower GPT-2 log-perplexities than well-formed outputs: across 300 randomly sampled validation inputs, degenerate outputs (12 out of 300) had a mean perplexity of −346.46-346.46, compared to −86.747-86.747 for non-degenerate outputs.

    In practical downstream pipelines, TAILOR mitigates this behavior by combining heuristic string-matching filters (e.g., detecting sanatate) with GPT-2 perplexity threshold cutoffs, retaining the top ∼75%\sim 75\% of valid generations.

Coverage note — All primary contributions—including the TAILOR model design, unlikelihood training scheme, perturbation operations, intrinsic evaluation, contrast set creation across four tasks, lexical diversity/artifact analyses, NLI data augmentation on HANS, StylePTB style transfer, and degenerate generation filtering—are covered. Single-transfer style transfer rows from Table 11a and individual sentence walk-through examples were omitted as standalone knowls in favor of the summarized tables and general operational definitions.

References

  1. 1.Ekin Akyürek, Afra Feyza Akyürek, and Jacob Andreas. 2021. Learning to recombine and resample data for compositional generalization. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  2. 2.Jacob Andreas. 2020. Good-enough compositional data augmentation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7556–7566, Online. Association for Computational Linguistics.
  3. 3.Laura Banarescu, Claire Bonial, Shu Cai, Madalina Georgescu, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Martha Palmer, and Nathan Schneider. 2013. Abstract Meaning Representation for sembanking. In Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse, pages 178–186, Sofia, Bulgaria. Association for Computational Linguistics.
  4. 4.Yu Bao, Hao Zhou, Shujian Huang, Lei Li, Lili Mou, Olga Vechtomova, Xin-yu Dai, and Jiajun Chen. 2019. Generating sentences from disentangled syntactic and semantic spaces. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6008–6019, Florence, Italy. Association for Computational Linguistics.
  5. 5.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  6. 6.Mingda Chen, Qingming Tang, Sam Wiseman, and Kevin Gimpel. 2019. Controllable paraphrase generation with a syntactic exemplar. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5972–5984, Florence, Italy. Association for Computational Linguistics.
  7. 7.Shizhe Chen, Qin Jin, Peng Wang, and Qi Wu. 2020. Say as you wish: Fine-grained control of image caption generation with abstract scene graphs. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 9959–9968. IEEE.
  8. 8.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2924–2936, Minneapolis, Minnesota. Association for Computational Linguistics.
  9. 9.Marco Damonte and Shay B. Cohen. 2019. Structural neural encoders for AMR-to-text generation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3649–3658, Minneapolis, Minnesota. Association for Computational Linguistics.
  10. 10.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020. Plug and play language models: A simple approach to controlled text generation. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  11. 11.Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020. Evaluating models’ local decision boundaries via contrast sets. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1307–1323, Online. Association for Computational Linguistics.
  12. 12.Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018. AllenNLP: A deep semantic natural language processing platform. In Proceedings of Workshop for NLP Open Source Software (NLP-OSS), pages 1–6, Melbourne, Australia. Association for Computational Linguistics.
  13. 13.Matt Gardner, William Merrill, Jesse Dodge, Matthew Peters, Alexis Ross, Sameer Singh, and Noah A. Smith. 2021. Competency problems: On finding and removing artifacts in language data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1801–1813, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  14. 14.Jan Hajič, Massimiliano Ciaramita, Richard Johansson, Daisuke Kawahara, Maria Antònia Martí, Lluís Màrquez, Adam Meyers, Joakim Nivre, Sebastian Padó, Jan Štěpánek, Pavel Stranák, Mihai Surdeanu, Nianwen Xue, and Yi Zhang. 2009. The CoNLL-2009 shared task: Syntactic and semantic dependencies in multiple languages. In Proceedings of the Thirteenth Conference on Computational Natural Language Learning (CoNLL 2009): Shared Task, pages 1–18, Boulder, Colorado. Association for Computational Linguistics.
  15. 15.Chris Hokamp and Qun Liu. 2017. Lexically constrained decoding for sequence generation using grid beam search. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1535–1546, Vancouver, Canada. Association for Computational Linguistics.
  16. 16.Zhiting Hu, Zichao Yang, Xiaodan Liang, Ruslan Salakhutdinov, and Eric P. Xing. 2017. Toward controlled generation of text. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pages 1587–1596. PMLR.
  17. 17.Kuan-Hao Huang and Kai-Wei Chang. 2021. Generating syntactically controlled paraphrases without using annotated parallel pairs. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1022–1033, Online. Association for Computational Linguistics.
  18. 18.Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018. Adversarial example generation with syntactically controlled paraphrase networks. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1875–1885, New Orleans, Louisiana. Association for Computational Linguistics.
  19. 19.Divyansh Kaushik, Eduard H. Hovy, and Zachary Chase Lipton. 2020. Learning the difference that makes A difference with counterfactually-augmented data. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  20. 20.N. Keskar, Bryan McCann, L. Varshney, Caiming Xiong, and R. Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. ArXiv, abs/1909.05858.
  21. 21.Ashutosh Kumar, Kabir Ahuja, Raghuram Vadapalli, and Partha Talukdar. 2020. Syntax-guided controlled generation of paraphrases. Transactions of the Association for Computational Linguistics, 8:329–345.
  22. 22.Kenton Lee, Kelvin Guu, Luheng He, Timothy Dozat, and Hyung Won Chung. 2021. Neural data augmentation via example extrapolation. ArXiv, abs/2102.01335.
  23. 23.Chuanrong Li, Lin Shengshuo, Zeyu Liu, Xinyi Wu, Xuhui Zhou, and Shane Steinert-Threlkeld. 2020. Linguistically-informed transformations (LIT): A method for automatically generating contrast sets. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 126–135, Online. Association for Computational Linguistics.
  24. 24.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  25. 25.Yiwei Lyu, Paul Pu Liang, Hai Pham, Eduard Hovy, Barnabás Póczos, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2021. StylePTB: A compositional benchmark for fine-grained controllable text style transfer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2116–2138, Online. Association for Computational Linguistics.
  26. 26.Bill MacCartney and Christopher D Manning. 2014. Natural logic and natural language inference. In Computing meaning, pages 129–147. Springer.
  27. 27.Aman Madaan, Amrith Setlur, Tanmay Parekh, Barnabas Poczos, Graham Neubig, Yiming Yang, Ruslan Salakhutdinov, Alan W Black, and Shrimai Prabhumoye. 2020a. Politeness transfer: A tag and generate approach. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1869–1881, Online. Association for Computational Linguistics.
  28. 28.Nishtha Madaan, Inkit Padhi, Naveen Panwar, and Diptikalyan Saha. 2020b. Generate your counterfactuals: Towards controlled counterfactual generation for text. ArXiv preprint, abs/2012.04698.
  29. 29.Manuel Mager, Ramón Fernandez Astudillo, Tahira Naseem, Md Arafat Sultan, Young-Suk Lee, Radu Florian, and Salim Roukos. 2020. GPT-too: A language-model-first approach for AMR-to-text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1846–1852, Online. Association for Computational Linguistics.
  30. 30.Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428–3448, Florence, Italy. Association for Computational Linguistics.
  31. 31.George A Miller. 1998. WordNet: An electronic lexical database. MIT press.
  32. 32.Junghyun Min, R. Thomas McCoy, Dipanjan Das, Emily Pitler, and Tal Linzen. 2020. Syntactic data augmentation increases robustness to inference heuristics. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2339–2352, Online. Association for Computational Linguistics.
  33. 33.Qiang Ning, Hao Wu, and Dan Roth. 2018. A multi-axis annotation scheme for event temporal relations. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1318–1328, Melbourne, Australia. Association for Computational Linguistics.
  34. 34.Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajič, Christopher D. Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman. 2016. Universal Dependencies v1: A multilingual treebank collection. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 1659–1666, Portorož, Slovenia. European Language Resources Association (ELRA).
  35. 35.Jiefu Ou, Nathaniel Weir, Anton Belyy, Felix Yu, and Benjamin Van Durme. 2021. InFillmore: Frameguided language generation with bidirectional context. In *Proceedings of SEM 2021: The Tenth Joint Conference on Lexical and Computational Semantics, pages 129–142, Online. Association for Computational Linguistics.
  36. 36.Martha Palmer, Daniel Gildea, and Paul Kingsbury. 2005. The Proposition Bank: An annotated corpus of semantic roles. Computational Linguistics, 31(1):71–106.
  37. 37.Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Jie Ma, Alessandro Achille, Rishita Anubhai, Cícero Nogueira dos Santos, Bing Xiang, and Stefano Soatto. 2021. Structured prediction as translation between augmented natural languages. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  38. 38.Hao Peng, Ankur Parikh, Manaal Faruqui, Bhuwan Dhingra, and Dipanjan Das. 2019. Text generation with exemplar-based adaptive decoding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2555–2565, Minneapolis, Minnesota. Association for Computational Linguistics.
  39. 39.Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Hwee Tou Ng, Anders Björkelund, Olga Uryupina, Yuchen Zhang, and Zhi Zhong. 2013. Towards robust linguistic analysis using OntoNotes. In Proceedings of the Seventeenth Conference on Computational Natural Language Learning, pages 143–152, Sofia, Bulgaria. Association for Computational Linguistics.
  40. 40.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  41. 41.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  42. 42.Machel Reid and Victor Zhong. 2021. LEWIS: Levenshtein editing for unsupervised text style transfer. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3932–3944, Online. Association for Computational Linguistics.
  43. 43.Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019. Are red roses red? evaluating consistency of question-answering models. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6174–6184, Florence, Italy. Association for Computational Linguistics.
  44. 44.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online. Association for Computational Linguistics.
  45. 45.Alexis Ross, Ana Marasović, and Matthew Peters. 2021. Explaining NLP models via minimal contrastive editing (MiCE). In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3840–3852, Online. Association for Computational Linguistics.
  46. 46.Benjamin Schiller, Johannes Daxenberger, and Iryna Gurevych. 2021. Aspect-controlled neural argument generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 380–396, Online. Association for Computational Linguistics.
  47. 47.Lei Sha, Patrick Hohenecker, and Thomas Lukasiewicz. 2021. Controlling text edition by changing answers of specific questions. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1288–1299, Online. Association for Computational Linguistics.
  48. 48.Shikhar Sharma, Layla El Asri, Hannes Schulz, and Jeremie Zumer. 2017. Relevance of unsupervised metrics in task-oriented dialogue for evaluating natural language generation. ArXiv preprint, abs/1706.09799.
  49. 49.Peng Shi and Jimmy Lin. 2019. Simple bert models for relation extraction and semantic role labeling. ArXiv preprint, abs/1904.05255.
  50. 50.Jiao Sun, Xuezhe Ma, and Nanyun Peng. 2021. AESOP: Paraphrase generation with adaptive syntactic control. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5176–5189, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  51. 51.Damien Teney, Ehsan Abbasnedjad, and Anton van den Hengel. 2020. Learning what makes a difference from counterfactual examples and gradient supervision. ArXiv preprint, abs/2004.09034.
  52. 52.Chantal van Son, Oana Inel, Roser Morante, Lora Aroyo, and Piek Vossen. 2018. Resource interoperability for sustainable benchmarking: The case of events. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  53. 53.Nathaniel Weir, João Sedoc, and Benjamin Van Durme. 2020. COD3S: Diverse generation with discrete semantic signatures. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5199–5211, Online. Association for Computational Linguistics.
  54. 54.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020. Neural text generation with unlikelihood training. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  55. 55.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  56. 56.Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2019. Errudite: Scalable, reproducible, and testable error analysis. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 747–763, Florence, Italy. Association for Computational Linguistics.
  57. 57.Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2021. Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6707–6723, Online. Association for Computational Linguistics.
  58. 58.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase adversaries from word scrambling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1298–1308, Minneapolis, Minnesota. Association for Computational Linguistics.

Citation

MLA
Ross, A., et al. “Tailor: Generating and Perturbing Text with Semantic Controls”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 3194–213, https://doi.org/10.18653/v1/2022.acl-long.228.
APA
Ross, A., Wu, T., Peng, H., Peters, M. E., & Gardner, M. (2022). Tailor: Generating and Perturbing Text with Semantic Controls. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3194–3213. https://doi.org/10.18653/v1/2022.acl-long.228
Chicago
Ross, A., T. Wu, H. Peng, M. E. Peters, and M. Gardner. 2022. “Tailor: Generating and Perturbing Text with Semantic Controls”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3194–3213. https://doi.org/10.18653/v1/2022.acl-long.228.
Harvard
Ross, A. et al. (2022) “Tailor: Generating and Perturbing Text with Semantic Controls”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3194–3213. Available at: https://doi.org/10.18653/v1/2022.acl-long.228.
Vancouver
1. Ross A, Wu T, Peng H, Peters ME, Gardner M (2022) Tailor: Generating and Perturbing Text with Semantic Controls. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3194–3213

BibTeX

@inproceedings{ross-etal-2022-tailor,
    title = "Tailor: Generating and Perturbing Text with Semantic Controls",
    author = "Ross, Alexis  and
      Wu, Tongshuang  and
      Peng, Hao  and
      Peters, Matthew  and
      Gardner, Matt",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.228/",
    doi = "10.18653/v1/2022.acl-long.228",
    pages = "3194--3213"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/