Tailor: A Soft-Prompt-Based Approach to Attribute-Based Controlled Text Generation
Kexin YangDayiheng LiuWenqiang LeiBaosong YangMingfeng XueBoxing ChenJun Xie
Proposes a parameter-efficient framework that controls language models for single- and multi-attribute text generation by learning lightweight continuous prompts and prompt connectors while adding only 0.08% extra parameters to a frozen GPT-2.
Controlling the output of large language models to satisfy specific attributes, such as tone, emotion, or topic, is crucial for modern automated writing systems. However, current techniques present significant operational trade-offs. Fully fine-tuning a base model requires extensive computing and storage costs because a separate copy of the model must be maintained for every attribute. Conversely, guiding models at run-time with external attribute classifiers slows down text generation and frequently damages output fluency.
The article introduces Tailor, a lightweight continuous soft-prompt framework designed to steer a frozen GPT-2 model across single-attribute and multi-attribute generation tasks. The primary objective is to demonstrate that parameter-efficient soft prompts—continuous task-specific input vectors—can steer generation accurately while retaining fluency, boosting inference speed, and enabling multi-attribute control without fine-tuning full models.
The researchers evaluated Tailor through empirical experiments across single-attribute and multi-attribute generation benchmarks using Yelp restaurant reviews and cross-domain datasets (SST-2 and AG News). In Tailor, each attribute is trained as a single continuous vector prefix while keeping the underlying base model parameters fixed. For multi-attribute tasks, the authors tested a non-training combination approach—utilizing an attention mask (MAP mask) and a re-indexing position sequence (RP sequence) to avoid cross-attention and order sensitivity—alongside a training-based Multi-Attribute Prompt connector (MAP connector) trained using pseudo-labeled attribute data. The evaluation benchmarked correctness, fluency, perplexity, text diversity, inference speed, and human quality ratings against standard fine-tuning, adapter modules, and classifier-guided methods like GeDi and PPLM.
The evaluation yielded several key findings. First, Tailor achieves highly competitive attribute control using only 0.08% of the trainable parameters needed for full fine-tuning. In multi-attribute generation, Tailor's connector approach achieved an average correctness score of 87.15%, significantly outperforming full fine-tuning at 69.80% and adapters at 69.10%. Second, Tailor operates dramatically faster during generation than classifier-guided alternatives, processing a sample in 0.758 seconds compared to 1.680 seconds for GeDi and 15.553 seconds for PPLM (roughly a 20-fold speedup over PPLM). Third, the non-training approach effectively eliminated position sensitivity during prompt concatenation, boosting multi-attribute correctness from 76.20% to 78.82%. Finally, Tailor demonstrated strong generalization in low-data regimes and on previously unseen attribute pairings, such as positive sentiment combined with Mexican food topics.
These findings indicate that organizations can implement highly customizable, multi-attribute text generation without incurring high infrastructure, storage, or computational costs. By freezing the underlying base model and training only modular soft prompts, engineering teams reduce retraining overhead and runtime latency while preserving text quality and fluency.
Decision-makers and practitioners deploying controllable language models should adopt modular soft-prompt strategies like Tailor for multi-attribute generation workflows. If zero extra training is preferred, the non-training mask and position re-indexing approach serves as an effective plug-and-play solution. Where maximum attribute fidelity is required, investing minimal compute to train a prompt connector provides the best performance trade-off.
The primary limitation of the article is its focus on two-attribute combinations and evaluations centered mainly on the GPT-2 base architecture. While confidence in the reported performance metrics is high given the combined automatic and human evaluations, further testing across diverse model architectures and scaling to combinations of three or more simultaneous attributes are necessary before deploying in broader production settings.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). Read this foundational account of learning continuous soft prompts while freezing the base model to understand the parameter-efficient prompting strategy Tailor adapts for attribute control.
- Paper: P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks, Xiao Liu et al. (2022). Its P-Tuning v2 approach establishes how continuous prompts can be applied across model layers, providing useful context for Tailor’s prompt-based adaptation of a frozen language model.
No sufficiently relevant recommendations were found.
