A Recipe for Arbitrary Text Style Transfer with Large Language Models
Emily ReifDaphne IppolitoAnn YuanAndy CoenenChris Callison-BurchJason Wei
Proposes an augmented zero-shot prompting technique that enables large language models to perform arbitrary text style transfer using only natural language instructions without fine-tuning or target-style exemplars.
Text style transfer modifies the tone, mood, or stylistic elements of a passage while retaining its core meaning. Traditional approaches depend heavily on large labeled datasets of parallel text or specialized exemplars for every target style, which makes adapting to arbitrary or non-standard styles difficult and resource-intensive. As demand grows for flexible AI writing assistants, there is a clear operational need for flexible, training-free text transformation methods.
The article evaluates a prompting technique called augmented zero-shot learning. The main objective is to demonstrate that large language models can perform text style transfer across arbitrary, user-specified styles using natural language instructions without fine-tuning or style-specific examples.
The researchers framed style transfer as a generalized sentence rewriting task. Instead of using examples of the specific target style, they primed large language models (primarily LaMDA with 137 billion parameters, alongside GPT-3 variants) with a single multi-task prompt containing diverse sentence rewrites wrapped in consistent formatting delimiters. They evaluated performance across six non-standard styles (such as adding metaphors or adjusting melodrama) and standard style benchmarks (sentiment and formality). Assessment relied on 3,600 human evaluation ratings covering transfer strength, semantic preservation, and fluency, supported by automatic metric evaluations and a user study involving 30 creative writers.
The findings show that augmented zero-shot prompting performs comparably to human-written text and supervised baselines across standard and atypical styles. Compared to standard zero-shot prompting, which failed to return a valid response 25.4% of the time, the augmented technique drastically reduced unparseable outputs to 0.6%. On sentiment benchmarks, the method achieved 90.6% accuracy, approaching the 94.3% accuracy of five-shot prompting while significantly improving text fluency. However, the model produced lower overlap scores (BLEU) against target human references because it frequently elaborated on text rather than making literal word substitutions.
These results demonstrate that organizations can implement flexible text-editing and style-transfer capabilities without the high costs, timelines, and technical overhead of curating specialized training datasets or fine-tuning models. However, the method offers less fine-grained constraint than models trained on fixed data, as large language models may hallucinate information or lean naturally toward formal or melodramatic phrasing.
Decision-makers should consider augmented zero-shot prompting for open-ended creative tools and interactive writing aids where user flexibility is paramount. When precise lexical preservation is required, teams should implement candidate selection filtering or specialized supervised pipelines. Further testing is recommended to explore broader safety guardrails and systematic evaluations of model behavior across different base architectures.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). Establishes that zero-shot natural language prompt framing can steer large language models effectively without relying on runtime few-shot demonstrations.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Introduces the paradigm of prompting large pretrained language models to perform downstream text tasks without weight updates.
- Paper: Multitask Prompted Training Enables Zero-Shot Task Generalization, Victor Sanh et al. (2021). Demonstrates how formatting diverse tasks as natural language prompts enables robust zero-shot generalization across unseen text transformations.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). Provides a comprehensive taxonomy and theoretical foundation for prompt-based learning and natural language instruction design.
- Paper: CTRL: A Conditional Transformer Language Model for Controllable Generation, Nitish Shirish Keskar et al. (2019). Lays foundational groundwork for controllable text generation and style steering using conditioning signals.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Automates the generation and optimization of natural language prompt instructions using language models themselves, scaling beyond manual prompt design.
- Paper: Guiding Large Language Models via Directional Stimulus Prompting, Zekun Li et al. (2023). Extends prompt-guided text rewriting by using a small policy model to generate instance-specific directional stimulus prompts for black-box language models.
- Paper: Composable Text Controls in Latent Space with ODEs, Guangyi Liu et al. (2023). Advances multi-attribute controllable text editing and style manipulation through continuous latent representations and energy-based guidance.
- Paper: Ask Me Anything: A simple strategy for prompting language models, Simran Arora et al. (2023). Develops multi-step functional prompt chains and weak supervision to reliably aggregate and improve zero-shot language model outputs.
- Paper: UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation, Daixuan Cheng et al. (2023). Proposes universal demonstration retrieval to systematically enhance zero-shot task evaluation and performance across diverse language models.
