A Recipe for Arbitrary Text Style Transfer with Large Language Models

Emily ReifDaphne IppolitoAnn YuanAndy CoenenChris Callison-BurchJason Wei

article2022ACL162 citationsArea Chair Award

Proposes an augmented zero-shot prompting technique that enables large language models to perform arbitrary text style transfer using only natural language instructions without fine-tuning or target-style exemplars.

Listen

Text style transfer modifies the tone, mood, or stylistic elements of a passage while retaining its core meaning. Traditional approaches depend heavily on large labeled datasets of parallel text or specialized exemplars for every target style, which makes adapting to arbitrary or non-standard styles difficult and resource-intensive. As demand grows for flexible AI writing assistants, there is a clear operational need for flexible, training-free text transformation methods.

The article evaluates a prompting technique called augmented zero-shot learning. The main objective is to demonstrate that large language models can perform text style transfer across arbitrary, user-specified styles using natural language instructions without fine-tuning or style-specific examples.

The researchers framed style transfer as a generalized sentence rewriting task. Instead of using examples of the specific target style, they primed large language models (primarily LaMDA with 137 billion parameters, alongside GPT-3 variants) with a single multi-task prompt containing diverse sentence rewrites wrapped in consistent formatting delimiters. They evaluated performance across six non-standard styles (such as adding metaphors or adjusting melodrama) and standard style benchmarks (sentiment and formality). Assessment relied on 3,600 human evaluation ratings covering transfer strength, semantic preservation, and fluency, supported by automatic metric evaluations and a user study involving 30 creative writers.

The findings show that augmented zero-shot prompting performs comparably to human-written text and supervised baselines across standard and atypical styles. Compared to standard zero-shot prompting, which failed to return a valid response 25.4% of the time, the augmented technique drastically reduced unparseable outputs to 0.6%. On sentiment benchmarks, the method achieved 90.6% accuracy, approaching the 94.3% accuracy of five-shot prompting while significantly improving text fluency. However, the model produced lower overlap scores (BLEU) against target human references because it frequently elaborated on text rather than making literal word substitutions.

These results demonstrate that organizations can implement flexible text-editing and style-transfer capabilities without the high costs, timelines, and technical overhead of curating specialized training datasets or fine-tuning models. However, the method offers less fine-grained constraint than models trained on fixed data, as large language models may hallucinate information or lean naturally toward formal or melodramatic phrasing.

Decision-makers should consider augmented zero-shot prompting for open-ended creative tools and interactive writing aids where user flexibility is paramount. When precise lexical preservation is required, teams should implement candidate selection filtering or specialized supervised pipelines. Further testing is recommended to explore broader safety guardrails and systematic evaluations of model behavior across different base architectures.

arXiv: 2109.03910
Cover for A Recipe for Arbitrary Text Style Transfer with Large Language Models

Abstract

In this paper, we leverage large language models (LMs) to perform zero-shot text style transfer. We present a prompting method that we call augmented zero-shot learning, which frames style transfer as a sentence rewriting task and requires only a natural language instruction, without model fine-tuning or exemplars in the target style. Augmented zero-shot learning is simple and demonstrates promising results not just on standard style transfer tasks such as sentiment, but also on natural language transformations such as "make this melodramatic" or "insert a metaphor.

Table of Contents

  • 1 Introduction
  • 2 Augmented zero-shot prompting
  • 3 Experimental Setup
  • 4 Results
  • 4.1 Non-Standard Styles
  • 4.2 Standard Styles
  • 5 Potential of Arbitrary Styles
  • 6 Limitations and Failure Modes
  • 7 Conclusions
  • References
  • Appendix
  • A Prompt Selection
  • B Low BLEU for LLM Outputs
  • C Further Related Work

Knowls

  1. Knowl 1 — Augmented zero-shot prompting for arbitrary style transfer

    model/method

    Augmented zero-shot learning is an inference-only prompting method that makes a large language model rewrite an input sentence according to a natural-language style instruction without fine-tuning, labeled training data, or exemplars written in the requested target style.

    The prompt contains a fixed set of demonstrations showing varied sentence-rewriting operations—such as making text scarier, more intense, more descriptive, or inserting a specified word—using a consistent input/output template. The requested source sentence and arbitrary instruction are then appended in the same format, and the language model generates the rewrite as its continuation. For example, the same prompt can request that a sentence become more melodramatic, contain a metaphor, or include the word “balloon.” The demonstrations are therefore task-general rather than specific to the requested transformation; they mainly encourage a valid, parseable rewrite while preserving the flexibility of natural-language instructions.

  2. Knowl 2 — Large-language-model configurations and decoding procedure

    experimental setup

    The main experiments use LaMDA, a left-to-right decoder-only Transformer with 137 billion non-embedding parameters. The pretrained LaMDA model, called LLM, was trained on 1.95 billion public web documents, tokenized into 2.49 trillion byte-pair-encoding tokens with a 32,000-token SentencePiece vocabulary. A second model, LLM-Dialog, is the final LaMDA model further fine-tuned on a curated high-quality conversational subset.

    LaMDA generations use top-k=40k=40 decoding. The method is also evaluated with GPT-3 models of different sizes, including ada, curie, and davinci; GPT-3 decoding uses nucleus sampling with probability threshold p=0.6p=0.6. For LLM-Dialog, the demonstrations and request are formatted as alternating conversational turns rather than ordinary text-completion examples. Curly braces delimit the source and generated rewrite so that the output can be extracted automatically.

  3. Knowl 3 — Human evaluation on non-standard style transformations

    empirical result

    The method was evaluated on six transformations that lack standard labeled classifiers: make a sentence more comic, more melodramatic, more descriptive, include a metaphor, include the word “park,” and include the word “balloon.” The inputs were 50 randomly selected, grammatical sentences from the Reddit Writing Prompts validation set, excluding sentences that already clearly exhibited one of the requested styles.

    Six professional English-fluent raters evaluated tuples consisting of an input sentence, a target instruction, and an output sentence. Each tuple received three independent ratings after a calibration exercise, producing 3,600 ratings. Transfer strength and semantic preservation were scored from 1 to 100, while fluency was judged categorically. The comparison systems were direct zero-shot prompting, a paraphrase control using the same generic rewriting demonstrations but the instruction “paraphrased,” and human-written transformations.

    Augmented zero-shot outputs were rated almost as highly as the human-written transformations across transfer strength, semantic preservation, and fluency. Direct zero-shot prompting failed to produce a valid response 25.4% of the time, compared with 0.6% for augmented zero-shot prompting. The augmented outputs averaged 107 characters, versus 66 characters for the inputs and 82 characters for human paraphrases. The requested word was inserted successfully in 85% of the “park” and “balloon” cases. Relative to the original sentences, augmented outputs were 252% longer for “more descriptive” and 146% longer for “include a metaphor”; the corresponding human transformations were 165% and 146% longer.

  4. Knowl 4 — Human evaluation on sentiment and formality transfer

    empirical result

    Augmented zero-shot learning was also tested on the standard transformations “more positive,” “more negative,” “more formal,” and “more informal,” using the Yelp Polarity dataset for sentiment and the Grammarly Yahoo Answers Formality Corpus for formality. Human evaluation used the same three dimensions—transfer strength, semantic preservation, and fluency—and compared augmented zero-shot outputs with direct zero-shot prompting, the paraphrase control, human-written transformations, Unsup. MT, and Dual RL.

    For 50 sentences per style, six outputs or baselines were independently rated by three raters, yielding 3,000 ratings. The augmented zero-shot outputs were judged comparably to both human-written transformations and the two prior trained style-transfer systems, Unsup. MT and Dual RL, across the evaluated styles. The result indicates that the inference-only prompting method can approach the human-evaluation quality of task-specific systems even on conventional style-transfer benchmarks.

  5. Knowl 5 — Automatic sentiment-transfer comparison with supervised and inference-only systems

    data/table

    On the Yelp sentiment-transfer task, transfer strength was measured by classifier accuracy, semantic similarity to human references by BLEU, and fluency by GPT-2 perplexity. Lower perplexity is better. The results compare supervised style-transfer systems with inference-only prompting using LaMDA and GPT-3.

    System Acc BLEU PPL
    Cross-alignment 73.4 17.6 812
    Backtrans 90.5 5.1 424
    Multidecoder 50.3 27.7 1,703
    Delete-only 81.4 28.6 606
    Delete-retrieve 86.2 31.1 948
    Unpaired RL 52.2 37.2 2,750
    Dual RL 85.9 55.1 982
    Style transformer 82.1 55.2 935
    GPT-3 ada, augmented zero-shot 31.5 39.0 283
    GPT-3 curie, augmented zero-shot 53.0 48.3 207
    GPT-3 davinci, augmented zero-shot 74.1 43.8 231
    LLM, zero-shot 69.7 28.6 397
    LLM, five-shot 83.2 19.8 240
    LLM, augmented zero-shot 79.6 16.1 173
    LLM-Dialog, zero-shot 59.1 17.6 138
    LLM-Dialog, five-shot 94.3 13.6 126
    LLM-Dialog, augmented zero-shot 90.6 10.4 79

    Augmented zero-shot prompting gives substantially better accuracy and lower perplexity than vanilla zero-shot prompting for both LaMDA configurations, and it approaches five-shot accuracy without target-task demonstrations. Larger GPT-3 models perform better than smaller ones, with accuracy increasing from 31.5 for ada to 53.0 for curie and 74.1 for davinci. Augmented zero-shot outputs have lower BLEU than most supervised systems despite strong classifier accuracy, which the paper attributes to generating semantically appropriate rewrites with wording different from the human references.

  6. Knowl 6 — User study of free-form rewriting requests

    experimental setup

    To examine whether users would request styles beyond standard sentiment or formality transformations, the authors built an AI-assisted story-writing editor with a free-form “rewrite as” control powered by augmented zero-shot learning. Thirty members of a creative-writing group each wrote a story of 100–300 words and could request arbitrary rewrites of selected passages. The study collected 333 rewrite requests.

    The requests included stylistic and semantic goals such as making text less angsty, more Dickensian, more magical, more suspenseful, more technical, more whimsical, about mining, about vegetables, or less diabolical. Examples of accepted model suggestions transformed prose to be about mining, more interesting, more descriptive, and more emotional. The diversity of requests provides qualitative evidence that a natural-language instruction interface exposes substantially broader style control than a fixed menu of predefined style labels.

  7. Knowl 7 — Candidate selection improves lexical overlap but changes other metrics

    empirical result

    Augmented zero-shot outputs often preserve meaning while using wording different from human reference rewrites, which produces low BLEU even when human semantic-preservation ratings are high. Each generation run produced 16 possible continuations. A candidate-selection variant chose the continuation with the highest BLEU against the original source sentence, favoring lexical overlap with the input.

    The resulting Yelp sentiment-transfer metrics were:

    Model Setting Acc BLEU PPL
    LLM-128B Zero-shot 69.7 28.6 397
    LLM-128B Zero-shot + candidate selection 31.4 61.5 354
    LLM-128B Five-shot 83.2 19.8 240
    LLM-128B Five-shot + candidate selection 61.5 55.6 306
    LLM-128B Augmented zero-shot 79.6 16.1 173
    LLM-128B Augmented zero-shot + candidate selection 65.0 49.3 292
    LLM-128B-dialog Zero-shot 59.1 17.6 138
    LLM-128B-dialog Zero-shot + candidate selection 46.8 24.2 166
    LLM-128B-dialog Five-shot 94.3 13.6 126
    LLM-128B-dialog Five-shot + candidate selection 81.3 47.6 345
    LLM-128B-dialog Augmented zero-shot 90.6 10.4 79
    LLM-128B-dialog Augmented zero-shot + candidate selection 73.7 40.6 184

    Candidate selection substantially raises BLEU—for example, from 16.1 to 49.3 for LLM-128B augmented zero-shot and from 10.4 to 40.6 for LLM-128B-dialog augmented zero-shot—but reduces sentiment accuracy and worsens perplexity. Thus, the low BLEU of the original generations reflects lexical divergence from references rather than necessarily poor semantic preservation.

  8. Knowl 8 — Sensitivity to sentiment-instruction wording

    data/table

    The performance of augmented zero-shot learning varies with the natural-language wording of the sentiment instruction. Four equivalent instruction pairs were tested: “more positive/negative,” “happier/sadder,” “more optimistic/pessimistic,” and “more cheerful/miserable.” Accuracy is sentiment-classifier accuracy, BLEU measures overlap with human reference rewrites, and PPL is GPT-2 perplexity.

    Model Prompt wording Acc BLEU PPL
    LLM more positive/negative 76.3 14.8 180
    LLM happier/sadder 62.6 15.5 173
    LLM more optimistic/pessimistic 69.7 14.1 143
    LLM more cheerful/miserable 74.5 15.7 186
    LLM-Dialog more positive/negative 90.5 10.4 79
    LLM-Dialog happier/sadder 85.9 9.6 90
    LLM-Dialog more optimistic/pessimistic 85.8 10.2 79
    LLM-Dialog more cheerful/miserable 88.8 11.4 93

    The four phrasings produce different scores, but their results are broadly comparable. This shows that augmented zero-shot behavior depends somewhat on prompt wording even when the intended transformation is unchanged.

  9. Knowl 9 — Operational failure modes of prompted style transfer

    limitation

    Large-language-model style transfer can fail even when the requested transformation is understandable. The generated response may be unparseable: instead of returning only a rewrite, the model may produce conversational praise such as “Sounds like you are a great writer!” or provide a recommendation embedded in prose such as “a good rewrite might be to say that the dress is pretty.” Curly-brace delimiters were used to reduce this problem, and responses without a valid delimited answer were treated as invalid.

    The model also frequently hallucinates new content. Such additions can be useful for creative writing but are undesirable in applications such as summarization, where unsupported information is harmful. In addition, even the paraphrase control tends to drift toward certain styles, including formality and melodrama, indicating that the language model has inherent stylistic preferences that can interfere with fine-grained control.

  10. Knowl 10 — Reliability, controllability, and safety limitations

    limitation

    For style-transfer tasks with available training data, task-specific supervised or fine-tuned systems are expected to be more reliable and offer finer control over the properties of the output. The paper observes this trade-off in the lower BLEU scores of augmented zero-shot prompting despite comparable transfer accuracy in sentiment evaluation. The inference-only method therefore exchanges training-data requirements and broad stylistic coverage for less precise control.

    The approach also inherits the barriers, biases, and safety risks of the underlying large language models. Arbitrary instructions can request harmful transformations such as making text more racist, sexist, or incendiary. The authors view probing such requests as potentially useful for exposing model boundaries and failure modes, but the method itself does not remove the underlying safety concerns.

Coverage note — No substantial contributed material was omitted; illustrative prompt examples and related-work discussion were excluded because they do not add standalone contribution beyond the method, evaluations, user study, and stated limitations.

References

  1. 1.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623, New York, NY, USA. Association for Computing Machinery.
  2. 2.Gwern Branwen. 2020. GPT-3 creative fiction.
  3. 3.Eleftheria Briakou, Sweta Agrawal, Ke Zhang, Joel R. Tetreault, and Marine Carpuat. 2021. A review of human evaluation for style transfer. CoRR, abs/2106.04747.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. CoRR, abs/2005.14165.
  5. 5.Ning Dai, Jianze Liang, Xipeng Qiu, and Xuanjing Huang. 2019. Style transformer: Unpaired text style transfer without disentangled latent representation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5997–6007, Florence, Italy. Association for Computational Linguistics.
  6. 6.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  7. 7.Zhenxin Fu, Xiaoye Tan, Nanyun Peng, Dongyan Zhao, and Rui Yan. 2018. Style transfer in text: Exploration and evaluation. In Proceedings of the AAAI Conference on Artificial Intelligence.
  8. 8.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. In International Conference on Learning Representations.
  9. 9.Zhiqiang Hu, Roy Ka-Wei Lee, and Charu C. Aggarwal. 2020. Text style transfer: A review and experiment evaluation. CoRR, abs/2010.12742.
  10. 10.Di Jin, Zhijing Jin, Zhiting Hu, Olga Vechtomova, and Rada Mihalcea. 2020. Deep learning for text style transfer: A survey. CoRR, abs/2011.00416.
  11. 11.Zhijing Jin, Di Jin, Jonas Mueller, Nicholas Matthews, and Enrico Santus. 2019. IMaT: Unsupervised text attribute transfer via iterative matching and translation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3097–3109, Hong Kong, China. Association for Computational Linguistics.
  12. 12.Kalpesh Krishna, John Wieting, and Mohit Iyyer. 2020. Reformulating unsupervised style transfer as paraphrase generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 737–762, Online. Association for Computational Linguistics.
  13. 13.Taku Kudo and John Richardson. 2018. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. CoRR, abs/1808.06226.
  14. 14.Juncen Li, Robin Jia, He He, and Percy Liang. 2018. Delete, retrieve, generate: a simple approach to sentiment and style transfer. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1865–1874, New Orleans, Louisiana. Association for Computational Linguistics.
  15. 15.Dayiheng Liu, Jie Fu, Yidan Zhang, Chris Pal, and Jiancheng Lv. 2020. Revision in continuous space: Unsupervised text style transfer without adversarial learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8376–8383.
  16. 16.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021a. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  17. 17.Ruibo Liu, Chenyan Jia, and Soroush Vosoughi. 2021b. A transformer-based framework for neutralizing and reversing the political polarity of news articles. Proc. ACM Hum.-Comput. Interact., 5(CSCW1).
  18. 18.Fuli Luo, Peng Li, Jie Zhou, Pengcheng Yang, Baobao Chang, Xu Sun, and Zhifang Sui. 2019. A dual reinforcement learning framework for unsupervised text style transfer. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI 2019, Macao, China, August 10-16, 2019, pages 5116–5122. ijcai.org.
  19. 19.Aman Madaan, Amrith Setlur, Tanmay Parekh, Barnabas Poczos, Graham Neubig, Yiming Yang, Ruslan Salakhutdinov, Alan W Black, and Shrimai Prabhumoye. 2020. Politeness transfer: A tag and generate approach. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1869–1881, Online. Association for Computational Linguistics.
  20. 20.Remi Mir, Bjarke Felbo, Nick Obradovich, and Iyad Rahwan. 2019. Evaluating style transfer for text. CoRR, abs/1904.02295.
  21. 21.Shrimai Prabhumoye, Yulia Tsvetkov, Ruslan Salakhutdinov, and Alan W Black. 2018. Style transfer through back-translation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 866–876, Melbourne, Australia. Association for Computational Linguistics.
  22. 22.Raul Puri and Bryan Catanzaro. 2019. Zero-shot text classification with generative language models. arXiv preprint arXiv:1912.10165.
  23. 23.Sudha Rao and Joel Tetreault. 2018. Dear sir or madam, may I introduce the GYAFC dataset: Corpus, benchmarks and metrics for formality style transfer. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 129–140, New Orleans, Louisiana. Association for Computational Linguistics.
  24. 24.Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm.
  25. 25.Parker Riley, Noah Constant, Mandy Guo, Girish Kumar, David C. Uthus, and Zarana Parekh. 2021. Textsettr: Label-free text style extraction and tunable targeted restyling. Proceedings of the Annual Meeting of the Association of Computational Linguistics (ACL).
  26. 26.Keisuke Sakaguchi and Benjamin Van Durme. 2018. Efficient online scalar annotation with bounded support. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 208–218, Melbourne, Australia. Association for Computational Linguistics.
  27. 27.Timo Schick and Hinrich Schütze. 2021. It’s not just size that matters: Small language models are also few-shot learners. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2339–2352, Online. Association for Computational Linguistics.
  28. 28.Tianxiao Shen, Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2017. Style transfer from non-parallel text by cross-alignment. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  29. 29.Akhilesh Sudhakar, Bhargav Upadhyay, and Arjun Maheswaran. 2019. “Transforming” delete, retrieve, generate approach for controlled text style transfer. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3269–3279, Hong Kong, China. Association for Computational Linguistics.
  30. 30.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  31. 31.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. CoRR, abs/1706.03762.
  32. 32.Orion Weller, Nicholas Lourie, Matt Gardner, and Matthew E. Peters. 2020. Learning from task descriptions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1361–1375, Online. Association for Computational Linguistics.
  33. 33.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  34. 34.Jingjing Xu, Xu Sun, Qi Zeng, Xiaodong Zhang, Xuancheng Ren, Houfeng Wang, and Wenjie Li. 2018. Unpaired sentiment-to-sentiment translation: A cycled reinforcement learning approach. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 979–988, Melbourne, Australia. Association for Computational Linguistics.
  35. 35.Peng Xu, Yanshuai Cao, and Jackie Chi Kit Cheung. 2020. On variational learning of controllable representations for text without supervision. Proceedings of the International Conference on Machine Learning (ICML), abs/1905.11975.
  36. 36.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. Proceedings of the Conference on Neural Information Processing Systems.
  37. 37.Zhemin Zhu, Delphine Bernhard, and Iryna Gurevych. 2010. A monolingual tree-based translation model for sentence simplification. In Proceedings of the 23rd International Conference on Computational Linguistics (COLING 2010), pages 1353–1361, Beijing, China. Coling 2010 Organizing Committee.

Citation

MLA
Reif, E., et al. “A Recipe for Arbitrary Text Style Transfer with Large Language Models”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2022, pp. 837–48, https://doi.org/10.18653/v1/2022.acl-short.94.
APA
Reif, E., Ippolito, D., Yuan, A., Coenen, A., Callison-Burch, C., & Wei, J. (2022). A Recipe for Arbitrary Text Style Transfer with Large Language Models. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 837–848. https://doi.org/10.18653/v1/2022.acl-short.94
Chicago
Reif, E., D. Ippolito, A. Yuan, A. Coenen, C. Callison-Burch, and J. Wei. 2022. “A Recipe for Arbitrary Text Style Transfer with Large Language Models”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 837–48. https://doi.org/10.18653/v1/2022.acl-short.94.
Harvard
Reif, E. et al. (2022) “A Recipe for Arbitrary Text Style Transfer with Large Language Models”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp. 837–848. Available at: https://doi.org/10.18653/v1/2022.acl-short.94.
Vancouver
1. Reif E, Ippolito D, Yuan A, Coenen A, Callison-Burch C, Wei J (2022) A Recipe for Arbitrary Text Style Transfer with Large Language Models. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp 837–848

BibTeX

@inproceedings{reif-etal-2022-recipe,
    title = "A Recipe for Arbitrary Text Style Transfer with Large Language Models",
    author = "Reif, Emily  and
      Ippolito, Daphne  and
      Yuan, Ann  and
      Coenen, Andy  and
      Callison-Burch, Chris  and
      Wei, Jason",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-short.94/",
    doi = "10.18653/v1/2022.acl-short.94",
    pages = "837--848"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/