Fine-Grained Controllable Text Generation Using Non-Residual Prompting

Fredrik CarlssonJoey ÖhmanFangyu LiuSeverine VerlindenJoakim NivreMagnus Sahlgren

article2022ACL65 citations

Proposes a non-residual attention architecture that enables fine-grained steering of causal language models at arbitrary decoding steps without degrading model representations or requiring labeled training data.

Listen

Modern causal language models generate remarkably fluent text, but steering them to meet specific constraints—such as including essential keywords, following length limits, or maintaining narrative context—remains a major operational hurdle. Existing techniques force an undesirable compromise between high-level prompt instructions, which lose effectiveness over longer passages, and token-level decoding rules, which disrupt generation flow and limit overall output quality. This lack of reliable control limits the adoption of generative language models in sensitive, high-precision business workflows such as automated reporting, guided drafting, and factual data-to-text generation.

To address this limitation, the article introduces Non-Residual Prompting (NRP), an encoder-decoder architecture that allows independent prompt instructions to steer text generation at arbitrary points in time. The primary objective is to demonstrate that pre-trained language models can achieve precise, fine-grained control without degrading text quality or requiring massive computational resources for retraining.

To evaluate this approach, the researchers converted a standard pre-trained GPT-2 Large model into the NRP architecture through a multi-phase, self-supervised training routine that kept the original base model weights frozen. They evaluated performance across standard benchmark tasks and introduced Contextualized CommonGen (C2GEN), a new evaluation dataset requiring models to incorporate specified target words while maintaining relevance to a preceding three-sentence context. Performance was assessed through automatic linguistic metrics alongside rigorous human evaluations measuring common sense and contextual consistency.

Across multiple experiments, NRP substantially outperformed existing baselines in control precision and language fluency. In free-text generation on CommonGen, NRP achieved a 98.4% target word inclusion rate compared to 72.2% for standard prompting and 13.3% for plug-and-play decoding methods, while also maintaining lower perplexity scores that indicate superior fluency. On the contextualized C2GEN benchmark, standard prompting degraded sharply to a 57.0% inclusion rate as it lost track of instructions, whereas NRP maintained a 96.9% inclusion rate in free text and an 81.0% inclusion rate in single-sentence generation. Human evaluations confirmed that NRP maintained solid common sense adherence and high contextual relevance across all settings, while alternative approaches either collapsed in text quality or failed to incorporate the required content.

These findings demonstrate that organizations do not need to choose between strict factual adherence and natural language fluency. By isolating prompt signals so they do not degrade the internal states of the base model over time, NRP enables modular, fine-grained control over length, vocabulary, and context. This significantly mitigates hallucination and omission risks in production applications while keeping computational and deployment costs low, as the underlying language model requires no fine-tuning.

Decision-makers exploring automated text workflows should consider modular non-residual architectures over traditional rigid decoding heuristics or standard long-prompt templates. Organizations should pilot NRP in targeted editorial and document-generation settings where factual completeness is mandatory. Future development should expand the framework into multi-task prompt libraries and evaluate larger modern base models with varied positional encoding schemes to broaden applicability.

While the results provide high confidence in the architecture's effectiveness, the study's scope was focused primarily on word inclusion and sentence length controls using GPT-2-scale architectures. Human evaluations showed low inter-annotator agreement on nuanced common-sense judgments, and models explicitly trained on knowledge graphs still retained a slight edge in domain-specific logic. Deployments in mission-critical environments should therefore retain human-in-the-loop validation while further domain-specific prompt tuning is conducted.

Carlsson et al (2022).pdf
  • Paper: Guiding Large Language Models via Directional Stimulus Prompting, Zekun Li et al. (2023). This paper extends controllable text generation by using an auxiliary policy model to generate instance-specific directional stimulus prompts for fine-grained guidance without updating the target model.
  • Paper: Diffusion-LM Improves Controllable Text Generation, Xiang Lisa Li et al. (2022). This work explores continuous diffusion models to overcome the rigid left-to-right control limitations of autoregressive models, offering an alternative paradigm for fine-grained syntactic and semantic control.
  • Paper: Composable Text Controls in Latent Space with ODEs, Guangyi Liu et al. (2023). This study advances fine-grained, composable text generation by steering continuous latent representations via ordinary differential equations rather than sequential prompt injection.
  • Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). This work investigates iterative self-refinement to enforce complex constraints during generation, contrasting inference-time iterative correction with continuous prompt steering.
Cover for Fine-Grained Controllable Text Generation Using Non-Residual Prompting

Abstract

The introduction of immensely large causal language models (CLMs) has rejuvenated the interest in open-ended text generation. However, controlling the generative process for these Transformer-based models is at large an unsolved problem. Earlier work has explored either plug-and-play decoding strategies or more powerful but blunt approaches such as prompting. There hence currently exists a trade-off between fine-grained control and the capability for more expressive high-level instructions. To alleviate this trade-off, we propose an encoder-decoder architecture that enables intermediate text prompts at arbitrary time steps. We propose a resource-efficient method for converting a pre-trained CLM into this architecture and demonstrate its potential in various experiments, including the novel task of contextualized word inclusion. Our method provides strong results in multiple experimental settings, proving itself to be both expressive and versatile.¹

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Controllable Text Generation
  • 2.2 Evaluation of Generated Text
  • 3 Model Architecture
  • 3.1 Non-Residual Attention
  • 3.2 Position Invariant Transformation
  • 4 Training Procedure
  • 4.1 Initialization
  • 4.2 Pre-training
  • 4.3 Fine-tuning
  • 5 Contextualized CommonGen Dataset (C2GEN)
  • 6 Word Inclusion Experiments
  • 6.1 Model Configurations
  • 6.2 Evaluation Metrics
  • 6.3 Quantitative Results
  • 6.4 Qualitative Results
  • 7 Sentence Length Experiments
  • 8 Discussion and Future Work
  • 9 Conclusion
  • Acknowledgements
  • Ethical Considerations
  • References
  • A Training Details
  • A.1 Training Data
  • A.2 Training Settings
  • A.3 Randomized Prompt Template
  • B Inference Details
  • B.1 Experiment Configurations
  • B.2 Prompting Schema
  • B.3 Keyword2Text Incorporation
  • C Architectural Motivation
  • C.1 Residual vs Non-Residual Prompts
  • C.2 Positional Invariant Transformation
  • C.3 Results Without Positional Invariant Transformation
  • C.4 Cross-Attention vs Non-Residual Attention
  • D Extended Related Work
  • D.1 Decoding Strategies
  • D.2 Training Strategies
  • D.3 GPT Prompting
  • E Evaluation Metrics
  • E.1 String Matching as Evaluation Metrics
  • E.2 Automatic Metrics
  • E.3 Human Evaluation Metrics
  • F Dataset C2GEN
  • G Qualitative Examples
  • H GPT Prompt Templates

Knowls

  1. Knowl 1 — Non-residual attention for independently timed prompts

    model/method

    The paper introduces Non-Residual Prompting (NRP), an encoder-decoder architecture for applying textual instructions at arbitrary generation steps of a pretrained causal language model (CLM). The CLM maintains two information streams: a textual stream that performs ordinary causal self-attention without seeing prompts, and a non-residual stream that produces the next-token distribution by attending to the current token, previous textual key-values, and the prompt model’s key-values. For a prompt SPS_P, a generated prefix SCLM={w1,…,wn}S_{\mathrm{CLM}}=\{w_1,\ldots,w_n\}, and textual key-values KVTi<nKV_T^{i<n} from earlier positions i<ni<n, the computation is

    KVP=PromptModel⁡(SP),KVTn=CLM⁡(wn∣KVTi<n),p(wn+1)=CLM⁡(wn∣KVP,KVTi<n).KV_P=\operatorname{PromptModel}(S_P),\qquad KV_T^n=\operatorname{CLM}(w_n\mid KV_T^{i<n}),\qquad p(w_{n+1})=\operatorname{CLM}(w_n\mid KV_P,KV_T^{i<n}).

    Here, wnw_n is the token at generation step nn, KVTnKV_T^n is the textual stream’s current key-value state, KVPKV_P is the prompt model’s key-value output, and p(wn+1)p(w_{n+1}) is the probability distribution for the next token. Non-residual key-values are never cached as part of the textual history, so a prompt can affect later generation only through the token it causes to be generated at its own step. Consequently, different prompts can be applied at different time steps without allowing earlier instructions to persist through the CLM’s internal textual state. Because the prompt model and CLM can compute in parallel rather than waiting for a fully encoded prompt as in conventional cross-attention, the paper argues that the architecture has a theoretical inference-speed advantage of up to a factor of two.

  2. Knowl 2 — Position-invariant prompt transformation

    model/method

    To make a prompt applicable at any generation position despite the CLM’s positional encoding, NRP adds a learned position-invariant transformation CC. If a prompt of length nn produces prompt key-values KVP⋆={kv1,kv2,…,kvn}KV_P^\star=\{kv^1,kv^2,\ldots,kv^n\}, where kvikv^i contains the key-values for prompt token ii across all LL transformer layers, the transformed prompt is

    KVP={kv1+C,kv2+C,…,kvn+C}.KV_P=\{kv^1+C,kv^2+C,\ldots,kv^n+C\}.

    The tensor CC has one learned parameter for every key-value parameter of the CLM and is added pointwise and identically to every prompt position. The transformation is trained after the prompt model so that the same instruction can be supplied at different points in a generation, while avoiding direct modification of the CLM’s absolute positional encodings. During optional task-specific fine-tuning, CC is temporarily removed, and it is reinserted after the prompt model has been fine-tuned.

  3. Knowl 3 — Resource-efficient conversion of a pretrained CLM

    model/method

    The paper converts a pretrained CLM into an NRP system without updating the original CLM and without requiring labeled task data. First, the prompt model is initialized by cloning the pretrained CLM. Second, the prompt model is trained on a word-inclusion objective using single Wikipedia sentences of 5–32 GPT-2-tokenized tokens; each prompt contains a target sentence length and 3–6 sampled unique non-stop words, and the position-invariant transformation is disabled. Third, only the transformation CC is trained on packed sequences of multiple sentences totaling up to 128 tokens, with an independently generated prompt for each sentence and attention restricted to the relevant prompt. Fourth, an optional task-specific fine-tuning stage updates only the prompt model on a target dataset without CC, after which the learned transformation is restored. All stages use teacher-forced causal language modeling, and the CLM remains frozen throughout.

    Both pretraining stages use batches of 1,280 samples, a maximum learning rate of 10−410^{-4}, linear warm-up over the first 500 updates, and early stopping based on word-inclusion coverage on a held-out validation set. The first stage validates on CommonGen; the second uses packed Wikipedia sentences. Prompts are procedurally randomized: each contains three demonstrations sampled from a held-out set of 50,000 sentences, with randomized instruction phrases, separators, delimiters, and formatting. The staged procedure reduces memory use because the optimizer state for the prompt model is not needed while training CC.

  4. Knowl 4 — Contextualized CommonGen and the C2GEN task

    data/table

    C2GEN extends CommonGen from unconstrained word inclusion to contextualized word inclusion. Each example supplies a set of 3–5 target words and a human-written context of three sentences; the system must generate a commonsense continuation that includes the target words and remains relevant to the context. The contexts were written by native-English Mechanical Turk annotators so that a subsequent sentence would naturally use the target words, and they were manually checked for quality.

    C2GEN contains 1,483 concept sets: 494 with three target words, 496 with four, and 493 with five. It has 1,122 unique concepts, 6,835 unique concept pairs, and 6,959 unique concept triples. The average context-sentence length is 43.92±10.3243.92\pm10.32 words; the contexts have GPT-2 XL perplexity 21.93 and Self-BLEU 16.63. For comparison, the corresponding CommonGen test set contains 1,497 concept sets, with 747 four-word and 750 five-word sets. C2GEN was constructed with a more balanced target-set-size distribution and deduplicated after filtering, making it slightly smaller than the CommonGen test set.

  5. Knowl 5 — Inference and evaluation protocol

    experimental setup

    The main NRP system pairs a newly trained prompt model with GPT-2 Large. Word-inclusion experiments use two generation settings: free text of exactly 32 tokens, and generation of one sentence. Decoding uses beam size 4 and a repetition penalty of 1.25; context-free generations are initialized with the word “The.” The prompted sentence lengths are selected on held-out data: 10 tokens for free text and 15 tokens for single-sentence generation. For each example, one prompt contains all target words and the desired length. NRP attends to that instruction until all target words have been generated; it then falls back to ordinary CLM generation. In free-text generation, the instruction remains active if the sentence ends before all targets are included. KG-BART and POINTER are not evaluated on C2GEN because they do not support the required contextualized setting.

    The paper also evaluates NRP combined with a modified Keyword2Text decoder. For each not-yet-included target word, the method multiplies its CLM logit by a time-dependent factor, also applying the factor to its lemmas; once a target word is included, all of its lemmas receive zero sampling probability, and the modifier is disabled after all targets are present. With T∈[0,1]T\in[0,1] denoting the fraction of the generation completed, maximum increase α\alpha, and curve parameter λ\lambda, the modifier is

    m(T)=1+αeλTeλ.m(T)=1+\alpha\frac{e^{\lambda T}}{e^\lambda}.

    All reported NRP-plus-Keyword2Text experiments use α=0.5\alpha=0.5 and λ=5.5\lambda=5.5.

    Automatic evaluation reports lemmatized target-word coverage, GPT-2 XL perplexity, and Self-BLEU-5; higher coverage and lower perplexity and Self-BLEU-5 are preferred, although perplexity and Self-BLEU are affected by sequence length. Human common-sense and, for C2GEN, contextual-relevance scores are averages over 100 generated texts per method, with five native-English judges per text. Judges select “Yes,” “Partly,” or “No,” mapped to 11, 0.50.5, or 00.

  6. Knowl 6 — NRP results on CommonGen

    data/table

    On CommonGen, NRP provides substantially stronger word inclusion than prompted GPT-2, PPLM, and Keyword2Text in free-text generation while retaining comparatively good fluency. In single-sentence generation, NRP is less complete than POINTER and KG-BART on coverage but is more balanced in fluency, diversity, and common sense. The table reports coverage (Cov, percent of target words included), perplexity (Ppl, lower is better), Self-BLEU-5 (lower indicates more diversity), human common-sense score (Sense, higher is better), and generated length (Len for single-sentence outputs). A dash indicates that a baseline was not evaluated in that setting.

    Could not parse LaTeX table

    The NRP-plus-Keyword2Text combination raises coverage slightly over NRP but also raises Self-BLEU-5, indicating reduced diversity. NRP’s free-text coverage is 98.4%, compared with 72.2% for prompted GPT-2 and 13.3% for PPLM; its perplexity is also lower than the other free-text baselines except for the small difference from PPLM’s 17.2. In single-sentence generation, NRP reaches 93.0% coverage with a perplexity of 24.0, while POINTER and KG-BART reach 98.0% and 97.2% coverage but have perplexities of 51.9 and 37.0, respectively.

  7. Knowl 7 — NRP results on contextualized word inclusion

    data/table

    On C2GEN, NRP is the only evaluated approach that combines strong target-word inclusion with contextual generation in both free-text and single-sentence settings. The table reports coverage (Cov, percent of target words included), GPT-2 XL perplexity (Ppl, lower is better), Self-BLEU-5 (lower indicates more diversity), human common-sense score (Sense), contextual-relevance score (Ctx), and generated length (Len for single-sentence outputs). Human scores are on a 00–11 scale; higher Sense and Ctx are better.

    Could not parse LaTeX table

    In free-text C2GEN generation, NRP obtains 96.9% coverage and perplexity 10.0, compared with 57.0% and 12.5 for prompted GPT-2, 19.1% and 12.5 for PPLM, and 93.9% and 18.0 for Keyword2Text. In single-sentence generation, NRP reaches 81.0% coverage and 81.2 contextual relevance, whereas prompted GPT-2 reaches 56.6% coverage and 74.3 contextual relevance. Adding Keyword2Text increases NRP coverage to 98.6% in free text and 82.1% in single sentences, with only small changes to contextual relevance and diversity. The methods have similar contextual-relevance scores in free-text generation, but NRP is markedly better at satisfying the word-inclusion constraint.

  8. Knowl 8 — Prompted sentence length controls generation style

    empirical result

    Including a target sentence length in the NRP pretraining objective gives the model a controllable stylistic and planning signal. LPL_P denotes the prompted length and LGL_G the resulting generated length. The generated content changes as the requested length changes, rather than merely being truncated or padded.

    Could not parse LaTeX table

    Across CommonGen validation examples, the mean generated-length offset remains above 0 and below 1 token, so NRP generally produces slightly longer sentences than requested. The standard deviation grows for both very short and very long requested lengths, reflecting the length distribution of the Wikipedia pretraining data. The model prioritizes linguistic quality over exact length matching.

  9. Knowl 9 — The position-invariant transformation is necessary with context

    empirical result

    An ablation of the position-invariant transformation shows that it is especially important when NRP must operate over context. The comparison uses the same NRP prompt model with or without the learned transformation. Metrics are coverage (Cov, higher is better), GPT-2 XL perplexity (Ppl, lower is better), and Self-BLEU-5 (lower is more diverse).

    Could not parse LaTeX table

    Removing the transformation causes only moderate degradation on context-free CommonGen but makes contextualized generation collapse: C2GEN coverage falls from 96.9% to 6.2% in free text and from 81.0% to 1.6% in single sentences, while perplexity rises to 39,796 and 41,905. The paper also reports that directly shifting GPT-2’s absolute positional inputs by one position produces repetitive degenerate continuations, motivating the learned additive transformation instead.

  10. Knowl 10 — Scope and limitations of the demonstrated method

    limitation

    The paper’s training procedure is primarily a resource-constrained approximation to the authors’ preferred approach of training the prompt model directly on long-context data without a separate positional transformation. The experiments therefore do not establish that the staged procedure is optimal, and alternative positional-encoding schemes remain unexplored. The authors also did not explicitly train NRP for common-sense reasoning; methods designed specifically for that objective, such as KG-BART, score better on the CommonGen common-sense metric in single-sentence generation. The demonstrated task is word inclusion, so the broader claim that non-residual prompts generalize to other controllable-generation tasks remains prospective rather than experimentally established. Finally, the inference procedure uses one fixed prompt containing all target words rather than adaptively updating the instruction as words are generated, and the paper acknowledges substantial remaining room for improvement on both CommonGen and C2GEN.

Coverage note — The appendix’s detailed prompt templates, baseline-specific implementation descriptions, qualitative example catalog, and annotation-agreement diagnostics were omitted because they provide supporting detail rather than additional load-bearing contributions.

References

  1. 1.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. Spice: Semantic propositional image caption evaluation. In European conference on computer vision, pages 382–398. Springer.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  3. 3.Anya Belz, Simon Mille, and David M. Howcroft. 2020. Disentangling the properties of human evaluation methods: A classification system to support comparability, meta-evaluation and reproducibility testing. In Proceedings of the 13th International Conference on Natural Language Generation, pages 183–194, Dublin, Ireland. Association for Computational Linguistics.
  4. 4.Shikha Bordia and Samuel R. Bowman. 2019. Identifying and reducing gender bias in word-level language models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop, pages 7–15, Minneapolis, Minnesota. Association for Computational Linguistics.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. Technical report, OpenAI.
  6. 6.Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020. Evaluation of text generation: A survey. arXiv preprint arXiv:2006.14799.
  7. 7.Jordan Clive, Kris Cao, and Marek Rei. 2021. Control prefixes for text generation. arXiv preprint arXiv:2110.08329.
  8. 8.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations.
  9. 9.Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017. The WebNLG challenge: Generating text from RDF data. In Proceedings of the 10th International Conference on Natural Language Generation, pages 124–133, Santiago de Compostela, Spain. Association for Computational Linguistics.
  10. 10.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh D Dhole, et al. 2021. The gem benchmark: Natural language generation, its evaluation and metrics. arXiv preprint arXiv:2102.01672.
  11. 11.Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020. spaCy: Industrial-strength Natural Language Processing in Python.
  12. 12.David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020. Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions. In Proceedings of the 13th International Conference on Natural Language Generation, pages 169–182, Dublin, Ireland. Association for Computational Linguistics.
  13. 13.Brad Jascob. 2019. lemminflect: A python module for english lemmatization and inflection.
  14. 14.Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation.
  15. 15.Muhammad Khalifa, Hady Elsahar, and Marc Dymetman. 2021. A distributional approach to controlled text generation. In International Conference on Learning Representations.
  16. 16.Rik Koncel-Kedziorski, Dhanush Bekal, Yi Luan, Mirella Lapata, and Hannaneh Hajishirzi. 2019. Text Generation from Knowledge Graphs with Graph Transformers. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2284–2293, Minneapolis, Minnesota. Association for Computational Linguistics.
  17. 17.Leo Leppanen, Myriam Munezero, Mark Granroth-Wilding, and Hannu Toivonen. 2017. Data-driven news generation for automated journalism. In Proceedings of the 10th International Conference on Natural Language Generation, pages 188–197, Santiago de Compostela, Spain. Association for Computational Linguistics.
  18. 18.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  19. 19.Boyang Li, Stephen Lee-Urban, George Johnston, and Mark Riedl. 2013. Story generation with crowdsourced plot graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 27, pages 598–604.
  20. 20.Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. CommonGen: A constrained text generation challenge for generative commonsense reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1823–1840, Online. Association for Computational Linguistics.
  21. 21.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  22. 22.Ye Liu, Yao Wan, Lifang He, Hao Peng, and Philip S Yu. 2021. Kg-bart: Knowledge graph-augmented bart for generative commonsense reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 6418–6425.
  23. 23.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  24. 24.Damian Pascual, Beni Egressy, Clara Meister, Ryan Cotterell, and Roger Wattenhofer. 2021. A plug-and-play method for controlled text generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3973–3997, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  25. 25.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  26. 26.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018a. Improving Language Understanding by Generative Pre-Training. Technical report, OpenAI.
  27. 27.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2018b. Language models are unsupervised multitask learners. Technical report, OpenAI.
  28. 28.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  29. 29.Emily Reif, Daphne Ippolito, Ann Yuan, Andy Coenen, Chris Callison-Burch, and Jason Wei. 2021. A recipe for arbitrary text style transfer with large language models. arXiv preprint arXiv:2109.03910.
  30. 30.Mark Riedl. 2021. An introduction to ai story generation. The Gradient.
  31. 31.Alexis Ross, Tongshuang Wu, Hao Peng, Matthew E Peters, and Matt Gardner. 2021. Tailor: Generating and perturbing text with semantic controls. arXiv preprint arXiv:2107.07150.
  32. 32.Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019. Best practices for the human evaluation of automatically generated text. In Proceedings of the 12th International Conference on Natural Language Generation, pages 355–368, Tokyo, Japan. Association for Computational Linguistics.
  33. 33.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  34. 34.Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4566–4575.
  35. 35.Ronald J. Williams and David Zipser. 1989. A learning algorithm for continually running fully recurrent neural networks. Neural Computation, 1(2):270–280.
  36. 36.Lili Yao, Nanyun Peng, Ralph Weischedel, Kevin Knight, Dongyan Zhao, and Rui Yan. 2019. Plan-and-write: Towards better automatic storytelling. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 7378–7385.
  37. 37.Yizhe Zhang, Guoyin Wang, Chunyuan Li, Zhe Gan, Chris Brockett, and Bill Dolan. 2020. POINTER: Constrained progressive text generation via insertion-based generative pre-training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8649–8670, Online. Association for Computational Linguistics.
  38. 38.Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018. Texygen: A benchmarking platform for text generation. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR ’18, page 1097–1100, New York, NY, USA. Association for Computing Machinery.
  39. 39.Xu Zou, Da Yin, Qingyang Zhong, Hongxia Yang, Zhilin Yang, and Jie Tang. 2021. Controllable generation from pre-trained language models via inverse prompting. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pages 2450–2460.

Citation

MLA
Carlsson, F., et al. “Fine-Grained Controllable Text Generation Using Non-Residual Prompting”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 6837–57, https://doi.org/10.18653/v1/2022.acl-long.471.
APA
Carlsson, F., Öhman, J., Liu, F., Verlinden, S., Nivre, J., & Sahlgren, M. (2022). Fine-Grained Controllable Text Generation Using Non-Residual Prompting. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6837–6857. https://doi.org/10.18653/v1/2022.acl-long.471
Chicago
Carlsson, F., J. Öhman, F. Liu, S. Verlinden, J. Nivre, and M. Sahlgren. 2022. “Fine-Grained Controllable Text Generation Using Non-Residual Prompting”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6837–57. https://doi.org/10.18653/v1/2022.acl-long.471.
Harvard
Carlsson, F. et al. (2022) “Fine-Grained Controllable Text Generation Using Non-Residual Prompting”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6837–6857. Available at: https://doi.org/10.18653/v1/2022.acl-long.471.
Vancouver
1. Carlsson F, Öhman J, Liu F, Verlinden S, Nivre J, Sahlgren M (2022) Fine-Grained Controllable Text Generation Using Non-Residual Prompting. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6837–6857

BibTeX

@inproceedings{carlsson-etal-2022-fine,
    title = "Fine-Grained Controllable Text Generation Using Non-Residual Prompting",
    author = {Carlsson, Fredrik  and
      {\"O}hman, Joey  and
      Liu, Fangyu  and
      Verlinden, Severine  and
      Nivre, Joakim  and
      Sahlgren, Magnus},
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.471/",
    doi = "10.18653/v1/2022.acl-long.471",
    pages = "6837--6857"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/