Fine-Grained Controllable Text Generation Using Non-Residual Prompting
Fredrik CarlssonJoey ÖhmanFangyu LiuSeverine VerlindenJoakim NivreMagnus Sahlgren
Proposes a non-residual attention architecture that enables fine-grained steering of causal language models at arbitrary decoding steps without degrading model representations or requiring labeled training data.
Modern causal language models generate remarkably fluent text, but steering them to meet specific constraints—such as including essential keywords, following length limits, or maintaining narrative context—remains a major operational hurdle. Existing techniques force an undesirable compromise between high-level prompt instructions, which lose effectiveness over longer passages, and token-level decoding rules, which disrupt generation flow and limit overall output quality. This lack of reliable control limits the adoption of generative language models in sensitive, high-precision business workflows such as automated reporting, guided drafting, and factual data-to-text generation.
To address this limitation, the article introduces Non-Residual Prompting (NRP), an encoder-decoder architecture that allows independent prompt instructions to steer text generation at arbitrary points in time. The primary objective is to demonstrate that pre-trained language models can achieve precise, fine-grained control without degrading text quality or requiring massive computational resources for retraining.
To evaluate this approach, the researchers converted a standard pre-trained GPT-2 Large model into the NRP architecture through a multi-phase, self-supervised training routine that kept the original base model weights frozen. They evaluated performance across standard benchmark tasks and introduced Contextualized CommonGen (C2GEN), a new evaluation dataset requiring models to incorporate specified target words while maintaining relevance to a preceding three-sentence context. Performance was assessed through automatic linguistic metrics alongside rigorous human evaluations measuring common sense and contextual consistency.
Across multiple experiments, NRP substantially outperformed existing baselines in control precision and language fluency. In free-text generation on CommonGen, NRP achieved a 98.4% target word inclusion rate compared to 72.2% for standard prompting and 13.3% for plug-and-play decoding methods, while also maintaining lower perplexity scores that indicate superior fluency. On the contextualized C2GEN benchmark, standard prompting degraded sharply to a 57.0% inclusion rate as it lost track of instructions, whereas NRP maintained a 96.9% inclusion rate in free text and an 81.0% inclusion rate in single-sentence generation. Human evaluations confirmed that NRP maintained solid common sense adherence and high contextual relevance across all settings, while alternative approaches either collapsed in text quality or failed to incorporate the required content.
These findings demonstrate that organizations do not need to choose between strict factual adherence and natural language fluency. By isolating prompt signals so they do not degrade the internal states of the base model over time, NRP enables modular, fine-grained control over length, vocabulary, and context. This significantly mitigates hallucination and omission risks in production applications while keeping computational and deployment costs low, as the underlying language model requires no fine-tuning.
Decision-makers exploring automated text workflows should consider modular non-residual architectures over traditional rigid decoding heuristics or standard long-prompt templates. Organizations should pilot NRP in targeted editorial and document-generation settings where factual completeness is mandatory. Future development should expand the framework into multi-task prompt libraries and evaluate larger modern base models with varied positional encoding schemes to broaden applicability.
While the results provide high confidence in the architecture's effectiveness, the study's scope was focused primarily on word inclusion and sentence length controls using GPT-2-scale architectures. Human evaluations showed low inter-annotator agreement on nuanced common-sense judgments, and models explicitly trained on knowledge graphs still retained a slight edge in domain-specific logic. Deployments in mission-critical environments should therefore retain human-in-the-loop validation while further domain-specific prompt tuning is conducted.
- Paper: CTRL: A Conditional Transformer Language Model for Controllable Generation, Nitish Shirish Keskar et al. (2019). This paper establishes the foundational concept of conditional text generation using control codes in causal language models, providing the core problem framing that non-residual prompting seeks to improve.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). This work introduces continuous soft prompt tuning with frozen model weights, laying the technical foundation for parameter-efficient adaptation methods leveraged by non-residual prompting.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). This paper demonstrates using continuous prompt embeddings optimized via neural network encoders to steer frozen language models, directly informing modular encoder-based prompt architectures.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). This survey provides a comprehensive taxonomy of discrete and continuous prompting mechanisms, framing the trade-offs between prompt-based steering and decoding interventions.
- Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). This chapter diagnoses degeneration issues in neural text generation and establishes standard sampling and perplexity evaluation protocols used to assess fluency in controlled generation.
- Paper: Guiding Large Language Models via Directional Stimulus Prompting, Zekun Li et al. (2023). This paper extends controllable text generation by using an auxiliary policy model to generate instance-specific directional stimulus prompts for fine-grained guidance without updating the target model.
- Paper: Diffusion-LM Improves Controllable Text Generation, Xiang Lisa Li et al. (2022). This work explores continuous diffusion models to overcome the rigid left-to-right control limitations of autoregressive models, offering an alternative paradigm for fine-grained syntactic and semantic control.
- Paper: Composable Text Controls in Latent Space with ODEs, Guangyi Liu et al. (2023). This study advances fine-grained, composable text generation by steering continuous latent representations via ordinary differential equations rather than sequential prompt injection.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). This work investigates iterative self-refinement to enforce complex constraints during generation, contrasting inference-time iterative correction with continuous prompt steering.
