RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs
Afra Feyza AkyürekEkin AkyürekAshwin KalyanPeter ClarkDerry Tanti WijayaNiket Tandon
Presents RL4F, a multi-agent framework that trains a compact critique generator with reinforcement learning to produce natural language feedback that guides frozen, black-box language models like GPT-3 to correct their own output errors without requiring fine-tuning.
Large language models often make factual, logical, or structural errors when generating text. While models can refine their predictions when provided with natural language feedback, obtaining human-written critiques in real time is too costly and slow for operational use. Furthermore, many enterprise applications rely on closed, third-party black-box models that cannot be directly modified or fine-tuned. To address these constraints, the article introduces RL4F (Reinforcement Learning for Feedback Generation), a collaborative system where a smaller, dedicated critique model is trained to provide plain-language feedback that guides a fixed, frozen downstream task model to correct its own errors.
The framework pairs a relatively lightweight critique model (a 770-million-parameter T5 network) with a large, fixed language model (175-billion-parameter GPT-3). The critique model is first warm-started on example critiques and then optimized using Proximal Policy Optimization, a reinforcement learning method. Training directly rewards the critic when its generated feedback steers the large language model to produce a more accurate final answer. This setup was evaluated across three distinct domains: topic-based summarization, everyday action planning (the Interscript benchmark), and a synthetic word alphabetization task.
The evaluation demonstrates that the reinforced critique approach consistently outperforms standard supervised training, retrieval baselines, and self-refinement prompting. In action planning and summarization, RL4F achieved relative text quality and similarity improvements of up to 10% over alternative automated feedback methods, narrowing the gap toward upper-bound performance achieved with human feedback. In alphabetization, the framework raised exact match accuracy to 66.1%, outperforming supervised feedback (38.9%) by more than 27 absolute percentage points and full model fine-tuning (55.3%). Moreover, the critique model scaled effectively with parameter size and produced additional accuracy gains when applied iteratively across multiple rounds of feedback without corrupting previously correct answers.
These results provide a practical, cost-effective blueprint for improving the reliability of deployed AI systems. Instead of undertaking expensive full-model fine-tuning or paying for continuous human monitoring, organizations can deploy a compact, specialized critic model to act as a lightweight external adapter. This multi-agent strategy reduces operational computing expenses, preserves the general capabilities of the primary system, and maintains interpretable, human-readable communication between models.
Organizations seeking to improve output quality in automated workflows should consider adopting external, reward-tuned feedback models for high-stakes tasks such as multi-step planning and content summarization. Before full production deployment, teams should conduct domain-specific pilot testing to monitor for potential semantic drift—where the critic might invent non-standard phrasing—and establish fallback rules for ambiguous model outputs. Future development should focus on extending the architecture into an ensemble framework capable of integrating feedback from multiple automated models and human experts simultaneously.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). This foundational preference-optimization work introduces reward-model-guided reinforcement learning for language models, the training logic RL4F adapts to optimize a critic for a fixed downstream model.
- Paper: Learning to summarize from human feedback, Nisan Stiennon et al. (2020). Its human-feedback summarization pipeline establishes reward modeling and reinforcement learning for language generation, directly contextualizing RL4F’s use of downstream task performance to train feedback generation.
- Paper: Guiding Large Language Models via Directional Stimulus Prompting, Zekun Li et al. (2023). Directional Stimulus Prompting trains a small policy with reinforcement learning against a frozen black-box model’s task performance, a close precursor to RL4F’s feedback generator for an unmodifiable target.
- Paper: Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback, Yafu Li et al. (2025). It carries feedback-driven repair into test-time preference optimization, using critiques and iterative revisions to align model outputs without changing model weights.
- Paper: Self-Generated Critiques Boost Reward Modeling for Language Models, Yue Yu 0009 et al. (2025). It extends learned natural-language critique generation into reward modeling, jointly training critiques and preference scores to improve model evaluation and correction.
