Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
Yafu LiXuyang HuXiaoye QuLinjie LiYu Cheng
Introduces Test-Time Preference Optimization (TPO), an inference-time alignment method that uses iterative textual critiques to guide model outputs, enabling standard supervised models to surpass fine-tuned aligned models across safety and reasoning tasks without parameter updates.
Deploying large language models that remain consistently safe, helpful, and compliant with user preferences typically requires intensive training-time alignment, such as Reinforcement Learning from Human Feedback or Direct Preference Optimization. However, updating model weights via offline retraining is computationally expensive, time-consuming, and inflexible when adapting quickly to evolving requirements or shifting real-world data distributions. Consequently, organizations face significant operational friction and infrastructure costs when trying to maintain alignment across diverse deployment settings.
The article demonstrates that large language models can be successfully aligned with human preferences entirely during inference—on the fly—without updating internal model parameters. It introduces Test-Time Preference Optimization, a framework that leverages the model's inherent reasoning and instruction-following capabilities to iteratively interpret numerical reward model feedback, generate plain-language critiques and revision instructions, and refine its own outputs before returning the final response.
To evaluate this approach, the authors tested the method across several standard benchmarks covering instruction following, preference alignment, safety compliance, and mathematics using both unaligned base models and previously aligned models, such as Llama-3.1-70B and Mistral-Small-22B. Operationally, the framework samples multiple candidate responses per prompt, evaluates them with an external scoring model to identify the best and worst candidates, prompts the core model to produce a textual evaluation comparing the pair, generates actionable suggestions for improvement, and drafts refined candidate outputs over several iterations.
The results show that just two rounds of test-time refinement allow an unaligned model to match or outperform models aligned through full training-time optimization. On the Arena-Hard benchmark, the unaligned 70-billion-parameter model using this method achieved a win rate of 70.5%, surpassing even a 405-billion-parameter instruct model. Furthermore, applying the method to a compact 22-billion-parameter instruct model achieved a 53.4% length-controlled win rate on AlpacaEval 2, rivaling major commercial models like GPT-4-Turbo. In addition to higher benchmark accuracy, the process significantly reduced variance across generation scores, leading to greater output stability and safety. In terms of resource efficiency, the test-time refinement consumed around 9.3 PFLOPs per query—less than 0.01% of the compute required to train an aligned model offline.
These findings imply that organizations can bypass expensive and rigid post-training retraining pipelines for preference alignment, shifting compute expenditures from upfront capital-heavy training to flexible, per-query inference. This approach reduces time-to-deployment and enables rapid customization of model behavior to new safety or policy guidelines simply by updating the reward criteria. However, because each query requires multi-step candidate generation and critique loops, inference latency will increase, presenting a direct trade-off between real-time responsiveness and output quality.
Engineering and product leaders should consider deploying this test-time alignment framework in high-stakes, asynchronous, or complex query pipelines where quality, safety, and deep reasoning outweigh sub-second response requirements. Teams should run initial pilot benchmarks comparing search width and depth configurations to establish the ideal latency-accuracy balance for their specific application. Future work should focus on refining prompt protocols for specialized domain tasks and investigating whether smaller models can be fine-tuned to handle iterative critique processing more effectively.
Confidence in these findings is strong across capable medium-to-large models, supported by consistent gains across diverse benchmark domains. However, readers should note an important boundary condition: the method relies strictly on the underlying model's base instruction-following competence. When evaluated on smaller, weaker architectures (such as an 8-billion-parameter model), alignment degraded during iterative refinement because the model failed to reliably interpret and execute the textual critique instructions.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). This foundational study establishes the preference-model-and-RLHF pipeline that TPO replaces at inference time with textual feedback.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). DPO provides the preference-optimization framework and reward-model connection that clarifies TPO’s contrast with parameter-updating alignment.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). SELF-REFINE introduces iterative textual feedback and response revision, the core inference-time refinement pattern that TPO adapts to preference rewards.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). InstructGPT explains the RLHF alignment baseline whose inference-time alternative TPO evaluates.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). This RLHF work supplies the helpfulness-and-harmlessness preference-training setup that makes TPO’s test-time alternative intelligible.
- Paper: Deep reinforcement learning from human preferences, Paul F. Christiano et al. (2017). This early preference-learning paper explains how human comparisons become rewards for optimizing behavior, a premise behind TPO’s textual reward signals.
- Paper: Training-Free Group Relative Policy Optimization, Yuzheng Cai et al. (2025). Training-Free GRPO carries inference-only optimization forward by turning rollout comparisons into textual insights that guide later responses without weight updates.
- Paper: TTPO: Test-Time Policy Optimization, Aozhe Wang et al. (2026). TTPO extends test-time policy optimization to label-free math reasoning, using groups of sampled trajectories to derive selective learning signals.
