Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment
Rui YangXiaoman PanFeng LuoShuang QiuHan ZhongDong YuJianshu Chen
Presents Rewards-in-Context, an efficient alignment method that conditions foundation models on multiple reward values via supervised fine-tuning, achieving Pareto-optimal multi-objective adaptation at inference time using only ten percent of the compute required by reinforcement learning baselines.
Aligning foundation models with human preferences is critical for creating safe and helpful artificial intelligence systems. However, human preferences are inherently diverse, multidimensional, and frequently conflicting, such as the tension between providing helpful answers and avoiding harmful content. Existing methods that rely on reinforcement learning across multiple objectives are computationally expensive, training-unstable, and struggle to dynamically adapt to varying user preferences without training separate models or interpolating model weights.
The article demonstrates and evaluates a novel framework called Rewards-in-Context (RiC). The primary objective is to achieve scalable, Pareto-optimal alignment across multiple conflicting objectives using only standard supervised fine-tuning on a single foundation model, while allowing real-time adjustment to user preferences during deployment.
To achieve this, the article employs a three-stage approach. In the offline phase, the system labels training prompts with normalized scores from multiple reward models and performs supervised fine-tuning. In the online phase, the model generates additional candidate responses targeting high-performing frontier trade-offs, which are filtered via multi-objective rejection sampling to expand optimal training data. During the inference phase, the system uses an analytically derived preference-to-reward mathematical mapping that dynamically adjusts the conditioned reward values in the input prompt based on user-specified priorities. The framework was evaluated across language generation tasks (using a 7-billion-parameter LLaMA 2 model on dialogue and summarization datasets) and text-to-image tasks (using Stable Diffusion).
Key findings show that the proposed approach consistently outperforms traditional multi-objective reinforcement learning, weight interpolation baselines, and direct preference optimization methods by achieving a superior empirical Pareto frontier. In terms of resource efficiency, the method required only about 10% of the GPU hours utilized by standard multi-objective reinforcement learning baselines and roughly 25% of weight-averaging baselines. Furthermore, the framework retained foundational model capabilities—such as factual faithfulness in summarization—that baseline reinforcement learning methods degraded due to catastrophic forgetting. Tests also confirmed that the framework successfully scales across multiple model sizes (1B to 7B parameters) and generalizes to three simultaneous objectives as well as multimodal image generation.
These findings indicate that organizations can significantly cut compute expenditures and operational complexity by replacing unstable reinforcement learning pipelines with multi-reward supervised conditioning. The ability to dynamically tune model behavior at inference time reduces the risk of deploying rigid models and allows fine-grained customization for safety, compliance, and user preferences without retraining.
Senior leaders should consider piloting reward-conditioned fine-tuning frameworks for multi-attribute alignment initiatives to decrease development cycle times and compute costs. However, teams must implement rigorous input filtering and guardrails, as conditioning models on arbitrary reward prompts introduces the risk that malicious actors could intentionally demand harmful outputs. Additionally, practitioners should exercise caution when applying this method to objectives that are strongly positively correlated, as the model may over-focus on a single dimension. Further work should explore context-aware dynamic preference mappings and test performance on larger-scale foundational models.
- Paper: Multi-Task Learning as Multi-Objective Optimization, Ozan Sener et al. (2018). It formulates multi-task deep learning as multi-objective optimization to achieve Pareto optimality, providing foundational principles for the multi-objective Pareto-alignment problem addressed in Rewards-in-Context.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). It establishes the standard reinforcement learning from human feedback pipeline for aligning language models along multiple helpfulness and harmlessness criteria, framing the multi-objective alignment problem that Rewards-in-Context simplifies via supervised fine-tuning.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). It introduces direct preference optimization to bypass complex multi-stage reinforcement learning pipelines, motivating simpler alternatives to RL-based alignment.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). It establishes the foundational methodology of fitting reward models to human preference data to guide model fine-tuning.
- Paper: Scaling Laws for Reward Model Overoptimization, Leo Gao et al. (2023). It analyzes the degradation and overoptimization risks when policies are heavily fine-tuned against proxy reward models, highlighting the instability of standard RL fine-tuning.
- Paper: GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization, Shih-Yang Liu et al. (2026). It tackles multi-reward policy optimization by decoupling and normalizing reward signals during reinforcement learning, offering a complementary multi-objective alignment strategy.
- Paper: A General Framework for Inference-time Scaling and Steering of Diffusion Models, Raghav Singhal et al. (2025). It provides a general inference-time framework for steering diffusion models according to reward functions without retraining, extending dynamic reward-guided generation concepts.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). It presents a theoretical unification of direct preference alignment methods, contextualizing non-RL alignment paradigms.
- Paper: SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training, Tianzhe Chu et al. (2025). It investigates how post-training methods generalize compared to memorizing data, complementing findings on supervised fine-tuning and alignment efficiency across foundation models.
