Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment

Rui YangXiaoman PanFeng LuoShuang QiuHan ZhongDong YuJianshu Chen

article2024ICML138 citations

Presents Rewards-in-Context, an efficient alignment method that conditions foundation models on multiple reward values via supervised fine-tuning, achieving Pareto-optimal multi-objective adaptation at inference time using only ten percent of the compute required by reinforcement learning baselines.

Listen

Aligning foundation models with human preferences is critical for creating safe and helpful artificial intelligence systems. However, human preferences are inherently diverse, multidimensional, and frequently conflicting, such as the tension between providing helpful answers and avoiding harmful content. Existing methods that rely on reinforcement learning across multiple objectives are computationally expensive, training-unstable, and struggle to dynamically adapt to varying user preferences without training separate models or interpolating model weights.

The article demonstrates and evaluates a novel framework called Rewards-in-Context (RiC). The primary objective is to achieve scalable, Pareto-optimal alignment across multiple conflicting objectives using only standard supervised fine-tuning on a single foundation model, while allowing real-time adjustment to user preferences during deployment.

To achieve this, the article employs a three-stage approach. In the offline phase, the system labels training prompts with normalized scores from multiple reward models and performs supervised fine-tuning. In the online phase, the model generates additional candidate responses targeting high-performing frontier trade-offs, which are filtered via multi-objective rejection sampling to expand optimal training data. During the inference phase, the system uses an analytically derived preference-to-reward mathematical mapping that dynamically adjusts the conditioned reward values in the input prompt based on user-specified priorities. The framework was evaluated across language generation tasks (using a 7-billion-parameter LLaMA 2 model on dialogue and summarization datasets) and text-to-image tasks (using Stable Diffusion).

Key findings show that the proposed approach consistently outperforms traditional multi-objective reinforcement learning, weight interpolation baselines, and direct preference optimization methods by achieving a superior empirical Pareto frontier. In terms of resource efficiency, the method required only about 10% of the GPU hours utilized by standard multi-objective reinforcement learning baselines and roughly 25% of weight-averaging baselines. Furthermore, the framework retained foundational model capabilities—such as factual faithfulness in summarization—that baseline reinforcement learning methods degraded due to catastrophic forgetting. Tests also confirmed that the framework successfully scales across multiple model sizes (1B to 7B parameters) and generalizes to three simultaneous objectives as well as multimodal image generation.

These findings indicate that organizations can significantly cut compute expenditures and operational complexity by replacing unstable reinforcement learning pipelines with multi-reward supervised conditioning. The ability to dynamically tune model behavior at inference time reduces the risk of deploying rigid models and allows fine-grained customization for safety, compliance, and user preferences without retraining.

Senior leaders should consider piloting reward-conditioned fine-tuning frameworks for multi-attribute alignment initiatives to decrease development cycle times and compute costs. However, teams must implement rigorous input filtering and guardrails, as conditioning models on arbitrary reward prompts introduces the risk that malicious actors could intentionally demand harmful outputs. Additionally, practitioners should exercise caution when applying this method to objectives that are strongly positively correlated, as the model may over-focus on a single dimension. Further work should explore context-aware dynamic preference mappings and test performance on larger-scale foundational models.

Cover for Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment

Abstract

We consider the problem of multi-objective alignment of foundation models with human preferences, which is a critical step towards helpful and harmless AI systems. However, it is generally costly and unstable to fine-tune large foundation models using reinforcement learning (RL), and the multi-dimensionality, heterogeneity, and conflicting nature of human preferences further complicate the alignment process. In this paper, we introduce Rewards-in-Context (RiC), which conditions the response of a foundation model on multiple rewards in its prompt context and applies supervised fine-tuning for alignment. The salient features of RiC are simplicity and adaptivity, as it only requires supervised fine-tuning of a single foundation model and supports dynamic adjustment for user preferences during inference time. Inspired by the analytical solution of an abstracted convex optimization problem, our dynamic inference-time adjustment method approaches the Pareto-optimal solution for multiple objectives. Empirical evidence demonstrates the efficacy of our method in aligning both Large Language Models (LLMs) and diffusion models to accommodate diverse rewards with only around 10% GPU hours compared with multi-objective RL baseline.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 3. RiC Algorithm
  • 3.1. Offline Training
  • 3.2. Online Training
  • 3.3. Inference Stage
  • 3.4. Determining the Preference-to-Reward Mapping
  • 4. Experiments
  • 4.1. Experimental Setups
  • 4.2. Experiments on Text Generation
  • 4.3. Ablations
  • 4.4. Text-to-image Generation
  • 5. Related Work
  • 6. Conclusion
  • Acknowledgment
  • Impact Statement
  • References
  • A. Determining Preference-to-Reward Mapping
  • A.1. Reduction of Optimization Problem (4)
  • A.2. Proof of Theorem 3.1
  • B. Implementation Details
  • Basic information
  • C. Additional Experiments
  • C.1. Text-to-image Generation: diffusion models with RiC
  • C.2. Scaling with Model Size
  • C.3. Comparison with MODPO
  • C.4. Additional Ablation Study
  • C.5. Positively Correlated Rewards
  • C.6. Examples in the Text Generation Tasks
  • D. Additional Related Work

Knowls

  1. Knowl 1 — Rewards-in-Context (RiC) Alignment Framework

    model/method

    Rewards-in-Context (RiC) is a framework for multi-objective alignment of foundation models (such as large language models and diffusion models) with diverse human preferences. Instead of running computationally heavy reinforcement learning across multiple preference weightings, RiC trains a single base model using supervised fine-tuning (SFT) conditioned on reward values prepended to the prompt context, coupled with dynamic preference adjustment at inference time.

    The RiC framework operates in three stages:

    1. Offline Training Stage: Prompt-response pairs from a training dataset are scored using NN distinct reward models. The obtained reward scores are normalized and appended to the input prompts using delimiter markers. The foundation model is then trained via standard autoregressive supervised fine-tuning to predict the response conditioned on the prompt and its reward values.

    2. Online Training Stage: The offline-trained model generates synthetic responses for training prompts conditioned on target reward combinations sampled near the empirical Pareto frontier. Generated pairs are filtered using Multi-Objective Rejection Sampling (MORS) and combined with offline demonstration data to further fine-tune the model via conditional SFT.

    3. Inference Stage: Given user-specified preference weights w=[w1,…,wN]w = [w_1, \dots, w_N] on the unit simplex, RiC uses a closed-form preference-to-reward mapping f(w)=[f1(w),…,fN(w)]f(w) = [f_1(w), \dots, f_N(w)] to dynamically compute conditioning reward values [R1,…,RN][R_1, \dots, R_N]. These values are inserted into the prompt context to steer the foundation model toward the corresponding Pareto-optimal response.

  2. Knowl 2 — Closed-Form Preference-to-Reward Mapping via Regularized Convex Optimization

    theoretical result

    Let w=[w1,…,wN]Tw = [w_1, \dots, w_N]^T be a user preference vector on the NN-simplex satisfying wi≥0w_i \ge 0 for all i∈{1,…,N}i \in \{1, \dots, N\} and ∑i=1Nwi=1\sum_{i=1}^N w_i = 1. Let ζ={ζi}i=1N\zeta = \{\zeta_i\}_{i=1}^N denote the permutation of indices sorting ww in descending order, such that wζ1≥wζ2≥⋯≥wζNw_{\zeta_1} \ge w_{\zeta_2} \ge \dots \ge w_{\zeta_N}.

    To determine the prompt conditioning rewards R=[R1,…,RN]TR = [R_1, \dots, R_N]^T, the mapping is formulated as the constrained optimization problem:

    max⁡{Ri}i=1N∑i=1Nwi⋅ϕi(Ri)s.t.[ϕ1(R1),…,ϕN(RN)]∈Cregλ,1≥ϕζ1(Rζ1)≥⋯≥ϕζN(RζN)≥0\max_{\{R_i\}_{i=1}^N} \sum_{i=1}^N w_i \cdot \phi_i(R_i) \quad \text{s.t.} \quad [\phi_1(R_1), \dots, \phi_N(R_N)] \in \mathcal{C}_{\text{reg}}^\lambda, \quad 1 \ge \phi_{\zeta_1}(R_{\zeta_1}) \ge \dots \ge \phi_{\zeta_N}(R_{\zeta_N}) \ge 0

    where ϕi:[rimin⁡,rimax⁡]→[0,1]\phi_i : [r_i^{\min}, r_i^{\max}] \to [0, 1] maps desired reward values to normalized expected scores, and the regularization set is defined for p>1p > 1 (including p=∞p = \infty) by:

    Cregλ:={x∈RN:∥λ⊙x∥p≤1,  λ≽1}\mathcal{C}_{\text{reg}}^\lambda := \left\{ x \in \mathbb{R}^N : \|\lambda \odot x\|_p \le 1, \; \lambda \succcurlyeq 1 \right\}

    where λ=[λ1,…,λN]T\lambda = [\lambda_1, \dots, \lambda_N]^T are reweighting hyperparameters with λi≥1\lambda_i \ge 1, and ⊙\odot is the element-wise product. Assuming λ\lambda satisfies wζ11/p/λζ1≥wζ21/p/λζ2≥⋯≥wζN1/p/λζN≥0w_{\zeta_1}^{1/p}/\lambda_{\zeta_1} \ge w_{\zeta_2}^{1/p}/\lambda_{\zeta_2} \ge \dots \ge w_{\zeta_N}^{1/p}/\lambda_{\zeta_N} \ge 0, the optimal solution is given by Ri∗=ϕi−1(zi∗)R_i^* = \phi_i^{-1}(z_i^*) where:

    1. For 1<p<∞1 < p < \infty:

    zi∗=wi1p−1λipp−1[∑j=1N(wjλj)pp−1]−1pz_i^* = \frac{w_i^{\frac{1}{p-1}}}{\lambda_i^{\frac{p}{p-1}}} \left[ \sum_{j=1}^N \left( \frac{w_j}{\lambda_j} \right)^{\frac{p}{p-1}} \right]^{-\frac{1}{p}}

    1. For p=∞p = \infty:

    zi∗=1λiz_i^* = \frac{1}{\lambda_i}

  3. Knowl 3 — Practical Dynamic Preference-to-Reward Formulations

    model/method

    In practical implementations of Rewards-in-Context (RiC), the reward normalization function is approximated by a linear scaling ϕi(x)=(x−rimin⁡)/(rimax⁡−rimin⁡)\phi_i(x) = (x - r_i^{\min}) / (r_i^{\max} - r_i^{\min}), giving the inverse mapping ϕi−1(y)=y(rimax⁡−rimin⁡)+rimin⁡\phi_i^{-1}(y) = y(r_i^{\max} - r_i^{\min}) + r_i^{\min}, where rimin⁡r_i^{\min} and rimax⁡r_i^{\max} represent the empirical minimum and maximum values of the ii-th reward dimension in the training dataset.

    Using the closed-form solution Ri∗=ϕi−1(zi∗)R_i^* = \phi_i^{-1}(z_i^*), specific practical mapping functions fi(w)f_i(w) are defined:

    1. Finite ℓp\ell_p-norm (1<p<∞1 < p < \infty, typically p=2p=2) with uniform scaling λ1=⋯=λN=1\lambda_1 = \dots = \lambda_N = 1:

    fi(wi)=(rimax⁡−rimin⁡)wi1p−1(∑j=1Nwjpp−1)−1p+rimin⁡f_i(w_i) = (r_i^{\max} - r_i^{\min}) w_i^{\frac{1}{p-1}} \left( \sum_{j=1}^N w_j^{\frac{p}{p-1}} \right)^{-\frac{1}{p}} + r_i^{\min}

    1. ℓ∞\ell_\infty-norm with adaptive scaling: Setting λi=1\lambda_i = 1 when wi≥1/Nw_i \ge 1/N and λi=1/(Nwi)\lambda_i = 1 / (N w_i) when wi<1/Nw_i < 1/N provides a piecewise linear mapping that assigns the maximum reward to high-preference dimensions while dynamically scaling remaining rewards:

    fi(wi)={rimax⁡,wi≥1NNwi(rimax⁡−rimin⁡)+rimin⁡,wi<1Nf_i(w_i) = \begin{cases} r_i^{\max}, & w_i \ge \frac{1}{N} \\ N w_i (r_i^{\max} - r_i^{\min}) + r_i^{\min}, & w_i < \frac{1}{N} \end{cases}

    Both p=2p=2 and p=∞p=\infty variants empirically establish effective Pareto frontiers, outperforming naive unregularized linear mapping fi(wi)=wi(rimax⁡−rimin⁡)+rimin⁡f_i(w_i) = w_i(r_i^{\max} - r_i^{\min}) + r_i^{\min}.

  4. Knowl 4 — Offline Multi-Reward Conditional Supervised Fine-Tuning

    model/method

    The offline training phase in RiC equips a foundation model policy πθ\pi_\theta to ground its outputs in explicit numerical reward specifications across NN objectives without requiring reinforcement learning or preference pairs.

    Given an uncurated dataset of prompt-response pairs {(x,y)}∼D\{(x, y)\} \sim \mathcal{D}, each sample is evaluated with NN reward models to obtain scores r1(x,y),…,rN(x,y)r_1(x, y), \dots, r_N(x, y). Each reward dimension is normalized across the dataset to have a mean of zero and a standard deviation of one, and rounded to one decimal place. The prompt is then augmented by prepending reward tags:

    x' = \text{"### Input:"} \{x\} \; \langle R_1 \rangle \; r_1 \dots \langle R_N \rangle \; r_N

    where ⟨R1⟩,…,⟨RN⟩\langle R_1 \rangle, \dots, \langle R_N \rangle are designated textual delimiter tokens. The policy parameters θ\theta are optimized using standard autoregressive cross-entropy loss over the target response tokens y=(y1,…,y∣y∣)y = (y_1, \dots, y_{|y|}):

    Loffline(θ)=−∑t=1∣y∣log⁡πθ(yt∣x,r1(x,y),…,rN(x,y),y<t)\mathcal{L}_{\text{offline}}(\theta) = - \sum_{t=1}^{|y|} \log \pi_\theta\left(y_t \mid x, r_1(x, y), \dots, r_N(x, y), y_{<t}\right)

    This objective trains the model on both positive and negative responses across diverse reward trade-offs without explicit data filtering.

  5. Knowl 5 — Multi-Objective Rejection Sampling (MORS) and Online Data Augmentation

    algorithm

    Because offline datasets typically exhibit sample scarcity along the empirical Pareto boundary, RiC utilizes an online augmentation stage where the offline-trained policy generates candidate responses near the Pareto frontier, filtered by Multi-Objective Rejection Sampling (MORS).

    Input: Offline policy πθ\pi_\theta, training prompts Dprompts\mathcal{D}_{\text{prompts}}, reward models {ri}i=1N\{r_i\}_{i=1}^N, quantile thresholds {r~i}i=1N\{\tilde{r}_i\}_{i=1}^N (e.g., 0.7-quantile), online iterations KK, sample size per iteration SS
    Output: Fine-tuned multi-objective policy πθ\pi_\theta
    for iteration k=1,…,Kk = 1, \dots, K do
        Initialize online buffer Bonline←∅\mathcal{B}_{\text{online}} \leftarrow \emptyset
        for step s=1,…,Ss = 1, \dots, S do
            Sample prompt x∼Dpromptsx \sim \mathcal{D}_{\text{prompts}}
            Select random target dimension j∈{1,…,N}j \in \{1, \dots, N\}
            Set desired target rewards: Ri←rimax⁡R_i \leftarrow r_i^{\max} for all i≠ji \ne j, and sample Rj∈[rjmin⁡,rjmax⁡]R_j \in [r_j^{\min}, r_j^{\max}]
            Construct prompt x' \leftarrow \text{"### Input:"} \{x\} \; \langle R_1 \rangle R_1 \dots \langle R_N \rangle R_N
            Generate response y∼πθ(⋅∣x′)y \sim \pi_\theta(\cdot \mid x')
            Compute actual reward scores: r^i←ri(x,y)\hat{r}_i \leftarrow r_i(x, y) for i=1,…,Ni = 1, \dots, N
            if not (r^1≤r~1∧r^2≤r~2∧⋯∧r^N≤r~N)(\hat{r}_1 \le \tilde{r}_1 \land \hat{r}_2 \le \tilde{r}_2 \land \dots \land \hat{r}_N \le \tilde{r}_N) then
                Add relabeled sample (x,r^1,…,r^N,y)(x, \hat{r}_1, \dots, \hat{r}_N, y) to Bonline\mathcal{B}_{\text{online}}
            end if
        end for
        Combine Bonline\mathcal{B}_{\text{online}} with offline regularization data (e.g., 0.5 ratio of original dataset samples passing MORS)
        Fine-tune πθ\pi_\theta on combined data via conditional SFT loss
    end for
    return πθ\pi_\theta
  6. Knowl 6 — Multi-Objective Alignment Performance and Preservation of Base Capabilities

    empirical result

    In evaluations on Large Language Models (LLaMA-2 7B) across two-objective and three-objective text generation tasks, RiC achieves empirical Pareto frontiers that consistently dominate Multi-Objective Reinforcement Learning from Human Feedback (MORLHF with PPO), Rewarded Soups (weight interpolation), and Multi-Objective Direct Preference Optimization (MODPO):

    1. Helpful Assistant Task (HH-RLHF): For 'harmless' vs. 'helpful' and 'humor' vs. 'helpful', RiC forms an outer Pareto boundary compared to MORLHF and Rewarded Soups. When scaled to three simultaneous objectives ('harmless', 'helpful', 'humor'), RiC yields the most balanced multi-reward performance under uniform preference weights w=[1/3,1/3,1/3]w = [1/3, 1/3, 1/3] and spans a wider 3D non-dominated frontier.

    2. Reddit Summary Task (openai/summarize_from_feedback): Optimizing preference rewards ('pref1' or 'pref2') against summary 'faithfulness' reveals a catastrophic forgetting vulnerability in baseline RLHF pipelines. Base LLaMA-2 achieves high faithfulness scores (>0.5>0.5) but poor preference scores (<−2.5<-2.5). Standard SFT and MORLHF/Rewarded Soups/MODPO optimize preference rewards at the cost of severely degrading faithfulness (dropping to negative values). In contrast, RiC restores base model capability via multi-reward conditioning, simultaneously achieving high preference scores (>−0.5>-0.5) and preserving high faithfulness (>0.5>0.5).

  7. Knowl 7 — Computational Efficiency Comparison of Alignment Algorithms

    data/table

    On the Helpful Assistant task aligning two reward models (N=2N=2) across M=5M=5 evaluated preference vectors using LLaMA-2 7B on NVIDIA Tesla V100 32 GB GPUs, RiC requires only a fraction of the computational time demanded by reinforcement learning baselines:

    Method GPU hours
    MORLHF 1,477.1
    Rewarded Soups 622.7
    RiC w/o online 54.0
    RiC w/ online iter1 103.6
    RiC w/ online iter2 153.2

    RiC with two online iterations achieves superior alignment using only 10.4%10.4\% of the GPU hours required by MORLHF (153.2153.2 hours vs. 1,477.11,477.1 hours) and 24.6%24.6\% of the GPU hours required by Rewarded Soups (622.7622.7 hours). Pure offline RiC requires only 3.7%3.7\% (54.054.0 hours) of MORLHF training time.

  8. Knowl 8 — Multi-Objective Alignment for Text-to-Image Latent Diffusion Models

    empirical result

    RiC extends directly to text-to-image foundation models. When applied to Stable Diffusion v1.5 (1B parameters) trained on a 120k subset of LAION-5B with two conflicting reward objectives—image aesthetic quality (LAION aesthetic predictor v2) and image compressibility (file size minimization under JPEG compression):

    1. Offline multi-reward conditional fine-tuning embeds numerical aesthetic and compressibility scores into text prompts.
    2. Adjusting the inference-time preference weight w1∈[0,1]w_1 \in [0, 1] across test captions (COCO Karpathy test set) produces a continuous Pareto trade-off: setting high aesthetic preference (w1=1.0w_1 = 1.0) yields visually detailed, high-scoring images at larger file sizes, while decreasing w1→0.0w_1 \to 0.0 smoothly reduces image complexity and file size (increasing compressibility score from −9.96-9.96 to +1.37+1.37) at the expense of aesthetic score (decreasing from −0.05-0.05 to −1.72-1.72).
    3. The base Stable Diffusion v1.5 model cannot dynamically steer between these attributes, whereas RiC spans the full non-dominated frontier.
  9. Knowl 9 — Model Size Scaling of Multi-Objective Alignment

    empirical result

    When evaluating RiC across varying base foundation model sizes on the Helpful Assistant 'harmless' vs. 'helpful' alignment task:

    • TinyLlama (1B parameters)
    • OpenLLaMA (3B parameters)
    • LLaMA-2 (7B parameters)

    All model scales successfully learn multi-reward conditioning and respond dynamically to preference changes. Furthermore, the empirical Pareto frontier scales monotonically with model capacity: larger base models consistently attain superior frontiers that dominate those of smaller models across all preference weightings.

  10. Knowl 10 — Limitation in Disentangling Strongly Positively Correlated Rewards

    limitation

    RiC struggles to form a well-defined empirical Pareto frontier when the target reward functions are strongly positively correlated rather than in trade-off conflict.

    When evaluated on pairs of positively correlated rewards:

    1. 'deberta-v1' vs. 'deberta-v2' reward models on the HH-RLHF dataset (Pearson's r=0.60r = 0.60)
    2. 'Verbosity' (character count) vs. 'Complexity' (Flesch-Kincaid grade level) on the Reddit Summary dataset (Pearson's r=0.35r = 0.35)

    the supervised policy tends to capture the shared underlying correlation, primarily optimizing one dominant reward dimension while neglecting fine-grained control of the other. The resulting outputs exhibit monotonic shifts along one reward dimension rather than spanning a distinct, controllable multi-objective trade-off surface.

Coverage note — Omitted specific Hugging Face model checkpoint URLs, specific LoRA rank/alpha hyperparameter tables for baselines, and repetitive qualitative text generation prompt-response transcripts, as they are standard configuration choices and illustrative examples.

References

  1. 1.Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W. Hindsight experience replay. Advances in neural information processing systems, 30, 2017.
  2. 2.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  3. 3.Barrett, L. and Narayanan, S. Learning all optimal policies with multiple criteria. In Proceedings of the 25th international conference on Machine learning, pp. 41–47, 2008.
  4. 4.Black, K., Janner, M., Du, Y., Kostrikov, I., and Levine, S. Training diffusion models with reinforcement learning. arXiv preprint arXiv:2305.13301, 2023.
  5. 5.Boyd, S. P. and Vandenberghe, L. Convex optimization. Cambridge university press, 2004.
  6. 6.Brandfonbrener, D., Bietti, A., Buckman, J., Laroche, R., and Bruna, J. When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35: 1542–1553, 2022.
  7. 7.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  8. 8.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 9650–9660, 2021.
  9. 9.Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34:15084–15097, 2021.
  10. 10.Chen, S., Hou, Y., Cui, Y., Che, W., Liu, T., and Yu, X. Recall and learn: Fine-tuning deep pretrained language models with less forgetting. arXiv preprint arXiv:2004.12651, 2020.
  11. 11.Chen, X., Zhong, H., Yang, Z., Wang, Z., and Wang, L. Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation. In International Conference on Machine Learning, pp. 3773–3793. PMLR, 2022.
  12. 12.Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  13. 13.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  14. 14.Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T. Raft: Reward ranked fine-tuning for generative foundation model alignment. arXiv preprint arXiv:2304.06767, 2023a.
  15. 15.Dong, Y., Wang, Z., Sreedhar, M. N., Wu, X., and Kuchaiev, O. Steerlm: Attribute conditioned sft as an (user-steerable) alternative to rlhf. arXiv preprint arXiv:2310.05344, 2023b.
  16. 16.Emmons, S., Eysenbach, B., Kostrikov, I., and Levine, S. Rvs: What is essential for offline rl via supervised learning? arXiv preprint arXiv:2112.10751, 2021.
  17. 17.Fang, M., Zhou, T., Du, Y., Han, L., and Zhang, Z. Curriculum-guided hindsight experience replay. Advances in neural information processing systems, 32, 2019.
  18. 18.Ghosh, D., Gupta, A., Reddy, A., Fu, J., Devin, C., Eysenbach, B., and Levine, S. Learning to reach goals via iterated supervised learning. arXiv preprint arXiv:1912.06088, 2019.
  19. 19.Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998, 2023.
  20. 20.Hayes, C. F., Radulescu, R., Bargiacchi, E., Källström, J., Macfarlane, M., Reymond, M., Verstraeten, T., Zintgraf, L. M., Dazeley, R., Heintz, F., et al. A practical guide to multi-objective reinforcement learning and planning. Autonomous Agents and Multi-Agent Systems, 36(1):26, 2022.
  21. 21.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  22. 22.Hu, J., Tao, L., Yang, J., and Zhou, C. Aligning language models with offline reinforcement learning from human feedback. arXiv preprint arXiv:2308.12050, 2023.
  23. 23.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  24. 24.Korbak, T., Elsahar, H., Kruszewski, G., and Dymetman, M. On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting. Advances in Neural Information Processing Systems, 35:16203–16220, 2022.
  25. 25.Kumar, A., Peng, X. B., and Levine, S. Reward-conditioned policies. arXiv preprint arXiv:1912.13465, 2019.
  26. 26.Langley, P. Crafting papers on machine learning. In Langley, P. (ed.), Proceedings of the 17th International Conference on Machine Learning (ICML 2000), pp. 1207–1216, Stanford, CA, 2000. Morgan Kaufmann.
  27. 27.Li, A., Pinto, L., and Abbeel, P. Generalized hindsight for reinforcement learning. Advances in neural information processing systems, 33:7754–7767, 2020a.
  28. 28.Li, K., Zhang, T., and Wang, R. Deep reinforcement learning for multiobjective optimization. IEEE transactions on cybernetics, 51(6):3103–3114, 2020b.
  29. 29.Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J. Statistical rejection sampling improves preference optimization. arXiv preprint arXiv:2309.06657, 2023.
  30. 30.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  31. 31.Lu, X., Welleck, S., Hessel, J., Jiang, L., Qin, L., West, P., Ammanabrolu, P., and Choi, Y. Quark: Controllable text generation with reinforced unlearning. Advances in neural information processing systems, 35:27591–27609, 2022.
  32. 32.Luo, F., Xiang, J., Zhang, J., Han, X., and Yang, W. Image super-resolution via latent diffusion: A sampling-space mixture of experts and frequency-augmented decoder approach. arXiv preprint arXiv:2310.12004, 2023.
  33. 33.Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, Z. D., Tang, Y., Geist, M., Mesnard, T., Michi, A., et al. Nash learning from human feedback. arXiv preprint arXiv:2312.00886, 2023.
  34. 34.Murray, N., Marchesotti, L., and Perronnin, F. Ava: A large-scale database for aesthetic visual analysis. In 2012 IEEE conference on computer vision and pattern recognition, pp. 2408–2415. IEEE, 2012.
  35. 35.Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  36. 36.OpenAI, R. Gpt-4 technical report. arxiv 2303.08774. View in Article, 2:13, 2023.
  37. 37.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  38. 38.Pacchiano, A., Saha, A., and Lee, J. Dueling rl: reinforcement learning with trajectory preferences. arXiv preprint arXiv:2111.04850, 2021.
  39. 39.Peng, B., Li, C., He, P., Galley, M., and Gao, J. Instruction tuning with gpt-4. arXiv preprint arXiv:2304.03277, 2023.
  40. 40.Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023.
  41. 41.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training. 2018.
  42. 42.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019.
  43. 43.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023.
  44. 44.Rame, A., Couairon, G., Shukor, M., Dancette, C., Gaya, J.-B., Soulier, L., and Cord, M. Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards. arXiv preprint arXiv:2306.04488, 2023.
  45. 45.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2):3, 2022.
  46. 46.Ramnath, S., Joshi, B., Hallinan, S., Lu, X., Li, L. H., Chan, A., Hessel, J., Choi, Y., and Ren, X. Tailoring self-rationalizers with multi-reward distillation. arXiv preprint arXiv:2311.02805, 2023.
  47. 47.Roijers, D. M., Vamplew, P., Whiteson, S., and Dazeley, R. A survey of multi-objective sequential decision-making. Journal of Artificial Intelligence Research, 48:67–113, 2013.
  48. 48.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022.
  49. 49.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems, 35: 36479–36494, 2022.
  50. 50.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35: 25278–25294, 2022.
  51. 51.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  52. 52.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33: 3008–3021, 2020.
  53. 53.Sun, H. Reinforcement learning in the era of llms: What is essential? what is needed? an rl perspective on rlhf, prompting, and beyond. arXiv preprint arXiv:2310.06147, 2023.
  54. 54.Sun, H., Li, Z., Liu, X., Zhou, B., and Lin, D. Policy continuation with hindsight inverse dynamics. Advances in Neural Information Processing Systems, 32, 2019.
  55. 55.Sun, H., Hüyük, A., and van der Schaar, M. Query-dependent prompt evaluation and optimization with offline inverse rl. In The Twelfth International Conference on Learning Representations, 2023.
  56. 56.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  57. 57.Vamplew, P., Dazeley, R., Foale, C., Firmin, S., and Mummery, J. Human-aligned artificial intelligence is a multi-objective problem. Ethics and Information Technology, 20:27–40, 2018.
  58. 58.Van Moffaert, K. and Nowé, A. Multi-objective reinforcement learning using sets of pareto dominating policies. The Journal of Machine Learning Research, 15(1):3483–3512, 2014.
  59. 59.von Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., and Wolf, T. Diffusers: State-of-the-art diffusion models, 2022.
  60. 60.von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., and Huang, S. Trl: Transformer reinforcement learning. https://github.com/huggingface/trl, 2020.
  61. 61.Wang, C., Jiang, Y., Yang, C., Liu, H., and Chen, Y. Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints. arXiv preprint arXiv:2309.16240, 2023a.
  62. 62.Wang, Z., Dong, Y., Zeng, J., Adams, V., Sreedhar, M. N., Egert, D., Delalleau, O., Scowcroft, J. P., Kant, N., Swope, A., et al. Helpsteer: Multi-attribute helpfulness dataset for steerlm. arXiv preprint arXiv:2311.09528, 2023b.
  63. 63.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pp. 38–45, 2020.
  64. 64.Wu, Z., Hu, Y., Shi, W., Dziri, N., Suhr, A., Ammanabrolu, P., Smith, N. A., Ostendorf, M., and Hajishirzi, H. Fine-grained human feedback gives better rewards for language model training. arXiv preprint arXiv:2306.01693, 2023.
  65. 65.Xiong, W., Dong, H., Ye, C., Zhong, H., Jiang, N., and Zhang, T. Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf. arXiv preprint arXiv:2312.11456, 2023.
  66. 66.Yang, R., Fang, M., Han, L., Du, Y., Luo, F., and Li, X. Mher: Model-based hindsight experience replay. arXiv preprint arXiv:2107.00306, 2021.
  67. 67.Yang, R., Lu, Y., Li, W., Sun, H., Fang, M., Du, Y., Li, X., Han, L., and Zhang, C. Rethinking goal-conditioned supervised learning and its connection to offline rl. arXiv preprint arXiv:2202.04478, 2022.
  68. 68.Yang, R., Yong, L., Ma, X., Hu, H., Zhang, C., and Zhang, T. What is essential for unseen goal generalization of offline goal-conditioned rl? In International Conference on Machine Learning, pp. 39543–39571. PMLR, 2023.
  69. 69.Yang, R., Zhong, H., Xu, J., Zhang, A., Zhang, C., Han, L., and Zhang, T. Towards robust offline reinforcement learning under diverse data corruption. In The Twelfth International Conference on Learning Representations, 2024.
  70. 70.Ye, H., Zhang, J., Liu, S., Han, X., and Yang, W. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv preprint arXiv:2308.06721, 2023.
  71. 71.Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F. Rrhf: Rank responses to align language models with human feedback without tears. arXiv preprint arXiv:2304.05302, 2023.
  72. 72.Zhang, R., Han, J., Zhou, A., Hu, X., Yan, S., Lu, P., Li, H., Gao, P., and Qiao, Y. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023a.
  73. 73.Zhang, T., Liu, F., Wong, J., Abbeel, P., and Gonzalez, J. E. The wisdom of hindsight makes language models better instruction followers. arXiv preprint arXiv:2302.05206, 2023b.
  74. 74.Zhou, Z., Liu, J., Yang, C., Shao, J., Liu, Y., Yue, X., Ouyang, W., and Qiao, Y. Beyond one-preference-for-all: Multi-objective direct preference optimization. arXiv preprint arXiv:2310.03708, 2023.
  75. 75.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Yang, R., et al. “Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment”. arXiv, 2024, http://arxiv.org/abs/2402.10207v6.
APA
Yang, R., Pan, X., Luo, F., Qiu, S., Zhong, H., Yu, D., & Chen, J. (2024). Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment. arXiv. http://arxiv.org/abs/2402.10207v6
Chicago
Yang, R., X. Pan, F. Luo, et al. 2024. “Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment”. arXiv. http://arxiv.org/abs/2402.10207v6.
Harvard
Yang, R. et al. (2024) “Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.10207v6.
Vancouver
1. Yang R, Pan X, Luo F, Qiu S, Zhong H, Yu D, Chen J (2024) Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment. arXiv

BibTeX

@article{yang2024rewards,
  title = {Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment},
  author = {Yang, Rui and Pan, Xiaoman and Luo, Feng and Qiu, Shuang and Zhong, Han and Yu, Dong and Chen, Jianshu},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.10207v6},
  eprint = {2402.10207}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/