Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Zixiang ChenYihe DengHuizhuo YuanKaixuan JiQuanquan Gu

article2024ICML513 citations

Proposes Self-Play Fine-Tuning (SPIN), a method that allows fine-tuned language models to continuously improve by competing against their past iterations, achieving superior performance to direct preference optimization without needing additional human labels or AI-generated feedback.

Listen

Modern large language models require extensive post-training alignment to generate helpful and accurate responses. Standard fine-tuning approaches, such as supervised fine-tuning and reinforcement learning from human or artificial intelligence feedback, face diminishing returns and depend on costly, labor-intensive datasets. Continued supervised training on existing demonstration data often leads to performance plateaus or model degradation. Consequently, organizations face increasing data acquisition costs to achieve incremental gains in model capabilities.

To address this challenge, the article introduces Self-Play Fine-Tuning (SPIN), a method designed to convert a weaker language model into a stronger one without requiring new human annotations, preference labels, or external reward models. The article evaluates how effectively a model can iteratively improve itself by treating fine-tuning as a two-player game against previous iterations of itself.

The authors conducted experimental evaluations using the open-source zephyr-7b-sft-full model (a 7-billion parameter model based on Mistral-7B) fine-tuned on subsets of the UltraChat dataset (50,000 to 100,000 samples). In this setup, the model from the prior iteration generates synthetic responses to existing prompts, and the updated model is trained to distinguish between these self-generated responses and the high-quality human demonstrations. Performance was evaluated across standard benchmarks, including the HuggingFace Open LLM Leaderboard (covering reasoning, truthfulness, and mathematics), MT-Bench, and Big-Bench tasks.

The investigation produced several key findings. First, self-play fine-tuning steadily boosted model accuracy across multiple iterations, raising the average benchmark score from 58.14% to 63.16%, with notable improvements exceeding 10% on mathematical problem solving (GSM8k) and truthfulness (TruthfulQA). Second, without introducing new human or external feedback, the model at its initial iteration achieved performance comparable to direct preference optimization trained on an additional 62,000 GPT-4 preference pairs, and it outperformed this baseline in subsequent iterations. Third, multi-turn iterative training proved necessary; simply training for additional epochs on the original supervised data or within a single iteration led to plateaued or degraded performance. Finally, self-play fine-tuning remained complementary to subsequent preference optimization, yielding an additional boost when combined with external preference datasets.

These findings indicate that organizations can extract significantly more performance from existing demonstration data while substantially reducing reliance on costly external data collection or proprietary supervisor models. The approach shortens development cycles and lowers annotation costs. Furthermore, the synthetic data generation introduces minimal computational overhead relative to training time, making the pipeline cost-effective and operationally efficient.

For teams maintaining fine-tuning pipelines, the article supports adopting iterative self-play between the initial supervised fine-tuning stage and any downstream preference optimization. Practitioners should plan for two to three self-play iterations, as gains diminish as the model approaches the demonstration data distribution. If additional preference data is available, it can be applied after self-play training to maximize performance.

The primary limitation of this approach is that the target data distribution remains fixed to the initial human demonstration dataset, which acts as a theoretical performance ceiling. Future work is required to explore dynamic target distributions that might enable models to surpass human-level demonstrations and to optimize synthetic data volume to further reduce computational requirements.

arXiv: 2401.01335
Cover for Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Abstract

Harnessing the power of human-annotated data through Supervised Fine-Tuning (SFT) is pivotal for advancing Large Language Models (LLMs). In this paper, we delve into the prospect of growing a strong LLM out of a weak one without the need for acquiring additional human-annotated data. We propose a new fine-tuning method called Self-Play fIne-tuNing (SPIN), which starts from a supervised fine-tuned model. At the heart of SPIN lies a self-play mechanism, where the LLM refines its capability by playing against instances of itself. More specifically, the LLM generates its own training data from its previous iterations, refining its policy by discerning these self-generated responses from those obtained from human-annotated data. Our method progressively elevates the LLM from a nascent model to a formidable one, unlocking the full potential of human-annotated demonstration data for SFT. Theoretically, we prove that the global optimum to the training objective function of our method is achieved only when the LLM policy aligns with the target data distribution. Empirically, we evaluate our method on several benchmark datasets including the HuggingFace Open LLM Leaderboard, MT-Bench, and datasets from Big-Bench. Our results show that SPIN can significantly improve the LLM's performance across a variety of benchmarks and even outperform models trained through direct preference optimization (DPO) supplemented with extra GPT-4 preference data. This sheds light on the promise of self-play, enabling the achievement of human-level performance in LLMs without the need for expert opponents. Codes are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Problem Setting and Preliminaries
  • 3.1 Supervised Fine-Tuning
  • 3.2 RL Fine-Tuning
  • 4 Method
  • 4.1 Self-Play Fine-Tuning (SPIN)
  • 4.2 Comparison between SPIN and DPO
  • 5 Theoretical Analysis
  • 6 Experiments
  • 6.1 Experiment Setup
  • 6.2 SPIN Effectively Improves Benchmark Performance
  • 6.3 Ablation Studies
  • 7 Conclusion and Discussion
  • References
  • A Further Related Work
  • B Experiments
  • B.1 Hyperparameters and Implementation Details
  • B.2 Training Overhead
  • B.3 Additional Experiment Result for SPIN+DPO
  • B.4 Further Experiment Results
  • B.5 Generation Examples
  • C Proof of Theorems in Section
  • C.1 Proof of Theorem
  • C.2 Proof Theorem

Knowls

  1. Knowl 1 — Self-Play Fine-Tuning (SPIN)

    model/method

    Self-Play Fine-Tuning (SPIN) is an alignment framework that converts a supervised fine-tuned (SFT) large language model (LLM) into a stronger model using exclusively the existing human-annotated SFT dataset SSFT={(xi,yi)}i=1N\mathcal{S}_{\text{SFT}} = \{(x_i, y_i)\}_{i=1}^N, without acquiring new human annotations, preference labels, or external AI feedback.

    SPIN formulates post-SFT alignment as a two-player game between two instances of the same model:

    1. Opponent Player: The model checkpoint from iteration tt, denoted by policy pθt(⋅∣x)p_{\theta_t}(\cdot|x), generates synthetic candidate responses y′y' for the prompts xx in the dataset.
    2. Main Player: A new model instance pθt+1(⋅∣x)p_{\theta_{t+1}}(\cdot|x) that is trained to discriminate between target human responses y∼pdata(⋅∣x)y \sim p_{\text{data}}(\cdot|x) and the self-generated responses y′∼pθt(⋅∣x)y' \sim p_{\theta_t}(\cdot|x) produced by the opponent.

    The main player acts as a discriminator optimizing an integral probability metric (IPM) value gap function ft+1(x,y)f_{t+1}(x, y). When the function class is parameterized directly by the LLM policy space as ft+1(x,y)=λlog⁡pθt+1(y∣x)pθt(y∣x)f_{t+1}(x, y) = \lambda \log \frac{p_{\theta_{t+1}}(y|x)}{p_{\theta_t}(y|x)} (where λ>0\lambda > 0 is a Kullback-Leibler regularization parameter), the opponent's best response to maximize ft+1f_{t+1} subject to a KL penalty against pθtp_{\theta_t} exactly coincides with pθt+1p_{\theta_{t+1}}. Across iterations t=0,1,…,T−1t = 0, 1, \dots, T-1, the model updates its parameters to make its generation distribution indistinguishable from pdatap_{\text{data}}.

  2. Knowl 2 — End-to-End SPIN Training Objective

    equation

    At iteration t+1t+1 of Self-Play Fine-Tuning (SPIN), given the fixed model parameters θt\theta_t from the previous iteration, the updated model parameters θt+1\theta_{t+1} are computed by minimizing the end-to-end objective LSPIN(θ,θt)\mathcal{L}_{\text{SPIN}}(\theta, \theta_t):

    LSPIN(θ,θt)=Ex∼q(⋅), y∼pdata(⋅∣x), y′∼pθt(⋅∣x)[ℓ(λlog⁡pθ(y∣x)pθt(y∣x)−λlog⁡pθ(y′∣x)pθt(y′∣x))]\mathcal{L}_{\text{SPIN}}(\theta, \theta_t) = \mathbb{E}_{x \sim q(\cdot), \, y \sim p_{\text{data}}(\cdot|x), \, y' \sim p_{\theta_t}(\cdot|x)} \left[ \ell \left( \lambda \log \frac{p_\theta(y|x)}{p_{\theta_t}(y|x)} - \lambda \log \frac{p_\theta(y'|x)}{p_{\theta_t}(y'|x)} \right) \right]

    where:

    • xx denotes the input prompt sampled from marginal prompt distribution q(⋅)q(\cdot).
    • yy is the ground-truth demonstration response from the target data distribution pdata(⋅∣x)p_{\text{data}}(\cdot|x).
    • y′y' is the synthetic response sampled autoregressively from the opponent policy pθt(⋅∣x)p_{\theta_t}(\cdot|x).
    • pθ(y∣x)=∏j=1mpθ(yj∣x,y<j)p_\theta(y|x) = \prod_{j=1}^m p_\theta(y_j|x, y_{<j}) is the conditional sequence probability evaluated under policy pθp_\theta.
    • λ>0\lambda > 0 is a regularization hyperparameter governing the step size and penalizing deviation from pθtp_{\theta_t}.
    • ℓ:R→R\ell: \mathbb{R} \to \mathbb{R} is a monotonically decreasing, convex loss function, commonly instantiated as the logistic loss ℓ(u)=log⁡(1+exp⁡(−u))\ell(u) = \log(1 + \exp(-u)).
  3. Knowl 3 — Self-Play Fine-Tuning Algorithm

    algorithm

    Self-Play Fine-Tuning (SPIN) operates iteratively over TT outer rounds. In each round, the current checkpoint pθtp_{\theta_t} generates synthetic completions for every prompt in the SFT dataset, and the model weights are updated by minimizing the empirical SPIN loss paired with ground-truth completions.

    Input: SFT dataset (xi,yi)i=1N{(x_i, y_i)}_{i=1}^N, initial fine-tuned model pθ0p_{\theta_0}, regularization parameter λ\lambda, total iterations TT, convex decreasing loss ℓ\ell
    Output: Aligned model parameters θT\theta_T
    for t=0,…,T−1t = 0, \ldots, T-1 do
        for i=1,…,Ni = 1, \ldots, N do
            Sample synthetic response yi′∼pθt(⋅∣xi)y'_i \sim p_{\theta_t}(\cdot|x_i)
        end for
        Compute updated parameters:
        θt+1←arg⁡min⁡θ∈Θ∑i=1Nℓ(λlog⁡pθ(yi∣xi)pθt(yi∣xi)−λlog⁡pθ(yi′∣xi)pθt(yi′∣xi))\theta_{t+1} \leftarrow \arg\min_{\theta \in \Theta} \sum_{i=1}^N \ell \left( \lambda \log \frac{p_\theta(y_i|x_i)}{p_{\theta_t}(y_i|x_i)} - \lambda \log \frac{p_\theta(y'_i|x_i)}{p_{\theta_t}(y'_i|x_i)} \right)
    end for
    return θT\theta_T

    In multi-iteration deployments (t≥1t \ge 1), the synthetic dataset can be formed by combining the NN synthetic completions generated at iteration tt with the NN synthetic completions from iteration t−1t-1, resulting in a synthetic buffer of size 2N2N.

  4. Knowl 4 — Global Optimality and Convergence of SPIN

    theoretical result

    Let the loss function ℓ:R→R\ell: \mathbb{R} \to \mathbb{R} be convex, monotonically non-increasing (ℓ′(u)≤0\ell'(u) \le 0 for all uu), and satisfy ℓ′(0)<0\ell'(0) < 0. Assume there exists some parameter θ∈Θ\theta \in \Theta such that pθ(⋅∣x)=pdata(⋅∣x)p_\theta(\cdot|x) = p_{\text{data}}(\cdot|x) almost surely.

    Then the following necessary and sufficient conditions characterize the global minimum of the objective LSPIN(θ,θt)\mathcal{L}_{\text{SPIN}}(\theta, \theta_t):

    • Sufficiency: If pθt(⋅∣x)=pdata(⋅∣x)p_{\theta_t}(\cdot|x) = p_{\text{data}}(\cdot|x), then θt\theta_t is the global minimum of LSPIN(θ,θt)\mathcal{L}_{\text{SPIN}}(\theta, \theta_t) for any λ≥0\lambda \ge 0.
    • Necessity: If pθt(⋅∣x)≠pdata(⋅∣x)p_{\theta_t}(\cdot|x) \ne p_{\text{data}}(\cdot|x), there exists a choice of λ>0\lambda > 0 such that θt\theta_t is not the global minimum of LSPIN(θ,θt)\mathcal{L}_{\text{SPIN}}(\theta, \theta_t).

    Consequently, the self-play update process reaches a stationary point if and only if the LLM's predictive distribution exactly aligns with the ground-truth target data distribution pdatap_{\text{data}}.

  5. Knowl 5 — Exact Policy Distribution Update Under Logistic Loss

    theoretical result

    When the loss function in SPIN is the logistic loss ℓ(u)=log⁡(1+exp⁡(−u))\ell(u) = \log(1 + \exp(-u)) and the function pθt(y∣x)(pdata(y∣x)pθt(y∣x))1/λp_{\theta_t}(y|x) \left( \frac{p_{\text{data}}(y|x)}{p_{\theta_t}(y|x)} \right)^{1/\lambda} lies within the model space {pθ(y∣x)∣θ∈Θ}\{p_\theta(y|x) \mid \theta \in \Theta\}, the global minimizer θt+1\theta_{t+1} of LSPIN(θ,θt)\mathcal{L}_{\text{SPIN}}(\theta, \theta_t) updates the opponent policy distribution according to:

    pθt+1(y∣x)∝pθt(y∣x)(pdata(y∣x)pθt(y∣x))1/λp_{\theta_{t+1}}(y|x) \propto p_{\theta_t}(y|x) \left( \frac{p_{\text{data}}(y|x)}{p_{\theta_t}(y|x)} \right)^{1/\lambda}

    This closed-form relation implies:

    • If pθt(y∣x)<pdata(y∣x)p_{\theta_t}(y|x) < p_{\text{data}}(y|x), the density ratio exceeds 1, increasing the probability pθt+1(y∣x)p_{\theta_{t+1}}(y|x).
    • If pθt(y∣x)>pdata(y∣x)p_{\theta_t}(y|x) > p_{\text{data}}(y|x), the density ratio is less than 1, decreasing the probability pθt+1(y∣x)p_{\theta_{t+1}}(y|x).
    • The shift magnitude is governed by 1/λ1/\lambda: smaller λ\lambda induces larger distribution shifts, whereas larger λ\lambda stabilizes learning as pθtp_{\theta_t} approaches pdatap_{\text{data}}.
  6. Knowl 6 — Loss Function Conditions for SPIN

    assumption

    The theoretical convergence guarantees of SPIN require the loss function ℓ:R→R\ell: \mathbb{R} \to \mathbb{R} to satisfy three conditions:

    1. Monotonic decrease: ℓ′(t)≤0\ell'(t) \le 0 for all t∈Rt \in \mathbb{R}.
    2. Strict negativity at origin: ℓ′(0)<0\ell'(0) < 0.
    3. Convexity: ℓ(t)\ell(t) is convex over R\mathbb{R}.

    Valid loss functions meeting these conditions include the logistic loss ℓ(t)=log⁡(1+exp⁡(−t))\ell(t) = \log(1 + \exp(-t)), the exponential loss ℓ(t)=exp⁡(−t)\ell(t) = \exp(-t), the hinge loss ℓ(t)=max⁡(0,1−t)\ell(t) = \max(0, 1 - t), and the correlation loss ℓ(t)=1−t\ell(t) = 1 - t.

  7. Knowl 7 — Open LLM Leaderboard Evaluation Across SPIN Iterations

    data/table

    SPIN was evaluated across iterations 0 through 3 starting from zephyr-7b-sft-full (Mistral-7B fine-tuned on UltraChat200k) using 50k prompts sampled from UltraChat200k on the HuggingFace Open LLM Leaderboard benchmark tasks (ARC Challenge 25-shot acc_norm, TruthfulQA 0-shot mc2, Winogrande 5-shot acc, GSM8k 5-shot acc, HellaSwag 10-shot acc_norm, MMLU 5-shot acc).

    Model ARC TruthfulQA Winogrande GSM8k HellaSwag MMLU Average
    zephyr-7b-sft-full 60.41 43.73 74.19 26.76 82.85 60.92 58.14
    SPIN iteration 0 63.40 49.18 72.69 35.10 84.38 60.03 60.80 (+2.66)
    SPIN iteration 1 65.19 55.17 72.30 35.78 84.96 59.34 62.12 (+1.32)
    SPIN iteration 2 65.96 54.91 73.56 38.06 85.41 59.93 62.97 (+0.85)
    SPIN iteration 3 65.87 54.90 73.72 38.97 85.54 59.99 63.16 (+0.19)

    The evaluation demonstrates monotonic improvement in average accuracy from 58.14% to 63.16% (+5.02% overall), driven by substantial gains in TruthfulQA (+11.17%) and GSM8k (+12.21%). Incremental gains diminish over successive iterations (+2.66% →\to +1.32% →\to +0.85% →\to +0.19%), matching the theoretical convergence behavior towards the limiting data distribution.

  8. Knowl 8 — SPIN Evaluation on MT-Bench and Big-Bench Tasks

    data/table

    To evaluate multi-turn dialog and out-of-distribution reasoning, SPIN was assessed on MT-Bench (score out of 10 evaluated with GPT-4 judge), Big-Bench Hard subtasks (Causal Judgment, Formal Fallacies, Sports Understanding under few-shot chain-of-thought accuracy), and OpenBookQA (1-shot acc_norm).

    Model MT-Bench BB-Causal BB-Formal BB-Sports OpenBookQA
    zephyr-7b-sft-full 5.94 56.15 49.6 96.0 45.4
    SPIN iteration 0 6.46 (+0.52) 57.75 51.6 95.2 46.8
    SPIN iteration 1 6.65 (+0.19) 58.82 51.2 95.2 47.2
    SPIN iteration 2 6.78 (+0.13) 59.36 51.2 94.4 47.6

    SPIN increases MT-Bench performance from 5.94 to 6.78 across iterations, exceeding vicuna-13b-v1.5 (6.57). General reasoning metrics improve or remain steady without exhibiting task degradation.

  9. Knowl 9 — Iterative Self-Play Outperforms Extended SFT and Multi-Epoch Training

    empirical result

    Ablation experiments show that SPIN's iterative generation and training mechanism breaks the performance barriers of repeated SFT and extended single-iteration training:

    1. Failure of continued SFT: Applying an additional epoch of standard SFT on zephyr-7b-sft-full using UltraChat200k decreases average Open LLM Leaderboard accuracy from 58.14% down to 57.23% (ARC drops from 60.41% to 57.76%, GSM8k drops from 26.76% to 25.85%). Fine-tuning Mistral-7B with SFT over 3 full epochs peaks at 59.82% at epoch 2 and declines to 59.27% at epoch 3.
    2. Necessity of multi-iteration updates: Training SPIN at iteration 0 for up to 5 epochs achieves most gains within the first 2 epochs (~60.80% average accuracy) and plateaus below 61.0%. It cannot achieve the 62.12% score unlocked by transitioning to iteration 1 (which generates new synthetic responses from the iteration 0 model).
    3. Data scaling at iteration 0: Evaluating SPIN on prompt subsets of sizes 14k, 26k, and 50k for 1 epoch shows monotonic average score scaling from 59.55% to 60.16% to 60.83%.
  10. Knowl 10 — Comparison and Combination of SPIN with Direct Preference Optimization

    empirical result

    SPIN was compared with Direct Preference Optimization (DPO) on the Open LLM Leaderboard:

    • zephyr-7b-dpo-full (or zephyr-7b-beta), trained from zephyr-7b-sft-full using 62k pairwise preference pairs annotated with GPT-4 from UltraFeedback, achieves an average leaderboard score of 61.31%.
    • SPIN, utilizing exclusively 50k prompts from the existing SFT dataset without external annotations, reaches 60.80% at iteration 0, 62.12% at iteration 1 (surpassing DPO), and 63.16% at iteration 3.

    Furthermore, applying 2 epochs of DPO on the 62k UltraFeedback dataset starting from the SPIN iteration 3 checkpoint achieves an average score of 64.05% (+0.89% over SPIN iteration 3), with TruthfulQA rising from 54.90% to 60.07% (+5.17%) and Winogrande rising from 73.72% to 78.06% (+4.34%), showing that SPIN and preference-based RL fine-tuning are complementary.

  11. Knowl 11 — Implementation and Training Setup for SPIN

    experimental setup

    The experimental configuration for evaluating SPIN on zephyr-7b-sft-full (Mistral-7B base) is as follows:

    • Dataset: 50,000 prompts randomly sampled from the first round of conversations in UltraChat200k.
    • Synthetic data: 50k model responses generated at iteration 0; at iterations 1, 2, and 3, 50k newly generated responses are concatenated with the 50k responses from the previous iteration to form 100k synthetic examples.
    • Optimization: RMSProp optimizer without weight decay, batch size of 64, 10% linear warmup, bfloat16 precision, maximum sequence length of 2048 tokens.
    • Learning rate and schedule: Peak learning rate of 5×10−75 \times 10^{-7} for iterations 0 and 1, decayed to 1×10−71 \times 10^{-7} for iterations 2 and 3. Models are trained for 2 epochs per iteration.
    • Regularization parameter: β=0.1\beta = 0.1 for iterations 0, 1, and 2, increased to β=5.0\beta = 5.0 at iteration 3.
    • Computation: Implemented with DeepSpeed ZeRO-3 and FlashAttention-2 on 8 ×\times NVIDIA A100 (80GB) GPUs. Synthetic response generation takes 1.45 hours (6.69s per 64 examples), and fine-tuning takes 4.32 hours for iteration 0 (50k samples) and 8.64 hours for iterations 1–3 (100k samples).
  12. Knowl 12 — Limitations of Fixed Target Distribution and Synthetic Generation Cost

    limitation

    SPIN exhibits two primary limitations:

    1. Bounded by target human demonstration distribution: Theoretical guarantees demonstrate convergence to the underlying data distribution pdatap_{\text{data}}. Because pdatap_{\text{data}} is fixed to the human-annotated SFT corpus, the model performance is constrained by this target distribution ceiling and cannot autonomously attain super-human capabilities without dynamic or evolving target distributions.
    2. Inference overhead of synthetic generation: Each self-play iteration requires generating autoregressive responses across the dataset (e.g. 1.45 hours for 50,000 sequences on 8 ×\times A100 GPUs), adding computational cost prior to model optimization.

Coverage note — No substantial contributed material was omitted; the extracted knowls comprehensively capture the SPIN formulation, objective function, algorithm, theoretical guarantees, empirical results across benchmarks and iterations, ablation analyses, combination with DPO, implementation details, and limitations.

References

  1. 1.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  2. 2.Anthony, T., Tian, Z., and Barber, D. Thinking fast and slow with deep learning and tree search. Advances in neural information processing systems, 30, 2017.
  3. 3.Arjovsky, M., Chintala, S., and Bottou, L. Wasserstein generative adversarial networks. In International conference on machine learning, pp. 214–223. PMLR, 2017.
  4. 4.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  5. 5.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022a.
  6. 6.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022b.
  7. 7.Bansal, T., Pachocki, J., Sidor, S., Sutskever, I., and Mordatch, I. Emergent complexity via multi-agent competition. In International Conference on Learning Representations, 2018.
  8. 8.Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T. Open llm leaderboard, 2023.
  9. 9.bench authors, B. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856.
  10. 10.Bengio, Y., Louradour, J., Collobert, R., and Weston, J. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pp. 41–48, 2009.
  11. 11.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  12. 12.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  13. 13.Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. 2023.
  14. 14.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  15. 15.Cheng, P., Yang, Y., Li, J., Dai, Y., and Du, N. Adversarial preference optimization, 2023.
  16. 16.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023.
  17. 17.Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  18. 18.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  19. 19.Cirik, V., Hovy, E., and Morency, L.-P. Visualizing and understanding curriculum learning for long short-term memory networks. arXiv preprint arXiv:1611.06204, 2016.
  20. 20.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  21. 21.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  22. 22.Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M. Ultrafeedback: Boosting language models with high-quality feedback, 2023.
  23. 23.Dao, T. Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv preprint arXiv:2307.08691, 2023.
  24. 24.Deng, Y., Zhang, W., Chen, Z., and Gu, Q. Rephrase and respond: Let large language models ask better questions for themselves. arXiv preprint arXiv:2311.04205, 2023.
  25. 25.Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B. Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:2305.14233, 2023.
  26. 26.Frei, S., Zou, D., Chen, Z., and Gu, Q. Self-training converts weak learners to strong learners in mixture models. In International Conference on Artificial Intelligence and Statistics, pp. 8003–8021. PMLR, 2022.
  27. 27.Freund, Y. Boosting a weak learning algorithm by majority. Information and computation, 121(2):256–285, 1995.
  28. 28.Freund, Y. and Schapire, R. E. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of computer and system sciences, 55(1):119–139, 1997.
  29. 29.Gao, L., Schulman, J., and Hilton, J. Scaling laws for reward model overoptimization. In International Conference on Machine Learning, pp. 10835–10866. PMLR, 2023a.
  30. 30.Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, 12 2023b.
  31. 31.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial nets. Advances in neural information processing systems, 27, 2014.
  32. 32.Grandvalet, Y. and Bengio, Y. Semi-supervised learning by entropy minimization. Advances in neural information processing systems, 17, 2004.
  33. 33.Gugger, S., Debut, L., Wolf, T., Schmid, P., Mueller, Z., Mangrulkar, S., Sun, M., and Bossan, B. Accelerate: Training and inference at scale made simple, efficient and adaptable., 2022.
  34. 34.Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C. Improved training of wasserstein gans. Advances in neural information processing systems, 30, 2017.
  35. 35.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  36. 36.Hernandez-Leal, P., Kartal, B., and Taylor, M. E. Is multiagent deep reinforcement learning the answer or the question? a brief survey. learning, 21:22, 2018.
  37. 37.Hinton, G., Srivastava, N., and Swersky, K. Neural networks for machine learning lecture 6a overview of mini-batch gradient descent. Cited on, 14(8):2, 2012.
  38. 38.Ho, J. and Ermon, S. Generative adversarial imitation learning. Advances in neural information processing systems, 29, 2016.
  39. 39.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  40. 40.Jolicoeur-Martineau, A. The relativistic discriminator: a key element missing from standard gan. arXiv preprint arXiv:1807.00734, 2018.
  41. 41.Josifoski, M., Sakota, M., Peyrard, M., and West, R. Exploiting asymmetry for synthetic training data generation: Synthie and the case of information extraction. arXiv preprint arXiv:2303.04132, 2023.
  42. 42.Kearns, M. and Valiant, L. Cryptographic limitations on learning boolean formulae and finite automata. Journal of the ACM (JACM), 41(1):67–95, 1994.
  43. 43.Kou, Y., Chen, Z., Cao, Y., and Gu, Q. How does semi-supervised learning with pseudo-labelers work? a case study. In The Eleventh International Conference on Learning Representations, 2022.
  44. 44.Kumar, M., Packer, B., and Koller, D. Self-paced learning for latent variable models. Advances in neural information processing systems, 23, 2010.
  45. 45.Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T. A unified game-theoretic approach to multiagent reinforcement learning. Advances in neural information processing systems, 30, 2017.
  46. 46.Lee, D.-H. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In ICML Challenges in Representation Learning Workshop, 2013.
  47. 47.Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267, 2023.
  48. 48.Lee, Y. J. and Grauman, K. Learning the easy things first: Self-paced visual category discovery. In CVPR 2011, pp. 1721–1728. IEEE, 2011.
  49. 49.Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al. Solving quantitative reasoning problems with language models. Advances in Neural Information Processing Systems, 35:3843–3857, 2022.
  50. 50.Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022.
  51. 51.Li, Y., Bubeck, S., Eldan, R., Giorno, A. D., Gunasekar, S., and Lee, Y. T. Textbooks are all you need ii: phi-1.5 technical report, 2023.
  52. 52.Lin, S., Hilton, J., and Evans, O. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958, 2021.
  53. 53.Liu, B., Bubeck, S., Eldan, R., Kulkarni, J., Li, Y., Nguyen, A., Ward, R., and Zhang, Y. Tinygsm: achieving> 80% on gsm8k with small language models. arXiv preprint arXiv:2312.09241, 2023.
  54. 54.Liu, C., He, S., Liu, K., Zhao, J., et al. Curriculum learning for natural answer generation. In IJCAI, pp. 4223–4229, 2018.
  55. 55.Liu, F., Ge, S., and Wu, X. Competence-based multimodal curriculum learning for medical report generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 3001–3012, 2021.
  56. 56.Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. arXiv preprint arXiv:2308.09583, 2023.
  57. 57.Mao, X., Li, Q., Xie, H., Lau, R. Y., Wang, Z., and Paul Smolley, S. Least squares generative adversarial networks. In Proceedings of the IEEE international conference on computer vision, pp. 2794–2802, 2017.
  58. 58.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2381–2391, 2018.
  59. 59.Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H. Cross-task generalization via natural language crowdsourcing instructions. arXiv preprint arXiv:2104.08773, 2021.
  60. 60.Mroueh, Y. and Sercu, T. Fisher gan. Advances in neural information processing systems, 30, 2017.
  61. 61.Müller, A. Integral probability metrics and their generating classes of functions. Advances in applied probability, 29 (2):429–443, 1997.
  62. 62.Muller, P., Omidshafiei, S., Rowland, M., Tuyls, K., Perolat, J., Liu, S., Hennes, D., Marris, L., Lanctot, M., Hughes, E., et al. A generalized training approach for multiagent learning. arXiv preprint arXiv:1909.12823, 2019.
  63. 63.OpenAI. Gpt-4 technical report, 2023.
  64. 64.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  65. 65.Prasad, A., Stengel-Eskin, E., and Bansal, M. Rephrase, augment, reason: Visual grounding of questions for vision-language models. arXiv preprint arXiv:2310.05861, 2023.
  66. 66.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  67. 67.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023.
  68. 68.Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–16. IEEE, 2020.
  69. 69.Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950, 2023.
  70. 70.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  71. 71.Samuel, A. L. Some studies in machine learning using the game of checkers. IBM Journal of research and development, 3(3):210–229, 1959.
  72. 72.Samuel, A. L. Some studies in machine learning using the game of checkers. IBM Journal of research and development, 44(1.2):206–226, 2000.
  73. 73.Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J. Self-critiquing models for assisting human evaluators. arXiv preprint arXiv:2206.05802, 2022.
  74. 74.Schapire, R. E. The strength of weak learnability. Machine learning, 5:197–227, 1990.
  75. 75.Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. Mastering chess and shogi by self-play with a general reinforcement learning algorithm. arXiv preprint arXiv:1712.01815, 2017a.
  76. 76.Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. Mastering the game of go without human knowledge. nature, 550(7676):354–359, 2017b.
  77. 77.Singh, A., Co-Reyes, J. D., Agarwal, R., Anand, A., Patil, P., Liu, P. J., Harrison, J., Lee, J., Xu, K., Parisi, A., et al. Beyond human data: Scaling self-training for problem-solving with language models. arXiv preprint arXiv:2312.06585, 2023.
  78. 78.Soviany, P., Ionescu, R. T., Rota, P., and Sebe, N. Curriculum learning: A survey. International Journal of Computer Vision, 130(6):1526–1565, 2022.
  79. 79.Spitkovsky, V. I., Alshawi, H., and Jurafsky, D. Baby steps: How “less is more” in unsupervised dependency parsing. In NIPS 2009 Workshop on Grammar Induction, Representation of Language and Language Learning, 2009.
  80. 80.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33: 3008–3021, 2020.
  81. 81.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: An instruction-following llama model, 2023.
  82. 82.Tesauro, G. et al. Temporal difference learning and td-gammon. Communications of the ACM, 38(3):58–68, 1995.
  83. 83.Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  84. 84.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  85. 85.Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., et al. Zephyr: Direct distillation of lm alignment. arXiv preprint arXiv:2310.16944, 2023a.
  86. 86.Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rush, A. M., and Wolf, T. The alignment handbook, 2023b.
  87. 87.Vapnik, V. The nature of statistical learning theory. Springer science & business media, 1999.
  88. 88.Victor, S., Albert, W., Colin, R., Stephen, B., Lintang, S., Zaid, A., Antoine, C., Arnaud, S., Arun, R., Manan, D., et al. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations, 2022.
  89. 89.Vinyals, O., Babuschkin, I., Chung, J., Mathieu, M., Jaderberg, M., Czarnecki, W., Dudzik, A., Huang, A., Georgiev, P., Powell, R., Ewalds, T., Horgan, D., Kroiss, M., Danihelka, I., Agapiou, J., Oh, J., Dalibard, V., Choi, D., Sifre, L., Sulsky, Y., Vezhnevets, S., Molloy, J., Cai, T., Budden, D., Paine, T., Gulcehre, C., Wang, Z., Pfaff, T., Pohlen, T., Yogatama, D., Cohen, J., McKinney, K., Smith, O., Schaul, T., Lillicrap, T., Apps, C., Kavukcuoglu, K., Hassabis, D., and Silver, D. AlphaStar: Mastering the Real-Time Strategy Game StarCraft II, 2019.
  90. 90.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35: 24824–24837, 2022.
  91. 91.Wu, J., Liang, Y., Akbari, H., Wang, Z., Yu, C., et al. Scaling multimodal pre-training via cross-modality gradient harmonization. Advances in Neural Information Processing Systems, 35:36161–36173, 2022.
  92. 92.Xu, J., Lee, A., Sukhbaatar, S., and Weston, J. Some things are more cringe than others: Preference optimization with the pairwise cringe loss. arXiv preprint arXiv:2312.16682, 2023.
  93. 93.Yang, Y., Singh, A. K., Elhoushi, M., Mahmoud, A., Tirumala, K., Gloeckle, F., Rozière, B., Wu, C.-J., Morcos, A. S., and Ardalani, N. Decoding data quality via synthetic corruptions: Embedding-guided pruning of code data. arXiv preprint arXiv:2312.02418, 2023.
  94. 94.Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284, 2023.
  95. 95.Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J. Self-rewarding language models. arXiv preprint arXiv:2401.10020, 2024.
  96. 96.Yuan, Z., Yuan, H., Li, C., Dong, G., Tan, C., and Zhou, C. Scaling relationship on learning mathematical reasoning with large language models. arXiv preprint arXiv:2308.01825, 2023.
  97. 97.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  98. 98.Zhang, D., Meng, D., Li, C., Jiang, L., Zhao, Q., and Han, J. A self-paced multiple-instance learning framework for co-saliency detection. In Proceedings of the IEEE international conference on computer vision, pp. 594–602, 2015.
  99. 99.Zhang, X., Kumar, G., Khayrallah, H., Murray, K., Gwinnup, J., Martindale, M. J., McNamee, P., Duh, K., and Carpuat, M. An empirical exploration of curriculum learning for neural machine translation. arXiv preprint arXiv:1811.00739, 2018.
  100. 100.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023.
  101. 101.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Chen, Z., et al. “Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models”. arXiv, 2024, http://arxiv.org/abs/2401.01335v3.
APA
Chen, Z., Deng, Y., Yuan, H., Ji, K., & Gu, Q. (2024). Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models. arXiv. http://arxiv.org/abs/2401.01335v3
Chicago
Chen, Z., Y. Deng, H. Yuan, K. Ji, and Q. Gu. 2024. “Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models”. arXiv. http://arxiv.org/abs/2401.01335v3.
Harvard
Chen, Z. et al. (2024) “Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.01335v3.
Vancouver
1. Chen Z, Deng Y, Yuan H, Ji K, Gu Q (2024) Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models. arXiv

BibTeX

@article{chen2024self,
  title = {Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models},
  author = {Chen, Zixiang and Deng, Yihe and Yuan, Huizhuo and Ji, Kaixuan and Gu, Quanquan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.01335v3},
  eprint = {2401.01335}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/