Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Collin BurnsPavel IzmailovJan Hendrik KirchnerBowen BakerLeo GaoLeopold AschenbrennerYining ChenAdrien EcoffetManas JoglekarJan Leike

article2024ICML546 citations

Demonstrates that strong pretrained language models can generalize beyond imperfect supervision from weaker models, providing an empirical methodology and techniques to study how humans might align superhuman AI systems.

Listen

As artificial intelligence approaches superhuman performance, models will increasingly execute complex tasks that humans cannot reliably evaluate, such as reviewing millions of lines of intricate code. Standard alignment techniques like reinforcement learning from human feedback rely on human evaluators to steer model behavior, but this framework breaks down when supervisors are less capable than the models they oversee. The fundamental challenge is understanding whether weak supervision can reliably control and elicit the full capabilities of much stronger models.

The main objective of the article is to establish an empirical framework to study this dynamic by using weak language models to supervise significantly stronger pretrained models. Specifically, it evaluates how effectively weak supervision elicits latent capabilities across natural language processing tasks, chess puzzles, and human preference reward modeling, while testing targeted methods to close the gap between weak supervision and true capability.

The researchers conducted an extensive empirical study using models from the GPT-4 family across compute disparities spanning up to seven orders of magnitude. The setup involved three steps: training a small model on ground truth to act as a weak supervisor, finetuning a large student model solely on labels generated by that weak supervisor, and comparing the student's performance against a ceiling model trained on ground truth. The primary metric evaluated was the Performance Gap Recovered, which quantifies the fraction of the performance difference between the weak supervisor and the ground truth ceiling that the student achieves.

The findings establish that strong pretrained models naturally outperform their weak supervisors across nearly all tasks when naively finetuned. On natural language benchmarks, naive finetuning recovers roughly 20% to over 50% of the performance gap, with recovery improving as student compute grows. However, naive finetuning alone is insufficient to recover full model capabilities and performs poorly on complex tasks, recovering only around 10% of the gap on human preference reward modeling. Importantly, targeted methods significantly enhance recovery: introducing an auxiliary confidence loss term increases gap recovery on language tasks to nearly 80%, bootstrapping through intermediate model sizes prevents performance plateaus on chess puzzles, and unsupervised generative finetuning raises reward modeling gap recovery by 10% to 20%.

These results indicate that current standard alignment protocols will likely scale poorly to superhuman systems if applied naively, posing safety and reliability risks. However, the findings also demonstrate that weak-to-strong generalization is empirically tractable today. Because strong pretrained models already contain latent task representations, weak supervision acts to elicit existing knowledge rather than teach new skills, allowing students to avoid imitating many supervisor mistakes.

To build toward reliable alignment, researchers should develop scalable oversight methods that enforce internal model consistency, investigate better early stopping criteria to prevent student overfitting to weak errors, and extend these techniques to generative workflows. Further analysis is necessary to determine how reward models trained via weak supervision perform when subjected to strong reinforcement learning optimization pressure.

These findings should be interpreted with caution due to several experimental limitations. Current language models are not explicitly trained to imitate human supervisors, meaning future superhuman models might mimic human errors more readily than observed here. Additionally, some benchmark capabilities may have been present during pretraining, whereas future superhuman tasks may rely on purely latent knowledge. Confidence in the viability of the weak-to-strong framework is high, but additional research on non-imitative losses and diverse error structures is required before applying these techniques in high-stakes deployment environments.

Cover for Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Abstract

Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of humans to supervise model behavior—for example, to evaluate whether a model faithfully followed instructions or generated safe outputs. However, future superhuman models will behave in complex ways too difficult for humans to reliably evaluate; humans will only be able to weakly supervise superhuman models. We study an analogy to this problem: can weak model supervision elicit the full capabilities of a much stronger model? We test this using a range of pretrained language models in the GPT-4 family on natural language processing (NLP), chess, and reward modeling tasks. We find that when we naively finetune strong pretrained models on labels generated by a weak model, they consistently perform better than their weak supervisors, a phenomenon we call weak-to-strong generalization. However, we are still far from recovering the full capabilities of strong models with naive finetuning alone, suggesting that techniques like RLHF may scale poorly to superhuman models without further work. We find that simple methods can often significantly improve weak-to-strong generalization: for example, when finetuning GPT-4 with a GPT-2-level supervisor and an auxiliary confidence loss, we can recover close to GPT-3.5-level performance on NLP tasks. Our results suggest that it is feasible to make empirical progress today on a fundamental challenge of aligning superhuman models.

Table of Contents

  • 1. Introduction
  • 2. Methodology
  • 3. Main Results
  • 3.1. Tasks
  • 3.2. Naive baseline
  • 3.3. Improving Generalization is Tractable
  • 3.3.1. CONFIDENCE LOSS
  • 3.3.2. BOOTSTRAPPING
  • 3.3.3. UNSUPERVISED FINETUNING
  • 4. Overview of Additional Results
  • 4.1. Imitation
  • 4.2. Salience
  • 4.3. Other experiments
  • 5. Related Work
  • 6. Discussion
  • 6.1. Remaining disanalogies
  • 6.2. Conclusion
  • Impact Statement
  • Acknowledgements
  • References
  • A. Understanding Weak-to-Strong Generalization
  • A.1. Understanding imitation
  • A.1.1. OVERFITTING TO WEAK SUPERVISION
  • A.1.2. STUDENT-SUPERVISOR AGREEMENT
  • A.1.3. INVERSE SCALING FOR IMITATING THE SUPERVISOR
  • A.2. Saliency in the strong model representations
  • A.2.1. ELICITING STRONG MODEL KNOWLEDGE WITH PROMPTING
  • A.2.2. FINETUNING ON WEAK SUPERVISION TO INCREASE CONCEPT SALIENCY
  • B. Future Work
  • B.1. Concrete Problems: Analogous Setups
  • B.2. Concrete Problems: Scalable Methods
  • B.3. Concrete Problems: Scientific Understanding
  • C. Further experimental details
  • C.1. NLP Tasks
  • C.2. Chess Puzzles
  • C.3. ChatGPT Reward Modeling
  • C.4. Auxiliary Confidence Loss
  • D. Additional results on methods
  • E. Easy-to-hard generalization
  • E.1. Chess puzzles
  • E.2. NLP tasks: difficulty thresholding
  • E.3. GPT-4 predicted difficulty
  • F. Other weak-to-strong settings
  • F.1. Self-supervised vision models
  • F.2. Linear probing
  • G. The effects of weak label structure
  • G.1. Synthetic experiments on simulation difficulty
  • G.2. Different weak error structure means different generalization
  • G.3. Making imitation trivial
  • H. How should we empirically study aligning superhuman models, methodologically?
  • I. How weak-to-strong generalization fits into alignment
  • I.1. High-level plan
  • I.2. Eliciting key alignment-relevant capabilities with weak-to-strong generalization
  • I.3. Alignment plan assumptions

Knowls

  1. Knowl 1 — Performance Gap Recovered Metric

    equation

    The Performance Gap Recovered (PGR) metric quantifies the degree to which a strong student model trained with weak supervision recovers the capability gap between the weak supervisor and a strong model trained on ground-truth supervision:

    PGR=Pw2s−PweakPstrong ceiling−Pweak\text{PGR} = \frac{P_{\text{w2s}} - P_{\text{weak}}}{P_{\text{strong ceiling}} - P_{\text{weak}}}

    where:

    • PweakP_{\text{weak}} is the test performance of the weak supervisor model trained directly on ground-truth labels.
    • Pw2sP_{\text{w2s}} is the weak-to-strong performance, measured as the test performance of the strong student model finetuned on labels generated by the weak supervisor.
    • Pstrong ceilingP_{\text{strong ceiling}} is the ceiling test performance of the strong student model finetuned directly on ground-truth labels.

    A PGR=1\text{PGR} = 1 indicates that the weak-supervised student achieves performance identical to ground-truth finetuning (perfect weak-to-strong generalization), PGR=0\text{PGR} = 0 indicates that the strong student performs no better than the weak supervisor, and PGR<0\text{PGR} < 0 indicates performance degradation below the weak supervisor.

  2. Knowl 2 — Weak-to-Strong Generalization Experimental Setup

    experimental setup

    The weak-to-strong generalization experimental framework evaluates whether strong pretrained models can elicit their full capabilities when supervised only by weaker models. Pretrained base language models from the GPT-4 family spanning 7 orders of magnitude (OOMs) of pretraining compute are evaluated across three primary domains:

    1. Natural Language Processing (NLP) Benchmarks: 22 popular classification datasets (including ethics, commonsense reasoning, natural language inference, and sentiment analysis) formatted as binary classification tasks with balanced classes. The weak supervisor outputs soft probability labels on training examples.
    2. Chess Puzzles: Predicting the optimal first move from chess positions extracted from lichess.org. Weak supervision is generated via greedy decoding (T=0T=0) of the weak model.
    3. ChatGPT Reward Modeling (RM): Pairwise comparison prediction between assistant responses on human-assistant dialogs using proprietary ChatGPT reward modeling data. Weak labels correspond to weak model preference probabilities: yw=σ(Mw(d,c2)−Mw(d,c1))y_w = \sigma(M_w(d, c_2) - M_w(d, c_1)), where Mw(d,c)M_w(d, c) is the predicted reward logit for completion cc given dialog context dd, and σ\sigma is the sigmoid function.

    Weak supervisors are trained on the first half of a dataset, and their predictions on the held-out second half serve as training labels for the strong student.

  3. Knowl 3 — Baseline Performance and Scaling of Naive Weak-to-Strong Generalization

    empirical result

    Naively finetuning strong pretrained models on labels generated by weak supervisors consistently yields positive weak-to-strong generalization across tasks, but with significant variation across task domains and compute gaps:

    • NLP Classification Benchmarks: Naive finetuning recovers more than 20% to over 50% of the performance gap (PGR >0.20> 0.20 to >0.50> 0.50). Furthermore, PGR exhibits positive scaling: for a fixed weak supervisor, larger strong student models generally achieve higher PGR.
    • Chess Puzzles: Naive finetuning achieves moderate generalization for small compute disparities (PGR exceeding 40%), but exhibits negative scaling where PGR decreases as the student model compute increases relative to a fixed weak supervisor.
    • ChatGPT Reward Modeling: Naive generalization is poor; the student model achieves a PGR of roughly 10% to 20% across most compute disparities, showing that naive supervised finetuning fails to bridge the supervisor-student capability gap on complex preference modeling.
  4. Knowl 4 — Auxiliary Confidence Loss for Weak-to-Strong Generalization

    model/method

    To prevent the strong student model from blindly imitating the mistakes of the weak supervisor and to encourage it to rely on its internal pretrained representations, training optimizes a confidence-regularized loss function:

    Lconf(f)=(1−α)⋅CE(f(x),fw(x))+α⋅CE(f(x),f^t(x))L_{\text{conf}}(f) = (1 - \alpha) \cdot \text{CE}(f(x), f_w(x)) + \alpha \cdot \text{CE}(f(x), \hat{f}_t(x))

    which can be equivalently expressed as a self-bootstrapping cross-entropy objective:

    Lconf(f)=CE(f(x),(1−α)⋅fw(x)+α⋅f^t(x))L_{\text{conf}}(f) = \text{CE}(f(x), (1 - \alpha) \cdot f_w(x) + \alpha \cdot \hat{f}_t(x))

    where:

    • f(x)∈[0,1]f(x) \in [0, 1] is the strong student model's predicted probability distribution on input xx.
    • fw(x)∈[0,1]f_w(x) \in [0, 1] is the predictive distribution provided by the weak supervisor.
    • CE(⋅,⋅)\text{CE}(\cdot, \cdot) is the standard cross-entropy loss.
    • f^t(x)=I[f(x)>t]∈{0,1}\hat{f}_t(x) = \mathbb{I}[f(x) > t] \in \{0, 1\} represents hardened pseudo-labels from the student, where I\mathbb{I} is the indicator function.
    • tt is an adaptive threshold set dynamically so that f(x)>tf(x) > t holds for exactly 50% of the examples in each batch.
    • α\alpha is a weighting hyperparameter linearly warmed up from 00 to αmax⁡\alpha_{\max} over the first 20% of training, with αmax⁡=0.75\alpha_{\max} = 0.75 for the largest strong students and αmax⁡=0.5\alpha_{\max} = 0.5 for smaller models.

    On 22 NLP classification datasets, applying this auxiliary confidence loss increases the median PGR for a GPT-4 strong student supervised by a GPT-2-level supervisor from ~25% to nearly 80%.

  5. Knowl 5 — Iterative Bootstrapping of Supervision Across Intermediate Model Scales

    model/method

    Bootstrapping mitigates performance degradation over large compute disparities by chaining weak-to-strong finetuning across a geometric progression of model sizes M1→M2→⋯→MnM_1 \to M_2 \to \dots \to M_n:

    1. Train the initial weak model M1M_1 on ground-truth supervision.
    2. Use M1M_1 to generate weak labels on an unlabeled dataset and finetune intermediate student model M2M_2.
    3. Use the newly trained M2M_2 to generate labels for model M3M_3, repeating the process sequentially until reaching the largest model MnM_n.

    In chess puzzle experiments utilizing three sequential stages (with two intermediate model sizes between the initial weak supervisor and the final strong student), bootstrapping eliminates the performance saturation observed in single-step naive finetuning, maintaining monotonic accuracy improvements and higher PGR as student compute scales.

  6. Knowl 6 — Unsupervised Task-Relevant Generative Finetuning for Representation Salience

    model/method

    When a target alignment capability (such as preference comparison in reward modeling) is not sufficiently salient in base pretrained models, an unsupervised language modeling finetuning stage on in-domain data prior to weak supervision improves downstream generalization.

    In the ChatGPT reward modeling setting, the strong base model is first finetuned using the standard autoregressive language modeling objective on unlabeled ChatGPT dialogs (including both high- and low-quality responses without human preference ratings or ground-truth comparison labels). When subsequent weak-to-strong preference finetuning is applied, this generative pre-finetuning increases the PGR by approximately 10% to 20% relative to an adjusted strong ceiling model that also underwent generative finetuning.

  7. Knowl 7 — Dynamics of Overfitting to Weak Supervisor Errors and Early Stopping

    empirical result

    Strong models trained on weak supervision overfit to the weak supervisor's systematic errors early in training, frequently within less than a single epoch. For large compute disparities between supervisor and student:

    • Ground-truth test accuracy peaks early in training and subsequently declines as training continues, even in the absence of classic sample-level memorization.
    • Using an oracle early stopping criterion evaluated on ground-truth labels increases PGR by approximately 5 percentage points in reward modeling and 15 percentage points on NLP tasks compared to end-of-training checkpoints.
    • Oracle early stopping outperforms standard early stopping based on weak validation labels by roughly 10 percentage points of PGR on NLP tasks.
    • The auxiliary confidence loss stabilizes training dynamics and substantially mitigates this error-overfitting, narrowing the gap between weak validation early stopping and oracle ground-truth early stopping to ~5% PGR.
  8. Knowl 8 — Inverse Scaling in Student Imitation of Supervisor Errors

    empirical result

    Measuring student-supervisor agreement—the fraction of test examples where the strong student's prediction matches the weak supervisor's prediction—reveals an inverse scaling trend: larger strong student models agree less with supervisor errors than smaller strong models, despite having greater capacity to fit the training labels and being trained to convergence without early stopping.

    Across NLP, chess, and reward modeling tasks, naive student-supervisor agreement exceeds the supervisor's ground-truth accuracy (confirming that the student learns some supervisor mistakes). Applying the auxiliary confidence loss significantly reduces agreement on incorrect supervisor predictions, allowing the strong model to override flawed weak supervision.

  9. Knowl 9 — Linear Representation Emergence and Task Saliency under Weak Supervision

    empirical result

    On NLP classification benchmarks, evaluating linear probes trained on frozen final activations demonstrates how weak supervision affects the geometry of learned representations:

    • A linear probe trained directly on ground-truth labels from the frozen base model achieves 72% accuracy, significantly underperforming full ground-truth finetuning (82% accuracy), indicating that the task concept is initially non-linearly represented.
    • Finetuning the base model on weak supervisor labels increases the linear separability of the underlying ground-truth concept: a ground-truth linear probe trained on the representations of a weak-finetuned model achieves 78% accuracy.
    • This weak finetuning plus linear probe approach closes 60% of the accuracy gap between frozen linear probing and full ground-truth finetuning, outperforming naive weak-to-strong end-to-end finetuning.
  10. Knowl 10 — Weak-to-Strong Generalization in Vision Without Pretraining Leakage

    data/table

    To test weak-to-strong generalization in a setting free from potential natural language pretraining contamination (where tasks might have appeared in pretraining data), linear classification probes were trained on frozen representations of self-supervised vision models (DINO ResNet-50 and DINO ViT-B/8) using labels generated by an AlexNet weak supervisor on ImageNet validation data (40k training examples, 10k evaluation examples):

    Model Top-1 Accuracy (%) PGR (%)
    AlexNet (weak supervisor) 56.6 -
    DINO ResNet-50 (strong ceiling) 63.7 -
    DINO ViT-B/8 (strong ceiling) 74.9 -
    AlexNet →\to DINO ResNet-50 60.7 57.8
    AlexNet →\to DINO ViT-B/8 64.2 41.5

    The strong student models outperform the AlexNet supervisor by 4.1% and 7.6% top-1 accuracy, recovering 57.8% and 41.5% of the performance gap respectively, demonstrating that weak supervision can elicit latent capabilities in purely self-supervised models without direct supervised pretraining leakage.

  11. Knowl 11 — Impact of Weak Label Error Structure on Generalization and Student Imitation

    empirical result

    The efficacy of weak-to-strong generalization depends critically on the structural properties and simulability of weak supervisor errors, holding overall error rate constant:

    • Unsimulatable Errors (Random Noise): When weak errors are generated by flipping labels uniformly at random, the strong student easily denoises the signal, resulting in high weak-to-strong accuracy and high PGR.
    • Simulatable / Systematic Errors: When errors are structured according to simple heuristics (such as truncating feature sets in linear probing or flipping labels based on input string length), strong students rapidly overfit to the systematic error pattern, yielding low or negative PGR under naive finetuning.
    • Confidence Loss Mitigation: In scenarios with structured, easily learnable weak errors, applying the auxiliary confidence loss substantially reduces imitation of systematic mistakes and restores positive generalization.
  12. Knowl 12 — Structural Disanalogies in Model-Supervised Alignment Frameworks

    limitation

    Using small language models to supervise large language models introduces two primary disanalogies when serving as an empirical testbed for human supervision of superhuman AI:

    1. Imitation Saliency (The Human Simulator Problem): Future superhuman models trained extensively on human web corpora will possess sophisticated internal models of human psychology and human mistakes. Consequently, superhuman models may have a strong inductive bias to predict and simulate human errors when given human supervision, whereas current strong base models are not explicitly pretrained to simulate smaller language models.
    2. Pretraining Leakage of Task Representations: Current NLP and chess benchmarks may be partially represented in large-scale pretraining datasets. Weak supervision in these setups may simply elicit capabilities that were observed during pretraining, whereas superhuman alignment tasks may require eliciting fully latent capabilities developed entirely via reinforcement learning or self-supervised mechanisms.

Coverage note — Omitted speculative alignment governance roadmaps (Appendix I.1-I.3), task-specific few-shot prompt strings (Table 2), and negative results on alternative regularization heuristics (LoRA rank sweeps, weight decay, dropout, weight averaging/EMA, LP-FT, and rephrasing data augmentations from Appendix D) which did not provide consistent improvements.

References

  1. 1.Lichess puzzle database. URL https://database.lichess.org/#puzzles.
  2. 2.Arazo, E., Ortego, D., Albert, P., O’Connor, N., and McGuinness, K. Unsupervised label noise modeling and loss correction. In International conference on machine learning, pp. 312–321. PMLR, 2019.
  3. 3.Atkeson, C. G. and Schaal, S. Robot learning from demonstration. In ICML, volume 97, pp. 12–20. Citeseer, 1997.
  4. 4.Awadalla, A., Wortsman, M., Ilharco, G., Min, S., Magnusson, I., Hajishirzi, H., and Schmidt, L. Exploring the landscape of distributional robustness for question answering models. arXiv preprint arXiv:2210.12517, 2022.
  5. 5.Bach, S. H., He, B., Ratner, A., and Re, C. Learning the structure of generative models without labeled data. In International Conference on Machine Learning, pp. 273–282. PMLR, 2017.
  6. 6.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022a.
  7. 7.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073, 2022b.
  8. 8.Bain, M. and Sammut, C. A framework for behavioural cloning. In Machine Intelligence 15, pp. 103–129, 1995.
  9. 9.Ben Zhou, Daniel Khashabi, Q. N. and Roth, D. “going on a vacation” takes longer than “going for a walk”: A study of temporal commonsense understanding. In EMNLP, 2019.
  10. 10.Bengio, Y., Hinton, G., Yao, A., Song, D., Abbeel, P., Harari, Y. N., Zhang, Y.-Q., Xue, L., Shalev-Shwartz, S., Hadfield, G., et al. Managing ai risks in an era of rapid progress. arXiv preprint arXiv:2310.17688, 2023.
  11. 11.Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., and Raffel, C. A. Mixmatch: A holistic approach to semi-supervised learning. Advances in neural information processing systems, 32, 2019.
  12. 12.Beyer, L., Zhai, X., Royer, A., Markeeva, L., Anil, R., and Kolesnikov, A. Knowledge distillation: A good teacher is patient and consistent. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10925–10934, 2022.
  13. 13.Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W. Language models can explain neurons in language models. OpenAI Blog, 2023. URL https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html.
  14. 14.Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y. Piqa: Reasoning about physical commonsense in natural language. In Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020.
  15. 15.Bowman, S. Artificial sandwiching: When can we test scalable alignment protocols without humans? AI Alignment Forum, 2022. URL https://www.alignmentforum.org/posts/nekLYqbCEBDEfbLzF/artificial-sandwiching-when-can-we-test-scalable-alignment.
  16. 16.Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., Lukosuite, K., Askell, A., Jones, A., Chen, A., et al. Measuring progress on scalable oversight for large language models. arXiv preprint arXiv:2211.03540, 2022.
  17. 17.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  18. 18.Burns, C., Ye, H., Klein, D., and Steinhardt, J. Discovering latent knowledge in language models without supervision. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=ETKGuby0hcs.
  19. 19.CAIS. Statement on ai risk, 2022. URL https://www.safe.ai/statement-on-ai-risk.
  20. 20.Carlsmith, J. Scheming ais: Will ais fake alignment during training in order to get power? arXiv preprint arXiv:2311.08379, 2023.
  21. 21.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 9650–9660, 2021.
  22. 22.Cha, J., Chun, S., Lee, K., Cho, H.-C., Park, S., Lee, Y., and Park, S. Swad: Domain generalization by seeking flat minima. Advances in Neural Information Processing Systems, 34:22405–22418, 2021.
  23. 23.Chen, T., Kornblith, S., Swersky, K., Norouzi, M., and Hinton, G. E. Big self-supervised models are strong semi-supervised learners. Advances in neural information processing systems, 33:22243–22255, 2020a.
  24. 24.Chen, Y., Wei, C., Kumar, A., and Ma, T. Self-training avoids using spurious features under domain shift. Advances in Neural Information Processing Systems, 33: 21061–21071, 2020b.
  25. 25.Christiano, P. Approval-directed bootstrapping. AI Alignment Forum, 2018. URL https://www.alignmentforum.org/posts/6x7oExXi32ot6HjJv/approval-directed-bootstrapping.
  26. 26.Christiano, P. Capability amplification. AI Alignment Forum, 2019. URL https://www.alignmentforum.org/posts/t3AJW5jP3sk36aGoC/capability-amplification-1.
  27. 27.Christiano, P., Shlegeris, B., and Amodei, D. Supervising strong learners by amplifying weak experts. arXiv preprint arXiv:1810.08575, 2018.
  28. 28.Christiano, P., Cotra, A., and Xu, M. Eliciting latent knowledge. Technical report, Technical report, ARC, 2022. URL https://docs. google. com/document/d . . . , 2022.
  29. 29.Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  30. 30.Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. Boolq: Exploring the surprising difficulty of natural yes/no questions. In NAACL, 2019.
  31. 31.Cotra, A. The case for aligning narrowly superhuman models. AI Alignment Forum, 2021. URL https://www.alignmentforum.org/posts/PZtsoaoSLpKjjbMqM/the-case-for-aligning-narrowly-superhuman-models.
  32. 32.Dai, A. M. and Le, Q. V. Semi-supervised sequence learning. Advances in neural information processing systems, 28, 2015.
  33. 33.Demski, A. and Garrabrant, S. Embedded agency. arXiv preprint arXiv:1902.09469, 2019.
  34. 34.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  35. 35.Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021. https://transformer-circuits.pub/2021/framework/index.html.
  36. 36.Evans, O., Cotton-Barratt, O., Finnveden, L., Bales, A., Balwit, A., Wills, P., Righetti, L., and Saunders, W. Truthful ai: Developing and governing ai that does not lie. arXiv preprint arXiv:2110.06674, 2021.
  37. 37.Frenay, B. and Verleysen, M. Classification in the presence of label noise: a survey. IEEE transactions on neural networks and learning systems, 25(5):845–869, 2013.
  38. 38.French, G., Mackiewicz, M., and Fisher, M. Self-ensembling for visual domain adaptation. arXiv preprint arXiv:1706.05208, 2017.
  39. 39.Furlanello, T., Lipton, Z., Tschannen, M., Itti, L., and Anandkumar, A. Born again neural networks. In International Conference on Machine Learning, pp. 1607–1616. PMLR, 2018.
  40. 40.Gao, L., Schulman, J., and Hilton, J. Scaling laws for reward model overoptimization. In International Conference on Machine Learning, pp. 10835–10866. PMLR, 2023.
  41. 41.Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375, 2022.
  42. 42.Gou, J., Yu, B., Maybank, S. J., and Tao, D. Knowledge distillation: A survey. International Journal of Computer Vision, 129:1789–1819, 2021.
  43. 43.Grandvalet, Y. and Bengio, Y. Semi-supervised learning by entropy minimization. Advances in neural information processing systems, 17, 2004.
  44. 44.Guo, S., Huang, W., Zhang, H., Zhuang, C., Dong, D., Scott, M. R., and Huang, D. Curriculumnet: Weakly supervised learning from large-scale web images. In Proceedings of the European Conference on Computer Vision (ECCV), September 2018.
  45. 45.Guo, S., Zhang, B., Liu, T., Liu, T., Khalman, M., Llinares-Lopez, F., Rame, A., Mesnard, T., Zhao, Y., Piot, B., Ferret, J., and Blondel, M. Direct language model alignment from online ai feedback. ArXiv, abs/2402.04792, 2024. URL https://api.semanticscholar.org/CorpusID:267522951.
  46. 46.Han, B., Yao, Q., Yu, X., Niu, G., Xu, M., Hu, W., Tsang, I., and Sugiyama, M. Co-teaching: Robust training of deep neural networks with extremely noisy labels. Advances in neural information processing systems, 31, 2018.
  47. 47.Hase, P., Bansal, M., Clark, P., and Wiegreffe, S. The unreasonable effectiveness of easy training data for hard tasks. ArXiv, abs/2401.06751, 2024. URL https://api.semanticscholar.org/CorpusID:266977266.
  48. 48.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  49. 49.Hendrycks, D., Mazeika, M., Wilson, D., and Gimpel, K. Using trusted data to train deep networks on labels corrupted by severe noise. Advances in neural information processing systems, 31, 2018.
  50. 50.Hendrycks, D., Lee, K., and Mazeika, M. Using pre-training can improve model robustness and uncertainty. In International conference on machine learning, pp. 2712–2721. PMLR, 2019.
  51. 51.Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., and Steinhardt, J. Aligning ai with shared human values. arXiv preprint arXiv:2008.02275, 2020a.
  52. 52.Hendrycks, D., Liu, X., Wallace, E., Dziedzic, A., Krishnan, R., and Song, D. Pretrained transformers improve out-of-distribution robustness. arXiv preprint arXiv:2004.06100, 2020b.
  53. 53.Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. Measuring mathematical problem solving with the math dataset. Sort, 2(4):0–6, 2021.
  54. 54.Hinton, G., Vinyals, O., and Dean, J. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.
  55. 55.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
  56. 56.Huang, L., Bras, R. L., Bhagavatula, C., and Choi, Y. Cosmos qa: Machine reading comprehension with contextual commonsense reasoning. arXiv preprint arXiv:1909.00277, 2019.
  57. 57.Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S. Risks from learned optimization in advanced machine learning systems. arXiv preprint arXiv:1906.01820, 2019.
  58. 58.Irving, G., Christiano, P., and Amodei, D. Ai safety via debate. arXiv preprint arXiv:1805.00899, 2018.
  59. 59.Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G. Averaging weights leads to wider optima and better generalization. arXiv preprint arXiv:1803.05407, 2018.
  60. 60.Khashabi, D., Chaturvedi, S., Roth, M., Upadhyay, S., and Roth, D. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 252–262, 2018.
  61. 61.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  62. 62.Kingma, D. P., Mohamed, S., Jimenez Rezende, D., and Welling, M. Semi-supervised learning with deep generative models. Advances in neural information processing systems, 27, 2014.
  63. 63.Kirichenko, P., Izmailov, P., and Wilson, A. G. Last layer re-training is sufficient for robustness to spurious correlations. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=Zb6c8A-Fghk.
  64. 64.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012.
  65. 65.Krogh, A. and Hertz, J. A simple weight decay can improve generalization. Advances in neural information processing systems, 4, 1991.
  66. 66.Kumar, A., Raghunathan, A., Jones, R. M., Ma, T., and Liang, P. Fine-tuning can distort pretrained features and underperform out-of-distribution. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=UYneFzXSJWh.
  67. 67.Laine, S. and Aila, T. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242, 2016.
  68. 68.Lee, D.-H. et al. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on challenges in representation learning, ICML, volume 3, pp. 896. Atlanta, 2013.
  69. 69.Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267, 2023.
  70. 70.Lee, Y., Chen, A. S., Tajwar, F., Kumar, A., Yao, H., Liang, P., and Finn, C. Surgical fine-tuning improves adaptation to distribution shifts. In The Eleventh International Conference on Learning Representations, 2022a.
  71. 71.Lee, Y., Yao, H., and Finn, C. Diversify and disambiguate: Learning from underspecified data. arXiv preprint arXiv:2202.03418, 2022b.
  72. 72.Leike, J. and Sutskever, I. Introducing superalignment. OpenAI Blog, 2023. URL https://openai.com/blog/introducing-superalignment.
  73. 73.Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871, 2018.
  74. 74.Li, J., Socher, R., and Hoi, S. C. Dividemix: Learning with noisy labels as semi-supervised learning. arXiv preprint arXiv:2002.07394, 2020.
  75. 75.Li, K., Patel, O., Viegas, F., Pfister, H., and Wattenberg, M. Inference-time intervention: Eliciting truthful answers from a language model. arXiv preprint arXiv:2306.03341, 2023.
  76. 76.Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023.
  77. 77.Liu, Z., Xu, Y., Xu, Y., Qian, Q., Li, H., Jin, R., Ji, X., and Chan, A. B. An empirical study on distribution shift robustness from the perspective of pre-training and data augmentation. arXiv preprint arXiv:2205.12753, 2022.
  78. 78.Ma, X., Huang, H., Wang, Y., Romano, S., Erfani, S., and Bailey, J. Normalized loss functions for deep learning with noisy labels. In International conference on machine learning, pp. 6543–6553. PMLR, 2020.
  79. 79.McKenzie, I. R., Lyzhov, A., Pieler, M., Parrish, A., Mueller, A., Prabhu, A., McLean, E., Kirtland, A., Ross, A., Liu, A., et al. Inverse scaling: When bigger isn’t better. arXiv preprint arXiv:2306.09479, 2023.
  80. 80.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? a new dataset for open book question answering. In EMNLP, 2018.
  81. 81.Ngo, R., Chan, L., and Mindermann, S. The alignment problem from a deep learning perspective. arXiv preprint arXiv:2209.00626, 2022.
  82. 82.Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., and Kiela, D. Adversarial nli: A new benchmark for natural language understanding. arXiv preprint arXiv:1910.14599, 2019.
  83. 83.Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., and Mordvintsev, A. The building blocks of interpretability. Distill, 2018. doi: 10.23915/distill.00010. https://distill.pub/2018/building-blocks.
  84. 84.OpenAI. GPT-4 technical report. arXiv prepreint arXiv:2303.08774, 2023.
  85. 85.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  86. 86.Pacchiardi, L., Chan, A. J., Mindermann, S., Moscovitz, I., Pan, A. Y., Gal, Y., Evans, O., and Brauner, J. How to catch an ai liar: Lie detection in black-box llms by asking unrelated questions. arXiv preprint arXiv:2309.15840, 2023.
  87. 87.Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Edouard Duchesnay. Scikit-learn: Machine learning in python. Journal of Machine Learning Research, 12(85):2825–2830, 2011. URL http://jmlr.org/papers/v12/pedregosa11a.html.
  88. 88.Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G. Red teaming language models with language models. arXiv preprint arXiv:2202.03286, 2022a.
  89. 89.Perez, E., Ringer, S., Lukoˇsiute, K., Nguyen, K., Chen, E., ˙ Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. Discovering language model behaviors with modelwritten evaluations. arXiv preprint arXiv:2212.09251, 2022b.
  90. 90.Pilehvar, M. T. and Camacho-Collados, J. Wic: the word-incontext dataset for evaluating context-sensitive meaning representations. arXiv preprint arXiv:1808.09121, 2018.
  91. 91.Radford, A., Jozefowicz, R., and Sutskever, I. Learning to generate reviews and discovering sentiment. arXiv preprint arXiv:1704.01444, 2017.
  92. 92.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  93. 93.Ratner, A., Bach, S. H., Ehrenberg, H., Fries, J., Wu, S., and Re, C. Snorkel: Rapid training data creation with weak supervision. In Proceedings of the VLDB Endowment. International Conference on Very Large Data Bases, volume 11, pp. 269. NIH Public Access, 2017.
  94. 94.Reed, S., Lee, H., Anguelov, D., Szegedy, C., Erhan, D., and Rabinovich, A. Training deep neural networks on noisy labels with bootstrapping. arXiv preprint arXiv:1412.6596, 2014.
  95. 95.Ribeiro, M. T., Singh, S., and Guestrin, C. ” why should i trust you?” explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, pp. 1135–1144, 2016.
  96. 96.Roger, F., Greenblatt, R., Nadeau, M., Shlegeris, B., and Thomas, N. Measurement tampering detection benchmark. arXiv preprint arXiv:2308.15605, 2023.
  97. 97.Rogers, A., Kovaleva, O., Downey, M., and Rumshisky, A. Getting closer to ai complete question answering: A set of prerequisite real tasks. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 8722–8731, 2020.
  98. 98.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115: 211–252, 2015.
  99. 99.Sap, M., Rashkin, H., Chen, D., LeBras, R., and Choi, Y. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728, 2019.
  100. 100.Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J. Self-critiquing models for assisting human evaluators. arXiv preprint arXiv:2206.05802, 2022.
  101. 101.Schwarzschild, A., Borgnia, E., Gupta, A., Bansal, A., Emam, Z., Huang, F., Goldblum, M., and Goldstein, T. Datasets for studying generalization from easy to hard examples. arXiv preprint arXiv:2108.06011, 2021a.
  102. 102.Schwarzschild, A., Borgnia, E., Gupta, A., Huang, F., Vishkin, U., Goldblum, M., and Goldstein, T. Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks. Advances in Neural Information Processing Systems, 34:6695–6706, 2021b.
  103. 103.Shu, R., Bui, H., Narui, H., and Ermon, S. A dirt-t approach to unsupervised domain adaptation. In International Conference on Learning Representations, 2018.
  104. 104.Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., and Anderson, R. The curse of recursion: Training on generated data makes models forget. ArXiv, abs/2305.17493, 2023. URL https://api.semanticscholar.org/CorpusID:258987240.
  105. 105.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pp. 1631–1642, 2013.
  106. 106.Song, H., Kim, M., Park, D., Shin, Y., and Lee, J.-G. Learning from noisy labels with deep neural networks: A survey. IEEE Transactions on Neural Networks and Learning Systems, 2022.
  107. 107.Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014.
  108. 108.Stanton, S., Izmailov, P., Kirichenko, P., Alemi, A. A., and Wilson, A. G. Does knowledge distillation really work? Advances in Neural Information Processing Systems, 34: 6906–6919, 2021.
  109. 109.Steinhardt, J. Ai forecasting: One year in, 2022. URL https://bounded-regret.ghost.io/ai-forecasting-one-year-in/.
  110. 110.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33: 3008–3021, 2020.
  111. 111.Sun, K., Yu, D., Chen, J., Yu, D., Choi, Y., and Cardie, C. Dream: A challenge data set and models for dialoguebased reading comprehension. Transactions of the Association for Computational Linguistics, 7:217–231, 2019.
  112. 112.Tafjord, O., Gardner, M., Lin, K., and Clark, P. ”quartz: An open-domain dataset of qualitative relationship questions”. 2019.
  113. 113.Tarvainen, A. and Valpola, H. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. Advances in neural information processing systems, 30, 2017.
  114. 114.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  115. 115.Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems, 32, 2019.
  116. 116.Warstadt, A., Singh, A., and Bowman, S. R. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641, 2019.
  117. 117.Wei, C., Shen, K., Chen, Y., and Ma, T. Theoretical analysis of self-training with deep networks on unlabeled data. In International Conference on Learning Representations, 2020.
  118. 118.Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. In International Conference on Learning Representations, 2021.
  119. 119.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022.
  120. 120.Welbl, J., Liu, N. F., and Gardner, M. Crowdsourcing multiple choice science questions. arXiv preprint arXiv:1707.06209, 2017.
  121. 121.Worley, G. S. Bootstrapped alignment. AI Alignment Forum, 2021. URL https://www.alignmentforum.org/posts/teCsd4Aqg9KDxkaC9/bootstrapped-alignment.
  122. 122.Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In International Conference on Machine Learning, pp. 23965–23998. PMLR, 2022a.
  123. 123.Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Lopes, R. G., Hajishirzi, H., Farhadi, A., Namkoong, H., et al. Robust fine-tuning of zero-shot models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7959–7971, 2022b.
  124. 124.Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862, 2021.
  125. 125.Xie, Q., Luong, M.-T., Hovy, E., and Le, Q. V. Self-training with noisy student improves imagenet classification. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10687–10698, 2020.
  126. 126.Yi, K. and Wu, J. Probabilistic end-to-end noise correction for learning with noisy labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019.
  127. 127.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
  128. 128.Zhang, H., Diao, S., Lin, Y., Fung, Y. R., Lian, Q., Wang, X., Chen, Y., Ji, H., and Zhang, T. R-tuning: Instructing large language models to say ‘i don’t know’. 2023. URL https://api.semanticscholar.org/CorpusID:265220839.
  129. 129.Zhang, Y., Baldridge, J., and He, L. PAWS: Paraphrase Adversaries from Word Scrambling. In Proc. of NAACL, 2019.
  130. 130.Zhang, Z. and Sabuncu, M. Generalized cross entropy loss for training deep neural networks with noisy labels. Advances in neural information processing systems, 31, 2018.

Citation

MLA
Burns, C., et al. “Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision”. arXiv, 2023, http://arxiv.org/abs/2312.09390v1.
APA
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., Sutskever, I., & Wu, J. (2023). Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. arXiv. http://arxiv.org/abs/2312.09390v1
Chicago
Burns, C., P. Izmailov, J. H. Kirchner, et al. 2023. “Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision”. arXiv. http://arxiv.org/abs/2312.09390v1.
Harvard
Burns, C. et al. (2023) “Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.09390v1.
Vancouver
1. Burns C, Izmailov P, Kirchner JH, et al (2023) Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. arXiv

BibTeX

@article{burns2023weak,
  title = {Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision},
  author = {Burns, Collin and Izmailov, Pavel and Kirchner, Jan Hendrik and Baker, Bowen and Gao, Leo and Aschenbrenner, Leopold and Chen, Yining and Ecoffet, Adrien and Joglekar, Manas and Leike, Jan and Sutskever, Ilya and Wu, Jeff},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.09390v1},
  eprint = {2312.09390}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/