Aligning Large Language Models through Synthetic Feedback

Sungdong KimSanghwan BaeJamin ShinSoyoung KangDonghyun KwakKang Min YooMinjoon Seo

article2023EMNLP95 citations

Presents an alignment learning framework that trains language models using synthetic feedback derived from contrasting different model sizes and prompt configurations, eliminating reliance on human annotations or proprietary APIs while outperforming models like Alpaca and Dolly-v2.

Listen

Training large language models to produce helpful, honest, and harmless responses traditionally requires substantial investments in manual human feedback or costly API distillation from proprietary models like ChatGPT. As organizations seek cost-effective, self-contained AI deployment pipelines, relying on external proprietary platforms creates data privacy risks, vendor dependencies, and high ongoing operational expenses.

The article demonstrates an end-to-end framework that aligns open-source language models using purely synthetic feedback. The primary objective is to show that foundation models can be aligned effectively with human values without requiring extensive manual annotations or proprietary model outputs.

The approach operates in three main steps using open-source base models. First, researchers generate synthetic comparison data by contrasting responses from models of varying sizes and prompt qualities based on empirical heuristics (e.g., larger, well-prompted models generally outperform smaller, less-prompted ones), followed by automated heuristic and community-model filtering to remove noise. Second, the authors train a synthetic reward model on 13,000 synthetic pairs and use it in a guided self-play simulation to create 20,000 high-quality synthetic demonstrations for supervised fine-tuning. Third, the resulting model, named ALMoST, is further refined using reinforcement learning against the synthetic reward model.

Evaluation shows that ALMoST consistently outperforms open-source models trained on human annotations or distilled from proprietary systems. In human preference studies, ALMoST-7B won 55.0% of head-to-head comparisons against Alpaca (distilled from InstructGPT) and 58.8% against Dolly-v2 (trained on human annotations). On standardized alignment benchmarks measuring helpfulness, harmlessness, and honesty, ALMoST achieved a 68.8% overall accuracy, exceeding Dolly-v2 (52.0%) and Alpaca (62.9%). Ablation analyses revealed that prompt design and heuristic length filtering were the most critical factors in training an effective synthetic reward model, with filtering alone preventing a 10 percentage point drop in reward model accuracy.

These findings indicate that organizations can establish competitive, safe language models entirely in-house using open-source foundations. This significantly reduces data annotation costs, shortens development timelines, and mitigates regulatory and compliance risks associated with transmitting internal data to third-party proprietary APIs.

Organizations developing custom language models should transition toward synthetic alignment pipelines with robust data-filtering rules and carefully engineered prompts rather than relying exclusively on costly manual labeling. However, testing also revealed evidence of an alignment tax, where general knowledge and language understanding benchmarks (such as MMLU) degraded post-reinforcement learning. Decision-makers should cautiously pilot this synthetic pipeline on task-specific applications, monitoring domain performance trade-offs, and explore scaling to larger model sizes where alignment tax effects are typically less severe.

arXiv: 2305.13735naver-ai/ALMoST
Cover for Aligning Large Language Models through Synthetic Feedback

Abstract

Aligning large language models (LLMs) to human values has become increasingly important as it enables sophisticated steering of LLMs. However, it requires significant human demonstrations and feedback or distillation from proprietary LLMs such as ChatGPT. In this work, we propose a novel alignment learning framework with synthetic feedback not dependent on extensive human annotations and proprietary LLMs. First, we perform reward modeling (RM) with synthetic feedback by contrasting responses from vanilla LLMs with various sizes and prompts. Then, we use the RM to simulate high-quality demonstrations to train a supervised policy and further optimize the model with reinforcement learning. Our resulting model, Aligned Language Model with Synthetic Training dataset (ALMoST), outperforms recent open-sourced models, which are trained on the outputs of InstructGPT or human-annotated demonstrations, in alignment benchmarks. In human evaluation, our model is preferred to Alpaca and Dolly-v2, 55.0% and 58.5% of the time, respectively. Further analyses demonstrate the efficacy and importance of synthetic feedback in our framework 1.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Step 1: Reward Modeling with Synthetic Feedback
  • 2.2 Step 2: Supervised Fine-Tuning
  • 2.3 Step 3: Reinforcement Learning from Synthetic Feedback (RLSF)
  • 3 Evaluating Alignment of ALMoST
  • 3.1 Dataset
  • 3.2 Baselines
  • 3.3 Evaluation Results
  • 4 Analysis
  • 4.1 Probing Main Hypothesis
  • 4.2 RM Evaluation
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Examples of alignment prompts
  • B Examples of synthetic comparisons
  • C Details of Synthetic Datasets
  • C.1 Details of Initial Query Mining
  • C.2 Sampling Configurations
  • C.3 Data Statistics
  • D More Training Details
  • D.1 RM
  • D.2 SFT
  • D.3 PPO
  • E Evaluation on Vicuna Questions
  • E.1 Details of Human Evaluation
  • E.2 All results of GPT-4 evaluation
  • F Alignment Tax
  • G Examples of RMSP vs. Self-Play
  • H Qualitative Examples
  • I Evaluation Prompts
  • TruthfulQA

Knowls

  1. Knowl 1 — Three-stage alignment using synthetic feedback

    model/method

    ALMoST aligns a language model without extensive human preference labels or outputs from proprietary aligned models by chaining three stages: (1) train a reward model (RM) on synthetic comparisons of responses from differently configured vanilla LLaMA models; (2) use the RM to select responses during simulated conversations, then supervised-fine-tune a LLaMA-7B policy on those demonstrations; and (3) further optimize that policy with PPO using the synthetic RM as its reward signal. The comparison-generation stage supplies preferences, RM-guided self-play supplies demonstrations, and reinforcement learning from synthetic feedback (RLSF) further optimizes the policy.

  2. Knowl 2 — Generating synthetic response comparisons

    model/method

    Synthetic comparisons are formed by applying empirical quality priors to responses from vanilla LLaMA models: larger models, more few-shot demonstrations, and better demonstrations are presumed to yield better responses. The five ranked configurations were LLaMA-30B-Faithful-3shot, LLaMA-30B-HHH-5shot, LLaMA-13B-HHH-3shot, LLaMA-7B-HHH-3shot, and LLaMA-7B-HHH-1shot, ordered from best to worst; Faithful is a manually designed prompt, while HHH refers to the alignment prompt based on Helpful, Harmless, and Honest conversations. For each query, the resulting order supplies chosen-versus-rejected response pairs. The queries were mined by prompting LLaMA-30B with 10 examples (7 fixed and 3 previously generated), starting from 10 manually written seed queries, to produce 10,000 initial queries. Queries containing terms such as image, graph, picture, or video were removed, as were queries with maximum Rouge-L overlap greater than 0.5 with an existing query. Response generation used top-p 0.9, temperature 1.0, and a 384-token maximum; query mining used temperature 1.2. Post-validation left 13,687 comparison instances.

  3. Knowl 3 — Filtering and training the synthetic reward model

    model/method

    The synthetic RM is trained on ranked pairs in which a chosen response should score above a rejected response. To reduce errors caused by stochastic generations, the authors apply two filters. The heuristic filter (HF) removes responses containing or beginning with undesirable cues such as “I don’t know” or “well”, and retains comparisons when the chosen response is longer than the rejected response or exceeds a length threshold of M−S/2M-S/2, where MM and SS are the mean and standard deviation of character lengths among responses generated for that query. This is intended to screen out abnormally short generations without simply training on longer-is-better preferences. A second, As-is RM, trained on 20,000 paired examples from StackExchange, is used to retain a synthetic comparison only when it agrees with the synthetic label. The reward model minimizes J(θ)=−E(x,yj,yk)∼D[logσ(rθ(x,yj)−rθ(x,yk))]J(θ)=-E_{(x,y_j,y_k)∼D}[log σ(r_θ(x,y_j)-r_θ(x,y_k))], where xx is a query, yjy_j is the chosen response, yky_k is the rejected response, DD is the synthetic comparison dataset, rθ(x,y)r_θ(x,y) is the scalar score from the RM with parameters θθ, and σσ is the logistic function. The RM starts from LLaMA-7B and is trained for one epoch with learning rate 10−510^{-5}, batch size 64, and maximum sequence length 1,024.

  4. Knowl 4 — Reward-Model-guided self-play for demonstrations

    model/method

    Reward-Model-guided Self-Play (RMSP) simulates assistant-user conversations while using the synthetic RM to select assistant responses. Starting from an initial query, LLaMA-30B-Faithful-3shot plays the assistant and LLaMA-30B-User-3shot plays a user instructed to ask follow-up questions when an answer is insufficient. At each assistant turn, the assistant model samples N=4N=4 candidate responses for the current context; the RM scores them, and the highest-scoring response is used in the simulated conversation. The two roles take turns up to a maximum of two turns. Applying this procedure to the mined queries produced 19,752 demonstrations, each comprising a prompt-response pair.

  5. Knowl 5 — Supervised fine-tuning and reinforcement learning

    model/method

    The 19,752 RMSP demonstrations are used to supervised-fine-tune LLaMA-7B for three epochs, with batch size 128, learning rate 2×10−52×10^{-5}, and maximum sequence length 512; the input format uses Human: and Assistant: prefixes. The resulting policy is then optimized with PPO on 20,000 distinct demonstration prompts for 80,000 episodes. The PPO setup uses batch size 512, minibatches of 32, four inner epochs per sample, a 128-token maximum response, sampling temperature 1, clip ratio 0.2, discount factor 1, and AdamW with initial learning rate 10−610^{-6}, β1=0.9β_1=0.9, and β2=0.95β_2=0.95. Its objective is expected synthetic reward with a KL penalty relative to the SFT policy: Ex∼D,y∼πϕ(⋅∣x)[rθ(x,y)−λlog(πϕ(y∣x)/ρ(y∣x))]E_{x∼D,y∼π_ϕ(·|x)}[r_θ(x,y)-λ log(π_ϕ(y|x)/ρ(y|x))], where πϕπ_ϕ is the PPO policy, ρρ is the fixed SFT reference policy, rθr_θ is the synthetic RM score, and the experimental KL coefficient is λ=0.05λ=0.05.

  6. Knowl 6 — Zero-shot alignment benchmark results

    empirical result

    On zero-shot Static HHH alignment and TruthfulQA MC1, ALMoST’s LLaMA-7B policies outperform the open-source human- or proprietary-model-distilled baselines listed below on the aggregate HHH and TruthfulQA results, though Vicuna scores higher on several measures. Values are accuracy percentages; Static HHH values are reported in the order Helpful, Harmless, Honest, Other, All, followed by TruthfulQA MC1. Dolly-v2-12B: 67.8, 46.6, 50.7, 62.8, 56.6, 15.2; OpenAssistant-v4-12B: 59.3, 56.9, 47.5, 69.8, 57.5, 23.3; Vicuna-13B: 78.0, 89.7, 70.5, 81.4, 79.6, 63.3; Dolly-v2-7B: 69.5, 41.4, 45.9, 51.2, 52.0, 24.2; Alpaca-7B: 71.2, 53.4, 62.3, 65.1, 62.9, 19.5; Vicuna-7B: 79.7, 72.4, 70.5, 76.7, 74.7, 52.5; ALMoST-SFT: 79.7, 56.9, 65.6, 69.8, 67.8, 31.5; ALMoST-PPO: 81.4, 60.3, 62.3, 72.1, 68.8, 38.0; and ALMoST-RM: 74.6, 67.2, 78.7, 86.0, 76.0, 54.8. Thus ALMoST-PPO attains the highest Helpful accuracy in this comparison, while ALMoST-RM attains the highest aggregate HHH and TruthfulQA MC1 scores among the listed 7B models.

  7. Knowl 7 — Human preferences on Vicuna evaluation questions

    empirical result

    In human A/B evaluations on 80 Vicuna questions, three workers judged each pair of answers for helpfulness, harmlessness, and honesty based on answer content. ALMoST-PPO was preferred over Alpaca-7B in 55.0% of comparisons (12.5% ties; 32.5% losses) and over Dolly-v2-7B in 58.8% (26.3% ties; 15.0% losses). Against ALMoST-SFT, its win/tie/loss rates were 37.5%/36.3%/26.3%; against Vicuna-7B, they were 25.0%/25.0%/50.0%. The overall inter-rater agreement was moderate (Fleiss’ κ=0.41κ=0.41).

  8. Knowl 8 — Synthetic reward model performance and filtering ablation

    empirical result

    On the Helpful-base split of HH-RLHF, the synthetic RM reached 65.2% accuracy using 13,687 synthetic comparisons. This matches the accuracy of a model trained on the 11,738-instance single-turn subset of Helpful-base (65.2%), and is about 90% of the 71.8% accuracy obtained by full fine-tuning on 43,835 Helpful-base instances. A zero-shot RM trained on 25,057 StackExchange pairs reached 63.7%; the random and always-choose-the-longer-response baselines reached 50.0% and 59.4%, respectively. The post-validation ablation shows that both filters contribute: using only the As-is RM gave 63.3%, using only HF gave 55.5%, and using both gave 65.2%. In particular, removing HF substantially reduced accuracy, while the full synthetic RM exceeded the length-only baseline.

  9. Knowl 9 — Evidence for prompt quality and RM-guided selection

    empirical result

    GPT-4 pairwise evaluations of prompted response generators on Vicuna questions, each compared with Alpaca-7B, supported the synthetic ranking priors but showed that prompt quality mattered strongly. Win/tie/loss counts were: LLaMA-7B-HHH-1shot 12/15/133; 7B-HHH-3shot 12/14/134; 7B-HHH-5shot 15/17/128; 13B-HHH-3shot 17/17/126; 30B-HHH-5shot 39/21/100; 7B-Faithful-3shot 52/14/94; 13B-Faithful-3shot 59/19/82; and 30B-Faithful-3shot 77/13/70. Faithful-3shot 7B therefore beat HHH-5shot 30B in this comparison. Separately, the supervised policy trained from RMSP demonstrations scored 67.8 on Static HHH and had a 54.3% GPT-4 win rate against Alpaca-7B; ordinary Self-Play with the Faithful prompt scored 66.0 and won 40.0%, while Self-Play with HHH scored 61.3 and won 15.0%. These results indicate that both RM-based rejection sampling and the Faithful prompt improved the evaluated alignment outcomes over the tested alternatives.

  10. Knowl 10 — Observed alignment tax and evaluation limits

    limitation

    The authors report that alignment with PPO reduced performance on zero-shot capability benchmarks relative to the unaligned LLaMA-7B, consistent with an alignment tax. MMLU aggregate accuracy was 31.3 for LLaMA-7B, 31.5 for ALMoST-SFT, and 28.9 for ALMoST-PPO; LAMBADA accuracy was 72.1, 68.0, and 65.4, respectively. The paper also notes that its alignment evaluations do not cover all relevant model behaviors and that the approach primarily targets values such as helpfulness. It leaves broader evaluation and mitigation of this capability trade-off for future work.

Coverage note — The paper’s detailed prompt examples, qualitative response comparisons, and full per-baseline GPT-4 evaluation breakdown are omitted because they illustrate or extend the methods and results already captured here rather than add distinct core contributions.

References

  1. 1.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861.
  2. 2.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022a. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  3. 3.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022b. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  4. 4.Edward Beeching, Younes Belkada, Kashif Rasul, Lewis Tunstall, Leandro von Werra, Nazneen Rajani, and Nathan Lambert. 2023. Stackllama: An rl fine-tuned llama model for stack exchange question and answering.
  5. 5.Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023. Pythia: A suite for analyzing large language models across training and scaling.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. volume 33, pages 1877–1901.
  7. 7.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  8. 8.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  9. 9.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  10. 10.DataBricks. 2023. Dolly. https://github.com/databrickslabs/dolly.
  11. 11.Amelia Glaese, Nat McAleese, Maja Tr˛ebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Sona Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving. 2022. Improving alignment of dialogue agents via targeted human judgements.
  12. 12.Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. The false promise of imitating proprietary llms. arXiv preprint arXiv:2305.15717.
  13. 13.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  14. 14.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751.
  15. 15.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, et al. 2023. Openassistant conversations–democratizing large language model alignment. arXiv preprint arXiv:2304.07327.
  16. 16.Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, and Ethan Perez. 2023. Pre-training language models with human preferences.
  17. 17.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  18. 18.Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958.
  19. 19.Hao Liu, Carmelo Sferrazza, and Pieter Abbeel. 2023. Chain of hindsight aligns language models with feedback.
  20. 20.Ruibo Liu, Ge Zhang, Xinyu Feng, and Soroush Vosoughi. 2022. Aligning generative language models with human values. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 241–252, Seattle, United States. Association for Computational Linguistics.
  21. 21.OpenAI. 2023. Gpt-4 technical report.
  22. 22.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  23. 23.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031.
  24. 24.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction tuning with gpt-4. arXiv preprint arXiv:2304.03277.
  25. 25.Jérémy Scheurer, Jon Ander Campos, Tomasz Korbak, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez. 2023. Training language models with language feedback at scale.
  26. 26.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
  27. 27.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2022. Learning to summarize from human feedback.
  28. 28.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  29. 29.Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. 2023. Principle-driven self-alignment of language models from scratch with minimal human supervision. arXiv preprint arXiv:2305.03047.
  30. 30.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  31. 31.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models.
  32. 32.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560.
  33. 33.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (EMNLP demo), pages 38–45.
  34. 34.Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. 2023. Rrhf: Rank responses to align language models with human feedback without tears.
  35. 35.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2023. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206.
  36. 36.Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2020. Fine-tuning language models from human preferences.

Citation

MLA
Kim, S., et al. “Aligning Large Language Models Through Synthetic Feedback”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 13677–700, https://doi.org/10.18653/v1/2023.emnlp-main.844.
APA
Kim, S., Bae, S., Shin, J., Kang, S., Kwak, D., Yoo, K., & Seo, M. (2023). Aligning Large Language Models through Synthetic Feedback. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13677–13700. https://doi.org/10.18653/v1/2023.emnlp-main.844
Chicago
Kim, S., S. Bae, J. Shin, et al. 2023. “Aligning Large Language Models Through Synthetic Feedback”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13677–700. https://doi.org/10.18653/v1/2023.emnlp-main.844.
Harvard
Kim, S. et al. (2023) “Aligning Large Language Models through Synthetic Feedback”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 13677–13700. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.844.
Vancouver
1. Kim S, Bae S, Shin J, Kang S, Kwak D, Yoo K, Seo M (2023) Aligning Large Language Models through Synthetic Feedback. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 13677–13700

BibTeX

@inproceedings{kim-etal-2023-aligning,
    title = "Aligning Large Language Models through Synthetic Feedback",
    author = "Kim, Sungdong  and
      Bae, Sanghwan  and
      Shin, Jamin  and
      Kang, Soyoung  and
      Kwak, Donghyun  and
      Yoo, Kang  and
      Seo, Minjoon",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.844/",
    doi = "10.18653/v1/2023.emnlp-main.844",
    pages = "13677--13700"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/