Large Language Models Can Self-Improve

Jiaxin HuangShixiang GuLe HouYuexin WuXuezhi WangHongkun YuJiawei Han

article2023EMNLP826 citations

Demonstrates that large language models can boost their reasoning capabilities on unlabeled datasets by fine-tuning on their own high-confidence, chain-of-thought solutions selected via majority voting.

Listen

Improving large language models typically requires massive volumes of expensive, human-annotated data. This requirement creates major bottlenecks for organizations working in specialized domains or low-resource settings where labeled training data is scarce. Meanwhile, humans can enhance their problem-solving abilities purely through internal reflection and practice. The article addresses whether large language models can likewise self-improve their reasoning capabilities using only unlabeled questions, removing the dependency on human-provided answers.

The article demonstrates and evaluates a self-training framework called Language Model Self-Improved (LMSI). The objective is to verify whether a model can generate its own step-by-step reasoning paths, filter out incorrect outputs without human supervision, and use those self-generated solutions to fine-tune itself into a more capable system.

The researchers evaluated this framework using a 540-billion-parameter language model across multiple arithmetic, commonsense, and natural language inference benchmarks. The approach prompts the model with a few examples to generate multiple diverse reasoning paths for each unlabeled question. It then applies majority voting—selecting the most frequent final answer—to identify high-confidence solutions. Finally, the model is fine-tuned on these self-selected reasoning paths formatted across multiple prompt styles. Credibility is further established through out-of-domain testing, ablation studies on data formatting, and experiments where the model also self-generates training questions and prompt demonstrations from scratch.

The evaluation revealed several critical findings. First, self-improvement without ground-truth answers significantly boosted performance across all tested tasks; for example, accuracy on the GSM8K math benchmark rose from 74.4% to 82.1%, while natural language inference on ANLI-A3 improved from 63.4% to 67.9%. Second, the approach generalized to out-of-domain tasks, raising performance across six completely unseen benchmarks by up to 10.4 percentage points. Third, when human prompts and questions were absent, the model self-generated questions and reasoning demonstrations, achieving a state-of-the-art zero-shot accuracy of 74.2% on GSM8K. Fourth, knowledge from the self-improved 540-billion model was successfully transferred to smaller systems; an 8-billion model distilled this way reached 33.4% accuracy, outperforming an un-fine-tuned 62-billion model (29.7%), while a distilled 62-billion model (57.4%) surpassed the base 540-billion model (56.5%). Finally, post-improvement inference required far fewer sampled paths—using just 5 paths after self-improvement exceeded the performance of using 32 paths on the base model.

These findings indicate that large models already contain latent reasoning capabilities that can be activated without costly manual labeling. For enterprise deployments, this drastically lowers data curation costs and mitigates compliance risks tied to manual annotation. Furthermore, the ability to distill reasoning into smaller models and reduce test-time sampling provides direct paths toward slashing operational compute expenses, lowering server latency, and improving environmental efficiency.

Organizations developing or deploying large language models should consider adopting self-improvement and distillation pipelines for reasoning-heavy workflows rather than immediately funding large data-annotation efforts. When implementing this workflow, practitioners should train on mixed prompt formats to avoid overfitting to specific instruction styles, tune sampling temperatures upward during evaluation, and test distilled smaller architectures to reduce operational costs. Before broad deployment, teams should conduct internal pilots to measure return on investment and assess task-specific accuracy.

The primary limitation of this method is its dependency on initial model scale; smaller baseline architectures lack the calibration and in-context reasoning needed to reliably generate high-confidence pseudo-labels. Consequently, organizations should apply this self-improvement technique directly to high-capacity frontier models, using smaller models primarily as downstream recipients of distilled knowledge.

arXiv: 2210.11610
Cover for Large Language Models Can Self-Improve

Abstract

Large Language Models (LLMs) have achieved excellent performances in various tasks. However, fine-tuning an LLM requires extensive supervision. Human, on the other hand, may improve their reasoning abilities by self-thinking without external inputs. In this work, we demonstrate that an LLM is also capable of self-improving with only unlabeled datasets. We use a pre-trained LLM to generate "high-confidence" rationale-augmented answers for unlabeled questions using Chain-of-Though (CoT) prompting and self-consistency, and fine-tune the LLM using those self-generated solutions as target outputs. We show that without any ground truth label, our approach significantly improves the general reasoning ability of PaLM 540B model (74.4%→82.1% on GSM8K, 90.0%→94.4% on OpenBookQA, and 63.4%→67.9% on ANLI-A3) and can also be adapted to extreme low-resource cases where even training questions and CoT prompts are limited. We conduct ablation studies and show that fine-tuning on diverse reasoning paths is critical for self-improvement.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Generating and Filtering Multiple Reasoning Paths
  • 3.2 Training with Mixed Formats
  • 3.3 Generating Questions and Prompts
  • 4 Experimental Setup
  • 5 Experiments and Results
  • 5.1 Main Results
  • 5.2 Pushing the limit of self-improvements
  • 5.3 Distillation to smaller models
  • 5.4 Hyperparameter Studies
  • 6 Conclusions
  • Limitations
  • Acknowledgments
  • References
  • A Appendix
  • A.1 Results on UL2 model
  • A.2 Chain-of-Thought Prompts for Each Dataset

Knowls

  1. Knowl 1 — Language Model Self-Improvement (LMSI) Framework

    algorithm

    Language Model Self-Improvement (LMSI) is an unsupervised framework that enhances the reasoning capabilities of a pre-trained large language model (LLM) MM using an unlabeled question set Dtrain={xi}i=1D\mathcal{D}_{\text{train}} = \{x_i\}_{i=1}^D and task-specific few-shot Chain-of-Thought (CoT) demonstration exemplars.

    Input: Pre-trained LLM MM, unlabeled dataset Dtrain={xi}i=1D\mathcal{D}_{\text{train}} = \{x_i\}_{i=1}^D, few-shot CoT exemplars, sampling temperature T>0T > 0, number of decoded paths mm, batch size BB, learning rate η\eta, training steps SS
    Output: Fine-tuned LLM MLMSIM_{\text{LMSI}}
    Initialize training dataset Dself=∅\mathcal{D}_{\text{self}} = \emptyset
    for each question xi∈Dtrainx_i \in \mathcal{D}_{\text{train}} do
        Prompt MM with few-shot CoT exemplars and xix_i
        Sample mm independent reasoning paths and answer pairs {(rij,yij)}j=1m\{(r_{ij}, y_{ij})\}_{j=1}^m using temperature TT
        Determine the majority-voted answer:
            y~i=arg⁡max⁡y∑j=1mI(yij=y)\tilde{y}_i = \arg\max_y \sum_{j=1}^m \mathbb{I}(y_{ij} = y)
        Select consistent reasoning paths: r~i={rij∣yij=y~i}\tilde{r}_i = \{r_{ij} \mid y_{ij} = \tilde{y}_i\}
        for each path r∈r~ir \in \tilde{r}_i do
            Generate 4 mixed-format pairs (uk,vk)k=14(u_k, v_k)_{k=1}^4 from (xi,r,y~i)(x_i, r, \tilde{y}_i)
            Add {(uk,vk)}k=14\{(u_k, v_k)\}_{k=1}^4 to Dself\mathcal{D}_{\text{self}}
        end for
    end for
    Fine-tune MM on Dself\mathcal{D}_{\text{self}} for SS steps with batch size BB and learning rate η\eta using standard cross-entropy loss to produce MLMSIM_{\text{LMSI}}
    return MLMSIM_{\text{LMSI}}

    For PaLM 540B, the generation hyperparameters are m=32m = 32 reasoning paths, sampling temperature T=0.7T = 0.7, and maximum decoding length of 256 tokens. The fine-tuning parameters are S=10,000S = 10{,}000 steps, learning rate η=5×10−5\eta = 5 \times 10^{-5}, and batch size B=32B = 32.

  2. Knowl 2 — Mixed-Format Training Augmentation for Self-Improvement

    model/method

    To prevent the fine-tuned language model from overfitting to specific prompt styles, answer patterns, or demonstration formats, LMSI augments every selected question-reasoning-answer triplet (x,r,y~)(x, r, \tilde{y}) into four distinct training input-output pairs:

    1. Few-Shot Chain-of-Thought (Format 1):
      • Input: [Few-Shot CoT Exemplars]\n[Question]\nA:
      • Target Output: [Reasoning Path r]\nThe answer is [\tilde{y}].
    2. Few-Shot Standard Direct Answer (Format 2):
      • Input: [Few-Shot Standard Exemplars]\n[Question]\nA:
      • Target Output: The answer is [\tilde{y}].
    3. Zero-Shot Chain-of-Thought (Format 3):
      • Input: [Question]\nA: Let's think step by step.
      • Target Output: [Reasoning Path r]\nThe answer is [\tilde{y}].
    4. Zero-Shot Standard Direct Answer (Format 4):
      • Input: [Question]\nA:
      • Target Output: The answer is [\tilde{y}].

    The few-shot standard exemplars in Format 2 use the identical questions and final answers as the few-shot CoT exemplars, but with all intermediate reasoning chains removed.

  3. Knowl 3 — In-Domain Reasoning Performance on PaLM 540B

    data/table

    The LMSI approach was evaluated on PaLM 540B across six reasoning benchmark datasets spanning arithmetic reasoning (GSM8K, DROP), commonsense reasoning (ARC-Challenge, OpenBookQA), and natural language inference (ANLI-A2, ANLI-A3). For each dataset, testing was performed under three evaluation regimes: Standard Prompting, single-path CoT Prompting, and multi-path Self-Consistency (SC).

    Prompting Method Model GSM8K DROP ARC-c OpenBookQA ANLI-A2 ANLI-A3
    Standard-Prompting w/o LMSI 17.9 60.0 87.1 84.4 55.8 55.8
    w. LMSI 32.2 (+14.3) 71.7 (+11.7) 87.2 (+0.1) 92.0 (+7.6) 64.8 (+9.0) 66.9 (+11.1)
    CoT-Prompting w/o LMSI 56.5 70.6 85.2 86.4 58.9 60.6
    w. LMSI 73.5 (+17.0) 76.2 (+5.6) 88.3 (+3.1) 93.0 (+6.6) 65.3 (+6.4) 67.3 (+6.7)
    Self-Consistency w/o LMSI 74.4 78.2 88.7 90.0 64.5 63.4
    w. LMSI 82.1 (+7.7) 83.0 (+4.8) 89.8 (+1.1) 94.4 (+4.4) 66.5 (+2.0) 67.9 (+4.5)

    Self-training on self-generated solutions without ground truth annotations improves all evaluation metrics. Notably, single-path CoT-Prompting on the LMSI model achieves performance comparable to or exceeding multi-path Self-Consistency on the pre-trained base model (e.g., 73.5% vs. 74.4% on GSM8K).

  4. Knowl 4 — Out-of-Domain Multi-Task Generalization of LMSI

    data/table

    To evaluate cross-task generalization, a single PaLM 540B model was fine-tuned on the combined unlabeled training data of six in-domain tasks (GSM8K, DROP, ARC-c, OpenBookQA, ANLI-A2, and ANLI-A3) and evaluated using CoT-prompting on six unseen out-of-domain (OOD) tasks spanning arithmetic (AQUA, SVAMP), commonsense reasoning (StrategyQA), and natural language inference (ANLI-A1, RTE, MNLI-M/MM).

    Model AQUA SVAMP StrategyQA ANLI-A1 RTE MNLI-M/MM
    w/o LMSI 35.8 79.0 75.3 68.8 79.1 72.0 / 74.0
    w. LMSI 39.0 (+3.2) 82.8 (+3.8) 77.8 (+2.5) 79.2 (+10.4) 80.1 (+1.0) 81.8 / 82.2 (+9.8 / +8.2)

    Multi-task self-training consistently improves accuracy across all unseen datasets, demonstrating that LMSI enhances broad underlying reasoning capabilities rather than simply memorizing dataset-specific idiosyncrasies.

  5. Knowl 5 — Impact of Training Format Augmentation on Self-Improvement

    data/table

    Ablation experiments on GSM8K using PaLM 540B illustrate the contribution of different format compositions during self-training.

    Training Configuration Std. Prompting CoT Prompting
    Pre-trained Base Model (w/o LMSI) 17.9 56.5
    LMSI without CoT formats (Direct formats only) 23.6 (+5.7) 61.6 (+5.1)
    LMSI with only few-shot CoT formats 29.2 (+11.3) 69.4 (+12.9)
    LMSI with all 4 mixed formats 32.2 (+14.3) 73.5 (+17.0)

    While training exclusively on direct answer formats yields minor gains and training exclusively on few-shot CoT formats risks overfitting to the prompt style, combining all four formats (few-shot CoT, few-shot direct, zero-shot CoT, and zero-shot direct) yields the largest improvements across both standard and CoT evaluation.

  6. Knowl 6 — Self-Generation of Questions and Few-Shot Prompts in Low-Resource Settings

    model/method

    In extreme low-resource regimes where training questions or human-curated CoT demonstration exemplars are missing, the LLM can bootstrap its own training pipeline:

    1. Unsupervised Question Generation: From a minimal pool of seed questions (e.g., 10 examples), subsets of questions are concatenated in randomized order to prompt the LLM to generate new candidate questions. Generated questions are filtered by self-consistency, retaining only those for which the LLM yields a highly confident consensus answer.
    2. Unsupervised Prompt Generation ("Few-Shot w/ Step-by-Step"): When no human-curated few-shot demonstrations exist, the LLM is prompted in a zero-shot fashion using the prefix A: Let's think step by step. on target questions. The generated reasoning chains are then collected and repurposed as few-shot CoT prompt templates for inference and self-training.
  7. Knowl 7 — Zero-Shot GSM8K Reasoning via Self-Generated Demonstrations

    empirical result

    On GSM8K, generating few-shot demonstration exemplars using zero-shot Let's think step by step greedy decoding ("Few-Shot w/ Step-by-Step") and applying multi-path self-consistency achieves 74.2% accuracy at 40 sampled paths on PaLM 540B.

    This substantially outperforms standard zero-shot Step-by-Step with self-consistency (53.8% at 10 paths, 70.1% at 40 paths) and nearly matches the performance of human-written few-shot CoT prompting (74.4% at 40 paths). The performance gain occurs because self-generating prompt templates enables constructing multiple diverse 4-shot prompt templates (e.g., 20 templates totaling 80 unique generated exemplars at 40 paths) to drive multi-path decoding, surpassing the prompt diversity of the 8 fixed human-written exemplars.

  8. Knowl 8 — Knowledge Distillation of LMSI Reasoning Data to Smaller Models

    data/table

    The self-training data generated by PaLM 540B (without ground truth supervision) can be used to fine-tune smaller language models (PaLM 8B and PaLM 62B).

    Model Configuration PaLM 8B PaLM 62B PaLM 540B
    Base Model (w/o LMSI) 5.0 29.7 56.5
    Distilled from PaLM 540B LMSI Data 33.4 (+28.4) 57.4 (+27.7) –

    Distillation using LMSI-generated rationales allows smaller models to outperform the base model of the next larger tier: distilled PaLM 8B (33.4%) surpasses the pre-trained base PaLM 62B (29.7%), and distilled PaLM 62B (57.4%) surpasses the pre-trained base PaLM 540B (56.5%).

  9. Knowl 9 — Sampling Temperature Shift and Path Efficiency Following Self-Improvement

    empirical result

    Fine-tuning on self-generated high-confidence reasoning paths reduces the entropy of the language model's output probability distribution. Consequently:

    1. Optimal Temperature Shift: The optimal sampling temperature for multi-path Self-Consistency increases from T=0.7T = 0.7 for the pre-trained PaLM 540B base model to T=1.2T = 1.2 for the LMSI fine-tuned model on GSM8K and DROP.
    2. Inference Efficiency: For the LMSI fine-tuned model, Self-Consistency accuracy largely saturates around m=15m = 15 sampled reasoning paths. Furthermore, sampling only m=5m = 5 paths with the LMSI model achieves higher accuracy on GSM8K than sampling m=32m = 32 paths on the base model without LMSI.
  10. Knowl 10 — LMSI Performance on UL2 20B Encoder-Decoder Model

    data/table

    LMSI was evaluated on the 20B parameter UL2 encoder-decoder model using m=40m = 40 sampled paths for majority voting, fine-tuning for 10k steps with learning rate 5×10−55 \times 10^{-5}, and evaluating with T=0.5T = 0.5 (base) and T=0.7T = 0.7 (LMSI).

    Model Method GSM8K DROP ARC-c OpenBookQA ANLI-A2 ANLI-A3
    w/o LMSI CoT 5.4 / 7.1 11.1 / 16.8 49.9 53.6 35.9 33.8
    SC 6.4 / 9.9 16.8 / 26.5 54.7 54.0 37.4 36.8
    LMSI CoT 6.1 / 8.6 11.4 / 17.1 50.9 53.8 35.4 34.4
    SC 7.9 / 10.2 18.1 / 28.1 54.9 55.2 38.1 37.4

    (Note: For GSM8K and DROP, scores indicate exact match accuracy / accuracy after equation-correction postprocessing).

    While LMSI consistently yields positive gains on UL2, the relative improvements are markedly smaller than those observed on the 540B PaLM model, indicating that self-improvement benefits scale with base model capacity.

  11. Knowl 11 — Model Scale and Confidence Calibration Requirements of LMSI

    limitation

    The effectiveness of LMSI relies on two emergent properties of large language models:

    1. In-Context Chain-of-Thought Reasoning Capability: Demonstration-driven intermediate reasoning only functions effectively in models with sufficient scale (e.g., while GPT-3 175B achieves 46.9% on GSM8K few-shot CoT, a 6B GPT-J achieves only 3.1%).
    2. Calibration of Self-Consistency Confidence: Majority voting filters out erroneous self-generated reasoning paths because model calibration improves with parameter scale (>10B> 10\text{B} parameters). In large models, correct answers correlate strongly with high vote consistency, ensuring high pseudo-label purity; in smaller models with poor calibration, noisy and incorrect majority predictions degrade fine-tuning.

Coverage note — The raw prompt texts listed in Appendix Tables 9-17 were omitted as they are standard few-shot exemplars adopted directly from prior work (Wei et al., 2022c; Wang et al., 2022b; Zhou et al., 2022), whereas their structural role in LMSI is fully covered.

References

  1. 1.Massih-Reza Amini, Vasilii Feofanov, Loic Pauletto, Emilie Devijver, and Yury Maximov. 2022. Self-training: A survey.
  2. 2.Jimmy Ba and Rich Caruana. 2014. Do deep nets really need to be deep? Advances in neural information processing systems, 27.
  3. 3.BIG bench collaboration. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. ArXiv, abs/2206.04615.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Neurips.
  5. 5.Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018. e-snli: Natural language inference with natural language explanations. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 9539–9549. Curran Associates, Inc.
  6. 6.Xiaokang Chen, Yuhui Yuan, Gang Zeng, and Jingdong Wang. 2021. Semi-supervised semantic segmentation with cross pseudo supervision. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek B Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Oliveira Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. ArXiv, abs/2204.02311.
  8. 8.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Adams Yu, Albert Webson, Xinyun Chen, Gaurav Mishra, Zhuyun Dai, Shixiang Shane Gu, Mirac Suzgun, Vincent Zhao, Aakanksha Chowdhery, Sharan Narang, Yanping Huang, Andrew Dai, Hongkun Yu, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. In arxiv.
  9. 9.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. ArXiv, abs/1803.05457.
  10. 10.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. ArXiv, abs/2110.14168.
  11. 11.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The pascal recognising textual entailment challenge. In MLCW.
  12. 12.Kumar Shri dhar, Alessandro Stolfo, and Mrinmaya Sachan. 2023. Distilling reasoning capabilities into smaller language models. In ACL.
  13. 13.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In NAACL.
  14. 14.Jacob Eisenstein, Daniel Andor, Bernd Bohnet, Michael Collins, and David Mimno. 2022. Honest students from untrusted teachers: Learning an interpretable question-answering pipeline from a pretrained language model. arXiv preprint arXiv:2210.02498.
  15. 15.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  16. 16.Junxian He, Jiatao Gu, Jiajun Shen, and Marc’Aurelio Ranzato. 2020. Revisiting self-training for neural sequence generation. In International Conference on Learning Representations.
  17. 17.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7).
  18. 18.Namgyu Ho, Laura Schmid, and Se-Young Yun. 2023. Large language models are reasoning teachers. ArXiv.
  19. 19.Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi. 2022. Maieutic prompting: Logically consistent reasoning with recursive explanations.
  20. 20.Saurav Kadavath, Tom Conerly, Amanda Askell, T. J. Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zachary Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, John Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom B. Brown, Jack Clark, Nicholas Joseph, Benjamin Mann, Sam McCandlish, Christopher Olah, and Jared Kaplan. 2022. Language models (mostly) know what they know. ArXiv, abs/2207.05221.
  21. 21.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Neural Information Processing Systems (NeurIPS).
  22. 22.Anoop Korattikara Balan, Vivek Rathod, Kevin P Murphy, and Max Welling. 2015. Bayesian dark knowledge. Advances in neural information processing systems, 28.
  23. 23.Andrew K. Lampinen, Ishita Dasgupta, Stephanie C. Y. Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L. McClelland, Jane X. Wang, and Felix Hill. 2022. Can language models learn from explanations in context?
  24. 24.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Venkatesh Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022. Solving quantitative reasoning problems with language models. ArXiv, abs/2206.14858.
  25. 25.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022a. On the advance of making language models better reasoners.
  26. 26.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022b. On the advance of making language models better reasoners. ArXiv, abs/2206.02336.
  27. 27.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let’s verify step by step. ArXiv, abs/2305.20050.
  28. 28.Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017a. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).
  29. 29.Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017b. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In ACL.
  30. 30.Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn. 2022. Teaching small language models to reason. ArXiv, abs/2212.08410.
  31. 31.Yu Meng, Jiaxin Huang, Yu Zhang, and Jiawei Han. 2022. Generating training data with language models: Towards zero-shot language understanding. ArXiv, abs/2202.04538.
  32. 32.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In EMNLP.
  33. 33.Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020. Wt5?! training text-to-text models to explain their predictions.
  34. 34.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019. Adversarial nli: A new benchmark for natural language understanding. ArXiv, abs/1910.14599.
  35. 35.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  36. 36.Arkil Patel, S. Bhattamishra, and Navin Goyal. 2021. Are nlp models really able to solve simple math word problems? In NAACL.
  37. 37.Danish Pruthi, Rachit Bansal, Bhuwan Dhingra, Livio Baldini Soares, Michael Collins, Zachary C. Lipton, Graham Neubig, and William W. Cohen. 2022. Evaluating Explanations: How Much Do Explanations from the Teacher Aid Students? Transactions of the Association for Computational Linguistics, 10:359–375.
  38. 38.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text. In EMNLP.
  39. 39.Aruni RoyChowdhury, Prithvijit Chakrabarty, Ashish Singh, SouYoung Jin, Huaizu Jiang, Liangliang Cao, and Erik G. Learned-Miller. 2019. Automatic adaptation of object detectors to new domains using self-training. In CVPR, pages 780–790.
  40. 40.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2022. Multitask prompted training enables zero-shot task generalization. In ICLR.
  41. 41.Timo Schick and Hinrich Schütze. 2020a. Exploiting cloze-questions for few-shot text classification and natural language inference. In Conference of the European Chapter of the Association for Computational Linguistics.
  42. 42.Timo Schick and Hinrich Schütze. 2020b. It’s not just size that matters: Small language models are also few-shot learners. ArXiv, abs/2009.07118.
  43. 43.Charlie Snell, Dan Klein, and Ruiqi Zhong. 2022. Learning by distilling context. arXiv preprint arXiv:2209.15189.
  44. 44.Yi Tay, Mostafa Dehghani, Vinh Quang Tran, Xavier García, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Huaixiu Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler. 2022. Ul2: Unifying language learning paradigms.
  45. 45.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. ArXiv, abs/1905.00537.
  46. 46.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. In BlackboxNLP@EMNLP.
  47. 47.Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, and Huan Sun. 2022a. Towards understanding chain-of-thought prompting: An empirical study of what matters. In Annual Meeting of the Association for Computational Linguistics.
  48. 48.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022b. Rationale-augmented ensembles in language models. ArXiv, abs/2207.00747.
  49. 49.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022c. Self-consistency improves chain of thought reasoning in language models. ArXiv, abs/2203.11171.
  50. 50.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  51. 51.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed Huai hsin Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022a. Emergent abilities of large language models. ArXiv, abs/2206.07682.
  52. 52.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022b. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  53. 53.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Brian Ichter, Fei Xia, Quoc Le, and Denny Zhou. 2022c. Chain of thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35.
  54. 54.Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In NAACL.
  55. 55.Zhiheng Xi, Senjie Jin, Yuhao Zhou, Rui Zheng, Songyang Gao, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023. Self-polish: Enhance reasoning in large language models via problem refinement.
  56. 56.Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V. Le. 2020. Self-training with noisy student improves imagenet classification. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695.
  57. 57.Xi Ye and Greg Durrett. 2022. The unreliability of explanations in few-shot in-context learning.
  58. 58.Kang Min Yoo, Dongju Park, Jaewook Kang, Sang-Woo Lee, and Woomyeong Park. 2021. Gpt3mix: Leveraging large-scale language models for text augmentation. In EMNLP Findings.
  59. 59.Omar Zaidan, Jason Eisner, and Christine Piatko. 2007. Using “annotator rationales” to improve machine learning for text categorization. NAACL.
  60. 60.Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. 2022. Star: Bootstrapping reasoning with reasoning.
  61. 61.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alexander J. Smola. 2022. Automatic chain of thought prompting in large language models. ArXiv, abs/2210.03493.
  62. 62.Chuanyang Zheng, Zhengying Liu, Enze Xie, Zhenguo Li, and Yu Li. 2023. Progressive-hint prompting improves reasoning in large language models. ArXiv, abs/2304.09797.
  63. 63.Denny Zhou, Nathanael Scharli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. ArXiv, abs/2205.10625.

Citation

MLA
Huang, J., et al. “Large Language Models Can Self-Improve”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1051–68, https://doi.org/10.18653/v1/2023.emnlp-main.67.
APA
Huang, J., Gu, S., Hou, L., Wu, Y., Wang, X., Yu, H., & Han, J. (2023). Large Language Models Can Self-Improve. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1051–1068. https://doi.org/10.18653/v1/2023.emnlp-main.67
Chicago
Huang, J., S. Gu, L. Hou, et al. 2023. “Large Language Models Can Self-Improve”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1051–68. https://doi.org/10.18653/v1/2023.emnlp-main.67.
Harvard
Huang, J. et al. (2023) “Large Language Models Can Self-Improve”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1051–1068. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.67.
Vancouver
1. Huang J, Gu S, Hou L, Wu Y, Wang X, Yu H, Han J (2023) Large Language Models Can Self-Improve. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1051–1068

BibTeX

@inproceedings{huang-etal-2023-large,
    title = "Large Language Models Can Self-Improve",
    author = "Huang, Jiaxin  and
      Gu, Shixiang  and
      Hou, Le  and
      Wu, Yuexin  and
      Wang, Xuezhi  and
      Yu, Hongkun  and
      Han, Jiawei",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.67/",
    doi = "10.18653/v1/2023.emnlp-main.67",
    pages = "1051--1068"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/