Large Language Models Can Self-Improve
Jiaxin HuangShixiang GuLe HouYuexin WuXuezhi WangHongkun YuJiawei Han
Demonstrates that large language models can boost their reasoning capabilities on unlabeled datasets by fine-tuning on their own high-confidence, chain-of-thought solutions selected via majority voting.
Improving large language models typically requires massive volumes of expensive, human-annotated data. This requirement creates major bottlenecks for organizations working in specialized domains or low-resource settings where labeled training data is scarce. Meanwhile, humans can enhance their problem-solving abilities purely through internal reflection and practice. The article addresses whether large language models can likewise self-improve their reasoning capabilities using only unlabeled questions, removing the dependency on human-provided answers.
The article demonstrates and evaluates a self-training framework called Language Model Self-Improved (LMSI). The objective is to verify whether a model can generate its own step-by-step reasoning paths, filter out incorrect outputs without human supervision, and use those self-generated solutions to fine-tune itself into a more capable system.
The researchers evaluated this framework using a 540-billion-parameter language model across multiple arithmetic, commonsense, and natural language inference benchmarks. The approach prompts the model with a few examples to generate multiple diverse reasoning paths for each unlabeled question. It then applies majority voting—selecting the most frequent final answer—to identify high-confidence solutions. Finally, the model is fine-tuned on these self-selected reasoning paths formatted across multiple prompt styles. Credibility is further established through out-of-domain testing, ablation studies on data formatting, and experiments where the model also self-generates training questions and prompt demonstrations from scratch.
The evaluation revealed several critical findings. First, self-improvement without ground-truth answers significantly boosted performance across all tested tasks; for example, accuracy on the GSM8K math benchmark rose from 74.4% to 82.1%, while natural language inference on ANLI-A3 improved from 63.4% to 67.9%. Second, the approach generalized to out-of-domain tasks, raising performance across six completely unseen benchmarks by up to 10.4 percentage points. Third, when human prompts and questions were absent, the model self-generated questions and reasoning demonstrations, achieving a state-of-the-art zero-shot accuracy of 74.2% on GSM8K. Fourth, knowledge from the self-improved 540-billion model was successfully transferred to smaller systems; an 8-billion model distilled this way reached 33.4% accuracy, outperforming an un-fine-tuned 62-billion model (29.7%), while a distilled 62-billion model (57.4%) surpassed the base 540-billion model (56.5%). Finally, post-improvement inference required far fewer sampled paths—using just 5 paths after self-improvement exceeded the performance of using 32 paths on the base model.
These findings indicate that large models already contain latent reasoning capabilities that can be activated without costly manual labeling. For enterprise deployments, this drastically lowers data curation costs and mitigates compliance risks tied to manual annotation. Furthermore, the ability to distill reasoning into smaller models and reduce test-time sampling provides direct paths toward slashing operational compute expenses, lowering server latency, and improving environmental efficiency.
Organizations developing or deploying large language models should consider adopting self-improvement and distillation pipelines for reasoning-heavy workflows rather than immediately funding large data-annotation efforts. When implementing this workflow, practitioners should train on mixed prompt formats to avoid overfitting to specific instruction styles, tune sampling temperatures upward during evaluation, and test distilled smaller architectures to reduce operational costs. Before broad deployment, teams should conduct internal pilots to measure return on investment and assess task-specific accuracy.
The primary limitation of this method is its dependency on initial model scale; smaller baseline architectures lack the calibration and in-context reasoning needed to reliably generate high-confidence pseudo-labels. Consequently, organizations should apply this self-improvement technique directly to high-capacity frontier models, using smaller models primarily as downstream recipients of distilled knowledge.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Introduces the self-consistency decoding mechanism over diverse chain-of-thought reasoning paths that the source directly relies on to sample high-confidence pseudo-labels.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Establishes chain-of-thought prompting as the foundational mechanism for eliciting step-by-step reasoning in large language models upon which self-improvement builds.
- Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). Pioneers the bootstrapping paradigm (STaR) of fine-tuning language models on their own generated rationales, providing the direct conceptual predecessor to unsupervised self-improvement.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Demonstrates zero-shot step-by-step reasoning elicitation, essential context for prompting models to reason over unlabeled inputs without ground-truth supervision.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Presents problem decomposition strategies that underpin multi-step reasoning capabilities leveraged during self-generated rationale creation.
- Paper: On-Policy Self-Distillation without Any Supervision, Yijiang Li et al.. Extends unsupervised self-improvement by using on-policy token-level self-distillation from consensus rationales without requiring external labels.
- Paper: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek-AI et al. (2025). Advances autonomous model self-improvement from supervised fine-tuning on self-consistent paths to large-scale reinforcement learning with self-verification.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). Generalizes the self-taught reasoning paradigm to meta-level cognitive search and trial-and-error verification beyond standard chain-of-thought generation.
- Paper: Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability, Shobhita Sundaram et al. (2026). Continues the self-improvement paradigm by training models to generate their own synthetic curriculum when self-consistency on hard problems fails.
- Paper: Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation, Xiaoying Zhang et al. (2024). Applies self-evaluation and preference optimization principles to autonomously align model factuality and mitigate hallucinations without human supervision.
- Paper: Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models, Xingyu Chen et al. (2025). Critiques and refines extended reasoning models produced by self-improvement frameworks by curbing overthinking on simpler queries.
