The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning
Seungone KimSe June JooDoyoung KimJoel JangSeonghyeon YeJamin ShinMinjoon Seo
Presents the CoT Collection, an instruction-tuning dataset of 1.84 million step-by-step rationales across 1,060 tasks that enables language models under 100 billion parameters to significantly improve their zero-shot and few-shot reasoning capabilities on unseen tasks.
Large language models excel at complex reasoning when prompted to explain their step-by-step thinking before providing a final answer. However, smaller models with fewer than 100 billion parameters struggle with this capability, limiting their practical deployment in cost-sensitive and compute-constrained environments. Prior efforts to address this gap primarily focused on single-task tuning or used narrow rationale datasets, leaving smaller models incapable of generalizing step-by-step reasoning across diverse, unfamiliar tasks.
The article demonstrates that fine-tuning smaller language models on a vast, diverse collection of step-by-step explanations significantly enhances their zero-shot reasoning and few-shot adaptation capabilities on unseen tasks.
To achieve this, the authors created the CoT Collection, an instruction-tuning dataset containing 1.84 million explanations generated across 1,060 tasks using a large proprietary language model. They systematically filtered out low-quality and inconsistent explanations, ensuring high dataset quality. Using this data, they fine-tuned standard 3-billion and 11-billion parameter models (Flan-T5) to create specialized reasoning models called CoT-T5, subsequently evaluating their performance across major academic benchmarks, domain-specific adaptation tasks, and multilingual settings.
The key findings show significant performance improvements across multiple settings. First, in zero-shot evaluations on the challenging BIG-Bench Hard benchmark, CoT-T5 improved accuracy by 4.34% for the 3-billion parameter model and 2.60% for the 11-billion parameter model compared to baseline Flan-T5 models. Second, task diversity proved far more critical than instance volume; fine-tuning on only 10,000 instances spread across 1,060 diverse tasks yielded better reasoning performance than training on 180,000 instances from only 9 tasks. Third, in few-shot domain adaptation across medical and legal benchmarks, fine-tuning CoT-T5 with parameter-efficient techniques outperformed proprietary large models like ChatGPT by 13.98% and Claude by 8.11%. Finally, multilingual experiments indicated that applying modest amounts of translated explanation data (60,000 to 80,000 instances) enabled smaller models to escape near-zero performance and achieve 2x to 10x accuracy gains on non-English reasoning tasks.
These findings demonstrate that organizations do not necessarily need massive, computationally expensive proprietary models to achieve strong reasoning performance on specialized tasks. Fine-tuning compact, open-source models with diverse step-by-step reasoning data drastically reduces inference costs and latency while matching or exceeding the capabilities of commercial cloud models. Furthermore, parameter-efficient fine-tuning allows these models to retain their core reasoning abilities while adapting to proprietary domains.
For practitioners seeking to deploy cost-effective reasoning systems, the recommended approach is to adopt parameter-efficient fine-tuning on compact models pre-trained on diverse rationale collections. When preparing training datasets, organizations should prioritize broad task variety over raw volume within a few tasks. For multilingual applications, introducing small sets of targeted translation data provides a viable adaptation path.
Certain limitations should be noted. The underlying models were evaluated on structured academic benchmarks and are not optimized for open-ended conversational chat applications. Additionally, multilingual evaluations focused on direct per-language adaptation rather than cross-lingual transfer, and reliance on proprietary models to generate training rationales introduces potential reproducibility dependencies. Nevertheless, the experimental results provide high confidence that diverse explanation data effectively equips compact models with robust multi-step reasoning capabilities.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). FLAN establishes the instruction-tuning framework and zero-shot generalization that the CoT Collection expands with reasoning rationales.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This paper establishes how chain-of-thought demonstrations elicit step-by-step reasoning, the capability the source transfers into smaller models through fine-tuning.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Its zero-shot chain-of-thought results introduce the reasoning behavior that the source seeks to teach smaller language models.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). This paper defines BIG-Bench Hard and tests chain-of-thought prompting on it, preparing readers for the source’s evaluation of CoT fine-tuning on BBH.
- Paper: Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes, Cheng-Yu Hsieh et al. (2023). Building on rationale-based supervision for compact models, it distills reasoning steps as an additional training signal while reducing the examples needed.
- Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). It continues the effort to teach small models multi-step reasoning by distilling teacher-generated chains across arithmetic, commonsense, and symbolic tasks.
- Paper: Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor, Or Honovich et al. (2023). It extends synthetic instruction-data creation with a pipeline that uses language models to generate and diversify training tasks with little human input.
- Paper: MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning, Zhiyang Xu et al. (2023). It applies instruction tuning to multimodal tasks, extending the source’s approach to zero-shot generalization beyond text-only language models.
- Paper: Scaling Instruction-Finetuned Language Models, Hyung Won Chung et al. (2024). It broadens the instruction-tuning picture by examining how task diversity, model scale, and chain-of-thought data jointly affect generalization.
