The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning

Seungone KimSe June JooDoyoung KimJoel JangSeonghyeon YeJamin ShinMinjoon Seo

article2023EMNLP190 citations

Presents the CoT Collection, an instruction-tuning dataset of 1.84 million step-by-step rationales across 1,060 tasks that enables language models under 100 billion parameters to significantly improve their zero-shot and few-shot reasoning capabilities on unseen tasks.

Listen

Large language models excel at complex reasoning when prompted to explain their step-by-step thinking before providing a final answer. However, smaller models with fewer than 100 billion parameters struggle with this capability, limiting their practical deployment in cost-sensitive and compute-constrained environments. Prior efforts to address this gap primarily focused on single-task tuning or used narrow rationale datasets, leaving smaller models incapable of generalizing step-by-step reasoning across diverse, unfamiliar tasks.

The article demonstrates that fine-tuning smaller language models on a vast, diverse collection of step-by-step explanations significantly enhances their zero-shot reasoning and few-shot adaptation capabilities on unseen tasks.

To achieve this, the authors created the CoT Collection, an instruction-tuning dataset containing 1.84 million explanations generated across 1,060 tasks using a large proprietary language model. They systematically filtered out low-quality and inconsistent explanations, ensuring high dataset quality. Using this data, they fine-tuned standard 3-billion and 11-billion parameter models (Flan-T5) to create specialized reasoning models called CoT-T5, subsequently evaluating their performance across major academic benchmarks, domain-specific adaptation tasks, and multilingual settings.

The key findings show significant performance improvements across multiple settings. First, in zero-shot evaluations on the challenging BIG-Bench Hard benchmark, CoT-T5 improved accuracy by 4.34% for the 3-billion parameter model and 2.60% for the 11-billion parameter model compared to baseline Flan-T5 models. Second, task diversity proved far more critical than instance volume; fine-tuning on only 10,000 instances spread across 1,060 diverse tasks yielded better reasoning performance than training on 180,000 instances from only 9 tasks. Third, in few-shot domain adaptation across medical and legal benchmarks, fine-tuning CoT-T5 with parameter-efficient techniques outperformed proprietary large models like ChatGPT by 13.98% and Claude by 8.11%. Finally, multilingual experiments indicated that applying modest amounts of translated explanation data (60,000 to 80,000 instances) enabled smaller models to escape near-zero performance and achieve 2x to 10x accuracy gains on non-English reasoning tasks.

These findings demonstrate that organizations do not necessarily need massive, computationally expensive proprietary models to achieve strong reasoning performance on specialized tasks. Fine-tuning compact, open-source models with diverse step-by-step reasoning data drastically reduces inference costs and latency while matching or exceeding the capabilities of commercial cloud models. Furthermore, parameter-efficient fine-tuning allows these models to retain their core reasoning abilities while adapting to proprietary domains.

For practitioners seeking to deploy cost-effective reasoning systems, the recommended approach is to adopt parameter-efficient fine-tuning on compact models pre-trained on diverse rationale collections. When preparing training datasets, organizations should prioritize broad task variety over raw volume within a few tasks. For multilingual applications, introducing small sets of targeted translation data provides a viable adaptation path.

Certain limitations should be noted. The underlying models were evaluated on structured academic benchmarks and are not optimized for open-ended conversational chat applications. Additionally, multilingual evaluations focused on direct per-language adaptation rather than cross-lingual transfer, and reliance on proprietary models to generate training rationales introduces potential reproducibility dependencies. Nevertheless, the experimental results provide high confidence that diverse explanation data effectively equips compact models with robust multi-step reasoning capabilities.

Cover for The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning

Abstract

Language models (LMs) with less than 100B parameters are known to perform poorly on chain-of-thought (CoT) reasoning in contrast to large LMs when solving unseen tasks. In this work, we aim to equip smaller LMs with the step-by-step reasoning capability by instruction tuning with CoT rationales. In order to achieve this goal, we first introduce a new instruction-tuning dataset called the CoT Collection, which augments the existing Flan Collection (including only 9 CoT tasks) with additional 1.84 million rationales across 1,060 tasks. We show that CoT fine-tuning Flan-T5 (3B & 11B) with CoT Collection enables smaller LMs to have better CoT capabilities on unseen tasks. On the BIG-Bench-Hard (BBH) benchmark, we report an average improvement of +4.34% (Flan-T5 3B) and +2.60% (Flan-T5 11B), in terms of zero-shot task accuracy. Furthermore, we show that instruction tuning with CoT Collection allows LMs to possess stronger few-shot learning capabilities on 4 domain-specific tasks, resulting in an improvement of +2.24% (Flan-T5 3B) and +2.37% (Flan-T5 11B), even outperforming ChatGPT utilizing demonstrations until the max length by a +13.98% margin. Our code, the CoT Collection data, and model checkpoints are publicly available¹.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Chain-of-Thought (CoT) Prompting
  • 2.2 Improving Zero-shot Generalization
  • 2.3 Improving Few-Shot Learning
  • 3 The COT Collection
  • 4 Experiments
  • 4.1 Evaluation
  • 4.2 Zero-shot Generalization
  • 4.3 Few-shot Generalization
  • 5 Analysis of CoT Fine-tuning
  • 5.1 Scaling the number of tasks & instances
  • 5.2 In-domain Task Accuracy of CoT-T5
  • 6 Conclusion
  • Acknowledgments
  • Limitations
  • References
  • Appendix A Analysis of COT Collection
  • Appendix B Filtering COT Collection

Knowls

  1. Knowl 1 — COT COLLECTION expands rationale supervision to 1,060 instruction tasks

    data/table

    COT COLLECTION is an instruction-tuning dataset containing 1.84 million chain-of-thought (CoT) rationales across 1,060 tasks drawn from the Flan Collection, a collection of 1,836 NLP tasks. The authors selected tasks from P3, Super-NaturalInstructions (SNI), Flan, and additional dialogue and code sources. When the same dataset appeared in multiple sources, they prioritized P3, then SNI, then Flan.

    The selection excluded tasks whose long target outputs would cause the rationale-plus-answer target to exceed 512 tokens; tasks that were not publicly available; tasks with mismatched input and output records; and tasks for which preliminary rationale generation produced very short, uninformative explanations (including some sentiment, sentence-completion, coreference, and word-disambiguation tasks). The resulting collection was designed to supply rationale supervision across substantially more tasks than the limited public CoT datasets available to the authors.

  2. Knowl 2 — Rationales are generated with task-family demonstrations and filtered before inclusion

    model/method

    For each selected task instance, the input comprises an instruction and instance, and the known ground-truth answer is used to elicit a rationale from OpenAI Codex through in-context learning. To avoid writing separate demonstrations for every task, the authors grouped tasks into 26 format-based families and prepared 6–8 demonstrations per family. Three authors developed these demonstrations: two wrote candidate rationales for sampled instances and a third compared the candidates. In the demonstrations, the answer is placed before the rationale; the authors report that this ordering helped the generator focus on producing a rationale rather than solving the task itself.

    The authors generated five candidate rationales per instance and filtered out candidates that did not contain the ground-truth answer as a whitespace-separated token, exceeded the 512-token combined rationale-and-answer limit, duplicated an earlier rationale, or contained repetitive sentences. They also applied trigger-based filtering to remove outputs that degenerated into code. Filtering was particularly relevant to arithmetic tasks, where the generator sometimes produced inconsistent or otherwise poor rationales.

  3. Knowl 3 — CoT-T5 is trained to generate a rationale before its answer

    model/method

    CoT-T5 is obtained by fine-tuning Flan-T5 on COT COLLECTION. For an input instruction and instance, the target sequence consists of the rationale followed by the answer. The phrase “Let’s think step by step” is included during training and evaluation to cue rationale generation.

    The authors trained for one epoch. CoT-T5-3B used batch size 64, learning rate 5×10−55\times10^{-5}, and AdamW; CoT-T5-11B used batch size 8, learning rate 10−410^{-4}, and Adafactor. Both settings used gradient accumulation of 8 and eight 80-GB A100 GPUs; the reported training times were one day for 3B and seven days for 11B. For CoT evaluation, generation was constrained to produce a rationale of at least eight tokens. Classification evaluation then applied verbalizers after an “[ANSWER]” indicator, while generation evaluation extracted the answer after that indicator. During evaluation, the authors used nucleus sampling with p=0.8p=0.8 and a three-gram repetition constraint.

  4. Knowl 4 — CoT-T5 improves zero-shot performance on BIG-Bench Hard

    empirical result

    The authors evaluated models zero-shot on all 27 BIG-Bench Hard (BBH) datasets, including generation tasks, using direct evaluation and CoT evaluation. The following are the reported CoT, direct, and total-average scores, respectively:

    • Flan-T5-3B: 34.06, 37.14, 35.60; CoT-T5-3B: 38.40, 36.18, 37.29.
    • Flan-T5-11B: 38.57, 40.99, 39.78; CoT-T5-11B: 42.20, 42.56, 42.38.
    • T5-3B fine-tuned on COT COLLECTION without Flan instruction tuning: 37.95, 35.52, 36.74; the corresponding 11B model: 40.02, 38.76, 39.54.

    Thus, CoT-T5-3B’s CoT score was 4.34 points above Flan-T5-3B’s, while CoT-T5-11B’s total average was 2.60 points above Flan-T5-11B’s. The direct score fell by 0.96 points for the 3B pair but rose by 1.57 points for the 11B pair. The results also show that rationale fine-tuning improved the T5-LM baselines’ CoT scores over the corresponding Flan-T5 models, while combining Flan instruction tuning with COT COLLECTION yielded the strongest total averages for the two model sizes.

  5. Knowl 5 — CoT fine-tuning improves few-shot adaptation on legal and medical tasks

    empirical result

    The few-shot evaluation used LEDGAR and Case Hold (legal) and MedNLI and PubMedQA (medical), with 64 randomly sampled training instances per dataset. Rationales for those instances were augmented using the collection procedure, and results are averages over three random seeds. Scores below are accuracy in the order LEDGAR, Case Hold, MedNLI, PubMedQA, and total average.

    • Flan-T5-3B, full fine-tuning: 52.60, 61.40, 66.82, 66.28, 61.78; Flan-T5-3B, full CoT fine-tuning: 53.60, 58.80, 65.89, 65.89, 61.05; CoT-T5-3B, full CoT fine-tuning: 51.90, 60.60, 67.16, 68.12, 61.95.
    • Flan-T5-3B, LoRA fine-tuning: 53.20, 58.80, 61.60, 67.18, 60.19; Flan-T5-3B, LoRA CoT fine-tuning: 51.20, 61.60, 62.59, 66.06, 60.36; CoT-T5-3B, LoRA CoT fine-tuning: 54.80, 63.60, 68.00, 69.66, 64.02.
    • Flan-T5-11B, LoRA fine-tuning: 55.30, 64.90, 75.91, 70.25, 66.59; Flan-T5-11B, LoRA CoT fine-tuning: 52.10, 65.50, 71.63, 71.60, 65.21; CoT-T5-11B, LoRA CoT fine-tuning: 56.10, 68.30, 78.02, 73.42, 68.96.
    • Claude with in-context demonstrations: 55.70, 57.20, 75.94, 54.58, 60.85; ChatGPT with in-context demonstrations: 51.70, 32.10, 70.53, 65.59, 54.98.

    The best reported average was 68.96 for CoT-T5-11B with LoRA CoT fine-tuning; CoT-T5-3B with LoRA CoT fine-tuning averaged 64.02. LoRA used rank 4 and 1,000 steps, training 2.35 million parameters at 3B scale or 4.72 million at 11B scale. The proprietary-model in-context baselines used demonstrations up to their maximum input lengths (4,000 tokens for ChatGPT and 9,000 for Claude).

  6. Knowl 6 — CoT fine-tuning on 163 tasks improves P3 evaluation accuracy

    empirical result

    To test whether CoT fine-tuning required the full 1,060-task collection, the authors fine-tuned T5-LM-3B and T0-3B on the COT COLLECTION subset corresponding to T0’s P3 training tasks: 644,000 instances across 163 tasks. T0 had originally been trained on 12 million instances, about 18.63 times as many. Evaluation used 11 P3 datasets, with Flan-T5 and CoT-T5 excluded because their training tasks overlapped the evaluation tasks.

    T0-3B’s reported total-average accuracy was 51.43. After CoT fine-tuning, T0-3B scored 59.67 with direct evaluation and 60.08 with CoT evaluation; the latter is an 8.65-point increase over the T0-3B baseline. T5-3B fine-tuned on the subset scored 56.99 with direct evaluation and 59.94 with CoT evaluation. These results demonstrate gains in this setup using a substantially smaller training set than T0’s original training data.

  7. Knowl 7 — Language-specific CoT fine-tuning raises MGSM scores in five languages

    empirical result

    The authors translated 60,000–80,000 COT COLLECTION instances into Korean, Russian, French, Chinese, and Japanese, then separately fine-tuned mT5-3.7B and mT0-3.7B on each target language. Evaluation was zero-shot with CoT prompting on the corresponding MGSM language subset; GPT-3’s comparison scores used six-shot prompting. Scores are listed in the order Korean, Russian, French, Chinese, Japanese.

    • mT5-3.7B before fine-tuning: 0.0, 1.2, 2.0, 0.8, 0.8; after language-specific CoT fine-tuning: 3.2, 6.8, 9.6, 6.0, 7.6.
    • mT0-3.7B before fine-tuning: 0.0, 4.8, 7.2, 1.6, 2.4; after language-specific CoT fine-tuning: 7.6, 10.4, 15.6, 11.2, 11.0.

    The adapted models improved in every reported language, including languages for which the untuned models were at or near zero. This experiment tests adaptation within a single target language; it does not establish transfer of CoT ability between languages.

  8. Knowl 8 — Task breadth mattered more than instance count in the BBH scaling comparison

    empirical result

    The page 9 scaling plot compares Flan-T5 models evaluated with CoT prompting on BBH after training with different amounts and breadths of rationale data. Training with the existing nine CoT tasks and 180,000 instances produced a score of 40.33. Training with COT COLLECTION across 1,060 tasks produced 42.04 with 10,000 instances, 44.76 with 100,000 instances, and 47.60 with the full 1.84 million instances. In this comparison, the 10,000-instance, 1,060-task condition exceeded the 180,000-instance, nine-task condition, supporting the authors’ conclusion that breadth of task coverage was especially valuable in this experiment.

  9. Knowl 9 — CoT fine-tuning did not reduce accuracy on the tested in-domain tasks

    empirical result

    The authors compared Flan-T5-11B with CoT-T5-11B using CoT evaluation on five tasks used during training: ANLI-R1, ANLI-R2, ANLI-R3, RTE, and WinoGrande. The page 9 bar chart reports higher accuracy for CoT-T5-11B than for Flan-T5-11B on each of the five tasks. This is evidence against performance loss on these particular in-domain tests, not a general demonstration that catastrophic forgetting is absent: the authors note that the evaluated tasks had already been used to train Flan-T5 and CoT-T5 and that results could differ for additional tasks.

  10. Knowl 10 — The demonstrated scope does not establish chat capability or cross-lingual CoT transfer

    limitation

    CoT-T5 was fine-tuned on academic instruction and benchmark tasks rather than long-form dialogue, so the paper does not establish that it can serve as a chat model or reproduce the dialogue behavior of models trained on long-form conversations. The primary model also inherits Flan-T5’s language limitations; the multilingual experiments show gains from separate, target-language fine-tuning but do not test cross-lingual transfer. Finally, rationale augmentation depended on OpenAI Codex, which was no longer supported when the paper was written. The authors identify reproducibility and rationale quality as concerns and note that other generators and improved prompting could change the resulting data.

Coverage note — The rationale-quality and diversity analyses—including ROSCOE comparisons with human rationales and other language-model generators—are omitted because they are supporting diagnostics rather than load-bearing methods or downstream findings; the dataset, augmentation procedure, principal evaluations, analyses, and stated limitations are included.

References

  1. 1.Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021. Explanations for commonsenseqa: New dataset and models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3050–3065.
  2. 2.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2357–2367.
  3. 3.Anthropic. 2023. Claude. https://www.anthropic.com/index/introducing-claude.
  4. 4.Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q Tran, Dara Bahri, Jianmo Ni, et al. 2021. Ext5: Towards extreme multi-task scaling for transfer learning. arXiv preprint arXiv:2111.10952.
  5. 5.Akari Asai, Mohammadreza Salehi, Matthew E Peters, and Hannaneh Hajishirzi. 2022. Attentional mixtures of soft prompt tuning for parameter-efficient multi-task knowledge sharing. arXiv preprint arXiv:2205.11961.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  7. 7.Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018. e-snli: Natural language inference with natural language explanations. Advances in Neural Information Processing Systems, 31.
  8. 8.Sanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che, Ting Liu, and Xiangzhan Yu. 2020. Recall and learn: Fine-tuning deep pretrained language models with less forgetting. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7870–7881.
  9. 9.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  10. 10.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  11. 11.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  12. 12.Yao Fu, Hao Peng, Litu Ou, Ashish Sabharwal, and Tushar Khot. 2023. Specializing smaller language models towards multi-step reasoning. arXiv preprint arXiv:2301.12726.
  13. 13.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  14. 14.Olga Golovneva, Moya Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2022. Roscoe: A suite of metrics for scoring step-by-step reasoning. arXiv preprint arXiv:2212.07919.
  15. 15.Google. 2023. Bard. https://blog.google/technology/ai/bard-google-ai-search-updates/.
  16. 16.Namgyu Ho, Laura Schmid, and Se-Young Yun. 2022. Large language models are reasoning teachers. arXiv preprint arXiv:2212.10071.
  17. 17.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751.
  18. 18.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022. Unnatural instructions: Tuning language models with (almost) no human labor. arXiv preprint arXiv:2212.09689.
  19. 19.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  20. 20.Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Dániel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, et al. 2022. Opt-iml: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint arXiv:2212.12017.
  21. 21.Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2023. Exploring the benefits of training expert language models over instruction tuning. arXiv preprint arXiv:2302.03202.
  22. 22.Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, and Minjoon Seo. 2021. Towards continual knowledge learning of language models. arXiv preprint arXiv:2110.03215.
  23. 23.Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019. Pubmedqa: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2567–2577.
  24. 24.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  25. 25.Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal. 2020. Qasc: A dataset for question answering via sentence composition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8082–8090.
  26. 26.Hyunwoo Kim, Jack Hessel, Liwei Jiang, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Le Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, et al. 2022. Soda: Million-scale dialogue distillation with social commonsense contextualization. arXiv preprint arXiv:2212.10465.
  27. 27.Seungone Kim, Se June Joo, Yul Jang, Hyungjoo Chae, and Jinyoung Yeo. 2023. Cotever: Chain of thought prompting annotation toolkit for explanation verification. arXiv preprint arXiv:2303.03628.
  28. 28.Nikita Kitaev, Steven Cao, and Dan Klein. 2019. Multilingual constituency parsing with self-attention and pre-training. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3499–3505.
  29. 29.Nikita Kitaev and Dan Klein. 2018. Constituency parsing with a self-attentive encoder. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2676–2686.
  30. 30.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.
  31. 31.Matthew Lamm, Jennimaria Palomaki, Chris Alberti, Daniel Andor, Eunsol Choi, Livio Baldini Soares, and Michael Collins. 2021. Qed: A framework and dataset for explanations in question answering. Transactions of the Association for computational Linguistics, 9:790–806.
  32. 32.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059.
  33. 33.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, et al. 2021. Datasets: A community library for natural language processing. arXiv preprint arXiv:2109.02846.
  34. 34.Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, and Yangqiu Song. 2023. Multi-step jailbreaking privacy attacks on chatgpt. arXiv preprint arXiv:2304.05197.
  35. 35.Alisa Liu, Swabha Swayamdipta, Noah A Smith, and Yejin Choi. 2022a. Wanli: Worker and ai collaboration for natural language inference dataset creation. arXiv preprint arXiv:2201.05955.
  36. 36.Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel. 2022b. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. Advances in Neural Information Processing Systems, 35:1950–1965.
  37. 37.Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022c. P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 61–68.
  38. 38.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021. Gpt understands, too. arXiv preprint arXiv:2103.10385.
  39. 39.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. 2023. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688.
  40. 40.David Mhlanga. 2023. Open ai in education, the responsible and ethical use of chatgpt towards lifelong learning. Education, the Responsible and Ethical Use of ChatGPT Towards Lifelong Learning (February 11, 2023).
  41. 41.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022. Metaicl: Learning to learn in context. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2791–2809.
  42. 42.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487.
  43. 43.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022. Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786.
  44. 44.Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. 2023. Orca: Progressive learning from complex explanation traces of gpt-4. arXiv preprint arXiv:2306.02707.
  45. 45.Yasumasa Onoe, Michael JQ Zhang, Eunsol Choi, and Greg Durrett. 2021. Creak: A dataset for commonsense reasoning over entity knowledge. arXiv preprint arXiv:2109.01653.
  46. 46.OpenAI. 2022. ChatGPT. https://openai.com/blog/chatgpt.
  47. 47.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  48. 48.Xiaoman Pan, Wenlin Yao, Hongming Zhang, Dian Yu, Dong Yu, and Jianshu Chen. 2022. Knowledge-in-context: Towards knowledgeable semi-parametric language models. arXiv preprint arXiv:2210.16433.
  49. 49.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  50. 50.Alexey Romanov and Chaitanya Shivade. 2018. Lessons from natural language inference in the clinical domain. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1586–1596.
  51. 51.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  52. 52.Timo Schick and Hinrich Schütze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269.
  53. 53.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. 2022. Language models are multilingual chain-of-thought reasoners. arXiv preprint arXiv:2210.03057.
  54. 54.Kumar Shridhar, Alessandro Stolfo, and Mrinmaya Sachan. 2022. Distilling multi-step reasoning capabilities of large language models into smaller models via semantic decompositions. arXiv preprint arXiv:2212.00193.
  55. 55.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. 2022. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261.
  56. 56.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  57. 57.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier García, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. 2022. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131.
  58. 58.Don Tuggener, Pius Von Däniken, Thomas Peetz, and Mark Cieliebak. 2020. Ledgar: a large-scale multi-label corpus for text classification of legal provisions in contracts. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 1235–1241.
  59. 59.Cunxiang Wang, Shuailong Liang, Yue Zhang, Xiaonan Li, and Tian Gao. 2019. Does it make sense? and why? a pilot study for sense making and explanation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4020–4026.
  60. 60.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022a. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  61. 61.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A Smith, Iz Beltagy, et al. 2023. How far can camels go? exploring the state of instruction tuning on open resources. arXiv preprint arXiv:2306.04751.
  62. 62.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. 2022b. Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. URL https://arxiv. org/abs/2204.07705.
  63. 63.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  64. 64.Jason Wei, Yi Tay, and Quoc V Le. 2022a. Inverse scaling can become u-shaped. arXiv preprint arXiv:2211.02011.
  65. 65.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  66. 66.Peter West, Chandra Bhagavatula, Jack Hessel, Jena Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2022. Symbolic knowledge distillation: from general language models to commonsense models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4602–4625.
  67. 67.Haike Xu, Zongyu Lin, Jing Zhou, Yanan Zheng, and Zhilin Yang. 2022. A universal discriminator for zero-shot generalization. arXiv preprint arXiv:2211.08099.
  68. 68.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mt5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498.
  69. 69.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601.
  70. 70.Michihiro Yasunaga and Percy Liang. 2020. Graph-based, self-supervised program repair from diagnostic feedback. In International Conference on Machine Learning, pages 10799–10808. PMLR.
  71. 71.Seonghyeon Ye, Hyeonbin Hwang, Sohee Yang, Hyeongu Yun, Yireun Kim, and Minjoon Seo. 2023. In-context instruction learning. arXiv preprint arXiv:2302.14691.
  72. 72.Seonghyeon Ye, Doyoung Kim, Joel Jang, Joongbo Shin, and Minjoon Seo. 2022. Guess the instruction! making language models stronger zero-shot learners. arXiv preprint arXiv:2210.02969.
  73. 73.Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2023. Mammoth: Building math generalist models through hybrid instruction tuning. arXiv preprint arXiv:2309.05653.
  74. 74.Eric Zelikman, Yuhuai Wu, and Noah D Goodman. 2022. Star: Bootstrapping reasoning with reasoning. arXiv preprint arXiv:2203.14465.
  75. 75.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.
  76. 76.Lucia Zheng, Neel Guha, Brandon R Anderson, Peter Henderson, and Daniel E Ho. 2021. When does pretraining help? assessing self-supervised learning for law and the casehold dataset of 53,000+ legal holdings. In Proceedings of the eighteenth international conference on artificial intelligence and law, pages 159–168.
  77. 77.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625.

Citation

MLA
Kim, S., et al. “The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 12685–708, https://doi.org/10.18653/v1/2023.emnlp-main.782.
APA
Kim, S., Joo, S., Kim, D., Jang, J., Ye, S., Shin, J., & Seo, M. (2023). The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12685–12708. https://doi.org/10.18653/v1/2023.emnlp-main.782
Chicago
Kim, S., S. Joo, D. Kim, et al. 2023. “The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12685–708. https://doi.org/10.18653/v1/2023.emnlp-main.782.
Harvard
Kim, S. et al. (2023) “The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 12685–12708. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.782.
Vancouver
1. Kim S, Joo S, Kim D, Jang J, Ye S, Shin J, Seo M (2023) The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 12685–12708

BibTeX

@inproceedings{kim-etal-2023-cot,
    title = "The {C}o{T} Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning",
    author = "Kim, Seungone  and
      Joo, Se  and
      Kim, Doyoung  and
      Jang, Joel  and
      Ye, Seonghyeon  and
      Shin, Jamin  and
      Seo, Minjoon",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.782/",
    doi = "10.18653/v1/2023.emnlp-main.782",
    pages = "12685--12708"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/