Teaching Small Language Models to Reason

Lucie Charlotte MagisterJonathan MallinsonJakub AdámekEric MalmiAliaksei Severyn

article2023ACL354 citations

Demonstrates that fine-tuning smaller language models on chain-of-thought rationales distilled from giant teacher models significantly improves their reasoning capabilities across arithmetic, commonsense, and symbolic benchmarks.

Listen

Large language models can solve complex multi-step problems using chain-of-thought prompting, which encourages them to generate intermediate reasoning steps before providing an answer. However, this capability typically emerges only in massive models with tens or hundreds of billions of parameters. Smaller models under 10 billion parameters often fail to generate logical intermediate reasoning and can even experience decreased accuracy when prompted this way. Because deploying giant models is computationally expensive and difficult to scale in production environments, transferring these reasoning abilities to smaller, more efficient models has become an important priority.

The article demonstrates a two-step knowledge distillation pipeline that transfers multi-step reasoning capabilities from large teacher models to substantially smaller student models. The researchers evaluate this approach across arithmetic, commonsense, and symbolic reasoning benchmarks to determine whether smaller models can successfully learn to reason without requiring real-time prompting tricks.

The approach consists of generating step-by-step reasoning solutions for existing training datasets using large teacher models—specifically PaLM with 540 billion parameters and GPT-3 with 175 billion parameters. Crucially, the authors supply the final target answer in the prompt to help the teacher correct minor errors in its chain of thought, and they filter out any remaining incorrect rationales. Smaller student models from the T5 family, ranging from 220 million to 11 billion parameters, are then fine-tuned on these generated reasoning sequences using teacher forcing.

The findings show substantial performance improvements across multiple domains. On the GSM8K math reasoning benchmark, the 11-billion-parameter T5 model increased its accuracy from 8.11% to 21.99% when trained on PaLM-generated reasoning, and reached 38.21% when paired with an external calculator. Similarly, on the MAWPS math dataset, accuracy rose from 54.15% to 70.41% (and 88.22% with a calculator). In ablation studies, a compact T5 base model with 44 times fewer parameters matched the baseline performance of the much larger T5 model when trained on reasoning data. The process also proved highly data-efficient: training on just 20% of the reasoning dataset yielded an 11.22% accuracy, outperforming the full baseline dataset.

These results demonstrate that organizations can significantly reduce model size, operational costs, and latency while retaining advanced multi-step reasoning performance. Smaller models fine-tuned on reasoning data can handle complex logic tasks without paying the computational overhead of running massive frontier models at inference time. However, improvements in commonsense reasoning were more modest—increasing from 68.12% to 71.98% on StrategyQA—because smaller architectures have limited internal memory to store broad factual knowledge.

Organizations seeking to deploy cost-effective reasoning systems should adopt this fine-tuning strategy: annotate datasets using large models conditioned on known target answers, filter out invalid reasoning paths, and train smaller models directly on the step-by-step solutions. To maximize mathematical accuracy, small models should be paired with external tool execution, such as automated calculators.

Confidence in these findings is strong across arithmetic tasks and standard text-based reasoning pipelines, with consistent results across multiple teacher architectures. However, decision-makers should note key limitations: the evaluations were conducted exclusively on English-language benchmarks, evaluated tasks independently rather than in multi-task configurations, and showed limited out-of-distribution generalizability on certain symbolic sequence tasks.

arXiv: 2212.08410
Cover for Teaching Small Language Models to Reason

Abstract

Chain of thought prompting successfully improves the reasoning capabilities of large language models, achieving state of the art results on a range of datasets. However, these reasoning capabilities only appear to emerge in models with at least tens of billions of parameters. In this paper, we explore the transfer of such reasoning capabilities to smaller models via knowledge distillation, also investigating model and dataset size trade-off. Specifically, we finetune a student model on the chain of thought outputs generated by a larger teacher model. Our experiments show that the proposed method improves task performance across arithmetic, commonsense and symbolic reasoning datasets. For example, the accuracy of T5 XXL on GSM8K improves from 8.11% to 21.99% and 18.42% when finetuned on PaLM 540B and GPT-3 175B generated chains of thought, respectively.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 4 Experimental Setup
  • 4.1 Benchmarks and Metrics
  • 4.1.1 Arithmetic Reasoning
  • 4.1.2 Commonsense Reasoning
  • 4.1.3 Symbolic Reasoning
  • 4.2 Baselines and setup
  • 5 Results
  • 5.1 Arithmetic reasoning
  • 5.1.1 Ablation study on generating chain-of-thought data
  • 5.2 Commonsense reasoning
  • 5.3 Symbolic reasoning
  • 5.4 Replicating Results using different Teacher Models
  • 5.5 Ablation study on model size
  • 5.6 Ablation study on dataset size
  • 6 Discussion
  • 7 Conclusion
  • 8 Limitations
  • 9 Ethical Considerations
  • References
  • A Dataset Usage and Licenses
  • A.1 Arithmetic Reasoning
  • A.2 Commonsense Reasoning
  • A.3 Symbolic Reasoning
  • B Computational Resources

Knowls

  1. Knowl 1 — Two-Step Chain-of-Thought Knowledge Distillation Pipeline

    model/method

    The Chain-of-Thought (CoT) knowledge distillation framework transfers multi-step reasoning capabilities from large language models (LLMs with ≥100B\ge 100\text{B} parameters, such as PaLM 540B or GPT-3 175B) to smaller student language models (such as T5 variants). The pipeline consists of two stages:

    1. Teacher CoT Data Generation and Filtering: An existing supervised dataset of question-answer pairs (x,y)(x, y) is annotated with CoT reasoning generated by a teacher LLM using few-shot prompting (typically 8 exemplars). To improve the quality of generated reasoning traces, the few-shot prompt is modified by providing the ground-truth target answer yy immediately after question xx before eliciting the intermediate CoT. Any candidate generation whose final answer does not match the ground-truth target yy is discarded to avoid training the student on flawed demonstrations.
    2. Student Model Finetuning: A student model is finetuned via teacher forcing where the input is the question xx, and the target sequence is the concatenated chain of thought and answer [CoT;y][\text{CoT}; y]. After finetuning, the student model generates full step-by-step reasoning and answers directly at inference time without requiring few-shot exemplars.
  2. Knowl 2 — Arithmetic Reasoning Distillation Performance Across Math Word Problem Benchmarks

    empirical result

    Finetuning T5 XXL (11 billion parameters) on Chain-of-Thought (CoT) data generated by PaLM 540B substantially improves accuracy across arithmetic reasoning benchmarks compared to a standard baseline T5 XXL model finetuned directly on target answers.

    Dataset Baseline T5 XXL CoT Finetuned T5 XXL CoT 8-Shot PaLM 540B
    Acc. (%) Acc. (%) Acc. w/ Calc (%) Acc. (%) Acc. w/ Calc (%)
    GSM8K 8.11 21.99 38.21 56.90 58.60
    MAWPS 54.15 70.41 88.22 93.00 93.66
    ASDiv 39.64 42.12 60.73 73.90 72.60

    On GSM8K, distillation improves raw accuracy from 8.11%8.11\% to 21.99%21.99\% (and to 38.21%38.21\% when arithmetic calculations are evaluated with an external calculator). On MAWPS, accuracy increases from 54.15%54.15\% to 70.41%70.41\% (88.22%88.22\% with calculator, nearing 8-shot PaLM 540B's 93.66%93.66\%). On ASDiv, accuracy improves from 39.64%39.64\% to 42.12%42.12\% (60.73%60.73\% with calculator). The large gains obtained with an external calculator show that the distilled student model learns the underlying logic and problem decomposition, while remaining limited by raw calculation accuracy.

  3. Knowl 3 — Target-Conditioned Prompting for High-Quality Teacher CoT Generation

    empirical result

    Conditioning a teacher LLM on the ground-truth target answer during few-shot CoT generation significantly enhances the yield and correctness of the generated intermediate reasoning.

    When prompting PaLM 540B with 8 few-shot exemplars on the GSM8K training set:

    • Generating CoT without providing the target answer achieves a generation accuracy of 59.98%59.98\%.
    • Providing the ground-truth target in the prompt immediately following the question increases generation accuracy to 79.37%79.37\% (an absolute increase of +19.39%+19.39\%).

    An analysis of the differences between reasoning paths generated with and without target conditioning shows that the performance gain arises primarily from the teacher LLM self-correcting intermediate reasoning paths that previously contained a wrong or missing arithmetic/logical step, rather than simply copying the given answer.

  4. Knowl 4 — Scaling Behavior of Student Model Size in CoT Distillation

    empirical result

    Evaluating T5 student models of varying parameter sizes on the GSM8K dataset shows that CoT distillation allows much smaller models to match or exceed the performance of substantially larger standard baselines.

    Evaluation Mode T5 Small (60M) T5 Base (220M) T5 Large (770M) T5 XL (3B) T5 XXL (11B)
    Baseline (No CoT) 3.34% 4.17% 4.40% 6.67% 8.11%
    CoT Finetuned 7.05% 7.96% 9.70% 12.13% 21.99%
    CoT w/ Calculator 13.12% 20.17% 22.21% 24.34% 38.21%

    Key findings include:

    1. T5 Base (220M220\text{M} parameters, 44×44\times smaller than T5 XXL) achieves 7.96%7.96\% with CoT finetuning, matching the baseline T5 XXL (11B11\text{B} parameters, 8.11%8.11\%).
    2. When evaluated with an external calculator to bypass low-level arithmetic calculation errors, even T5 Small (60M60\text{M} parameters) achieves 13.12%13.12\%, outperforming the baseline T5 XXL model without CoT (8.11%8.11\%).
    3. Reasoning accuracy scales consistently with student parameter size, reaching 21.99%21.99\% (unassisted) and 38.21%38.21\% (calculator-assisted) on T5 XXL.
  5. Knowl 5 — Sample Efficiency and Dataset Size Trade-off in CoT Distillation

    empirical result

    Finetuning student models on multi-step CoT reasoning paths exhibits substantially higher sample efficiency than direct-target supervised finetuning. On GSM8K, T5 XXL was finetuned on subsets of the PaLM 540B-generated CoT data:

    Training Data Proportion Number of Examples Accuracy (%) Accuracy w/ Calc (%)
    4% 213 6.29 12.28
    20% 1067 11.22 20.47
    100% 5337 21.99 38.21

    With only 20%20\% of the CoT data (1,0671{,}067 examples), the model achieves 11.22%11.22\% unassisted accuracy and 20.47%20.47\% calculator-assisted accuracy, exceeding the baseline T5 XXL model trained on 100%100\% of the direct-target data (6,7256{,}725 examples, 8.11%8.11\% accuracy). This represents an approximate 6×6\times data efficiency improvement.

  6. Knowl 6 — Cross-Teacher Generalization and Comparison with Ground-Truth CoT

    empirical result

    CoT knowledge distillation is effective across different teacher architectures (PaLM 540B and GPT-3 175B) and yields results competitive with or superior to human-authored (golden) CoTs on GSM8K and StrategyQA.

    Task / Metric Base Task Original CoT CoT Finetuned T5 XXL CoT 8-Shot Teacher
    Baseline Finetuned PaLM 540B GPT-3 175B PaLM 540B GPT-3 175B
    GSM8K Acc. (%) 8.11 19.94 21.99 18.42 56.9 46.9
    GSM8K w/ Calc (%) - 26.99 38.21 33.06 58.6 49.6
    GSM8K Training Size 6725 6725 5337 5298 - -
    StrategyQA Acc. (%) 68.12 71.98 67.15 63.77 77.8 65.4
    StrategyQA Training Size 1648 1648 1319 1319 - -

    On GSM8K, distilling from PaLM 540B achieves 21.99%21.99\% (38.21%38.21\% with calculator), outperforming distillation from GPT-3 175B (18.42%18.42\%) and direct finetuning on original human CoTs (19.94%19.94\%). On StrategyQA, finetuning on original CoTs achieves 71.98%71.98\%, outperforming PaLM distillation (67.15%67.15\%) and GPT-3 distillation (63.77%63.77\%), partly because filtering teacher errors reduces the usable dataset size (1,3191{,}319 vs 1,6481{,}648 examples).

  7. Knowl 7 — Length Generalization in Symbolic Reasoning Tasks via CoT Distillation

    empirical result

    Evaluating student models finetuned on length-2 symbolic reasoning tasks on out-of-distribution (OOD) sequence lengths (lengths 3 and 4) reveals task-dependent generalization behavior:

    Task OOD Length Baseline T5 XXL (%) CoT Finetuned T5 XXL (%) PaLM 540B 8-Shot (%)
    Last Letter Concat. 3 0.00 0.00 94.80
    Last Letter Concat. 4 0.00 0.00 63.00
    Coinflip 3 13.10 86.70 98.60
    Coinflip 4 73.80 70.50 90.20

    On Coinflip state tracking, CoT finetuning improves OOD length-3 accuracy from 13.10%13.10\% to 86.70%86.70\%, but drops slightly to 70.50%70.50\% on length-4 sequences (where baseline direct finetuning achieves 73.80%73.80\%). On Last Letter Concatenation, both baseline finetuning and CoT finetuning fail completely (0.00%0.00\% accuracy on lengths 3 and 4), despite the teacher PaLM 540B achieving 94.80%94.80\% and 63.00%63.00\%.

  8. Knowl 8 — External Calculator Execution for Equation Correction in Arithmetic Reasoning

    model/method

    To isolate the model's structural reasoning quality from basic arithmetic execution errors during mathematical evaluation, an automated external calculator verification routine is applied to generated CoTs:

    1. The routine scans the generated text sequentially to identify mathematical equations of the form LHS=RHS\text{LHS} = \text{RHS}.
    2. For each detected equation, the arithmetic expression on the left-hand side (LHS\text{LHS}) is computed by an external calculator.
    3. The original right-hand side (RHS\text{RHS}) in the generated text is replaced with the computed result.
    4. Downstream occurrences of the erroneous number are updated in subsequent equations to prevent arithmetic calculation errors from compounding through the reasoning chain (for instance, transforming 5 + 5 = 11. 11 * 2 = 22 into 5 + 5 = 10. 10 * 2 = 20).
    5. The final answer extracted from the corrected chain is compared against the benchmark target.
  9. Knowl 9 — Factual Knowledge Bottleneck in Commonsense Reasoning Distillation

    empirical result

    Applying CoT distillation to commonsense question answering on the StrategyQA benchmark yields more limited gains than on arithmetic benchmarks. Baseline direct finetuning of T5 XXL achieves 68.12%68.12\% accuracy; finetuning on PaLM 540B CoT data yields 67.15%67.15\%, GPT-3 175B CoT data yields 63.77%63.77\%, and ground-truth human CoTs yields 71.98%71.98\%.

    StrategyQA requires multi-step reasoning over implicit factual knowledge. Because smaller student models (such as T5 XXL with 11B11\text{B} parameters) possess substantially lower memorization capacity than ≥100B\ge 100\text{B} teacher LLMs, the lack of parametric world knowledge acts as a bottleneck that prevents the full transfer of commonsense reasoning capabilities.

  10. Knowl 10 — Limitations of CoT Distillation Framework

    limitation

    The CoT knowledge distillation framework exhibits several key limitations:

    1. Language and Task Setup: Experiments are conducted exclusively in English and evaluate single tasks in isolation, omitting multilingual evaluation and multitask fine-tuning setups.
    2. Dependency on Proprietary LLMs and Inference Compute: Generating large-scale, high-quality CoT training sets requires access to proprietary or non-public LLMs exceeding 100B100\text{B} parameters and requires substantial computational resources to run teacher inference over full datasets.
    3. Absence of Advanced Decoding Strategies: The framework relies on standard greedy/few-shot prompting during teacher data generation and student inference, without incorporating sampling-based reasoning methods such as self-consistency majority voting.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.BIG-bench collaboration. 2021. Beyond the imitation game: Measuring and extrapolating the capabilities of language models. In preparation.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  4. 4.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  5. 5.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  6. 6.Jacob Eisenstein, Daniel Andor, Bernd Bohnet, Michael Collins, and David Mimno. 2022. Honest students from untrusted teachers: Learning an interpretable question-answering pipeline from a pretrained language model. arXiv preprint arXiv:2210.02498.
  7. 7.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462.
  8. 8.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021a. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  9. 9.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021b. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  10. 10.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7).
  11. 11.Namgyu Ho, Laura Schmid, and Se-Young Yun. 2022. Large language models are reasoning teachers. arXiv preprint arXiv:2212.10071.
  12. 12.Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2022. Large language models can self-improve. arXiv preprint arXiv:2210.11610.
  13. 13.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022. Survey of hallucination in natural language generation. ACM Computing Surveys.
  14. 14.Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. Mawps: A math word problem repository. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1152–1157.
  15. 15.Sarah Kreps, R Miles McCain, and Miles Brundage. 2022. All the news that’s fit to fabricate: Aigenerated text as a tool of media misinformation. Journal of Experimental Political Science, 9(1):104–117.
  16. 16.Shiyang Li, Jianshu Chen, Yelong Shen, Zhiyu Chen, Xinlu Zhang, Zekun Li, Hong Wang, Jing Qian, Baolin Peng, Yi Mao, et al. 2022. Explanations from large language models make small reasoners better. arXiv preprint arXiv:2210.06726.
  17. 17.Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2021. A diverse corpus for evaluating and developing english math word problem solvers. arXiv preprint arXiv:2106.15772.
  18. 18.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with frank: A benchmark for factuality metrics. arXiv preprint arXiv:2104.13346.
  19. 19.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  20. 20.Kumar Shridhar, Alessandro Stolfo, and Mrinmaya Sachan. 2022. Distilling multi-step reasoning capabilities of large language models into smaller models via semantic decompositions. arXiv preprint arXiv:2212.00193.
  21. 21.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. 2022. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131.
  22. 22.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  23. 23.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  24. 24.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  25. 25.Ronald J Williams and David Zipser. 1989. A learning algorithm for continually running fully recurrent neural networks. Neural computation, 1(2):270–280.

Citation

MLA
Magister, L. C., et al. “Teaching Small Language Models to Reason”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2023, pp. 1773–81, https://doi.org/10.18653/v1/2023.acl-short.151.
APA
Magister, L. C., Mallinson, J., Adamek, J., Malmi, E., & Severyn, A. (2023). Teaching Small Language Models to Reason. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 1773–1781. https://doi.org/10.18653/v1/2023.acl-short.151
Chicago
Magister, L. C., J. Mallinson, J. Adamek, E. Malmi, and A. Severyn. 2023. “Teaching Small Language Models to Reason”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 1773–81. https://doi.org/10.18653/v1/2023.acl-short.151.
Harvard
Magister, L.C. et al. (2023) “Teaching Small Language Models to Reason”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp. 1773–1781. Available at: https://doi.org/10.18653/v1/2023.acl-short.151.
Vancouver
1. Magister LC, Mallinson J, Adamek J, Malmi E, Severyn A (2023) Teaching Small Language Models to Reason. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp 1773–1781

BibTeX

@inproceedings{magister-etal-2023-teaching,
    title = "Teaching Small Language Models to Reason",
    author = "Magister, Lucie Charlotte  and
      Mallinson, Jonathan  and
      Adamek, Jakub  and
      Malmi, Eric  and
      Severyn, Aliaksei",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-short.151/",
    doi = "10.18653/v1/2023.acl-short.151",
    pages = "1773--1781"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/