Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective

Yiyao YuYuxiang ZhangDongdong ZhangXiao LiangHengyuan ZhangXingxing ZhangMahmoud KhademiHany AwadallaJunjie WangYujiu Yang

article2025ACL42 citations

Introduces Chain-of-Reasoning, a multi-paradigm framework that synthesizes natural language, algorithmic, and symbolic logic to produce a 7B model capable of outperforming GPT-4o by 41% on mathematical theorem proving.

Listen

Large language models have achieved substantial success in basic mathematical reasoning, but existing systems struggle with comprehensive mathematical problem solving across diverse tasks. Current approaches predominantly optimize a single reasoning paradigm—such as natural language explanation, algorithmic programming, or formal symbolic logic. This single-paradigm specialization restricts model upper-bound performance, incurs high computational search costs at test time, and undermines cross-task generalization, causing models specialized in arithmetic to fail at formal theorem proving and vice versa.

The article demonstrates that unifying natural language reasoning, symbolic reasoning in formal proof languages like Lean 4, and algorithmic reasoning via Python execution into a cohesive framework enables superior cross-task mathematical performance. To evaluate this hypothesis, the authors introduced Chain-of-Reasoning (CoR), constructed a Multi-Paradigm Mathematical (MPM) training dataset comprising 167,412 multi-paradigm reasoning paths across 82,770 problems, and developed a Progressive Paradigm Training strategy. They used this process to train CoR-Math-7B, fine-tuned from DeepSeekMath-Base-7B, alongside an inference technique called Sequential Multi-Paradigm Sampling that systematically expands reasoning paths across paradigms.

The findings establish substantial performance advantages across standard benchmarks. In formal theorem proving on the miniF2F benchmark, CoR-Math-7B achieved a 66.0% accuracy in a zero-shot setting, representing a 41.0% absolute improvement over GPT-4o's few-shot performance and surpassing specialized few-shot proof search systems while utilizing far fewer generated candidate paths. In arithmetic computation, the model achieved 66.7% zero-shot accuracy on the MATH benchmark, outperforming GPT-4 by 24.2% and reinforcement-learning-based baseline DeepSeekMath-RL-7B by 15.0%. Ablation evaluations revealed that chaining paradigms sequentially—specifically natural language followed by symbolic reasoning, then algorithmic execution—significantly improved outcomes by enabling structured decomposition and cross-paradigm self-correction.

These results demonstrate that multi-paradigm collaboration provides higher resource efficiency and accuracy than scaling single-paradigm search spaces. By allowing models to autonomously or instructionally switch reasoning modes, organizations can lower inference costs and reduce the amount of fine-tuning data required to achieve high-accuracy mathematical reasoning. The findings suggest that future development in AI reasoning should prioritize multi-medium integration over purely increasing parameter size or single-paradigm search budgets.

Organizations evaluating this approach should consider implementing multi-paradigm reasoning pipelines for complex numerical and logical automation, while ensuring pre-deployment security audits are conducted on external execution environments such as Python interpreters. Future research should expand multi-paradigm evaluations to broader enterprise tasks and test-time search budgets on large-scale arithmetic datasets.

Readers should note certain limitations: performance in symbolic reasoning remains constrained by the relative scarcity of formal proof corpora in pre-training data, and fixed model context windows can occasionally truncate long reasoning chains, leading to syntax or completion errors. Nonetheless, the high confidence of the findings across multiple standard benchmarks supports multi-paradigm reasoning as an effective architecture for complex problem solving.

No sufficiently relevant recommendations were found.

Cover for Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective

Abstract

Large Language Models (LLMs) have made notable progress in mathematical reasoning, yet often rely on single-paradigm reasoning, limiting their effectiveness across diverse tasks. We introduce Chain-of-Reasoning (CoR), a novel unified framework integrating multiple reasoning paradigms--Natural Language Reasoning (NLR), Algorithmic Reasoning (AR), and Symbolic Reasoning (SR)--to enable synergistic collaboration. CoR generates multiple potential answers via different reasoning paradigms and synthesizes them into a coherent final solution. We propose a Progressive Paradigm Training (PPT) strategy for models to progressively master these paradigms, leading to CoR-Math-7B. Experimental results demonstrate that CoR-Math-7B significantly outperforms current SOTA models, achieving up to a 41.0% absolute improvement over GPT-4o in theorem proving and a 15.0% improvement over RL-based methods on the MATH benchmark in arithmetic tasks. These results show the enhanced mathematical comprehension ability of our model, enabling zero-shot generalization across tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Chain-of-Reasoning Framework
  • 3.1 Overview
  • 3.2 Collecting Dataset
  • 3.3 Training
  • 3.4 Inference
  • 4 Experimental Settings
  • 4.1 Evaluation Setup
  • 4.2 Implementation Details
  • 4.3 Baselines
  • 5 Main Results
  • 5.1 Comparisons with General-purpose Mathematical Models
  • 5.2 Comparisons with Theorem Proving Experts
  • 5.3 Comparisons with Arithmetic Experts
  • 6 Ablation Study
  • 6.1 Impact of Stages in PPT Method
  • 6.2 Order of the Reasoning Paradigms
  • 6.3 Impact of Model Scales
  • 6.4 Multi- vs. Single-Paradigm Reasoning
  • 7 Conclusions
  • References
  • A Discussion on Reasoning Hierarchy: Paradigms, Paths, and Steps
  • B Experiment Details
  • B.1 The detail of the Universal Text Template
  • B.2 Example Prompts for Dataset Enhancement
  • B.3 Details of Training Settings
  • B.4 Benchmarks
  • B.5 Details of Metrics
  • B.6 Experimental Details of MATH and GSM8K benchmarks
  • C Additional Analysis
  • C.1 The Risk of Data Leakage
  • C.2 Different Evaluation Strategies on Arithmetic Benchmarks
  • C.3 Ethical Considerations
  • D Case Studies
  • D.1 Cases of Instruction-free Reasoning
  • D.2 Qualitative Analysis of Error Cases

Knowls

  1. Knowl 1 — CoR chains distinct reasoning paradigms into one solution

    model/method

    Chain-of-Reasoning (CoR) treats natural-language reasoning (NLR), symbolic reasoning (SR), and algorithmic reasoning (AR) as distinct knowledge media that can be applied successively to the same mathematical problem. Given a problem xx, each paradigm's reasoning trace τi\tau_i is generated conditioned on xx and the traces generated earlier in the chain; a final answer yy is then synthesized from the problem and the paradigm outputs. The traces for SR and AR include their tool interactions: SR uses the Lean prover, while AR uses a Python compiler. The framework's central distinction from tool-assisted single-paradigm reasoning is that multiple paradigms contribute complete reasoning paths, rather than one paradigm doing the reasoning while another only assists with a subproblem.

  2. Knowl 2 — MPM supplies verified multi-paradigm training paths

    data/table

    The Multi-Paradigm Math (MPM) dataset contains 82,770 mathematical problems and 167,412 multi-paradigm reasoning solutions. Its construction begins with Numina-TIR and Lean-Workbook examples: samples lacking a corresponding solution are removed, and language models such as GPT-4o are used to generate missing paradigms and refine existing ones. The resulting MPM-raw collection contains about 285,000 samples. Each symbolic proof is then checked by Lean; successful samples are retained, while failed proofs are sent with prover errors to DeepSeek-Prover-V1.5 for revision and resubmission, for at most 64 attempts. The final collection includes only samples whose formal proof passes Lean verification, with natural-language and algorithmic reasoning also manually reviewed.

  3. Knowl 3 — Progressive Paradigm Training builds capabilities in stages

    model/method

    Progressive Paradigm Training (PPT) fine-tunes a model in three successive stages, adding reasoning media rather than merely increasing task difficulty within one medium. Stage 1 trains on Numina-CoT* with NLR; stage 2 uses Numina-TIR* with NLR and AR; stage 3 uses MPM with NLR, AR, and SR. The default CoR-Math-7B backbone is DeepSeekMath-Base 7B; the paper also trains Llama-3.1 8B with PPT. Across stages the learning rate is 2×10−52\times10^{-5} and warm-up ratio is 1%; stage 1 runs for 3 epochs, stages 2 and 3 for 4 epochs each. Maximum sequence length is 2,048 tokens in stages 1–2 and 4,096 in stage 3, followed by an annealing phase on high-quality MPM samples. In the reported stage ablation, stage 1 raised Llama-3.1-8B's MATH and GSM8K performance by 47.9 and 61.0 points, respectively; stage 2 yielded smaller immediate gains, while adding the third paradigm further improved results.

  4. Knowl 4 — Prompted inference adapts reasoning depth and samples across paradigms

    model/method

    At inference, prompts select a reasoning sequence appropriate to the task. For theorem proving, CoR first produces NLR and then SR, and the Lean proof is extracted as the answer. For arithmetic, it uses NLR, then SR, then AR, followed by a summary of the result. The reported main results use instruction-followed reasoning; the authors also show examples in which the model switches paradigms without explicit instructions.

    Sequential Multi-Paradigm Sampling (SMPS) expands candidates at the paradigm level. For a two-paradigm chain and input problem xx, sample JJ first-paradigm paths τ1j\tau_{1j} from P(τ1∣x)P(\tau_1\mid x). For each jj, sample KK second-paradigm paths τ2k\tau_{2k} conditioned on xx and τ1j\tau_{1j}, then generate answers conditioned on xx, τ1j\tau_{1j}, and τ2k\tau_{2k}. This yields JKJ K candidate responses. In miniF2F experiments with NLR and SR, the paper reports SMPS sample budgets as N=NNLR×NSRN=N_{\mathrm{NLR}}\times N_{\mathrm{SR}}.

  5. Knowl 5 — CoR reaches 66.0% on miniF2F with zero-shot multi-paradigm reasoning

    empirical result

    On the Lean-4 version of miniF2F, CoR-Math-7B obtains 66.0% pass@N with an NLR-by-SR SMPS budget of 128×128128\times128, without demonstrations. Its reported scores are 52.9% at 128×1128\times1 and 59.4% at 32×10032\times100. The paper reports 25.0% for GPT-4o at pass@128 and describes CoR's 66.0% as a 41.0-point absolute improvement; these are not matched inference budgets or prompting conditions. CoR also exceeds the reported 63.5% for DeepSeek-Prover-V1.5-RL + RMaxTS at 32×640032\times6400, and the 65.9% for InternLM2.5-StepProver-BF+CG at 256×32×600256\times32\times600. Thus, the results show strong miniF2F performance under the paper's zero-shot setup, while the comparisons draw on baselines with different search budgets and settings.

  6. Knowl 6 — Arithmetic results are competitive, but do not lead every benchmark

    empirical result

    In zero-shot arithmetic evaluation, CoR-Math-7B scores 66.7% on MATH and 88.7% on GSM8K (Pass@1), and solves 34/40 AMC2023 and 12/30 AIME2024 problems (Maj@64). On MATH, 66.7% is 15.0 points above DeepSeekMath-RL-7B's 51.7% and 11.4 points above NuminaMath-7B-TIR's 55.3%. However, Qwen2.5-Math-7B-Instruct scores higher than CoR on both MATH (83.6%) and GSM8K (95.2%). The paper reports 1,098k supervised fine-tuning examples for CoR-Math-7B, versus 3,026k for Qwen2.5-Math-7B-Instruct; this indicates lower reported SFT data volume, not better accuracy than Qwen on those benchmarks.

  7. Knowl 7 — Controlled comparisons favor multi-paradigm fine-tuning

    empirical result

    A controlled ablation fine-tuned DeepSeekMath-7B-Base separately on NLR-, AR-, or SR-only paths from MPM, and compared those models with CoR using the same base model and 10,000 training samples per paradigm. On GSM8K and MATH, the best single-paradigm scores were 75.8% (AR) and 37.6% (AR), while CoR scored 88.7% and 66.7%, gains of 12.9 and 29.1 points. On AMC2023, CoR solved 34/40 versus the best single-paradigm result of 14/40; on AIME2024, it solved 12/30 versus 1/30. On miniF2F, CoR scored 52.9% versus 44.3% for SR alone, an 8.6-point difference. These results support the benefit of combining the paradigms under this training comparison.

  8. Knowl 8 — The order of paradigms changes arithmetic accuracy

    empirical result

    In a zero-shot CoR-Math-7B arithmetic ablation that changed only the prompt order, NLR→SR→AR achieved 66.7% on MATH and 88.7% on GSM8K, while NLR→AR→SR achieved 49.9% and 84.2%, respectively. NLR was kept first because it aligns with language-model pretraining. The results show that performance depends on how paradigms are connected, not only on whether the model has been trained on each one. The authors suggest that SR may structure a problem into substeps before AR performs calculations.

  9. Knowl 9 — CoR transfers across model families and parameter scales

    empirical result

    The paper applies CoR training to Qwen2.5-Math and Llama-3.1 models at multiple sizes and evaluates zero-shot MATH, GSM8K, and miniF2F performance. For Qwen2.5-Math at 1.5B parameters, base-to-CoR scores change from 34.0% to 57.6% on MATH, 39.3% to 84.5% on GSM8K, and 0.0% to 51.6% on miniF2F; at 7B, the corresponding changes are 51.8% to 64.7%, 90.0% to 90.0%, and 0.0% to 52.5%. For Llama-3.1 at 8B, scores change from 4.2% to 58.2%, 6.2% to 84.0%, and 25.8% to 53.3%; at 70B, from 16.8% to 70.7%, 20.5% to 90.0%, and 22.5% to 56.2%. These results demonstrate that the framework was applied beyond the default DeepSeekMath backbone and that its effect varies by benchmark and model size.

  10. Knowl 10 — Evaluation alignment and incomplete outputs remain limitations

    limitation

    The authors note that zero-shot evaluation is difficult to align with prior work, especially on miniF2F, where many comparison systems use few-shot settings. They also do not apply SMPS to arithmetic benchmarks because of resource constraints and because competing methods generally generate within a single paradigm. A preliminary random-sample error analysis identifies incomplete proofs or code syntax errors as the largest category: among 200 arithmetic errors, 85 were incomplete proof/code errors, 56 computational errors, 33 comprehension or premise errors, and 26 logical fallacies; among 50 theorem-proving errors, the respective counts were 33, 9, 2, and 6. The paper additionally notes that the backbone's fixed context window can truncate long reasoning chains and contribute to incomplete outputs. The error counts describe the sampled cases, not estimated population-wide rates.

Coverage note — Detailed prompt exemplars, the training–test similarity filtering procedure, and evaluation-set statistics are omitted as supporting implementation details rather than separate central contributions.

References

  1. 1.AI-MO. 2024a. Aime 2023. https://huggingface.co/datasets/AI-MO/aimo-validation-aime/.
  2. 2.AI-MO. 2024b. Amc 2023. https://huggingface.co/datasets/AI-MO/aimo-validation-amc/.
  3. 3.Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen Marcus McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck. 2024. Llemma: An open language model for mathematics. In ICLR. OpenReview.net.
  4. 4.Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. 2024. Graph of thoughts: Solving elaborate problems with large language models. In AAAI, pages 17682–17690. AAAI Press.
  5. 5.Jürgen Branke, Kalyanmoy Deb, Kaisa Miettinen, and Roman Slowinski, editors. 2008. Multiobjective Optimization, Interactive and Evolutionary Approaches [outcome of Dagstuhl seminars], volume 5252 of Lecture Notes in Computer Science. Springer.
  6. 6.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. 2023. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. Trans. Mach. Learn. Res., 2023.
  7. 7.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. CoRR, abs/2110.14168.
  8. 8.Tri Dao. 2024. Flashattention-2: Faster attention with better parallelism and work partitioning. In ICLR. OpenReview.net.
  9. 9.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurélien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Rozière, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Graeme Nail, Grégoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel M. Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, and et al. 2024. The llama 3 herd of models. CoRR, abs/2407.21783.
  10. 10.Edward A Feigenbaum, Julian Feldman, et al. 1963. Computers and thought, volume 37. New York McGraw-Hill.
  11. 11.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. PAL: program-aided language models. In ICML, volume 202 of Proceedings of Machine Learning Research, pages 10764–10799. PMLR.
  12. 12.Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2024. Tora: A tool-integrated reasoning agent for mathematical problem solving. In ICLR. OpenReview.net.
  13. 13.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the MATH dataset. In NeurIPS Datasets and Benchmarks.
  14. 14.Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014. Learning to solve arithmetic word problems with verb categorization. In EMNLP, pages 523–533. ACL.
  15. 15.Yinya Huang, Xiaohan Lin, Zhengying Liu, Qingxing Cao, Huajian Xin, Haiming Wang, Zhenguo Li, Linqi Song, and Xiaodan Liang. 2024. MUSTARD: mastering uniform synthesis of theorem and proof data. In ICLR. OpenReview.net.
  16. 16.Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander Madry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, Alex Nichol, Alex Paino, Alex Renzin, Alex Tachard Passos, Alexander Kirillov, Alexi Christakis, Alexis Conneau, Ali Kamali, Allan Jabri, Allison Moyer, Allison Tam, Amadou Crookes, Amin Tootoonchian, Ananya Kumar, Andrea Vallone, Andrej Karpathy, Andrew Braunstein, Andrew Cann, Andrew Codispoti, Andrew Galu, Andrew Kondrich, Andrew Tulloch, Andrey Mishchenko, Angela Baek, Angela Jiang, Antoine Pelisse, Antonia Woodford, Anuj Gosalia, Arka Dhar, Ashley Pantuliano, Avi Nayak, Avital Oliver, Barret Zoph, Behrooz Ghorbani, Ben Leimberger, Ben Rossen, Ben Sokolowsky, Ben Wang, Benjamin Zweig, Beth Hoover, Blake Samic, Bob McGrew, Bobby Spero, Bogo Giertler, Bowen Cheng, Brad Lightcap, Brandon Walkin, Brendan Quinn, Brian Guarraci, Brian Hsu, Bright Kellogg, Brydon Eastman, Camillo Lugaresi, Carroll L. Wainwright, Cary Bassin, Cary Hudson, Casey Chu, Chad Nelson, Chak Li, Chan Jun Shern, Channing Conger, Charlotte Barette, Chelsea Voss, Chen Ding, Cheng Lu, Chong Zhang, Chris Beaumont, Chris Hallacy, Chris Koch, Christian Gibson, Christina Kim, Christine Choi, Christine McLeavey, Christopher Hesse, Claudia Fischer, Clemens Winter, Coley Czarnecki, Colin Jarvis, Colin Wei, Constantin Koumouzelis, and Dane Sherburn. 2024. Gpt-4o system card. CoRR, abs/2410.21276.
  17. 17.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b. CoRR, abs/2310.06825.
  18. 18.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024. Mixtral of experts. CoRR, abs/2401.04088.
  19. 19.Haein Kong, Yongsu Ahn, Sangyub Lee, and Yunho Maeng. 2024. Gender bias in llm-generated interview responses. CoRR, abs/2410.20739.
  20. 20.Guillaume Lample, Timothée Lacroix, Marie-Anne Lachaux, Aurélien Rodriguez, Amaury Hayat, Thibaut Lavril, Gabriel Ebner, and Xavier Martinet. 2022. Hypertree proof search for neural theorem proving. In NeurIPS.
  21. 21.Chen Li, Weiqi Wang, Jingcheng Hu, Yixuan Wei, Nanning Zheng, Han Hu, Zheng Zhang, and Houwen Peng. 2024. Common 7b language models already possess strong math capabilities. CoRR, abs/2403.04706.
  22. 22.Jia LI, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Costa Huang, Kashif Rasul, Longhui Yu, Albert Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Yann Fleureau, Guillaume Lample, and Stanislas Polu. 2024. NuminaMath. https://github.com/project-numina/aimo-progress-prize.
  23. 23.Haohan Lin, Zhiqing Sun, Yiming Yang, and Sean Welleck. 2024a. Lean-star: Learning to interleave thinking and proving. CoRR, abs/2407.10040.
  24. 24.Zicheng Lin, Tian Liang, Jiahao Xu, Xing Wang, Ruilin Luo, Chufan Shi, Siheng Li, Yujiu Yang, and Zhaopeng Tu. 2024b. Critical tokens matter: Token-level contrastive estimation enhence llm’s reasoning capability. arXiv preprint arXiv:2411.19943.
  25. 25.Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. CoRR, abs/2308.09583.
  26. 26.Ruilin Luo, Zhuofan Zheng, Yifan Wang, Yiyao Yu, Xinzhe Ni, Zicheng Lin, Jin Zeng, and Yujiu Yang. 2025. Ursa: Understanding and verifying chain-of-thought reasoning in multimodal mathematics. arXiv preprint arXiv:2501.04686.
  27. 27.Frederic P. Miller, Agnes F. Vandome, and John McBrewster. 2009. Levenshtein Distance: Information theory, Computer science, String (computer science), String metric, Damerau?Levenshtein distance, Spell checker, Hamming distance. Alpha Press.
  28. 28.OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  29. 29.OpenAI. 2024. Openai o1 system card. Preprint, arXiv:2412.16720.
  30. 30.Stanislas Polu and Ilya Sutskever. 2020. Generative language modeling for automated theorem proving. CoRR, abs/2009.03393.
  31. 31.Jiahao Qiu, Yifu Lu, Yifan Zeng, Jiacheng Guo, Jiayi Geng, Huazheng Wang, Kaixuan Huang, Yue Wu, and Mengdi Wang. 2024. Treebon: Enhancing inference-time alignment with speculative tree-search and best-of-n sampling. CoRR, abs/2410.16033.
  32. 32.Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. Zero: memory optimizations toward training trillion parameter models. In SC, page 20. IEEE/ACM.
  33. 33.Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton-Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023. Code llama: Open foundation models for code. CoRR, abs/2308.12950.
  34. 34.Bilgehan Sel, Ahmad Al-Tawaha, Vanshaj Khattar, Ruoxi Jia, and Ming Jin. 2024. Algorithm of thoughts: Enhancing exploration of ideas in large language models. In ICML. OpenReview.net.
  35. 35.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. CoRR, abs/2402.03300.
  36. 36.Yuxuan Tong, Xiwen Zhang, Rui Wang, Ruidong Wu, and Junxian He. 2024. Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving. CoRR, abs/2407.13690.
  37. 37.Shaobo Wang, Xiangqi Jin, Ziming Wang, Jize Wang, Jiajun Zhang, Kaixin Li, Zichen Wen, Zhong Li, Conghui He, Xuming Hu, and Linfeng Zhang. 2025. Data whisperer: Efficient data selection for task-specific llm fine-tuning via few-shot in-context learning. Annual Meeting of the Association for Computational Linguistics.
  38. 38.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In ICLR. OpenReview.net.
  39. 39.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS.
  40. 40.Sean Welleck and Rahul Saha. 2023. LLMSTEP: LLM proofstep suggestions in lean. CoRR, abs/2310.18457.
  41. 41.Zijian Wu, Suozhi Huang, Zhejian Zhou, Huaiyuan Ying, Jiayu Wang, Dahua Lin, and Kai Chen. 2024. Internlm2.5-stepprover: Advancing automated theorem proving via expert iteration on large-scale LEAN problems. CoRR, abs/2410.15700.
  42. 42.Huajian Xin, Z. Z. Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, Wenjun Gao, Qihao Zhu, Dejian Yang, Zhibin Gou, Z. F. Wu, Fuli Luo, and Chong Ruan. 2024. Deepseek-prover-v1.5: Harnessing proof assistant feedback for reinforcement learning and monte-carlo tree search. CoRR, abs/2408.08152.
  43. 43.Yunfan Xiong, Ruoyu Zhang, Yanzeng Li, Tianhao Wu, and Lei Zou. 2024. Dyspec: Faster speculative decoding with dynamic token tree structure. CoRR, abs/2410.11744.
  44. 44.An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, Keming Lu, Mingfeng Xue, Runji Lin, Tianyu Liu, Xingzhang Ren, and Zhenru Zhang. 2024. Qwen2.5-math technical report: Toward mathematical expert model via self-improvement. CoRR, abs/2409.12122.
  45. 45.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. In NeurIPS.
  46. 46.Huaiyuan Ying, Zijian Wu, Yihan Geng, Jiayu Wang, Dahua Lin, and Kai Chen. 2024a. Lean workbook: A large-scale lean problem set formalized from natural language math problems. CoRR, abs/2406.03847.
  47. 47.Huaiyuan Ying, Shuo Zhang, Linyang Li, Zhejian Zhou, Yunfan Shao, Zhaoye Fei, Yichuan Ma, Jiawei Hong, Kuikun Liu, Ziyi Wang, Yudong Wang, Zijian Wu, Shuaibin Li, Fengzhe Zhou, Hongwei Liu, Songyang Zhang, Wenwei Zhang, Hang Yan, Xipeng Qiu, Jiayu Wang, Kai Chen, and Dahua Lin. 2024b. Internlm-math: Open math large language models toward verifiable reasoning. Preprint, arXiv:2402.06332.
  48. 48.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2024. Metamath: Bootstrap your own mathematical questions for large language models. In ICLR. OpenReview.net.
  49. 49.Yiyao Yu, Junjie Wang, Yuxiang Zhang, Lin Zhang, Yujiu Yang, and Tetsuya Sakai. 2023. EALM: introducing multidimensional ethical alignment in conversational information retrieval. In SIGIR-AP, pages 32–39. ACM.
  50. 50.Yu-Xiang Zhang, Junjie Wang, Xinyu Zhu, Tetsuya Sakai, and Hayato Yamana. 2024. SSR: solving named entity recognition problems via a single-stream reasoner. ACM Trans. Inf. Syst., 42(5):138:1–138:28.
  51. 51.Jinman Zhao and Xueyan Zhang. 2024. Large language model is not a (multilingual) compositional relation reasoner. In First Conference on Language Modeling.
  52. 52.Kunhao Zheng, Jesse Michael Han, and Stanislas Polu. 2022. minif2f: a cross-system benchmark for formal olympiad-level mathematics. In ICLR. OpenReview.net.
  53. 53.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V. Le, and Ed H. Chi. 2023. Least-to-most prompting enables complex reasoning in large language models. In ICLR. OpenReview.net.
  54. 54.Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Ruyi Gan, Jiaxing Zhang, and Yujiu Yang. 2023. Solving math word problems via cooperative reasoning induced language models. In ACL (1), pages 4471–4485. Association for Computational Linguistics.

Citation

MLA
Yu, Y., et al. “Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective”. arXiv, 2025, http://arxiv.org/abs/2501.11110v4.
APA
Yu, Y., Zhang, Y., Zhang, D., Liang, X., Zhang, H., Zhang, X., Yang, Z., Khademi, M., Awadalla, H., Wang, J., Yang, Y., & Wei, F. (2025). Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective. arXiv. http://arxiv.org/abs/2501.11110v4
Chicago
Yu, Y., Y. Zhang, D. Zhang, et al. 2025. “Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective”. arXiv. http://arxiv.org/abs/2501.11110v4.
Harvard
Yu, Y. et al. (2025) “Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.11110v4.
Vancouver
1. Yu Y, Zhang Y, Zhang D, et al (2025) Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective. arXiv

BibTeX

@article{yu2025chain,
  title = {Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective},
  author = {Yu, Yiyao and Zhang, Yuxiang and Zhang, Dongdong and Liang, Xiao and Zhang, Hengyuan and Zhang, Xingxing and Yang, Ziyi and Khademi, Mahmoud and Awadalla, Hany and Wang, Junjie and Yang, Yujiu and Wei, Furu},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.11110v4},
  eprint = {2501.11110}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/