Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective
Yiyao YuYuxiang ZhangDongdong ZhangXiao LiangHengyuan ZhangXingxing ZhangMahmoud KhademiHany AwadallaJunjie WangYujiu Yang
Introduces Chain-of-Reasoning, a multi-paradigm framework that synthesizes natural language, algorithmic, and symbolic logic to produce a 7B model capable of outperforming GPT-4o by 41% on mathematical theorem proving.
Large language models have achieved substantial success in basic mathematical reasoning, but existing systems struggle with comprehensive mathematical problem solving across diverse tasks. Current approaches predominantly optimize a single reasoning paradigm—such as natural language explanation, algorithmic programming, or formal symbolic logic. This single-paradigm specialization restricts model upper-bound performance, incurs high computational search costs at test time, and undermines cross-task generalization, causing models specialized in arithmetic to fail at formal theorem proving and vice versa.
The article demonstrates that unifying natural language reasoning, symbolic reasoning in formal proof languages like Lean 4, and algorithmic reasoning via Python execution into a cohesive framework enables superior cross-task mathematical performance. To evaluate this hypothesis, the authors introduced Chain-of-Reasoning (CoR), constructed a Multi-Paradigm Mathematical (MPM) training dataset comprising 167,412 multi-paradigm reasoning paths across 82,770 problems, and developed a Progressive Paradigm Training strategy. They used this process to train CoR-Math-7B, fine-tuned from DeepSeekMath-Base-7B, alongside an inference technique called Sequential Multi-Paradigm Sampling that systematically expands reasoning paths across paradigms.
The findings establish substantial performance advantages across standard benchmarks. In formal theorem proving on the miniF2F benchmark, CoR-Math-7B achieved a 66.0% accuracy in a zero-shot setting, representing a 41.0% absolute improvement over GPT-4o's few-shot performance and surpassing specialized few-shot proof search systems while utilizing far fewer generated candidate paths. In arithmetic computation, the model achieved 66.7% zero-shot accuracy on the MATH benchmark, outperforming GPT-4 by 24.2% and reinforcement-learning-based baseline DeepSeekMath-RL-7B by 15.0%. Ablation evaluations revealed that chaining paradigms sequentially—specifically natural language followed by symbolic reasoning, then algorithmic execution—significantly improved outcomes by enabling structured decomposition and cross-paradigm self-correction.
These results demonstrate that multi-paradigm collaboration provides higher resource efficiency and accuracy than scaling single-paradigm search spaces. By allowing models to autonomously or instructionally switch reasoning modes, organizations can lower inference costs and reduce the amount of fine-tuning data required to achieve high-accuracy mathematical reasoning. The findings suggest that future development in AI reasoning should prioritize multi-medium integration over purely increasing parameter size or single-paradigm search budgets.
Organizations evaluating this approach should consider implementing multi-paradigm reasoning pipelines for complex numerical and logical automation, while ensuring pre-deployment security audits are conducted on external execution environments such as Python interpreters. Future research should expand multi-paradigm evaluations to broader enterprise tasks and test-time search budgets on large-scale arithmetic datasets.
Readers should note certain limitations: performance in symbolic reasoning remains constrained by the relative scarcity of formal proof corpora in pre-training data, and fixed model context windows can occasionally truncate long reasoning chains, leading to syntax or completion errors. Nonetheless, the high confidence of the findings across multiple standard benchmarks supports multi-paradigm reasoning as an effective architecture for complex problem solving.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Chain-of-Thought Prompting establishes the natural-language step-by-step reasoning paradigm that CoR combines with algorithmic and symbolic approaches.
- Paper: PAL: Program-aided Language Models, Luyu Gao et al. (2023). PAL shows how language-model reasoning can delegate computation to executable programs, providing a concrete foundation for CoR’s algorithmic paradigm.
- Paper: LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers, Theo Olausson et al. (2023). LINC demonstrates how language models can pair natural-language interpretation with formal logic and a symbolic prover, clarifying CoR’s symbolic-reasoning component.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). Program of Thoughts separates linguistic problem-solving from executable computation, making the relationship between CoR’s natural-language and algorithmic paradigms easier to follow.
No sufficiently relevant recommendations were found.
