ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering
Zhiyu ChenShiyang LiCharese SmileyZhiqiang MaSameena ShahWilliam Yang Wang
Introduces a large-scale financial question answering benchmark to evaluate how language models execute multi-turn, chained numerical reasoning over complex financial reports.
Modern artificial intelligence models have achieved remarkable success at general language pattern matching, but automating complex, multi-step numerical analysis in specialized domains remains a persistent bottleneck. In corporate finance, analysts rarely ask isolated questions; instead, they conduct dynamic conversations over complex financial reports, forming sequential reasoning chains that build upon prior context. The article introduces and evaluates a new benchmark dataset called CONVFINQA to investigate how effectively artificial intelligence systems can navigate multi-turn numerical reasoning in conversational financial question answering.
To rigorously evaluate machine performance, the article developed a dataset comprising 3,892 multi-turn conversations and 14,115 questions grounded in real-world corporate financial filings containing both text and structured tables. The conversation flows were constructed through a two-step framework that systematically decomposed and integrated multi-hop calculations into single-turn queries, which professional financial annotators then authored into realistic dialogue. The researchers benchmarked two primary modeling approaches against human performance: fully trained specialized neural symbolic pipelines that explicitly retrieve data and generate executable mathematical programs, and few-shot prompting techniques utilizing large-scale language models like GPT-3.
The findings establish that complex conversational numerical reasoning is far from solved. Human financial experts achieved an execution accuracy of 89.44%, whereas general non-expert crowd workers achieved only 46.90%. The best-performing neural symbolic model reached 68.90% execution accuracy when retrieving its own facts and 77.32% when supplied with perfect retrieval facts. In contrast, few-shot prompting methods using GPT-3 reached a maximum accuracy of only 45.15% to 50.30%, performing on par with non-expert crowd workers despite receiving perfect factual context. Across all systems, performance declined sharply as conversation length and reasoning dependency increased, with accuracy dropping substantially on hybrid multi-topic conversations and later conversation turns.
These results indicate that general-purpose large language models struggle to manage the sequential conversational context and specialized domain logic required for professional financial reasoning. Rather than understanding the multi-step conversational structure, prompt-based models frequently relied on superficial pattern imitation or defaulted to simple internal arithmetic, leading to compounding errors across long reasoning chains. For decision-makers evaluating automation in finance, deploying off-the-shelf generative language models presents significant operational and accuracy risks. Specialized architectures that pair domain-specific retrieval with formal symbolic program execution remain substantially more reliable.
The article recommends that organizations developing conversational financial tools focus on specialized, domain-tailored neural symbolic architectures rather than relying solely on prompted generalist models. Current systems should serve to assist human analysts rather than replace expert judgment. Further research is necessary to explore broader conversational structures, enhance deep financial domain knowledge within language encoders, and evaluate emerging larger foundation models with advanced prompt engineering.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces the foundational chain-of-thought prompting methodology that ConvFinQA directly adapts and evaluates across multi-turn conversational numerical reasoning chains.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Provides the foundational subproblem decomposition framework used to conceptualize and construct multi-hop conversational question chains.
- Paper: DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs, Dheeru Dua et al. (2019). Establishes standard discrete and arithmetic reasoning tasks over text passages that precede and motivate multi-turn financial question answering.
- Paper: CoQA: A Conversational Question Answering Challenge, Siva Reddy et al. (2018). Pioneers conversational question answering benchmarks requiring contextual dialogue tracking that ConvFinQA extends into the numerical finance domain.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Demonstrates step-by-step reasoning capabilities in zero-shot large language models, providing crucial baseline prompting context for the experiments in ConvFinQA.
- Paper: NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks, Swaroop Mishra et al. (2022). Benchmarks multi-task numerical and arithmetic reasoning in language models, framing the difficulty of mathematical problem-solving in natural language.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). Directly tackles ConvFinQA and financial QA challenges by disentangling reasoning from calculation through executable program generation.
- Paper: Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning, Pan Lu et al. (2023). Advances tabular mathematical reasoning by using reinforcement learning to dynamically optimize demonstration prompts over structured and semi-structured data.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). Develops a general iterative reading-then-reasoning framework enabling large language models to systematically query and reason over structured tables and graphs.
- Paper: TableBench: A Comprehensive and Complex Benchmark for Table Question Answering, Xianjie Wu et al. (2025). Expands on complex tabular reasoning benchmarks across diverse corporate workflows and multi-step analytical capabilities.
- Paper: ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness, Archiki Prasad et al. (2023). Proposes a fine-grained evaluation metric to assess the correctness and informativeness of multi-step intermediate reasoning chains.
- Paper: Large Language Models Can Be Easily Distracted by Irrelevant Context, Freda Shi et al. (2023). Analyzes the vulnerability of language models to distracting context during multi-step arithmetic reasoning, extending the error analysis seen in conversational QA.
