ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering

Zhiyu ChenShiyang LiCharese SmileyZhiqiang MaSameena ShahWilliam Yang Wang

article2022EMNLP212 citations

Introduces a large-scale financial question answering benchmark to evaluate how language models execute multi-turn, chained numerical reasoning over complex financial reports.

Listen

Modern artificial intelligence models have achieved remarkable success at general language pattern matching, but automating complex, multi-step numerical analysis in specialized domains remains a persistent bottleneck. In corporate finance, analysts rarely ask isolated questions; instead, they conduct dynamic conversations over complex financial reports, forming sequential reasoning chains that build upon prior context. The article introduces and evaluates a new benchmark dataset called CONVFINQA to investigate how effectively artificial intelligence systems can navigate multi-turn numerical reasoning in conversational financial question answering.

To rigorously evaluate machine performance, the article developed a dataset comprising 3,892 multi-turn conversations and 14,115 questions grounded in real-world corporate financial filings containing both text and structured tables. The conversation flows were constructed through a two-step framework that systematically decomposed and integrated multi-hop calculations into single-turn queries, which professional financial annotators then authored into realistic dialogue. The researchers benchmarked two primary modeling approaches against human performance: fully trained specialized neural symbolic pipelines that explicitly retrieve data and generate executable mathematical programs, and few-shot prompting techniques utilizing large-scale language models like GPT-3.

The findings establish that complex conversational numerical reasoning is far from solved. Human financial experts achieved an execution accuracy of 89.44%, whereas general non-expert crowd workers achieved only 46.90%. The best-performing neural symbolic model reached 68.90% execution accuracy when retrieving its own facts and 77.32% when supplied with perfect retrieval facts. In contrast, few-shot prompting methods using GPT-3 reached a maximum accuracy of only 45.15% to 50.30%, performing on par with non-expert crowd workers despite receiving perfect factual context. Across all systems, performance declined sharply as conversation length and reasoning dependency increased, with accuracy dropping substantially on hybrid multi-topic conversations and later conversation turns.

These results indicate that general-purpose large language models struggle to manage the sequential conversational context and specialized domain logic required for professional financial reasoning. Rather than understanding the multi-step conversational structure, prompt-based models frequently relied on superficial pattern imitation or defaulted to simple internal arithmetic, leading to compounding errors across long reasoning chains. For decision-makers evaluating automation in finance, deploying off-the-shelf generative language models presents significant operational and accuracy risks. Specialized architectures that pair domain-specific retrieval with formal symbolic program execution remain substantially more reliable.

The article recommends that organizations developing conversational financial tools focus on specialized, domain-tailored neural symbolic architectures rather than relying solely on prompted generalist models. Current systems should serve to assist human analysts rather than replace expert judgment. Further research is necessary to explore broader conversational structures, enhance deep financial domain knowledge within language encoders, and evaluate emerging larger foundation models with advanced prompt engineering.

arXiv: 2210.03849czyssrs/ConvFinQA
Cover for ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering

Abstract

With the recent advance in large pre-trained language models, researchers have achieved record performances in NLP tasks that mostly focus on language pattern matching. The community is experiencing the shift of the challenge from how to model language to the imitation of complex reasoning abilities like human beings. In this work, we investigate the application domain of finance that involves real-world, complex numerical reasoning. We propose a new large-scale dataset, CONVFINQA, aiming to study the chain of numerical reasoning in conversational question answering. Our dataset poses great challenge in modeling long-range, complex numerical reasoning paths in real-world conversations. We conduct comprehensive experiments and analyses with both the neural symbolic methods and the prompting-based methods, to provide insights into the reasoning mechanisms of these two divisions. We believe our new dataset should serve as a valuable resource to push forward the exploration of real-world, complex reasoning tasks as the next research focus. Our dataset and code is publicly available¹.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Task Formulation
  • 4 The CONVFINQA Dataset
  • 4.1 Dataset Construction
  • 4.2 Dataset Analysis
  • 5 Experiments on Neural Symbolic Approaches
  • 5.1 Methods and Main Results
  • 5.2 Performance Breakdown
  • 5.3 Analyses and Findings
  • 6 Experiments on Prompting-Based Approaches
  • 6.1 Methods and Main Results
  • 6.2 Performance Breakdown
  • 6.3 Analyses and Findings
  • 7 Conclusion and Discussion
  • 8 Limitations
  • 9 Ethical Considerations
  • Acknowledgment
  • References
  • Appendix A: Operation Definitions
  • Appendix B: Annotation Interface
  • Appendix C: Experiment Details
  • Appendix D: Prompt Details

Knowls

  1. Knowl 1 — Task Formulation and Evaluation Metrics in Conversational Finance Question Answering

    equation

    In the CONVFINQA setting, a system is given a financial report containing both unstructured text TT and a structured table BB, along with a sequence of conversational questions (Q0,Q1,…,Qn)(Q_0, Q_1, \dots, Q_n) where question QiQ_i (i>0i > 0) may depend on previous conversational context (Q0,…,Qi−1)(Q_0, \dots, Q_{i-1}) and their intermediate resolutions. The goal is to generate an executable reasoning program GG to yield the final numerical or boolean answer AA for the latest turn QnQ_n.

    The conditional probability of obtaining the correct answer AA is modeled by summing over all valid latent reasoning programs Gi∈{Gi}G_i \in \{G_i\} that evaluate to AA:

    P(A∣T,B,Qn)=∑GiP(Gi∣T,B,Q0,Q1,…,Qn−1)P(A \mid T, B, Q_n) = \sum_{G_i} P(G_i \mid T, B, Q_0, Q_1, \dots, Q_{n-1})

    A reasoning program GG is structured as a sequence of domain-specific operational clauses:

    op1[args1],op2[args2],…,opk[argsk]\text{op}_1[\text{args}_1], \text{op}_2[\text{args}_2], \dots, \text{op}_k[\text{args}_k]

    where each operation opj\text{op}_j takes arguments argsj\text{args}_j composed of constants, numbers extracted from TT or BB, or references to prior step outputs denoted by #m (indicating the output of the mm-th operation clause).

    Performance is evaluated under two primary metrics:

    1. Execution Accuracy (Exe Acc): Measures whether the evaluated mathematical result of the generated program matches the ground-truth numerical/boolean answer.
    2. Program Accuracy (Prog Acc): Measures whether the predicted program is structurally equivalent to the ground-truth reasoning program.
  2. Knowl 2 — Domain-Specific Language Operations for Financial Reasoning

    data/table

    The domain-specific language (DSL) for financial question answering in CONVFINQA consists of six core operations designed to express mathematical reasoning over numerical facts in financial reports:

    Name Arguments Output Description
    add number1, number2 number Addition: number1+number2\text{number1} + \text{number2}
    subtract number1, number2 number Subtraction: number1−number2\text{number1} - \text{number2}
    multiply number1, number2 number Multiplication: number1⋅number2\text{number1} \cdot \text{number2}
    divide number1, number2 number Division: number1/number2\text{number1} / \text{number2}
    exp number1, number2 number Exponentiation: number1number2\text{number1}^{\text{number2}}
    greater number1, number2 bool Comparison: number1>number2\text{number1} > \text{number2}

    Arguments can be raw numbers from the input financial tables/texts, fixed constant values (such as 100 or 1), or reference pointers (e.g., #0, #1) representing the return values of earlier operations in multi-step reasoning programs.

  3. Knowl 3 — Simulation and Composition Framework for Conversational Reasoning Datasets

    model/method

    CONVFINQA constructs multi-turn conversational reasoning chains from financial reports via a two-stage process:

    1. Conversational QA Flow Simulation:

      • Type I (Simple conversations): A multi-hop reasoning program from the FinQA dataset is decomposed into single-step operations. When a step introduces a new numerical value, a number-selection query turn is randomly inserted prior to that calculation step to simulate surface fact queries.
      • Type II (Hybrid conversations): Two distinct multi-hop reasoning programs associated with the same financial report are decomposed, augmented with surface number queries, and concatenated into a unified skeleton to simulate multi-turn conversations that bridge distinct but related financial topics with cross-turn dependencies.
    2. Expert Question Composition:

      • Financial experts (CPAs, MBAs) review the simulated conversation skeletons and underlying financial reports.
      • Annotators realize each skeleton turn into natural language questions.
      • Annotators are instructed to skip redundant steps and compress artificial turns using contextual references and coreferences to reflect natural dialogue flow, or discard invalid skeletons entirely.
  4. Knowl 4 — CONVFINQA Dataset Statistics and Properties

    data/table

    The CONVFINQA dataset contains 3,892 multi-turn conversations encompassing 14,115 questions across 2,066 financial report pages. The dataset is split into 3,037 train, 421 development, and 434 test conversations. It contains 2,715 Type I simple conversations and 1,177 Type II hybrid conversations.

    Statistic Value
    Total Conversations 3,892
    Total Questions 14,115
    Report Pages 2,066
    Vocabulary Size 20k
    Average Questions per Conversation 3.67
    Average Question Length (tokens) 10.59
    Average Sentences in Input Text 23.65
    Average Rows in Input Table 6.39
    Average Total Tokens in Inputs (Text Table) 675.61
    Maximum Total Tokens in Inputs (Text Table) 2,338.00

    Key structural properties of the reasoning tasks:

    • Dependency Distance: Over 60% of questions require viewing previous dialogue turns. In hybrid conversations, 65.0% of turns in the second question set depend directly on the first question set.
    • Reasoning Complexity: 34.73% of questions are single number selection queries, 35.10% require 1-step calculation programs, 25.41% require 2-step programs, and 4.75% require ≥3\ge 3 steps.
    • Evidence Modality: 59.18% of questions draw supporting evidence solely from tables, 25.56% solely from text, and 15.26% require combining evidence from both text and tables.
    • Operation Frequency: Subtraction accounts for 40.49% of operations, division 33.43%, addition 18.80%, and multiplication 6.92%.
  5. Knowl 5 — Expert vs. Crowd Worker Human Performance Baseline on CONVFINQA

    empirical result

    Human evaluation on 200 sampled questions from CONVFINQA establishes the performance upper bound and illustrates the requirement for specialized domain knowledge:

    • Domain Experts (Finance professionals, CPAs, MBAs) achieved an average Execution Accuracy of 89.44%89.44\% and a Program Accuracy of 86.34%86.34\%, with inter-annotator agreement exceeding 85.0%85.0\% on both metrics.
    • General Crowd Workers (Amazon Mechanical Turk) achieved an Execution Accuracy of only 46.90%46.90\% and a Program Accuracy of 45.52%45.52\%, with agreement rates below 60.0%60.0\%.

    This large performance gap (42.54%42.54\% in execution accuracy) demonstrates that conversational numerical reasoning over financial filings requires specialized domain expertise rather than basic general-domain comprehension.

  6. Knowl 6 — Neural Symbolic Model Performance on CONVFINQA

    empirical result

    Neural symbolic methods were evaluated on the CONVFINQA test set using a two-stage retriever-generator pipeline. The retriever retrieves supporting textual and tabular facts using conversation history up to the current turn (achieving 86.38%86.38\% recall for the top 3 facts). The generator generates structured reasoning programs given the retrieved evidence and conversational history.

    Model Execution Accuracy (%) Program Accuracy (%)
    GPT-2 (medium) 58.19 57.00
    T5 (large) 58.66 57.05
    FinQANet (BERT-base) 55.03 54.57
    FinQANet (BERT-large) 61.14 60.55
    FinQANet (RoBERTa-base) 64.95 64.16
    FinQANet (RoBERTa-large) 68.90 68.24
    FinQANet-Gold (RoBERTa-large) 77.32 76.46
    Human Expert 89.44 86.34
    General Crowd 46.90 45.52

    FinQANet (RoBERTa-large) achieves the best performance (68.90%68.90\% Execution Accuracy), outperforming pure generative language models (GPT-2, T5) due to structure-constrained program decoding. Supplying gold retrieved facts (FinQANet-Gold) increases execution accuracy to 77.32%77.32\%, indicating that retriever errors account for an 8.42%8.42\% performance gap.

  7. Knowl 7 — Performance Breakdown and Error Modes of Neural Symbolic Models

    empirical result

    Performance breakdown of the top-performing neural symbolic model (FinQANet with RoBERTa-large) across conversation types and turn positions reveals distinct reasoning bottlenecks:

    Subset / Category Execution Accuracy (%) Program Accuracy (%)
    Full Test Set 68.90 68.24
    By Turn Type
    Number Selection Questions 82.54 82.34
    Program / Calculation Questions 62.14 61.26
    By Conversation Complexity
    Simple Conversations (Type I) 72.37 72.00
    Hybrid Conversations (Type II) 60.99 59.70
    – Hybrid Conversations (First Part) 68.11 66.54
    – Hybrid Conversations (Second Part) 52.38 51.43

    Key behavioral patterns and failure modes:

    1. Number Selection vs. Calculation: The model performs substantially better on single number retrieval (82.54%82.54\%) than on generating multi-step programs (62.14%62.14\%).
    2. Long-Range Context Dependency: Accuracy declines steadily as the turn index nn increases within a conversation. The second half of hybrid conversations suffers the lowest accuracy (52.38%52.38\%) because models struggle to determine when to attend to or discard earlier turns across topic shifts.
    3. Domain Knowledge Deficits: Missing domain-specific accounting knowledge causes retrieval errors, erroneous value bindings, and invalid mathematical operations in multi-step reasoning chains.
  8. Knowl 8 — Prompting-Based Few-Shot Evaluation of GPT-3 on Conversational Finance QA

    empirical result

    Few-shot prompting experiments using OpenAI's GPT-3 (text-davinci-002) evaluated four prompt formats on the CONVFINQA test set using gold retrieved evidence and 10 exemplars (averaged across 3 distinct exemplar sets):

    Prompting Strategy Execution Accuracy (%) Program Accuracy (%)
    Answer-only 24.09 ±\pm 0.61 –
    Program-original 40.81 ±\pm 4.68 36.62 ±\pm 4.22
    Program-normal 45.15 ±\pm 2.77 38.88 ±\pm 2.57
    Chain of Thought (CoT) 40.63 ±\pm 1.25 33.84 ±\pm 2.19
    Human Expert Baseline 89.44 86.34
    General Crowd Baseline 46.90 45.52

    Key observations:

    • Program-normal outperforms Program-original: Standard infix arithmetic expressions (e.g., a1−a2a_1 - a_2) align better with pre-training data than domain-specific prefix notations (e.g., subtract(a1, a2)), where GPT-3 frequently makes syntactic errors.
    • CoT Prompting degrades performance: Chain of Thought prompting (40.63%40.63\%) underperforms direct program generation (45.15%45.15\%), as natural language reasoning steps frequently cause the model to hallucinate or drift from the conversational context.
    • Comparison to Supervised Baselines: Even with gold retrieved evidence provided in the prompt, GPT-3 (45.15%45.15\%) falls significantly behind supervised FinQANet-Gold (77.32%77.32\%) and human experts (89.44%89.44\%).
  9. Knowl 9 — Behavioral Failure Modes and Turn Breakdown of GPT-3 Prompting

    empirical result

    Detailed evaluation of the best-performing GPT-3 configuration (Program-normal with gold retrieved context) highlights core deficiencies in conversational reasoning:

    Category / Turn Type Execution Accuracy (%) Program Accuracy (%)
    Full Test Set 48.85 42.14
    Number Selection Questions 35.32 34.72
    Program / Calculation Questions 55.56 45.82
    Simple Conversations 52.22 46.64
    Hybrid Conversations 41.16 31.90
    – Hybrid Conversations (First Part) 56.30 48.03
    – Hybrid Conversations (Second Part) 22.85 12.38

    Major findings on prompted large language model behavior:

    1. Inability to Resolve Conversational Context: Unlike supervised models, GPT-3 performs worse on number selection turns (35.32%35.32\%) than program turns (55.56%55.56\%). It struggles with conversational anaphora (e.g., querying values for "the subsequent year" results in re-selecting the previous year's value).
    2. Second-Turn Failure: GPT-3 accuracy collapses on the second turn of conversations where questions directly depend on resolving references to Turn 1.
    3. Exemplar Over-Fitting and Superficial Mimicry: GPT-3 often blindly copies the reasoning program structure of in-context exemplars while ignoring the actual numbers and text of the input question.
    4. Exemplar Scaling Plateau: Increasing in-context exemplars from 5 to 20 yields minor gains (43.52%43.52\% to 50.30%50.30\% Exe Acc), plateauing at 25 exemplars (49.90%49.90\%).
  10. Knowl 10 — Limitations of CONVFINQA Dataset and Study

    limitation

    The authors identify two key limitations in their benchmark and experimental setup:

    1. Synthetic Scope of Dialogue Flows: Conversation skeletons are constructed through the decomposition of single FinQA programs (Type I) or the concatenation and decomposition of pairs of FinQA programs (Type II). While expert re-annotation ensures natural textual phrasing, this simulation mechanism does not cover all open-ended, exploratory, or multi-party conversation patterns encountered in real-world financial workflows.
    2. Prompting Exploration Constraints: Due to API usage costs and model accessibility, few-shot prompting experiments were restricted to OpenAI's GPT-3 (text-davinci-002) evaluated with gold retrieved contexts on the test set, omitting extensive automated prompt optimization or testing on other contemporary large models such as PaLM.

Coverage note — None was omitted; all contributed material—including dataset creation methodologies, dataset statistics, human baselines, neural symbolic experimental benchmarks, prompting results, breakdown analyses, and stated limitations—is fully captured.

References

  1. 1.Md. Shad Akhtar, Abhishek Kumar, Deepanway Ghosal, Asif Ekbal, and Pushpak Bhattacharyya. 2017. A multilayer perceptron based ensemble technique for fine-grained financial sentiment analysis. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017, pages 540–546. Association for Computational Linguistics.
  2. 2.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 2357–2367. Association for Computational Linguistics.
  3. 3.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  4. 4.Xinyun Chen, Chen Liang, Adams Wei Yu, Denny Zhou, Dawn Song, and Quoc V. Le. 2020. Neural symbolic reader: Scalable integration of distributed and symbolic representations for reading comprehension. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  5. 5.Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan R. Routledge, and William Yang Wang. 2021. Finqa: A dataset of numerical reasoning over financial data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 3697–3711. Association for Computational Linguistics.
  6. 6.Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wentau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. Quac: Question answering in context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 2174–2184. Association for Computational Linguistics.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. CoRR, abs/2204.02311.
  8. 8.Min-Yuh Day and Chia-Chou Lee. 2016. Deep learning for financial sentiment analysis on finance news providers. In 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining, ASONAM 2016, San Francisco, CA, USA, August 18-21, 2016, pages 1127–1134. IEEE Computer Society.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics.
  10. 10.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACLHLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 2368–2378. Association for Computational Linguistics.
  11. 11.Jingguang Han, Utsab Barman, Jer Hayes, Jinhua Du, Edward Burgin, and Dadong Wan. 2018. Nextgen AML: distributed deep learning based language technologies to augment anti money laundering investigation. In Proceedings of ACL 2018, Melbourne, Australia, July 15-20, 2018, System Demonstrations, pages 37–42. Association for Computational Linguistics.
  12. 12.Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. 2017. Search-based neural structured learning for sequential question answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1821–1831. Association for Computational Linguistics.
  13. 13.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  14. 14.Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. MAWPS: A math word problem repository. In NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016, pages 1152–1157. The Association for Computational Linguistics.
  15. 15.Chen Liang, Jonathan Berant, Quoc V. Le, Kenneth D. Forbus, and Ni Lao. 2017. Neural symbolic machines: Learning semantic parsers on freebase with weak supervision. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 23–33. Association for Computational Linguistics.
  16. 16.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  17. 17.Zhuang Liu, Degen Huang, Kaiyu Huang, Zhuang Li, and Jun Zhao. 2020. Finbert: A pre-trained financial language representation model for financial text mining. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI 2020, pages 4513–4519. ijcai.org.
  18. 18.Armineh Nourbakhsh and Grace Bang. 2019. A framework for anomaly detection using language modeling, and its applications to finance. CoRR, abs/1908.09156.
  19. 19.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  20. 20.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  21. 21.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100, 000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pages 2383–2392. The Association for Computational Linguistics.
  22. 22.Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. Coqa: A conversational question answering challenge. Trans. Assoc. Comput. Linguistics, 7:249–266.
  23. 23.Amrita Saha, Vardaan Pahuja, Mitesh M. Khapra, Karthik Sankaranarayanan, and Sarath Chandar. 2018. Complex sequential question answering: Towards learning to converse over linked question answer pairs with a knowledge graph. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, pages 705–713. AAAI Press.
  24. 24.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M. Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2021. Multitask prompted training enables zero-shot task generalization. CoRR, abs/2110.08207.
  25. 25.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008.
  26. 26.Weikang Wang, Jiajun Zhang, Qian Li, Chengqing Zong, and Zhifei Li. 2019. Are you for real? detecting identity fraud via dialogue interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 1762–1771. Association for Computational Linguistics.
  27. 27.Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei, Xudong Lin, Shuohang Wang, Ziyi Yang, Chenguang Zhu, Derek Hoiem, Shih-Fu Chang, Mohit Bansal, and Heng Ji. 2022. Language models with image descriptors are strong few-shot video-language learners. CoRR, abs/2205.10747.
  28. 28.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. CoRR, abs/2201.11903.
  29. 29.Munazza Zaib, Wei Emma Zhang, Quan Z. Sheng, Adnan Mahmood, and Yang Zhang. 2021. Conversational question answering: A survey. CoRR, abs/2106.00874.
  30. 30.Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 3277–3287. Association for Computational Linguistics.

Citation

MLA
Chen, Z., et al. “ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 6279–92, https://doi.org/10.18653/v1/2022.emnlp-main.421.
APA
Chen, Z., Li, S., Smiley, C., Ma, Z., Shah, S., & Wang, W. Y. (2022). ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6279–6292. https://doi.org/10.18653/v1/2022.emnlp-main.421
Chicago
Chen, Z., S. Li, C. Smiley, Z. Ma, S. Shah, and W. Y. Wang. 2022. “ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6279–92. https://doi.org/10.18653/v1/2022.emnlp-main.421.
Harvard
Chen, Z. et al. (2022) “ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6279–6292. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.421.
Vancouver
1. Chen Z, Li S, Smiley C, Ma Z, Shah S, Wang WY (2022) ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6279–6292

BibTeX

@inproceedings{chen-etal-2022-convfinqa,
    title = "{C}onv{F}in{QA}: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering",
    author = "Chen, Zhiyu  and
      Li, Shiyang  and
      Smiley, Charese  and
      Ma, Zhiqiang  and
      Shah, Sameena  and
      Wang, William Yang",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.421/",
    doi = "10.18653/v1/2022.emnlp-main.421",
    pages = "6279--6292"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/