Exploring Chain of Thought Style Prompting for Text-to-SQL

Chang-Yu TaiZiru ChenTianshu ZhangXiang DengHuan Sun

article2023EMNLP128 citations

Presents a single-pass question-decomposition prompting method for text-to-SQL parsing that reduces error propagation and outperforms multi-turn iterative prompting on the Spider benchmarks without requiring expensive intermediate execution calls.

Listen

Converting everyday human language into database queries—known as text-to-SQL parsing—is essential for building intuitive AI database assistants. Traditional approaches require large, expensive labeled datasets to train specialized models. While large language models can perform this task using only a few examples in a prompt, their accuracy remains limited when dealing with complex, multi-step relational reasoning.

The article systematically evaluates how chain-of-thought prompting strategies can improve the reasoning capabilities of large language models for text-to-SQL parsing. Specifically, it demonstrates how decomposing questions into simpler steps in a single prompt outperforms both standard prompting and iterative problem-solving techniques.

To conduct this evaluation, the researchers adapted two existing prompting paradigms: standard chain-of-thought, which mimics SQL query execution steps, and least-to-most prompting, which iteratively generates and solves sub-questions across multiple stages. Identifying critical shortcomings in these methods, they designed a novel approach called Question Decomposition (QDecomp) and an enhanced variant, QDecomp+InterCOL, which incorporates relevant table and column names into each decomposed step. The evaluation was carried out primarily using OpenAI's Codex model on cross-domain benchmarks, including Spider and Spider Realistic, alongside three single-domain datasets (GeoQuery, IMDB, and Yelp), measuring performance with strict test-suite execution accuracy.

The findings establish that the proposed QDecomp+InterCOL approach delivers superior performance. On the Spider development and Spider Realistic benchmarks, QDecomp+InterCOL achieved test-suite accuracies of 68.4% and 56.5%, representing absolute improvements of 5.2 and 6.5 percentage points over standard prompting without reasoning. Furthermore, it outperformed least-to-most prompting by 2.4 and 1.5 percentage points. By contrast, traditional chain-of-thought prompting fell behind standard baseline prompting (56.8% versus 63.2%), as generating granular SQL execution steps introduced severe reasoning errors that derailed the final query. The experiments also revealed that concise API-style schema descriptions performed on par with lengthy database dumps while consuming far fewer tokens.

These results demonstrate that multi-step iterative prompting is unnecessary and computationally inefficient for text-to-SQL tasks. Iterative pipelines and overly detailed reasoning steps amplify error propagation, where an early mistake compounds and ruins the final database query. By prompting the model to decompose questions and predict intermediate database schema elements in a single generation pass, organizations can achieve higher query accuracy with lower computational overhead and reduced latency.

Organizations developing database agents should adopt single-pass question decomposition prompts enriched with intermediate schema targets rather than complex iterative prompting frameworks. For production systems, teams can integrate this decomposed natural-language output into interactive semantic parsing interfaces, allowing end users to review and correct intermediate sub-questions before final query execution.

Readers should note that the empirical results rely primarily on OpenAI's Codex model and a fixed set of academic benchmarks. Future evaluations should validate these single-pass decomposition strategies on newer models, such as GPT-4, and assess resilience against real-world database perturbations before large-scale enterprise deployment.

arXiv: 2305.14215

No sufficiently relevant recommendations were found.

Cover for Exploring Chain of Thought Style Prompting for Text-to-SQL

Abstract

In-context learning with large language models (LLMs) has recently caught increasing attention due to its superior few-shot performance on various tasks. However, its performance on text-to-SQL parsing still has much room for improvement. In this paper, we hypothesize that a crucial aspect of LLMs to improve for text-to-SQL parsing is their multi-step reasoning ability. Thus, we systematically study how to enhance LLMs’ reasoning ability through chain of thought (CoT) style prompting, including the original chain-of-thought prompting (Wei et al., 2022b) and least-to-most prompting (Zhou et al., 2023). Our experiments demonstrate that iterative prompting as in Zhou et al. (2023) may be unnecessary for text-to-SQL parsing, and using detailed reasoning steps tends to have more error propagation issues. Based on these findings, we propose a new CoT-style prompting method for text-to-SQL parsing. It brings 5.2 and 6.5 point absolute gains on the Spider development set and the Spider Realistic set, respectively, compared to the standard prompting method without reasoning steps; 2.4 and 1.5 point absolute gains, compared to the least-to-most prompting method¹.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Prompting for Multi-Step Reasoning in Text-to-SQL
  • 3.1 Chain-of-Thought Prompting
  • 3.2 Least-to-Most Prompting
  • 3.3 Question Decomposition Prompting
  • 4 Experimental Setup
  • 4.1 Datasets
  • 4.2 In-context Example Selection
  • 4.3 Prompt Formats
  • 4.4 Evaluation Metric
  • 5 Results and Analysis
  • 5.1 Main Results
  • 5.2 Error Analysis
  • 5.3 Robustness to Prompt Design
  • 5.4 Results on Other Text-to-SQL Datasets
  • 6 Conclusion and Future Work
  • Limitations
  • Acknowledgements
  • References
  • A Example Prompts

Knowls

  1. Knowl 1 — Question decomposition prompting generates sub-questions and SQL in one pass

    model/method

    Question-decomposition prompting (QDecomp) adapts chain-of-thought prompting to text-to-SQL by asking a large language model (LLM) to produce a sequence of simpler sub-questions for the original natural-language question, followed by the SQL query, all in one generation pass. The intermediate steps are question reformulations, not descriptions of SQL execution and not intermediate SQL queries. The decomposition follows the problem-reduction principles used for least-to-most prompting: split multi-sentence questions; further split at conjunctions and prepositions where appropriate; and remove words or phrases from earlier sub-questions if they reveal information that should only appear later. The final SQL is generated after the complete sequence of sub-questions, allowing the model to retain the context of the full request.

  2. Knowl 2 — InterCOL adds incremental table and column grounding to decomposed questions

    model/method

    QDecomp+InterCOL augments each QDecomp sub-question with the database table and column names associated with that step. In the in-context examples, those names are annotated from the corresponding sub-question's SQL parse. If the parse uses *, one column is sampled at random from a table in that sub-query. A table-column pair already annotated at an earlier step is omitted from later annotations; if a step has no pair to annotate, one pair is randomly selected from a preceding step. The prompt therefore demonstrates how schema information is introduced incrementally as the question is decomposed, while the final SQL is still generated once, after the sequence of sub-questions.

  3. Knowl 3 — Chain-of-thought and least-to-most adaptations provide distinct text-to-SQL baselines

    model/method

    The chain-of-thought (CoT) adaptation represents the SQL query through natural-language descriptions of its clauses, ordered by the query's logical execution sequence; the LLM generates those descriptions and then the final SQL in one pass. The least-to-most adaptation instead has two stages. In problem reduction, the LLM decomposes the question using sentence boundaries, conjunctions or prepositions, and removal of information that would leak into an earlier sub-question. In problem solving, the LLM receives one sub-question at a time and produces SQL incrementally, using the prior SQL as context; the original question is the final sub-question. These adaptations let the experiments compare clause-by-clause reasoning, iterative question decomposition and solving, and the single-pass QDecomp approach.

  4. Knowl 4 — Evaluation uses Codex, Spider benchmarks, and test-suite execution accuracy

    experimental setup

    The experiments used OpenAI Codex code-davinci-002, accessed through the API from January through March 2023, with greedy decoding and temperature 0. Spider has 7,000 training question-query pairs and 1,034 development pairs across 200 databases and 138 domains; the unavailable Spider test set was not evaluated. Spider Realistic contains 508 questions adapted from Spider development questions by removing or paraphrasing explicit column-name mentions. Main experiments used eight in-context examples selected at random from Spider training data, the API Docs schema representation (table and column names), and five different example sets per method; reported means include standard deviations where provided. Test-suite accuracy compares predicted and gold query execution across many synthesized databases, reducing false positives possible when execution is compared on only one database. Standard execution accuracy is also reported parenthetically in the main comparison.

  5. Knowl 5 — QDecomp+InterCOL leads the main Spider comparisons

    data/table

    The following results compare prompting methods using eight API Docs examples on Spider development and Spider Realistic. The Spider development columns report test-suite accuracy by query difficulty and overall; Spider Realistic reports overall test-suite accuracy. Parentheses give standard execution accuracy. Overall means and standard deviations are reproduced as reported. QDecomp+InterCOL improves over standard prompting by 5.2 and 6.5 absolute percentage points in test-suite accuracy on Spider development and Spider Realistic, respectively, and exceeds least-to-most prompting by 2.4 and 1.5 points. With extra-hard-only examples (G3), QDecomp+InterCOL reaches 78.2% standard execution accuracy on Spider development. G3 results on Spider Realistic were unavailable.

    Method Easy Medium Hard Extra Hard Spider Dev Overall TS (EX) Spider Realistic Overall TS (EX)
    Standard 86.8 65.3 50.3 36.0 63.2±2.5163.2\pm2.51 (68.7±4.0868.7\pm4.08) 51.0±4.2951.0\pm4.29 (62.5±4.0162.5\pm4.01)
    Chain-of-Thought 73.9 64.5 44.6 23.4 56.8±5.8356.8\pm5.83 (53.9±7.2153.9\pm7.21) 50.3±4.9450.3\pm4.94 (53.4±9.1953.4\pm9.19)
    Least-to-Most 88.1 68.7 52.9 39.5 66.0±2.4866.0\pm2.48 (68.9±3.4468.9\pm3.44) 55.0±2.5155.0\pm2.51 (63.3±2.7363.3\pm2.73)
    Least-to-Most (G3) 80.3 64.6 52.8 45.3 63.3±1.9563.3\pm1.95 (73.8±1.7273.8\pm1.72) –
    QDecomp 89.8 71.3 53.1 38.6 67.4±1.8967.4\pm1.89 (70.7±2.8070.7\pm2.80) 55.8±2.0155.8\pm2.01 (65.8±2.2965.8\pm2.29)
    QDecomp + InterCOL 89.6 74.1 52.4 38.1 68.4±2.0568.4\pm2.05 (69.7±5.8269.7\pm5.82) 56.5±2.0556.5\pm2.05 (63.3±4.1963.3\pm4.19)
    QDecomp + InterCOL (G3) 88.7 71.1 56.8 45.7 68.8±1.1668.8\pm1.16 (78.2±1.0778.2\pm1.07) –

    Here, TS denotes test-suite accuracy, EX denotes standard execution accuracy, and G3 denotes selection of examples from the extra-hard difficulty level only.

  6. Knowl 6 — Detailed intermediate reasoning is associated with error propagation

    empirical result

    On Spider development, component matching accuracy evaluates SELECT, WHERE, GROUP BY, ORDER BY, and KEYWORDS (SQL keywords, operators, and column names). A component counts as correct when it matches exactly or when the full query has test-suite accuracy 1. QDecomp variants perform better than the other reasoning prompts across these five components. CoT performs worse than standard prompting overall: its detailed clause-by-clause explanations can contain schema or query-structure errors that the generated SQL then faithfully reproduces. Least-to-most improves on CoT but can make an incorrect intermediate SQL parse, particularly for difficult clauses such as JOIN, GROUP BY, and ORDER BY; later steps can inherit those mistakes. QDecomp avoids generating intermediate SQL, and its less detailed natural-language steps reduce opportunities for errors to accumulate.

    Method SELECT WHERE GROUP BY ORDER BY KEYWORDS
    Standard 89.8 66.1 74.7 83.0 84.2
    Chain-of-Thought 83.5 70.7 67.1 72.8 76.9
    Least-to-Most 90.0 70.7 72.5 82.4 84.3
    QDecomp 91.2 70.7 77.2 85.1 86.4
    QDecomp + InterCOL 91.4 72.4 76.6 85.3 86.0

    Entries are component matching accuracies in percent.

  7. Knowl 7 — One-pass decomposition makes iterative SQL solving unnecessary in these experiments

    empirical result

    QDecomp and QDecomp+InterCOL generate decomposed questions and the final SQL in one pass, yet outperform the iterative least-to-most approach on both Spider development and Spider Realistic test-suite accuracy in the main eight-example comparison. The results therefore indicate that iterative prompting is not necessary for improving text-to-SQL reasoning in these experiments. This comparison does not establish that iteration is universally unhelpful; the authors also note that the prompting methods differ in more than whether generation is iterative. Iterative least-to-most additionally incurs the cost of repeated model calls to solve the sub-questions.

  8. Knowl 8 — Example difficulty affects performance, with QDecomp+InterCOL remaining competitive

    empirical result

    The researchers varied the difficulty composition of eight in-context examples on Spider development. G1 samples equally across easy, medium, hard, and extra-hard queries; G2 samples equally from hard and extra-hard queries; G3 uses only extra-hard queries. QDecomp+InterCOL achieved its best easy and medium results with G1 and its best hard and extra-hard results with G3, consistent with the authors' conjecture that examples help the model perform at similar difficulty levels. Across all methods and selection strategies, QDecomp+InterCOL has the highest overall score in each of the four settings in the comparison below. Its results by difficulty are also shown; the reported standard deviations are included where available.

    QDecomp+InterCOL selection Easy Medium Hard Extra Hard Overall TS
    Random 89.6 74.1 52.4 38.1 68.4±2.0568.4\pm2.05
    G1 89.8 75.6 51.7 38.8 69.0±2.1869.0\pm2.18
    G2 87.4 72.2 50.4 39.4 66.9±2.3166.9\pm2.31
    G3 88.7 71.1 56.8 45.7 68.8±1.1668.8\pm1.16

    All entries are Spider development test-suite accuracy in percent; the second comparison reports overall accuracy.

  9. Knowl 9 — More in-context examples generally improve accuracy, with QDecomp robust at lower counts

    empirical result

    On Spider development, test-suite accuracy generally rises as the number of in-context examples increases from one to eight. QDecomp and QDecomp+InterCOL outperform standard prompting at each reported nonzero count, while least-to-most is below standard prompting at one and four examples. The proposed decomposition methods have no zero-shot result because the prompt needs at least one example to teach the step-by-step format. Preliminary experiments found only minor gains from adding examples beyond eight, so eight examples were used in the main comparisons.

    Method 0 examples 1 example 4 examples 8 examples
    Standard 59.6 62.0 63.9 63.2
    Least-to-Most – 59.2 62.1 65.8
    QDecomp – 63.1 66.6 67.4
    QDecomp + InterCOL – 61.4 66.5 68.4

    Entries are test-suite accuracy in percent.

  10. Knowl 10 — API Docs and richer schema examples yield similar four-shot results

    empirical result

    A four-example Spider development comparison tested API Docs prompts, which list table and column names, against Create Table + Select 3 prompts, which add SQL table definitions, types, foreign keys, and up to three example rows per table. The richer format does not consistently improve accuracy: its standard-prompting score is slightly higher, while QDecomp and QDecomp+InterCOL scores are slightly lower. QDecomp is the highest-scoring method under Create Table + Select 3. The authors use API Docs as the default because it is more compact and thus more efficient under prompt-length limits.

    Method API Docs Create Table + Select 3
    Standard 63.9 64.1
    Least-to-Most 62.1 63.8
    QDecomp 66.6 66.2
    QDecomp + InterCOL 66.5 64.3

    Entries are four-shot test-suite accuracy in percent.

  11. Knowl 11 — QDecomp variants lead on three additional text-to-SQL datasets

    data/table

    The authors tested four-shot prompting on GeoQuery, IMDB, and Yelp, whose database schemas and SQL queries they describe as more complex than those in Spider. QDecomp or QDecomp+InterCOL is the best-scoring method on each dataset, and QDecomp+InterCOL has the highest macro-average. Least-to-most falls below standard prompting on IMDB and Yelp. The authors attribute such failures in part to iterative error propagation: a sub-question may permit many plausible schema choices before later question details become available, and an incorrect early SQL parse can carry into subsequent steps.

    Method GeoQuery IMDB Yelp Macro-average
    Standard 60.99 73.28 45.31 59.86
    Least-to-Most 60.99 58.78 36.72 52.16
    QDecomp 64.84 77.86 48.44 63.71
    QDecomp + InterCOL 75.82 73.28 49.22 66.11

    Entries are four-shot test-suite accuracy in percent.

  12. Knowl 12 — The evaluation is limited to one LLM and a defined set of robustness variations

    limitation

    The experiments use Codex (code-davinci-002) only, so they do not establish whether the relative effects of CoT, least-to-most, QDecomp, and QDecomp+InterCOL hold for other or newer LLMs. Robustness analyses vary in-context example selection, example count, and prompt format, but do not comprehensively test changes to databases, natural-language questions, or SQL perturbations. The authors identify both cross-model evaluation and broader robustness testing as needed follow-up work.

Coverage note — No substantial contributed material was deliberately omitted; introductory background, related work, acknowledgements, and prompt-gallery examples were excluded because they add no distinct contribution beyond the methods and findings captured here.

References

  1. 1.Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, et al. 2023. Do as i can, not as i say: Grounding language in robotic affordances. In Conference on Robot Learning, pages 287–318. PMLR.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Shuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang, Jiarong Jiang, Joseph Lilien, Steve Ash, William Yang Wang, Zhiguo Wang, Vittorio Castelli, Patrick Ng, and Bing Xiang. 2023a. Dr.spider: A diagnostic evaluation benchmark towards text-to-SQL robustness. In The Eleventh International Conference on Learning Representations.
  4. 4.Shuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang, Jiarong Jiang, Joseph Lilien, et al. 2023b. Dr. spider: A diagnostic evaluation benchmark towards text-to-sql robustness. arXiv preprint arXiv:2301.08881.
  5. 5.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  6. 6.Shijie Chen, Ziru Chen, Huan Sun, and Yu Su. 2023a. Error detection for text-to-sql semantic parsing.
  7. 7.Ziru Chen, Shijie Chen, Michael White, Raymond Mooney, Ali Payani, Jayanth Srinivasa, Yu Su, and Huan Sun. 2023b. Text-to-SQL error correction with language models of code. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1359–1372, Toronto, Canada. Association for Computational Linguistics.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  9. 9.Deborah A Dahl, Madeleine Bates, Michael K Brown, William M Fisher, Kate Hunicke-Smith, David S Pallett, Christine Pao, Alexander Rudnicky, and Elizabeth Shriberg. 1994. Expanding the scope of the atis task: The atis-3 corpus. In Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8-11, 1994.
  10. 10.Xiang Deng, Ahmed Hassan, Christopher Meek, Oleksandr Polozov, Huan Sun, and Matthew Richardson. 2021. Structure-grounded pretraining for text-to-sql. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1337–1350.
  11. 11.Ahmed Elgohary, Saghar Hosseini, and Ahmed Hassan Awadallah. 2020. Speak to your parser: Interactive text-to-SQL with natural language feedback. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2065–2077, Online. Association for Computational Linguistics.
  12. 12.Ahmed Elgohary, Christopher Meek, Matthew Richardson, Adam Fourney, Gonzalo Ramos, and Ahmed Hassan Awadallah. 2021. NL-EDIT: Correcting semantic parse errors through natural language interaction. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5599–5610, Online. Association for Computational Linguistics.
  13. 13.Nitish Gupta and Mike Lewis. 2018. Neural compositional denotational semantics for question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2152–2161.
  14. 14.SU Hongjin, Jungo Kasai, Chen Henry Wu, Weijia Shi, Tianlu Wang, Jiayi Xin, Rui Zhang, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, et al. 2023. Selective annotation makes language models better few-shot learners. In The Eleventh International Conference on Learning Representations.
  15. 15.Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, Jayant Krishnamurthy, and Luke Zettlemoyer. 2017. Learning a neural semantic parser from user feedback. In 55th Annual Meeting of the Association for Computational Linguistics.
  16. 16.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems.
  17. 17.Yuntao Li, Bei Chen, Qian Liu, Yan Gao, Jian-Guang Lou, Yan Zhang, and Dongmei Zhang. 2020. “what do you mean by that?” a parser-independent interactive approach for enhancing text-to-SQL. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6913–6922, Online. Association for Computational Linguistics.
  18. 18.Aiwei Liu, Xuming Hu, Lijie Wen, and Philip S Yu. 2023a. A comprehensive evaluation of chatgpt’s zero-shot text-to-sql capability. arXiv preprint arXiv:2303.13547.
  19. 19.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023b. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1–35.
  20. 20.Joan C. Miller and Clifford J. Maloney. 1963. Systematic mistake analysis of digital computer programs. Commun. ACM, 6(2):58–63.
  21. 21.Sewon Min, Victor Zhong, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2019. Multi-hop reading comprehension through question decomposition and rescoring. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6097–6109.
  22. 22.Lingbo Mo, Ashley Lewis, Huan Sun, and Michael White. 2022. Towards transparent interactive semantic parsing via step-by-step correction. In Findings of the Association for Computational Linguistics: ACL 2022, pages 322–342, Dublin, Ireland. Association for Computational Linguistics.
  23. 23.Arpit Narechania, Adam Fourney, Bongshin Lee, and Gonzalo Ramos. 2021. Diy: Assessing the correctness of natural language to sql systems. In 26th International Conference on Intelligent User Interfaces, pages 597–607.
  24. 24.Ansong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov, Wen-Tau Yih, Sida Wang, and Xi Victoria Lin. 2023. LEVER: Learning to verify language-to-code generation with execution. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 26106–26128. PMLR.
  25. 25.Mohammadreza Pourreza and Davood Rafiei. 2023. Din-sql: Decomposed in-context learning of text-to-sql with self-correction.
  26. 26.Jiexing Qi, Jingyao Tang, Ziwei He, Xiangpeng Wan, Yu Cheng, Chenghu Zhou, Xinbing Wang, Quanshi Zhang, and Zhouhan Lin. 2022. RASAT: Integrating relational structures into pretrained Seq2Seq model for text-to-SQL. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3215–3229, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  27. 27.Nitarshan Rajkumar, Raymond Li, and Dzmitry Bahdanau. 2022. Evaluating the text-to-sql capabilities of large language models. arXiv preprint arXiv:2204.00498.
  28. 28.Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. 2021. Picard: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9895–9901.
  29. 29.Peter Shaw, Ming-Wei Chang, Panupong Pasupat, and Kristina Toutanova. 2021. Compositional generalization and natural language variation: Can a semantic parsing approach handle both? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 922–938.
  30. 30.Bailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov, and Matthew Richardson. 2020. Rat-sql: Relation-aware schema encoding and linking for text-to-sql parsers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7567–7578.
  31. 31.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022a. Emergent abilities of large language models. Transactions on Machine Learning Research.
  32. 32.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed H Chi, Quoc V Le, Denny Zhou, et al. 2022b. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  33. 33.Tomer Wolfson, Daniel Deutch, and Jonathan Berant. 2022. Weakly supervised text-to-sql parsing through question decomposition. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 2528–2542.
  34. 34.Tomer Wolfson, Mor Geva, Ankit Gupta, Matt Gardner, Yoav Goldberg, Daniel Deutch, and Jonathan Berant. 2020. Break it down: A question understanding benchmark. Transactions of the Association for Computational Linguistics, 8:183–198.
  35. 35.Navid Yaghmazadeh, Yuepeng Wang, Isil Dillig, and Thomas Dillig. 2017. Sqlizer: query synthesis from natural language. Proceedings of the ACM on Programming Languages, 1(OOPSLA):1–26.
  36. 36.Ziyu Yao, Yu Su, Huan Sun, and Wen-tau Yih. 2019. Model-based interactive semantic parsing: A unified framework and a text-to-SQL case study. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5447–5458, Hong Kong, China. Association for Computational Linguistics.
  37. 37.Ziyu Yao, Yiqi Tang, Wen-tau Yih, Huan Sun, and Yu Su. 2020. An imitation game for learning semantic parsers from user interaction. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6883–6902, Online. Association for Computational Linguistics.
  38. 38.Tao Yu, Chien-Sheng Wu, Xi Victoria Lin, Yi Chern Tan, Xinyi Yang, Dragomir Radev, Caiming Xiong, et al. 2021. Grappa: Grammar-augmented pre-training for table semantic parsing. In International Conference on Learning Representations.
  39. 39.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3911–3921.
  40. 40.John M Zelle and Raymond J Mooney. 1996. Learning to parse database queries using inductive logic programming. In Proceedings of the national conference on artificial intelligence, pages 1050–1055.
  41. 41.Jichuan Zeng, Xi Victoria Lin, Steven C.H. Hoi, Richard Socher, Caiming Xiong, Michael Lyu, and Irwin King. 2020. Photon: A robust cross-domain text-to-SQL system. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 204–214, Online. Association for Computational Linguistics.
  42. 42.Ruiqi Zhong, Tao Yu, and Dan Klein. 2020. Semantic evaluation for text-to-sql with distilled test suites. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 396–411.
  43. 43.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H. Chi. 2023. Least-to-most prompting enables complex reasoning in large language models. In The Eleventh International Conference on Learning Representations.

Citation

MLA
Tai, C.-Y., et al. “Exploring Chain of Thought Style Prompting for Text-to-SQL”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 5376–93, https://doi.org/10.18653/v1/2023.emnlp-main.327.
APA
Tai, C.-Y., Chen, Z., Zhang, T., Deng, X., & Sun, H. (2023). Exploring Chain of Thought Style Prompting for Text-to-SQL. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 5376–5393. https://doi.org/10.18653/v1/2023.emnlp-main.327
Chicago
Tai, C.-Y., Z. Chen, T. Zhang, X. Deng, and H. Sun. 2023. “Exploring Chain of Thought Style Prompting for Text-to-SQL”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 5376–93. https://doi.org/10.18653/v1/2023.emnlp-main.327.
Harvard
Tai, C.-Y. et al. (2023) “Exploring Chain of Thought Style Prompting for Text-to-SQL”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 5376–5393. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.327.
Vancouver
1. Tai C-Y, Chen Z, Zhang T, Deng X, Sun H (2023) Exploring Chain of Thought Style Prompting for Text-to-SQL. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 5376–5393

BibTeX

@inproceedings{tai-etal-2023-exploring,
    title = "Exploring Chain of Thought Style Prompting for Text-to-{SQL}",
    author = "Tai, Chang-Yu  and
      Chen, Ziru  and
      Zhang, Tianshu  and
      Deng, Xiang  and
      Sun, Huan",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.327/",
    doi = "10.18653/v1/2023.emnlp-main.327",
    pages = "5376--5393"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/