Exploring Chain of Thought Style Prompting for Text-to-SQL
Chang-Yu TaiZiru ChenTianshu ZhangXiang DengHuan Sun
Presents a single-pass question-decomposition prompting method for text-to-SQL parsing that reduces error propagation and outperforms multi-turn iterative prompting on the Spider benchmarks without requiring expensive intermediate execution calls.
Converting everyday human language into database queries—known as text-to-SQL parsing—is essential for building intuitive AI database assistants. Traditional approaches require large, expensive labeled datasets to train specialized models. While large language models can perform this task using only a few examples in a prompt, their accuracy remains limited when dealing with complex, multi-step relational reasoning.
The article systematically evaluates how chain-of-thought prompting strategies can improve the reasoning capabilities of large language models for text-to-SQL parsing. Specifically, it demonstrates how decomposing questions into simpler steps in a single prompt outperforms both standard prompting and iterative problem-solving techniques.
To conduct this evaluation, the researchers adapted two existing prompting paradigms: standard chain-of-thought, which mimics SQL query execution steps, and least-to-most prompting, which iteratively generates and solves sub-questions across multiple stages. Identifying critical shortcomings in these methods, they designed a novel approach called Question Decomposition (QDecomp) and an enhanced variant, QDecomp+InterCOL, which incorporates relevant table and column names into each decomposed step. The evaluation was carried out primarily using OpenAI's Codex model on cross-domain benchmarks, including Spider and Spider Realistic, alongside three single-domain datasets (GeoQuery, IMDB, and Yelp), measuring performance with strict test-suite execution accuracy.
The findings establish that the proposed QDecomp+InterCOL approach delivers superior performance. On the Spider development and Spider Realistic benchmarks, QDecomp+InterCOL achieved test-suite accuracies of 68.4% and 56.5%, representing absolute improvements of 5.2 and 6.5 percentage points over standard prompting without reasoning. Furthermore, it outperformed least-to-most prompting by 2.4 and 1.5 percentage points. By contrast, traditional chain-of-thought prompting fell behind standard baseline prompting (56.8% versus 63.2%), as generating granular SQL execution steps introduced severe reasoning errors that derailed the final query. The experiments also revealed that concise API-style schema descriptions performed on par with lengthy database dumps while consuming far fewer tokens.
These results demonstrate that multi-step iterative prompting is unnecessary and computationally inefficient for text-to-SQL tasks. Iterative pipelines and overly detailed reasoning steps amplify error propagation, where an early mistake compounds and ruins the final database query. By prompting the model to decompose questions and predict intermediate database schema elements in a single generation pass, organizations can achieve higher query accuracy with lower computational overhead and reduced latency.
Organizations developing database agents should adopt single-pass question decomposition prompts enriched with intermediate schema targets rather than complex iterative prompting frameworks. For production systems, teams can integrate this decomposed natural-language output into interactive semantic parsing interfaces, allowing end users to review and correct intermediate sub-questions before final query execution.
Readers should note that the empirical results rely primarily on OpenAI's Codex model and a fixed set of academic benchmarks. Future evaluations should validate these single-pass decomposition strategies on newer models, such as GPT-4, and assess resilience against real-world database perturbations before large-scale enterprise deployment.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Wei et al. introduce the few-shot chain-of-thought method that this paper tests and adapts for text-to-SQL.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Zhou et al. establish least-to-most prompting, a key comparison method whose iterative decomposition the source evaluates for text-to-SQL.
No sufficiently relevant recommendations were found.
