DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction

Mohammadreza PourrezaDavood Rafiei

article2023NeurIPS627 citations

Proposes a decomposed in-context learning method with self-correction that breaks text-to-SQL generation into modular sub-tasks, enabling prompting-based large language models to outperform heavily fine-tuned baselines on the Spider and BIRD benchmarks.

Listen

Enabling business users to query relational databases via natural language has long been a key operational goal, but existing large language models (LLMs) have historically lagged behind heavily fine-tuned, specialized systems on complex database benchmarks. While standard zero-shot and few-shot prompting techniques avoid the massive compute costs and training data requirements of fine-tuning, they often fail on intricate database queries involving complex joins, nested subqueries, and table schema alignments.

The article evaluates whether breaking down the natural language-to-SQL translation process into smaller sub-problems can bridge this performance gap. The authors propose and test DIN-SQL (Decomposed In-Context Learning of Text-to-SQL with Self-Correction), a modular framework that relies entirely on few-shot in-context learning without requiring model fine-tuning.

The framework consists of four sequential stages executed through prompting: schema linking (matching natural language references to database tables and columns), query classification and decomposition (categorizing queries as easy, non-nested, or nested complex and breaking down sub-problems), specialized SQL generation using intermediate representations like NatSQL, and zero-shot self-correction to repair minor syntax or keyword errors. The authors evaluated DIN-SQL across multiple LLMs—primarily GPT-4, CodeX Davinci, and CodeX Cushman—on two challenging cross-domain benchmarks, Spider and BIRD.

The evaluation yielded several key findings. First, task decomposition consistently improved standard few-shot execution accuracy across all tested language models by roughly 10%. Second, DIN-SQL paired with GPT-4 achieved a state-of-the-art execution accuracy of 85.3% on the hidden test set of the Spider benchmark, outperforming previous top fine-tuned models (79.9%) by over 5%. Third, on the complex BIRD benchmark, the approach set a new state-of-the-art execution accuracy of 55.9% on the holdout test set and improved execution efficiency scores by 9% over baseline GPT-4 models. Finally, ablation analysis revealed that query classification and schema linking provided the largest accuracy gains, particularly on hard and extra-hard query categories where baseline LLMs traditionally struggle.

These findings indicate that organizations can achieve state-of-the-art database interface performance without investing significant capital and engineering time into training and maintaining customized, fine-tuned models. Prompt-based decomposition substantially reduces implementation friction and domain-adaptation overhead. However, decision-makers must weigh performance against operational trade-offs: the multi-step reasoning pipeline introduces higher per-query inference costs (approximately $0.50 per query using GPT-4 at the time of the study) and longer response latency (around 60 seconds per query).

Organizations planning to deploy natural language database interfaces should adopt modular, decomposed prompting strategies rather than monolithic prompts when dealing with multi-table, complex relational databases. For production rollouts, teams should select self-correction strategies tailored to the underlying model: smaller models benefit from explicit bug-hunting prompts, whereas more capable models like GPT-4 perform best with gentle verification prompts. As a next step, organizations should pilot this workflow on internal schemas and explore automated demonstration selection to optimize latency and operational costs before enterprise-wide integration.

Confidence in these findings is high across standard academic benchmarks, though readers should note certain limitations. The demonstrations used in the study were manually constructed and fixed per query class, and schema ambiguity remains the leading source of residual errors. Real-world enterprise databases with extensive domain jargon or ambiguous table naming may require additional automated schema-linking safeguards or human-in-the-loop validation.

Cover for DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction

Abstract

There is currently a significant gap between the performance of fine-tuned models and prompting approaches using Large Language Models (LLMs) on the challenging task of text-to-SQL, as evaluated on datasets such as Spider. To improve the performance of LLMs in the reasoning process, we study how decomposing the task into smaller sub-tasks can be effective. In particular, we show that breaking down the generation problem into sub-problems and feeding the solutions of those sub-problems into LLMs can be an effective approach for significantly improving their performance. Our experiments with three LLMs show that this approach consistently improves their simple few-shot performance by roughly 10%, pushing the accuracy of LLMs towards SOTA or surpassing it. On the holdout test set of Spider, the SOTA, in terms of execution accuracy, was 79.9 and the new SOTA at the time of this writing using our approach is 85.3. Our approach with in-context learning beats many heavily fine-tuned models by at least 5%. Additionally, when evaluated on the BIRD benchmark, our approach achieved an execution accuracy of 55.9%, setting a new SOTA on its holdout test set.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Few-shot Error Analysis
  • 4 Methodology
  • 4.1 Schema Linking Module
  • 4.2 Classification & Decomposition Module
  • 4.3 SQL Generation Module
  • 4.4 Self-correction Module
  • 5 Experiments
  • 5.1 Models
  • 5.2 Hyperparameter
  • 5.3 Dataset
  • 5.4 Metrics
  • 5.5 Results
  • 5.5.1 Test set results
  • 5.5.2 Development set results
  • 5.5.3 Error improvements
  • 5.6 Ablation study
  • 6 Conclusions
  • 7 Limitations
  • References
  • A Prompts
  • A.1 Zero-shot prompting
  • A.2 Few-shot prompting
  • A.3 Schema linking prompt
  • A.4 Classification & decomposition prompt
  • A.5 SQL generation
  • A.5.1 Easy Class
  • A.5.2 Non-Nested Complex
  • A.5.3 Nested Complex
  • A.6 Self-correction prompts
  • A.6.1 Generic self-correction prompt
  • A.6.2 Gentle self-correction prompt

Knowls

  1. Knowl 1 — DIN-SQL: Decomposed In-Context Learning Framework for Text-to-SQL

    model/method

    DIN-SQL (Decomposed In-Context Learning of Text-to-SQL) is a multi-step prompting framework that breaks the declarative natural language to SQL generation task into smaller, structured sub-problems for Large Language Models (LLMs). Rather than generating complex SQL queries in a single zero-shot or standard few-shot prompt, DIN-SQL executes a four-module pipeline:

    1. Schema Linking Module: Identifies database schema elements (tables, columns, foreign keys) and literal condition values referenced in the natural language question.
    2. Query Classification and Decomposition Module: Categorizes the query based on structural complexity into one of three classes—Easy, Non-Nested Complex, or Nested Complex—and decomposes nested queries into intermediate sub-questions while specifying required table joins.
    3. SQL Generation Module: Applies class-specific prompt demonstrations. For complex queries, it bridges the gap between natural language and SQL syntax using an intermediate representation (NatSQL) and solves intermediate sub-queries before synthesizing the final SQL query.
    4. Self-Correction Module: Refines the generated SQL query in a zero-shot setting by prompting the LLM to inspect and correct subtle syntax and semantic errors (e.g., missing predicates, incorrect aggregations, or omitted DISTINCT/DESC keywords).

    All modules are implemented entirely through few-shot or zero-shot in-context prompting without fine-tuning model parameters.

  2. Knowl 2 — Error Categorization of Standard Few-Shot LLM Prompting in Text-to-SQL

    empirical result

    A manual error analysis was conducted on 500 randomly sampled failed queries from the Spider training set generated by CodeX Davinci under standard few-shot prompting. The execution failures were classified into six primary categories:

    • Schema Linking (37%): The model fails to identify the correct table names, column names, or condition entities mentioned in the natural language query, or confuses aggregate functions with existing columns sharing names like average.
    • JOIN (21%): The model fails to identify all required tables or selects incorrect foreign keys when joining tables.
    • GROUP BY (13%): The model either fails to recognize the requirement for grouping or selects incorrect columns in the GROUP BY clause.
    • Nested and Set Operations (13%): The model fails to recognize nested query structures or uses incorrect set operators (e.g., EXCEPT, UNION, INTERSECT, IN, NOT IN).
    • Invalid SQL (3%): The generated SQL contains syntactic errors that prevent execution.
    • Miscellaneous / Other (13%): Queries containing extraneous or missing predicates, incorrect DISTINCT or DESC keywords, missing WHERE clauses, or redundant aggregation functions.
  3. Knowl 3 — Schema Linking Module in DIN-SQL

    model/method

    The Schema Linking Module in DIN-SQL is designed to resolve schema entity references and extract query constants before SQL syntax generation. It uses a few-shot prompt containing 10 demonstrations drawn from the Spider benchmark training set, formatted with chain-of-thought prompting initiated by the phrase "Let's think step by step."

    Given a database schema definition, foreign key relations, and a natural language question QQ, the module outputs:

    1. The relevant column names mapped to their respective tables (e.g., table.column).
    2. The necessary foreign key relations required to connect the referenced tables (e.g., table1.fk = table2.pk).
    3. The set of extracted literal constants and cell values (e.g., strings, numbers) appearing in the question.
    4. A compiled summary of schema links denoted as S=[columns,foreign keys,literals]S = [\text{columns}, \text{foreign keys}, \text{literals}].

    This extracted schema link list SS is directly forwarded as an input conditioning context to all downstream DIN-SQL modules.

  4. Knowl 4 — Query Classification and Decomposition Module in DIN-SQL

    model/method

    The Query Classification and Decomposition Module categorizes natural language questions according to the structural complexity of their target SQL queries and breaks down complex hierarchical queries into sub-problems.

    Queries are classified into three disjoint complexity classes:

    • EASY: Single-table queries that require no JOIN operations and no sub-queries.
    • NON-NESTED COMPLEX: Queries that require joining multiple tables via JOIN clauses but do not involve sub-queries or set operations.
    • NESTED COMPLEX: Queries that require nested sub-queries, set operations (INTERSECT, UNION, EXCEPT, IN, NOT IN), and potentially multi-table joins.

    For queries classified as NESTED COMPLEX, the module also generates explicit intermediate natural language sub-questions Q1,Q2,…,QkQ_1, Q_2, \dots, Q_k corresponding to the procedural components or uncorrelated sub-queries needed to construct the complete query. It also outputs the explicit set of database tables required for the joins.

  5. Knowl 5 — Class-Specific SQL Generation with NatSQL Intermediate Representation

    model/method

    To prevent performance degradation caused by overly complex chain-of-thought prompts on simple queries while providing sufficient reasoning steps for complex queries, DIN-SQL employs three distinct prompt formats tailored to the predicted query class:

    1. Easy Class Prompt Format: ⟨Qj,Sj,Aj⟩\langle Q_j, S_j, A_j \rangle where QjQ_j is the natural language question, SjS_j is the schema link representation, and AjA_j is the target SQL statement.

    2. Non-Nested Complex Class Prompt Format: ⟨Qj,Sj,Ij,Aj⟩\langle Q_j, S_j, I_j, A_j \rangle where IjI_j represents an intermediate query in NatSQL grammar. NatSQL removes explicit JOIN ON, FROM, and GROUP BY operators and merges HAVING and WHERE clauses, bridging the structural mismatch between natural language and SQL.

    3. Nested Complex Class Prompt Format: ⟨Qj,Sj,⟨Qj1,Aj1,…,Qjk,Ajk⟩,Ij,Aj⟩\langle Q_j, S_j, \langle Q_{j1}, A_{j1}, \dots, Q_{jk}, A_{jk} \rangle, I_j, A_j \rangle where QjiQ_{ji} and AjiA_{ji} are the ii-th sub-question and its corresponding intermediate SQL sub-query solution (for i∈{1,…,k}i \in \{1, \dots, k\}), followed by the intermediate NatSQL representation IjI_j and the final synthesized SQL statement AjA_j.

  6. Knowl 6 — Self-Correction Strategies for LLM-Generated SQL

    model/method

    The Self-Correction Module in DIN-SQL repairs minor syntax and semantic errors in candidate SQL queries (such as omitted DISTINCT, missing DESC, invalid grouping, or extraneous predicates) under a zero-shot prompting setup without executing the query on a database engine.

    DIN-SQL specifies two distinct self-correction prompt variants based on model capability:

    • Generic Self-Correction: Explicitly prompts the LLM to identify and fix errors in a given "BUGGY SQL" string. This approach is effective for smaller LLMs (such as CodeX Davinci), where generation bugs are frequent.
    • Gentle Self-Correction: Does not assume the query contains a bug; instead, it presents the generated SQL alongside a checklist of instructions and hints (e.g., verifying SELECT columns, checking JOIN foreign keys, checking GROUP BY necessity, and ensuring DESC/DISTINCT accuracy). This prevents more capable models (such as GPT-4) from hallucinating bugs in already correct queries.

    By default, DIN-SQL uses the gentle prompt for GPT-4 and the generic prompt for CodeX models.

  7. Knowl 7 — DIN-SQL Performance on the Spider Benchmark

    data/table

    DIN-SQL was evaluated on the cross-domain Spider benchmark, which comprises 10,181 natural language questions and 5,693 unique SQL queries across 200 databases. Performance was measured using Execution Accuracy (EX) and Exact Set Match Accuracy (EM).

    On the Spider holdout test set (2,147 examples across 34 unseen databases), DIN-SQL with GPT-4 achieved state-of-the-art execution accuracy among published methods without utilizing database content:

    Model EX (%) EM (%)
    DIN-SQL + GPT-4 (Ours) 85.3 60.0
    RESDSQL-3B + NatSQL (DB content used) 79.9 72.0
    DIN-SQL + CodeX Davinci (Ours) 78.2 57.0
    Graphix-3B + PICARD (DB content used) 77.6 74.0
    SHiP + PICARD (DB content used) 76.6 73.1
    N-best Rerankers + PICARD (DB content used) 75.9 72.2
    RASAT + PICARD (DB content used) 75.5 70.9
    T5-3B + PICARD (DB content used) 75.1 71.9
    RATSQL + GAP + NatSQL (DB content used) 73.3 68.7
    RYANSQL v2 + BERT - 60.6
    SmBoP + BART - 60.5

    On the Spider development set (1,034 examples), DIN-SQL achieved consistent improvements over zero-shot and few-shot baselines across model sizes:

    Prompting Method Model EX (%) EM (%)
    DIN-SQL GPT-4 74.2 60.1
    DIN-SQL CodeX Davinci 69.9 57.2
    DIN-SQL CodeX Cushman 47.6 35.7
    Few-shot baseline GPT-4 67.4 54.3
    Few-shot baseline CodeX Davinci 61.5 50.2
    Few-shot baseline CodeX Cushman 43.1 30.9
    Zero-shot baseline GPT-4 64.9 40.4
    Zero-shot (Liu et al., 2023a) ChatGPT 60.1 -
    Zero-shot (Rajkumar et al., 2022) CodeX Davinci 47.5 -
    Zero-shot w/ DB content CodeX Davinci 55.1 -
    Zero-shot w/ DB content CodeX Cushman 53.0 -
    Zero-shot w/ DB content GPT-3 21.7 -

    Across query difficulty tiers on the development set, DIN-SQL + GPT-4 achieved EX scores of 91.1% (Easy), 79.8% (Medium), 64.9% (Hard), and 43.4% (Extra Hard), compared to standard few-shot GPT-4 EX scores of 86.7%, 73.1%, 59.2%, and 31.9%.

  8. Knowl 8 — DIN-SQL Performance on the BIRD Benchmark

    data/table

    The BIRD benchmark contains 12,751 question-SQL pairs over 95 large-scale databases (33.4 GB) spanning 37 domains, incorporating external knowledge hints (numeric reasoning, domain knowledge, synonyms, value illustration) and sample table rows. Models are evaluated on Execution Accuracy (EX) and Valid Efficiency Score (VES), which penalizes inefficient execution time among correct queries.

    On the BIRD holdout test set:

    Model VES (%) EX (%)
    DIN-SQL + GPT-4 (Ours) 59.44 55.90
    GPT-4 baseline 60.77 54.89
    Claude-2 - 49.02
    ChatGPT + CoT 56.56 40.08
    ChatGPT 51.40 39.30
    CodeX 41.60 36.47
    PaLM-2 - 33.04
    T5-3B 27.80 24.05
    T5-Large 25.00 20.94
    T5-Base 14.70 12.89

    On the BIRD development set:

    Model VES (%) EX (%)
    DIN-SQL + GPT-4 (Ours) 58.79 50.72
    GPT-4 baseline 49.77 46.35
    Claude-2 - 42.70
    ChatGPT + CoT 42.30 36.64
    ChatGPT 43.81 37.22
    CodeX 43.41 34.35
    PaLM-2 - 27.38
    T5-3B 25.57 23.34
    T5-Large 22.74 19.75
    T5-Base 12.90 11.54

    On the development set, DIN-SQL + GPT-4 established a new state of the art, outperforming the baseline GPT-4 model by 4.37% in EX and by 9.02% in VES.

  9. Knowl 9 — Ablation Study of DIN-SQL Modules

    data/table

    An ablation study evaluated the contribution of individual DIN-SQL components (Schema Linking, Query Classification, Self-Correction, and Self-Correction prompting style) across difficulty tiers on the Spider development set using Execution Accuracy (EX):

    Prompting Configuration Model Easy Medium Hard Extra All
    DIN-SQL (generic self-corr) CodeX Davinci 89.1 75.6 58.0 38.6 69.9
    DIN-SQL (gentle self-corr) CodeX Davinci 87.5 76.9 51.7 36.1 68.7
    DIN-SQL w/o self-corr CodeX Davinci 83.9 75.4 52.3 36.1 67.3
    DIN-SQL w/o schema linking CodeX Davinci 87.3 70.6 57.6 27.1 65.9
    DIN-SQL w/o classification (simple few-shot) CodeX Davinci 87.9 68.2 51.7 27.1 63.1
    DIN-SQL w/o classification (decomposed COT) CodeX Davinci 84.2 71.2 54.3 38.6 68.2
    DIN-SQL (gentle self-corr) GPT-4 91.1 79.8 64.9 43.4 74.2
    DIN-SQL (generic self-corr) GPT-4 89.9 76.5 59.2 34.3 70.0
    DIN-SQL w/o self-corr GPT-4 91.1 79.1 63.2 41.6 73.3

    Key takeaways from the ablation data include:

    1. Removing schema linking causes the largest drop in overall EX for CodeX Davinci (from 69.9% to 65.9%), particularly on Extra Hard queries (from 38.6% to 27.1%).
    2. Applying decomposed Chain-of-Thought (COT) to all queries without classification degrades Easy query performance (84.2% vs. 89.1%), whereas using simple few-shot prompting across all queries severely harms Hard (51.7%) and Extra Hard (27.1%) performance.
    3. Generic self-correction benefits CodeX Davinci (+2.6% over no correction) but hurts GPT-4 (-3.3% relative to gentle correction), whereas gentle self-correction improves GPT-4 from 73.3% to 74.2%.
  10. Knowl 10 — Computational Cost and Latency Limitations of DIN-SQL

    limitation

    The multi-step, decomposed architecture of DIN-SQL introduces notable computational overhead and execution latency compared to single-pass prompting approaches:

    • Inference Latency: When evaluated on a natural language query from the Spider dataset using GPT-4 via the OpenAI API, the sequential four-stage prompting pipeline incurs an average latency of approximately 60 seconds per query.
    • Monetary Cost: Processing a single question through all intermediate LLM calls (schema linking, classification/decomposition, class-tailored SQL generation with NatSQL, and zero-shot self-correction) costs approximately $0.50 USD per query when using GPT-4.
    • Static In-Context Demonstrations: The demonstration examples used in the prompt templates are manually constructed and fixed per complexity class rather than dynamically retrieved or adapted to specific target database domains at a finer granularity.

Coverage note — None was omitted. All primary contributions—including the four-stage DIN-SQL decomposition methodology, the few-shot error analysis, class-specific prompting formats, self-correction prompt design, comprehensive benchmark evaluations on Spider and BIRD, module ablation studies, and operational limitations—are fully captured as standalone knowls.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  2. 2.Ruichu Cai, Jinjie Yuan, Boyan Xu, and Zhifeng Hao. Sadga: Structure-aware dual graph aggregation network for text-to-sql. Advances in Neural Information Processing Systems, 34:7664–7676, 2021.
  3. 3.Ruisheng Cao, Lu Chen, Zhi Chen, Yanbin Zhao, Su Zhu, and Kai Yu. Lgesql: line graph enhanced text-to-sql model with mixed local and non-local relations. arXiv preprint arXiv:2106.01093, 2021.
  4. 4.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  5. 5.Wenhu Chen. Large language models are few (1)-shot table reasoners. arXiv preprint arXiv:2210.06710, 2022.
  6. 6.DongHyun Choi, Myeong Cheol Shin, EungGyun Kim, and Dong Ryeol Shin. Ryansql: Recursively applying sketch-based slot fillings for complex text-to-sql in cross-domain databases. Computational Linguistics, 47(2):309–332, 2021.
  7. 7.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  9. 9.Li Dong and Mirella Lapata. Language to logical form with neural attention. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 33–43, 2016.
  10. 10.Yujian Gan, Xinyun Chen, Jinxia Xie, Matthew Purver, John R Woodward, John Drake, and Qiaofu Zhang. Natural sql: Making sql easier to infer from natural language specifications. arXiv preprint arXiv:2109.05153, 2021.
  11. 11.Alex Graves and Alex Graves. Long short-term memory. Supervised sequence labelling with recurrent neural networks, pages 37–45, 2012.
  12. 12.Jiaqi Guo, Zecheng Zhan, Yan Gao, Yan Xiao, Jian-Guang Lou, Ting Liu, and Dongmei Zhang. Towards complex text-to-sql in cross-domain database with intermediate representation. arXiv preprint arXiv:1905.08205, 2019.
  13. 13.Zhixin Guo, Minyxuan Yan, Jiexing Qi, Jianping Zhou, Ziwei He, Zhouhan Lin, Guanjie Zheng, and Xinbing Wang. Few-shot table-to-text generation with prompt planning and knowledge memorization. arXiv preprint arXiv:2302.04415, 2023.
  14. 14.Jonathan Herzig, Paweł Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos. Tapas: Weakly supervised table parsing via pre-training. arXiv preprint arXiv:2004.02349, 2020.
  15. 15.Junyang Huang, Yongbo Wang, Yongliang Wang, Yang Dong, and Yanghua Xiao. Relation aware semi-autoregressive semantic parsing for nl2sql. arXiv preprint arXiv:2108.00804, 2021.
  16. 16.Binyuan Hui, Xiang Shi, Ruiying Geng, Binhua Li, Yongbin Li, Jian Sun, and Xiaodan Zhu. Improving text-to-sql with schema dependency learning. arXiv preprint arXiv:2103.04399, 2021.
  17. 17.Wonseok Hwang, Jinyeong Yim, Seunghyun Park, and Minjoon Seo. A comprehensive exploration on wikisql with table-aware word contextualization. arXiv preprint arXiv:1902.01069, 2019.
  18. 18.Rohit Kate. Transforming meaning representation grammars to improve semantic parsing. In CoNLL 2008: Proceedings of the Twelfth Conference on Computational Natural Language Learning, pages 33–40, 2008.
  19. 19.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. Decomposed prompting: A modular approach for solving complex tasks. arXiv preprint arXiv:2210.02406, 2022.
  20. 20.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916, 2022.
  21. 21.Brenden Lake and Marco Baroni. Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In International conference on machine learning, pages 2873–2882. PMLR, 2018.
  22. 22.Wenqiang Lei, Weixin Wang, Zhixin Ma, Tian Gan, Wei Lu, Min-Yen Kan, and Tat-Seng Chua. Re-examining the role of schema linking in text-to-sql. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6943–6954, 2020.
  23. 23.Fei Li and Hosagrahar V Jagadish. Constructing an interactive natural language interface for relational databases. Proceedings of the VLDB Endowment, 8(1):73–84, 2014.
  24. 24.Haoyang Li, Jing Zhang, Cuiping Li, and Hong Chen. Decoupling the skeleton parsing and schema linking for text-to-sql. arXiv preprint arXiv:2302.05965, 2023a.
  25. 25.Jinyang Li, Binyuan Hui, Reynold Cheng, Bowen Qin, Chenhao Ma, Nan Huo, Fei Huang, Wenyu Du, Luo Si, and Yongbin Li. Graphix-t5: Mixing pre-trained transformers with graph-aware layers for text-to-sql parsing. arXiv preprint arXiv:2301.07507, 2023b.
  26. 26.Jinyang Li, Binyuan Hui, Ge Qu, Binhua Li, Jiaxi Yang, Bowen Li, Bailin Wang, Bowen Qin, Rongyu Cao, Ruiying Geng, Nan Huo, Chenhao Ma, Kevin C. C. Chang, Fei Huang, Reynold Cheng, and Yongbin Li. Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls, 2023c.
  27. 27.Yunyao Li, Huahai Yang, and HV Jagadish. Nalix: A generic natural language search environment for xml data. ACM Transactions on database systems (TODS), 32(4):30–es, 2007.
  28. 28.Aiwei Liu, Xuming Hu, Lijie Wen, and Philip S Yu. A comprehensive evaluation of chatgpt’s zero-shot text-to-sql capability. arXiv preprint arXiv:2303.13547, 2023a.
  29. 29.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1–35, 2023b.
  30. 30.Ana-Maria Popescu, Oren Etzioni, and Henry Kautz. Towards a theory of natural language interfaces to databases. In Proceedings of the 8th international conference on Intelligent user interfaces, pages 149–157, 2003.
  31. 31.Ana-Maria Popescu, Alex Armanasu, Oren Etzioni, David Ko, and Alexander Yates. Modern natural language interfaces to databases: Composing statistical parsing with semantic tractability. In COLING 2004: Proceedings of the 20th International Conference on Computational Linguistics, pages 141–147, 2004.
  32. 32.Jiexing Qi, Jingyao Tang, Ziwei He, Xiangpeng Wan, Chenghu Zhou, Xinbing Wang, Quanshi Zhang, and Zhouhan Lin. Rasat: Integrating relational structures into pretrained seq2seq model for text-to-sql. arXiv preprint arXiv:2205.06983, 2022.
  33. 33.Bowen Qin, Binyuan Hui, Lihan Wang, Min Yang, Jinyang Li, Binhua Li, Ruiying Geng, Rongyu Cao, Jian Sun, Luo Si, et al. A survey on text-to-sql parsing: Concepts, methods, and future directions. arXiv preprint arXiv:2208.13629, 2022.
  34. 34.Nitarshan Rajkumar, Raymond Li, and Dzmitry Bahdanau. Evaluating the text-to-sql capabilities of large language models. arXiv preprint arXiv:2204.00498, 2022.
  35. 35.Ohad Rubin and Jonathan Berant. Smbop: Semi-autoregressive bottom-up semantic parsing. arXiv preprint arXiv:2010.12412, 2020.
  36. 36.Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. Picard: Parsing incrementally for constrained auto-regressive decoding from language models. arXiv preprint arXiv:2109.05093, 2021.
  37. 37.Niculae Stratica, Leila Kosseim, and Bipin C Desai. Using semantic templates for a natural language interface to the cindi virtual library. Data & Knowledge Engineering, 55(1):4–19, 2005.
  38. 38.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with neural networks. Advances in neural information processing systems, 27, 2014.
  39. 39.Bailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov, and Matthew Richardson. Ratsql: Relation-aware schema encoding and linking for text-to-sql parsers. arXiv preprint arXiv:1911.04942, 2019.
  40. 40.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  41. 41.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022a.
  42. 42.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022b.
  43. 43.Xiaojun Xu, Chang Liu, and Dawn Song. Sqlnet: Generating structured queries from natural language without reinforcement learning. arXiv preprint arXiv:1711.04436, 2017.
  44. 44.Kuan Xuan, Yongbo Wang, Yongliang Wang, Zujie Wen, and Yang Dong. Sead: end-to-end text-to-sql generation with schema-aware denoising. arXiv preprint arXiv:2105.07911, 2021.
  45. 45.Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. Tabert: Pretraining for joint understanding of textual and tabular data. arXiv preprint arXiv:2005.08314, 2020.
  46. 46.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. arXiv preprint arXiv:1809.08887, 2018.
  47. 47.Tao Yu, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang, Yi Chern Tan, Xinyi Yang, Dragomir Radev, Richard Socher, and Caiming Xiong. Grappa: Grammar-augmented pre-training for table semantic parsing. arXiv preprint arXiv:2009.13845, 2020.
  48. 48.Lu Zeng, Sree Hari Krishnan Parthasarathi, and Dilek Hakkani-Tur. N-best hypotheses reranking for text-to-sql systems. arXiv preprint arXiv:2210.10668, 2022.
  49. 49.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493, 2022.
  50. 50.Yiyun Zhao, Jiarong Jiang, Yiqun Hu, Wuwei Lan, Henry Zhu, Anuj Chauhan, Alexander Li, Lin Pan, Jun Wang, Chung-Wei Hang, et al. Importance of synthesizing high-quality data for text-to-sql parsing. arXiv preprint arXiv:2212.08785, 2022.
  51. 51.Ruiqi Zhong, Tao Yu, and Dan Klein. Semantic evaluation for text-to-sql with distilled test suite. In The 2020 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2020.
  52. 52.Victor Zhong, Caiming Xiong, and Richard Socher. Seq2sql: Generating structured queries from natural language using reinforcement learning. arXiv preprint arXiv:1709.00103, 2017.
  53. 53.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625, 2022.

Citation

MLA
Pourreza, M., and D. Rafiei. “DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction”. arXiv, 2023, http://arxiv.org/abs/2304.11015v3.
APA
Pourreza, M., & Rafiei, D. (2023). DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction. arXiv. http://arxiv.org/abs/2304.11015v3
Chicago
Pourreza, M., and D. Rafiei. 2023. “DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction”. arXiv. http://arxiv.org/abs/2304.11015v3.
Harvard
Pourreza, M. and Rafiei, D. (2023) “DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.11015v3.
Vancouver
1. Pourreza M, Rafiei D (2023) DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction. arXiv

BibTeX

@article{pourreza2023din,
  title = {DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction},
  author = {Pourreza, Mohammadreza and Rafiei, Davood},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.11015v3},
  eprint = {2304.11015}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors