Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs

Jinyang LiBinyuan HuiGe QuJiaxi YangBinhua LiBowen LiBailin WangBowen QinRuiying GengNan Huo

article2023NeurIPS841 citations

Introduces BIRD, a large-scale text-to-SQL benchmark across 95 multi-domain databases that reveals substantial performance gaps in large language models by evaluating real-world challenges such as dirty database values, external knowledge reasoning, and query execution efficiency.

Listen

Organizations increasingly seek to enable non-technical users to query massive corporate databases using natural language. While recent advances in artificial intelligence have shown high accuracy on academic benchmarks, existing evaluations focus predominantly on clean database schemas rather than the messy, large-scale data environments encountered in real-world business operations. This leaves significant uncertainty about whether leading automated systems can reliably serve as standalone database interfaces without risking costly errors or processing inefficiencies.

To bridge this gap, the article introduces BIRD, a large-scale evaluation benchmark grounded in massive, realistic databases across 37 diverse domains. The authors developed a dataset comprising 12,751 question-query pairs spanning 95 relational databases totaling 33.4 GB. Built using rigorous double-blind crowdsourcing and expert oversight, the benchmark introduces two critical dimensions: the necessity of external business knowledge to interpret ambiguous values, and a dedicated Valid Efficiency Score to assess the runtime efficiency of generated queries alongside their execution accuracy.

The findings show that current automated systems struggle significantly when confronted with realistic data environments. The top-performing system, GPT-4, achieved only a 54.89% execution accuracy on the hidden test set even when supplied with necessary domain knowledge, falling far short of the 92.96% achieved by human experts. Smaller fine-tuned language models scored under 25%. Furthermore, when external domain knowledge was omitted, automated accuracy plunged by 10 to 20 percentage points across all systems. An in-depth error analysis revealed that over 82% of model failures stemmed from misidentifying table structures or misunderstanding the actual underlying database contents. On the operational front, optimizing query structures through two-stage rewriting reduced query execution time by 77.75%, while adding targeted indexes reduced execution time by 87.27%.

These results demonstrate that contemporary language models are not yet dependable as fully autonomous interfaces for enterprise databases. Deploying them without oversight introduces substantial business risks, including hallucinated schema items, inaccurate data retrievals, and expensive, unoptimized queries that can strain computational resources. The significant performance drop in the absence of external context underscores that raw language comprehension is insufficient; reliable enterprise deployment fundamentally requires integrating structured business logic and automated query optimization.

Stakeholders are advised against deploying fully autonomous natural language database interfaces for high-stakes operational decisions at this stage. Instead, organizations should pursue human-in-the-loop workflows where automated queries assist rather than replace technical analysts. Technical roadmaps should prioritize developing structured knowledge-grounding frameworks, incorporating query-optimization layers before execution, and establishing standardized indexing to prevent performance bottlenecks. Although the evaluation is currently constrained to SQLite implementations, confidence in the findings remains high due to the strict double-blind methodology and the vast scale of tested domains.

Cover for Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs

Abstract

Text-to-SQL parsing, which aims at converting natural language instructions into executable SQLs, has gained increasing attention in recent years. In particular, Codex and ChatGPT have shown impressive results in this task. However, most of the prevalent benchmarks, i.e., Spider, and WikiSQL, focus on database schema with few rows of database contents leaving the gap between academic study and real-world applications. To mitigate this gap, we present Bird, a big benchmark for large-scale database grounded in text-to-SQL tasks, containing 12,751 pairs of text-to-SQL data and 95 databases with a total size of 33.4 GB, spanning 37 professional domains. Our emphasis on database values highlights the new challenges of dirty database contents, external knowledge between NL questions and database contents, and SQL efficiency, particularly in the context of massive databases. To solve these problems, text-to-SQL models must feature database value comprehension in addition to semantic parsing. The experimental results demonstrate the significance of database values in generating accurate text-to-SQLs for big databases. Furthermore, even the most effective text-to-SQL models, i.e. ChatGPT, only achieves 40.08% in execution accuracy, which is still far from the human result of 92.96%, proving that challenges still stand. Besides, we also provide an efficiency analysis to offer insights into generating text-to-efficient-SQLs that are beneficial to industries. We believe that BIRD will contribute to advancing real-world applications of text-to-SQL research. The leaderboard and source code are available: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Task Formulation & Annotations
  • 3 Dataset Construction
  • 3.1 Annotation Entrance
  • 3.2 Database Source
  • 3.3 Question Annotation
  • 3.4 SQL Annotation
  • 4 Data Statistics
  • 5 Evaluation Metrics
  • 6 Experiments
  • 6.1 Baseline Models
  • 6.2 Execution Accuracy Analysis
  • 6.3 Baseline Performance on Spider
  • 6.4 Efficiency Analysis
  • 6.5 Knowledge Evidence Analysis
  • 6.6 More Analysis
  • 7 Related Work
  • 8 Limitation and Future work
  • 9 Conclusion
  • References
  • A Datasheet for Datasets
  • A.1 Motivation
  • A.2 Composition
  • A.3 Collection Process
  • A.4 Preprocessing/cleaning/labeling
  • A.5 Uses
  • A.6 Distribution
  • A.7 Maintenance
  • B Appendix
  • B.1 Text-to-SQL Difficulty
  • B.2 Annotation Entrance
  • B.3 Question Distribution
  • B.4 Experiment Details
  • B.5 Efficiency Analysis Details
  • B.6 Error Analysis Details
  • B.7 Evaluation Details
  • B.8 VES Details
  • B.9 Human Performance Collection
  • B.10 Distribution of Open-source Databases
  • B.11 SQL Function Taxonomy
  • B.12 Keyword Statistic
  • B.13 Study about Text-to-SQL Models

Knowls

  1. Knowl 1 — The BIRD Benchmark for Large-Scale Database-Grounded Text-to-SQL

    definition

    BIRD (BIg Bench for LaRge-Scale Database Grounded Text-to-SQLs) is a cross-domain text-to-SQL benchmark designed to evaluate semantic parsing and database value comprehension on massive, dirty, and realistic relational databases.

    The benchmark comprises:

    • 12,751 text-to-SQL query pairs.
    • 95 databases spanning 37 professional domains (e.g., financial, healthcare, blockchain, sports, education) with a total volume of 33.4 GB in SQLite.
    • An average of 7.3 tables per database and 549,000 rows per database.
    • Database sources: 32% from Kaggle competition datasets, 48% from the CTU Prague Relational Learning Repository, and 20% synthesized and standardized open tabular datasets.
    • Data splits: 9,428 examples across 69 databases for training, 1,534 examples across 11 databases for development, and 1,789 examples across 15 newly designed databases for a concealed test set to prevent pre-training contamination.

    Each database is accompanied by a database description file documenting expanded table and column names, data types, and value descriptions to enable models to reason over abbreviated schemas and noisy value distributions.

  2. Knowl 2 — Valid Efficiency Score (VES) Metric for SQL Execution Efficiency

    equation

    The Valid Efficiency Score (VES) evaluates text-to-SQL parsers by rewarding execution efficiency only for predicted queries that produce semantically correct results.

    Formally, given NN evaluation instances, ground-truth SQL queries YnY_n, and predicted SQL queries Y^n\hat{Y}_n:

    VES=1N∑n=1N1(Vn,V^n)⋅R(Yn,Y^n)VES = \frac{1}{N} \sum_{n=1}^{N} \mathbb{1}(V_n, \hat{V}_n) \cdot R(Y_n, \hat{Y}_n)

    where 1(Vn,V^n)\mathbb{1}(V_n, \hat{V}_n) is the correctness indicator:

    1(V,V^)={1,if V=V^0,if V≠V^\mathbb{1}(V, \hat{V}) = \begin{cases} 1, & \text{if } V = \hat{V} \\ 0, & \text{if } V \neq \hat{V} \end{cases}

    Here, VnV_n and V^n\hat{V}_n represent the execution result sets of YnY_n and Y^n\hat{Y}_n, respectively. Set equality is evaluated using HashSet to disregard row order unless explicitly constrained. R(Yn,Y^n)R(Y_n, \hat{Y}_n) denotes the relative execution efficiency ratio:

    R(Yn,Y^n)=E(Yn)E(Y^n)R(Y_n, \hat{Y}_n) = \sqrt{\frac{E(Y_n)}{E(\hat{Y}_n)}}

    where E(Y)∈(ϵ,30 s)E(Y) \in (\epsilon, 30\,\text{s}) measures the execution running time of query YY, where ϵ>0\epsilon > 0 prevents division by zero. Execution time E(Y)E(Y) is computed as the average over 100 executions on identical CPU hardware after discarding statistical outliers outside [μ−3σ,μ+3σ][\mu - 3\sigma, \mu + 3\sigma], where μ\mu is the mean execution time and σ\sigma is the standard deviation. A higher R(Yn,Y^n)R(Y_n, \hat{Y}_n) indicates that the predicted query executes faster than the human-annotated ground-truth query. Queries with incorrect results (1=0\mathbb{1} = 0) receive an efficiency score of 0.

  3. Knowl 3 — Knowledge-Grounded Text-to-SQL Formulation and Knowledge Taxonomy

    definition

    Knowledge-grounded text-to-SQL maps a natural language question QQ, a relational database schema D=⟨C,T⟩D = \langle C, T \rangle with columns CC and tables TT, and external knowledge evidence KK to an executable SQL query YY:

    Y=f(Q,D,K∣θ)Y = f(Q, D, K \mid \theta)

    where f(⋅∣θ)f(\cdot \mid \theta) is a neural semantic parser parameterized by θ\theta.

    The external knowledge evidence KK provides domain context required to interpret vague question phrases and map them to database values. In the BIRD benchmark, KK is classified into four categories:

    1. Numeric Reasoning Knowledge (24.5% of questions): Arithmetic operations (addition, subtraction, division, multiplication) and compositional formulas (e.g., calculating percentages, win rates, or financial ratios).
    2. Domain Knowledge (23.6% of questions): Industry-specific business logic, rules, and indicators (e.g., medical laboratory threshold values or loan eligibility criteria).
    3. Synonym Knowledge (7.2% of questions): Semantic alignments between natural language words and specific schema or data tokens (e.g., aligning 'female' with 'F').
    4. Value Illustration (70.1% of questions): Explicit mappings linking question entities to column values, abbreviations, and composite filtering conditions (e.g., specifying that the basketball position 'center' corresponds to column condition pos = 'C').
  4. Knowl 4 — Multi-Dimensional Difficulty Rating Scheme for Text-to-SQL

    definition

    The BIRD benchmark classifies instances into difficulty tiers by assessing four distinct dimensions on a 1–3 discrete scale:

    1. Question Understanding (Scale 1–3): Assesses semantic ambiguity in user intent (1=straightforward1 = \text{straightforward}, 2=clear but requires thought2 = \text{clear but requires thought}, 3=highly ambiguous3 = \text{highly ambiguous}).
    2. Knowledge Reasoning (Scale 1–3): Assesses the depth of external domain or mathematical knowledge required (1=no external knowledge1 = \text{no external knowledge}, 2=straightforward external hint2 = \text{straightforward external hint}, 3=extensive domain reasoning3 = \text{extensive domain reasoning}).
    3. Data Complexity (Scale 1–3): Assesses relational structure, schema size, and value noise (1=simple schema and data1 = \text{simple schema and data}, 2=complex schema resolved via descriptions2 = \text{complex schema resolved via descriptions}, 3=highly complex schema and values3 = \text{highly complex schema and values}).
    4. SQL Complexity (Scale 1–3): Assesses target SQL syntax and clause depth (1=simple query without multiple keywords1 = \text{simple query without multiple keywords}, 2=moderate complexity2 = \text{moderate complexity}, 3=complex query with numerous joins, subqueries, and window functions3 = \text{complex query with numerous joins, subqueries, and window functions}).

    Each dimension contributes equally to an aggregated score, categorizing the dataset into three balanced partitions: Simple (30%), Moderate (60%), and Challenging (10%).

  5. Knowl 5 — Double-Blind Quality-Controlled Text-to-SQL Annotation Pipeline

    algorithm

    The data construction methodology utilizes a four-step double-blind annotation and verification workflow to ensure high-fidelity ground-truth SQL queries and eliminate annotation bias.

    Input: Relational database DD, database description documentation MM, candidate annotators AQ,ASA_Q, A_S
    Output: Benchmark dataset $\mathcal{D} = \{(Q_i, D, K_i, Y_i)\}
    Screen candidate question crowdworkers AQA_Q on domain exams (retain if valid score ≥8/10\ge 8/10)
    Screen candidate SQL annotators ASA_S on query translation tests (retain if score ≥9/10\ge 9/10)
    for each database DD with description file MM do
        Qset←Q_{set} \leftarrow GenerateNaturalQuestions(AQ,D,MA_Q, D, M)
        for each question Q∈QsetQ \in Q_{set} do
            Assign QQ independently to two annotators a1,a2∈ASa_1, a_2 \in A_S
            Y1←WriteSQL(a1,Q,D,M)Y_1 \leftarrow \text{WriteSQL}(a_1, Q, D, M)
            Y2←WriteSQL(a2,Q,D,M)Y_2 \leftarrow \text{WriteSQL}(a_2, Q, D, M)
            V1←ExecuteSQL(Y1,D)V_1 \leftarrow \text{ExecuteSQL}(Y_1, D)
            V2←ExecuteSQL(Y2,D)V_2 \leftarrow \text{ExecuteSQL}(Y_2, D)
            if V1=V2V_1 = V_2 and V1≠∅V_1 \neq \emptyset and V1≠NULLV_1 \neq \text{NULL} then
                Y∗←SelectMoreEfficientQuery(Y1,Y2)Y^* \leftarrow \text{SelectMoreEfficientQuery}(Y_1, Y_2)
                K←ExtractKnowledgeEvidence(Q,D,Y∗)K \leftarrow \text{ExtractKnowledgeEvidence}(Q, D, Y^*)
                Add (Q,D,K,Y∗)(Q, D, K, Y^*) to D\mathcal{D}
            else
                (Q′,K′,Y∗)←ExpertAdjudication(Q,Y1,Y2,D,M)(Q', K', Y^*) \leftarrow \text{ExpertAdjudication}(Q, Y_1, Y_2, D, M)
                Ensure ExecuteSQL(Y∗,D)≠∅\text{ExecuteSQL}(Y^*, D) \neq \emptyset
                Add (Q′,D,K′,Y∗)(Q', D, K', Y^*) to D\mathcal{D}
  6. Knowl 6 — Execution Accuracy Benchmark Results on BIRD

    data/table

    Execution Accuracy (EX) evaluates the proportion of generated SQL queries that yield execution result sets identical to ground-truth queries on the database engine. The benchmark compares fine-tuned (FT) seq2seq models and in-context learning (ICL) large language models in zero-shot settings, evaluated with and without ground-truth external knowledge evidence (KK).

    Models Development Data Testing Data
    w/o knowledge w/ knowledge w/o knowledge w/ knowledge
    FT-based
    T5-Base 6.32% 11.54% (+5.22) 7.06% 12.89% (+5.83)
    T5-Large 9.71% 19.75% (+10.04) 10.38% 20.94% (+10.56)
    T5-3B 10.37% 23.34% (+12.97) 11.17% 24.05% (+12.88)
    ICL-based
    PaLM-2 18.77% 27.38% (+8.61) 24.71% 33.04% (+8.33)
    Codex 25.42% 34.35% (+8.93) 24.86% 36.47% (+11.61)
    ChatGPT 24.05% 37.22% (+13.17) 26.77% 39.30% (+12.53)
    ChatGPT + CoT 25.88% 36.64% (+10.76) 28.95% 40.08% (+11.24)
    Claude-2 28.29% 42.70% (+14.41) 34.60% 49.02% (+14.42)
    GPT-4 30.90% 46.35% (+15.45) 34.88% 54.89% (+20.01)
    GPT-4 + DIN-SQL – 50.72% – 55.90%
    Human Performance – – 72.37% 92.96% (+20.59)

    Key takeaways from these results:

    1. State-of-the-art LLMs such as GPT-4 (54.89% on the test set) and GPT-4 + DIN-SQL (55.90%) fall substantially short of human performance (92.96% with knowledge).
    2. External knowledge grounding provides consistent, substantial performance improvements across all model architectures and scales (e.g., +20.01% for GPT-4 and +12.88% for T5-3B on the test set).
    3. ICL-based large models outperform fine-tuned smaller language models (<5B<5\text{B} parameters) in both zero-shot generalization and knowledge grounding.
  7. Knowl 7 — Valid Efficiency Score (VES) Benchmark Results on BIRD

    data/table

    The Valid Efficiency Score (VES) assesses the relative execution speed of correctly executed SQL queries generated by baseline models compared to human-optimized ground-truth SQL queries.

    Models Development Data Testing Data
    w/o knowledge w/ knowledge w/o knowledge w/ knowledge
    FT-based
    T5-Base 7.78 12.90 (+5.12) 8.97 14.71 (+5.74)
    T5-Large 9.90 22.74 (+12.84) 12.25 25.00 (+12.75)
    T5-3B 13.62 25.57 (+11.95) 15.17 27.80 (+12.63)
    ICL-based
    PaLM-2 20.82 28.64 (+7.82) 31.32 38.41 (+7.09)
    Codex 33.37 43.41 (+10.04) 35.40 41.60 (+6.20)
    ChatGPT 27.97 43.81 (+15.84) 36.68 51.40 (+14.72)
    ChatGPT + CoT 32.33 42.30 (+9.97) 49.69 56.56 (+6.87)
    Claude-2 32.75 45.28 (+12.53) 39.32 55.77 (+16.45)
    GPT-4 34.60 49.77 (+15.17) 40.20 60.77 (+20.57)
    GPT-4 + DIN-SQL – 58.79 – 59.44
    Human Performance – – 70.36 90.27 (+19.91)

    Key takeaways:

    1. VES strictly correlates with Execution Accuracy (EX) because queries that fail to generate correct results receive an efficiency score of 0.
    2. GPT-4 achieves the highest individual model VES of 60.77 on the test set with knowledge, while human expert performance reaches 90.27 with knowledge.
    3. Incorporating external knowledge evidence boosts VES across all models (e.g., raising test VES from 40.20 to 60.77 for GPT-4).
  8. Knowl 8 — Cross-Benchmark Performance Discrepancy: SPIDER vs. BIRD

    empirical result

    When evaluated under identical value-knowledge prompting conditions on development sets, state-of-the-art text-to-SQL models exhibit severe degradation when moving from SPIDER to BIRD:

    • DIN-SQL: 82.8% on SPIDER vs. 50.7% on BIRD (-32.1%)
    • ChatGPT (gpt-3.5-turbo): 72.1% on SPIDER vs. 37.2% on BIRD (-34.9%)
    • Codex (code-davinci-002): 74.1% on SPIDER vs. 34.4% on BIRD (-39.7%)
    • T5-3B: 71.5% on SPIDER vs. 23.3% on BIRD (-48.2%)
    • T5-Large: 69.3% on SPIDER vs. 19.8% on BIRD (-49.5%)
    • T5-Base: 57.9% on SPIDER vs. 11.5% on BIRD (-46.4%)

    This gap demonstrates that existing semantic parsers are overfit to schema-heavy, value-lean academic benchmarks and struggle with BIRD's large-scale databases containing noisy data types, multi-table joins, and ungrounded domain values.

  9. Knowl 9 — Error Categorization of LLM-Generated SQL Queries

    empirical result

    A manual evaluation of 500 randomly sampled incorrect SQL queries produced by ChatGPT (gpt-3.5-turbo) reveals four primary error distributions:

    1. Wrong Schema Linking (41.6%): The model understands the question intent but incorrectly maps semantic entities to unrelated table columns or foreign tables (e.g., retrieving Street, City, Zip when only StreetAbr was requested, or selecting from the wrong table).
    2. Misunderstanding Database Content (40.8%): The model hallucinates non-existent schema objects (e.g., querying non-existent table lap_records instead of lapTimes) or predicts erroneous column values and predicate conditions due to massive database sizes.
    3. Misunderstanding Knowledge Evidence (17.6%): The model fails to translate human knowledge annotations into proper SQL syntax, frequently copying arithmetic pseudocode verbatim into the SQL query (e.g., generating DIVIDE(SUM(spent), COUNT(spent)) instead of using standard SQL arithmetic operators).
    4. Syntax and Dialect Errors (3.0%): The model generates invalid SQL keywords or mixes SQL dialects (e.g., invoking MySQL's YEAR() function inside SQLite instead of STRFTIME('%Y', ...)).
  10. Knowl 10 — SQL Execution Efficiency Optimization Strategies: Query Rewriting and Indexing

    empirical result

    Experiments on large-scale databases demonstrate two distinct mechanisms to achieve substantial execution time savings:

    1. Two-Stage Rule-Based Query Rewriting: Rewriting semantically correct SQL queries according to query rewrite rules yields an average execution time reduction of 77.75% across sampled development queries. Key rewrite rules include:
      • Converting WHERE ... IN (SELECT ...) subqueries to explicit INNER JOIN operations, enabling the query engine to perform join and filter operations simultaneously without materializing intermediate tables (e.g., 99.92% time savings).
      • Replacing COUNT(*) with COUNT(primary_key) on NOT NULL columns to avoid scanning all columns (e.g., 67.93% time savings).
      • Replacing ORDER BY column DESC LIMIT 1 with a subquery condition WHERE column = (SELECT MAX(column) FROM ...) to bypass full-table sort operations in unindexed environments (e.g., 62.39% time savings).
    2. Database Indexing: Adding single-column indexes and unique indexes on frequently joined foreign key and primary key columns (e.g., CREATE INDEX account_district_id_index ON account(district_id)) without altering the generated SQL text yields an execution time reduction of 87.27% by avoiding full table scans.

Coverage note — Omitted minor prompt variations, detailed sub-word cloud frequency counts, and standard open-source license summaries that do not affect the main methodological or empirical contributions.

References

  1. 1.Peter Alsberg. Space and time savings through large data base compression and dynamic restructuring. Proceedings of the IEEE, 63:1114–1122, 1975.
  2. 2.Anthropic. Introducing Claude. 2023. URL https://www.anthropic.com/index/introducing-claude.
  3. 3.Ruichu Cai, Boyan Xu, Zhenjie Zhang, Xiaoyan Yang, Zijian Li, and Zhihao Liang. An encoder-decoder framework translating natural language to database queries. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence (IJCAI-18), page 3977–3983, 2018.
  4. 4.Zefeng Cai, Xiangyu Li, Binyuan Hui, Min Yang, Bowen Li, Binhua Li, Zheng Cao, Weijie Li, Fei Huang, Luo Si, and Yongbin Li. STAR: SQL guided pre-training for context-dependent text-to-SQL parsing. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 1235–1247, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics.
  5. 5.Ruisheng Cao, Lu Chen, Zhi Chen, Yanbin Zhao, Su Zhu, and Kai Yu. LGESQL: Line graph enhanced text-to-SQL model with mixed local and non-local relations. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2541–2555, Online, August 2021. Association for Computational Linguistics.
  6. 6.Shuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang, Jiarong Jiang, Joseph Lilien, Steve Ash, William Yang Wang, Zhiguo Wang, Vittorio Castelli, Patrick Ng, and Bing Xiang. Dr.spider: A diagnostic evaluation benchmark towards text-to-SQL robustness. In The Eleventh International Conference on Learning Representations, 2023.
  7. 7.Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. FinQA: A dataset of numerical reasoning over financial data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3697–3711, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling language modeling with pathways. ArXiv, abs/2204.02311, 2022.
  9. 9.Deborah A. Dahl, Madeleine Bates, Michael Brown, William Fisher, Kate Hunicke-Smith, David Pallett, Christine Pao, Alexander Rudnicky, and Elizabeth Shriberg. Expanding the scope of the ATIS task: The ATIS-3 corpus. In Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8-11, 1994, 1994.
  10. 10.Longxu Dou, Yan Gao, Xuqi Liu, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Min-Yen Kan, and Jian-Guang Lou. Towards knowledge-intensive text-to-SQL semantic parsing with formulaic knowledge. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5240–5253, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics.
  11. 11.Yujian Gan, Xinyun Chen, Qiuping Huang, Matthew Purver, John R. Woodward, Jinxia Xie, and Pengsheng Huang. Towards robustness of text-to-sql models against synonym substitution. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 2505–2515. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.acl-long.195.
  12. 12.Jiaqi Guo, Zecheng Zhan, Yan Gao, Yan Xiao, Jian-Guang Lou, Ting Liu, and Dongmei Zhang. Towards complex text-to-SQL in cross-domain database with intermediate representation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4524–4535, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1444.
  13. 13.Moshe Hazoom, Vibhor Malik, and Ben Bogin. Text-to-SQL in the wild: A naturally-occurring dataset based on stack exchange data. In Proceedings of the 1st Workshop on Natural Language Processing for Programming (NLP4Prog 2021), pages 77–87, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.nlp4prog-1.9.
  14. 14.Joseph M. Hellerstein. Quantitative data cleaning for large databases. In Proceedings of the 2008 ACM SIGMOD international conference on Management of data, page 1197–1200, 2008.
  15. 15.Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang. Are large pre-trained language models leaking your personal information? In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 2038–2047, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics.
  16. 16.Binyuan Hui, Ruiying Geng, Qiyu Ren, Binhua Li, Yongbin Li, Jian Sun, Fei Huang, Luo Si, Pengfei Zhu, and Xiaodan Zhu. Dynamic hybrid relation exploration network for cross-domain context-dependent semantic parsing. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pages 13116–13124. AAAI Press, 2021.
  17. 17.Binyuan Hui, Xiang Shi, Ruiying Geng, Binhua Li, Yongbin Li, Jian Sun, and Xiaodan Zhu. Improving text-to-sql with schema dependency learning. In arXiv:2103.04399, 2021.
  18. 18.Binyuan Hui, Ruiying Geng, Lihan Wang, Bowen Qin, Yanyang Li, Bowen Li, Jian Sun, and Yongbin Li. S2SQL: Injecting syntax to question-schema interaction graph encoder for text-to-SQL parsers. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1254–1262, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.findings-acl.99.
  19. 19.Ihab F. Ilyas and Xu Chu. Trends in cleaning relational data: Consistency and deduplication. Foundations and Trends in Databases, 5:281–393, 2015.
  20. 20.Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, Jayant Krishnamurthy, and Luke Zettlemoyer. Learning a neural semantic parser from user feedback. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 963–973, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1089.
  21. 21.Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems, volume 35, pages 22199–22213, 2022.
  22. 22.Jan Kossmann, Stefan Halfpap, Marcel Jankrift, and Rainer Schlosser. Magic mirror in my hand, which is the best in the land? an experimental evaluation of index selection algorithms. Proceedings of the VLDB Endowment, 13(12):2382–2395, 2020.
  23. 23.Prerna S. Kulkarni and Jagdish W. Bakal. Survey on data cleaning. In 2014 International Conference on Advances in Computing, Communications and Informatics (ICACCI), page 2361–2366, 2014.
  24. 24.Chia-Hsuan Lee, Oleksandr Polozov, and Matthew Richardson. KaggleDBQA: Realistic evaluation of text-to-SQL parsers. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2261–2273, Online, August 2021. Association for Computational Linguistics.
  25. 25.Gyubok Lee, Hyeonji Hwang, Seongsu Bae, Yeonsu Kwon, Woncheol Shin, Seongjun Yang, Minjoon Seo, Jong-Yeup Kim, and Edward Choi. Ehrsql: A practical text-to-sql benchmark for electronic health records. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 15589–15601. Curran Associates, Inc., 2022.
  26. 26.Dandan Li, Lu Han, and Yi Ding. Sql query optimization methods of relational database system. In 2010 Second International Conference on Computer Engineering and Applications, volume 1, pages 557–560. IEEE, 2010.
  27. 27.Jinyang Li, Binyuan Hui, Reynold Cheng, Bowen Qin, Chenhao Ma, Nan Huo, Fei Huang, Wenyu Du, Luo Si, and Yongbin Li. Graphix-t5: Mixing pre-trained transformers with graph-aware layers for text-to-sql parsing. ArXiv, abs/2301.07507, 2023.
  28. 28.Tanzim Mahmud, KM Azharul Hasan, Mahtab Ahmed, and Thwoi Hla Ching Chak. A rule based approach for nlp based query processing. In 2015 2nd International Conference on Electrical Information and Communication Technologies (EICT), pages 78–82. IEEE, 2015.
  29. 29.Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ramakanth Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, Edouard Grave, Yann LeCun, and Thomas Scialom. Augmented language models: a survey. ArXiv, abs/2302.07842, 2023.
  30. 30.Vamsi Krishna Myalapalli and ASN Chakravarthy. Revamping sql queries for cost based optimization. In 2016 International Conference on Circuits, Controls, Communications and Computing (I4C), pages 1–6. IEEE, 2016.
  31. 31.Paulo H. Oliveira, Daniel dos Santos Kaster, Caetano Traina, and Ihab F. Ilyas. Batchwise probabilistic incremental data cleaning. ArXiv, abs/2011.04730, 2020.
  32. 32.OpenAI. Gpt-4 technical report. ArXiv, abs/2303.08774, 2023.
  33. 33.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155, 2022.
  34. 34.Hamid Pirahesh, Joseph M. Hellerstein, and Waqar Hasan. Extensible/rule based query rewrite optimization in starburst. In Proceedings of the ACM SIGMOD International Conference on Management of Data, pages 39–48, 1992.
  35. 35.Mohammadreza Pourreza and Davood Rafiei. DIN-SQL: decomposed in-context learning of text-to-sql with self-correction. CoRR, abs/2304.11015, 2023. doi: 10.48550/arXiv.2304.11015.
  36. 36.Jiexing Qi, Jingyao Tang, Ziwei He, Xiangpeng Wan, Yu Cheng, Chenghu Zhou, Xinbing Wang, Quanshi Zhang, and Zhouhan Lin. RASAT: Integrating relational structures into pretrained Seq2Seq model for text-to-SQL. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3215–3229, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics.
  37. 37.Bowen Qin, Binyuan Hui, Lihan Wang, Min Yang, Jinyang Li, Binhua Li, Ruiying Geng, Rongyu Cao, Jian Sun, Luo Si, Fei Huang, and Yongbin Li. A survey on text-to-sql parsing: Concepts, methods, and future directions. In arXiv:2208.13629, 2022.
  38. 38.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020.
  39. 39.Nitarshan Rajkumar, Raymond Li, and Dzmitry Bahdanau. Evaluating the text-to-sql capabilities of large language models. ArXiv, abs/2204.00498, 2022.
  40. 40.Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9895–9901, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics.
  41. 41.Peter Shaw, Ming-Wei Chang, Panupong Pasupat, and Kristina Toutanova. Compositional generalization and natural language variation: Can a semantic parsing approach handle both? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 922–938, Online, August 2021. Association for Computational Linguistics.
  42. 42.Alon Talmor, Ori Yoran, Ronan Le Bras, Chandra Bhagavatula, Yoav Goldberg, Yejin Choi, and Jonathan Berant. Commonsenseqa 2.0: Exposing the limits of ai through gamification. In J. Vanschoren and S. Yeung, editors, Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1. Curran, 2021.
  43. 43.Bailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov, and Matthew Richardson. RAT-SQL: Relation-aware schema encoding and linking for text-to-SQL parsers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7567–7578, Online, July 2020. Association for Computational Linguistics.
  44. 44.Lihan Wang, Bowen Qin, Binyuan Hui, Bowen Li, Min Yang, Bailin Wang, Binhua Li, Jian Sun, Fei Huang, Luo Si, and Yongbin Li. Proton: Probing schema linking information from pre-trained language models for text-to-sql parsing. In Aidong Zhang and Huzefa Rangwala, editors, KDD ’22: The 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Washington, DC, USA, August 14 - 18, 2022, pages 1889–1898. ACM, 2022. doi: 10.1145/3534678.3539305.
  45. 45.Lijie Wang, Ao Zhang, Kun Wu, Ke Sun, Zhenghua Li, Hua Wu, Min Zhang, and Haifeng Wang. DuSQL: A large-scale and pragmatic Chinese text-to-SQL dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6923–6935, Online, November 2020. Association for Computational Linguistics.
  46. 46.Ping Wang, Tian Shi, and Chandan K. Reddy. Text-to-sql generation for question answering on electronic medical records. Proceedings of The Web Conference 2020, 2020.
  47. 47.Zhaoguo Wang, Zhou Zhou, Yicun Yang, Haoran Ding, Gansen Hu, Ding Ding, Chuzhe Tang, Haibo Chen, and Jinyang Li. Wetune: Automatic discovery and verification of query rewrite rules. In Proceedings of the 2022 International Conference on Management of Data, pages 94–107, 2022.
  48. 48.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Huai hsin Chi, F. Xia, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. ArXiv, abs/2201.11903, 2022.
  49. 49.Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. UnifiedSKG: Unifying and multi-tasking structured knowledge grounding with text-to-text language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 602–631, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics.
  50. 50.Xiaojun Xu, Chang Liu, and Dawn Song. Sqlnet: Generating structured queries from natural language without reinforcement learning. ArXiv preprint, 2017.
  51. 51.Navid Yaghmazadeh, Yuepeng Wang, Isil Dillig, and Thomas Dillig. Sqlizer: query synthesis from natural language. Proceedings of the ACM on Programming Languages, 1(OOPSLA): 1–26, 2017.
  52. 52.Tao Yu, Zifan Li, Zilin Zhang, Rui Zhang, and Dragomir Radev. TypeSQL: Knowledge-based type-aware neural text-to-SQL generation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), page 588–594, 2018.
  53. 53.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, page 3911–3921, 2018.
  54. 54.Tao Yu, Rui Zhang, Alex Polozov, Christopher Meek, and Ahmed Hassan Awadallah. Score: Pre-training for context representation in conversational semantic parsing. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021.
  55. 55.John M. Zelle and Raymond J. Mooney. Learning to parse database queries using inductive logic programming. In Proceedings of the Fourteenth National Conference on Artificial Intelligence and Ninth Conference on Innovative Applications of Artificial Intelligence, pages 1050–1055, 1996.
  56. 56.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. Opt: Open pre-trained transformer language models. ArXiv, abs/2205.01068, 2022.
  57. 57.Chen Zhao, Yu Su, Adam Pauls, and Emmanouil Antonios Platanios. Bridging the generalization gap in text-to-SQL parsing with schema expansion. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5568–5578, Dublin, Ireland, May 2022. Association for Computational Linguistics.
  58. 58.Victor Zhong, Caiming Xiong, and Richard Socher. Seq2SQL: Generating structured queries from natural language using reinforcement learning. In CoRR abs/1709.00103, 2017.
  59. 59.Victor Zhong, Mike Lewis, Sida I. Wang, and Luke Zettlemoyer. Grounded adaptation for zero-shot executable semantic parsing. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6869–6882. Association for Computational Linguistics, 2020. doi: 10.18653/v1/2020.emnlp-main.558.
  60. 60.Rong Zhou. Research on key performance index prediction of distributed database based on machine learning algorithm. In Proceedings of the 2nd International Conference on Cognitive Based Information Processing and Applications (CIPA 2022) Volume 2, pages 563–567. Springer, 2023.
  61. 61.Xuanhe Zhou, Guoliang Li, Chengliang Chai, and Jianhua Feng. A learned query rewrite system using monte carlo tree search. Proceedings of the VLDB Endowment, 15(1):46–58, 2021. doi: 10.14778/3485450.3485456.
  62. 62.Xuanhe Zhou, Chengliang Chai, Guoliang Li, and Ji Sun. Database meets artificial intelligence: A survey. IEEE Transactions on Knowledge and Data Engineering, 34(3):1096–1116, 2022. doi: 10.1109/TKDE.2020.2994641.

Citation

MLA
Li, J., et al. “Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs”. arXiv, 2023, http://arxiv.org/abs/2305.03111v3.
APA
Li, J., Hui, B., Qu, G., Yang, J., Li, B., Li, B., Wang, B., Qin, B., Cao, R., Geng, R., Huo, N., Zhou, X., Ma, C., Li, G., Chang, K. C. C., Huang, F., Cheng, R., & Li, Y. (2023). Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs. arXiv. http://arxiv.org/abs/2305.03111v3
Chicago
Li, J., B. Hui, G. Qu, et al. 2023. “Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs”. arXiv. http://arxiv.org/abs/2305.03111v3.
Harvard
Li, J. et al. (2023) “Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.03111v3.
Vancouver
1. Li J, Hui B, Qu G, et al (2023) Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs. arXiv

BibTeX

@article{li2023can,
  title = {Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs},
  author = {Li, Jinyang and Hui, Binyuan and Qu, Ge and Yang, Jiaxi and Li, Binhua and Li, Bowen and Wang, Bailin and Qin, Bowen and Cao, Rongyu and Geng, Ruiying and Huo, Nan and Zhou, Xuanhe and Ma, Chenhao and Li, Guoliang and Chang, Kevin C. C. and Huang, Fei and Cheng, Reynold and Li, Yongbin},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.03111v3},
  eprint = {2305.03111}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors