Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs
Jinyang LiBinyuan HuiGe QuJiaxi YangBinhua LiBowen LiBailin WangBowen QinRuiying GengNan Huo
Introduces BIRD, a large-scale text-to-SQL benchmark across 95 multi-domain databases that reveals substantial performance gaps in large language models by evaluating real-world challenges such as dirty database values, external knowledge reasoning, and query execution efficiency.
Organizations increasingly seek to enable non-technical users to query massive corporate databases using natural language. While recent advances in artificial intelligence have shown high accuracy on academic benchmarks, existing evaluations focus predominantly on clean database schemas rather than the messy, large-scale data environments encountered in real-world business operations. This leaves significant uncertainty about whether leading automated systems can reliably serve as standalone database interfaces without risking costly errors or processing inefficiencies.
To bridge this gap, the article introduces BIRD, a large-scale evaluation benchmark grounded in massive, realistic databases across 37 diverse domains. The authors developed a dataset comprising 12,751 question-query pairs spanning 95 relational databases totaling 33.4 GB. Built using rigorous double-blind crowdsourcing and expert oversight, the benchmark introduces two critical dimensions: the necessity of external business knowledge to interpret ambiguous values, and a dedicated Valid Efficiency Score to assess the runtime efficiency of generated queries alongside their execution accuracy.
The findings show that current automated systems struggle significantly when confronted with realistic data environments. The top-performing system, GPT-4, achieved only a 54.89% execution accuracy on the hidden test set even when supplied with necessary domain knowledge, falling far short of the 92.96% achieved by human experts. Smaller fine-tuned language models scored under 25%. Furthermore, when external domain knowledge was omitted, automated accuracy plunged by 10 to 20 percentage points across all systems. An in-depth error analysis revealed that over 82% of model failures stemmed from misidentifying table structures or misunderstanding the actual underlying database contents. On the operational front, optimizing query structures through two-stage rewriting reduced query execution time by 77.75%, while adding targeted indexes reduced execution time by 87.27%.
These results demonstrate that contemporary language models are not yet dependable as fully autonomous interfaces for enterprise databases. Deploying them without oversight introduces substantial business risks, including hallucinated schema items, inaccurate data retrievals, and expensive, unoptimized queries that can strain computational resources. The significant performance drop in the absence of external context underscores that raw language comprehension is insufficient; reliable enterprise deployment fundamentally requires integrating structured business logic and automated query optimization.
Stakeholders are advised against deploying fully autonomous natural language database interfaces for high-stakes operational decisions at this stage. Instead, organizations should pursue human-in-the-loop workflows where automated queries assist rather than replace technical analysts. Technical roadmaps should prioritize developing structured knowledge-grounding frameworks, incorporating query-optimization layers before execution, and establishing standardized indexing to prevent performance bottlenecks. Although the evaluation is currently constrained to SQLite implementations, confidence in the findings remains high due to the strict double-blind methodology and the vast scale of tested domains.
- Paper: Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task, Tao Yu et al. (2018). Spider established the foundational cross-domain text-to-SQL benchmark whose clean schema assumptions and limitations directly motivated the creation of the realistic, large-scale BIRD benchmark.
- Paper: Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning, Victor Zhong et al. (2017). Seq2SQL introduced large-scale semantic parsing benchmarks and execution-guided evaluation for relational databases, providing key methodological foundations for modern text-to-SQL tasks.
- Paper: RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQL, Haoyang Li et al. (2023). RESDSQL demonstrates decoupling schema linking from skeleton parsing to address schema complexity in text-to-SQL, representing a key specialized baseline and structural approach examined in large-scale database querying.
- Paper: Graphix-T5: Mixing Pre-trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing, Jinyang Li et al. (2023). Graphix-T5 establishes how pre-trained transformers handle relational database schemas, providing essential context on the architectural difficulties models face when navigating complex database structures.
- Paper: DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction, Mohammadreza Pourreza et al. (2023). DIN-SQL directly builds upon the findings of the BIRD benchmark by designing a decomposed prompting framework with self-correction specifically evaluated against BIRD's challenging database scenarios.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). StructGPT extends the challenge of grounding large language models in large-scale structured databases and tables by developing a specialized reading-then-reasoning interface framework.
- Paper: Rethinking Tabular Data Understanding with Large Language Models, Tianyang Liu et al. (2024). This work explores the structural reasoning vulnerabilities and symbolic execution boundaries of LLMs over tabular databases highlighted by enterprise benchmarking studies.
- Paper: TableBench: A Comprehensive and Complex Benchmark for Table Question Answering, Xianjie Wu et al. (2025). TableBench advances realistic tabular evaluation beyond SQL generation into multi-step numerical analysis, fact checking, and data visualization across professional enterprise domains.
