RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQL
Haoyang LiJing ZhangCuiping LiHong Chen
Proposes a text-to-SQL framework that decouples schema linking from query generation by pre-filtering relevant database elements with a ranking cross-encoder and generating SQL skeletons before full queries to guide decoding.
Modern data management systems rely heavily on relational databases, but non-technical users often struggle to query them effectively because they lack expertise in Structured Query Language (SQL). Automated text-to-SQL systems address this gap by translating everyday natural language questions into database queries. However, standard sequence-to-sequence language models face significant performance hurdles because they attempt to simultaneously determine which database tables and columns to reference (schema linking) and assemble the structural query keywords (skeleton parsing). This coupling creates excessive complexity and leaves models vulnerable to variations in phrasing and database structures.
The article demonstrates and evaluates a novel framework called RESDSQL, which decouples schema linking from skeleton parsing to improve both query generation accuracy and overall system robustness. The authors evaluate this approach across standard cross-domain benchmarks and challenging stress-test scenarios that simulate real-world usage.
To achieve this separation, the approach implements a two-stage process. First, an independent cross-encoder model identifies, ranks, and filters database tables and columns, feeding only the most relevant items into the main translation model. This cross-encoder incorporates a column-enhancement mechanism to infer table names when users mention only column names, along with an imbalanced-data loss function. Second, the sequence-to-sequence decoder generates the overall SQL skeleton before generating the final, complete SQL query, allowing the high-level structural framework to guide the detailed output.
Empirical evaluations demonstrate substantial performance gains across key metrics. First, on the challenging Spider benchmark test set, RESDSQL set a new state-of-the-art result by achieving 79.9% execution accuracy, outperforming the previous top baseline of 75.5% by 4.4 percentage points. Second, the framework proves highly parameter-efficient: its base configuration outmatched standard baseline models that were roughly 13 times larger. Third, ablation analyses show that filtering and ranking schema items is the primary performance driver, improving exact match accuracy by 4.5 percentage points and execution accuracy by 7.8 percentage points. Finally, across three robustness benchmarks featuring synonyms, domain knowledge paraphrases, and omitted column names, RESDSQL consistently surpassed leading competitors by wide margins, achieving up to 81.9% execution accuracy on realistic perturbation tests.
These findings indicate that separating structural query planning from schema item selection substantially reduces model training difficulty and operational noise. For organizations deploying natural language interfaces to databases, this translates directly to lower computational infrastructure costs, higher query reliability, and greater resilience against ambiguous or imperfect user inputs without requiring complex graph architectures.
Organizations developing or deploying natural language query interfaces should consider adopting decoupled pre-filtering and two-step skeleton decoding architectures over monolithic models. Teams should also select filtering thresholds (such as the number of top tables and columns retained) based on their specific schema complexity, balancing the trade-off between missing required fields and introducing unnecessary noise.
Confidence in these findings is high given the extensive validation across multiple public benchmark datasets. However, decision-makers should note that the system relies on fixed selection thresholds for database elements and requires normalized SQL training conventions. In operational settings with highly specialized enterprise jargon or vastly larger enterprise schemas, targeted pilot testing remains advisable.
- Paper: Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task, Tao Yu et al. (2018). Introduces the Spider benchmark and formalizes the cross-domain Text-to-SQL challenge upon which RESDSQL is developed and evaluated.
- Paper: Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning, Victor Zhong et al. (2017). Establishes foundational modular neural semantic parsing techniques for mapping natural language questions into structured SQL components.
- Paper: Natural Language to Code Translation with Execution, Freda Shi et al. (2022). Demonstrates execution-guided decoding and candidate selection strategies for natural language to SQL code generation.
- Paper: TIARA: Multi-grained Retrieval for Robust Question Answering over Large Knowledge Base, Yiheng Shu et al. (2022). Pioneers multi-stage retrieval and cross-encoder schema ranking coupled with constrained sequence-to-sequence generation for structured query synthesis.
- Paper: DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction, Mohammadreza Pourreza et al. (2023). Extends the decoupled schema linking and modular query parsing concept to in-context learning with large language models and self-correction mechanisms.
- Paper: Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs, Jinyang Li et al. (2023). Introduces BIRD, a large-scale database-grounded benchmark that tests text-to-SQL models on real-world dirty database contents and complex schema linking challenges.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). Generalizes decoupled reading, retrieval, and reasoning strategies to structured reasoning over databases and knowledge graphs using foundation models.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). Investigates self-debugging and iterative error-correction for generated SQL and code directly on the Spider benchmark.