Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval
Peter Baile ChenYi ZhangDan Roth
Proposes a join-aware table retrieval framework using mixed-integer programming to re-rank candidate tables by jointly evaluating query-table relevance and table joinability, improving end-to-end question-answering accuracy over multi-table databases.
Modern question-answering systems and retrieval-augmented generation pipelines increasingly depend on tabular data to answer complex business and analytical queries. In real-world enterprise databases and data lakes, information is typically normalized across multiple separate tables rather than centralized in a single flat table. Existing retrieval methods often assume that an answer resides in a single table or that multi-table relationships can be deduced entirely from the text of the user prompt. When systems fail to account for how tables connect, they retrieve disjointed, unjoinable data, causing language models to generate incorrect answers or hallucinate facts.
The article evaluates and demonstrates a join-aware multi-table retrieval framework designed to overcome these limitations. The objective is to determine whether incorporating inferred table-to-table join relationships alongside standard query relevance during the retrieval phase improves the accuracy of both table selection and downstream question answering.
To address this challenge, the authors formulated table retrieval as a re-ranking optimization task solved through mixed-integer linear programming. The method decomposes user queries into fine-grained concept-attribute pairs and evaluates candidate tables across three dimensions: coarse query-table relevance, fine-grained sub-query coverage, and table-to-table compatibility. Compatibility is inferred by evaluating schema similarity, instance overlap, and primary-to-foreign key likelihoods, while network flow constraints ensure the retrieved tables form a fully connected join graph. The approach was evaluated on multi-table queries from two benchmark datasets, Spider (443 queries across 81 tables) and Bird (1,095 queries across 77 tables), against standard dense retrieval baselines.
The analysis shows that join-aware re-ranking consistently outperforms traditional table retrieval baselines. In top-2 table retrieval, the proposed method improved retrieval F1 scores by up to 5.6% on Spider and up to 10.7% on Bird compared to baseline methods. Downstream question answering accuracy improved noticeably: feeding the top-5 re-ranked tables to a language model yielded up to a 5.4% increase in end-to-end execution accuracy over baseline inputs. Furthermore, the performance gains were most pronounced on complex queries involving three or more tables and bridging tables, where table-to-table relevance provided the largest relative boost in retrieval quality.
These findings indicate that retrieval systems cannot treat tables as isolated text passages; inferring relational structure during retrieval is essential for accurate multi-table reasoning. In practice, adopting join-aware retrieval reduces downstream execution errors and mitigates the risk of feeding irrelevant context to costly language models. The results also show that when databases already possess documented primary and foreign key constraints, integrating these gold constraints pushes retrieval and answering performance even higher.
Organizations developing retrieval-augmented generation or automated SQL generation over relational data should incorporate relational compatibility constraints into their retrieval pipelines. For existing deployments, teams should evaluate whether retrieval quality degrades on complex, multi-table questions and consider a staged pilot using re-ranking optimization before expanding query complexity.
The approach currently relies on mixed-integer programming solvers, which can face scalability challenges when applied to massive data lakes or extremely large schema collections. The evaluation also assumes single-column join keys, leaving multi-column compound keys and non-relational table connections for future work. Nonetheless, the evidence strongly supports that join-aware re-ranking provides significant accuracy gains across multi-table question-answering tasks.
- Paper: Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task, Tao Yu et al. (2018). Spider establishes the cross-domain, multi-table text-to-SQL benchmark used by the source, making its schema and query challenges essential context for the evaluation.
- Paper: Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs, Jinyang Li et al. (2023). BIRD introduces the realistic database benchmark used in the source, so its task design and evaluation setting clarify what the reported retrieval gains address.
No sufficiently relevant recommendations were found.
