Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval

Peter Baile ChenYi ZhangDan Roth

article2024ACL53 citations

Proposes a join-aware table retrieval framework using mixed-integer programming to re-rank candidate tables by jointly evaluating query-table relevance and table joinability, improving end-to-end question-answering accuracy over multi-table databases.

Listen

Modern question-answering systems and retrieval-augmented generation pipelines increasingly depend on tabular data to answer complex business and analytical queries. In real-world enterprise databases and data lakes, information is typically normalized across multiple separate tables rather than centralized in a single flat table. Existing retrieval methods often assume that an answer resides in a single table or that multi-table relationships can be deduced entirely from the text of the user prompt. When systems fail to account for how tables connect, they retrieve disjointed, unjoinable data, causing language models to generate incorrect answers or hallucinate facts.

The article evaluates and demonstrates a join-aware multi-table retrieval framework designed to overcome these limitations. The objective is to determine whether incorporating inferred table-to-table join relationships alongside standard query relevance during the retrieval phase improves the accuracy of both table selection and downstream question answering.

To address this challenge, the authors formulated table retrieval as a re-ranking optimization task solved through mixed-integer linear programming. The method decomposes user queries into fine-grained concept-attribute pairs and evaluates candidate tables across three dimensions: coarse query-table relevance, fine-grained sub-query coverage, and table-to-table compatibility. Compatibility is inferred by evaluating schema similarity, instance overlap, and primary-to-foreign key likelihoods, while network flow constraints ensure the retrieved tables form a fully connected join graph. The approach was evaluated on multi-table queries from two benchmark datasets, Spider (443 queries across 81 tables) and Bird (1,095 queries across 77 tables), against standard dense retrieval baselines.

The analysis shows that join-aware re-ranking consistently outperforms traditional table retrieval baselines. In top-2 table retrieval, the proposed method improved retrieval F1 scores by up to 5.6% on Spider and up to 10.7% on Bird compared to baseline methods. Downstream question answering accuracy improved noticeably: feeding the top-5 re-ranked tables to a language model yielded up to a 5.4% increase in end-to-end execution accuracy over baseline inputs. Furthermore, the performance gains were most pronounced on complex queries involving three or more tables and bridging tables, where table-to-table relevance provided the largest relative boost in retrieval quality.

These findings indicate that retrieval systems cannot treat tables as isolated text passages; inferring relational structure during retrieval is essential for accurate multi-table reasoning. In practice, adopting join-aware retrieval reduces downstream execution errors and mitigates the risk of feeding irrelevant context to costly language models. The results also show that when databases already possess documented primary and foreign key constraints, integrating these gold constraints pushes retrieval and answering performance even higher.

Organizations developing retrieval-augmented generation or automated SQL generation over relational data should incorporate relational compatibility constraints into their retrieval pipelines. For existing deployments, teams should evaluate whether retrieval quality degrades on complex, multi-table questions and consider a staged pilot using re-ranking optimization before expanding query complexity.

The approach currently relies on mixed-integer programming solvers, which can face scalability challenges when applied to massive data lakes or extremely large schema collections. The evaluation also assumes single-column join keys, leaving multi-column compound keys and non-relational table connections for future work. Nonetheless, the evidence strongly supports that join-aware re-ranking provides significant accuracy gains across multi-table question-answering tasks.

No sufficiently relevant recommendations were found.

Cover for Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval

Abstract

Retrieving relevant tables containing the necessary information to accurately answer a given question over tables is critical to open-domain question-answering (QA) systems. Previous methods assume the answer to such a question can be found either in a single table or multiple tables identified through question decomposition or rewriting. However, neither of these approaches is sufficient, as many questions require retrieving multiple tables and joining them through a join plan that cannot be discerned from the user query itself. If the join plan is not considered in the retrieval stage, the subsequent steps of reasoning and answering based on those retrieved tables are likely to be incorrect. To address this problem, we introduce a method that uncovers useful join relations for any query and database during table retrieval. We use a novel re-ranking method formulated as a mixed-integer program that considers not only table-query relevance but also table-table relevance that requires inferring join relationships. Our method outperforms the state-of-the-art approaches for table retrieval by up to 9.3% in F1 score and for end-to-end QA by up to 5.4% in accuracy.

Table of Contents

  • 1 Introduction
  • 2 Problem Description
  • 3 Join-aware Multi-Table Retrieval
  • 3.1 Query-Table Relevance
  • 3.2 Table Re-ranking with Table-Table Relevance
  • 3.2.1 Sub-query Coverage
  • 3.2.2 Connectedness
  • 3.3 Table-Table Relationship Inference
  • 4 Evaluation
  • 4.1 Experimental Settings
  • 4.2 Table Retrieval Performances
  • 4.3 End-to-end Performances
  • 4.4 Discussion
  • 5 Related Work
  • 6 Conclusion
  • 7 Limitations
  • References
  • A Prompts
  • A.1 Decomposition
  • A.2 SQL generation
  • B MIP formulation
  • B.1 Coverage quality
  • B.2 Connectedness
  • C Dataset processing

Knowls

  1. Knowl 1 — Join-aware multi-table retrieval returns a connected table expression

    definition

    Given a query QQ and a corpus CC of tables, join-aware multi-table retrieval returns a ranked expression of selected tables together with join conditions, rather than returning a single table. The returned expression must contain the information needed to answer QQ, and its tables must be joinable into a connected set. A join condition identifies a column in one table and a column in another table; for example, joining Client Info.id to Disp.client_id specifies the pair of columns used to connect those tables. The central retrieval challenge is to choose tables that are both relevant to the query and mutually compatible, even when the required join plan cannot be inferred from the wording of the query alone.

  2. Knowl 2 — MIP jointly optimizes query relevance and table compatibility

    model/method

    The proposed re-ranker selects exactly KK tables from a candidate corpus by maximizing coarse query-to-table relevance, fine-grained query-to-column relevance, and pairwise table-join compatibility. Let i,ji,j index candidate tables, k,lk,l index columns, and qq index query sub-queries. The binary variable bib_i indicates whether table ii is selected; dqikd_{qik} indicates whether column kk of table ii covers sub-query qq; and cijklc_{ij}^{kl} indicates whether column kk of table ii is selected to join with column ll of table jj. The input scores rir_i, rqikr_{qik}, and ωijkl\omega_{ij}^{kl} respectively denote coarse table relevance, fine-grained sub-query/column relevance, and column-pair joinability. The number of selected tables KK is an input integer, and ∣Q∣|Q| is the number of sub-queries.

    max⁡b,d,c  ∑iribi+∑q,i,krqikdqik+∑i,j,k,lωijklcijkl\max_{b,d,c}\; \sum_i r_i b_i + \sum_{q,i,k} r_{qik}d_{qik} + \sum_{i,j,k,l}\omega_{ij}^{kl}c_{ij}^{kl}

    All decision variables are binary. The model imposes ∑ibi=K\sum_i b_i=K and ∑i,j,k,lcijkl≤K−1\sum_{i,j,k,l}c_{ij}^{kl}\le K-1; it permits a join only between selected tables, with 2(cijkl+cjilk)≤bi+bj2(c_{ij}^{kl}+c_{ji}^{lk})\le b_i+b_j. At most one column pair may be selected as the join between any table pair, so ∑k,lcijkl≤1\sum_{k,l}c_{ij}^{kl}\le1. A sub-query may be assigned to at most one column in any one table, so ∑kdqik≤1\sum_k d_{qik}\le1, and relevance assignments are allowed only for selected tables. The complete system is designed to favor a set that covers query information while also providing compatible join paths.

  3. Knowl 3 — Fine-grained relevance aligns query concepts and attributes to columns

    model/method

    To identify which table columns cover different parts of a complex query, the method decomposes the query into sub-queries represented as concept–attribute pairs. For example, a question about the identifier of a trip that began at the station with the greatest dock count can be represented by the pairs (trip,id)(\text{trip},\text{id}) and (station,dock count)(\text{station},\text{dock count}). GPT-3.5 Turbo produces this decomposition using a prompted format with five in-context examples in the retrieval experiments. A bi-encoder then scores the semantic similarity between each sub-query qq and each candidate column ckc_k in table TiT_i; this is the fine-grained relevance score rqikr_{qik}. The re-ranker also uses a coarse query-to-table relevance score from the initial retrieval model, so that individual tables can be selected for relevance to only part of a multi-part query.

  4. Knowl 4 — Joinability combines schema, instance overlap, and key likelihood

    equation

    For columns ckc_k in table TiT_i and clc_l in table TjT_j, the proposed joinability score combines column relevance with an estimate of whether the columns satisfy a key–foreign-key relationship:

    ωijkl=(e(ck,cl)+j(ck,cl))max⁡{u(ck),u(cl)}.\omega_{ij}^{kl} = \bigl(e(c_k,c_l)+j(c_k,c_l)\bigr)\max\{u(c_k),u(c_l)\}.

    Here, j(ck,cl)j(c_k,c_l) is the Jaccard similarity of the columns’ instance-value sets: the number of shared distinct values divided by the number of distinct values in their union. The schema similarity e(ck,cl)e(c_k,c_l) is a weighted sum of semantic similarities between schema segments, including the column headers, table names, and other columns in the respective tables; segment representations are encoded with a pretrained embedding model and compared by cosine similarity. The uniqueness u(c)u(c) of a column is the number of its unique values divided by the number of instances in its table. Taking the maximum uniqueness score reflects the method’s assumption that at least one side of a valid key–foreign-key join should be a primary key. When gold key–foreign-key constraints are supplied in the experiments, the compatibility score for corresponding column pairs is set to 1.

  5. Knowl 5 — Coverage terms discourage spreading weak matches across columns

    model/method

    The fine-grained relevance sum alone can reward assigning a query sub-query weakly to many columns. The model therefore adds a binary variable dqd_q to indicate whether sub-query qq is covered at least once and adds a coverage reward α∑qdq\alpha\sum_q d_q to the objective, where α\alpha controls the importance of covering sub-queries. Coverage requires at least one column assignment for that sub-query, expressed as dq≤∑i,kdqikd_q\le\sum_{i,k}d_{qik}. The total number of sub-query-to-column assignments is capped at the number of sub-queries, ∑q,i,kdqik≤∣Q∣\sum_{q,i,k}d_{qik}\le |Q|, encouraging stronger, less diffuse mappings. As characterized by the paper, a large α\alpha prioritizes covering sub-queries and favors strong mappings, whereas a lower α\alpha allows more emphasis on selecting multiple tables with weaker mappings.

  6. Knowl 6 — A flow constraint enforces connectedness of selected tables

    model/method

    The selected tables must form a connected subgraph whose nodes are tables and whose edges are selected join relationships. The model enforces this by augmenting the compatibility graph with a source and a sink. It chooses one selected table as the root, sends KK units of flow from the source through the selected-table graph, and requires each of the KK selected tables to send one unit to the sink. Flow can traverse a table-to-table edge only when the corresponding column-pair join has been selected. Thus all selected tables must be reachable from the root through chosen joins; disconnected but individually relevant tables cannot satisfy the flow requirement.

  7. Knowl 7 — Evaluation uses aggregated Spider and Bird corpora with inferred joins

    experimental setup

    The evaluation adapts the Spider and Bird text-to-SQL datasets to open-domain retrieval by aggregating tables across their topic-specific databases into one corpus per dataset. Queries whose gold SQL uses only one table are excluded, as are queries removed after human inspection because their gold joins produce incorrect instances. The resulting evaluation sets contain 443 queries and 81 tables for Spider, and 1,095 queries and 77 tables for Bird. Experiments are run both without gold key–foreign-key constraints and with the original constraints supplied. The baselines are Contriever-msmarco, which ranks flattened tables by query/table embedding cosine similarity, and DTR, implemented by fine-tuning TAPAS-large with an 8:2 training/validation split. Retrieval is measured with precision, recall, and F1 at different values of KK. For downstream evaluation, GPT-3.5 Turbo 1106 at temperature 0 generates SQL from retrieved tables, and execution accuracy is computed against the gold SQL results.

  8. Knowl 8 — Join-aware reranking improves retrieval, especially at small K and for complex queries

    empirical result

    The retrieval results grid reported on page 7 compares F1 for the original retrievers, full join-aware reranking (JAR-F), reranking with gold key–foreign-key constraints (JAR-G), and an ablation without table–table relevance (JAR-D). The values below are F1 scores; each triplet is for K=2,5,10K=2,5,10, respectively.

    Dataset Method K=2K=2 K=5K=5 K=10K=10
    Spider DTR 78.9 57.1 34.9
    Spider Contriever 74.0 55.5 33.8
    Spider JAR-F (DTR) 84.5 58.3 35.0
    Spider JAR-F (Contriever) 80.5 55.7 33.9
    Spider JAR-G (DTR) 89.6 59.1 35.1
    Spider JAR-G (Contriever) 87.1 56.3 33.8
    Spider JAR-D (DTR) 83.9 57.6 35.0
    Spider JAR-D (Contriever) 79.5 55.5 33.8
    Bird DTR 61.3 50.9 34.8
    Bird Contriever 61.6 50.8 34.3
    Bird JAR-F (DTR) 72.0 55.1 35.3
    Bird JAR-F (Contriever) 69.5 55.2 34.9
    Bird JAR-G (DTR) 73.8 56.0 35.8
    Bird JAR-G (Contriever) 72.7 55.3 35.3
    Bird JAR-D (DTR) 70.7 54.2 35.1
    Bird JAR-D (Contriever) 68.8 53.7 34.5

    Full reranking improves over the corresponding baseline at every reported dataset and KK combination; for example, JAR-F (DTR) gains 10.7 F1 points over DTR on Bird at K=2K=2 (72.0 versus 61.3). JAR-G generally improves further when gold join constraints are available. JAR-F also generally exceeds JAR-D, showing that table–table compatibility adds value beyond fine-grained query–table relevance. The page 8 breakdown by query size shows that this additional value is larger for queries requiring at least three tables: the JAR-F versus JAR-D F1 gains are 0.9 and 1.5 points on Spider and Bird with DTR, and 1.0 and 2.8 points with Contriever.

  9. Knowl 9 — Improved retrieval raises SQL execution accuracy and table selection quality

    empirical result

    In the end-to-end evaluation, GPT-3.5 Turbo receives retrieved tables and generates SQL. The execution-accuracy results reported on page 8 show that full join-aware reranking with five input tables exceeds the corresponding baseline supplied with five, ten, or twenty tables in most comparisons. Entries are execution accuracy; the baseline rows specify the number of input tables, while JAR-F and JAR-G use five.

    Retriever Input Spider Bird
    DTR Top-5 46.3 30.4
    DTR Top-10 47.4 34.8
    DTR Top-20 48.1 31.5
    JAR-F (DTR) Top-5 50.6 35.8
    JAR-G (DTR) Top-5 52.2 36.9
    Contriever Top-5 43.5 29.7
    Contriever Top-10 43.1 30.6
    Contriever Top-20 44.7 32.9
    JAR-F (Contriever) Top-5 47.4 35.1
    JAR-G (Contriever) Top-5 48.3 36.2

    The paper also evaluates the tables selected by the SQL-generating LLM, reporting top-5 table-selection F1. JAR-F (DTR) obtains 77.0 on Spider and 65.3 on Bird, compared with 75.2 and 58.8 for top-5 DTR; JAR-F (Contriever) obtains 74.7 and 64.7, compared with 71.3 and 55.2 for top-5 Contriever. These results indicate that reranking improves both executed answer accuracy and the LLM’s selection of tables for its final SQL.

  10. Knowl 10 — The reranking formulation has scalability and relationship-scope limitations

    limitation

    The paper identifies possible scalability problems and sensitivity to the input data as limitations of the mixed-integer-programming re-ranker. Its join model assumes single-column joins and focuses on key–foreign-key relationships; compound-key joins are left for future work. The authors also note that real questions may depend on other kinds of table connections, such as matching values of the same column across tables. They identify more scalable optimization and broader relationship types as directions for further investigation.

Coverage note — No substantial contributed material was omitted; prompt exemplars and individual dataset-cleaning examples were left out because they are illustrative implementation details rather than separate contributions.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  2. 2.Wenhu Chen, Ming-Wei Chang, Eva Schlinger, William Wang, and William W Cohen. 2020a. Open question answering over tables and text. arXiv preprint arXiv:2010.10439.
  3. 3.Zhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu, and Brian D Davison. 2020b. Table search using a deep contextualized language model. In Proceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval, pages 589–598.
  4. 4.Tianji Cong, James Gale, Jason Frantz, HV Jagadish, and Çagatay Demiralp. 2022. Warpgate: A semantic join discovery system for cloud data warehouse. arXiv preprint arXiv:2212.14155.
  5. 5.Shimon Even and R Endre Tarjan. 1975. Network flow and testing graph connectivity. SIAM journal on computing, 4(4):507–518.
  6. 6.Dawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun, Yichen Qian, Bolin Ding, and Jingren Zhou. 2024. Text-to-sql empowered by large language models: A benchmark evaluation. Proc. VLDB Endow., 17(5):1132–1145.
  7. 7.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
  8. 8.Jonathan Herzig, Thomas Müller, Syrine Krichene, and Julian Eisenschlos. 2021. Open domain question answering over tables via dense retrieval. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 512–519, Online. Association for Computational Linguistics.
  9. 9.Junjie Huang, Wanjun Zhong, Qian Liu, Ming Gong, Daxin Jiang, and Nan Duan. 2022. Mixed-modality representation learning and pre-training for joint table-and-text retrieval in openqa. arXiv preprint arXiv:2210.05197.
  10. 10.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised dense information retrieval with contrastive learning. Trans. Mach. Learn. Res., 2022.
  11. 11.Chia-Hsuan Lee, Oleksandr Polozov, and Matthew Richardson. 2021. KaggleDBQA: Realistic evaluation of text-to-SQL parsers. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2261–2273, Online. Association for Computational Linguistics.
  12. 12.Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. ArXiv, abs/2005.11401.
  13. 13.Jinyang Li, Binyuan Hui, Ge Qu, Binhua Li, Jiaxi Yang, Bowen Li, Bailin Wang, Bowen Qin, Rongyu Cao, Ruiying Geng, Nan Huo, Xuanhe Zhou, Chenhao Ma, Guoliang Li, Kevin C. C. Chang, Fei Huang, Reynold Cheng, and Yongbin Li. 2023a. Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls.
  14. 14.Jinyang Li, Binyuan Hui, Ge Qu, Binhua Li, Jiaxi Yang, Bowen Li, Bailin Wang, Bowen Qin, Rongyu Cao, Ruiying Geng, et al. 2023b. Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls. arXiv preprint arXiv:2305.03111.
  15. 15.Aiwei Liu, Xuming Hu, Lijie Wen, and Philip S. Yu. 2023. A comprehensive evaluation of chatgpt’s zero-shot text-to-sql capability.
  16. 16.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  17. 17.Linyong Nan, Yilun Zhao, Weijin Zou, Narutatsu Ri, Jaesung Tae, Ellen Zhang, Arman Cohan, and Dragomir Radev. 2023. Enhancing few-shot text-to-sql capabilities of large language models: A study on prompt design strategies.
  18. 18.Fatemeh Nargesian, Erkang Zhu, Renée J Miller, Ken Q Pu, and Patricia C Arocena. 2019. Data lake management: challenges and opportunities. Proceedings of the VLDB Endowment, 12(12):1986–1989.
  19. 19.Feifei Pan, Mustafa Canim, Michael Glass, Alfio Gliozzo, and James Hendler. 2022. End-to-end table question answering via retrieval-augmented generation. arXiv preprint arXiv:2203.16714.
  20. 20.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761.
  21. 21.Michael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Wen-tau Yih, and Sebastian Riedel. 2021. Joint verification and reranking for open fact checking over tables. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6787–6799, Online. Association for Computational Linguistics.
  22. 22.Tianshu Wang, Hongyu Lin, Xianpei Han, Le Sun, Xiaoyang Chen, Hao Wang, and Zhenyu Zeng. 2024. Db copilot: Scaling natural language querying to massive databases.
  23. 23.Zhiruo Wang, Zhengbao Jiang, Eric Nyberg, and Graham Neubig. 2022. Table retrieval may not necessitate table-specific model design. In Proceedings of the Workshop on Structured and Unstructured Knowledge Integration (SUKI), pages 36–46, Seattle, USA. Association for Computational Linguistics.
  24. 24.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018a. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3911–3921, Brussels, Belgium. Association for Computational Linguistics.
  25. 25.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. 2018b. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. arXiv preprint arXiv:1809.08887.
  26. 26.Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang. 2022. A survey of knowledge-enhanced text generation. ACM Computing Surveys, 54(11s):1–38.
  27. 27.Yi Zhang and Zachary G. Ives. 2020. Finding related tables in data lakes for interactive data science. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, SIGMOD ’20, page 1951–1966, New York, NY, USA. Association for Computing Machinery.
  28. 28.Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2sql: Generating structured queries from natural language using reinforcement learning. ArXiv, abs/1709.00103.
  29. 29.Chunting Zhou, Graham Neubig, Jiatao Gu, Mona Diab, Francisco Guzmán, Luke Zettlemoyer, and Marjan Ghazvininejad. 2021. Detecting hallucinated content in conditional neural sequence generation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1393–1404, Online. Association for Computational Linguistics.
  30. 30.Erkang Zhu, Dong Deng, Fatemeh Nargesian, and Renée J. Miller. 2019. Josie: Overlap set similarity search for finding joinable tables in data lakes. In Proceedings of the 2019 International Conference on Management of Data, SIGMOD ’19, page 847–864, New York, NY, USA. Association for Computing Machinery.

Citation

MLA
Chen, P. B., et al. “Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 2687–99, https://doi.org/10.18653/v1/2024.acl-long.148.
APA
Chen, P. B., Zhang, Y., & Roth, D. (2024). Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2687–2699. https://doi.org/10.18653/v1/2024.acl-long.148
Chicago
Chen, P. B., Y. Zhang, and D. Roth. 2024. “Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2687–99. https://doi.org/10.18653/v1/2024.acl-long.148.
Harvard
Chen, P.B., Zhang, Y. and Roth, D. (2024) “Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2687–2699. Available at: https://doi.org/10.18653/v1/2024.acl-long.148.
Vancouver
1. Chen PB, Zhang Y, Roth D (2024) Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2687–2699

BibTeX

@inproceedings{chen-etal-2024-table,
    title = "Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval",
    author = "Chen, Peter Baile  and
      Zhang, Yi  and
      Roth, Dan",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.148/",
    doi = "10.18653/v1/2024.acl-long.148",
    pages = "2687--2699"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/