Rethinking Tabular Data Understanding with Large Language Models

Tianyang LiuFei WangMuhao Chen

article2024NAACL51 citations

Demonstrates how structural variations degrade tabular reasoning in large language models and introduces table normalization alongside a mixed self-consistency framework that unites textual and symbolic reasoning to reach state-of-the-art accuracy on WikiTableQuestions.

Listen

Modern organizations increasingly rely on large language models to automate data analysis and decision-making from structured tabular data. However, the extent to which these models can reliably comprehend and reason over structured tables remains underexplored. The article evaluates the robustness of language models against structural table changes, compares the effectiveness of textual versus code-based symbolic reasoning, and tests whether aggregating multiple reasoning pathways can improve overall accuracy.

The researchers conducted an extensive empirical evaluation using GPT-3.5 across more than 3,300 perturbed configurations and full benchmarks from the standard WikiTableQuestions dataset, along with supplementary testing on the TabFact dataset. They evaluated models across four table variations, including original tables, row-shuffled tables, transposed tables, and combinations of both, using two zero-shot reasoning approaches: step-by-step natural language prompting and dynamic code execution via a Python shell agent.

The findings reveal several critical insights into model performance. First, language models are highly vulnerable to structural changes, with direct table transposition causing performance to drop by about 14% in textual prompting and plunging by roughly 78% in code-based reasoning. Second, models struggle with direct transposition detection and execution, but a proposed two-stage content-aware normalization method effectively eliminates structural vulnerabilities without degrading baseline performance. Third, textual reasoning slightly outperforms code-based reasoning on standard tasks, with textual methods excelling in semantic comprehension while code execution performs better on precise counting and filtering. Finally, combining both reasoning pathways via a mixed self-consistency voting mechanism achieved a new state-of-the-art accuracy of 73.6% on the WikiTableQuestions benchmark.

These findings demonstrate that adopting automated table normalization and hybrid reasoning strategies significantly reduces operational risk, minimizes calculation errors, and improves data processing accuracy. Organizations deploying language models should not rely solely on single-prompt or purely code-based frameworks. Instead, implementations should combine automated structural preprocessing to standardize table orientations and leverage ensemble voting across both textual and programmatic reasoning paths, while taking care to avoid automated row reordering if queries depend on original row sequences.

While the study provides strong empirical support for these strategies, stakeholders should note certain limitations. The evaluations relied on GPT-3.5 and Wikipedia-sourced tables, which may introduce domain-specific biases or data overlap, and all methods experienced noticeable accuracy declines as table length increased. Further pilots and evaluations using newer model architectures and enterprise-specific datasets are recommended before full-scale deployment.

arXiv: 2312.16702Leolty/tablellm
Cover for Rethinking Tabular Data Understanding with Large Language Models

Abstract

Large Language Models (LLMs) have shown to be capable of various tasks, yet their capability in interpreting and reasoning over tabular data remains an underexplored area. In this context, this study investigates from three core perspectives: the robustness of LLMs to structural perturbations in tables, the comparative analysis of textual and symbolic reasoning on tables, and the potential of boosting model performance through the aggregation of multiple reasoning pathways. We discover that structural variance of tables presenting the same content reveals a notable performance decline, particularly in symbolic reasoning tasks. This prompts the proposal of a method for table structure normalization. Moreover, textual reasoning slightly edges out symbolic reasoning, and a detailed error analysis reveals that each exhibits different strengths depending on the specific tasks. Notably, the aggregation of textual and symbolic reasoning pathways, bolstered by a mix self-consistency mechanism, resulted in achieving SOTA performance, with an accuracy of 73.6% on WIKITABLEQUESTIONS, representing a substantial advancement over previous existing table processing paradigms of LLMs.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 3.1 Problem Definition
  • 3.2 Experimental Setup
  • 4 LLM Robustness to Structural Perturbations
  • 4.1 Impacts of Table Perturbations on LLMs
  • 4.2 Limitations of Table Transposition with LLMs
  • 4.3 Table Structure Normalization
  • 5 Comparing Textual and Symbolic Reasoning
  • 5.1 Results
  • 5.2 Error Analysis
  • 6 Reasoning Aggregation
  • 6.1 Methods
  • 6.2 Overall Evaluation
  • 7 Conclusion
  • References
  • A Prompts
  • A.1 Prompt of Direct Prompting (DP).
  • A.2 Prompt of Python Agent.
  • A.3 Prompt of LLMs as Table Transposer
  • A.4 Prompt of LLMs as Table Transposition Detector
  • A.5 Prompt of Content-Aware Transposition Determination
  • A.6 Prompt of Resorting
  • A.7 Prompt of Self-Evaluation
  • B Analysis for LLMs as Table Transposer
  • B.1 Case Study
  • B.2 Analysis
  • C Impact of Table Size on WTQ Performance
  • D Error Case Study for WTQ
  • D.1 Table Misinterpretation
  • D.1.1 Counting Error
  • D.1.2 Locating Error
  • D.2 Coding Error
  • D.2.1 Attribute Noise Error
  • D.2.2 Special Row Misinterpretation Error
  • D.2.3 Incorrect Coding
  • D.3 Misalignment Issue
  • D.3.1 Answer Format Issue
  • D.3.2 Answer Deviation Error
  • D.4 Logical Inconsistency
  • D.4.1 Reasoning Conflict in DP
  • D.4.2 Reasoning Mistakes in PyAgent
  • D.5 Execution Issue
  • D.5.1 Interaction Bound or Looping Error
  • D.5.2 Non-Observable Action Error
  • D.6 Resorting Issue
  • E Analysis of Mix Self-Consistency
  • E.1 Ablation Study of Output Selection
  • E.2 Mechanics of Mix Self-Consistency in Output Selection
  • F Results of Mix Self-Consistency on TabFact

Knowls

  1. Knowl 1 — Mix Self-Consistency Framework for Tabular Reasoning

    model/method

    Mix Self-Consistency (Mix-SC) is a reasoning aggregation mechanism that integrates both textual reasoning (Direct Prompting, DP) and symbolic reasoning (Python Shell Agent, PyAgent) to answer questions over tabular data.

    Under positive decoding temperature (e.g., T=0.8T = 0.8), the method samples NDPN_{\text{DP}} trajectories from Direct Prompting (step-by-step textual chain-of-thought) and NPyAgentN_{\text{PyAgent}} trajectories from PyAgent (multi-step Python code generation and execution in a pandas shell). The total sampled outputs N=NDP+NPyAgentN = N_{\text{DP}} + N_{\text{PyAgent}} (with an equal 5+55+5 split totaling 10 outputs being the standard configuration) are aggregated, and the final answer is selected via majority voting.

    The framework exploits the property that LLM answer variance reflects task confidence: when a reasoning paradigm is well suited to the question (e.g., symbolic operations for counting/filtering, or textual reasoning for semantic nuances), its stochastic trajectories consistently converge to a shared prediction. Conversely, when a paradigm is poorly suited, its predictions disperse across different candidate answers. Aggregating outputs across paradigms allows dominant, high-confidence predictions to prevail. In the event of ties during majority voting, textual reasoning outputs are prioritized due to higher baseline accuracy.

  2. Knowl 2 — Table Structure Normalization Strategy

    model/method

    Table Structure Normalization (NORM) is a two-stage preprocessing pipeline designed to convert structurally perturbed tables (such as transposed column-tables or row-shuffled tables) into standardized, well-ordered row-tables before passing them to downstream reasoning models:

    1. Stage 1 (Content-Aware Transposition): Rather than relying on spatial layout parsing, NORM extracts the candidate header sets from the first row T0,∗T_{0,*} and the first column T∗,0T_{*,0} of a given table TT. An LLM semantically determines whether T0,∗T_{0,*} or T∗,0T_{*,0} better represents attribute headers. If T∗,0T_{*,0} is selected, the table is transposed via Ti,j⊤=Tj,iT^\top_{i,j} = T_{j,i} to place headers in the topmost row.

    2. Stage 2 (Row Reordering / Re-sorting): To recover logical ordering after arbitrary row shuffling, the LLM is provided the table column headings along with only the first three and last three rows (to prevent bias from existing messy row arrangements). The LLM recommends a sorting scheme (e.g., numerical, alphabetical, or chronological sorting on primary/secondary columns), which is then applied programmatically across the entire table.

    For downstream questions sensitive to specific original row indices or original sequence (e.g., "What is the last entry in the chart?"), the re-sorting stage is omitted (designated as NORM w/o Resort) to prevent altering sequence-dependent answers.

  3. Knowl 3 — Content-Aware Transposition Determination

    model/method

    Directly prompting LLMs to detect if a table requires transposition (f(T)→{0,1}f(T) \to \{0, 1\}) suffers from severe structural bias, where models default to recommending no transposition (achieving 93.35% accuracy on original row-tables but dropping to 32.54% on transposed column-tables).

    Content-Aware Transposition Determination replaces structural layout evaluation with semantic header identification. Given table TT with title τ\tau, first row vector T0,∗T_{0,*}, and first column vector T∗,0T_{*,0}, the model computes:

    f(T,T0,∗,T∗,0)→choice∈{T0,∗,T∗,0,None}f(T, T_{0,*}, T_{*,0}) \to \text{choice} \in \{T_{0,*}, T_{*,0}, \text{None}\}

    Selecting T0,∗T_{0,*} indicates that the table is already row-oriented, whereas selecting T∗,0T_{*,0} indicates that the table is column-oriented and must be transposed. This semantic approach achieves 97.39% accuracy on original tables and 94.77% accuracy on transposed tables using zero-shot GPT-3.5 on WikiTableQuestions, neutralizing the structural detection bias.

  4. Knowl 4 — Tabular QA Performance on WikiTableQuestions Benchmark

    data/table

    Exact Match Accuracy on the complete test set of WikiTableQuestions comparing fine-tuned PLMs and LLM-based prompting approaches. All reported LLM results use zero-shot or few-shot inference with GPT-3.5 or Codex:

    Method Accuracy (%)
    Fine-tuning Based Models
    TAPAS (Herzig et al., 2020) 48.8
    T5-3B (Xie et al., 2022) 49.3
    TAPEX (Liu et al., 2022) 57.5
    ReasTAP (Zhao et al., 2022) 58.7
    OmniTab (Jiang et al., 2022) 63.3
    LLM-Based Methods
    StructGPT (Jiang et al., 2023a) [GPT-3.5] 48.4
    BINDER (Cheng et al., 2023) [GPT-3.5] 55.5
    BINDER (Cheng et al., 2023) [Codex] 64.6
    LEVER (Ni et al., 2023) [Codex] 65.8
    Dater (Ye et al., 2023) [Codex] 65.9
    Mix Self-Consistency w/ NORM w/o Resort 73.6

    The Mix Self-Consistency method combining 5 Direct Prompting and 5 Python Agent trajectories alongside table structure normalization (without re-sorting) achieves a state-of-the-art accuracy of 73.6% on the WikiTableQuestions benchmark in a fully zero-shot setting.

  5. Knowl 5 — Impact of Structural Perturbations on LLM Tabular Reasoning

    data/table

    Evaluation of GPT-3.5 under four structural table configurations on a 421-table subset (837 QA pairs, totaling 3,348 evaluations) of WikiTableQuestions, using Direct Prompting (DP) and Python Shell Agent (PyAgent), with and without Table Structure Normalization (NORM):

    Perturbation Configuration DP DP + NORM PyAgent PyAgent + NORM
    Original (TT) 59.50 58.66 (-1.41%) 55.91 56.87 (+1.72%)
    Row-Shuffled (TΠT_{\Pi}) 52.21 (-12.25%) 58.66 (+12.35%) 47.91 (-14.31%) 57.11 (+19.20%)
    Transposed (T⊤T^\top) 51.14 (-14.05%) 58.30 (+14.00%) 12.45 (-77.73%) 55.44 (+346.02%)
    Transposed Shuffled (TΠ⊤T_{\Pi}^\top) 37.51 (-36.96%) 57.71 (+53.85%) 8.96 (-83.97%) 55.08 (+514.73%)

    Textual reasoning (DP) demonstrates greater inherent resilience to structural perturbations than symbolic reasoning (PyAgent), which drops drastically under transposition from 55.91% to 12.45% and 8.96%. Applying NORM mitigates structural degradation across all configurations, restoring performance to baseline levels.

  6. Knowl 6 — Comparison of Textual and Symbolic Tabular Reasoning

    empirical result

    A comparison between Direct Prompting (DP, zero-shot textual Chain-of-Thought) and Python Shell Agent (PyAgent, zero-shot symbolic tool-use interacting with a pandas dataframe up to 5 iterative steps) on WikiTableQuestions reveals distinct operational characteristics:

    1. Single-Attempt Accuracy: In single zero-shot passes on normalized tables, Direct Prompting marginally outperforms PyAgent (58.66% vs. 56.87%).
    2. Partial Table View Execution: PyAgent can execute queries when provided only partial views of the table in the prompt (e.g., omitting central rows and displaying only the first three and last three rows). This partial-view setup achieves 52.45% accuracy (a 4.42% drop compared to the 56.87% full-table setup), enabling symbolic LLM processing of large tables exceeding prompt token limits.
    3. Self-Consistency Scaling: Aggregating 10 sampled trajectories at temperature 0.80.8 improves DP from 58.66% to 66.39% (66.99% with NORM w/o resort) and PyAgent from 56.87% to 61.39% (62.84% with NORM w/o resort).
  7. Knowl 7 — Error Taxonomy of Direct Prompting and Symbolic Python Agent

    data/table

    Categorization of errors from 50 sampled incorrect outputs each for Direct Prompting (DP) and Python Shell Agent (PyAgent) on the WikiTableQuestions dataset:

    Error Type DP (%) PyAgent (%) Description
    Table Misinterpretation 42% –†^\dagger Misunderstanding cell contents, counting errors, or cell mislocalization.
    Coding Errors – 38% Flawed code, unhandled attribute variations, or counting summary rows.
    Misalignment Issue 24% 28% Conceptually correct answers violating output formatting constraints.
    Logical Inconsistency 20% 10% Reasoning contradictions between intermediate deductions and final answer.
    Execution Issue – 12% Tool iteration limit reached, looping errors, or non-observable actions.
    Resorting Issue 10% 8% Normalization sorting altering ground truth of position-dependent queries.

    †^\daggerTable misinterpretation in PyAgent is categorized under coding errors. Percentages do not sum to 100% due to uncategorized issues such as dataset annotation errors.

    Direct Prompting fails primarily on table misinterpretation (counting and spatial localization failures resulting from table linearization), whereas PyAgent fails primarily on programmatic oversights (such as failing to filter out aggregated Total rows or handling noisy string affixes).

  8. Knowl 8 — Failure Modes of LLMs as Direct Table Transposers

    empirical result

    When prompted to directly transpose tables (f(T)→T⊤f(T) \to T^\top or f(T⊤)→Tf(T^\top) \to T) without external programmatic tools, GPT-3.5 achieves only 53.68% accuracy when transposing row tables to column tables and 51.07% on inverse transposition.

    Transposition accuracy exhibits distinct dimension sensitivity: row-to-column transposition accuracy drops sharply as the number of rows increases, whereas column-to-row transposition accuracy drops as the number of columns increases.

    Qualitative analysis reveals that errors concentrate in columns or rows with homogeneous, repetitive, or missing data entries (e.g., repeated placeholder symbols such as '?'). In linearized textual representations, the LLM loses tracking of coordinate alignments across repetitive values, leading to cell misplacements and offset shifts.

  9. Knowl 9 — Self-Evaluation Selection Strategy for Hybrid Tabular Reasoning

    model/method

    Self-Evaluation is an LLM-guided selection heuristic designed to choose between a textual reasoning prediction (Answer A from Direct Prompting) and a symbolic reasoning prediction (Answer B from Python Shell Agent) without requiring multi-sample generation.

    The model evaluates both candidates using a two-stage evaluation prompt:

    1. Preliminary Evaluation: Assesses whether each candidate provides a direct, unambiguous response aligning with the query type, disregarding answers with extraneous formatting or off-target outputs.
    2. Task-Nature Evaluation: If both candidates are direct, the model classifies the task type. Symbolic reasoning (Answer B) is preferred for mathematical computation, row counting, and column localization tasks on extensive tables, provided execution logs contain no unhandled runtime errors. Textual reasoning (Answer A) is selected for semantic interpretation tasks.

    On the sampled WikiTableQuestions dataset, Self-Evaluation using only 2 reasoning paths achieves 64.22% accuracy (and 64.99% in self-evaluation settings), matching the performance of independent 10-path self-consistency runs of DP (66.39%) or PyAgent (61.39%).

  10. Knowl 10 — Tabular Reasoning Degradation Across Table Row Lengths

    empirical result

    Segmenting the WikiTableQuestions test set into 10 row-count bins (from [5,9][5, 9] rows to [43,518][43, 518] rows, each containing approximately 430 data points) demonstrates that reasoning accuracy decreases monotonically as table length increases across all methods:

    • Direct Prompting (1-path): Accuracy decreases from ≈63%\approx 63\% in the [5,9][5, 9] row range to <50%<50\% in the [43,518][43, 518] row range.
    • Python Shell Agent (1-path): Accuracy decreases from ≈61%\approx 61\% in small tables to ≈49%\approx 49\% in the largest table range.
    • Mix Self-Consistency (5 DP + 5 PyAgent): Accuracy decreases from ≈77%\approx 77\% on short tables to ≈67%\approx 67\% on tables with over 43 rows.

    This shared performance drop across both textual and symbolic paradigms reflects context distraction, attention dilution over large linearized token sequences, and increased probability of coding errors when filtering larger datasets.

  11. Knowl 11 — Fact Verification Performance on TabFact Benchmark

    data/table

    Evaluation of Mix Self-Consistency (5 DP + 5 PyAgent) without fine-tuning on a random subsample of 500 test instances from the TabFact fact-verification dataset at temperature T=0.8T = 0.8:

    Method Accuracy
    StructGPT (Jiang et al., 2023a) 0.708
    Dater (Ye et al., 2023) 0.874
    Mix Self-Consistency (5 DP + 5 PyAgent) 0.885

    The Mix Self-Consistency mechanism achieves 0.885 accuracy, outperforming both StructGPT (0.708) and Dater (0.874), demonstrating that aggregating textual and symbolic reasoning paths generalizes effectively to tabular fact verification without task-specific fine-tuning.

  12. Knowl 12 — Limitations of Tabular Data Understanding with LLMs

    limitation

    The methodology and experimental findings have three primary stated limitations:

    1. Model Scope: Due to API budgetary constraints, experiments were conducted exclusively on GPT-3.5 (gpt-3.5-turbo-0613 and gpt-3.5-turbo-16k-0613); performance and robustness patterns on larger foundation models such as GPT-4 or open-source tabular models were not evaluated.
    2. Training Data Contamination: Benchmark tables are sourced from Wikipedia, creating potential risks of memorization or data leakage where model answers might be recalled from pre-training rather than deduced from tabular reasoning.
    3. Index-Dependent Perturbations: Normalization via row re-sorting alters the ground truth answers for questions that explicitly depend on row ordering or ordinal positioning (such as queries asking for the first or last entry in the chart).

Coverage note — A detailed hyperparameter ablation across all 11 output allocation splits (ranging from 10 DP + 0 PyAgent to 0 DP + 10 PyAgent) was summarized in the Mix-SC knowl rather than given its own standalone knowl, as its key finding (5+5 being the optimal average configuration) is fully captured.

References

  1. 1.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  2. 2.Shuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang, Jiarong Jiang, Joseph Lilien, Steve Ash, William Yang Wang, Zhiguo Wang, Vittorio Castelli, Patrick Ng, and Bing Xiang. 2023. Dr.spider: A diagnostic evaluation benchmark towards text-to-sql robustness.
  3. 3.Harrison Chase. 2022. LangChain.
  4. 4.Wenhu Chen. 2023. Large language models are few(1)-shot table reasoners.
  5. 5.Wenhu Chen, Ming-Wei Chang, Eva Schlinger, William Yang Wang, and William W Cohen. 2020. Open question answering over tables and text. In International Conference on Learning Representations.
  6. 6.Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, and Tao Yu. 2023. Binding language models in symbolic languages. ICLR, abs/2210.02875.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.
  8. 8.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  9. 9.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models.
  10. 10.Leo Gao, John Schulman, and Jacob Hilton. 2022. Scaling laws for reward model overoptimization.
  11. 11.Zihui Gu, Ju Fan, Nan Tang, Preslav Nakov, Xiaoman Zhao, and Xiaoyong Du. 2022. PASTA: Table-operations aware fact verification via sentence-table cloze pre-training. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4971–4983, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  12. 12.Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. 2023. Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.
  13. 13.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Eisenschlos. 2020. TaPas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics.
  14. 14.Jinhao Jiang, Kun Zhou, Zican Dong, Keming Ye, Wayne Xin Zhao, and Ji-Rong Wen. 2023a. Structgpt: A general framework for large language model to reason on structured data.
  15. 15.Zhengbao Jiang, Yi Mao, Pengcheng He, Graham Neubig, and Weizhu Chen. 2022. OmniTab: Pretraining with natural and synthetic data for few-shot table-based question answering. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 932–942, Seattle, United States. Association for Computational Linguistics.
  16. 16.Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023b. Active retrieval augmented generation.
  17. 17.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023. Large language models are zero-shot reasoners.
  18. 18.Hongxin Li, Jingran Su, Yuntao Chen, Qing Li, and Zhaoxiang Zhang. 2023a. Sheetcopilot: Bringing software productivity to the next level through large language models.
  19. 19.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, Logesh Kumar Umapathi, Jian Zhu, Benjamin Lipkin, Muhtasham Oblokulov, Zhiruo Wang, Rudra Murthy, Jason Stillerman, Siva Sankalp Patel, Dmitry Abulkhanov, Marco Zocca, Manan Dey, Zhihan Zhang, Nour Fahmy, Urvashi Bhattacharyya, Wenhao Yu, Swayam Singh, Sasha Luccioni, Paulo Villegas, Maxim Kunakov, Fedor Zhdanov, Manuel Romero, Tony Lee, Nadav Timor, Jennifer Ding, Claire Schlesinger, Hailey Schoelkopf, Jan Ebert, Tri Dao, Mayank Mishra, Alex Gu, Jennifer Robinson, Carolyn Jane Anderson, Brendan Dolan-Gavitt, Danish Contractor, Siva Reddy, Daniel Fried, Dzmitry Bahdanau, Yacine Jernite, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. 2023b. Starcoder: may the source be with you!
  20. 20.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2023c. Making large language models better reasoners with step-aware verifier.
  21. 21.Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022. Tapex: Table pre-training via learning a neural sql executor.
  22. 22.Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Zhuang Li, Wen-Ding Li, Megan Risdal, Jia Li, Jian Zhu, Terry Yue Zhuo, Evgenii Zheltonozhskii, Nii Osae Osae Dade, Wenhao Yu, Lucas Krauß, Naman Jain, Yixuan Su, Xuanli He, Manan Dey, Edoardo Abati, Yekun Chai, Niklas Muennighoff, Xiangru Tang, Muhtasham Oblokulov, Christopher Akiki, Marc Marone, Chenghao Mou, Mayank Mishra, Alex Gu, Binyuan Hui, Tri Dao, Armel Zebaze, Olivier Dehaene, Nicolas Patry, Canwen Xu, Julian McAuley, Han Hu, Torsten Scholak, Sebastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, Mostofa Patwary, Nima Tajbakhsh, Yacine Jernite, Carlos Muñoz Ferrandis, Lingming Zhang, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. 2024. Starcoder 2 and the stack v2: The next generation.
  23. 23.Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, Edouard Grave, Yann LeCun, and Thomas Scialom. 2023. Augmented language models: a survey.
  24. 24.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2022. Webgpt: Browser-assisted question-answering with human feedback.
  25. 25.Ansong Ni, Srini Iyer, Dragomir Radev, Ves Stoyanov, Wen-tau Yih, Sida I Wang, and Xi Victoria Lin. 2023. Lever: Learning to verify language-to-code generation with execution. In Proceedings of the 40th International Conference on Machine Learning (ICML’23).
  26. 26.OpenAI. 2022. Codex.
  27. 27.OpenAI. 2023a. Chatgpt.
  28. 28.OpenAI. 2023b. Dall·e 3.
  29. 29.OpenAI. 2023c. Gpt-4 technical report.
  30. 30.Panupong Pasupat and Percy Liang. 2015. Compositional semantic parsing on semi-structured tables.
  31. 31.Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, and Denny Zhou. 2023. Large language models can be easily distracted by irrelevant context.
  32. 32.Significant Gravitas. 2023. Auto-GPT.
  33. 33.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2022. Learning to summarize from human feedback.
  34. 34.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. ArXiv preprint, abs/2302.13971.
  35. 35.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models.
  36. 36.Fei Wang, Kexuan Sun, Jay Pujara, Pedro Szekely, and Muhao Chen. 2021. Table-based fact verification with salience-aware learning. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4025–4036, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  37. 37.Fei Wang, Zhewei Xu, Pedro Szekely, and Muhao Chen. 2022. Robust (controlled) table-to-text generation with structure-aware equivariance learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  38. 38.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models.
  39. 39.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned language models are zero-shot learners.
  40. 40.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-thought prompting elicits reasoning in large language models.
  41. 41.Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2022. Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models. EMNLP.
  42. 42.Yunhu Ye, Binyuan Hui, Min Yang, Binhua Li, Fei Huang, and Yongbin Li. 2023. Large language models are versatile decomposers: Decompose evidence and questions for table-based reasoning.
  43. 43.Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020. TaBERT: Pretraining for joint understanding of textual and tabular data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8413–8426, Online. Association for Computational Linguistics.
  44. 44.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. Opt: Open pretrained transformer language models.
  45. 45.Tianshu Zhang, Xiang Yue, Yifei Li, and Huan Sun. 2023a. Tablellama: Towards open large generalist models for tables.
  46. 46.Wenqi Zhang, Yongliang Shen, Weiming Lu, and Yueting Zhuang. 2023b. Data-copilot: Bridging billions of data and humans with autonomous workflow.
  47. 47.Yilun Zhao, Linyong Nan, Zhenting Qi, Rui Zhang, and Dragomir Radev. 2022. ReasTAP: Injecting table reasoning skills during pre-training via synthetic reasoning examples. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9006–9018, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  48. 48.Yilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi, Wenlin Zhang, Xiangru Tang, Boyu Mi, and Dragomir Radev. 2023. Robut: A systematic study of table qa robustness against human-annotated adversarial perturbations.
  49. 49.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi. 2023. Least-to-most prompting enables complex reasoning in large language models.

Citation

MLA
Liu, T., et al. “Rethinking Tabular Data Understanding with Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 450–82, https://doi.org/10.18653/v1/2024.naacl-long.26.
APA
Liu, T., Wang, F., & Chen, M. (2024). Rethinking Tabular Data Understanding with Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 450–482. https://doi.org/10.18653/v1/2024.naacl-long.26
Chicago
Liu, T., F. Wang, and M. Chen. 2024. “Rethinking Tabular Data Understanding with Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 450–82. https://doi.org/10.18653/v1/2024.naacl-long.26.
Harvard
Liu, T., Wang, F. and Chen, M. (2024) “Rethinking Tabular Data Understanding with Large Language Models”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 450–482. Available at: https://doi.org/10.18653/v1/2024.naacl-long.26.
Vancouver
1. Liu T, Wang F, Chen M (2024) Rethinking Tabular Data Understanding with Large Language Models. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 450–482

BibTeX

@inproceedings{liu-etal-2024-rethinking,
    title = "Rethinking Tabular Data Understanding with Large Language Models",
    author = "Liu, Tianyang  and
      Wang, Fei  and
      Chen, Muhao",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.26/",
    doi = "10.18653/v1/2024.naacl-long.26",
    pages = "450--482"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/