HyTrel: Hypergraph-enhanced Tabular Data Representation Learning

Pei ChenSoumajyoti SarkarLeonard LausenBalasubramaniam SrinivasanSheng ZhaRuihong HuangGeorge Karypis

article2023NeurIPS52 citations

Proposes a hypergraph-based tabular language model that encodes structural inductive biases like row and column permutation invariance to improve representation learning across downstream table understanding tasks with minimal pretraining.

Listen

Tabular data is ubiquitous across enterprise databases, documents, and web pages, serving as a critical asset for business intelligence, knowledge extraction, and automated decision-making. Recent machine learning advances have applied language models to tables by flattening two-dimensional grids into one-dimensional text sequences. However, this sequential approach fails to account for fundamental tabular structures, specifically the principle that reordering rows or columns should not alter a table's underlying meaning, as well as the complex, multi-element relationships among cells, rows, and headers.

The article evaluates a novel tabular language model called HYTREL (Hypergraph-enhanced Tabular Data Representation Learning). The primary objective is to demonstrate that explicitly incorporating tabular structural properties—such as permutation invariance and hierarchical relationships—directly into model architecture yields more accurate, robust, and computationally efficient representations than standard sequence-based language models.

To achieve this, the approach models tables as hypergraphs, where individual cells act as nodes, and rows, columns, and the whole table act as multi-node connections known as hyperedges. The model uses a 12-layer structure-aware transformer architecture with set attention mechanisms that theoretically guarantee identical representations for permuted tables. The model was pretrained on 27 million public web tables using two self-supervised objectives: an ELECTRA-style cell corruption objective and a contrastive hypergraph learning objective. It was then fine-tuned and benchmarked against competitive baselines across four core table-understanding tasks: Column Type Annotation, Column Property Annotation, Table Type Detection, and Table Similarity Prediction.

The analysis produced several key findings. First, HYTREL consistently outperformed leading models such as TaBERT, TURL, and Doduo across all four downstream tasks, achieving top performance across classification and similarity metrics. Second, HYTREL demonstrated remarkable pretraining efficiency; even without any pretraining, its randomly initialized model achieved near state-of-the-art results, while baseline models degraded sharply without extensive pretraining. Third, HYTREL required only 5 pretraining epochs to reach peak performance, compared to 10 to 100 epochs required by conventional models. Fourth, theoretical and empirical tests confirmed that HYTREL generates zero representational distortion under row and column permutations, whereas sequential baselines exhibited significant sensitivity to order changes. Finally, HYTREL exhibited lower computational inference complexity (linear rather than quadratic relative to table elements) and faster inference times than sequential models.

These findings imply that treating tabular data according to its native geometry significantly lowers training overhead, compute costs, and execution latency while mitigating the risk of fragile, order-dependent errors. The results highlight that the choice of pretraining objective matters depending on the use case: the ELECTRA objective excelled at structural and semantic classification tasks, while contrastive pretraining proved superior for table similarity matching.

For technical leaders and practitioners, the article recommends adopting hypergraph-based tabular encoders when building automated data discovery, knowledge graph extraction, and cataloging systems. For processing exceptionally large enterprise tables, the authors recommend utilizing downsampling strategies, as empirical tests show downsampling drastically reduces memory consumption and training time without sacrificing downstream accuracy.

Decision-makers should note certain limitations and boundary conditions. HYTREL was evaluated strictly as a standalone tabular encoder; it is not currently configured out-of-the-box for joint text-table tasks, such as conversational table question answering or text generation, nor does it natively handle deeply nested hierarchical column headers or multi-table relational joins. Overall, confidence in HYTREL’s performance on single-table understanding tasks is high based on rigorous cross-validation and theoretical proofs, but extending it to conversational or generative workflows will require further architectural integration.

arXiv: 2307.08623
Cover for HyTrel: Hypergraph-enhanced Tabular Data Representation Learning

Abstract

Language models pretrained on large collections of tabular data have demonstrated their effectiveness in several downstream tasks. However, many of these models do not take into account the row/column permutation invariances, hierarchical structure, etc. that exist in tabular data. To alleviate these limitations, we propose HYTREL, a tabular language model, that captures the permutation invariances and three more structural properties of tabular data by using hypergraphs—where the table cells make up the nodes and the cells occurring jointly together in each row, column, and the entire table are used to form three different types of hyperedges. We show that HYTREL is maximally invariant under certain conditions for tabular data, i.e., two tables obtain the same representations via HYTREL iff the two tables are identical up to permutations. Our empirical results demonstrate that HYTREL consistently outperforms other competitive baselines on four downstream tasks with minimal pretraining, illustrating the advantages of incorporating the inductive biases associated with tabular data into the representations. Finally, our qualitative analyses showcase that HYTREL can assimilate the table structures to generate robust representations for the cells, rows, columns, and the entire table.1

Table of Contents

  • 1 Introduction
  • 2 HYTREL Model
  • 2.1 Formatter & Embedding Layer
  • 2.2 Hypergraph Encoder
  • 2.3 Invariances of the HYTREL Model
  • 2.4 Pretraining Heads
  • 3 Experiments
  • 3.1 Pre-training
  • 3.2 Fine-tuning
  • 3.3 Baselines
  • 3.4 Main Results
  • 4 Qualitative Analysis
  • 4.1 HYTREL Learns Permutation Robust Representations
  • 4.2 HYTREL Learns the Underlying Hierarchical Table Structure
  • 4.3 HYTREL Demonstrates Effective Pretraining by Capturing the Table Structures
  • 5 Related Work
  • 6 Limitations
  • 7 Conclusion
  • References
  • A Details of the Proofs
  • A.1 Preliminaries
  • A.2 Proofs
  • B Further Analysis of HYTREL
  • B.1 Effect of Input Table Size
  • B.2 Effect of Excessive Invariance
  • B.3 Inference Complexity of the Model
  • C Details of the Experiments
  • C.1 Pretraining
  • C.1.1 Model Size
  • C.1.2 Hyperparameters
  • C.2 Fine-tuning
  • C.2.1 Details of the Datasets
  • C.2.2 Hyperparameters
  • C.2.3 Details of the Baselines

Knowls

  1. Knowl 1 — Tabular Hypergraph Formulation and Embedding Initialization in HYTREL

    model/method

    HYTREL represents a table T=[M,H,R]\mathcal{T} = [M, H, R]—where MM is a caption, H=[h1,h2,…,hm]H = [h_1, h_2, \dots, h_m] denotes mm column headers, and R=[R1,R2,…,Rn]R = [R_1, R_2, \dots, R_n] denotes nn rows with cell values cijc_{ij}—as an undirected hypergraph G=(V,E)\mathcal{G} = (\mathcal{V}, \mathcal{E}).

    Each table cell cijc_{ij} is represented as a vertex vij∈Vv_{ij} \in \mathcal{V}, yielding ∣V∣=mn|\mathcal{V}| = mn nodes. The hyperedge set E\mathcal{E} consists of nn row hyperedges eire_i^r, mm column hyperedges ejce_j^c, and 11 global table hyperedge ete^t, giving ∣E∣=m+n+1|\mathcal{E}| = m + n + 1. Each cell node vijv_{ij} connects to exactly three incident hyperedges: its corresponding row hyperedge eire_i^r, column hyperedge ejce_j^c, and the table hyperedge ete^t. The graph topology is represented by an incidence matrix B∈{0,1}mn×(m+n+1)B \in \{0, 1\}^{mn \times (m+n+1)}, where Bij=1B_{ij} = 1 if node ii belongs to hyperedge jj and 00 otherwise.

    Embeddings are initialized into a hidden dimension FF as follows:

    • For each cell node vijv_{ij}, its feature vector Xvij,:∈RFX_{v_{ij}, :} \in \mathbb{R}^F is computed by feeding its constituent text tokens into an embedding layer and averaging their token embeddings.
    • For each column hyperedge ejce_j^c, its initial embedding Sejc,:∈RFS_{e_j^c, :} \in \mathbb{R}^F is obtained by averaging the token embeddings of header hjh_j.
    • For the table hyperedge ete^t, its initial embedding Set,:∈RFS_{e^t, :} \in \mathbb{R}^F is obtained by averaging the token embeddings of caption MM.
    • For each row hyperedge eire_i^r, lacking explicit semantic header text, embeddings Seir,:∈RFS_{e_i^r, :} \in \mathbb{R}^F are randomly initialized.
  2. Knowl 2 — HyperTrans Layer Architecture and Message Passing

    model/method

    HYTREL stacks 12 hypergraph-structure-aware transformer (HyperTrans) layers to alternate information propagation between cell nodes and hyperedges. Each HyperTrans layer tt comprises a Node-to-Hyperedge attention block (fV→Ef_{\mathcal{V} \to \mathcal{E}}), a Hyperedge Fusion block, and a Hyperedge-to-Node attention block (fE→Vf_{\mathcal{E} \to \mathcal{V}}).

    1. Node2Hyperedge Attention: Aggregates constituent cell node representations into intermediate hyperedge representations: S~e,:(t+1)=fV→E(Ke,X(t))\tilde{S}_{e,:}^{(t+1)} = f_{\mathcal{V} \to \mathcal{E}}\left(K_{e, X^{(t)}}\right) where Ke,X(t)={Xv,:(t):v∈e}K_{e, X^{(t)}} = \{X_{v,:}^{(t)} : v \in e\} denotes the set of hidden representations of nodes belonging to hyperedge ee.

    2. Hyperedge Fusion: Propagates and integrates hyperedge state from the previous layer via a multi-layer perceptron: Se,:(t+1)=MLP([Se,:(t);S~e,:(t+1)])S_{e,:}^{(t+1)} = \text{MLP}\left(\left[S_{e,:}^{(t)}; \tilde{S}_{e,:}^{(t+1)}\right]\right)

    3. Hyperedge2Node Attention: Aggregates updated hyperedge representations back to constituent cell nodes: Xv,:(t+1)=fE→V(Lv,S(t+1))X_{v,:}^{(t+1)} = f_{\mathcal{E} \to \mathcal{V}}\left(L_{v, S^{(t+1)}}\right) where Lv,S(t+1)={Se,:(t+1):v∈e}L_{v, S^{(t+1)}} = \{S_{e,:}^{(t+1)} : v \in e\} denotes the set of representations of hyperedges incident on node vv.

  3. Knowl 3 — Permutation-Invariant Set Multi-Head Attention in HyperAtt

    equation

    To preserve row and column permutation invariance over arbitrary sets of nodes or hyperedges, the attention block fV→Ef_{\mathcal{V} \to \mathcal{E}} (or fE→Vf_{\mathcal{E} \to \mathcal{V}}) in HyperTrans replaces standard pairwise self-attention with a set attention mechanism (SetMH):

    f(I)=LN(Y+FFN(Y))f(I) = \text{LN}(Y + \text{FFN}(Y))

    Y=LN(ω+SetMH(ω,I,I))Y = \text{LN}(\omega + \text{SetMH}(\omega, I, I))

    SetMH(ω,I,I)=∥i=1hOi\text{SetMH}(\omega, I, I) = \Vert_{i=1}^h O_i

    Oi=Softmax(ωi(IWiK)T)(IWiV)O_i = \text{Softmax}\left(\omega_i (I W_i^K)^T\right)(I W_i^V)

    where II represents the unordered set of input node or hyperedge representations, LN\text{LN} denotes Layer Normalization, FFN\text{FFN} is a position-wise feed-forward network, ∥\Vert represents vector concatenation across hh attention heads, ω=∥i=1hωi∈Rd\omega = \Vert_{i=1}^h \omega_i \in \mathbb{R}^d is a learnable query parameter vector, and WiK,WiVW_i^K, W_i^V are projection weight matrices for keys and values.

  4. Knowl 4 — Maximal Permutation Invariance of HYTREL Tabular Representations

    theoretical result

    Let a table T\mathcal{T} of size n×mn \times m have tasks whose target random variables remain unchanged under arbitrary actions of the direct product symmetric group Sn×SmS_n \times S_m acting independently on rows and columns. Let ϕ:T↦z∈Rd\phi: \mathcal{T} \mapsto z \in \mathbb{R}^d be a continuous target function satisfying ϕ(a⋅T)=ϕ(T)\phi(a \cdot \mathcal{T}) = \phi(\mathcal{T}) for all a∈Sn×Sma \in S_n \times S_m.

    Let g:B↦y∈Rkg: B \mapsto y \in \mathbb{R}^k be the HYTREL representation function operating on the hypergraph incidence matrix B∈{0,1}mn×(m+n+1)B \in \{0, 1\}^{mn \times (m+n+1)}.

    If there exists a bijective mapping between the space of tables and their corresponding hypergraph incidence matrices, then ϕ\phi is maximally invariant when modeled as g(B)g(B): ϕ(T1)=ϕ(T2)  ⟺  g(B1)=g(B2)\phi(\mathcal{T}_1) = \phi(\mathcal{T}_2) \iff g(B_1) = g(B_2) which holds if and only if T2\mathcal{T}_2 is identical to T1\mathcal{T}_1 up to independent permutations of rows and/or columns, where B1,B2B_1, B_2 are the hypergraph incidence matrices of T1,T2\mathcal{T}_1, \mathcal{T}_2.

  5. Knowl 5 — HYTREL Self-Supervised Pretraining Objectives

    model/method

    HYTREL is pretrained using two distinct self-supervised objectives:

    1. ELECTRA Head: 15%15\% of cell values and column headers in an input table are randomly replaced with values sampled based on frequency across the entire pretraining corpus. A binary classification head is trained with cross-entropy loss to predict whether each individual cell or header has been corrupted or preserved.

    2. Hypergraph Contrastive Head: For a given table hypergraph, two corrupted views are generated by randomly masking 30%30\% of node-hyperedge incidence connections. The two augmented views form a positive pair (q,k+)(q, k_+), and all other table views within the training batch serve as negative pairs {ki}i=1K\{k_i\}_{i=1}^K. The model optimizes the InfoNCE loss over table and column hyperedge representations: Lcontrastive=−log⁡exp⁡(q⋅k+/τ)∑i=0Kexp⁡(q⋅ki/τ)\mathcal{L}_{\text{contrastive}} = -\log \frac{\exp(q \cdot k_+ / \tau)}{\sum_{i=0}^K \exp(q \cdot k_i / \tau)} where τ=0.007\tau = 0.007 is the temperature hyperparameter.

  6. Knowl 6 — Downstream Performance of HYTREL on Tabular Understanding Benchmarks

    data/table

    Downstream evaluations were conducted across four tasks: Column Type Annotation (CTA) on TURL-CTA (Macro Precision/Recall/F1, %), Column Property Annotation (CPA) on TURL-CPA (Macro Precision/Recall/F1, %), Table Type Detection (TTD) on WDC Schema.org (Accuracy, %), and Table Similarity Prediction (TSP) on PMC (Macro Precision/Recall/F1 and Accuracy, %).

    Systems CTA (P/R/F1) CPA (P/R/F1) TTD (Acc) TSP (F1 / Acc)
    Sherlock 88.40 / 70.55 / 78.47 - - -
    BERTbase_{\text{base}} - 91.18 / 90.69 / 90.94 - -
    TURL + metadata 92.75 / 92.63 / 92.69 92.90 / 93.80 / 93.35 - -
    Doduo + metadata 93.25 / 92.34 / 92.79 91.20 / 94.50 / 92.82 - -
    TFIDF+GloVe+MLP - - - 84.47 / 85.06
    TabSim - - - 86.13 / 87.05
    TaBERTbase_{\text{base}} (K=1K=1) 91.40 / 89.49 / 90.43 92.31 / 90.42 / 91.36 93.11 86.18 / 87.35
    w/o Pretrain 90.00 / 85.50 / 87.70 89.74 / 68.74 / 77.84 85.04 40.30 / 63.45
    TaBERTbase_{\text{base}} (K=3K=3) 91.63 / 91.12 / 91.37 92.49 / 92.49 / 92.49 95.15 87.36 / 88.29
    w/o Pretrain 90.77 / 87.23 / 88.97 90.10 / 84.83 / 87.38 89.88 82.05 / 82.57
    HYTREL w/o Pretrain 92.92 / 92.50 / 92.71 92.85 / 91.50 / 92.17 93.84 87.30 / 88.38
    HYTREL w/ ELECTRA 92.85 / 94.21 / 93.53 92.88 / 94.07 / 93.48 95.81 87.32 / 88.29
    HYTREL w/ Contrastive 92.71 / 93.24 / 92.97 93.01 / 93.16 / 93.09 94.52 89.26 / 90.12

    HYTREL without pretraining achieves competitive performance comparable to fully pretrained TaBERT baselines. ELECTRA pretraining provides the best performance on CTA (93.53%93.53\% F1), CPA (93.48%93.48\% F1), and TTD (95.81%95.81\% accuracy), while contrastive pretraining provides the best performance on the cross-domain TSP biomedical benchmark (89.26%89.26\% F1, 90.12%90.12\% accuracy).

  7. Knowl 7 — Empirical Verification of Permutation Invariance and Representation Hierarchy

    empirical result

    When evaluated across 5,000 validation tables under row permutations, column permutations, and independent joint row-and-column permutations, the average L2L_2-norm distance between representations before and after permutations is 00 for HYTREL across all cells, rows, columns, and table hyperedges.

    In contrast, linearized models that rely on positional encodings (such as TaBERT with K=1K=1 and K=3K=3) exhibit non-zero L2L_2 distance shifts under all permutations, with column permutations causing a greater representation change than row permutations, and joint row-column permutations causing the largest degradation.

    t-SNE visual analysis confirms that HYTREL separates table elements into distinct structural clusters: cell representations group tightly and separate from column, row, and table hyperedge representations. ELECTRA pretraining yields complete cluster separation among cells, rows, columns, and tables, reflecting an explicit hierarchy.

  8. Knowl 8 — Performance Degradation from Excessive Invariance via Positional Encoding Removal

    empirical result

    True table permutation invariance requires independent row and column permutations (Sn×SmS_n \times S_m). Completely removing positional encodings from linearized tabular models (such as TaBERT) induces excessive permutation invariance (SmnS_{mn}), which treats the table as an unstructured set of cells and destroys row-column relationships.

    Evaluating TaBERT on Table Type Detection (TTD) demonstrates that removing positional encodings causes severe performance drops:

    • TaBERT (K=1K=1) with positional encodings achieves 94.11%94.11\% Dev / 93.44%93.44\% Test accuracy, but drops to 88.62%88.62\% Dev / 88.91%88.91\% Test accuracy without positional encodings.
    • TaBERT (K=3K=3) with positional encodings achieves 95.82%95.82\% Dev / 95.22%95.22\% Test accuracy, but drops to 92.80%92.80\% Dev / 92.44%92.44\% Test accuracy without positional encodings.
  9. Knowl 9 — Theoretical and Empirical Inference Complexity of HYTREL vs Linearized TaLMs

    theoretical result

    For a table with nn rows and mm columns (total cell count mnmn) and hidden dimension dd, the theoretical inference complexity per layer is:

    • Standard BERT-based Linearization:

      • Without linear transformation: O((mn)2⋅d)\mathcal{O}\left((mn)^2 \cdot d\right)
      • With linear transformation: O((mn)2⋅d+(mn)⋅d2)\mathcal{O}\left((mn)^2 \cdot d + (mn) \cdot d^2\right)
    • HYTREL SetMH Attention:

      • Without linear transformation: O(mnd)\mathcal{O}(mnd)
      • With linear transformation: O(mnd+(mn)⋅d2)\mathcal{O}\left(mnd + (mn) \cdot d^2\right)

    HYTREL reduces the self-attention computation from quadratic in table size (O((mn)2))(\mathcal{O}((mn)^2)) to linear (O(mn))(\mathcal{O}(mn)) by employing a single learnable query vector ω∈Rd\omega \in \mathbb{R}^d instead of an (mn)×d(mn) \times d query matrix.

    Empirically, on an NVIDIA A10 GPU with batch size 8, total inference times (including table graph construction) on validation sets are:

    • CTA (4,844 samples): TaBERT (K=3K=3) 99s vs HYTREL 21s
    • CPA (1,560 samples): TaBERT (K=3K=3) 21s vs HYTREL 7s
    • TTD (4,500 samples): TaBERT (K=3K=3) 131s vs HYTREL 33s
    • TSP (1,391 samples): TaBERT (K=3K=3) 93s vs HYTREL 18s
  10. Knowl 10 — Scalability to Arbitrary Table Sizes and Downsampling Dynamics

    data/table

    Unlike sequence-based tabular language models that are constrained by hard sequence position limits (e.g., 512 tokens), HYTREL accommodates arbitrarily large tables by expanding node sets connected to row/column hyperedges.

    Experiments on the Table Type Detection (TTD) dataset (average table length 157 rows) fine-tuned with HYTREL (ELECTRA) on an NVIDIA A10 GPU demonstrate accuracy and resource trade-offs across input size truncations:

    Size Limit (#rows, #cols) Dev Acc (%) Test Acc (%) Train Time (min/epoch) GPU Memory (GB/sample)
    (3, 2) 78.47 78.00 3 0.14
    (15, 10) 95.36 95.96 9 0.48
    (30, 20) 96.23 95.81 14 0.58
    (60, 40) 96.38 95.80 23 1.41
    (120, 80) 96.02 95.71 51 5.13
    (240, 160) 96.06 95.67 90 16.00

    Classification performance plateaus around size limits of (15,10)(15, 10) to (30,20)(30, 20) (95.81%95.81\% to 95.96%95.96\% test accuracy), showing that sub-sampling or downsampling large tables substantially reduces compute and memory overhead without degrading downstream prediction accuracy.

  11. Knowl 11 — Architectural and Structural Limitations of HYTREL

    limitation

    HYTREL possesses three key architectural and structural limitations:

    1. Encoder-Only Paradigm: It is designed solely as a tabular encoder and cannot directly perform joint tabular natural language tasks—such as table-based question answering (Table QA) or table-to-text generation—without introducing external text encoders or autoregressive decoders.
    2. Flat Column Schema Restriction: The hypergraph structure assumes flat single-level column headers and cannot natively model tables with nested or multi-level hierarchical column headers.
    3. Isolation from Cross-Table Relations: The formulation models each table as an isolated hypergraph and does not capture multi-table database interactions, cross-table foreign key links, or cross-table join operations.

Coverage note — None was omitted; the knowls comprehensively capture the hypergraph representation, HyperTrans architecture, SetMH attention equations, maximal invariance theoretical results, pretraining heads, downstream empirical benchmarks, permutation invariance/hierarchy analyses, excessive invariance ablation, complexity analysis, size scalability dynamics, and limitations.

References

  1. 1.Pengcheng Yin, Graham Neubig, Wen tau Yih, and Sebastian Riedel. TaBERT: Pretraining for joint understanding of textual and tabular data. In Annual Conference of the Association for Computational Linguistics (ACL), July 2020.
  2. 2.Jingfeng Yang, Aditya Gupta, Shyam Upadhyay, Luheng He, Rahul Goel, and Shachi Paul. Table-Former: Robust transformer modeling for table-text encoding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 528–537, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.40. URL https://aclanthology.org/2022.acl-long.40.
  3. 3.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Eisenschlos. TaPas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4320–4333, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.398. URL https://aclanthology.org/2020.acl-main.398.
  4. 4.Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. Turl: table understanding through representation learning. Proceedings of the VLDB Endowment, 14(3):307–319, 2020.
  5. 5.Hiroshi Iida, Dung Thai, Varun Manjunatha, and Mohit Iyyer. TABBIE: Pretrained representations of tabular data. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3446–3456, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.270. URL https://aclanthology.org/2021.naacl-main.270.
  6. 6.Zhiruo Wang, Haoyu Dong, Ran Jia, Jia Li, Zhiyi Fu, Shi Han, and Dongmei Zhang. Tuta: Tree-based transformers for generally structured table pre-training. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, KDD '21, page 1780–1790, New York, NY, USA, 2021a. Association for Computing Machinery. ISBN 9781450383325. doi: 10.1145/3447548.3467434. URL https://doi.org/10.1145/3447548.3467434.
  7. 7.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. ELECTRA: Pre-training text encoders as discriminators rather than generators. In ICLR, 2020. URL https://openreview.net/pdf?id=r1xMH1BtvB.
  8. 8.Tianxin Wei, Yuning You, Tianlong Chen, Yang Shen, Jingrui He, and Zhangyang Wang. Augmentations in hypergraph contrastive learning: Fabricated and generative. arXiv preprint arXiv:2210.03801, 2022.
  9. 9.Eli Chien, Chao Pan, Jianhao Peng, and Olgica Milenkovic. You are allset: A multiset function framework for hypergraph neural networks. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=hpBTIv2uy_E.
  10. 10.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
  11. 11.Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. ArXiv, abs/1607.06450, 2016.
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. doi: 10.1109/CVPR.2016.90.
  13. 13.Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola. Deep sets. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper/2017/file/f22e4747da1aa27e363d86d40ff442fe-Paper.pdf.
  14. 14.Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. Set transformer: A framework for attention-based permutation-invariant neural networks. In Proceedings of the 36th International Conference on Machine Learning, pages 3744–3753, 2019.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  16. 16.RI Tyshkevich and Vadim E Zverovich. Line hypergraphs. Discrete Mathematics, 161(1-3):265–283, 1996.
  17. 17.Clare Lyle, Mark van der Wilk, Marta Kwiatkowska, Yarin Gal, and Benjamin Bloem-Reddy. On the benefits of invariance in neural networks. arXiv preprint arXiv:2005.00178, 2020.
  18. 18.Aäron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. CoRR, abs/1807.03748, 2018. URL http://arxiv.org/abs/1807.03748.
  19. 19.Maryam Habibi, Johannes Starlinger, and Ulf Leser. Tabsim: A siamese neural network for accurate estimation of table similarity. pages 930–937, 12 2020. doi: 10.1109/BigData50022.2020.9378077.
  20. 20.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. TinyBERT: Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4163–4174, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.372. URL https://aclanthology.org/2020.findings-emnlp.372.
  21. 21.Yoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang, Çagatay Demiralp, Chen Chen, and Wang-Chiew Tan. Annotating Columns with Pre-trained Language Models. Association for Computing Machinery, 2022. ISBN 9781450392495. URL https://doi.org/10.1145/3514221.3517906.
  22. 22.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  23. 23.Xin Huang, Ashish Khetan, Milan Cvitkovic, and Zohar Karnin. Tabtransformer: Tabular data modeling using contextual embeddings. arXiv preprint arXiv:2012.06678, 2020.
  24. 24.Sercan Ö Arik and Tomas Pfister. Tabnet: Attentive interpretable tabular learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 6679–6687, 2021.
  25. 25.Gowthami Somepalli, Micah Goldblum, Avi Schwarzschild, C Bayan Bruss, and Tom Goldstein. Saint: Improved neural networks for tabular data via row attention and contrastive pre-training. arXiv preprint arXiv:2106.01342, 2021.
  26. 26.Yury Gorishniy, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. Revisiting deep learning models for tabular data. In NeurIPS, 2021.
  27. 27.Leo Grinsztajn, Edouard Oyallon, and Gael Varoquaux. Why do tree-based models still outperform deep learning on typical tabular data? In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. URL https://openreview.net/forum?id=Fp7__phQszn.
  28. 28.Zifeng Wang and Jimeng Sun. Transtab: Learning transferable tabular transformers across tables. In Advances in Neural Information Processing Systems, 2022.
  29. 29.Kounianhua Du, Weinan Zhang, Ruiwen Zhou, Yangkun Wang, Xilong Zhao, Jiarui Jin, Quan Gan, Zheng Zhang, and David Paul Wipf. Learning enhanced representations for tabular data via neighborhood propagation. In NeurIPS 2022, 2022. URL https://www.amazon.science/publications/learning-enhanced-representations-for-tabular-data-via-neighborhood-propagation.
  30. 30.Witold Wydmański, Oleksii Bulenok, and Marek Śmieja. Hypertab: Hypernetwork approach for deep learning on small tabular datasets. arXiv preprint arXiv:2304.03543, 2023.
  31. 31.Tennison Liu, Jeroen Berrevoets, Zhaozhi Qian, and Mihaela Van Der Schaar. Learning representations without compositional assumptions. In Proceedings of the 40th International Conference on Machine Learning, ICML'23. JMLR.org, 2023.
  32. 32.Julian Eisenschlos, Maharshi Gor, Thomas Müller, and William Cohen. MATE: Multi-view attention for table transformer efficiency. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7606–7619, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.600. URL https://aclanthology.org/2021.emnlp-main.600.
  33. 33.Sarthak Dash, Sugato Bagchi, Nandana Mihindukulasooriya, and Alfio Gliozzo. Permutation invariant strategy using transformer encoders for table understanding. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 788–800, 2022.
  34. 34.Thomas Mueller, Francesco Piccinno, Peter Shaw, Massimo Nicosia, and Yasemin Altun. Answering conversational questions on structured data without logical forms. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5902–5910, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1603. URL https://aclanthology.org/D19-1603.
  35. 35.Fei Wang, Kexuan Sun, Muhao Chen, Jay Pujara, and Pedro Szekely. Retrieving complex tables with multi-granular graph representation learning. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '21, page 1472–1482, New York, NY, USA, 2021b. Association for Computing Machinery. ISBN 9781450380379. doi: 10.1145/3404835.3462909. URL https://doi.org/10.1145/3404835.3462909.
  36. 36.Daheng Wang, Prashant Shiralkar, Colin Lockard, Binxuan Huang, Xin Luna Dong, and Meng Jiang. Tcn: Table convolutional network for web table interpretation. In Proceedings of the Web Conference 2021, pages 4020–4032, 2021c.
  37. 37.Peng Shi, Patrick Ng, Feng Nan, Henghui Zhu, Jun Wang, Jiarong Jiang, Alexander Hanbo Li, Rishav Chakravarti, Donald Weidner, Bing Xiang, et al. Generation-focused table-based intermediate pre-training for free-form question answering. 2022.
  38. 38.Jonathan Herzig, Thomas Müller, Syrine Krichene, and Julian Martin Eisenschlos. Open domain question answering over tables via dense retrieval. arXiv preprint arXiv:2103.12011, 2021.
  39. 39.Michael Glass, Mustafa Canim, Alfio Gliozzo, Saneem Chemmengath, Vishwajeet Kumar, Rishav Chakravarti, Avi Sil, Feifei Pan, Samarth Bharadwaj, and Nicolas Rodolfo Fauceglia. Capturing row and column semantics in transformer based question answering over tables. arXiv preprint arXiv:2104.08303, 2021.
  40. 40.Ankur P Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. Totto: A controlled table-to-text generation dataset. arXiv preprint arXiv:2004.14373, 2020.
  41. 41.Ori Yoran, Alon Talmor, and Jonathan Berant. Turning tables: Generating examples from semi-structured tables for endowing language models with reasoning skills. arXiv preprint arXiv:2107.07261, 2021.
  42. 42.Fei Wang, Zhewei Xu, Pedro Szekely, and Muhao Chen. Robust (controlled) table-to-text generation with structure-aware equivariance learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5037–5048, Seattle, United States, July 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.371. URL https://aclanthology.org/2022.naacl-main.371.
  43. 43.Ewa Andrejczuk, Julian Martin Eisenschlos, Francesco Piccinno, Syrine Krichene, and Yasemin Altun. Table-to-text generation and pre-training with tabt5. arXiv preprint arXiv:2210.09162, 2022.
  44. 44.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  45. 45.Sameer Agarwal, Kristin Branson, and Serge Belongie. Higher order learning with graphs. In Proceedings of the 23rd international conference on Machine learning, pages 17–24, 2006.
  46. 46.Naganand Yadati, Madhav Nimishakavi, Prateek Yadav, Vikram Nitin, Anand Louis, and Partha Talukdar. Hypergcn: A new method for training graph convolutional networks on hypergraphs. Advances in neural information processing systems, 32, 2019.
  47. 47.Devanshu Arya, Deepak K Gupta, Stevan Rudinac, and Marcel Worring. Hypersage: Generalizing inductive representation learning on hypergraphs. arXiv preprint arXiv:2010.04558, 2020.
  48. 48.László Babai, Paul Erdo˝s, and Stanley M. Selkow. Random graph isomorphism. SIAM Journal on Computing, 9(3):628–635, 1980. doi: 10.1137/0209047. URL https://doi.org/10.1137/0209047.
  49. 49.Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks? arXiv preprint arXiv:1810.00826, 2018.
  50. 50.Balasubramaniam Srinivasan, Da Zheng, and George Karypis. Learning over families of sets—hypergraph representation learning for higher order tasks. In Proceedings of the 2021 SIAM International Conference on Data Mining (SDM), pages 756–764. SIAM, 2021.
  51. 51.Mohamed Trabelsi, Zhiyu Chen, Shuo Zhang, Brian D Davison, and Jeff Heflin. StruBERT: Structure-aware bert for table search and matching. In Proceedings of the ACM Web Conference, WWW '22, 2022.
  52. 52.Madelon Hulsebos, Kevin Hu, Michiel Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çagatay Demiralp, and César Hidalgo. Sherlock: A deep learning approach to semantic data type detection. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. ACM, 2019.
  53. 53.Jeffrey Pennington, Richard Socher, and Christopher Manning. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar, October 2014. Association for Computational Linguistics. doi: 10.3115/v1/D14-1162. URL https://aclanthology.org/D14-1162.

Citation

MLA
Chen, P., et al. “HyTrel: Hypergraph-enhanced Tabular Data Representation Learning”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 32173–93, https://proceedings.neurips.cc/paper_files/paper/2023/file/66178beae8f12fcd48699de95acc1152-Paper-Conference.pdf.
APA
Chen, P., Sarkar, S., Lausen, L., Srinivasan, B., Zha, S., Huang, R., & Karypis, G. (2023). HyTrel: Hypergraph-enhanced Tabular Data Representation Learning. Advances in Neural Information Processing Systems, 36, 32173–32193. https://proceedings.neurips.cc/paper_files/paper/2023/file/66178beae8f12fcd48699de95acc1152-Paper-Conference.pdf
Chicago
Chen, P., S. Sarkar, L. Lausen, et al. 2023. “HyTrel: Hypergraph-enhanced Tabular Data Representation Learning”. Advances in Neural Information Processing Systems 36: 32173–93. https://proceedings.neurips.cc/paper_files/paper/2023/file/66178beae8f12fcd48699de95acc1152-Paper-Conference.pdf.
Harvard
Chen, P. et al. (2023) “HyTrel: Hypergraph-enhanced Tabular Data Representation Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 32173–32193. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/66178beae8f12fcd48699de95acc1152-Paper-Conference.pdf.
Vancouver
1. Chen P, Sarkar S, Lausen L, Srinivasan B, Zha S, Huang R, Karypis G (2023) HyTrel: Hypergraph-enhanced Tabular Data Representation Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 32173–32193

BibTeX

@inproceedings{chen2023hytrel,
  title = {HyTrel: Hypergraph-enhanced Tabular Data Representation Learning},
  author = {Chen, Pei and Sarkar, Soumajyoti and Lausen, Leonard and Srinivasan, Balasubramaniam and Zha, Sheng and Huang, Ruihong and Karypis, George},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {32173-32193},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/66178beae8f12fcd48699de95acc1152-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/