Graphix-T5: Mixing Pre-trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing

Jinyang LiBinyuan HuiReynold ChengBowen QinChenhao MaNan HuoFei HuangWenyu DuLuo SiYongbin Li

article2023AAAI161 citations

Proposes Graphix-T5, a text-to-SQL architecture that integrates relational graph neural network layers directly into pre-trained T5 encoder blocks to improve multi-hop reasoning over database schemas, achieving state-of-the-art cross-domain parsing accuracy while outperforming much larger standard models.

Listen

Non-technical users face significant barriers when querying relational databases because formulating Structured Query Language (SQL) statements requires specialized expertise. While automated text-to-SQL systems aim to translate plain language questions into executable database queries, existing models struggle to generalize across unfamiliar database domains. State-of-the-art language models, such as the text-to-text transformer model known as T5, excel at understanding natural language but often fail to capture the complex, multi-hop relational structures inherent in database schemas.

The article introduces and evaluates Graphix-T5, a novel architecture designed to augment standard pre-trained language models with graph-aware structural reasoning capabilities. The primary objective is to demonstrate that integrating structural graph mechanisms directly into the language model improves the accuracy and robustness of cross-domain database querying without losing the foundational strengths of pre-trained language understanding.

To achieve this, the authors designed specialized Graphix layers that combine standard transformer self-attention with relational graph neural network blocks. These layers explicitly map relationships among question tokens, tables, and columns, while employing an efficient "bridge node" strategy to connect questions and schemas without overwhelming the network with excessive noisy links. Crucially, the authors stacked these layers into the model encoder to maintain deep, layer-by-layer interactions between semantics and structure, rather than inserting an isolated graph module that disrupts model information flow. The approach was evaluated on four benchmark datasets (Spider, SYN, DK, and Realistic) across varying query difficulties and compositional splits using exact match and execution accuracy metrics.

The experimental findings show substantial performance improvements across all tested benchmarks. Most notably, the mid-sized Graphix-T5-large model achieved a 5.7 percentage point gain in exact match accuracy (reaching 72.7%) and a 6.6 percentage point gain in execution accuracy (reaching 75.9%) over the standard T5-large baseline on the Spider benchmark. This enhanced mid-sized model even surpassed the much larger baseline model (T5-3B) by 1.2% in exact match accuracy. When combined with constrained decoding techniques, Graphix-T5-3B established a new state-of-the-art result of 81.0% execution accuracy. The model demonstrated especially large gains on complex, multi-hop reasoning tasks, synonym variations, and queries missing explicit schema mentions, while ablation experiments confirmed that placing graph layers strictly in the encoder avoids the catastrophic forgetting observed in alternate architectures.

These findings indicate that structural inductive bias is critical for deploying reliable natural language interfaces over enterprise databases. By enabling smaller models to outperform significantly larger language models, this architectural approach can reduce computational and operational costs while delivering higher accuracy on complex queries. Furthermore, the robust performance under realistic and synonym-perturbed settings reduces the risk of systems generating syntactically valid but structurally incorrect queries that return flawed business data.

Organizations developing database query interfaces should consider integrating graph-aware structural layers into encoder architectures rather than relying purely on sequence-to-sequence scaling or disjoint graph modules. Where applicable, pairing this model with constrained decoding mechanisms yields the highest execution reliability. Further research is recommended to explore structural grounding across broader low-resource database settings and multi-turn conversational querying.

The conclusions are supported by thorough evaluations on standard industry benchmarks and isolated ablation studies. However, confidence should be tempered by the fact that evaluation was conducted primarily on curated academic benchmarks and single-turn query settings, meaning additional piloting in unstructured production environments and highly customized enterprise databases is advisable before widespread operational deployment.

arXiv: 2301.07507
Cover for Graphix-T5: Mixing Pre-trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing

Abstract

The task of text-to-SQL parsing, which aims at converting natural language questions into executable SQL queries, has garnered increasing attention in recent years. One of the major challenges in text-to-SQL parsing is domain generalization, i.e., how to generalize well to unseen databases. Recently, the pre-trained text-to-text transformer model, namely T5, though not specialized for text-to-SQL parsing, has achieved state-of-the-art performance on standard benchmarks targeting domain generalization. In this work, we explore ways to further augment the pre-trained T5 model with specialized components for text-to-SQL parsing. Such components are expected to introduce structural inductive bias into text-to-SQL parsers thus improving model's capacity on (potentially multi-hop) reasoning, which is critical for generating structure-rich SQLs. To this end, we propose a new architecture GRAPHIX-T5, a mixed model with the standard pre-trained transformer model augmented by specially-designed graph-aware layers. Extensive experiments and analysis demonstrate the effectiveness of GRAPHIX-T5 across four text-to-SQL benchmarks: SPIDER, SYN, REALISTIC and DK. GRAPHIX-T5 surpass all other T5-based parsers with a significant margin, achieving new state-of-the-art performance. Notably, GRAPHIX-T5-large reaches performance superior to the original T5-large by 5.7% on exact match (EM) accuracy and 6.6% on execution accuracy (EX). This even outperforms the T5-3B by 1.2% on EM and 1.5% on EX.

Table of Contents

  • 1 Introduction
  • 2 Task Formulation and Notations
  • 2.1 Task Definition
  • 2.2 Vanilla T5 Architecture
  • Model Inputs
  • Encoder-Decoder Training Mechanism
  • 3 Proposed Approach: Graphix-T5
  • 3.1 Model Inputs
  • Contextual Encoding
  • Graph Construction
  • No-Match Mode vs. Bridge Mode
  • 3.2 Graphix-Layer
  • 3.3 Graphix-T5
  • 3.4 Training
  • 4 Experiment
  • 4.1 Set up
  • 4.2 Overall Performance
  • 4.3 Ablation Study
  • 4.4 Case Study
  • 5 Related Works
  • 6 Conclusion
  • References
  • A Fine-grained Syntax Relations
  • B Leadboard Result

Knowls

  1. Knowl 1 — GRAPHIX-T5 Architecture for Cross-Domain Text-to-SQL Parsing

    model/method

    GRAPHIX-T5 is an encoder-decoder architecture for cross-domain text-to-SQL semantic parsing that integrates relational graph neural networks directly into each layer of a pre-trained sequence-to-sequence Transformer (specifically T5).

    In standard text-to-SQL parsing with T5, natural language questions and relational database schemas are linearized into a single token sequence xx, and the model is trained to output SQL tokens yy. GRAPHIX-T5 replaces the standard T5 encoder layers with a stack of GRAPHIX layers while preserving the pre-trained autoregressive T5 decoder. Each GRAPHIX layer contains two parallel processing components:

    1. A contextual semantic Transformer block initialized with the corresponding layer weights Θ\Theta of the pre-trained T5 encoder.
    2. A Relational Graph Attention Network (RGAT) block parameterized by Ψ\Psi (initialized randomly) operating over a question-schema heterogeneous graph G=⟨V,R⟩G = \langle V, R \rangle.

    The final representation of each layer is formed by fusing the contextualized semantic representation with the structural node representations. The encoder output connects to the pre-trained T5 decoder parameterized by Υ\Upsilon. The model is trained by maximizing the conditional log-likelihood:

    max⁡Θ,Υ,Ψ∑i=1∣y∣log⁡pΘ,Υ,Ψ(yi∣y1:i−1,x,G)\max_{\Theta, \Upsilon, \Psi} \sum_{i=1}^{|y|} \log p_{\Theta, \Upsilon, \Psi}(y_i \mid y_{1:i-1}, x, G)

    GRAPHIX layers are restricted to the encoder; adding relational graph layers to the autoregressive decoder disrupts causal masking because global graph linking provides representations of future tokens.

  2. Knowl 2 — Mathematical Formulation of the GRAPHIX Layer

    equation

    Let HS(l)={h1(l),…,hN(l)}H^{(l)}_S = \{h^{(l)}_1, \dots, h^{(l)}_N\} denote the input hidden representations to the ll-th GRAPHIX layer for a sequence of length NN. The semantic block applies Multi-Head Attention (MHA) and a Feed-Forward Network (FFN) initialized from T5:

    H^S(l)=MHA(HS(l))=Concat(head1,…,headh)WO\hat{H}^{(l)}_S = \text{MHA}(H^{(l)}_S) = \text{Concat}(\text{head}_1, \dots, \text{head}_h) W^O

    headi=softmax(QWiQ(KWiK)⊤dk)VWiV\text{head}_i = \text{softmax}\left(\frac{Q W_i^Q (K W_i^K)^\top}{\sqrt{d_k}}\right) V W_i^V

    H~S(l)=LayerNorm(H^S(l)+FFN(H^S(l)))\tilde{H}^{(l)}_S = \text{LayerNorm}\left(\hat{H}^{(l)}_S + \text{FFN}(\hat{H}^{(l)}_S)\right)

    where WiQ,WiK∈Rdm×dkW_i^Q, W_i^K \in \mathbb{R}^{d_m \times d_k}, WiV∈Rdm×dvW_i^V \in \mathbb{R}^{d_m \times d_v}, WO∈Rdmh×dmW^O \in \mathbb{R}^{d_m h \times d_m}, and dk=dv=dm/hd_k = d_v = d_m / h.

    In parallel, the structural representation is computed via Relational Graph Attention Network (RGAT) over the input node embeddings eiinite^{\text{init}}_i (initialized from the layer's semantic representations) and neighbor set Nei\mathcal{N}_{e_i}:

    α⃗ij=eiinitW~Q(ejinitW~K+ϕ(rij))⊤dz\vec{\alpha}_{ij} = \frac{e^{\text{init}}_i \tilde{W}^Q (e^{\text{init}}_j \tilde{W}^K + \phi(r_{ij}))^\top}{\sqrt{d_z}}

    αij=softmaxj(α⃗ij)=exp⁡(α⃗ij)∑k∈Neiexp⁡(α⃗ik)\alpha_{ij} = \text{softmax}_j(\vec{\alpha}_{ij}) = \frac{\exp(\vec{\alpha}_{ij})}{\sum_{k \in \mathcal{N}_{e_i}} \exp(\vec{\alpha}_{ik})}

    e^iinit=∑j∈Neiαij(ejinitW~V+ϕ(rij))\hat{e}^{\text{init}}_i = \sum_{j \in \mathcal{N}_{e_i}} \alpha_{ij} (e^{\text{init}}_j \tilde{W}^V + \phi(r_{ij}))

    e^i(l)=LayerNorm(eiinit+e^iinitW~O)\hat{e}^{(l)}_i = \text{LayerNorm}(e^{\text{init}}_i + \hat{e}^{\text{init}}_i \tilde{W}^O)

    e~i(l)=LayerNorm(e^i(l)+FFN(e^i(l)))\tilde{e}^{(l)}_i = \text{LayerNorm}(\hat{e}^{(l)}_i + \text{FFN}(\hat{e}^{(l)}_i))

    where W~Q,W~K,W~V,W~O∈Rd×d\tilde{W}^Q, \tilde{W}^K, \tilde{W}^V, \tilde{W}^O \in \mathbb{R}^{d \times d}, ϕ(rij)\phi(r_{ij}) is a learned dd-dimensional embedding for relation rijr_{ij}, and dzd_z is the scaling dimension. The structural output is collected as E~G(l)={e~1(l),…,e~N(l)}\tilde{E}^{(l)}_G = \{\tilde{e}^{(l)}_1, \dots, \tilde{e}^{(l)}_N\}.

    The combined representation output by the ll-th GRAPHIX layer is defined as the element-wise sum:

    H~M(l)=H~S(l)+E~G(l)\tilde{H}^{(l)}_M = \tilde{H}^{(l)}_S + \tilde{E}^{(l)}_G

  3. Knowl 3 — Heterogeneous Graph Construction and Bridge Node Mechanism

    model/method

    Given a natural language question Q={q1,…,q∣Q∣}Q = \{q_1, \dots, q_{|Q|}\} and database schema D=⟨C,T⟩D = \langle C, T \rangle with columns CC and tables TT, the unified input sequence is formatted as:

    x=[q1,…,q∣Q∣∣Dname∣t1:c1t1,…,c∣C∣t1∣⋯∣t∣T∣:c1t∣T∣,…,c∣C∣t∣T∣∣∗]x = [q_1, \dots, q_{|Q|} \mid D_{name} \mid t_1 : c^{t_1}_1, \dots, c^{t_1}_{|C|} \mid \dots \mid t_{|T|} : c^{t_{|T|}}_1, \dots, c^{t_{|T|}}_{|C|} \mid *]

    where ∗* is a special column token and DnameD_{name} is the database name.

    A heterogeneous graph G=⟨V,R⟩G = \langle V, R \rangle is built over the node set V=Q∪C∪TV = Q \cup C \cup T using three relation families:

    1. Schema relations: FOREIGN-KEY, PRIMARY-KEY, and SAME-TABLE capturing explicit database constraints.
    2. Schema linking relations: EXACT-MATCH, PARTIAL-MATCH, VALUE-MATCH, and BRIDGE.
    3. Question relations: MODIFIER and ARGUMENT capturing dependency relationships between question tokens.

    To handle question tokens and schema items that are semantically related but lack lexical matching, previous methods instantiate dense dummy NO-MATCH edges between all unmatched pairs, resulting in A×BA \times B edges for AA question tokens and BB schema items and causing over-smoothing. The Bridge Node mechanism connects the unlinked question tokens and schema nodes through the special token node ∗* as an intermediate hub. This reduces the number of required edges from A×BA \times B to A+BA + B, enabling multi-hop structural grounding while avoiding noisy neighborhood over-smoothing.

  4. Knowl 4 — Cross-Domain Text-to-SQL Parsing Performance on SPIDER

    data/table

    The table below reports Exact Match (EM) and Execution Accuracy (EX) percentages on the official SPIDER development set across baseline text-to-SQL models, vanilla T5 models, and GRAPHIX-T5 variants. Models marked with ♡\heartsuit do not predict values, and models marked with ♣\clubsuit use the PICARD constrained decoding procedure.

    Model EM (%) EX (%)
    RAT-SQL + BERT ♡\heartsuit 69.7 -
    RAT-SQL + Grappa ♡\heartsuit 73.9 -
    GAZP + BERT 59.1 59.2
    BRIDGE v2 + BERT 70.0 68.3
    NatSQL + GAP 73.7 75.0
    SMBOP + GRAPPA 74.7 75.0
    LGESQL + ELECTRA ♡\heartsuit 75.1 -
    S2^2SQL + ELECTRA ♡\heartsuit 76.4 -
    T5-large 67.0 69.3
    GRAPHIX-T5-large 72.7 75.9
    T5-large + PICARD ♣\clubsuit 69.1 72.9
    GRAPHIX-T5-large + PICARD ♣\clubsuit 76.6 80.5
    T5-3B 71.5 74.4
    GRAPHIX-T5-3B 75.6 78.2
    T5-3B + PICARD ♣\clubsuit 75.5 79.3
    GRAPHIX-T5-3B + PICARD ♣\clubsuit 77.1 81.0

    GRAPHIX-T5-large achieves +5.7% EM and +6.6% EX over vanilla T5-large. When combined with PICARD constrained decoding, GRAPHIX-T5-large reaches 76.6% EM and 80.5% EX, outperforming the larger vanilla T5-3B + PICARD (75.5% EM / 79.3% EX). GRAPHIX-T5-3B + PICARD achieves 77.1% EM and 81.0% EX.

  5. Knowl 5 — Robustness and Zero-Shot Evaluation on SYN, DK, and REALISTIC Benchmarks

    data/table

    The table below presents Exact Match (EM) accuracy percentages under zero-shot evaluation on three challenging robustness benchmarks: SPIDER-SYN (synonym substitutions in questions/schemas), SPIDER-DK (requiring domain knowledge reasoning), and REALISTIC (removal/alteration of explicit schema mentions).

    Model SYN (%) DK (%) REALISTIC (%)
    GNN 23.6 26.0 -
    IRNet 28.4 33.1 -
    RAT-SQL 33.6 35.8 -
    RAT-SQL + BERT 48.2 40.9 58.1
    RAT-SQL + Grappa 49.1 38.5 59.3
    LGESQL + ELECTRA 64.6 48.4 69.2
    T5-large 53.6 40.0 58.5
    GRAPHIX-T5-large 61.1 48.6 67.3
    T5-3B 58.0 46.9 62.0
    GRAPHIX-T5-3B 66.9 51.2 72.4

    GRAPHIX-T5-large improves upon vanilla T5-large by +7.5% on SYN, +8.6% on DK, and +8.8% on REALISTIC. GRAPHIX-T5-3B improves upon vanilla T5-3B by +8.9% on SYN, +4.3% on DK, and +10.4% on REALISTIC, establishing superior generalization when lexical surface cues are perturbed or absent.

  6. Knowl 6 — Compositional Generalization on SPIDER-SSP Benchmark

    data/table

    The table below shows Exact Match (EM) accuracy percentages on the SPIDER-SSP compositional generalization benchmark across three structural splits: Template (split by distinct query templates), Length (split by query length), and TMCD (Target Maximum Compound Divergence).

    Model Template (%) Length (%) TMCD (%)
    T5-base 59.3 49.0 60.9
    T5-3B 64.8 56.7 69.6
    NQG-T5-3B 64.7 56.7 69.5
    GRAPHIX-T5-3B 70.1 60.6 73.8

    While natural language grammar-augmented models like NQG-T5-3B fail to outperform vanilla T5-3B, GRAPHIX-T5-3B achieves absolute improvements of +5.4% on Template, +3.9% on Length, and +4.3% on TMCD over vanilla T5-3B, confirming that injecting structural relational inductive bias aids compositional SQL generalization.

  7. Knowl 7 — Architectural Ablation Studies on Graph Placement and Linking Strategy

    data/table

    Ablation experiments evaluated on the SPIDER development set using the T5-large base illustrate the impact of graph module placement and schema-linking designs:

    Configuration EM (%) EX (%)
    (a) RAT-SQL + BERT 69.7 -
    (b) Vanilla T5-large 67.0 69.3
    (c) GNN-T5-large (Encoder →\rightarrow GNN →\rightarrow Decoder) 51.6 54.5
    (d) GRAPHIX-T5-large (BRIDGE Mode) 72.7 75.9
    w/ NO-MATCH Mode 71.1 74.2
    w/ DOUBLE-GRAPH (Encoder + Decoder GNN) 72.0 74.7

    Key architectural conclusions include:

    1. Catastrophic forgetting in severed pipelines: Placing an isolated GNN between the T5 encoder and T5 decoder (GNN-T5) drops performance by ~20% compared to GRAPHIX-T5 (51.6% vs 72.7% EM) and maintains 0% accuracy over the initial thousands of training steps due to severed contextual information flow.
    2. Decoder GNN degradation: Adding GRAPHIX layers to both encoder and decoder (DOUBLE-GRAPH) decreases EM accuracy by 0.7% (72.0% vs 72.7%) because global graph attention leaks future token context into the autoregressive decoder.
    3. Linking efficiency: Using the BRIDGE token ∗* outperforms the quadratic NO-MATCH linking strategy by 1.6% EM (72.7% vs 71.1%) by avoiding over-smoothing from noisy dense edges.
  8. Knowl 8 — SQL Generation Performance by Query Difficulty

    data/table

    The table below breaks down Exact Match (EM) accuracy percentages across query difficulty levels (Easy, Medium, Hard, Extra-hard, and All) on the SPIDER, SYN, and REALISTIC datasets for T5 and GRAPHIX-T5 models.

    Model SPIDER SYN REALISTIC
    Easy Med Hard Extra All Easy Med Hard Extra All Easy Med Hard Extra All
    T5-large 85.5 70.9 55.2 41.6 67.0 69.0 56.8 46.3 30.2 53.6 79.8 68.0 44.4 28.9 58.5
    GRAPHIX-T5-large 89.9 78.7 59.8 44.0 72.6 75.8 67.5 50.6 33.1 61.1 88.1 77.3 50.5 40.2 67.3
    T5-3B 89.5 78.3 58.6 40.4 71.6 74.2 64.5 48.0 27.8 58.0 85.3 73.4 46.5 27.8 62.0
    GRAPHIX-T5-3B 91.9 81.6 61.5 50.0 75.6 80.6 73.1 52.9 44.6 66.9 93.6 85.7 52.5 41.2 72.4

    The performance gains of GRAPHIX-T5 over vanilla T5 are particularly pronounced on Hard and Extra-hard query subsets (e.g., on SPIDER Extra-hard, GRAPHIX-T5-3B achieves 50.0% vs 40.4% for T5-3B, a +9.6% absolute gain; on REALISTIC Extra-hard, 41.2% vs 27.8%, a +13.4% gain), illustrating that multi-hop structural reasoning is most critical for complex queries.

Coverage note — Qualitative case study examples illustrating specific SQL predictions from Figure 5 were omitted as their conceptual points (multi-hop path grounding and prevention of hallucinated column names) are fully captured in the methodology and quantitative knowls.

References

  1. 1.Bogin, B.; Berant, J.; and Gardner, M. 2019. Representing Schema Structure with Graph Neural Networks for Text-to-SQL Parsing. In Proc. of ACL.
  2. 2.Cai, R.; Xu, B.; Zhang, Z.; Yang, X.; Li, Z.; and Liang, Z. 2018. An Encoder-Decoder Framework Translating Natural Language to Database Queries. In Proc. of IJCAI.
  3. 3.Cai, R.; Yuan, J.; Xu, B.; and Hao, Z. 2021. SADGA: Structure-Aware Dual Graph Aggregation Network for Text-to-SQL. In Proc. of NeurIPS.
  4. 4.Cai, Z.; Li, X.; Hui, B.; Yang, M.; Li, B.; Li, B.; Cao, Z.; Li, W.; Huang, F.; Si, L.; and Li, Y. 2022. STAR: SQL Guided Pre-Training for Context-dependent Text-to-SQL Parsing. In Proc. of EMNLP Findings.
  5. 5.Cao, R.; Chen, L.; Chen, Z.; Zhao, Y.; Zhu, S.; and Yu, K. 2021. LGESQL: Line Graph Enhanced Text-to-SQL Model with Mixed Local and Non-Local Relations. In Proc. of ACL.
  6. 6.Chen, D.; Lin, Y.; Li, W.; Li, P.; Zhou, J.; and Sun, X. 2020a. Measuring and Relieving the Over-Smoothing Problem for Graph Neural Networks from the Topological View. In Proc. of AAAI.
  7. 7.Chen, X.; Meng, F.; Li, P.; Chen, F.; Xu, S.; Xu, B.; and Zhou, J. 2020b. Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue Generation. In Proc. of EMNLP.
  8. 8.Chen, Z.; Chen, L.; Zhao, Y.; Cao, R.; Xu, Z.; Zhu, S.; and Yu, K. 2021. ShadowGNN: Graph Projection Neural Network for Text-to-SQL Parser. In Proc. of NAACL.
  9. 9.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proc. of NAACL.
  10. 10.Fang, Y.; Sun, S.; Gan, Z.; Pillai, R.; Wang, S.; and Liu, J. 2020. Hierarchical Graph Network for Multi-hop Question Answering. In Proc. of EMNLP.
  11. 11.French, R. M. 1999. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences.
  12. 12.Gan, Y.; Chen, X.; Huang, Q.; Purver, M.; Woodward, J. R.; Xie, J.; and Huang, P. 2021a. Towards Robustness of Text-to-SQL Models against Synonym Substitution. In Proc. of ACL.
  13. 13.Gan, Y.; Chen, X.; and Purver, M. 2021. Exploring Under-explored Limitations of Cross-Domain Text-to-SQL Generalization. In Proc. of EMNLP.
  14. 14.Gan, Y.; Chen, X.; Xie, J.; Purver, M.; Woodward, J. R.; Drake, J.; and Zhang, Q. 2021b. Natural SQL: Making SQL Easier to Infer from Natural Language Specifications. In Proc. of EMNLP Findings.
  15. 15.Guo, J.; Zhan, Z.; Gao, Y.; Xiao, Y.; Lou, J.-G.; Liu, T.; and Zhang, D. 2019. Towards Complex Text-to-SQL in Cross-Domain Database with Intermediate Representation. In Proc. of ACL.
  16. 16.Hui, B.; Geng, R.; Ren, Q.; Li, B.; Li, Y.; Sun, J.; Huang, F.; Si, L.; Zhu, P.; and Zhu, X. 2021a. Dynamic Hybrid Relation Network for Cross-Domain Context-Dependent Semantic Parsing. In Proc. of AAAI.
  17. 17.Hui, B.; Geng, R.; Wang, L.; Qin, B.; Li, Y.; Li, B.; Sun, J.; and Li, Y. 2022. S2^2SQL: Injecting Syntax to Question-Schema Interaction Graph Encoder for Text-to-SQL Parsers. In Proc. of ACL Findings.
  18. 18.Hui, B.; Shi, X.; Geng, R.; Li, B.; Li, Y.; Sun, J.; and Zhu, X. 2021b. Improving Text-to-SQL with Schema Dependency Learning. In arXiv:2103.04399.
  19. 19.Qin, B.; Hui, B.; Wang, L.; Yang, M.; Li, J.; Li, B.; Geng, R.; Cao, R.; Sun, J.; Si, L.; Huang, F.; and Li, Y. 2022a. A Survey on Text-to-SQL Parsing: Concepts, Methods, and Future Directions. In arXiv:2208.13629.
  20. 20.Qin, B.; Wang, L.; Hui, B.; Geng, R.; Cao, Z.; Yang, M.; Sun, J.; and Li, Y. 2022b. Linking-Enhanced Pre-Training for Table Semantic Parsing. In arXiv:2111.09486.
  21. 21.Qin, B.; Wang, L.; Hui, B.; Li, B.; Wei, X.; Li, B.; Huang, F.; Si, L.; Yang, M.; and Li, Y. 2022c. SUN: Exploring Intrinsic Uncertainties in Text-to-SQL Parsers. In Proc. of COLING.
  22. 22.Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research.
  23. 23.Rubin, O.; and Berant, J. 2021. SmBoP: Semi-autoregressive Bottom-up Semantic Parsing. In Proc. of NAACL.
  24. 24.Scholak, T.; Schucher, N.; and Bahdanau, D. 2021. PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models. In Proc. of EMNLP.
  25. 25.Shaw, P.; Chang, M.-W.; Pasupat, P.; and Toutanova, K. 2021. Compositional Generalization and Natural Language Variation: Can a Semantic Parsing Approach Handle Both? In Proc. of ACL.
  26. 26.Shazeer, N.; and Stern, M. 2018. Adafactor: Adaptive Learning Rates with Sublinear Memory Cost. In Proc. of ICML.
  27. 27.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention is All you Need. In Proc. of NeurIPS.
  28. 28.Velickovic, P.; Cucurull, G.; Casanova, A.; Romero, A.; Lio,` P.; and Bengio, Y. 2018. Graph Attention Networks. In Proc. of ICLR.
  29. 29.Wang, B.; Shin, R.; Liu, X.; Polozov, O.; and Richardson, M. 2020a. RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers. In Proc. of ACL.
  30. 30.Wang, K.; Shen, W.; Yang, Y.; Quan, X.; and Wang, R. 2020b. Relational Graph Attention Network for Aspect-based Sentiment Analysis. In Proc. of ACL.
  31. 31.Wang, L.; Qin, B.; Hui, B.; Li, B.; Yang, M.; Wang, B.; Li, B.; Huang, F.; Si, L.; and Li, Y. 2022. Proton: Probing Schema Linking Information from Pre-trained Language Models for Text-to-SQL Parsing. In Proc. of KDD.
  32. 32.Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Le Scao, T.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. 2020. Transformers: State-of-the-Art Natural Language Processing. In Proc. of EMNLP.
  33. 33.Xie, T.; Wu, C. H.; Shi, P.; Zhong, R.; Scholak, T.; Yasunaga, M.; Wu, C.-S.; Zhong, M.; Yin, P.; Wang, S. I.; Zhong, V.; Wang, B.; Li, C.; Boyle, C.; Ni, A.; Yao, Z.; Radev, D.; Xiong, C.; Kong, L.; Zhang, R.; Smith, N. A.; Zettlemoyer, L.; and Yu, T. 2022. UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models. ArXiv preprint.
  34. 34.Xu, X.; Liu, C.; and Song, D. 2017. Sqlnet: Generating structured queries from natural language without reinforcement learning. ArXiv preprint.
  35. 35.Yaghmazadeh, N.; Wang, Y.; Dillig, I.; and Dillig, T. 2017. SQLizer: query synthesis from natural language. Proceedings of the ACM on Programming Languages.
  36. 36.Yu, T.; Li, Z.; Zhang, Z.; Zhang, R.; and Radev, D. 2018a. TypeSQL: Knowledge-Based Type-Aware Neural Text-to-SQL Generation. In Proc. of NAACL.
  37. 37.Yu, T.; Zhang, R.; Polozov, A.; Meek, C.; and Awadallah, A., Hassan. 2021. SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing. In Proc. of ICLR.
  38. 38.Yu, T.; Zhang, R.; Yang, K.; Yasunaga, M.; Wang, D.; Li, Z.; Ma, J.; Li, I.; Yao, Q.; Roman, S.; Zhang, Z.; and Radev, D. 2018b. Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. In Proc. of EMNLP.
  39. 39.Zelle, J. M.; and Mooney, R. J. 1996. Learning to Parse Database Queries Using Inductive Logic Programming. In Proc. of AAAI.
  40. 40.Zhong, V.; Lewis, M.; Wang, S. I.; and Zettlemoyer, L. 2020. Grounded Adaptation for Zero-shot Executable Semantic Parsing. In Proc. of EMNLP.

Citation

MLA
Li, J., et al. “Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing”. arXiv, 2023, http://arxiv.org/abs/2301.07507v1.
APA
Li, J., Hui, B., Cheng, R., Qin, B., Ma, C., Huo, N., Huang, F., Du, W., Si, L., & Li, Y. (2023). Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing. arXiv. http://arxiv.org/abs/2301.07507v1
Chicago
Li, J., B. Hui, R. Cheng, et al. 2023. “Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing”. arXiv. http://arxiv.org/abs/2301.07507v1.
Harvard
Li, J. et al. (2023) “Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.07507v1.
Vancouver
1. Li J, Hui B, Cheng R, Qin B, Ma C, Huo N, Huang F, Du W, Si L, Li Y (2023) Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing. arXiv

BibTeX

@article{li2023graphix,
  title = {Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing},
  author = {Li, Jinyang and Hui, Binyuan and Cheng, Reynold and Qin, Bowen and Ma, Chenhao and Huo, Nan and Huang, Fei and Du, Wenyu and Si, Luo and Li, Yongbin},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.07507v1},
  eprint = {2301.07507}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF