Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning

Saibo GengMartin JosifoskiMaxime PeyrardRobert West

article2023EMNLP159 citations

Proposes a unified grammar-constrained decoding framework using input-dependent grammars to enforce strict structural and vocabulary constraints on off-the-shelf language models during inference, outperforming unconstrained baselines and rivaling task-specific finetuned models across diverse structured prediction tasks.

Listen

Large language models excel at producing fluent, free-form text, but they frequently struggle with structured language processing tasks that require strict adherence to predefined syntax or restricted vocabularies, such as extracting facts or linking entities. Standard approaches typically rely on task-specific model finetuning, which requires substantial computational resources and large labeled datasets that are often unavailable in specialized domains.

The article investigates whether a technique called grammar-constrained decoding can serve as a universal framework to force off-the-shelf, pretrained language models to produce perfectly valid structured outputs during generation without requiring any model finetuning.

To evaluate this approach, the researchers formalized the output requirements of 14 common language processing tasks into formal grammars and introduced input-dependent grammars to dynamically restrict model outputs based on specific inputs. They tested this method across several model sizes (LLaMA and Vicuna ranging from 7 billion to 33 billion parameters) in few-shot prompt settings across three benchmark tasks: closed information extraction, entity disambiguation across six standard datasets, and constituency parsing on Penn Treebank data.

The evaluation revealed several key findings. First, grammar-constrained decoding dramatically improved performance over unconstrained models; for example, on closed information extraction, a grammar-constrained 33-billion-parameter model achieved an F1 score of 36.0, doubling the unconstrained baseline of 17.5 and surpassing a dedicated, finetuned baseline model. Second, in entity disambiguation, input-dependent grammars boosted average accuracy from 54.1% unconstrained to 80.3%, outperforming models trained solely on domain data. Third, in syntactic parsing, grammar constraints guaranteed 100% structurally valid parse trees, compared to 54% to 69% validity for unconstrained models, although overall syntactic accuracy remained below specialized, fully supervised parsers. Additionally, parsing overhead added minimal latency (1 to 4 milliseconds per token) for entity and parsing tasks, though grammars with millions of rules introduced larger delays.

These findings demonstrate that grammar constraints provide a rapid, cost-effective way to adapt large language models to complex structured workflows without expensive retraining pipelines. This capability significantly lowers implementation costs and reduces compliance and operational risks associated with invalid or hallucinated outputs in low-resource environments.

Organizations operating in data-scarce domains should consider grammar-constrained decoding as an immediate, lightweight alternative to model finetuning for semantic tasks like information extraction and entity linking, using the largest accessible models and most restrictive grammars possible. However, where extensive labeled datasets exist or where deep syntactic parsing is required, dedicated finetuned models remain preferable.

Confidence in these findings is high for semantic extraction and disambiguation tasks across open-weight models. Readers should note two practical limitations: the technique cannot be directly used with closed, commercial application programming interfaces (such as proprietary cloud models) that hide token probabilities, and extremely large knowledge-base grammars can add computational latency that may require further parser optimization before deployment in high-throughput production systems.

arXiv: 2305.13971epfl-dlab/GCD
Cover for Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning

Abstract

Despite their impressive performance, large language models (LMs) still struggle with reliably generating complex output structures when not finetuned to follow the required output format exactly. To address this issue, grammar-constrained decoding (GCD) can be used to control the generation of LMs, guaranteeing that the output follows a given structure. Most existing GCD methods are, however, limited to specific tasks, such as parsing or code generation. In this work, we demonstrate that formal grammars can describe the output space for a much wider range of tasks and argue that GCD can serve as a unified framework for structured NLP tasks in general. For increased flexibility, we introduce input-dependent grammars, which allow the grammar to depend on the input and thus enable the generation of different output structures for different inputs. We then empirically demonstrate the power and flexibility of GCD-enhanced LMs on (1) information extraction, (2) entity disambiguation, and (3) constituency parsing. Our results indicate that grammar-constrained LMs substantially outperform unconstrained LMs or even beat task-specific finetuned models. Grammar constraints thus hold great promise for harnessing off-the-shelf LMs for a wide range of structured NLP tasks, especially where training data is scarce or finetuning is expensive. Code and data: https://github.com/epfl-dlab/GCD.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 NLP tasks as formal languages
  • 2.2 Grammar-constrained decoding (GCD)
  • 2.3 Few-shot learning with GCD
  • 3 Experimental setup
  • 3.1 Closed information extraction (cIE)
  • 3.2 Entity disambiguation (ED)
  • 3.3 Constituency parsing (CP)
  • 3.4 LLMs and prompting
  • 4 Experimental results
  • 4.1 Closed information extraction (cIE)
  • 4.2 Entity disambiguation (ED)
  • 4.3 Constituency parsing (CP)
  • 4.4 Latency
  • 5 Likelihood misalignment in GCD
  • 6 Related work
  • 7 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Fine-tuning vs. GCD
  • B Grammatical Framework
  • B.1 Abstract Syntax vs. Concrete Syntax
  • B.2 Multilingual Grammars
  • B.3 Expressivity
  • C IE Task Settings
  • D ED Task Settings
  • E Prompt Construction
  • E.1 Information Extraction
  • E.2 Entity Disambiguation
  • E.3 Constituency Parsing
  • F Decoding Settings
  • G Additional Experimental Results
  • G.1 Information Extraction
  • H Latency
  • I Error Examples of Constituency Parsing

Knowls

  1. Knowl 1 — Grammar-Constrained Decoding Framework

    model/method

    Grammar-constrained decoding (GCD) controls the autoregressive generation of language models (LMs) to guarantee that decoded token sequences conform strictly to a predefined formal grammar GG. At decoding step tt, given the partially generated prefix sequence y<t=⟨y0,…,yt−1⟩y_{<t} = \langle y_0, \dots, y_{t-1} \rangle, an incremental parser (acting as a completion engine) evaluates y<ty_{<t} under GG and determines the set of grammatically valid next tokens At⊆ΣtokA_t \subseteq \Sigma_{\text{tok}}, where Σtok\Sigma_{\text{tok}} is the model vocabulary. The LM's unconstrained output probability distribution over the vocabulary, P(yt∣y<t,x)P(y_t \mid y_{<t}, x), is pruned by setting the logit scores of all disallowed tokens v∉Atv \notin A_t to −∞-\infty. GCD is compatible with standard decoding algorithms, including greedy search, beam search, and top-kk sampling, provided the host environment allows access to full vocabulary distributions at each decoding step.

  2. Knowl 2 — Input-Dependent Grammars for Structured NLP

    model/method

    Input-dependent grammars (IDG) extend standard context-free grammars by dynamically parameterizing terminal symbols or production rules based on the input instance x=⟨x0,…,xn−1⟩x = \langle x_0, \dots, x_{n-1} \rangle, rather than using a static grammar for all inputs.

    In entity disambiguation (ED), an IDG restricts generated entity names to a mention-specific candidate set α(x)\alpha(x) extracted for mention mm: S→ℓ m[α(x)] rS \to \ell \, m[\alpha(x)] \, r where ℓ\ell and rr are the left and right context.

    In constituency parsing (CP), an IDG enforces parse-tree completeness by requiring every word of xx to appear in its exact sequence order as leaf nodes within balanced brackets, using state-tracking non-terminals Bi,jB_{i,j} and Ci,jC_{i,j} where index ii tracks word position and index jj tracks open bracket depth: S→B0,0S \to B_{0,0} Bi,j→[α (Bi,j+1∣Ci,j+1)B_{i,j} \to [\alpha \, (B_{i,j+1} \mid C_{i,j+1}) Ci,j→xi (Ci+1,j∣Ei+1,j)C_{i,j} \to x_i \, (C_{i+1,j} \mid E_{i+1,j}) Ei,j+1→] (Ei,j∣Bi,j)E_{i,j+1} \to ] \, (E_{i,j} \mid B_{i,j}) En,0→εE_{n,0} \to \varepsilon with constituent label α∈{S,NP,VP,… }\alpha \in \{\text{S}, \text{NP}, \text{VP}, \dots\}.

  3. Knowl 3 — Token-Level Grammar Decoupling via Grammatical Framework

    model/method

    To prevent subword tokenization ambiguities (such as alternative BPE segmentations of structural characters like brackets) and tokenizer-specific dependencies, GCD decouples grammar definition into an abstract syntax and concrete syntaxes using Grammatical Framework (GF).

    1. Abstract Grammar (GG): Defined at the character/text level, describing the domain language structure independently of any language model tokenization scheme.
    2. Concrete Grammars (GtokG_{\text{tok}}): Automatically compiled from GG into token-level representations Gtok=(V,Σtok,Ptok,S)G_{\text{tok}} = (V, \Sigma_{\text{tok}}, P_{\text{tok}}, S) corresponding to specific LM vocabularies Σtok\Sigma_{\text{tok}}.

    An incremental PMCFG (Parallel Multiple Context-Free Grammar) parser evaluates the generated token stream against GtokG_{\text{tok}} during generation.

  4. Knowl 4 — Length-Normalized Scoring for Likelihood Misalignment Mitigation

    model/method

    When applying GCD to unfinetuned or few-shot language models on structured tasks (such as closed information extraction or constituency parsing), the empty string (or single end-of-sequence token $ with probability qq) frequently achieves the highest total sequence probability. This occurs because generating the full structured sequence requires multiple decoding steps, causing the cumulative probability to fall below qq.

    To resolve this likelihood misalignment without task finetuning, candidate sequences during constrained beam search are reranked using length-normalized sentence scoring: Scorenorm(y)=S(y)mα\text{Score}_{\text{norm}}(y) = \frac{S(y)}{m^\alpha} where S(y)=∑t=1mlog⁡P(yt∣y<t,x)S(y) = \sum_{t=1}^m \log P(y_t \mid y_{<t}, x) is the cumulative log-likelihood, mm is the total token count of candidate yy, and α\alpha is the length normalization parameter. Setting α≥2.5\alpha \ge 2.5 consistently prevents the empty sequence from dominating the beam search output in few-shot settings.

  5. Knowl 5 — Closed Information Extraction Benchmark Results

    data/table

    Few-shot evaluation (4-shot) on the SynthIE-text-small dataset (using 2.7M Wikidata entities and 888 relations) compares unconstrained decoding against grammar-constrained decoding across LLaMA and Vicuna models, and benchmarks them against the supervised GenIE T5-base model.

    Method Precision Recall F1
    Weakly supervised
    GenIE T5-base 49.6 ±\pm 0.3 26.8 ±\pm 0.2 34.8 ±\pm 0.2
    Few-shot unconstrained
    LLaMA-7B 10.2 ±\pm 0.5 14.3 ±\pm 0.7 11.9 ±\pm 0.5
    LLaMA-13B 10.3 ±\pm 0.6 17.0 ±\pm 0.9 12.9 ±\pm 0.6
    LLaMA-33B 14.1 ±\pm 1.0 23.1 ±\pm 1.4 17.5 ±\pm 1.0
    Vicuna-7B 12.5 ±\pm 0.2 16.7 ±\pm 0.1 14.3 ±\pm 0.2
    Vicuna-13B 13.4 ±\pm 0.2 15.2 ±\pm 0.2 14.4 ±\pm 0.2
    Few-shot constrained
    LLaMA-7B 27.9 ±\pm 0.6 20.2 ±\pm 0.5 23.5 ±\pm 0.5
    LLaMA-13B 36.2 ±\pm 0.7 26.5 ±\pm 0.5 30.6 ±\pm 0.5
    LLaMA-33B 39.3 ±\pm 0.9 33.2 ±\pm 0.8 36.0 ±\pm 0.7
    Vicuna-7B 25.4 ±\pm 0.5 15.8 ±\pm 0.3 19.5 ±\pm 0.3
    Vicuna-13B 38.7 ±\pm 1.0 19.8 ±\pm 0.8 26.1 ±\pm 0.8

    Constraining decoding to valid knowledge base entities and relations yields substantial gains over unconstrained baselines (e.g., LLaMA-33B F1 increases from 17.5 to 36.0), outperforming the supervised GenIE baseline.

  6. Knowl 6 — Entity Disambiguation Benchmark Results

    data/table

    Micro-accuracy across six entity disambiguation test datasets under 4-shot prompting, comparing unconstrained decoding, input-independent grammar (IIG, entire entity catalog), and input-dependent grammar (IDG, candidate entities per mention).

    Method AIDA MSNBC AQUAINT ACE2004 CWeb WIKI Avg.
    Supervised
    Le and Titov (2018) 89.6 92.2 90.7 88.1 78.2 81.7 86.8
    BLINK w/o cand. set 79.6 80.0 80.3 82.5 64.2 75.5 77.0
    BLINK 86.7 90.3 88.9 88.7 82.6 86.1 87.2
    GENRE (AIDA only) 88.6 88.1 77.1 82.3 71.9 71.7 80.0
    GENRE 93.3 94.3 89.9 90.1 77.3 87.4 88.8
    ReFinED 93.9 94.1 90.8 90.8 79.4 87.4 89.4
    Few-shot unconstrained
    LLaMA-7B 42.0 44.6 30.2 43.8 35.8 27.7 37.4
    LLaMA-13B 48.1 50.2 36.2 47.5 40.7 37.2 43.3
    LLaMA-33B 62.6 63.0 42.9 56.3 48.1 51.4 54.1
    Few-shot constrained (IIG)
    LLaMA-7B 56.3 57.3 61.6 54.6 50.5 47.0 54.5
    LLaMA-13B 51.8 57.3 53.3 50.8 48.2 39.7 50.6
    LLaMA-33B 69.8 73.3 74.9 71.7 61.6 57.6 68.2
    Few-shot constrained (IDG)
    LLaMA-7B 73.4 87.6 83.2 82.9 69.4 67.1 77.2
    LLaMA-13B 75.8 86.6 82.4 84.2 68.1 68.1 77.5
    LLaMA-33B 81.0 88.2 86.2 85.4 70.7 70.5 80.3

    IDG achieves an average accuracy of 80.3% with LLaMA-33B, outperforming both IIG (68.2%) and unconstrained prompting (54.1%), and exceeding GENRE trained solely on AIDA data (80.0%).

  7. Knowl 7 — Constituency Parsing Bracketing Performance and Structural Validity

    data/table

    Performance on Penn Treebank test sentences with gold parse trees shorter than 64 tokens under 8-shot prompting, comparing bracketing F1 and syntactic parse-tree validity percentage.

    Method F1 Validity (%)
    Bespoke supervised methods
    Vinyals et al. (2015a) 92.1 98.5
    Dyer et al. (2016) 93.3 100.0
    Kitaev and Klein (2018) 95.6 100.0
    Zhang et al. (2020) 95.7 100.0
    Few-shot unconstrained
    LLaMA-7B 28.1 54.3
    LLaMA-13B 42.8 69.4
    LLaMA-33B 42.9 64.2
    Few-shot constrained (IIG)
    LLaMA-7B 34.7 65.9
    LLaMA-13B 45.4 80.3
    LLaMA-33B 47.1 72.3
    Few-shot constrained (IDG)
    LLaMA-7B 45.8 100.0
    LLaMA-13B 53.4 100.0
    LLaMA-33B 54.6 100.0

    IDG constraints guarantee 100% parse-tree validity and improve LLaMA-33B bracketing F1 from 42.9 to 54.6, though performance remains below fully supervised parsers (>95%>95\% F1).

  8. Knowl 8 — Per-Token Decoding Latency and Overhead of GCD

    empirical result

    Unconstrained language model forward passes on an NVIDIA A100 GPU exhibit per-token latencies of 54 ms for LLaMA-7B, 87 ms for LLaMA-33B, and 136 ms for LLaMA-65B. The CPU-based incremental parsing overhead introduced by GCD varies by task grammar size:

    • Entity Disambiguation (IDG): 1 ms overhead per token.
    • Constituency Parsing (IDG): 4 ms overhead per token.
    • Closed Information Extraction (WikiNER, 279K entities, 158 relations): 69 ms average overhead per token.
    • Closed Information Extraction (REBEL, 5.9M entities, 857 relations): ~500 ms average overhead per token.
  9. Knowl 9 — Limitations of Grammar-Constrained Decoding

    limitation

    Grammar-constrained decoding possesses three primary limitations:

    1. Black-Box Incompatibility: GCD requires direct access to vocabulary-level logit distributions at each step to prune disallowed tokens. Consequently, it cannot be applied to proprietary API-only language models that only return generated text.
    2. Grammar Size Latency Scaling: For grammars comprising millions of production rules (such as comprehensive knowledge bases in REBEL with 5.9M entities), CPU-based incremental parser latency scales up to ~0.5 seconds per token, significantly exceeding base LM GPU inference latency.
    3. Syntactic vs. Semantic Gap: While GCD guarantees formal syntactic and structural validity, it does not correct the underlying linguistic or syntactic reasoning deficits of the LM (as evidenced by the performance gap between GCD-guided few-shot models and supervised parsers on constituency parsing).

Coverage note — Omitted specific prompt text variants and minor dataset ablation configurations (e.g., subject-collapsed vs. fully expanded linearization ablation in Table 7) to focus on the primary conceptual and empirical contributions.

References

  1. 1.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2017. Guided open vocabulary image captioning with constrained beam search. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 936–945, Copenhagen, Denmark. Association for Computational Linguistics.
  2. 2.Krasimir Angelov. 2009. Incremental parsing with parallel multiple context-free grammars. In Proceedings of the 12th Conference of the European Chapter of the ACL (EACL 2009), pages 69–76, Athens, Greece. Association for Computational Linguistics.
  3. 3.Ben Athiwaratkun, Cicero Nogueira dos Santos, Jason Krone, and Bing Xiang. 2020. Augmented natural language for generative sequence labeling. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 375–385, Online. Association for Computational Linguistics.
  4. 4.Tom Ayoola, Shubhi Tyagi, Joseph Fisher, Christos Christodoulopoulos, and Andrea Pierleoni. 2022. ReFinED: An efficient zero-shot-capable approach to end-to-end entity linking. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Track, pages 209–220, Hybrid: Seattle, Washington + Online. Association for Computational Linguistics.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  7. 7.Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. 2021. Autoregressive entity retrieval.
  8. 8.Daniel Deutsch, Shyam Upadhyay, and Dan Roth. 2019. A general-purpose algorithm for constrained sequential inference. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 482–492, Hong Kong, China. Association for Computational Linguistics.
  9. 9.Alexander Dunn, John Dagdelen, Nicholas Walker, Sanghoon Lee, Andrew S. Rosen, Gerbrand Ceder, Kristin Persson, and Anubhav Jain. 2022. Structured information extraction from complex scientific text with fine-tuned large language models.
  10. 10.Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016. Recurrent neural network grammars.
  11. 11.Evgeniy Gabrilovich, Michael Ringgaard, and Amarnag Subramanya. 2013. Facc1: Freebase annotation of clueweb corpora, version 1 (release date 2013-06-26, format version 1, correction level 0).
  12. 12.Octavian-Eugen Ganea and Thomas Hofmann. 2017. Deep joint entity disambiguation with local neural attention.
  13. 13.Zhaochen Guo and Denilson Barbosa. 2017. Robust named entity disambiguation with random walks. Semantic Web, 9:1–21.
  14. 14.Johannes Hoffart, Mohamed Amir Yosef, Ilaria Bordino, Hagen Fürstenau, Manfred Pinkal, Marc Spaniol, Bilyana Taneva, Stefan Thater, and Gerhard Weikum. 2011. Robust disambiguation of named entities in text. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, pages 782–792, Edinburgh, Scotland, UK. Association for Computational Linguistics.
  15. 15.Chris Hokamp and Qun Liu. 2017. Lexically constrained decoding for sequence generation using grid beam search. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1535–1546, Vancouver, Canada. Association for Computational Linguistics.
  16. 16.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models.
  17. 17.J. Edward Hu, Huda Khayrallah, Ryan Culkin, Patrick Xia, Tongfei Chen, Matt Post, and Benjamin Van Durme. 2019. Improved lexically constrained decoding for translation and monolingual rewriting. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 839–850, Minneapolis, Minnesota. Association for Computational Linguistics.
  18. 18.Pere-Lluís Huguet Cabot and Roberto Navigli. 2021. REBEL: Relation extraction by end-to-end language generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2370–2381, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  19. 19.Bernal Jimenez Gutierrez, Nikolas McNeal, Clayton Washington, You Chen, Lang Li, Huan Sun, and Yu Su. 2022. Thinking about GPT-3 in-context learning for biomedical IE? think again. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 4497–4512, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  20. 20.Martin Josifoski, Nicola De Cao, Maxime Peyrard, Fabio Petroni, and Robert West. 2022. GenIE: Generative information extraction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4626–4643, Seattle, United States. Association for Computational Linguistics.
  21. 21.Martin Josifoski, Marija Sakota, Maxime Peyrard, and Robert West. 2023. Exploiting asymmetry for synthetic training data generation: SynthIE and the case of information extraction. arXiv preprint arXiv:2303.04132.
  22. 22.Nikita Kitaev and Dan Klein. 2018. Constituency parsing with a self-attentive encoder.
  23. 23.Phong Le and Ivan Titov. 2018. Improving entity linking by modeling latent relations between mentions. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1595–1604, Melbourne, Australia. Association for Computational Linguistics.
  24. 24.Giovanni Paolini, Ben Athiwaratkun, Jason Krone, JIE MA, Alessandro Achille, Rishita Anubhai, Cícero Nogueira dos Santos, Bing Xiang, and Stefano Soatto. 2021. Structured prediction as translation between augmented natural languages. In ICLR 2021.
  25. 25.Gabriel Poesia, Oleksandr Polozov, Vu Le, Ashish Tiwari, Gustavo Soares, Christopher Meek, and Sumit Gulwani. 2022. Synchromesh: Reliable code generation from pre-trained language models.
  26. 26.Matt Post and David Vilar. 2018. Fast lexically constrained decoding with dynamic beam allocation for neural machine translation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1314–1324, New Orleans, Louisiana. Association for Computational Linguistics.
  27. 27.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  28. 28.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. CoRR, abs/1910.10683.
  29. 29.Aarne Ranta. 2019. Grammatical framework: an interlingual grammar formalism. In Proceedings of the 14th International Conference on Finite-State Methods and Natural Language Processing, pages 1–2, Dresden, Germany. Association for Computational Linguistics.
  30. 30.Subhro Roy, Sam Thomson, Tongfei Chen, Richard Shin, Adam Pauls, Jason Eisner, and Benjamin Van Durme. 2022. Benchclamp: A benchmark for evaluating language models on semantic parsing.
  31. 31.Timo Schick and Hinrich Schütze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269, Online. Association for Computational Linguistics.
  32. 32.Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. 2021. PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9895–9901, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  33. 33.Hiroyuki Seki, Takashi Matsumura, Mamoru Fujii, and Tadao Kasami. 1991. On multiple context-free grammars. Theoretical Computer Science, 88(2):191–229.
  34. 34.Satoshi Sekine and Michael Collins. 2008. Evalb: Bracket scoring program.
  35. 35.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  36. 36.Richard Shin, Christopher Lin, Sam Thomson, Charles Chen, Subhro Roy, Emmanouil Antonios Platanios, Adam Pauls, Dan Klein, Jason Eisner, and Benjamin Van Durme. 2021. Constrained language models yield few-shot semantic parsers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7699–7715, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  37. 37.Felix Stahlberg and Bill Byrne. 2019. On NMT search errors and model errors: Cat got your tongue? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3356–3362, Hong Kong, China. Association for Computational Linguistics.
  38. 38.Elias Stengel-Eskin, Kyle Rawlins, and Benjamin Van Durme. 2023. Zero and few-shot semantic parsing with ambiguous inputs.
  39. 39.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models.
  40. 40.Roy Tromble and Jason Eisner. 2006. A fast finite-state relaxation method for enforcing global constraints on sequence decoding. In Proceedings of the Human Language Technology Conference of the NAACL, Main Conference, pages 423–430, New York City, USA. Association for Computational Linguistics.
  41. 41.Oriol Vinyals, Lukasz Kaiser, Terry Koo, Slav Petrov, Ilya Sutskever, and Geoffrey Hinton. 2015a. Grammar as a foreign language.
  42. 42.Oriol Vinyals, Łukasz Kaiser, Terry Koo, Slav Petrov, Ilya Sutskever, and Geoffrey Hinton. 2015b. Grammar as a foreign language. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc.
  43. 43.Denny Vrandeˇci´c. 2012. Wikidata: A new platform for collaborative data collection. In Proceedings of the 21st International Conference on World Wide Web, WWW ’12 Companion, pages 1063–1064, New York, NY, USA. Association for Computing Machinery.
  44. 44.Bailin Wang, Zi Wang, Xuezhi Wang, Yuan Cao, Rif A. Saurous, and Yoon Kim. 2023. Grammar prompting for domain-specific language generation with large language models.
  45. 45.Ledell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel, and Luke Zettlemoyer. 2020. Scalable zero-shot entity linking with dense entity retrieval. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6397–6407, Online. Association for Computational Linguistics.
  46. 46.Pengcheng Yin and Graham Neubig. 2017. A syntactic neural model for general-purpose code generation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 440–450, Vancouver, Canada. Association for Computational Linguistics.
  47. 47.Yu Zhang, Houquan Zhou, and Zhenghua Li. 2020. Fast and accurate neural CRF constituency parsing. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization.

Citation

MLA
Geng, S., et al. “Grammar-Constrained Decoding for Structured NLP Tasks Without Finetuning”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 10932–52, https://doi.org/10.18653/v1/2023.emnlp-main.674.
APA
Geng, S., Josifoski, M., Peyrard, M., & West, R. (2023). Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10932–10952. https://doi.org/10.18653/v1/2023.emnlp-main.674
Chicago
Geng, S., M. Josifoski, M. Peyrard, and R. West. 2023. “Grammar-Constrained Decoding for Structured NLP Tasks Without Finetuning”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10932–52. https://doi.org/10.18653/v1/2023.emnlp-main.674.
Harvard
Geng, S. et al. (2023) “Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10932–10952. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.674.
Vancouver
1. Geng S, Josifoski M, Peyrard M, West R (2023) Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10932–10952

BibTeX

@inproceedings{geng-etal-2023-grammar,
    title = "Grammar-Constrained Decoding for Structured {NLP} Tasks without Finetuning",
    author = "Geng, Saibo  and
      Josifoski, Martin  and
      Peyrard, Maxime  and
      West, Robert",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.674/",
    doi = "10.18653/v1/2023.emnlp-main.674",
    pages = "10932--10952"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/