Built independently by an author, for readers. Read the story and support ChapterPal

keyword

terminology constraints

Terminology constraints are predefined lexical rules in natural language processing and text generation that mandate the inclusion or exclusion of specific words, phrases, or domain-specific terms in a model-generated output. Commonly applied in neural machine translation, structured data-to-text generation, and controllable text synthesis, these constraints ensure that specialized vocabulary, standardized nomenclature, or brand names are accurately preserved rather than replaced by generic alternatives. Generative language systems typically enforce terminology constraints during the decoding phase using specialized search heuristics, finite-state automata, or constrained optimization techniques to satisfy the required lexical rules without sacrificing the grammatical fluency and contextual coherence of the resulting text.

1 item

NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics

NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics

Ximing Lu, Sean Welleck, Peter West, Liwei Jiang, Jungo Kasai, Daniel Khashabi, Ronan Le Bras, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah A. Smith, Yejin Choi

OrganizationsAllen Institute for AIUniversity of Washington

Why you should read this

Proposes an A*-inspired decoding algorithm with efficient lookahead heuristics that enables autoregressive language models to satisfy complex lexical constraints and achieve state-of-the-art performance across multiple text generation benchmarks without task-specific training data.

The dominant paradigm for neural text generation is left-to-right decoding from autoregressive language models. Constrained or controllable generation under complex lexical constraints, however, requires foresight to plan ahead for feasible future paths. Drawing inspiration from the A* search algorithm, we propose NEUROLOGIC A★esque,1 a decoding algorithm that incorporates heuristic estimates of future cost. We develop lookahead heuristics that are efficient for large-scale language models, making our method a drop-in replacement for common techniques such as beam search and top-k sampling. To enable constrained generation, we build on NEUROLOGIC decoding (Lu et al., 2021), combining its flexibility in incorporating logical constraints with A★esque estimates of future constraint satisfaction. Our approach outperforms competitive baselines on five generation tasks, and achieves new state-of-the-art performance on table-to-text generation, constrained machine translation, and keyword-constrained generation. The improvements are particularly notable on tasks that require complex constraint satisfaction or in few-shot or zero-shot settings. NEUROLOGIC A★esque illustrates the power of decoding for improving and enabling new capabilities of large-scale language models.

Added

2026-09-28