Built independently by an author, for readers. Read the story and support ChapterPal

keyword

lexical constraints

Lexical constraints are explicit rules or requirements in natural language processing and text generation that dictate the inclusion, exclusion, or specific positioning of particular words, phrases, or terminology in the generated output. In controlled language generation tasks such as machine translation, summarization, and keyword-to-text generation, positive lexical constraints specify essential target terms that must appear in the final text, while negative lexical constraints mandate that certain forbidden words be avoided. Enforcing these constraints requires language models or specialized decoding algorithms to guide the sequence generation process to satisfy the prescribed vocabulary requirements while preserving grammatical correctness, fluency, and overall semantic coherence.

3 items

NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics

NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics

Ximing Lu, Sean Welleck, Peter West, Liwei Jiang, Jungo Kasai, Daniel Khashabi, Ronan Le Bras, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah A. Smith, Yejin Choi

OrganizationsAllen Institute for AIUniversity of Washington

Why you should read this

Proposes an A*-inspired decoding algorithm with efficient lookahead heuristics that enables autoregressive language models to satisfy complex lexical constraints and achieve state-of-the-art performance across multiple text generation benchmarks without task-specific training data.

The dominant paradigm for neural text generation is left-to-right decoding from autoregressive language models. Constrained or controllable generation under complex lexical constraints, however, requires foresight to plan ahead for feasible future paths. Drawing inspiration from the A* search algorithm, we propose NEUROLOGIC A★esque,1 a decoding algorithm that incorporates heuristic estimates of future cost. We develop lookahead heuristics that are efficient for large-scale language models, making our method a drop-in replacement for common techniques such as beam search and top-k sampling. To enable constrained generation, we build on NEUROLOGIC decoding (Lu et al., 2021), combining its flexibility in incorporating logical constraints with A★esque estimates of future constraint satisfaction. Our approach outperforms competitive baselines on five generation tasks, and achieves new state-of-the-art performance on table-to-text generation, constrained machine translation, and keyword-constrained generation. The improvements are particularly notable on tasks that require complex constraint satisfaction or in few-shot or zero-shot settings. NEUROLOGIC A★esque illustrates the power of decoding for improving and enabling new capabilities of large-scale language models.

Added

2026-09-28

Controlled Text Generation with Natural Language Instructions

Controlled Text Generation with Natural Language Instructions

Wangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell, Mrinmaya Sachan

OrganizationsETH Zurich

Why you should read this

Introduces INSTRUCTCTG, a training-time framework that verbalizes diverse lexical, syntactic, semantic, style, and length constraints into natural language prompts to control text generation efficiently without modifying decoding algorithms.

Large language models can be prompted to produce fluent output for a wide range of tasks without being specifically trained to do so. Nevertheless, it is notoriously difficult to control their generation in such a way that it satisfies user-specified constraints. In this paper, we present INSTRUCTCTG, a simple controlled text generation framework that incorporates different constraints by verbalizing them as natural language instructions. We annotate natural texts through a combination of off-the-shelf NLP tools and simple heuristics with the linguistic and extra-linguistic constraints they satisfy. Then, we verbalize the constraints into natural language instructions to form weakly supervised training data, i.e., we prepend the natural language verbalizations of the constraints in front of their corresponding natural language sentences. Next, we fine-tune a pre-trained language model on the augmented corpus. Compared to existing methods, INSTRUCTCTG is more flexible in terms of the types of constraints it allows the practitioner to use. It also does not require any modification of the decoding procedure. Finally, INSTRUCTCTG allows the model to adapt to new constraints without re-training through the use of in-context learning. Our code is available at https://github.com/MichaelZhouwang/InstructCTG.

Added

2026-09-26

Evaluating Large Language Models at Evaluating Instruction Following

Evaluating Large Language Models at Evaluating Instruction Following

Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, Danqi Chen

OrganizationsPrinceton UniversityTsinghua UniversityUniversity of Illinois Urbana-Champaign

Why you should read this

Introduces LLMBar, an adversarial meta-evaluation benchmark that exposes how large language model evaluators fail when assessing instruction following against deceptive surface qualities, and proposes prompting strategies to bring their performance closer to human judgment.

As research in large language models (LLMs) continues to accelerate, LLM-based evaluation has emerged as a scalable and cost-effective alternative to human evaluations for comparing the ever increasing list of models. This paper investigates the efficacy of these ``LLM evaluators'', particularly in using them to assess instruction following, a metric that gauges how closely generated text adheres to the given instruction. We introduce a challenging meta-evaluation benchmark, LLMBar, designed to test the ability of an LLM evaluator in discerning instruction-following outputs. The authors manually curated 419 pairs of outputs, one adhering to instructions while the other diverging, yet may possess deceptive qualities that mislead an LLM evaluator, e.g., a more engaging tone. Contrary to existing meta-evaluation, we discover that different evaluators (i.e., combinations of LLMs and prompts) exhibit distinct performance on LLMBar and even the highest-scoring ones have substantial room for improvement. We also present a novel suite of prompting strategies that further close the gap between LLM and human evaluators. With LLMBar, we hope to offer more insight into LLM evaluators and foster future research in developing better instruction-following models.

Added

2026-09-25