keyword
syntactic constraints
Syntactic constraints are structural rules and grammatical principles that govern how words, phrases, and clauses can be legitimately combined and ordered within a language to form well-formed sentences. These constraints restrict language by specifying mandatory grammatical relationships, such as subject-verb agreement, argument structure, parts of speech arrangements, and lexico-syntactic co-occurrence patterns between specific vocabulary items. In computational linguistics and natural language processing, syntactic constraints serve as explicit specifications or formal targets that guide parsing algorithms and text generation models, ensuring that produced or analyzed text conforms to required grammatical forms, stylistic templates, or structural dependencies.
2 items

Controlled Text Generation with Natural Language Instructions
Wangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell, Mrinmaya Sachan
Why you should read this
Introduces INSTRUCTCTG, a training-time framework that verbalizes diverse lexical, syntactic, semantic, style, and length constraints into natural language prompts to control text generation efficiently without modifying decoding algorithms.
Large language models can be prompted to produce fluent output for a wide range of tasks without being specifically trained to do so. Nevertheless, it is notoriously difficult to control their generation in such a way that it satisfies user-specified constraints. In this paper, we present INSTRUCTCTG, a simple controlled text generation framework that incorporates different constraints by verbalizing them as natural language instructions. We annotate natural texts through a combination of off-the-shelf NLP tools and simple heuristics with the linguistic and extra-linguistic constraints they satisfy. Then, we verbalize the constraints into natural language instructions to form weakly supervised training data, i.e., we prepend the natural language verbalizations of the constraints in front of their corresponding natural language sentences. Next, we fine-tune a pre-trained language model on the augmented corpus. Compared to existing methods, INSTRUCTCTG is more flexible in terms of the types of constraints it allows the practitioner to use. It also does not require any modification of the decoding procedure. Finally, INSTRUCTCTG allows the model to adapt to new constraints without re-training through the use of in-context learning. Our code is available at https://github.com/MichaelZhouwang/InstructCTG.
Added
2026-09-26

Word Association Norms, Mutual Information, and Lexicography
Kenneth Ward Church, Patrick Hanks
Why you should read this
Proposes a corpus-based mutual information metric that automates the extraction of semantic and syntactic word associations at scale, replacing expensive human subject testing for computational linguists and lexicographers.
The term word association is used in a very particular sense in the psycholinguistic literature. (Generally speaking, subjects respond quicker than normal to the word “nurse” if it follows a highly associated word such as “doctor.”) We will extend the term to provide the basis for a statistical description of a variety of interesting linguistic phenomena, ranging from semantic relations of the doctor/nurse type (content word/content word) to lexico-syntactic co-occurrence constraints between verbs and prepositions (content word/function word). This paper will propose a new objective measure based on the information theoretic notion of mutual information, for estimating word association norms from computer readable corpora. (The standard method of obtaining word association norms, testing a few thousand subjects on a few hundred words, is both costly and unreliable.) The proposed measure, the association ratio, estimates word association norms directly from computer readable corpora, making it possible to estimate norms for tens of thousands of words.
Added
2026-09-10
