NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics

Ximing LuSean WelleckPeter WestLiwei JiangJungo KasaiDaniel KhashabiRonan Le BrasLianhui QinYoungjae YuRowan Zellers

article2022NAACL211 citationsBest new method paper

Proposes an A*-inspired decoding algorithm with efficient lookahead heuristics that enables autoregressive language models to satisfy complex lexical constraints and achieve state-of-the-art performance across multiple text generation benchmarks without task-specific training data.

Listen

Modern neural language models typically generate text sequentially from left to right, picking each next word based solely on past context. While this process works well for open-ended writing, it often fails when outputs must satisfy strict business or domain constraints, such as incorporating mandatory terminology, adhering to structured data, or enforcing specific keywords. Standard methods lack foresight, frequently leading to awkward phrasing or failing to incorporate required content later in the sentence.

The article demonstrates and evaluates an advanced text decoding method named NEUROLOGIC Aesque (abbreviated as NEUROLOGIC). The main objective is to establish whether integrating forward-looking heuristic estimates—inspired by classical A* shortest-path search algorithms—into standard left-to-right generation can improve both content quality and strict adherence to constraints without requiring expensive model retraining.

The researchers evaluated the approach across five distinct language generation benchmarks covering commonsense sentence formulation, terminology-constrained German-English machine translation, structured table-to-text generation, keyword-constrained question writing, and unconstrained creative story continuation. The testing spanned both fully supervised models and resource-constrained few-shot and zero-shot scenarios using standard language model architectures (such as GPT-2 and Marian translation systems). The core technique evaluates short lookahead paths during generation to approximate future probabilities and the likelihood of satisfying pending constraints.

The findings establish that lookahead-guided decoding systematically outperforms standard decoding baselines across all evaluated tasks. On structured table-to-text tasks with very scarce training data (using only 0.1% of typical training instances), the method achieved 100% information coverage and increased automated quality scores substantially over conventional beam search. In commonsense and question generation, the method reached near-perfect to perfect constraint satisfaction while scoring significantly higher on human assessments of grammatical quality and scenario plausibility. Notably, an off-the-shelf, unsupervised language model using this decoding approach outperformed fully supervised baseline models on human preference metrics. Lookahead search also produced superior fluency and diversity in unconstrained creative generation.

These results demonstrate that significant improvements in model capability, reliability, and factual consistency can be achieved purely at inference time. Organizations can deploy standard, pre-trained language models directly into specialized domains without incurring the substantial timeline delays and computational costs required to train or fine-tune custom models. For compliance-heavy, high-precision applications like technical translation or business reporting from databases, this approach provides a dependable mechanism to enforce mandatory language constraints.

Decision-makers should consider adopting lookahead heuristic decoding as a drop-in replacement for conventional beam search or sampling in accuracy-critical pipelines. When implementing, technical teams should favor greedy lookahead approximations for latency-sensitive tasks, as they capture most of the quality gains at lower operational overhead compared to multi-branch lookahead search. Further pilot testing in targeted operational domains is recommended to benchmark production performance.

The primary operational limitation is increased latency: lookahead search evaluates future trajectories at each step, making it roughly an order of magnitude slower than simple beam search (e.g., around 19 seconds per sentence compared to 2 seconds for earlier constrained methods under test conditions). Additionally, the algorithm only supports constraints expressible as formal logical phrases (inclusion or exclusion). While confidence in the benchmark results is high, practitioners must remain cautious regarding standard neural risks, as constrained decoding can be steered to force biased or harmful phrases if inputs are unmonitored.

Cover for NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics

Abstract

The dominant paradigm for neural text generation is left-to-right decoding from autoregressive language models. Constrained or controllable generation under complex lexical constraints, however, requires foresight to plan ahead for feasible future paths.

Drawing inspiration from the A* search algorithm, we propose NEUROLOGIC A★esque,1 a decoding algorithm that incorporates heuristic estimates of future cost. We develop lookahead heuristics that are efficient for large-scale language models, making our method a drop-in replacement for common techniques such as beam search and top-k sampling. To enable constrained generation, we build on NEUROLOGIC decoding (Lu et al., 2021), combining its flexibility in incorporating logical constraints with A★esque estimates of future constraint satisfaction.

Our approach outperforms competitive baselines on five generation tasks, and achieves new state-of-the-art performance on table-to-text generation, constrained machine translation, and keyword-constrained generation. The improvements are particularly notable on tasks that require complex constraint satisfaction or in few-shot or zero-shot settings. NEUROLOGIC A★esque illustrates the power of decoding for improving and enabling new capabilities of large-scale language models.

Table of Contents

  • 1 Introduction
  • 2 NEUROLOGIC A ⋆ esque Decoding
  • 2.1 Decoding With A ⋆ esque Lookahead
  • 2.2 Unconstrained Generation with NEUROLOGIC ⋆
  • 2.3 NEUROLOGIC ⋆ for Constrained Generation
  • 3 Experiments
  • 3.1 Constrained Commonsense Generation
  • 3.2 Constrained Machine Translation
  • 3.3 Table-to-text Generation
  • 3.4 Constrained Question Generation
  • 3.5 Unconstrained Story Generation
  • 4 Related Work
  • 5 Conclusion
  • Acknowledgment
  • Broader Impact and Ethical Implications
  • References
  • A Further Experiments
  • A.1 Constrained Commonsense Generation
  • A.2 Unconstrained Story Generation
  • B Runtime
  • C Experimental Details
  • C.1 Off-the-Shelf Models
  • C.2 Model Training Details
  • C.2.1 COMMONGEN
  • C.2.2 Constrained Machine Translation
  • C.2.3 Table-to-text Generation
  • C.2.4 Unconstrained Story Generation
  • C.3 Generation Details
  • C.3.1 COMMONGEN
  • C.3.2 Constrained Machine Translation
  • C.3.3 Table-to-text Generation
  • C.3.4 Constrained Question Generation
  • C.3.5 Unconstrained Story Generation
  • C.4 Dataset Details
  • D Human Evaluation
  • E Qualitative Generation Examples
  • F Limitations and Risks.

Knowls

  1. Knowl 1 — A*-style lookahead decoding for autoregressive generation

    model/method

    NEUROLOGIC A*esque, abbreviated NEUROLOGICF, treats left-to-right neural text generation as discrete search. Given an input sequence xx, an autoregressive language model pθp_\theta, vocabulary VV, and a partial output prefix y≤ty_{\leq t}, the target sequence maximizes a model score plus an optional constraint score:

    pθ(y∣x)=∏t=1∣y∣pθ(yt∣y<t,x),F(y)=s(y)+H(y),s(y)=log⁡pθ(y∣x).p_\theta(y\mid x)=\prod_{t=1}^{|y|}p_\theta(y_t\mid y_{<t},x),\qquad F(y)=s(y)+H(y),\qquad s(y)=\log p_\theta(y\mid x).

    Here, H(y)=0H(y)=0 for unconstrained generation and measures constraint satisfaction when lexical constraints are provided. At step tt, the decoder expands each retained prefix with possible next tokens, forms candidate prefixes y≤t=y<t∘yty_{\leq t}=y_{<t}\mathbin{\circ}y_t, scores them, and keeps the best kk candidates. Unlike ordinary beam search, NEUROLOGICF scores a candidate using both its accumulated log-probability and the best heuristic estimate among a finite set Lℓ(y≤t)L_\ell(y_{\leq t}) of length-ℓ\ell continuations:

    Yt=TopK⁡y≤t∈Yt′[s(y≤t)+max⁡z∈Lℓ(y≤t)h(z)].Y_t=\operatorname{TopK}_{y_{\leq t}\in Y'_t}\left[s(y_{\leq t})+\max_{z\in L_\ell(y_{\leq t})}h(z)\right].

    The lookahead set is regenerated for every candidate prefix. Decoding repeats expansion, lookahead generation, scoring, and pruning until sequences terminate, normally at an end-of-sequence token or a maximum length. Beam search is recovered when the lookahead length and heuristic are both zero. The procedure preserves beam-search-style pruning while using an A*-like estimate of future return; the paper does not claim exact optimality for finite lookahead.

  2. Knowl 2 — Lookahead continuation strategies

    model/method

    NEUROLOGICF supports four ways to construct the finite lookahead set Lℓ(y≤t)L_\ell(y_{\leq t}), where ℓ\ell is the number of future tokens examined and VV is the language-model vocabulary.

    A greedy lookahead generates one continuation by selecting the highest-probability token at every future step. A soft lookahead also follows one trajectory, but feeds the model an expected token embedding instead of a single token. If st∈R∣V∣s_t\in\mathbb{R}^{|V|} is the logit vector at a future step, τ≥0\tau\geq 0 is a temperature, p~θ\widetilde p_\theta is the temperature-adjusted distribution, and E∈R∣V∣×dE\in\mathbb{R}^{|V|\times d} is the token-embedding matrix, the input embedding is

    p~θ(yt∣y<t)=softmax⁡(st/τ),et=Eyt∼p~θ(⋅∣y<t)[E(yt)].\widetilde p_\theta(y_t\mid y_{<t})=\operatorname{softmax}(s_t/\tau),\qquad e_t=\mathbb{E}_{y_t\sim\widetilde p_\theta(\cdot\mid y_{<t})}[E(y_t)].

    As τ→0\tau\to0, the soft lookahead approaches greedy decoding; as τ→∞\tau\to\infty, it approaches a uniform mixture of token embeddings. When scoring soft lookaheads, the decoder uses p~θ\widetilde p_\theta rather than the original token distribution.

    A beam lookahead runs beam search for ℓ\ell future steps and returns the top-kk continuations. A sampling lookahead independently samples multiple continuations token by token from the model distribution. Greedy and soft lookaheads are cheapest but explore one trajectory; beam and sampling lookaheads explore multiple possible futures at higher computational cost.

  3. Knowl 3 — Logical constraints and the underlying NEUROLOGIC search

    model/method

    The constrained version accepts lexical constraints written in conjunctive normal form (CNF). A literal is either a positive requirement that a phrase aa occur in the output, written D(a,y)D(a,y), or a negative requirement that it not occur, written ¬D(a,y)\neg D(a,y). Literals are grouped into clauses, and all clauses must be satisfied:

    C1∧C2∧⋯∧CM,C_1\land C_2\land\cdots\land C_M,

    where each clause CjC_j is a disjunction of literals. The underlying NEUROLOGIC search scores a partial prefix y≤ty_{\leq t} by its log-probability plus partial progress toward unsatisfied multi-token constraints:

    f(y≤t)=log⁡pθ(y≤t∣x)+λ1max⁡D(a,y≤t)∣a^∣∣a∣.f(y_{\leq t})=\log p_\theta(y_{\leq t}\mid x)+\lambda_1\max_{D(a,y_{\leq t})}\frac{|\hat a|}{|a|}.

    Here, aa is an unsatisfied constraint phrase, ∣a∣|a| is its token length, a^\hat a is the matching prefix of aa already present at the end of the generated sequence, and λ1\lambda_1 is a scaling factor. For example, if the current output ends in apple and the constraint is apple tree, the partial match is apple. The search also removes candidates that violate negative constraints, groups candidates to preserve diversity, and selects high-scoring candidates from the groups rather than simply taking the globally highest-scoring prefixes.

  4. Knowl 4 — Future constraint-satisfaction heuristic

    model/method

    NEUROLOGICF extends the logical-constraint search by estimating whether unsatisfied constraint phrases can be completed in the generated lookahead. For a candidate prefix y≤ty_{\leq t} followed by a length-ℓ\ell lookahead yt+1:t+ℓy_{t+1:t+\ell}, the future heuristic is

    hfuture(y≤t+ℓ)=λ2max⁡D(a,y≤t)log⁡pθ(D(a,yt+1:t+ℓ)∣x,y≤t),h_{\mathrm{future}}(y_{\leq t+\ell})=\lambda_2\max_{D(a,y_{\leq t})}\log p_\theta\big(D(a,y_{t+1:t+\ell})\mid x,y_{\leq t}\big),

    where λ2\lambda_2 scales the future constraint signal and the maximization ranges over tracked unsatisfied constraint phrases aa. The paper defines the probability of satisfying a phrase in the lookahead using its most probable matching subsequence:

    pθ(D(a,yt+1:t+ℓ)∣x,y≤t)=max⁡t′∈[t,t+ℓ]pθ(yt′:t′+∣a∣=a∣x,y<t′).p_\theta\big(D(a,y_{t+1:t+\ell})\mid x,y_{\leq t}\big)=\max_{t'\in[t,t+\ell]}p_\theta\big(y_{t':t'+|a|}=a\mid x,y_{<t'}\big).

    The constrained decoder adds this heuristic to the partial-prefix score while retaining NEUROLOGIC pruning and diversity grouping. Thus, selecting a token that begins a multi-token constraint can receive credit for the probability of completing the whole phrase, while selecting an unrelated token can still be preferred if its future lookahead makes a later unsatisfied constraint likely. The stated scope limitation is that the constrained decoder currently accepts constraints expressible in the described logical form.

  5. Knowl 5 — Unconstrained future-likelihood heuristic

    model/method

    For unconstrained generation, NEUROLOGICF sets the constraint term to zero and uses the likelihood of a generated lookahead as an estimate of future quality. For a length-ℓ\ell continuation yt+1:t+ℓy_{t+1:t+\ell} after prefix y≤ty_{\leq t}, the heuristic is

    h(y≤t+ℓ)=λlog⁡pθ(yt+1:t+ℓ∣y≤t,x),h(y_{\leq t+\ell})=\lambda\log p_\theta(y_{t+1:t+\ell}\mid y_{\leq t},x),

    where λ\lambda controls the relative weight of estimated future likelihood and already-generated history. Candidate prefixes are ranked by accumulated log-probability plus the maximum heuristic value among their lookahead continuations. The same mechanism can modify beam search or top-kk sampling: for top-kk sampling, the heuristic-adjusted token scores are renormalized into a sampling distribution.

  6. Knowl 6 — Constrained commonsense generation results

    data/table

    The COMMONGEN experiment required every supplied concept to appear under some morphological inflection. The supervised condition used GPT-2 fine-tuned on COMMONGEN; the unsupervised condition used off-the-shelf GPT-2. Automatic metrics are reported in the order ROUGE-L, BLEU-4, METEOR, CIDEr, SPICE, and concept Coverage percentage. Human metrics are 3-point scores for Quality, Plausibility, Concepts, and Overall.

    Supervised results:

    • CBS: 38.8, 20.6, 28.5, 12.9, 27.1, 97.6%; 2.27, 2.35, 2.51, 2.23.
    • GBS: 38.2, 18.4, 26.7, 11.7, 26.1, 97.4%; 2.06, 2.17, 2.29, 2.01.
    • DBA: 38.3, 18.7, 27.7, 12.4, 26.3, 97.5%; 2.23, 2.30, 2.43, 2.15.
    • NEUROLOGIC: 42.8, 26.7, 30.2, 14.7, 30.3, 97.7%; 2.54, 2.56, 2.67, 2.50.
    • NEUROLOGICF with greedy lookahead: 43.6, 28.2, 30.8, 15.2, 30.8, 97.8%; 2.66, 2.67, 2.73, 2.59.
    • NEUROLOGICF with sampling lookahead: 43.4, 27.9, 30.8, 15.3, 31.0, 97.7%; 2.64, 2.64, 2.74, 2.58.
    • NEUROLOGICF with beam lookahead: 43.2, 28.2, 30.7, 15.2, 31.0, 97.6%; 2.68, 2.67, 2.76, 2.60.

    Unsupervised results:

    • TSMH: 24.7, 2.2, 14.5, 3.6, 15.4, 71.5%; 1.85, 1.92, 1.95, 1.63.
    • NEUROLOGIC: 41.9, 24.7, 29.5, 14.4, 27.5, 96.7%; 2.64, 2.52, 2.68, 2.50.
    • NEUROLOGICF with greedy lookahead: 44.3, 28.6, 30.7, 15.6, 29.6, 97.1%; 2.78, 2.70, 2.77, 2.70.

    NEUROLOGICF improved both automatic and human quality over prior constrained decoders while preserving high concept coverage. In particular, the unsupervised NEUROLOGICF system exceeded every supervised baseline on the reported human metrics. Only greedy lookahead was tested in the unsupervised condition because of computational cost; among supervised variants, beam lookahead had the strongest human scores while greedy lookahead was the fastest.

  7. Knowl 7 — Constrained machine translation results

    data/table

    The constrained machine-translation experiment integrated terminology from the IATE database into 414 WMT17 English-German test sentences. BLEU and Term% measure translation quality and the percentage of supplied terminology terms generated. Results were obtained with both the two-layer Transformer used in earlier work and the off-the-shelf Marian MT system.

    For the two-layer Transformer, the reported pairs are BLEU/Term%: unconstrained 25.8/76.3; training by appending 26.0/92.9; training by repeating 26.0/94.5; Post and Vilar decoding 25.3/82.0; NEUROLOGIC 26.5/95.1; NEUROLOGICF with greedy lookahead 26.7/95.8; NEUROLOGICF with sampling lookahead 26.6/95.4; and NEUROLOGICF with beam lookahead 26.6/95.8. The two training-based methods were not applicable to Marian MT.

    For Marian MT, the corresponding pairs are: unconstrained 32.9/85.0; Post and Vilar decoding 33.0/94.3; NEUROLOGIC 33.4/97.1; NEUROLOGICF with greedy lookahead 33.7/97.2; NEUROLOGICF with sampling lookahead 33.7/97.2; and NEUROLOGICF with beam lookahead 33.6/97.2. Thus, NEUROLOGICF improved both BLEU and terminology coverage on both model families without retraining the underlying MT system.

    The two-layer Transformer results broken down by the number of terminology terms were as follows. With one term, beam search scored 25.4 BLEU and 79.6% Term%, NEUROLOGIC scored 26.2 and 95.2%, and NEUROLOGICF scored 26.3 and 95.8%. With two or more terms, beam search scored 28.1 and 85.0%, NEUROLOGIC scored 28.9 and 93.7%, and NEUROLOGICF scored 29.3 and 96.5%. The larger gain for multiple terms supports the claim that lookahead is especially useful for complex constraint sets.

  8. Knowl 8 — Few-shot table-to-text results

    data/table

    The E2ENLG experiment linearized structured tables as input and fine-tuned GPT-2 using only randomly sampled fractions of the training data. With 0.1% of the training data, the reported metrics were NIST, BLEU, METEOR, CIDEr, ROUGE, and information Coverage percentage.

    At 0.1% training data, beam search scored 3.82, 42.8, 32.6, 10.8, 57.8, and 73.6%; CBS scored 6.50, 42.3, 36.4, 13.0, 54.3, and 91.6%; GBS scored 6.26, 40.7, 36.7, 12.9, 54.2, and 94.1%; NEUROLOGIC scored 6.95, 47.6, 38.9, 16.3, 58.7, and 97.6%; NEUROLOGICF with greedy lookahead scored 7.11, 49.2, 40.0, 17.5, 60.0, and 100.0%; NEUROLOGICF with beam lookahead scored 7.01, 48.9, 40.0, 17.2, 59.8, and 99.9%; and NEUROLOGICF with sampling lookahead scored 7.11, 49.3, 40.1, 17.5, 60.0, and 100.0%. The lookahead systems therefore improved both text quality and factual information coverage, unlike CBS and GBS, which raised coverage while lowering several quality metrics.

    BLEU-4 as a function of the fine-tuning fraction was: TGen, 3.6, 27.9, 35.2, and 57.3; Template-GPT-2, 22.5, 47.8, 53.3, and 59.9; KGPT-Graph, 39.8, 53.3, 55.1, and 61.5; KGPT-Seq, 40.2, 53.0, 54.1, and 61.1; GPT-2, 42.8, 57.1, 56.8, and 61.1; GPT-2 plus NEUROLOGIC, 47.6, 56.9, 58.0, and 62.9; and GPT-2 plus greedy NEUROLOGICF, 49.2, 58.0, 58.4, and 63.4, respectively at 0.1%, 0.5%, 1%, and 5% of the training data. The relative benefit of NEUROLOGICF was largest in the lowest-data setting.

  9. Knowl 9 — Zero-shot constrained question generation results

    data/table

    The constrained question-generation experiment used off-the-shelf GPT-2 with no task-specific training data. Each system received keywords and had to produce an interrogative question. Automatic metrics are ROUGE, BLEU, METEOR, CIDEr, SPICE, and keyword Coverage percentage; human metrics are 3-point scores for Grammar, Fluency, Meaningfulness, and Overall.

    CGMH scored 28.8, 2.0, 18.0, 5.5, 21.5, 18.3%; 2.28, 2.34, 2.11, 2.02. TSMH scored 42.0, 4.3, 25.9, 10.4, 37.7, 92.7%; 2.35, 2.28, 2.37, 2.22. NEUROLOGIC scored 38.8, 11.2, 24.5, 18.0, 41.7, 90.6%; 2.78, 2.71, 2.49, 2.51. NEUROLOGICF with greedy lookahead scored 43.7, 14.7, 28.0, 20.9, 47.7, 100.0%; 2.83, 2.77, 2.74, 2.76. NEUROLOGICF with beam lookahead scored 42.9, 14.4, 27.8, 20.3, 46.9, 100.0%; 2.81, 2.86, 2.76, 2.75. NEUROLOGICF with sampling lookahead scored 43.5, 14.6, 28.2, 20.8, 47.8, 100.0%; 2.83, 2.75, 2.76, 2.73.

    NEUROLOGICF achieved perfect keyword coverage and exceeded the prior methods on the automatic metrics and overall human quality. The improvement over vanilla NEUROLOGIC was particularly large in this zero-shot setting, which combines absent supervision with logical constraints involving both keywords and question syntax.

  10. Knowl 10 — Unconstrained story-generation results and decoding cost

    empirical result

    On conditional RocStories generation, GPT-2 was fine-tuned on the story-training set and evaluated with perplexity, BLEU-1, BLEU-2, unique 2-, 3-, and 4-grams, and human Coherence and Overall scores. The reported results were:

    • Beam search: PPL 2.24; BLEU-1 33.7; BLEU-2 16.5; unique 2-grams 20.13k; unique 3-grams 34.09k; unique 4-grams 41.91k; coherence 2.46; overall 2.32.
    • Beam search plus greedy-lookahead NEUROLOGICF: PPL 2.11; BLEU-1 34.3; BLEU-2 16.7; unique 2-grams 20.63k; unique 3-grams 34.94k; unique 4-grams 43.02k; coherence 2.56; overall 2.57.
    • Beam search plus beam-lookahead NEUROLOGICF: PPL 2.14; BLEU-1 34.4; BLEU-2 16.8; unique 2-grams 20.68k; unique 3-grams 35.03k; unique 4-grams 43.12k; coherence 2.62; overall 2.63.
    • Beam search plus sampling-lookahead NEUROLOGICF: PPL 2.16; BLEU-1 34.4; BLEU-2 16.7; unique 2-grams 20.78k; unique 3-grams 35.41k; unique 4-grams 43.64k; coherence 2.59; overall 2.57.
    • Top-kk sampling: PPL 4.01; BLEU-1 31.4; BLEU-2 13.9; unique 2-grams 28.54k; unique 3-grams 48.36k; unique 4-grams 56.62k; coherence 2.23; overall 2.15.
    • Top-kk sampling plus greedy-lookahead NEUROLOGICF: PPL 3.68; BLEU-1 32.1; BLEU-2 14.3; unique 2-grams 28.47k; unique 3-grams 48.44k; unique 4-grams 56.63k; coherence 2.48; overall 2.47.
    • Top-kk sampling plus beam-lookahead NEUROLOGICF: PPL 3.75; BLEU-1 32.2; BLEU-2 14.4; unique 2-grams 28.53k; unique 3-grams 48.27k; unique 4-grams 56.36k; coherence 2.39; overall 2.34.
    • Top-kk sampling plus sampling-lookahead NEUROLOGICF: PPL 3.70; BLEU-1 32.0; BLEU-2 14.2; unique 2-grams 28.57k; unique 3-grams 48.04k; unique 4-grams 56.15k; coherence 2.47; overall 2.44.

    Lookahead improved fluency and human quality for both decoding families. With beam search it also increased diversity, while with top-kk sampling it preserved approximately the original diversity. Beam lookahead was strongest for beam search, and greedy lookahead was strongest for top-kk sampling. The method is substantially more expensive: on COMMONGEN with fine-tuned GPT2-L, runtime was 0.20 seconds per sentence for beam search, 2.04 seconds for NEUROLOGIC, and 19.24 seconds for NEUROLOGICF.

Coverage note — Detailed hyperparameter-sensitivity plots, training and decoding hyperparameter tables, qualitative examples, and human-evaluation templates were omitted because they support the reported method and results rather than adding separate load-bearing contributions.

References

  1. 1.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. Spice: Semantic propositional image caption evaluation. In European conference on computer vision, pages 382–398. Springer.
  2. 2.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2017. Guided open vocabulary image captioning with constrained beam search. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 936–945, Copenhagen, Denmark. Association for Computational Linguistics.
  3. 3.Michael Auli and Adam Lopez. 2011. Efficient CCG parsing: A* versus adaptive supertagging. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 1577–1585, Portland, Oregon, USA. Association for Computational Linguistics.
  4. 4.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72.
  5. 5.Emily Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT).
  6. 6.Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi. 2017. Findings of the 2017 conference on machine translation (WMT17). In Proceedings of the Second Conference on Machine Translation, pages 169–214, Copenhagen, Denmark. Association for Computational Linguistics.
  7. 7.T. Brown, B. Mann, Nick Ryder, Melanie Subbiah, J. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, G. Krüger, T. Henighan, R. Child, Aditya Ramesh, D. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, E. Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, J. Clark, Christopher Berner, Sam McCandlish, A. Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS).
  8. 8.Rajen Chatterjee, Matteo Negri, Marco Turchi, Marcello Federico, Lucia Specia, and Frédéric Blain. 2017. Guiding neural machine translation decoding with external knowledge. In Proceedings of the Second Conference on Machine Translation, pages 157–168, Copenhagen, Denmark. Association for Computational Linguistics.
  9. 9.Wenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen, and William Yang Wang. 2020a. Logical natural language generation from open-domain tables. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7929–7942, Online. Association for Computational Linguistics.
  10. 10.Wenhu Chen, Yu Su, Xifeng Yan, and William Yang Wang. 2020b. KGPT: Knowledge-grounded pre-training for data-to-text generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8635–8648, Online. Association for Computational Linguistics.
  11. 11.Yining Chen, Sorcha Gilroy, Andreas Maletti, Jonathan May, and Kevin Knight. 2018. Recurrent neural networks as weighted language recognizers. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2261–2271, New Orleans, Louisiana. Association for Computational Linguistics.
  12. 12.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations.
  13. 13.Georgiana Dinu, Prashant Mathur, Marcello Federico, and Yaser Al-Onaizan. 2019. Training neural machine translation to apply terminology constraints. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3063–3068, Florence, Italy. Association for Computational Linguistics.
  14. 14.Ondřej Dušek and Filip Jurčíček. 2016. Sequence-to-sequence generation for spoken dialogue via deep syntax trees and strings. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 45–51, Berlin, Germany. Association for Computational Linguistics.
  15. 15.Ondřej Dušek, Jekaterina Novikova, and Verena Rieser. 2018. Findings of the E2E NLG Challenge. In Proc. of the 11th International Conference on Natural Language Generation, pages 322–328, Tilburg, The Netherlands. Association for Computational Linguistics.
  16. 16.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  17. 17.Aria Haghighi, John DeNero, and Dan Klein. 2007. Approximate factoring for A* search. In Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Proceedings of the Main Conference, pages 412–419, Rochester, New York. Association for Computational Linguistics.
  18. 18.Peter E. Hart, Nils J. Nilsson, and Bertram Raphael. 1968. A formal basis for the heuristic determination of minimum cost paths. IEEE Transactions on Systems Science and Cybernetics, 4(2):100–107.
  19. 19.Eva Hasler, Adrià de Gispert, Gonzalo Iglesias, and Bill Byrne. 2018. Neural machine translation decoding with terminology constraints. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 506–512, New Orleans, Louisiana. Association for Computational Linguistics.
  20. 20.Chris Hokamp and Qun Liu. 2017. Lexically constrained decoding for sequence generation using grid beam search. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1535–1546, Vancouver, Canada. Association for Computational Linguistics.
  21. 21.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In International Conference on Learning Representations.
  22. 22.Mark Hopkins and Greg Langmead. 2009. Cube pruning as heuristic search. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 62–71.
  23. 23.J. Edward Hu, Huda Khayrallah, Ryan Culkin, Patrick Xia, Tongfei Chen, Matt Post, and Benjamin Van Durme. 2019. Improved lexically constrained decoding for translation and monolingual rewriting. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 839–850, Minneapolis, Minnesota. Association for Computational Linguistics.
  24. 24.Zhiting Hu, Zichao Yang, Xiaodan Liang, Ruslan Salakhutdinov, and Eric P Xing. 2017. Toward controlled generation of text. In International Conference on Machine Learning, pages 1587–1596. PMLR.
  25. 25.Daphne Ippolito, Reno Kriz, João Sedoc, Maria Kustikova, and Chris Callison-Burch. 2019. Comparison of diverse decoding methods from conditional language models. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3752–3762.
  26. 26.Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, André F. T. Martins, and Alexandra Birch. 2018. Marian: Fast neural machine translation in C++. In Proceedings of ACL 2018, System Demonstrations, pages 116–121, Melbourne, Australia. Association for Computational Linguistics.
  27. 27.Dan Klein and Christopher D. Manning. 2003. A* parsing: Fast exact Viterbi parse selection. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, pages 119–126.
  28. 28.Richard E Korf. 1985. Depth-first iterative-deepening: An optimal admissible tree search. Artificial intelligence, 27(1):97–109.
  29. 29.Klaus Krippendorff. 2007. Computing krippendorff’s alpha reliability. Departmental papers (ASC), page 43.
  30. 30.Rémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar, Jean-Baptiste Lespiau, Ioannis Antonoglou, Karen Simonyan, and Oriol Vinyals. 2021. Machine translation decoding beyond beam search. arXiv preprint arXiv:2104.05336.
  31. 31.Kenton Lee, Mike Lewis, and Luke Zettlemoyer. 2016. Global neural CCG parsing with optimality guarantees. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2366–2376, Austin, Texas. Association for Computational Linguistics.
  32. 32.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  33. 33.Bill Yuchen Lin, Ming Shen, Wangchunshu Zhou, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. Commongen: A constrained text generation challenge for generative commonsense reasoning. In Findings of EMNLP.
  34. 34.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. Text Summarization Branches Out.
  35. 35.Chin-Yew Lin and Eduard Hovy. 2003. Automatic evaluation of summaries using n-gram co-occurrence statistics. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, pages 150–157.
  36. 36.Ximing Lu, Peter West, Rowan Zellers, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. NeuroLogic decoding: (un)supervised neural text generation with predicate logic constraints. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4288–4299, Online. Association for Computational Linguistics.
  37. 37.Kris McGuffie and Alex Newhouse. 2020. The radicalization risks of gpt-3 and advanced neural language models. arXiv.
  38. 38.Clara Meister, Tim Vieira, and Ryan Cotterell. 2020. Best-first beam search. Transactions of the Association for Computational Linguistics, 8:795–809.
  39. 39.Ning Miao, Hao Zhou, Lili Mou, Rui Yan, and Lei Li. 2019. Cgmh: Constrained sentence generation by metropolis-hastings sampling. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6834–6842.
  40. 40.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. A corpus and cloze evaluation for deeper understanding of commonsense stories. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 839–849, San Diego, California. Association for Computational Linguistics.
  41. 41.Iftekhar Naim, Daniel Gildea, Walter Lasecki, and Jeffrey P Bigham. 2013. Text alignment for real-time crowd captioning. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 201–210.
  42. 42.Franz Josef Och, Nicola Ueffing, and Hermann Ney. 2001. An efficient a* search algorithm for statistical machine translation. In Proceedings of the ACL 2001 Workshop on Data-Driven Methods in Machine Translation.
  43. 43.Jayshree Pandya. 2019. The dual-use dilemma of artificial intelligence. Forbes Magazine.
  44. 44.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In ACL, pages 311–318.
  45. 45.Judea Pearl. 1984. Heuristics - intelligent search strategies for computer problem solving. In Addison-Wesley series in artificial intelligence.
  46. 46.Ira Pohl. 1970. First Results on the Effect of Error in Heuristic Search.
  47. 47.Matt Post and David Vilar. 2018a. Fast lexically constrained decoding with dynamic beam allocation for neural machine translation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1314–1324, New Orleans, Louisiana. Association for Computational Linguistics.
  48. 48.Matt Post and David Vilar. 2018b. Fast lexically constrained decoding with dynamic beam allocation for neural machine translation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1314–1324.
  49. 49.Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D Hwang, Ronan Le Bras, Antoine Bosselut, and Yejin Choi. 2020. Backpropagation-based decoding for unsupervised counterfactual and abductive reasoning. In EMNLP.
  50. 50.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.
  51. 51.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9.
  52. 52.Sheng Shen, Daniel Fried, Jacob Andreas, and Dan Klein. 2019. Pragmatically informative text generation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4060–4067, Minneapolis, Minnesota. Association for Computational Linguistics.
  53. 53.Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4566–4575.
  54. 54.Sean Welleck, Kianté Brantley, Hal Daumé Iii, and Kyunghyun Cho. 2019a. Non-monotonic sequential text generation. In International Conference on Machine Learning, pages 6716–6726. PMLR.
  55. 55.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2019b. Neural text generation with unlikelihood training. In International Conference on Learning Representations.
  56. 56.Peter West, Ximing Lu, Ari Holtzman, Chandra Bhagavatula, Jena D. Hwang, and Yejin Choi. 2021. Reflective decoding: Beyond unidirectional generation with off-the-shelf language models. In ACL/IJCNLP.
  57. 57.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP): System Demonstrations.
  58. 58.Hao Zhang and Daniel Gildea. 2006. Efficient search for inversion transduction grammar. In Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing, pages 224–231.
  59. 59.Maosen Zhang, Nan Jiang, Lei Li, and Yexiang Xue. 2020. Language generation via combinatorial constraint satisfaction: A tree search enhanced Monte-Carlo approach. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1286–1298, Online. Association for Computational Linguistics.
  60. 60.Renjie Zheng, Mingbo Ma, Baigong Zheng, Kaibo Liu, and Liang Huang. 2020. Opportunistic decoding with timely correction for simultaneous translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 437–442.

Citation

MLA
Lu, X., et al. “NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 780–99, https://doi.org/10.18653/v1/2022.naacl-main.57.
APA
Lu, X., Welleck, S., West, P., Jiang, L., Kasai, J., Khashabi, D., Bras, R. L., Qin, L., Yu, Y., Zellers, R., Smith, N. A., & Choi, Y. (2022). NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 780–799. https://doi.org/10.18653/v1/2022.naacl-main.57
Chicago
Lu, X., S. Welleck, P. West, et al. 2022. “NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 780–99. https://doi.org/10.18653/v1/2022.naacl-main.57.
Harvard
Lu, X. et al. (2022) “NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 780–799. Available at: https://doi.org/10.18653/v1/2022.naacl-main.57.
Vancouver
1. Lu X, Welleck S, West P, et al (2022) NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 780–799

BibTeX

@inproceedings{lu-etal-2022-neurologic,
    title = "{N}euro{L}ogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics",
    author = "Lu, Ximing  and
      Welleck, Sean  and
      West, Peter  and
      Jiang, Liwei  and
      Kasai, Jungo  and
      Khashabi, Daniel  and
      Le Bras, Ronan  and
      Qin, Lianhui  and
      Yu, Youngjae  and
      Zellers, Rowan  and
      Smith, Noah A.  and
      Choi, Yejin",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.57/",
    doi = "10.18653/v1/2022.naacl-main.57",
    pages = "780--799"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/